A Minute-Level Non-Invasive Load Decomposition Method and System

By introducing WaveNet network and multi-head self-attention mechanism in non-invasive load decomposition, the problems of reduced accuracy and inefficiency in minute-level meter data processing are solved in the prior art, and high-precision and efficient minute-level load decomposition are achieved.

CN115619589BActive Publication Date: 2025-06-13INFORMATION & COMM BRANCH OF STATE GRID JIANGSU ELECTRIC POWER +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211164490.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-09-23
Publication Date
2025-06-13
Estimated Expiration
2042-09-23

AI Technical Summary

Technical Problem

The existing non-invasive load decomposition method reduces accuracy, is costly and difficult to promote when facing meter data at minute levels or higher frequencies. The existing models are prone to information leakage, gradient vanishing and gradient explosion problems when processing long sequence data.

Method used

The WaveNet network and the multi-head self-attention mechanism are used for feature extraction, and the timing and receptive fields are ensured through causal convolution, hollow convolution and residual network. The features are extracted in parallel with the multi-head self-attention mechanism, and feature fusion is performed. Finally, regression is performed through the full connection layer to achieve minute-level non-invasive load decomposition.

Benefits of technology

The accuracy of minute-level load decomposition is improved, the efficiency and accuracy of existing models in high-frequency data processing is solved, and the training efficiency of the model and the accuracy of load decomposition are ensured.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115619589B_ABST
    Figure CN115619589B_ABST
Patent Text Reader

Abstract

The present invention provides a minute-level non-invasive load decomposition method and system, which uses a WaveNet network that can process data in parallel and extract global features and a multi-head self-attention mechanism to extract features from minute-level power load data. Since there will be missing power features in minute-level data compared with second-level data, in order to strengthen feature extraction, while using WaveNet to extract features, a multi-head self-attention mechanism is introduced to extract features in parallel, and the features extracted by the two are fused and then regression is performed, thereby improving the training efficiency of the load decomposition model while ensuring the accuracy of minute-level non-invasive load decomposition.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a minute-level non-intrusive load decomposition method and system, belonging to the technical field of smart grids. Background Art

[0002] Smart power consumption is an important part of the "grid-load" interactive service system that connects the grid side and the user side to realize a smart grid. It has a great impact on the efficient operation and construction planning of the entire grid. Residential users have the characteristics of diverse electricity consumption behaviors and a large number of users. Realizing flexible and interactive smart power consumption for residents is the main development goal of the power system. Load monitoring is the primary link to achieve smart power consumption for residents. By sampling and analyzing the total load data of residential electricity consumption, monitoring the detailed operation status of each electrical appliance inside the residential household, and obtaining data information such as the power consumption of each electrical appliance and the electricity consumption behavior, on the one hand, it can improve the accuracy of the grid load model, the safety of grid operation, and the scientific nature of power planning. On the other hand, it can effectively improve the energy-saving awareness of users and achieve energy conservation, emission reduction, and sustainable development of the whole society.

[0003] Traditional load monitoring usually adopts an intrusive method, that is, collectors are installed on each electrical appliance of the user to record its usage. The advantage of this method is that the monitoring data is accurate and reliable, but the disadvantages are poor practical operability, high implementation cost, and low user acceptance. Non-intrusive load monitoring decomposes the total load information of household users to obtain the load information of each electrical appliance, and then obtains electricity consumption information such as the energy consumption of electrical appliances and the electricity consumption behavior rules of users. Compared with intrusive load monitoring, non-intrusive load monitoring does not need to go deep into the interior of the user's home, and only installs collection equipment externally, which is relatively simple and convenient. However, in actual applications, this method mostly realizes through the transformation or addition of monitoring equipment, facing difficulties such as more investment, higher cost, wider scope, and greater promotion difficulty. Therefore, how to use the existing in-operation electricity meters and data collection environment of residents to carry out research on non-intrusive load decomposition methods based on minute-level electricity meter data has important application value.

[0004] According to the load decomposition principle, the existing non-intrusive load decomposition methods are mainly divided into load decomposition based on optimization algorithms and load decomposition based on pattern recognition.

[0005] (1) Load decomposition based on optimization algorithms

[0006] The optimization algorithm searches for the optimal match between the unknown total load and the known loads in the database by selecting linearly superimposable load characteristics and fitness functions. Common mathematical optimization algorithms include integer programming, differential evolution, particle swarm optimization algorithm, and genetic algorithm, etc. The non-intrusive home load decomposition algorithm based on the dynamic adaptive particle swarm algorithm adds the total harmonic distortion coefficient as a new feature to the objective function, and uses the dynamic adaptive particle swarm algorithm to improve the accuracy of load identification, with high reliability and robustness. The non-intrusive load decomposition algorithm based on affinity propagation clustering and genetic optimization uses discretization to obtain finite states and uses the genetic algorithm to achieve more accurate load decomposition and state recognition. The non-intrusive load decomposition algorithm considering the state probability factor and state correction adds the state probability factor to the objective function to improve the load decomposition accuracy in the case of similar powers and less prior information. The above algorithms are all studied based on signal characteristics, but the extraction of load characteristics is relatively complex, and the overall understanding of the signal is insufficient. Some scholars use signal separation technology for the load current signal, reducing the cumbersome feature extraction, and transforming the underdetermined problem into an optimal constraint problem to complete effective identification. However, the load decomposition based on the optimization algorithm has high requirements for the extraction of load characteristics, low algorithm solving efficiency, is easy to fall into local optimum, and the optimization algorithm decomposes the power into the discrete power of individual devices, which is only applicable to analyzing the on / off state of electrical appliances and cannot accurately decompose the total power into the continuous power values of individual powers.

[0007] (2) Load Decomposition Based on Pattern Recognition

[0008] Aiming at the limitations of the optimization algorithm, pattern recognition has gradually become a common algorithm for load decomposition. Pattern recognition is divided into unsupervised learning algorithms and supervised learning algorithms. Unsupervised learning algorithms do not need to label the training data in advance, but according to the context information of the target device, use algorithm models such as clustering to obtain the relationship between data and features, minimizing the cost of manual interaction. The Factorial Hidden Markov Model (FHMM) is a common unsupervised learning method for load decomposition. Based on the FHMM, there are various extended algorithms, such as the Conditional Factorial Hidden Markov, Factorial Hidden Semi-Markov Model, and Conditional Factorial Hidden Semi-Markov Model. Although the unsupervised learning algorithm minimizes the initial cost, the load decomposition effect is poor.

[0009] Supervised learning algorithms require labeling the training data and finding the relationship between data and features through model training, which can improve the accuracy of load decomposition. The k-nearest neighbor algorithm is simple and clear, effectively classifying highly complex things, but the recognition requires a large amount of computing time and memory, is sensitive to the composition of training samples, and the selection of the k value will affect the classification effect. The Support Vector Machine (SVM) and Gaussian Mixture Model (GMM) use GMM to describe the current waveform distribution and use SVM to classify the extracted power characteristics. The non-intrusive load monitoring algorithm based on k-NN and kernel Fisher discriminant uses the AdaBoost algorithm to streamline the training samples, combines k-NN and kernel Fisher, and takes into account both recognition accuracy and computational complexity.

[0010] Deep learning algorithms have the advantages of superior classification performance, fixed computational complexity, and no need for feature engineering, and have been widely used in non-intrusive load decomposition in recent years. In 2015, Jack Kelly first used three deep learning frameworks to handle non-intrusive load decomposition problems, namely Long Short-Term Memory (LSTM), "rectangular architecture", and Denoising AutoEncoders (DAE) to achieve load decomposition, and compared with the Combinatorial Optimization (CO) and FHMM algorithms. The results showed that the DAE and "rectangular" architectures were superior to the CO and FHMM algorithms, while the LSTM algorithm was suitable for electrical appliances in two states and performed poorly on multi-state devices. The Sequence to Point (Seq2Point) algorithm uses a convolutional neural network to learn the features of the target device. The input is the total load window data, and the output is the midpoint of the target device. The results show that the Seq2Point model performs better in terms of performance and computational cost. Most load decomposition methods adopt single-task learning methods, ignoring the dependencies between multiple devices. The UNet-NILM network can perform state detection and power estimation of multi-task electrical appliances, applying multi-label learning strategies and multi-objective quantile regression. Experimental evaluations on the UKDALE dataset show that this method has good performance compared with traditional single-task learning. The BERT4NILM model applies the Bidirectional Encoder Representations from Transformers (BERT) architecture in natural language processing to the field of load decomposition, and specifically designs and improves the objective function for non-intrusive load decomposition, and follows the sequence-to-sequence learning pattern. Experiments show that BERT4NILM performs well on two public datasets, UK-DALE and REDD, but the training speed is slow.

[0011] To sum up, non-intrusive load decomposition started earlier abroad, and the main research focused on the decomposition of high-frequency load data mainly based on active power, which needs to be achieved by transforming or adding monitoring devices in practical applications, with high costs and great difficulties in popularization. If it is directly used for the minute-level load data obtained through existing residential electricity meters in operation and data acquisition environments, the load decomposition accuracy will be reduced, especially for electrical appliances with shorter working cycles.

[0012] Existing load decomposition models are not applicable to second - level or millisecond - level power data. In practical applications, in order to collect second - level or millisecond - level high - frequency data, it is necessary to transform or add dedicated collection equipment, which faces difficulties such as high costs and large promotion difficulties. If the existing load decomposition models are directly used for the minute - level data obtained from in - use residential electricity meters and data collection environments, the load decomposition accuracy will decrease, especially for short - term load electrical equipment. At the same time, most of the existing load decomposition models use recurrent neural networks and their variants, which require serial data processing. If the input sequence is too long, it is easy to cause information leakage, and the efficiency is low, and problems such as gradient disappearance and gradient explosion are likely to occur. Summary of the Invention

[0013] The purpose of the present invention is to overcome the deficiencies in the prior art and provide a minute - level non - intrusive load decomposition method and system to improve the accuracy of minute - level load decomposition.

[0014] To achieve the above - mentioned purpose, the present invention is implemented by the following technical solutions:

[0015] In the first aspect, the present invention provides a minute - level non - intrusive load decomposition method, including the following steps:

[0016] Step S1: Obtain the minute - level total power sequence of the electricity meter;

[0017] Step S2: Input the minute - level total power sequence of the electricity meter into the minute - level non - intrusive load decomposition model to obtain the target single - electrical - appliance power type sequence, thereby realizing load decomposition;

[0018] The minute - level non - intrusive load decomposition model includes a feature extraction module and a regression module; the feature extraction module includes a WaveNet network and a multi - head self - attention mechanism; the regression module includes a fully - connected layer;

[0019] Inputting the minute - level total power sequence of the electricity meter into the minute - level non - intrusive load decomposition model to obtain the target single - electrical - appliance power type sequence includes:

[0020] Step S21: Input the minute - level total power sequence of the electricity meter into the WaveNet network for feature extraction to obtain a WaveNet feature vector.

[0021] Step S22: While using the WaveNet to extract features, introduce a multi - head self - attention mechanism to extract features from the minute - level total power sequence in parallel to obtain a self - attention feature vector.

[0022] Step S23: Fuse the WaveNet feature vector and the self - attention feature vector to obtain a fused feature vector;

[0023] Step S24: Input the fused feature vector into the fully connected layer, and perform regression through the fully connected layer to obtain the target single - electrical - appliance power type sequence, where the single - electrical - appliance power type is the label, thereby realizing load decomposition.

[0024] Further, the training method of the minute - level non - intrusive load decomposition model includes:

[0025] Obtain a training data set with artificial labels, which is composed of the minute - level total power sequence of the electric meter and the corresponding single - electrical - appliance power type sequence within the same time period, where the single - electrical - appliance power type is the label;

[0026] Use the minute - level total power sequence of the electric meter as the model input and the single - electrical - appliance power type as the output to train the model, and obtain a trained minute - level non - intrusive load decomposition model.

[0027] Further, the training data set is obtained by reading the electric meter data.

[0028] Further, the WaveNet network includes causal convolution, dilated convolution, and a residual network, as well as a gated structure composed of activation functions of tanh and sigmoid;

[0029] The causal convolution is used to perform parallel one - way convolution operations on the minute - level total power time - series data of the electric meter;

[0030] The residual network includes multiple residual blocks;

[0031] The dilated convolution is used to perform feature representation on the input power in each residual block;

[0032] The gated structure composed of the activation functions of tanh and sigmoid is used to perform selective non - linear transformation on the input.

[0033] Further, in step S21, inputting the minute - level total power sequence of the electric meter into the WaveNet network for feature extraction includes:

[0034] Use the minute - level total power sequence of the electric meter as the input, obtain a preliminary feature representation through causal convolution, and then enter multiple residual blocks for further feature extraction;

[0035] In each residual block, the input power is first subjected to feature representation through dilated convolution, and then enters the tanh activation function for non - linear feature transformation to obtain a non - linear transformation result. While performing the non - linear feature transformation, it enters the sigmoid activation function to generate a selective output. Finally, the non - linear transformation result and the selective output are subjected to Hadamard product; where the selective output is used to control the degree of non - linear transformation;

[0036] Among them, the selective output is used to control the degree of non-linear transformation;

[0037] The result of the Hadamard product is convolved by a 1×1 convolution to obtain the output of the gating structure. The output of the gating structure is superimposed on the original input through a residual connection and then enters the next residual block, and serves as part of the information of the final output through another residual connection mechanism;

[0038] The final output part sums the information obtained from multiple residual blocks, and then passes through two layers of 1×1 convolution kernels and the Relu activation function to obtain the final power information output, that is, the WaveNet feature vector.

[0039] Further, in step S22, the multi-head self-attention mechanism is used to extract features from the source sequence, including:

[0040] The minute-level total power sequence of the electric meter is linearly transformed through three randomly initialized weight matrices W Q 、W K and W V to obtain Q, K, and V; where Q is the Query matrix, K is the Key matrix, and V is the Value matrix.

[0041] Calculate the dot product of Q and K, and scale the dot product with the dimension d k of the row vector in K, then normalize the result of the previous step with the softmax function to obtain the attention score, and then weighted sum V with the attention score to obtain the attention value Attention, and its calculation formula is:

[0042]

[0043] In the formula, T is the transpose, Q is the Query matrix, K is the Key matrix, V is the Value matrix, d k is the dimension of the row vector in K. In the formula, the softmax function is

[0044]

[0045] In the formula, z is the input vector of the function, containing k values, and e is the natural constant

[0046] The multi-head self-attention mechanism maps Q, K, and V h times respectively in the dimension of d s using a learnable linear mapping matrix, d s =d k / h, then executes the attention function in parallel, and finally connects the h results and performs a linear mapping again to obtain the self-attention feature vector, and the calculation formula is as follows:

[0047] Multihead(Q, K, V) = Concat(head 1 ,..., head h )W o

[0048] head i = Attention(QW i Q , KW i K , VW i V ), i = 1, 2,..., h

[0049] where h represents the number of heads, Q is the Query matrix, K is the Key matrix, and V is the Value matrix.

[0050] The multi-head attention mechanism extends the model's ability to focus on different positions, with multiple linear transformations and concatenations, and the parameters of the linear transformations are learnable, improving the model's fitting ability.

[0051] Furthermore, in step S23, fusing the WaveNet feature vector and the self-attention feature vector to obtain a fused feature vector includes:

[0052] The feature fusion uses the method of vector concatenation to concatenate two groups of parallel output feature vectors to obtain a fused feature representation.

[0053] In a second aspect, the present invention provides a minute-level non-intrusive load decomposition system, including:

[0054] An input module: used to obtain the minute-level total power sequence of the electricity meter;

[0055] A decomposition module: used to input the minute-level total power sequence of the electricity meter into a minute-level non-intrusive load decomposition model to obtain a target single-appliance power type sequence, thereby realizing load decomposition;

[0056] The minute-level non-intrusive load decomposition model includes a feature extraction module and a regression module; the feature extraction module includes a WaveNet network and a multi-head self-attention mechanism; the regression module includes a fully connected layer;

[0057] The decomposition module includes:

[0058] A WaveNet network module: used to input the minute-level total power sequence of the electricity meter into the WaveNet network for feature extraction to obtain a WaveNet feature vector.

[0059] Multi-head self-attention module: It is used to extract features with WaveNet and introduce the multi-head self-attention mechanism to extract features from the minute-level total power sequence of the electricity meter in parallel, obtaining a self-attention feature vector.

[0060] Feature fusion module: It is used to fuse the WaveNet feature vector and the self-attention feature vector to obtain a fused feature vector.

[0061] Output module: It is used to input the fused feature vector into a fully connected layer, and perform regression through the fully connected layer to obtain the target single-electrical-appliance power type sequence, where the single-electrical-appliance power type is the label, thereby realizing load decomposition.

[0062] Compared with the prior art, the beneficial effects achieved by the present invention are as follows:

[0063] 1. The present invention can process data in parallel and use the WaveNet network that extracts global features to extract features from power load data. The multi-head self-attention mechanism is introduced to further strengthen the extraction of load features, thereby improving the accuracy of minute-level load decomposition; since there will be missing power features in minute-level data compared with second-level data, in order to strengthen the extraction of features, while using WaveNet to extract features, the multi-head self-attention mechanism is introduced to extract features in parallel, and the features extracted by the two are fused and then regressed, so as to improve the training efficiency of the load decomposition model while ensuring the accuracy of minute-level non-intrusive load decomposition.

[0064] 2. The WaveNet network of the present invention is composed of causal convolution, dilated convolution, and residual network. Causal convolution performs parallel unidirectional convolution operations on the minute-level total power time series data. Compared with ordinary convolution, it improves the training efficiency while ensuring the time series property; dilated convolution enables the convolution network to obtain a larger receptive field and capture multi-scale context information; the residual network is used to solve the problems of gradient vanishing and explosion.

[0065] 3. The results of load decomposition can effectively guide users to change their electricity consumption behaviors, promote the intelligence of the home energy-saving plan, and at the same time provide data support for the power department to formulate corresponding demand-side management, making home electricity consumption change towards a more energy-saving and efficient direction. Description of the Drawings

[0066] Figure 1 It is a schematic diagram of the WaveNet network structure;

[0067] Figure 2 It is a schematic diagram of the system;

[0068] Figure 3 It is a flowchart of the method of the present invention. Detailed Embodiments

[0069] The present invention will be further described below with reference to the accompanying drawings. The following embodiments are only used to more clearly illustrate the technical solution of the present invention and cannot be used to limit the protection scope of the present invention.

[0070] Embodiment 1:

[0071] This embodiment provides a minute-level non-intrusive load decomposition method. First, the WaveNet network is used to extract features from the minute-level total power load data of the electricity meter. Through causal convolution, parallel one-way convolution operations are performed on the minute-level time series data of the total power load of the electricity meter to ensure the data timeliness while performing parallel processing on the data, preventing information leakage and improving the training efficiency. Through dilated convolution, the convolutional network can obtain a larger receptive field and capture multi-scale context information. Through the residual network, the problems of gradient vanishing and gradient explosion are eliminated. Then, the multi-head self-attention mechanism is used to extract features from the minute-level total power load data of the electricity meter. Next, the feature vectors extracted by WaveNet and the multi-head self-attention mechanism are fused. Finally, the fully connected layer is used for regression to obtain the power data of a single electrical appliance in the corresponding time period.

[0072] Specifically, this method includes the following steps:

[0073] The present invention provides a minute-level non-intrusive load decomposition method, including the following steps:

[0074] Step S1: Obtain the minute-level total power sequence of the electricity meter;

[0075] Step S2: Input the minute-level total power sequence of the electricity meter into the WaveNet network for feature extraction to obtain the WaveNet feature vector.

[0076] Step S3: While using WaveNet to extract features, introduce the multi-head self-attention mechanism to parallelly extract features from the minute-level total power sequence of the electricity meter to obtain the self-attention feature vector.

[0077] Step S4: Fuse the WaveNet feature vector and the self-attention feature vector to obtain the fused feature vector.

[0078] Step S5: Input the fused feature vector into the fully connected layer, and perform regression through the fully connected layer to obtain the target single electrical appliance power type sequence, thereby realizing load decomposition.

[0079] Specifically, the construction method of the WaveNet network includes:

[0080] WaveNet is essentially a stack of multi-layer causal convolutions and dilated convolutions, using gated convolution operations with Sigmoid and Tanh activation functions. Between each layer, a residual network module is used to take the sum of the input and output of this layer as the input of the next layer, and skip connections are used to concatenate the outputs of each hidden layer. After passing through the Relu activation function, a one-dimensional convolution operation is performed to obtain the final output result.

[0081] The specific data processing flow of the WaveNet network is as follows:

[0082] The total power sequence of the residential household electricity meter is used as the input of WaveNet. First, it undergoes causal convolution to obtain a preliminary feature representation, and then enters multiple residual blocks for specific feature extraction. In each residual block, the input power first undergoes feature representation through dilated convolution, then enters the tanh activation function for non-linear feature transformation, and at the same time enters the sigmoid activation function to generate a selective output. Finally, the result of the non-linear transformation and the selective output are subjected to the Hadamard product. The Hadamard product is a type of operation on matrices. If A=(a ij ) and B=(b ij ) are two matrices of the same order, and if c ij = a ij × b ij , then the matrix C=(c ij ) is called the Hadamard product of A and B. Among them, the selective output is used to control the degree of non-linear transformation. The gated structure composed of the tanh and sigmoid activation functions is a common type of attention-like mechanism structure, which can perform selective non-linear transformation on the input.

[0083] The outputs of the gated structure are respectively superimposed with the original input through residual connections and enter the next residual block, and are used as part of the information of the final output through another residual connection mechanism. The final output part sums up the information obtained from multiple residual blocks to obtain the WaveNet feature vector.

[0084] WaveNet uses causal convolution to prevent data leakage. Causal convolution means that the output unit at each time t is only obtained from the units before time t through convolution operations.

[0085] To obtain a flexible receptive field and capture long-term historical information, a sufficiently large convolution kernel size or a sufficiently deep network is required. However, the increase in hidden layers easily leads to gradient vanishing or gradient explosion. Therefore, dilated convolution and residual networks are introduced.

[0086] Dilated convolution can enable the convolutional network to obtain a larger receptive field with the same number of hidden layers and convolution kernel size. The receptive field can be calculated by formula (4-1), where L is the number of layers and k is the filter size.

[0087] r = 2 L-1 k(4 - 1)

[0088] Atrous convolution is equivalent to introducing a fixed stride in two adjacent receptive fields to connect the upper and lower layers. The increase in the convolution rate can make the unit output connect to a wider range of input units, and the receptive field of the convolutional network is effectively expanded. The introduction of the atrous rate idea provides two methods to improve the receptive field of WaveNet. One is to select a larger filter size, and the other is to increase the atrous rate. Usually, the atrous rate increases exponentially with the number of layers of the convolutional network, which can ensure that the convolutional network uses the entire sequence as input and solve the problem of insufficient machine performance caused by the explosive growth of parameters in deep networks by reducing the number of parameters.

[0089] Residual network (Resnet) is used to solve the problem of gradient vanishing or gradient explosion in deep networks. The idea of the residual network is to add the output result to the input, and after passing through the activation function, it is used as the input of the next layer. Since the number of input and output units may not be the same, a one-dimensional convolution operation can be performed on the input and then added to the input. This method enables the network to learn not the direct global transformation of the input, but the potential identity relationship or residual term. Residual connections avoid gradient explosion or vanishing, enabling the network to transmit information in a cross-layer manner.

[0090] Specifically, in step S3, the multi-head self-attention mechanism is used to extract features from the source sequence, including:

[0091] To further enhance the load feature extraction ability, while using WaveNet to extract features, a multi-head self-attention mechanism is introduced to extract features in parallel.

[0092] First, the minute-level total power sequence of the electric meter is linearly transformed through three randomly initialized weight matrices Wq, Wk, and Wv to obtain Q, K, and V respectively. Then, the dot product of Q and K is calculated, and the dimension d of the row vector in K k is used to scale the dot product, and then the softmax function is used to normalize the result of the previous step to obtain the attention scores. Then, the attention values are obtained by weighted summing the Values with the attention scores.

[0093] Its calculation formula is:

[0094]

[0095] The softmax function is

[0096]

[0097] The multi-head self-attention mechanism maps Q, K, and V h times respectively in the dimension of d using learnable linear mapping matrices s of ds = d k / h, and then execute the attention function in parallel. Finally, connect the h results and perform a linear mapping again to obtain the self-attention feature vector. The calculation formula is as follows:

[0098] Multihead(Q, K, V) = Concat(head 1 ,..., head h )W o

[0099] head i = Attention(QW i Q , KW i K , VW i V ), i = 1, 2,..., h

[0100] Among them, h represents the number of heads, and W is the linear mapping matrix. The multi-head attention mechanism expands the model's ability to focus on different positions. Multiple linear transformations and concatenations are performed, and the parameters of the linear transformations are learnable, improving the model's fitting ability.

[0101] Specifically, in step S4, the features extracted by WaveNet and the multi-head self-attention mechanism are fused, including:

[0102] Feature fusion is performed by vector concatenation. Concatenate the two groups of parallel output feature vectors to obtain the fused feature representation.

[0103] The minute-level non-intrusive load decomposition model proposed by the present invention is composed of a feature extraction module and a regression module composed of a WaveNet network and a multi-head self-attention mechanism, as Figure 3 shown. The initial construction depends on the manually labeled training data set, which is composed of the minute-level total power sequence of the electricity meter and the corresponding single-appliance power type sequence within the same time period, where the single-appliance power type is the label. Use the minute-level total power sequence of the electricity meter as the model input and the single-appliance power type as the output to train the model. For this purpose, a system needs to be built and deployed based on this model. The system is as Figure 2 shown.

[0104] (1) Feature extraction module: This module is composed of WaveNet and a multi-head self-attention mechanism. WaveNet feature vectors and multi-head self-attention feature vectors are respectively extracted from the minute-level total power sequence of the electricity meter, and then feature fusion is performed.

[0105] (2) Regression module: This module consists of fully connected layers. By inputting the fused feature vector into this module, the target single electrical appliance power type sequence can be obtained.

[0106] The data processing flow is as follows: The minute-level total power sequence of the electric meter is respectively input into WaveNet and the multi-head self-attention mechanism for parallel feature extraction to obtain the WaveNet feature vector and the multi-head self-attention feature vector. Then, the two feature vectors are fused to obtain the fused feature, and the fused feature is input into the fully connected layer for regression to obtain the single electrical appliance power sequence. As Figure 3 shown.

[0107] When the existing load decomposition model is applied to minute-level electric meter data, the load decomposition accuracy will decrease, especially for short-term load electrical equipment. At the same time, data needs to be processed serially. If the input sequence is too long, it is easy to cause information leakage, and the efficiency is low, and problems such as gradient disappearance and gradient explosion are likely to occur. The present invention uses the WaveNet network and the multi-head attention mechanism that can process data in parallel and extract global features to extract features from the power load data. Among them, the WaveNet network consists of causal convolution, dilated convolution, and residual network. Causal convolution performs parallel one-way convolution operations on the minute-level total power time series data of the electric meter. Compared with ordinary convolution, it can improve the training efficiency while ensuring the time series. Dilated convolution can enable the convolutional network to obtain a larger receptive field and capture multi-scale context information. The residual network is used to solve the problems of gradient disappearance and explosion.

[0108] Since there will be missing power features in minute-level data compared with second-level data, in order to strengthen feature extraction, while using WaveNet to extract features, a multi-head self-attention mechanism is introduced to extract features in parallel. The features extracted by the two are fused and then regressed, so as to improve the training efficiency of the load decomposition model and ensure the accuracy of minute-level non-intrusive load decomposition at the same time;

[0109] The results of load decomposition can effectively guide users to change their electricity consumption behaviors, promote the intelligence of the home energy-saving plan, and at the same time provide data support for the power department to formulate corresponding demand-side management, so that home electricity consumption changes towards a more energy-saving and efficient direction.

[0110] Specifically, the model construction method includes:

[0111] Step 1: Obtain the minute-level total power sequence of the electric meter.

[0112] Step 2: Construct the WaveNet network

[0113] WaveNet is essentially a stack of multi-layer causal convolutions and dilated convolutions, using gated convolution operations with Sigmoid and Tanh activation functions. Between each layer, a residual network module is used to take the sum of the input and output of this layer as the input of the next layer, and skip connections are used to concatenate the outputs of each hidden layer. After passing through the Relu activation function, a one-dimensional convolution operation is performed to obtain the final output result. As Figure 1 shown.

[0114] WaveNet uses causal convolutions to prevent data leakage. Causal convolutions mean that the output units at each time t are only obtained from the units before time t through convolution operations. To obtain a flexible receptive field and capture long-term historical information, a sufficiently large convolution kernel size or a sufficiently deep network is required. However, the increase in hidden layers easily leads to gradient vanishing or gradient explosion. Therefore, dilated convolutions and residual networks are introduced.

[0115] Dilated convolutions can enable the convolutional network to obtain a larger receptive field with the same number of hidden layers and convolution kernel size. The receptive field can be calculated by the formula r = 2 L-1 k, where L is the number of layers and k is the filter size.

[0116] Dilated convolutions are equivalent to introducing a fixed stride between two adjacent receptive fields to connect the upper and lower layers. The increase in the dilation rate can make the unit output connect to a wider range of input units, and the receptive field of the convolutional network is effectively expanded. The introduction of the dilation rate idea provides two methods to improve the receptive field of WaveNet. One is to choose a larger filter size, and the other is to increase the dilation rate. Usually, the dilation rate increases exponentially with the number of layers of the convolutional network. In this way, it can ensure that the convolutional network uses the entire sequence as the input and solve the problem of insufficient machine performance caused by the explosive growth of parameters in deep networks by reducing the number of parameters.

[0117] Residual networks (Resnet) are used to solve the problem of gradient vanishing or gradient explosion in deep networks. The idea of residual networks is to add the output result to the input. After passing through the activation function, it is used as the input of the next layer. Since the number of units in the input and output may not be the same, a one-dimensional convolution operation can be performed on the input and then added to the input. This method enables the network to learn not directly the global transformation of the input, but the potential identity relationship or residual term. Residual connections avoid gradient explosion or vanishing, enabling the network to transmit information in a cross-layer manner.

[0118] Step 3: Use the multi-head self-attention mechanism to extract features from the source sequence

[0119] To further enhance the load feature extraction ability, while using WaveNet to extract features, a multi-head self-attention mechanism is introduced to extract features in parallel.

[0120] First, the total power sequence of the minute-level electricity meter is linearly transformed through three randomly initialized weight matrices \(W_q\), \(W_k\), and \(W_v\) to obtain \(Q\), \(K\), and \(V\). Then, calculate the dot product of \(Q\) and \(K\), and scale the dot product with the dimension \(d\) of the row vectors in \(K\). k Normalize the result of the previous step with the softmax function to obtain the attention scores, and then weighted sum \(V\) with the attention scores to obtain the attention values.

[0121] Its calculation formula is:

[0122]

[0123] The multi-head self-attention mechanism maps \(Q\), \(K\), and \(V\) \(h\) times respectively in the dimension of \(d\) with learnable linear mapping matrices. \(d = d / h\), and then execute the attention function in parallel. Finally, concatenate the \(h\) results and perform a linear mapping again to obtain the self-attention feature vector. The calculation formula is as follows: s in the dimension of \(d\) s \(d = d\) k / h, and then execute the attention function in parallel. Finally, concatenate the \(h\) results and perform a linear mapping again to obtain the self-attention feature vector. The calculation formula is as follows:

[0124] \(Multihead(Q, K, V)=Concat(head\) 1 ,..., head\) h )W o

[0125] head i \(=Attention(QW\) i Q , KW\) i K , VW\) i V ), \(i = 1, 2,..., h\)

[0126] where \(h\) represents the number of heads and \(W\) is the linear mapping matrix. The multi-head attention mechanism expands the model's ability to focus on different positions. Multiple linear transformations and concatenations, and the parameters of the linear transformations are learnable, which improves the model's fitting ability.

[0127] Step 4: Fuse the features extracted by WaveNet and the multi-head self-attention mechanism to construct a minute-level non-intrusive load decomposition model

[0128] Feature fusion adopts the method of vector concatenation. Concatenate the two groups of parallel output feature vectors to obtain the fused feature representation, and then input the fused feature vector into the fully connected layer for regression to obtain the target single-appliance power type sequence, thus completing the minute-level non-intrusive load decomposition.

[0129] Step 5: Complete the model deployment and perform data processing and training.

[0130] The initial construction of the minute-level non-intrusive load decomposition model proposed by the present invention relies on a manually labeled training dataset, which consists of the minute-level total power sequence of the electricity meter and the corresponding single-appliance power type sequence within the same time period, where the single-appliance power type is the label. The minute-level total power sequence of the electricity meter is used as the model input, and the single-appliance power type is used as the output to train the model. For this purpose, a system needs to be constructed and deployed based on this model. The system is as Figure 2 shown.

[0131] (1) Feature extraction module: This module consists of WaveNet and the multi-head self-attention mechanism. The WaveNet feature vector and the multi-head self-attention feature vector are respectively extracted from the minute-level total power sequence of the electricity meter by the two, and then feature fusion is performed.

[0132] (2) Regression module: This module consists of a fully connected layer. The fused feature vector is input into this module, and the target single-appliance power type sequence can be obtained.

[0133] The following further elaborates on the technical solution of the present invention in combination with Figure 3 :

[0134] (1) Using machine learning open-source architectures such as TensorFlow, SKlearn, and Numpy, construct a minute-level non-intrusive load decomposition model according to the model structure shown in the figure.

[0135] (2) While using WaveNet to extract features, introduce the multi-head self-attention mechanism to extract features in parallel.

[0136] (3) Concatenate the two groups of parallel output feature vectors to obtain the fused feature representation, and then input the fused feature vector into the fully connected layer for regression processing, and finally obtain the target single-appliance power type sequence.

[0137] Existing load decomposition models are all oriented to second-level or millisecond-level power data. In practical applications, in order to achieve high-frequency data acquisition, it is necessary to transform or add dedicated acquisition equipment, facing difficulties such as high costs and great promotion difficulties. If the existing load decomposition model is directly used for the minute-level data obtained based on the in-use electricity meters of residents and the data acquisition environment, the load decomposition accuracy will be reduced, especially for short-term load electrical equipment. At the same time, data needs to be processed serially. If the input sequence is too long, it is easy to cause information leakage, and the efficiency is low, and problems such as gradient disappearance and gradient explosion are likely to occur. To address these problems, the present invention introduces WaveNet and the multi-head self-attention mechanism to achieve non-intrusive load decomposition based on minute-level electricity meter data. Among them, the causal convolution of WaveNet can achieve sequence modeling and extract context information, ensuring temporality compared with ordinary convolution, and achieving parallel processing of sequences and global feature extraction compared with recurrent neural networks. The dilated convolution of WaveNet can enable the convolutional network to obtain a larger receptive field and capture multi-scale context information, and the residual network can solve the problems of gradient disappearance and gradient explosion. Since there are missing power features in minute-level data compared with second-level data, in order to strengthen feature extraction, while using WaveNet to extract features, the multi-head self-attention mechanism is introduced to extract features in parallel, and the features extracted by the two are fused and then regressed, so as to improve the training efficiency of the load decomposition model while ensuring the accuracy of minute-level non-intrusive load decomposition

[0138] Embodiment 2:

[0139] Provide a minute-level non-intrusive load decomposition system, including:

[0140] Input module: used to obtain the minute-level total power sequence of the electricity meter;

[0141] Decomposition module: used to input the minute-level total power sequence of the electricity meter into the minute-level non-intrusive load decomposition model to obtain the target single-electrical power type sequence, so as to achieve load decomposition;

[0142] The minute-level non-intrusive load decomposition model includes a feature extraction module and a regression module; the feature extraction module includes a WaveNet network and a multi-head self-attention mechanism; the regression module includes a fully connected layer;

[0143] The decomposition module includes:

[0144] WaveNet network module: used to input the minute-level total power sequence of the electricity meter into the WaveNet network for feature extraction to obtain a WaveNet feature vector.

[0145] Multi-head self-attention module: While extracting features using WaveNet, it introduces the multi-head self-attention mechanism to extract features from the minute-level total power sequence of the electricity meter in parallel, obtaining the self-attention feature vector.

[0146] Feature fusion module: It is used to fuse the WaveNet feature vector and the self-attention feature vector to obtain the fused feature vector;

[0147] Output module: It is used to input the fused feature vector into the fully connected layer, and through regression by the fully connected layer, obtain the target single-appliance power type sequence, where the single-appliance power type is the label, thereby realizing load decomposition.

[0148] The system of this embodiment can be used to implement the method described in Embodiment 1.

[0149] Embodiment 3:

[0150] This embodiment provides a minute-level non-intrusive load decomposition system, including a processor and a storage medium;

[0151] The storage medium is used to store instructions;

[0152] The processor is used to operate according to the instructions to execute the steps of the method described in Embodiment 1.

[0153] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0154] The present application is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each flow and / or block in the flowchart and / or block diagram can be implemented by computer program instructions, and the combination of the flows and / or blocks in the flowchart and / or block diagram can also be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate a device for realizing the specified function in Figure 1 one or more flows and / or blocks Figure 1 one or more blocks.

[0155] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing apparatus to function in a particular manner, such that the instructions stored in the computer-readable memory produce a manufacture including an instruction device that implements the functions specified in one or more of the processes and / or blocks Figure 1 one or more of the processes and / or blocks Figure 1 specified in the block or blocks.

[0156] These computer program instructions can also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer-implemented process, whereby the instructions executed on the computer or other programmable apparatus provide steps for implementing the functions specified in one or more of the processes and / or blocks Figure 1 one or more of the processes and / or blocks Figure 1 specified in the block or blocks.

[0157] The foregoing is only a preferred embodiment of the present invention, and it should be noted that for those of ordinary skill in the art, without departing from the technical principle of the present invention, several improvements and modifications can be made, and these improvements and modifications should also be regarded as the protection scope of the present invention.

Claims

1. A minute-level non-intrusive load decomposition method, characterized in that, it includes the following steps: Step S1: Obtain the minute-level total power sequence of the electric meter; Step S2: Input the minute-level total power sequence of the electric meter into the minute-level non-intrusive load decomposition model to obtain the target single-appliance power type sequence; The minute-level non-intrusive load decomposition model includes a feature extraction module and a regression module; the feature extraction module includes a WaveNet network and a multi-head self-attention mechanism; the regression module includes a fully connected layer; Inputting the minute-level total power sequence of the electric meter into the minute-level non-intrusive load decomposition model to obtain the target single-appliance power type sequence includes: Step S21: Input the minute-level total power sequence of the electric meter into the WaveNet network for feature extraction to obtain a WaveNet feature vector; Step S22: While using the WaveNet to extract features, introduce a multi-head self-attention mechanism to parallelly extract features from the minute-level total power sequence of the electric meter to obtain a self-attention feature vector; Step S23: Fuse the WaveNet feature vector and the self-attention feature vector to obtain a fused feature vector; Step S24: Input the fused feature vector into the fully connected layer, and perform regression through the fully connected layer to obtain the target single-appliance power type sequence, thereby realizing load decomposition; The WaveNet network includes a causal convolution, a dilated convolution, a residual network, and a gated structure composed of activation functions of tanh and sigmoid; The causal convolution is used to perform parallel one-way convolution operations on the minute-level total power time series data; The residual network includes multiple residual blocks; The dilated convolution is used to represent the features of the input power in each residual block; The gated structure composed of the activation functions of tanh and sigmoid is used to perform selective non-linear transformation on the input; In step S21, inputting the minute-level total power sequence of the electric meter into the WaveNet network for feature extraction includes: Taking the minute-level total power sequence of the electric meter as the input, obtaining a preliminary feature representation through causal convolution, and then entering multiple layers of residual blocks for further feature extraction; In each residual block, the input power first undergoes feature representation through dilated convolution, and then enters the tanh activation function for non-linear feature transformation to obtain a non-linear transformation result. While performing the non-linear feature transformation, it enters the sigmoid activation function to generate a selective output. Finally, the non-linear transformation result and the selective output are subjected to a Hadamard product; where the selective output is used to control the degree of non-linear transformation; The result of the Hadamard product is convolved with a 1×1 convolution to obtain the output of the gated structure. The output of the gated structure and the original input are superimposed through a residual connection and enter the next residual block, and are used as part of the information of the final output through another residual connection mechanism; The final output part sums the information obtained from multiple residual blocks, and then passes through two layers of 1×1 convolution kernels and the Relu activation function to obtain the final power information output, that is, the WaveNet feature vector.

2. The minute-level non-intrusive load decomposition method according to claim 1, characterized in that, the training method of the minute-level non-intrusive load decomposition model includes: obtaining a manually labeled training data set, which is composed of a minute-level total power sequence of an electric meter and a corresponding single electrical appliance power type sequence within the same time period, where the single electrical appliance power type is the label; using the minute-level total power sequence of the electric meter as the model input and the single electrical appliance power type as the output to train the model, and obtaining a trained minute-level non-intrusive load decomposition model.

3. The minute-level non-intrusive load decomposition method according to claim 2, characterized in that, the training data set is obtained by reading the electric meter data.

4. The minute-level non-intrusive load decomposition method according to claim 1, characterized in that, in step S22, the multi-head self-attention mechanism is used to extract features from the source sequence, including: The total power sequence of the minute-level electricity meter is linearly transformed through three randomly initialized weight matrices W Q , W K and W V to obtain Q, K, and V; where Q is the Query matrix, K is the Key matrix, and V is the Value matrix; Calculate the dot product of Q and K, and use the dimension d of the row vectors in K k Scale the dot product, then normalize the result of the previous step using the softmax function to obtain attention scores, and then weighted sum V using the attention scores Attention to obtain the attention value Attention. Its calculation formula is as follows: where T is the transpose, Q is the Query matrix, K is the Key matrix, V is the Value matrix, and d k is the dimension of the row vectors in K where the softmax function is where z is the input vector of the function, including k values, and e is the natural constant; The multi-head self-attention mechanism maps Q, K, and V h times respectively in the dimension of d s using learnable linear mapping matrices, where d s = d k / h. Then, the attention function is executed in parallel. Finally, the h results are concatenated and linearly mapped again to obtain the self-attention feature vector. The calculation formula is as follows: Multihead(Q, K, V) = Concat(head 1 ,..., head h )W O Among them, h represents the number of heads, Q is the Query matrix, K is the Key matrix, V is the Value matrix, and the linear mapping matrix 5. The minute-level non-intrusive load decomposition method according to claim 1, characterized in that, in step S23, the WaveNet feature vector and the self-attention feature vector are fused to obtain a fused feature vector, including: The feature fusion adopts the method of vector splicing, and the two groups of parallel output feature vectors are spliced to obtain a fused feature representation.

6. A minute-level non-intrusive load decomposition system, characterized in that, including: Input module: used to obtain the minute-level total power sequence of the electric meter; Decomposition module: used to input the minute-level total power sequence of the electric meter into the minute-level non-intrusive load decomposition model to obtain the target single electrical appliance power type sequence, so as to realize load decomposition; The minute-level non-intrusive load decomposition model includes a feature extraction module and a regression module; The feature extraction module includes a WaveNet network and a multi-head self-attention mechanism; the regression module includes a fully connected layer; The decomposition module includes: WaveNet network module: used to input the minute-level total power sequence of the electric meter into the WaveNet network for feature extraction to obtain a WaveNet feature vector; Multi-head self-attention module: used to introduce the multi-head self-attention mechanism to extract features from the minute-level total power sequence in parallel while using WaveNet to extract features, and obtain a self-attention feature vector; Feature fusion module: used to fuse the WaveNet feature vector and the self-attention feature vector to obtain a fused feature vector; Output module: used to input the fused feature vector into the fully connected layer, and perform regression through the fully connected layer to obtain the target single electrical appliance power type sequence, so as to realize load decomposition; The WaveNet network includes a causal convolution, a dilated convolution, a residual network, and a gated structure composed of activation functions of tanh and sigmoid; The causal convolution is used to perform parallel one-way convolution operations on the minute-level total power time series data of the electric meter; The residual network includes a plurality of residual blocks; The dilated convolution is used to represent the features of the input power in each residual block; The gating structure composed of the activation functions of tanh and sigmoid is used to perform selective non-linear transformation on the input; In step S21, the minute-level total power sequence of the electric meter is input into the WaveNet network for feature extraction, including: Taking the minute-level total power sequence of the electric meter as the input, obtaining a preliminary feature representation through causal convolution, and then entering multiple residual blocks for further feature extraction; In each residual block, the input power first performs feature representation through dilated convolution, and then enters the tanh activation function for non-linear feature transformation to obtain the non-linear transformation result. While performing the non-linear feature transformation, it enters the sigmoid activation function to generate a selective output. Finally, the non-linear transformation result and the selective output are subjected to Hadamard product; the selective output is used to control the degree of non-linear transformation; The result of the Hadamard product is convolved by 1×1 convolution to obtain the output of the gating structure. The output of the gating structure and the original input are superimposed through residual connection and enter the next residual block, and are used as part of the information of the final output through another residual connection mechanism; The final output part sums the information obtained from multiple residual blocks, and then passes through two layers of 1×1 convolution kernels and the Relu activation function to obtain the final power information output, that is, the WaveNet feature vector.

Citation Information

Patent Citations

  • End-to-end-based voice navigation method applied to power enterprise customer service

    CN111312228A

  • Non-invasive load decomposition method based on residual convolution and attention mechanism

    CN114722873A