A non-intrusive load decomposition method and system
By decomposing non-invasive loads based on identity mapping and receptive field amplification, the problem of low accuracy of load characteristics is solved, and the accuracy and authenticity of the decomposition results are improved.
Patent Information
- Application Number
- CN201910554344.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2019-06-25
- Publication Date
- 2025-06-17
- Estimated Expiration
- 2039-06-25
AI Technical Summary
Among the existing non-invasive load decomposition methods, the load characteristics are low and difficult to be applied to small electrical appliances and constantly changing electrical appliances.
By obtaining the total load power sequence of the target user, the pre-trained total power sequence and the power characteristic sequence of each load type are decomposed to obtain the power characteristic sequence of each load type. The training process includes identity mapping and receptive field amplification.
The accuracy and accuracy of load decomposition results are improved, the gradient vanishing problem is solved, and the receptive field is increased, making the load decomposition results more authentic.
Smart Images

Figure CN110445126B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of power load decomposition, and in particular to a non-intrusive load decomposition method and system. Background Art
[0002] In recent years, with the development of smart grids, power load decomposition technology has received increasing attention from scholars. Load monitoring mainly includes two types: intrusive and non-intrusive. Traditional load power monitoring is generally an intrusive method. The data obtained by the intrusive method is accurate and reliable, with low data noise, but it is difficult to implement and has low user acceptance. The emergence of non-intrusive load decomposition technology makes up for the defects in intrusive load decomposition, can provide instant power consumption information for the power grid and users, and is of great significance for power suppliers to achieve more refined demand-side response and for power users to understand their own power consumption and reduce power consumption costs. Non-intrusive load decomposition is the process of decomposing the total energy consumption of users into individual devices. However, this method has limitations and is not applicable to small appliances, constantly changing appliances, and appliances that are always in use. Non-intrusive load monitoring relies on a practical and widely applicable software system to collect data using existing software facilities.
[0003] Non-intrusive charge monitoring only installs monitoring equipment at the user's power inlet, collects the total voltage and current of the user, decomposes the total power consumption into the power consumption information of independent appliances through analysis, and then monitors the details of the user's power consumption information.
[0004] The research on non-intrusive charge monitoring mainly includes integer programming, sparse coding algorithm, hidden Markov model, deep long short-term memory network, etc. The sparse coding algorithm can improve the decomposition performance, but this method only targets datasets containing low-resolution data types. In traditional deep neural networks, as the number of network layers increases, the number of parameters increases, and the training difficulty of the model also increases. Most seriously, as the number of network layers increases, the problems of gradient disappearance and model degradation become more and more serious, resulting in low accuracy of the extracted load feature sequence. Summary of the Invention
[0005] In order to solve the problem of low accuracy of the extracted load features in the existing methods, the present invention provides a non-intrusive load decomposition method and system.
[0006] The technical solution provided by the present invention is as follows:
[0007] A non-intrusive load decomposition method, comprising:
[0008] Obtaining the total load power sequence of a target user within a period of time;
[0009] Based on the total load power sequence of the user and the set load types, according to the relationship between the pre-trained total power sequence and the power feature sequences of each load type, decompose the total load power to obtain the power feature sequences of each load type within the time period;
[0010] Among them, the load types are determined by the user's usage of household appliances;
[0011] The training process includes: based on the identity mapping and magnifying the receptive field in chronological order.
[0012] Preferably, the training of the relationship between the total power sequence and the power feature sequences of each load type includes:
[0013] Obtain the historical total load power sequences of multiple users;
[0014] Obtain the usage sequences of each load based on the types of electrical appliances used by the user and matching with the total load power sequence;
[0015] Process the total load power sequence and the usage sequences of each load to obtain training sample data and test sample data;
[0016] Based on the NILM network, train the training sample data to obtain the relationship between the power feature sequences of the load types.
[0017] Preferably, the process of processing the total load power sequence and the usage sequences of each load to obtain training sample data and test sample data includes:
[0018] Perform normalization processing on the usage sequences of each load. By using the maximum-minimum normalization method, map the results to the interval [0,1], and set the power thresholds of each load;
[0019] Decompose the total load power sequence and the normalized load sequences into initial training sequences and initial test sequences;
[0020] Perform sliding processing on the initial training sequences and initial test sequences through the set length and set step size to obtain training sample data and test sample data.
[0021] Preferably, the training sample data is obtained through the following formula:
[0022] Y m:m+n-1 =F(X m:m+n-1 )+e
[0023] Among them, X m:m+n-1 is the total power feature sequence in the training sample data, Y m:m+n-1 is the single-load power feature sequence in the training sample data, F is the mapping function, and e is the Gaussian noise vector;
[0024] The test sample data is obtained by the following formula:
[0025] Y' m:m+n-1 = F(X' m:m+n-1 ) + e
[0026] where X' m:m+n-1 is the total power feature sequence in the test sample data, Y' m:m+n-1 is the single load power feature sequence in the test sample data, F is the mapping function, and e is the Gaussian noise vector.
[0027] Preferably, based on the NILM network, training the training sample data to obtain the power feature sequence relationship of load types includes:
[0028] Step 1: Through the NILM network, perform an identity mapping on the total load power sequence in turn to obtain the total load power sequence with an enlarged receptive field;
[0029] Step 2: Decompose the total load power sequence with an enlarged receptive field into the power feature sequences corresponding to each load type;
[0030] Step 3: Repeat Step 1 and Step 2 until all the total load powers in the training sample data are identity mapped to the power feature sequences corresponding to their respective load types, obtaining the power feature sequence relationship of load types.
[0031] Preferably, through the NILM network, performing an identity mapping on the total load power sequence in turn to obtain the total load power sequence with an enlarged receptive field includes:
[0032] Performing an identity mapping on the total load power sequence through the residual block in the NILM network;
[0033] Based on the identity mapping, enlarging the receptive field of the total load power sequence through the convolution kernel and dilation rate in the residual block.
[0034] Preferably, the receptive field is enlarged by the following formula:
[0035] R = k + (k - 1)×(r - 1)
[0036] where R is the receptive field size of the source sequence, k is the size of the one-dimensional convolution kernel, and r is the dilation rate.
[0037] Preferably, based on the NILM network, training the training sample data to obtain the power feature sequence relationship of load types further includes:
[0038] Using the power feature sequence relationship to obtain the test result for the total load power sequence in the test sample data;
[0039] Compare the test results with the power characteristic sequences of each load in the test sample data to obtain the average error;
[0040] Correct the power characteristic sequence relationship according to the average error to obtain the final power characteristic sequence relationship of the load type.
[0041] Preferably, the average error is calculated by the following formula:
[0042]
[0043] where MAE is the average error, g t is the actual power consumed by the load in the test sample at time t, p t is the power of the load at time t obtained from the power characteristic sequence relationship, and T represents the number of time points.
[0044] Preferably, the decomposition of the total load power to obtain the power characteristic sequence of each load type includes:
[0045] Obtain the gear position characteristics, start-stop characteristics, and operation characteristics of each load type from the total load power through the power characteristic sequence relationship;
[0046] Based on the gear position characteristics, start-stop characteristics, and operation characteristics, and according to the power threshold, obtain the power characteristic sequences of each load type.
[0047] A non-intrusive load decomposition system includes:
[0048] Data acquisition module: Acquire the total load power sequence of the target user over a period of time;
[0049] Decomposition module: Based on the total load power sequence of the user and the set load types, decompose the total load power according to the pre-trained relationship between the total power sequence and the power characteristic sequences of each load type to obtain the power characteristic sequences of each load type during the time period;
[0050] where the load types are determined by the user's usage of household appliances;
[0051] The training process in the decomposition module includes: Based on the identity mapping and magnifying the receptive field in sequence.
[0052] Preferably, the decomposition module includes:
[0053] Multi-user data acquisition sub-module: Acquire the historical multi-user total load power sequence;
[0054] Use the data acquisition sub-module: Acquire the types of electrical appliances used by the user and each load usage sequence that matches the total load power sequence.
[0055] Sample data acquisition sub-module: Process the total load power sequence and each load usage sequence to obtain training sample data and test sample data.
[0056] Training sub-module: Based on the NILM network, train the training sample data to obtain the power feature sequence relationship of the load types.
[0057] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0058] The technical solution provided by the present invention includes: acquiring the total load power sequence of a target user over a period of time; based on the total load power sequence of the user and the set load types, according to the pre-trained relationship between the total power sequence and the power feature sequences of each load type, decomposing the total load power to obtain the power feature sequences of each load type within the time period; wherein, the load types are determined by the user's usage of household appliances; the training process includes: based on the identity mapping and magnifying the receptive field in time sequence. In the present invention, through the pre-trained relationship between the total power sequence and the power feature sequences of each load type, non-intrusive load decomposition is performed, improving the accuracy and precision of the load decomposition result; during the training process, the receptive field method of the total load power is performed in time sequence through the identity mapping, solving the problem of gradient disappearance during the load decomposition process and increasing the receptive field, making the load decomposition result more realistic. Description of the Drawings
[0059] Figure 1 It is a flowchart of a non-intrusive load decomposition method of the present invention;
[0060] Figure 2 It is a network structure diagram of the dilated residual network of the present invention;
[0061] Figure 3 It is a residual structure diagram of the present invention;
[0062] Figure 4 It is a traditional residual structure diagram;
[0063] Figure 5 It is a pre-activated residual network structure diagram of the present invention;
[0064] Figure 6 It is a receptive field effect diagram with a dilation rate of 1;
[0065] Figure 7 It is a receptive field effect diagram with a dilation rate of 2;
[0066] Figure 8It is the effect diagram of the receptive field with a porosity of 3. Detailed implementation mode
[0067] To better understand the present invention, the content of the present invention will be further described below in conjunction with the accompanying drawings of the specification and examples.
[0068] Example 1:
[0069] This example provides a non-invasive load decomposition method, and the method flow chart is as Figure 1 shown.
[0070] S1: Obtain the total load power sequence of the target user within a period of time.
[0071] Collect the long-term power consumption information of various electrical appliances of more than 600 families and the total power consumption of the entire family. The historical power consumption data of air conditioners, refrigerators, washing machines, dishwashers, and microwave ovens in the family are used for the load decomposition task. The main reasons for selecting these several electrical appliances are as follows: (1) The experimental amount of supervised learning on each electrical appliance in each family is huge and unnecessary, while experiments on some common and representative electrical appliances are necessary. (2) Some small power consumption electrical appliances included in the data set are easily affected by noise and are difficult to decompose, so research on such electrical appliances is not carried out. (3) Secondly, the power consumption of these 5 electrical appliances actually accounts for a relatively large proportion of the total consumption of the entire family, that is, decomposition is performed on common electrical appliances. (4) Finally, the historical power consumption of the 5 selected electrical appliances includes pattern decompositions from simple to complex.
[0072] S2: Based on the total load power sequence of the user and the set load types, according to the relationship between the pre-trained total power sequence and the power feature sequences of each load type, decompose the total load power to obtain the power feature sequences of each load type within the time period.
[0073] Data processing is carried out. Different evaluation indicators often have different dimensions and dimension units. In order to eliminate the influence of dimensions between indicators, data normalization processing is required. By adopting the maximum-minimum normalization method, the result is mapped between [0,1]. The normalization function is as follows:
[0074]
[0075] where x t is the normalized total power or the power consumption of the target electrical appliance at the end of time t, X max is the maximum value of the total power sequence, X min is the minimum value of the total power sequence, is the normalization result at time t.
[0076] Deep learning relies on the training of a large amount of data. Therefore, 80% of the data is used as the training sequence, where the training sequence includes the total power sequence X and the power sequence Y of a single electrical appliance. The remaining 20% of the data is used as the test sequence. However, the input of D-ResNet needs to be of a certain size, which requires sliding processing of the original training sequence and test sequence.
[0077] The sliding processing of the training sequence is carried out in an overlapping sliding manner to increase the training samples. Suppose the length of the training sequence is z, and a sliding window with length n and step size u (u < n) is slid on the original sequence to obtain training sample data. Similarly, in a non-overlapping sliding manner, assuming the length of the test sequence is h, test sample data can be obtained. The sequence-to-sequence (Seq2Seq) network maps the total power sample X m:m+n-1 to the corresponding power sample Y of a single electrical appliance m:m+n-1 . Among them, m represents the moment when the data starts to slide, n represents the size of the sliding window, and m:m + n - 1 represents the moment from when the sequence starts to slide to the end. Modeling from the total power data to the power data of a single electrical appliance can be expressed by a formula.
[0078] Y m:m+n-1 = F(X m:m+n-1 ) + e
[0079] In the formula, F represents the mapping function, and e represents the Gaussian noise vector. Then, the total power sequence and the sequence of a single electrical appliance can be corresponded one by one in the form of (x, y) and sent into the network for training, where x and y represent the samples slid from X and Y.
[0080] In the dilated residual network, it includes: dilated convolution.
[0081] Dilated convolution, also known as atrous convolution, can exponentially increase the receptive field without increasing the model parameters or computational complexity. Normally, pooling and downsampling operations will cause information loss, while dilated convolution can replace the pooling effect and multiply the receptive field, enabling each convolution output to contain information in a larger range. Dilated convolution introduces a dilation rate parameter to represent the dilation size.
[0082] The dilated convolution operation can be expressed as a formula, where i ∈ {0, 1, …, L l+1}, represents the dot product operation, x l , x l+1 are the input and output of the (l + 1)-th layer respectively, also known as feature maps. L l+1 is for x l+1The size, x(i) represents the corresponding value at position i on the feature map, c is the number of channels, k is the size of the convolutional kernel, s0 is the size of the convolutional kernel, p is the number of padding 0s, and r is the dilation rate.
[0083] At its essence, dilated convolution is also a dot product operation between the elements on the convolutional kernel and the elements on the feature map. The difference lies in that the convolutional method using dilated convolution performs the dot product operation by skipping some elements, thereby increasing the receptive field.
[0084] The formula for calculating the receptive field of one-dimensional dilated convolution can be expressed as:
[0085] R = k + (k - 1)×(r - 1)
[0086] Among them, R is the receptive field size of the source sequence, k is the size of the one-dimensional convolutional kernel, and r is the dilation rate.
[0087] Using dilated convolution increases the receptive field, captures more power data, and improves the accuracy of load decomposition.
[0088] It also includes a deep residual network.
[0089] The deep residual network solves the problems of performance degradation and gradient disappearance after the network becomes deeper, and can improve the model accuracy, borrowing the idea of cross-layer connection in high-speed networks. If it is simply difficult to use a forward neural network with multiple hidden layers stacked to fit an identity mapping function H(x) = x. However, by designing the network in the form of H(x) = F(x) + x, the function can be converted into a residual function F(x) = H(x) - x. When F(x) = 0, the mapping function H(x) = x is obtained. The residual network is such a structure, which passes the output of the previous layer to the next layer by adding an identity mapping. By introducing the residual structure in the network, a very deep neural network can be built with good performance.
[0090] Assume that the input of a forward neural network is x, and the expected output of the network is H(x). The residual structure replaces the expected output H(x) with F(x) + x by adding an identity transformation. These two expressions have the same training effect on the network, but the optimization difficulties are different. F(x) + x can be achieved through a forward neural network and a shortcut connection. The shortcut connection refers to connecting by crossing one or more layers of the network. The "shortcut connection" only performs an identity mapping and adds its output to the output of the stacked layer.
[0091] Each residual structure has two layers. x represents the input, W1 and W2 represent weight vectors. The expression of the forward neural network in the residual block is as follows, and σ represents the activation function ReLU:
[0092] F(x) = W2σ(W1x)
[0093] By adding an identity mapping, the output of the forward neural network in the left branch and the output of the identity mapping in the right branch are added together and activated by an activation function, thus obtaining the output of the residual structure:
[0094] y = F(x, {W i}) + x
[0095] where F(x, {W i}) represents the residual mapping function to be learned, and W i represents the hidden layer weight matrix. By performing a linear transformation W·x on the input x of the residual structure at the shortcut, the number of input feature maps can be made equal to the number of output feature maps. The specific expression is as follows:
[0096] H(x) = F(x, {W i}) + W s x
[0097] Because x l+1 and the previous layer's x l are purely linearly superimposed, we can obtain:
[0098]
[0099] where x L , x i represent the inputs of the L-th and i-th residual units, and the output is activated by an activation function. It can be seen that the input of the L-th deep residual unit can be expressed as the sum of the input of a certain shallow residual unit and all the complex mappings therein.
[0100]
[0101] From the above equation, it can be seen that during the error backpropagation process, the successive multiplications generated by the traditional error backpropagation network no longer exist, that is, the problem of gradient disappearance has been effectively solved.
[0102] The use of residual blocks deepens the network depth, can extract deep load characteristics in power data, and improves the accuracy of load decomposition.
[0103] During the training phase, a large amount of actual training data contains characteristics such as different gears, start-stop, and operating times, and NILM based on D-ResNet can learn rich electrical characteristics from the training of these data. NILM based on D-ResNet takes the total household power sample x as the input of D-ResNet and the power sample y of a single electrical appliance as the output. During the test phase, D-ResNet, through the mapping relationship between x and y established in the training phase, given an input sample, D-ResNet can decompose the corresponding power consumption of the target electrical appliance through the corresponding mapping relationship.
[0104] To preserve the temporal information as much as possible, no pooling operations are used in D-ResNet, and zero-padding is adopted in the convolution of each layer to make the temporal lengths of the input and output of the dilated convolution equal.
[0105] Using the pre-activated residual structure as the building block, residual units are constructed, then the residual units are built into residual blocks, and finally the residual blocks are stacked to construct the dilated residual convolutional neural network.
[0106] The deep residual network takes the total power sequence as the input. First, it passes through an ordinary convolutional layer to extract the shallow load features. After the convolutional operation, the temporal feature map enters the first residual block, where each residual unit contains 3 dilated convolutions, and operations such as batch normalization and ReLU activation are performed in each residual unit.
[0107] The dilation rates of the convolutions in the first residual block are all 1, and there is not much difference between the dilated convolution and the ordinary one-dimensional convolution in the first residual unit. As the network deepens, the ability of the residual network to extract deep features gradually increases. In the deep residual blocks, the dilation rate of the dilated convolution can be gradually increased to increase the receptive field of the convolutional layer and capture the sequence relationship in the long temporal feature map.
[0108] Each residual unit in the deep dilated residual network structure contains 3 dilated convolutions, and the convolution kernel size of the convolutional layer in each residual unit is 3×1. In the designed network, the sizes of the feature maps obtained after convolution and residual block extraction are the same. The number of residual units in the first residual block is 3, the number of residual units in the second residual block is 4, the number of residual units in the third residual block is 6, and the number of residual units in the fourth residual block is 3. The number of convolution kernels in the first residual block is 30, and the dilation rate is 1. The number of convolution kernels in the second residual block is 40, and the dilation rate is 2. The number of convolution kernels in the third and fourth residual blocks is 50, and the dilation rate is 3. Each residual block concatenates the feature maps obtained from the previous residual block.
[0109] It can be noticed that the number of convolution kernels used in the first residual block and the second residual block is different, so the number of feature maps input to the second residual block and the third residual block is different. Then, in the shortcut connection part, the number of feature maps does not match, and the addition operation cannot be completed. By adding a convolution operation with 40 convolution kernels in the shortcut connection part of the second residual block, the number of input channels and output channels in the second residual block is made equal, and the feature extraction of the entire residual block can proceed smoothly.
[0110] And so on, the convolutional kernel sizes in each residual unit of the hollow residual network are the same as those in residual block 1. The sequence information after feature extraction by multiple residual blocks is stretched into a one-dimensional vector, which is the extracted feature vector. A hidden layer with a number of nodes less than the dimension of the feature vector is used to reduce the dimension of the feature vector, and then the vector can be mapped to the target sequence through a fully connected layer.
[0111] After the prediction is completed, all data needs to be de-normalized to obtain the actual power consumption value. The formula for de-normalization is as follows:
[0112]
[0113] where is the power data obtained by de-normalization, and x pred is the power data predicted by the network.
[0114] Example 2:
[0115] This example provides an explanation of the residual structure used in a non-intrusive load decomposition method.
[0116] (1) Residual structure in the residual network: The residual network draws on the idea of cross-layer connection in high-speed networks. It may be difficult to simply use a forward neural network stacked with one or more hidden layers to fit an identity mapping H(x) = x. However, if the network is designed in the form of H(x) = F(x) + x, the function can be converted into a residual function F(x) = H(x) - x. When F(x) = 0, the mapping function H(x) = x is obtained. The residual network is such a structure that directly transmits the output of the previous layer to the next layer by adding an identity mapping. The residual structure replaces the expected output H(x) with F(x) + x by adding an identity transformation. The two expressions have the same effect on network training, but the optimization difficulty is different. F(x) + x can be implemented through a forward neural network and a shortcut connection. The shortcut connection refers to connecting across one or more layers of the network. The "shortcut connection" only performs an identity mapping and adds its output to the output of the stacked layer.
[0117] The deep residual network is stacked by residual blocks, and the residual blocks are composed of residual units. Each residual unit has a residual structure and includes a ReLU activation function, a convolutional layer, and a Batch Normalization layer.
[0118] (2) Model Structure of Dilated Convolutional Neural Network: In a regular convolutional neural network, the receptive field represents the size of the receptive range of different neurons inside the network for the original information. The larger the value of the receptive field of a neuron, the larger the range of the original information it can access, which also means it contains more global and higher-level semantic features. On the contrary, the smaller the value, the more local and detailed the features it contains. Dilated convolution, also known as Atrous convolution, can exponentially increase the receptive field without increasing the model parameters or computational complexity. Normally, pooling and downsampling operations will lead to information loss, while dilated convolution can replace the pooling function and multiply the receptive field, enabling each convolutional output to contain a larger range of information.
[0119] The schematic diagram of two-dimensional dilated convolution is as Figure 4 shown, while Figure 1 the dilated convolution in Figure 4 is one-dimensional dilated convolution, which is designed for time series data. The comparison diagram between dilated convolution and regular convolution is as
[0120] (3) Non-intrusive Load Decomposition: The power data is processed by overlapping sliding to increase the training samples. The total household power data is used as the input data of the dilated residual network, and the output of the entire network is decomposed into the power data of individual appliances.
[0121] The further design of the non-intrusive load decomposition method based on the dilated residual network lies in that each residual structure in step (1) has two layers. Let x represent the input, W1 and W2 represent weight vectors. The expression of the forward neural network in the residual block is as follows, where σ represents the activation function ReLU:
[0122] F(x) = W2σ(W1x) (1)
[0123] By adding an identity mapping, the output of the forward neural network and the output of the identity mapping are added together and activated by the activation function to obtain the output y of the residual structure:
[0124] y = F(x,{W i}) + x (2)
[0125] where F(x,{W i}) represents the residual mapping function to be learned, and W i represents the hidden layer weight matrix. By performing a linear transformation W·x on the input x of the residual structure at the shortcut, the number of input feature maps and the number of output feature maps can be made equal. The expression is as follows:
[0126] H(x) = F(x, {W i}) + W s x(3)
[0127] Because x l+1 and the previous layer's x l are purely linearly superimposed relationships, and by recursively deriving from Equation (3), the relationships between multiple residual blocks can be obtained:
[0128]
[0129] Among them, x L , x i represent the inputs of the L-th and i-th residual units, and the output is activated through an activation function. From Equation (3), it can be seen that the input of the L-th deep residual unit can be expressed as the sum of the input of a certain shallow residual unit and all complex mappings.
[0130]
[0131] It can be seen from Equation (5) that during the error backpropagation process, the successive multiplications generated by the traditional error backpropagation network no longer exist, that is, the problem of gradient disappearance has been effectively solved.
[0132] The further design of the non-intrusive load decomposition method based on the dilated residual network lies in the improved residual structure in step (1): Equation (3) holds on the condition that x l+1 = y l holds, that is, the input of the current residual unit is the output of the previous layer's residual unit, and no linear or non-linear transformation is performed between the residual units. It can be seen from the residual structure diagram that an activation function is used between the residual units, while x l+1 = y l holds on the condition that the activation function between the residual units needs to be removed, but the necessary activation is essential, which requires redesigning the original residual structure.
[0133] One idea for establishing the identity mapping is to move ReLU before the addition based on the original residual structure. However, simply moving ReLU before the addition will cause the output of the residual function F to be non-negative, while the output of a residual function should be in (-∞, +∞). This causes the signal during forward propagation to be monotonically increasing. This will affect the expressive ability and the results will become worse. We hope that the value of the residual function is in the interval (-∞, +∞).
[0134] ReLU is a non-linear activation function. The activation function is crucial for neural network technology, and the activation function increases the non-linear fitting ability of the neural network.
[0135] Addition is a simple element-wise addition. It is a simple addition operation.
[0136] To more clearly illustrate the principle of residual learning, the residual structure diagram of the present invention is as Figure 3 shown.
[0137] The receptive field effect diagram with a dilation rate of 1 is as Figure 6 shown; the receptive field effect diagram with a dilation rate of 2 is as Figure 7 shown; the receptive field effect diagram with a dilation rate of 3 is as Figure 8 shown; Figures 6 to 8 , which reflects the difference between ordinary convolution and dilated convolution, and applies dilated convolution to time series analysis. A one-dimensional dilated convolution structure is designed to increase the receptive field of the convolution kernel. The calculation of the receptive field sizes of two-dimensional dilated convolution and one-dimensional dilated convolution is shown in Equations (9) and (10).
[0138] The excellent performance of a neural network depends on the training of a large number of parameters. The training process involves two processes: forward propagation and backward propagation. Forward propagation can obtain the predicted value of the neural network. There is a certain error between this predicted value and the true label. The backward propagation process is to reduce the error between the predicted value and the true value through repeated iterations. The backpropagation algorithm is the core of the neural network. Using the gradient descent algorithm, the error can be effectively backpropagated to adjust the weight parameters in each layer of the neural network, making the predicted value approach the true value. The gradient descent algorithm needs to calculate the derivative of the error in each layer with respect to the input. The gradient of a shallow neural network will become smaller and even disappear through backpropagation. As the depth of the network increases, the phenomenon of gradient disappearance will become more obvious.
[0139] Understanding this part of knowledge requires a lot of knowledge of neural networks. To clearly explain this part, it needs to be proven through a large number of formula algorithms. Neural network technology has been widely applied in various fields, and this patent is a further exploration based on neural network technology, without discussing traditional neural networks in too much detail.
[0140] y l = f(y l-1 ) + F(f(y l-1 ), W l ) (6)
[0141] As can be known from the above, x l = f(y l-1 ), x l , y l respectively represent the input and output of the l-th residual unit, and y l-1 represents the output of the (l - 1)-th residual unit.
[0142] H(x) = F(x, {W i}) + W s x (7)
[0143] That is, we obtain:
[0144] x l+1 = x l + F(f(x l ), W l ) (8)
[0145] By readjusting the positions of the activation function and Batch Normalization in the residual structure, the residual structure satisfies both the condition for solving the vanishing gradient x l+1 = y l , and the output of the residual function is (-∞, +∞). The experiments by He Kaiming have shown that by moving the activation function ReLU to the branch of the residual unit, not only does it satisfy the condition of x l+1 = y l , but the shortcut connection branch is also not affected. Moreover, it is optimal among the known structures.
[0146] Moving the activation function ReLU to the branch of the residual unit, the pre-activated residual network structure diagram is as shown in Figure 5 , and the traditional residual structure diagram is as shown in Figure 4 . Referring to the description in Step 1 of the specific implementation manner, pre-activation operations are performed before each convolution in the pre-activated residual unit branch, and then matrix addition is carried out for combination. Such operations not only meet the activation requirements but also eliminate the need for additional activation functions outside the branch.
[0147] The further design of the non-intrusive load decomposition based on the dilated residual network lies in that the dilated convolution in step (2) introduces a dilation rate parameter to represent the size of dilation. When performing dilated convolution, the number of parameters of the convolution kernel remains unchanged, and the size of the receptive field increases exponentially with the increase of the "dilation rate". The calculation formula for the receptive field of two-dimensional dilated convolution can be expressed as:
[0148] R = (k + (k - 1) × (r - 1)) 2 (9)
[0149] In the formula, k is the size of the convolution kernel, r is the dilation rate, and R refers to the receptive field size of the convolution kernel.
[0150] Similarly, we designed one-dimensional dilated convolution, that is, performing dilated convolution operations on one dimension of the time series data. Accordingly, the calculation formula for the receptive field of one-dimensional dilated convolution can be expressed as:
[0151] R = k + (k - 1) × (r - 1) (10)
[0152] The further design of the non-intrusive load decomposition based on the dilated residual network lies in that the dilated residual network model includes: Each residual block in the dilated residual network structure consists of 3 one-dimensional residual units. Each residual unit contains 3 dilated convolutions, and operations such as batch normalization and ReLU activation are performed in each residual unit. The convolution kernel size of the convolutional layer in each residual unit is 3×1. In the network we designed, the feature map sizes obtained through convolution and residual blocks are the same. The numbers of residual units in the first, second, third, and fourth residual blocks are 3, 4, 6, and 3 respectively. The number of convolution kernels in the first residual block is 30, and the dilation rate is 1. The number of convolution kernels in the second residual block is 40, and the dilation rate is 2. The number of convolution kernels in the third and fourth residual blocks is 50, and the dilation rate is 3. Each residual block concatenates the feature map obtained from the previous residual block, and the feature extraction of the entire residual block can proceed smoothly.
[0153] In the present invention, by combining the dilated convolutional neural network and the deep residual network for non-intrusive load decomposition, it is ensured that the network has good generalization performance. The introduction of the residual structure in the residual network enables the construction of a very deep neural network with excellent performance. It solves the problems of performance degradation and vanishing gradients after the network depth becomes deeper, and makes the load decomposition result more accurate, improving the model accuracy. On the training set and the validation set, it is proved that the deeper the network, the smaller the error rate. By using the dilated convolutional network, a larger receptive field can be obtained compared to ordinary convolution, while not increasing the model parameters and computational complexity. On the premise of improving the accuracy, the training and testing speed of samples and the sample classification speed have been greatly improved under the same hardware conditions.
[0154] The network structure diagram of the dilated residual network constructed based on the above method is as Figure 2 shown.
[0155] Embodiment 3:
[0156] This embodiment provides a non-intrusive load decomposition system, including:
[0157] Data acquisition module: Acquire the total load power sequence of the target user for a period of time;
[0158] Decomposition module: Based on the total load power sequence of the user and the set load types, according to the pre-trained relationship between the total power sequence and the power feature sequences of each load type, decompose the total load power to obtain the power feature sequences of each load type within the time period;
[0159] wherein, the load types are determined by the user's usage of household appliances;
[0160] The training process in the decomposition module includes: Based on the identity mapping and magnifying the receptive field in time sequence.
[0161] The decomposition module includes:
[0162] Multi-user data acquisition sub-module: acquiring the historical multi-user load total power sequence;
[0163] Usage data acquisition sub-module: acquiring each load usage sequence based on the types of electrical appliances used by the user and matching with the load total power sequence;
[0164] Sample data acquisition sub-module: processing the load total power sequence and each load usage sequence to obtain training sample data and test sample data;
[0165] Training sub-module: based on the NILM network, training the training sample data to obtain the power feature sequence relationship of the load types.
[0166] The sample data acquisition sub-module includes:
[0167] Normalization unit: performing normalization processing on each load usage sequence, by adopting the maximum-minimum normalization method, mapping the result to the interval [0, 1], and setting the power threshold of each load;
[0168] Decomposition unit: decomposing the load total power sequence and the normalized load sequence into an initial training sequence and an initial test sequence;
[0169] Sliding processing unit: performing sliding processing on the initial training sequence and the initial test sequence through a set length and a set step size to obtain training sample data and test sample data.
[0170] The training sample data is obtained through the following formula in the sliding processing unit:
[0171] Y m:m+n-1 = F(X m:m+n-1 ) + e
[0172] where, X m:m+n-1 is the total power feature sequence in the training sample data, Y m:m+n-1 is the single-load power feature sequence in the training sample data, F is the mapping function, and e is the Gaussian noise vector;
[0173] The test sample data is obtained through the following formula in the sliding processing unit:
[0174] Y' m:m+n-1 = F(X' m:m+n-1 ) + e
[0175] where, X' m:m+n-1 is the total power feature sequence in the test sample data, Y' m:m+n-1For the single-load power feature sequence in the test sample data, F is the mapping function, and e is the Gaussian noise vector.
[0176] The training sub-module includes:
[0177] Identity mapping unit: Through the NILM network, the total load power sequence is sequentially subjected to identity mapping to obtain the total load power sequence with an enlarged receptive field;
[0178] Decomposition unit: Decompose the total load power sequence with an enlarged receptive field into the power feature sequences corresponding to each load type;
[0179] Recurrent unit: Repeatedly run the identity mapping unit and the decomposition unit until all the total load powers in the training sample data are identity mapped to the power feature sequences corresponding to their respective load types, obtaining the power feature sequence relationship of the load types.
[0180] The identity mapping unit includes:
[0181] Mapping sub-unit: Perform identity mapping on the total load power sequence through the residual block in the NILM network;
[0182] Receptive field amplification sub-unit: Based on the identity mapping, amplify the receptive field of the total load power sequence through the convolutional kernel and dilation rate in the residual block.
[0183] In the receptive field amplification sub-unit, the receptive field is amplified by the following formula:
[0184] R = k+(k - 1)×(r - 1)
[0185] Where, R is the receptive field size of the source sequence, k is the size of the one-dimensional convolutional kernel, and r is the dilation rate.
[0186] The training sub-module further includes:
[0187] Testing unit: Pass the total load power sequence in the test sample data through the power feature sequence relationship to obtain the test result;
[0188] Comparison unit: Compare the test result with the power feature sequences of each load in the test sample data to obtain the average error;
[0189] Calibration unit: Calibrate the power feature sequence relationship according to the average error to obtain the final power feature sequence relationship of the load types.
[0190] In the comparison unit, the average error is calculated by the following formula:
[0191]
[0192] wherein, MAE is the mean error, g t is the power actually consumed by the load in the test sample at time t, and p t is the power of the load at time t obtained from the power characteristic sequence relationship, and T represents the number of time points.
[0193] The decomposition module further includes:
[0194] An operating characteristic acquisition sub-module: obtaining the gear position characteristics, start-stop characteristics, and operating characteristics of each load type from the total load power through the power characteristic sequence relationship;
[0195] A power characteristic sequence acquisition sub-module for each load: obtaining the power characteristic sequences of each load type based on the gear position characteristics, start-stop characteristics, and operating characteristics and according to the power threshold.
[0196] Obviously, the described embodiments are some embodiments of the present invention, rather than all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0197] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memories, CD-ROMs, optical memories, etc.) containing computer-usable program codes.
[0198] The present application is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or block in the flowcharts and / or block diagrams, and the combination of processes and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate a device for implementing the functions specified in one process Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.
[0199] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing apparatus to operate in a particular manner, such that the instructions stored in the computer-readable memory produce a manufacture including an instruction device that implements the functions specified in one or more of the processes and / or blocks Figure 1 one or more of the processes and / or blocks Figure 1 specified in the block or blocks.
[0200] These computer program instructions can also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer-implemented process, whereby the instructions executed on the computer or other programmable apparatus provide steps for implementing the functions specified in one or more of the processes and / or blocks Figure 1 one or more of the processes and / or blocks Figure 1 specified in the block or blocks.
[0201] The above are only embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present invention are included within the scope of the claims of the present invention pending approval of the application.
Claims
1. A non-intrusive load decomposition method, characterized in that, Including: Obtain the total load power sequence of the target user over a period of time; Based on the total load power sequence of the user and the set load types, according to the pre-trained relationship between the total power sequence and the power feature sequences of each load type, decompose the total load power to obtain the power feature sequences of each load type within the time period; Among them, the load types are determined by the user's usage of household appliances; The training process includes: based on the identity mapping and magnifying the receptive field in sequence; The training of the relationship between the total power sequence and the power feature sequences of each load type includes: Obtain the historical total load power sequences of multiple users; Obtain the usage sequences of each load based on the types of electrical appliances used by the user and matching with the total load power sequence; Process the total load power sequence and the usage sequences of each load to obtain training sample data and test sample data; Based on the NILM network, train the training sample data to obtain the relationship of the power feature sequences of the load types; The training of obtaining the relationship of the power feature sequences of the load types by training the training sample data based on the NILM network includes: Step 1: Through the NILM network, perform identity mapping on the total load power sequence in sequence to obtain the total load power sequence with the receptive field magnified; Step 2: Decompose the total load power sequence with the receptive field magnified into the power feature sequences of the corresponding load types; Step 3: Repeat Step 1 and Step 2 until all the total load powers in the training sample data are identity mapped to the power feature sequences of their respective corresponding load types, obtaining the relationship of the power feature sequences of the load types; The process of performing identity mapping on the total load power sequence in sequence through the NILM network to obtain the total load power sequence with the receptive field magnified includes: Perform identity mapping on the total load power sequence through the residual block in the NILM network; Based on the identity mapping, magnify the receptive field of the total load power sequence through the convolutional kernel and dilation rate in the residual block; The receptive field is magnified by the following formula: R = k+(k - 1)×(r - 1) Where, R is the receptive field size of the source sequence, k is the size of the one-dimensional convolutional kernel, and r is the dilation rate.
2. The method according to claim 1, characterized in that, The process of processing the total load power sequence and the usage sequences of each load to obtain training sample data and test sample data includes: Perform normalization processing on the usage sequences of each load. By using the maximum-minimum normalization method, map the results to the interval [0,1], and set the power thresholds of each load; Decompose the total load power sequence and the normalized load sequence into initial training sequences and initial test sequences; Perform sliding processing on the initial training sequences and initial test sequences through the set length and set step size to obtain training sample data and test sample data.
3. The method according to claim 2, characterized in that, The training sample data is obtained by the following formula: Y m:m+n-1 = F(X m:m+n-1 ) + e Among them, X m:m+n-1 is the total power feature sequence in the training sample data, Y m:m+n-1 is the single load power feature sequence in the training sample data, F is the mapping function, and e is the Gaussian noise vector; The test sample data is obtained by the following formula: Y' m:m+n-1 = F(X' m:m+n-1 ) + e Among them, X' m:m+n-1 is the total power feature sequence in the test sample data, Y' m:m+n-1 is the single load power feature sequence in the test sample data, F is the mapping function, and e is the Gaussian noise vector.
4. The method according to claim 1, characterized in that, The training of obtaining the relationship of the power feature sequences of the load types by training the training sample data based on the NILM network further includes: Obtain a test result by means of the power characteristic sequence relationship for the total load power sequence in the test sample data; Compare the test result with the power characteristic sequences of each load in the test sample data to obtain an average error; Correct the power characteristic sequence relationship according to the average error to obtain the final power characteristic sequence relationship of the load types.
5. The method according to claim 4, characterized in that, The average error is calculated by the following formula: Among them, MAE is the mean absolute error, and g t is the actual power consumed by the load in the test sample at time t, and p t is the power of the load at time t obtained from the power characteristic sequence relationship, and T represents the number of time points.
6. The method according to claim 2, characterized in that, The decomposition of the total load power to obtain the power characteristic sequences of each load type includes: Obtain the gear position characteristics, start-stop characteristics, and operation characteristics of each load type by means of the power characteristic sequence relationship for the total load power; Based on the gear position characteristics, start-stop characteristics, and operation characteristics, obtain the power characteristic sequences of each load type according to the power threshold.
7. A system for implementing the non-intrusive load disaggregation method according to claim 1, characterized in that, The system includes: Data acquisition module: Acquire the total load power sequence of the target user over a period of time; Decomposition module: Based on the total load power sequence of the user and the set load types, decompose the total load power according to the pre-trained relationship between the total power sequence and the power characteristic sequences of each load type to obtain the power characteristic sequences of each load type during the period; wherein, the load types are determined by the user's usage of household appliances; The training process in the decomposition module includes: performing receptive field amplification based on the identity mapping and in chronological order.
8. The system according to claim 7, characterized in that, The decomposition module includes: Multi-user data acquisition sub-module: Acquire the historical total load power sequences of multiple users; Usage data acquisition sub-module: Acquire the usage sequences of each load based on the types of electrical appliances used by the user and matching the total load power sequence; Sample data acquisition sub-module: Process the total load power sequence and the usage sequences of each load to obtain training sample data and test sample data; Training sub-module: Based on the NILM network, train the training sample data to obtain the power characteristic sequence relationship of the load types.
Citation Information
Patent Citations
DFHSMM-based non-intrusion type electric power load monitoring method and system
CN106600074A
Non-intrusive household appliance load identification method
CN106646026A