Load monitoring method

By adopting a combination architecture of bidirectional time convolution network and channel attention network in non-invasive load monitoring technology, the problems of low recognition accuracy, gradient disappearance and gradient explosion in non-invasive load monitoring technology are solved, and more stable and efficient load monitoring is achieved.

CN119986197APending Publication Date: 2025-05-13SHANGHAI ENEINTEL TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510127103.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-30
Publication Date
2025-05-13

Smart Images

  • Figure CN119986197A_ABST
    Figure CN119986197A_ABST
Patent Text Reader

Abstract

The invention provides a load monitoring method. The load monitoring method specifically comprises a data acquisition step, a data preprocessing step, a model construction step and a training and testing step. According to the load monitoring method, the load characteristics of different electric appliances can be accurately identified, the load decomposition performance is improved, the technical problem of low identification precision of non-intrusive load monitoring in the prior art can be effectively solved, and the technical problems of gradient disappearance and gradient explosion in deep learning of the existing non-intrusive load monitoring technology are solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of deep learning, and in particular to a load monitoring method. Background Art

[0002] In the field of power monitoring, power load monitoring technologies are mainly divided into two types: Intrusive Load Monitoring (ILM) and Non-Intrusive Load Monitoring (NILM). Intrusive load monitoring technology embeds sensors in each electrical appliance to collect equipment operating parameters in real time to achieve accurate monitoring, but it is costly, complex to install and maintain, and may have an adverse effect on equipment performance and life. In contrast, non-intrusive load monitoring technology analyzes the overall power consumption data of smart meters, extracts features with the help of algorithms, and splits the power consumption behavior of electrical appliances without modifying equipment or installing additional sensors. This method is low-cost, easy to deploy, has good adaptability and scalability, and provides important technical support for large-scale power consumption monitoring and energy-saving management.

[0003] At present, non-invasive load monitoring technology has problems with low identification accuracy and lack of time series of load data. Since the identification results of non-invasive load monitoring technology for electrical appliances with low power and similar load curves are not ideal, non-invasive load monitoring technology sometimes has identification errors. At the same time, due to the lack of time series of load data of some electrical equipment, it will affect the accuracy of non-invasive load monitoring technology. In order to solve the accuracy problem of non-invasive load monitoring technology, deep neural networks are introduced into non-invasive load monitoring technology to improve the accuracy of non-invasive load monitoring. However, in the process of deep learning, the feature extraction step still faces problems such as gradient vanishing and gradient explosion, which limits the performance of deep learning models. Summary of the invention

[0004] The present invention provides a load monitoring method to solve the technical problem of low recognition accuracy of non-invasive load monitoring in the prior art, and to solve the technical problem of gradient vanishing and gradient exploding in deep learning of the prior art non-invasive load monitoring technology.

[0005] In order to solve the above problems, the present invention provides a load monitoring method, which specifically includes the following steps. A data acquisition step, non-invasively collecting the power sequence of a total power supply and the power sequence of each power-consuming device under the total power supply; a data preprocessing step, standardizing the power sequence of the total power supply and the power sequence of each power-consuming device obtained in the data acquisition step, so that the power sequences of each power-consuming device are comparable; a model construction step, building a complete network consisting of a bidirectional time convolution network and a channel attention network as a deep learning model, inputting the power sequence of the total power supply obtained in the data preprocessing step into a bidirectional time convolution network, and then inputting the output result of the bidirectional time convolution network into a channel attention network to obtain the predicted value of the power sequence of each power-consuming device; the bidirectional time convolution network includes a forward time convolution network layer and a reverse time convolution network layer;

[0006] The formula for the forward temporal convolutional network layer is:

[0007]

[0008] in is the forward output of the i-th layer at position j, w i,k is the kth filter of the i-th layer, is the output of the corresponding position of the previous layer, d i is the dilation factor of the i-th layer, δ is the ReLU activation function, and b i is bias;

[0009] The formula of the reverse time convolutional network layer is:

[0010]

[0011] in is the backward output of the i-th layer at position j, w i,k is the kth filter of the i-th layer, is the output of the corresponding position of the previous layer in the backward channel, d i is the expansion factor of the i-th layer, δ is the ReLu activation function, b i is bias;

[0012] The load monitoring method also includes a training and testing step, in which the predicted value of the power sequence of each electrical equipment obtained in the power prediction step is trained through a loss function to optimize the deep learning model in the model building step.

[0013] Furthermore, the channel attention network includes a channel compression step and a channel excitation step; the formula of the channel compression step is:

[0014]

[0015] where z i is the compressed output of channel i, x i,j is the input feature map at channel i position j, and H is the height of the input feature map;

[0016] The formula for the channel excitation step is:

[0017] s i =σ(W2δ(W1z i ))

[0018] Among them, s i is the weight factor of channel i, σ and δ are activation functions, W1 and W2 are learnable weights, and z i is the output of the channel compression step.

[0019] Furthermore, the output of the channel attention network is:

[0020] y i,j =s i ·x i,j

[0021] where y i,j is the output of the channel attention network at channel i and position j.

[0022] Furthermore, the data preprocessing step specifically includes the following steps: when there is any missing sequence in the power sequence of an electrical device or the power sequence of the total power supply, and the duration of the missing sequence is less than or equal to a preset time t, the missing sequence is filled by a linear interpolation method; the linear interpolation method specifically includes the following steps: setting the adjacent data points before and after the missing data point to (t1, P1) and (t2, P2) respectively, the time of the missing data point is t, and the filling value is The preset time t has a value range of 5 minutes to 15 minutes. If the duration of the missing sequence is greater than the preset time t, the missing sequence is discarded.

[0023] Furthermore, the data preprocessing step specifically includes the following steps: unifying the sampling frequency of the power sequence of each power-consuming device and the power sequence of the total power supply so that the time step of the power sequence of each power-consuming device is aligned.

[0024] Furthermore, the data preprocessing step specifically includes the following steps: standardizing the power sequence of the total power supply, and the standardization formula is:

[0025]

[0026] Among them, P (t) is the total power at time t, represents the mean of the total power supply, and σ represents the standard deviation of the power.

[0027] Furthermore, the training and testing step specifically includes the following steps: the output result of the power prediction step is optimized by combining the Huber loss function and the Quantile loss function, and the combination formula of the Huber loss function and the Quantile loss function is:

[0028] L(y,f(x))=(1-α)L δ (y,f(x))+αL q (y,f(x))

[0029] Among them, y is the true value, f(x) is the predicted value, and L δ is the Huber loss function, L q is the Quantile loss function, and α is a weight parameter used to weigh the importance of the Huber loss function and the Quantile loss function.

[0030] Furthermore, the Huber loss function is specifically:

[0031]

[0032] Among them, y is the true value, f(x) is the predicted value, δ is used to control the tolerance to outliers. When |yf(x)|≤δ, the loss function is the square error; when |yf(x)|>δ, the loss function becomes linear.

[0033] Furthermore, the Quantile loss function is specifically:

[0034]

[0035] Among them, y is the true value, f(x) is the predicted value, and q is the expected quantile, which determines the way the loss function weights the prediction error in different situations.

[0036] The present invention also includes a data processing device, the data processing device includes a memory, the memory is used to store executable program code. The data processing device also includes a processor, which is used to read the executable program code to run a computer program corresponding to the executable program code to perform the steps in the above load monitoring method.

[0037] The advantage of the present invention is that the load monitoring method provided by the present invention uses an architecture that combines a bidirectional temporal convolutional network and a channel attention network to process the total power supply data. The bidirectional temporal convolutional network can efficiently process time series data and accurately extract long-term dependencies, while successfully alleviating the problems of gradient vanishing and gradient explosion, thereby significantly improving the stability of model training in deep learning. The channel attention network effectively captures the differences in the characteristics of different electrical loads in the channel dimension by recalibrating the importance of features between channels, enabling the model to focus on key features and further enhance the ability to identify electrical appliances. In terms of loss functions, the present invention combines the advantages of Huber loss and Quantile loss, which not only has good robustness to outliers, but also can accurately capture the quantile differences in energy consumption of different electrical appliances, thereby comprehensively optimizing the prediction performance of the model.

[0038] The load monitoring method provided by the present invention can accurately identify the load characteristics of different electrical appliances, improve the load decomposition performance, and can effectively solve the technical problem of low recognition accuracy of non-invasive load monitoring in the prior art, and solve the technical problems of gradient vanishing and gradient exploding in deep learning of the existing non-invasive load monitoring technology. BRIEF DESCRIPTION OF THE DRAWINGS

[0039] Figure 1 is a flow chart of a load monitoring method in an embodiment of the present invention;

[0040] Figure 2 is a schematic diagram of a load monitoring method according to an embodiment of the present invention;

[0041] Figure 3 It is a flow chart of a bidirectional temporal convolutional network combined with a channel attention network in an embodiment of the present invention;

[0042] Figure 4 Schematic diagram of the channel attention network in an embodiment of the present invention. Specific embodiments

[0043] The following describes the preferred embodiments of the present invention with reference to the drawings in the specification to illustrate that the present invention can be implemented. These embodiments can fully introduce the technical content of the present invention to those skilled in the art, making the technical content of the present invention clearer and easier to understand. However, the present invention can be embodied in many different forms of embodiments, and the protection scope of the present invention is not limited to the embodiments mentioned in the text.

[0044] like Figure 1 as well as Figure 2 As shown, this embodiment provides a load monitoring method, and the load monitoring method specifically includes steps S1 to S4.

[0045] Step S1: Data collection step, using smart meters or sub-metering technology and a certain sampling period to collect the power sequence of a household's total meter (i.e., the power sequence of the total power supply) and the power sequence of each electrical device connected to the total meter. In practice, it is necessary to collect the power sequence of each electrical device multiple times, and determine the accuracy and stability of the power sequence of each electrical device by averaging to ensure that reliable raw data can be obtained, thereby ensuring that the deep learning model built later is accurate enough.

[0046] Step S2: a data preprocessing step, in which the power sequence of each electrical device and the power sequence of the total power supply obtained in the data collection step are standardized so that the power sequences of each electrical device are comparable;

[0047] Perform visual analysis on the collected data. Check whether there are any missing sequences in the power sequence of each electrical device and the power sequence of the total power supply. When there is any missing sequence in the power sequence of an electrical device or the power sequence of the total power supply, and the duration of the missing sequence is less than or equal to a preset time t, use linear interpolation to fill in the missing sequence; the linear interpolation method specifically includes the following steps: set the adjacent data points before and after the missing data point to (t1, P1) and (t2, P2), respectively, and the time of the missing data point is t, then the filling value is The filled missing data is used as a kind of noise data for the regularization training of the subsequent model to enhance the robustness of the model to data noise. For data whose missing time exceeds the preset time t, it is deleted from the data set in strict accordance with the data deletion rules, and the deleted data information is recorded in detail to evaluate the potential impact of data missing on the results in subsequent analysis. The preset time t ranges from 5 minutes to 15 minutes. In this embodiment, t is set to 10 minutes. If the value of t is too small, the sample size will be reduced, and if the value of t is too large, the accuracy of the sample will be reduced.

[0048] Then unify the sampling frequency of the power sequence of each electrical device and the power sequence of the total power supply so that the time step of the power sequence of each electrical device is aligned. In the specific implementation, there is a problem that the power sequences of multiple electrical appliances and a single electrical appliance sequence in the collected power sequences of each electrical device have different sampling frequencies. Resampling technology is used here to ensure the consistency of each power sequence. For the problem of inconsistent time steps in the power sequences of each electrical device, linear interpolation or nearest neighbor interpolation is performed on the power sequences of different steps based on the timestamp, so that the power data of each appliance is strictly aligned on the time axis.

[0049] The data preprocessing step specifically includes a standardization step, which standardizes the power sequence of the total power supply. At time t, the bus power P (t) It can be expressed as:

[0050]

[0051] Where N is the total number of load types; s i (t) is the state of load i at time point t, which is mainly divided into two types: 1 when on and 0 when off; p i (t) is the power of load i at time point t; ε(t) is the background noise at time point t.

[0052] The normalization formula is:

[0053]

[0054] Among them, P (t) is the total power at time t, represents the mean of the total power supply, and σ represents the standard deviation of the power.

[0055] Step S3: Power prediction step, such as Figure 3 As shown, a complete network consisting of a bidirectional time convolutional network and a channel attention network is built as a deep learning model, the power sequence of each electrical device obtained in the data preprocessing step is input into a bidirectional time convolutional network, and then the output result of the bidirectional time convolutional network is input into a channel attention network to obtain the predicted value of the power sequence of each electrical device.

[0056] The bidirectional temporal convolutional network consists of multiple convolutional layers, each with a dilated convolution to process data with different temporal receptive fields. The core of the bidirectional temporal convolutional network lies in its bidirectional structure. It consists of two independent unidirectional convolutional neural networks, one for processing past information of the time series and the other for processing future information. The outputs of these two unidirectional networks are then connected and fed into the subsequent fully connected layer for final prediction. This bidirectional design enables the model to make full use of the contextual information of the time series and improve the prediction accuracy. In order to improve the expressive power of the model, multiple convolutional layers are usually used, and the pooling layer is used to reduce the feature dimension, reduce the amount of calculation and prevent overfitting. The bidirectional temporal convolutional network can output the predicted values ​​of multiple time series. The output layer usually uses a linear activation function because the predicted values ​​are usually continuous values.

[0057] The bidirectional temporal convolutional network includes a forward temporal convolutional network layer and a reverse temporal convolutional network layer; the formula of the forward temporal convolutional network layer is:

[0058]

[0059] in is the forward output of the i-th layer at position j, w i,k is the kth filter of the i-th layer, is the output of the corresponding position of the previous layer, d i is the dilation factor of the i-th layer, δ is the ReLU activation function, and b i is bias.

[0060] The formula for the reverse time convolutional network layer is:

[0061]

[0062] in is the backward output of the i-th layer at position j, w i,k is the kth filter of the i-th layer, is the output of the corresponding position of the previous layer in the backward channel, d i is the expansion factor of the i-th layer, δ is the ReLu activation function, b i is bias.

[0063] like Figure 4 As shown in Figure 1, the channel attention network is a technology used in deep learning models to learn the importance of features between channels. The channel attention network includes a channel compression step and a channel excitation step. The formula for the channel compression step is:

[0064]

[0065] where z i is the compressed output of channel i, x i,j is the input feature map at position j of channel i, and H is the height of the input feature map; the channel compression step compresses the feature map of each channel into one value.

[0066] The formula for the channel excitation step is:

[0067] s i =σ(W2δ(W1z i ))

[0068] Among them, s i is the weight factor of channel i, σ and δ are activation functions, W1 and W2 are learnable weights, and z i is the output of the channel compression step. The channel excitation step enables the deep learning model to learn the weight of each channel

[0069] The output of the channel attention network is:

[0070] y i,j =s i ·x i,j

[0071] where y i,jis the output of the channel attention network at channel i and position j. The feature map is recalibrated by multiplying the input feature map by the weight factor of the corresponding channel, thereby enhancing the model's attention to important channel features and improving model performance.

[0072] Finally, the output of the bidirectional temporal convolutional network is given to the channel attention module to determine the importance of each channel. The formula is:

[0073] y=SE(BTCN(x)),

[0074] Among them, y is the output sequence, x is the input sequence, BTCN is a bidirectional temporal convolutional network, and SE is a channel attention module.

[0075] Step S4: training and testing step, the output result of the power prediction step is combined with the Huber loss function and the Quantile loss function to optimize the deep learning model.

[0076] The combined formula of Huber loss function and Quantile loss function is:

[0077] L(y,f(x))=(1-α)L δ (y,f(x))+αL q (y,f(x))

[0078] Among them, y is the true value, f(x) is the predicted value, and L δ is the Huber loss function, Lq is the Quantile loss function, and α is a weight parameter used to weigh the importance of the Huber loss function and the Quantile loss function.

[0079] The Huber loss function is specifically:

[0080]

[0081] Among them, y is the true value, f(x) is the predicted value, and δ is used to control the tolerance to outliers. When |yf(x)|≤δ, the loss function is the square error; when |yf(x)|>δ, the loss function becomes linear. The Huber loss function is robust to outliers: Since the Huber loss function uses the mean square error when the absolute error is small and the absolute error when the absolute error is large, the impact on outliers is relatively small and has a certain degree of robustness.

[0082] The Quantile loss function is specifically:

[0083]

[0084] Where y is the true value, f(x) is the predicted value, and q is the expected quantile, which determines how the loss function weights the prediction error in different situations. The main purpose of the Quantile loss function is to enable the model to predict an interval rather than a single point prediction. Unlike the traditional mean squared error loss, the quantile loss can capture the asymmetry of the distribution, especially when dealing with data with non-uniform error distribution.

[0085] The advantage of the present invention is that the load monitoring method provided by the present invention uses an architecture that combines a bidirectional temporal convolutional network and a channel attention network to process the total power supply data. The bidirectional temporal convolutional network can efficiently process time series data and accurately extract long-term dependencies, while successfully alleviating the problems of gradient vanishing and gradient explosion, thereby significantly improving the stability of model training in deep learning. The channel attention network effectively captures the differences in the characteristics of different electrical loads in the channel dimension by recalibrating the importance of features between channels, enabling the model to focus on key features and further enhance the ability to identify electrical appliances. In terms of loss functions, the present invention combines the advantages of Huber loss and Quantile loss, which not only has good robustness to outliers, but also can accurately capture the quantile differences in energy consumption of different electrical appliances, thereby comprehensively optimizing the prediction performance of the model.

[0086] The load monitoring method provided by the present invention can accurately identify the load characteristics of different electrical appliances, improve the load decomposition performance, and can effectively solve the technical problem of low recognition accuracy of non-invasive load monitoring in the prior art, and solve the technical problems of gradient vanishing and gradient exploding in deep learning of the existing non-invasive load monitoring technology.

[0087] The preferred specific embodiments of the present invention are described in detail above. It should be understood that a person skilled in the art can make many modifications and changes based on the concept of the present invention without creative work. Therefore, any technical solution that can be obtained by a person skilled in the art through logical analysis, reasoning or limited experiments based on the concept of the present invention on the basis of the prior art should be within the scope of protection determined by the claims.

Claims

1. A load monitoring method, characterized in that: The specific steps include: The data collection step is to non-invasively collect a power sequence of a total power supply and a power sequence of each power-consuming device under the total power supply; A data preprocessing step of standardizing the power sequence of the total power supply and the power sequence of each power-consuming device obtained in the data collection step so that the power sequences of the various power-consuming devices are comparable; A model construction step, building a complete network consisting of a bidirectional time convolution network and a channel attention network as a deep learning model, inputting the power sequence of the total power obtained in the data preprocessing step into the bidirectional time convolution network, and then inputting the output result of the bidirectional time convolution network into the channel attention network to obtain a predicted value of the power sequence of each electrical device; The bidirectional temporal convolutional network includes a forward temporal convolutional network layer and a reverse temporal convolutional network layer; the formula of the forward temporal convolutional network layer is: in is the forward output of the i-th layer at position j, w i,k is the kth filter of the i-th layer, is the output of the corresponding position of the previous layer, d i is the dilation factor of the i-th layer, δ is the ReLU activation function, and b i is bias; The formula of the reverse time convolutional network layer is: in is the backward output of the i-th layer at position j, w i,k is the kth filter of the i-th layer, is the output of the corresponding position of the previous layer in the backward channel, d i is the expansion factor of the i-th layer, δ is the ReLu activation function, b i is bias; The training and testing step trains the predicted value of the power sequence of each electrical device obtained in the power prediction step through a loss function to optimize the deep learning model in the model building step.

2. The load monitoring method according to claim 1, characterized in that: The channel attention network includes a channel compression step and a channel excitation step; The formula for the channel compression step is: where z i is the compressed output of channel i, is the input feature map at position j of channel i, and H is the height of the input feature map; the formula of the channel excitation step is: s i =σ(W2δ(W1z i )) Among them, s i is the weight factor of channel i, σ and δ are activation functions, W1 and W2 are learnable weights, and z i is the output of the channel compression step.

3. The load monitoring method according to claim 2, characterized in that: The output of the channel attention network is: y i,j =s i ·x i,j where y i,j is the output of the channel attention network at channel i and position j.

4. The load monitoring method according to claim 1, characterized in that: The data preprocessing step specifically includes the following steps: When there is any missing sequence in the power sequence of an electric device or the power sequence of the total power supply, and the duration of the missing sequence is less than or equal to a preset time t, a linear interpolation method is used to fill the missing sequence; the linear interpolation method specifically includes the following steps: Assume that the adjacent data points before and after the missing data point are (t1, P1) and (t2, P2), and the time of the missing data point is t, then the filling value is The preset time t ranges from 5 minutes to 15 minutes; If the duration of the missing sequence is greater than a preset time t, the missing sequence is discarded.

5. The load monitoring method according to claim 1, characterized in that: The data preprocessing step specifically includes the following steps: The sampling frequencies of the power sequence of each power-consuming device and the power sequence of the total power supply are unified so that the time steps of the power sequence of each power-consuming device and the power sequence of the total power supply are aligned.

6. The load monitoring method according to claim 1, characterized in that: The data preprocessing step specifically includes the following steps: The power series of the standardized total power supply is: Among them, P (t) is the total power at time t, represents the mean of the total power supply, and σ represents the standard deviation of the power.

7. The load monitoring method according to claim 1, characterized in that: The training and testing steps specifically include the following steps: The output result of the power prediction step is optimized by combining the Huber loss function and the Quantile loss function. The combination formula of the Huber loss function and the Quantile loss function is: L(y,f(x))=(1-α)L δ (y,f(x))+αL q (y,f(x)) Among them, y is the true value, f(x) is the predicted value, and L δ is the Huber loss function, L q is the Quantile loss function, and α is a weight parameter used to weigh the importance of the Huber loss function and the Quantile loss function.

8. The load monitoring method according to claim 7, characterized in that: The Huber loss function is specifically: Among them, y is the true value, f(x) is the predicted value, δ is used to control the tolerance to outliers. When |yf(x)|≤δ, the loss function is the square error; when |yf(x)|>δ, the loss function becomes linear.

9. The load monitoring method according to claim 7, characterized in that: The Quantile loss function is specifically: Among them, y is the true value, f(x) is the predicted value, and q is the expected quantile, which determines the way the loss function weights the prediction error in different situations.

10. A data processing device, characterized in that: include: A memory for storing executable program codes; as well as A processor is used to read the executable program code to run a computer program corresponding to the executable program code to execute the steps in the load monitoring method according to any one of claims 1 to 9.