Conv-BERT-BiLSTM-based non-intrusive load intelligent analysis method and device

By using the Conv-BERT-BiLSTM hybrid model, which combines data preprocessing and joint loss function, the problem of decomposing total power load and judging the status of electrical appliances in non-intrusive load monitoring is solved, achieving efficient and accurate load decomposition and status judgment.

CN120873455APending Publication Date: 2025-10-31国网新疆电力有限公司营销服务中心
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510943932.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-09
Publication Date
2025-10-31

AI Technical Summary

Technical Problem

Existing non-intrusive load monitoring technologies are unable to effectively capture the transient characteristics of dynamic loads. Single neural network models are unable to take into account the spatiotemporal characteristics. Furthermore, existing hybrid models have a large number of parameters, high training costs, and poor real-time performance and interpretability.

Method used

A Conv-BERT-BiLSTM hybrid model is adopted, which combines data preprocessing, convolutional layers, BERT layers, BiLSTM layers and transposed convolutional layers to achieve accurate decomposition of total power load and determination of electrical appliance status. The joint loss function of the flexibility factor is used to optimize the model performance.

Benefits of technology

It achieves efficient and accurate non-intrusive load decomposition and electrical appliance status judgment, improving the model's real-time performance and interpretability, and reducing training costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120873455A_ABST
    Figure CN120873455A_ABST
Patent Text Reader

Abstract

The invention discloses a Conv-BERT-BiLSTM (Conv-BERT-BiLSTM)-based non-intrusive load intelligent analysis method and device. The method comprises the following steps: step 1, preprocessing data; 2, a Conv-BERT-BiLSTM model is constructed, the model comprises a convolution layer, a BERT layer, a BiLSTM layer and a transposition convolution layer, the model outputs a single electric appliance energy consumption result decomposed from total power by receiving power total power load time sequence data and power data of each device, and meanwhile, the model judges the operation state of an electric appliance according to comparative analysis of a preset threshold value and real-time data; 3, training the model; and 4, analyzing the non-intrusive load: inputting the preprocessed total power time sequence data, outputting a power decomposition value of each electric appliance, and judging the state of the electric appliance according to the predicted power value. The power consumption value of each electric appliance can be decomposed according to a data set such as the total load of the electric power, and the on-off state of the electric appliance is judged.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to a non-intrusive intelligent load analysis method and device based on Conv-BERT-BiLSTM, belonging to the field of smart power technology. Background Technology

[0002] Current non-intrusive load monitoring technologies mainly include traditional methods based on steady-state feature analysis, as well as models based on machine learning and deep learning. Traditional methods rely on steady-state features, making it difficult to capture the transient characteristics of dynamic loads; machine learning models require manually designed features and have weak generalization ability. Single neural network models struggle to account for spatiotemporal feature correlations, and existing hybrid models have large parameter counts, high training costs, and poor real-time performance and interpretability. Deep learning-based non-intrusive load monitoring technologies mainly improve recognition accuracy through multimodal feature fusion or hybrid architectures, but the technical approach still focuses on single model optimization or shallow feature combination, and has not yet effectively addressed the challenges of identifying complex load dynamic characteristics and low-power devices.

[0003] Currently, the following issues still need to be addressed regarding non-invasive load monitoring:

[0004] (1) Electricity load data is usually complex multidimensional time series data, and the electricity consumption behavior of different devices has different patterns. How to perform effective data preprocessing, including data cleaning, missing value imputation, noise removal and feature extraction, to ensure data quality and model input adaptability is a key issue.

[0005] (2) Based on the hybrid model of Conv-BERT-BiLSTM, the model can accurately decompose the load power of individual electrical appliances from the total load power and realize the judgment of the opening and closing of electrical appliances.

[0006] (3) Appropriate optimization algorithms and parameter tuning methods should be selected to find the best combination of hyperparameters in order to improve model performance. Summary of the Invention

[0007] The purpose of this invention is to provide a non-intrusive intelligent load analysis method and device based on Conv-BERT-BiLSTM, which can decompose the power consumption value of each electrical appliance based on datasets such as total power load and determine the on / off status of the electrical appliances.

[0008] To achieve the above objectives, the technical solution adopted by the present invention is as follows:

[0009] A non-intrusive intelligent load analysis method based on Conv-BERT-BiLSTM includes the following steps:

[0010] Step 1: Data preprocessing, converting the raw electricity consumption data into a format that the model can recognize;

[0011] Step 2: Construct the Conv-BERT-BiLSTM model: This model includes convolutional layers, BERT layers, BiLSTM layers, and transposed convolutional layers. The model receives time-series data of total power load and power data of each device, and outputs the energy consumption results of individual electrical appliances decomposed from the total power. At the same time, the model determines the operating status of electrical appliances based on the comparison and analysis of preset thresholds and real-time data.

[0012] Step 3: Train the model;

[0013] Step 4: Analyze non-intrusive loads: After the model training is completed, input the preprocessed total power time series data, and output the power decomposition value of each appliance. Determine the appliance status based on the predicted power value.

[0014] Furthermore, in step 1, data preprocessing includes:

[0015] (1) Data merging and cleaning: Merge the main power data and the power data of each device by performing an inner join through time index to ensure that data is retained only at the time points where both the main power data and the device data exist, and ensure the time alignment of the data. After data merging, remove the missing values ​​that still exist after merging, and remove rows with a total power of zero or negative values.

[0016] (2) Data normalization: Perform normalization operation on the total power load data, and select mean standardization to normalize the data;

[0017] (3) Determine the window size: Use a window to combine continuous measurements into a sequence as the input to the model;

[0018] (4) Generate tags: Generate power value tags for electrical appliances based on their power data.

[0019] Furthermore, in step 2, the model structure is as follows:

[0020] Convolutional Layers: The convolutional layers extract temporal data features and transform dimensions through three layers of convolution and pooling operations. The first convolutional layer maps single-channel data to 64 channels, capturing local temporal dependencies, and residual convolution enhances expressive power and alleviates gradient vanishing. The second convolutional layer increases the number of channels to 128, capturing multi-scale features, and dilated convolution expands the receptive field. The third convolutional layer further increases the number of channels to 256, providing high-dimensional features for subsequent layers. The pooling layer downsamples the feature sequence, halving its length to reduce computational complexity, and the subsequent one-dimensional convolutional blocks perform feature transformations, preserving key feature information.

[0021] BERT layer: The Transformer block in BERT captures long-distance dependencies between different positions in the feature sequence through a multi-head self-attention mechanism, enhancing the model's ability to perceive global features;

[0022] BiLSTM layer: As a time series feature learning module, the BiLSTM layer captures the sequential dependencies in time series data through its bidirectional structure and learns the contextual features of time series data.

[0023] Transposed convolutional layer: The transposed convolutional layer upsamples the feature sequence through transposed convolution to restore its length; it further extracts features using residual convolutional blocks and channel attention mechanisms to enhance expressive power and highlight key information; 1×1 convolution fuses features and reduces dimensionality; finally, the fully connected layer maps the features to the prediction space, generates decomposition results, and outputs the decomposed power value of the appliance.

[0024] Furthermore, in the convolutional layer of the aforementioned steps:

[0025] In the first convolutional layer, the parameters of the 7×1 convolutional block are configured as follows: 1 input channel, 64 output channels, 7 kernels, 1 stride, 3 padding, and the padding mode is to copy edge values ​​for padding. The convolution operation captures local temporal dependencies in the temporal data.

[0026] Subsequently, residual convolutional blocks are introduced to enhance feature learning capabilities. The residual convolutional block contains two convolutional layers and one residual connection. The parameters of each convolutional layer are 64 channels and the kernel size is 5.

[0027] Finally, a channel attention mechanism is applied to highlight important feature channels. The channel attention mechanism obtains channel-level features through global average pooling and global max pooling, and then generates attention weights through a fully connected layer and a sigmoid activation function.

[0028] In the second convolutional layer, the parameters of the 5×1 convolutional block are configured as follows: 64 input channels, 128 output channels, kernel size of 5, stride of 1, padding of 2, and padding mode of copying edge values. After the 5×1 convolutional block, dilated convolution is used to expand the receptive field and capture dependencies over a longer time range. The parameters of the dilated convolution are 128 channels, kernel size of 3, and dilation rate of 2.

[0029] In the third convolutional layer, the parameters of the convolutional layer are 128 input channels, 256 output channels, 3 kernels, 1 stride, 1 padding, and the padding mode is to copy edge values ​​for padding. The convolution operation can further extract high-level features while keeping the sequence length unchanged.

[0030] After convolutional feature extraction, an average pooling layer is used to downsample the sequence, reducing the sequence length and retaining key features. The parameters of the pooling layer are a pooling window size of 2 and a stride of 2.

[0031] Pooling halves the sequence length while preserving important temporal features. To process the pooled features, a one-dimensional convolutional block is used for feature transformation. The parameters of this convolution are 256 input channels, 256 output channels, a kernel size of 3, a stride of 1, padding of 1, and padding mode of copying edge values.

[0032] Furthermore, the BERT layer uses a two-layer bidirectional Transformer encoder, with each layer containing two attention heads; the multi-head self-attention mechanism allows the model to pay attention to other words in the entire sequence while processing a word, thereby capturing global dependencies.

[0033] In the BiLSTM layer, the input is the power load feature sequence processed by BERT. When the feature sequence After being input into BiLSTM, the forward LSTM processes the features of each time step in chronological order, while the backward LSTM processes them in reverse order starting from the end of the sequence. Finally, the hidden states of the two directions are concatenated together. The output of BiLSTM is concatenated with the output of the BERT module to form the final feature representation.

[0034] Furthermore, in the transposed convolutional layer, upsampling is first performed using transposed convolution to restore the length of the feature sequence. The parameters of the transposed convolutional layer are 256 input channels, 256 output channels, a kernel size of 4, a stride of 2, and padding of 1. After the transposed convolution, a residual convolutional block is used to further enhance the feature representation, and a channel attention mechanism is applied to highlight important feature channels. After the residual convolutional block and the channel attention mechanism, a 1×1 convolutional layer is used for feature fusion and channel reduction. The parameters of this convolutional layer are 256 input channels, 128 output channels, and a kernel size of 1.

[0035] Furthermore, in step 3, the training process is as follows:

[0036] (1) Input: Input the preprocessed total power timing data;

[0037] (2) Conv-BERT-BiLSTM model processing: First, the convolutional layer is used to extract features from the input sequence and the pooling layer is used for downsampling. Then, BERT is used to capture long-term dependencies in the sequence and BiLSTM is used to further process the context features of the time series data. Finally, the output layer is upsampled through transposed convolution and feature fusion is performed. The normalized target electrical power prediction value is output through the fully connected layer.

[0038] (3) Denormalization of output power: The output result is the electrical power decomposition value presented in normalized form. Denormalization is performed on it to restore it to the actual measurable power value.

[0039] (4) Status judgment: Based on the preset power threshold, the decomposed power data is binarized to simplify the complex and ever-changing power data into two distinct status signals, providing information support for subsequent intelligent control and decision-making, and ensuring the efficient operation and precise regulation of the entire system.

[0040] (5) Loss calculation and weight update: During the training process, the predicted value is calculated through forward propagation. The on / off state of the electrical appliance is determined by the predicted value. The predicted decomposition value and state are compared with the true value and true state to calculate the loss function. The gradient of the loss with respect to each model parameter is calculated using the backpropagation algorithm. Finally, the Adam optimization algorithm is used to update the model weights according to the gradient, thereby optimizing the model performance.

[0041] Furthermore, a joint loss function with a flexibility factor is used to simultaneously optimize the performance of power decomposition and state classification. The joint total loss is a weighted sum of four types of losses: KL divergence loss, mean squared error loss, boundary loss, and L1 activation loss.

[0042] Furthermore, in step 4, the continuous energy consumption value is discretized into binary states, namely on and off states, based on the electrical appliance starting power threshold and combined with the state judgment algorithm. In the state judgment, if the device power is briefly lower than the lower power limit and the time below the limit does not exceed the minimum off time, it is not judged as being off. If the device power exceeds the set value for a shorter time than the minimum on time, it is not judged as being on.

[0043] A non-invasive intelligent load analysis device based on Conv-BERT-BiLSTM includes:

[0044] The data preprocessing module is used to convert the raw electricity consumption data into a format that the model can recognize;

[0045] The model training module is used to train the constructed Conv-BERT-BiLSTM model, which includes convolutional layers, BERT layers, BiLSTM layers and transposed convolutional layers. The model receives time-series data of total power load and power data of each device, and outputs the energy consumption results of individual appliances decomposed from the total power. At the same time, the model determines the operating status of appliances based on the comparison analysis between preset thresholds and real-time data.

[0046] A non-intrusive load analysis module is used to determine the status of electrical appliances based on predicted power values.

[0047] The beneficial effects of this invention are as follows:

[0048] (1) Using convolution to extract local features of power load data, BERT's long-distance dependency modeling and BiLSTM's temporal context learning capabilities, an intelligent analysis framework is constructed to achieve accurate and efficient non-intrusive load decomposition.

[0049] (2) Based on real power load data, the Conv-BERT-BiLSTM model was trained and tested to verify its effectiveness in load decomposition and electrical appliance status judgment.

[0050] (3) Using the joint loss function with flexible factors, multiple loss terms are considered, which can simultaneously optimize the performance of power decomposition and state classification. Attached Figure Description

[0051] Figure 1 This is a flowchart of the present invention.

[0052] Figure 2 This is a diagram of the Conv-BERT-BiLSTM model framework.

[0053] Figure 3 This is a diagram of the convolutional layer framework.

[0054] Figure 4 This is a diagram of the BERT layer structure.

[0055] Figure 5 This is a diagram of the BiLSTM layer framework.

[0056] Figure 6 This is a diagram of the transposed convolutional layer architecture.

[0057] Figure 7 Flowchart for training the Conv-BERT-BiLSTM model.

[0058] Figure 8 This is a flowchart for determining the status. Detailed Implementation

[0059] The technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention.

[0060] like Figure 1 A non-invasive intelligent load analysis method and device based on Conv-BERT-BiLSTM includes the following steps:

[0061] Step 1: Data Preprocessing

[0062] Given that household electricity load data encompasses power sampling values ​​and time information from multiple time points, the raw electricity data needs to be converted into a format recognizable by the model. By merging, cleaning, and normalizing the raw household electricity load data, noise can be effectively removed, data scales standardized, and potential patterns uncovered, thus providing high-quality input for model training. The data preprocessing of this invention mainly includes the following four steps:

[0063] (1) Data merging and cleaning. The total load data and the load data of each device were read. The data were resampled at a fixed frequency of 6 seconds, and the mean was used to fill the data within the sampling interval. A forward filling method was used to handle any missing values ​​that might appear after resampling. The main power data and the power data of each device were merged using an inner join through a time index, ensuring that data was retained only at time points where both the main power data and device data were available, thus guaranteeing data time alignment. After data merging, missing values ​​that still existed after merging were removed, and rows with zero or negative total power were removed. A low-power filtering strategy was adopted, forcing weak values ​​less than 5W to zero to eliminate sensor noise and standby power interference. Simultaneously, an upper limit was set on the power values, limiting the power of each channel to within 6000W to avoid abnormally high values ​​caused by lightning strikes or equipment failures interfering with model training.

[0064] (2) Data normalization. In order to avoid interference with model training due to different units, the present invention performs a normalization operation on the total power load data and selects the mean standardization method to normalize the data, as shown in formula (1).

[0065] (1)

[0066] in, This represents the mean. It represents the standard deviation.

[0067] (3) Determine the window size. To capture the temporal characteristics of the load data, a windowing technique is used to combine continuous measurements into a sequence as the input to the model. Based on the operating characteristics of the electrical appliances and the data sampling rate, a window size of 480 sampling points is selected. A sliding window method is used to divide the data into multiple windows, with a step size of 120 sampling points between windows.

[0068] (4) Generate tags. Generate power value tags for the appliances based on their power data. Generate on and off status tags for the appliances based on their power threshold, on / off time, and other characteristics.

[0069] Step 2: Construct the Conv-BERT-BiLSTM model

[0070] The Conv-BERT-BiLSTM model proposed in this invention integrates convolutional neural networks, BERT, and bidirectional long short-term memory networks, such as... Figure 2 As shown.

[0071] This model receives time-series data of total power load and power data of each device, and outputs the energy consumption results of individual appliances extracted from the total power. Simultaneously, the model determines the operating status of appliances based on a comparative analysis of preset thresholds and real-time data. The model mainly consists of the following four layers:

[0072] (1) Convolutional Layers: The convolutional layers extract temporal data features and transform dimensions through three convolutional and pooling operations. The first convolutional layer maps single-channel data to 64 channels, capturing local temporal dependencies. Residual convolutions enhance expressive power and alleviate gradient vanishing. The second convolutional layer increases the number of channels to 128, capturing multi-scale features, and dilated convolutions expand the receptive field. The third convolutional layer further increases the number of channels to 256, providing high-dimensional features for subsequent layers. Pooling layers downsample the feature sequence, halving its length to reduce computational complexity. Subsequent one-dimensional convolutional blocks perform feature transformations, preserving key feature information.

[0073] (2) BERT layer: The Transformer block in BERT captures long-distance dependencies between different positions in the feature sequence through a multi-head self-attention mechanism, enhancing the model's ability to perceive global features.

[0074] (3) BiLSTM layer: The BiLSTM layer, as a time series feature learning module, effectively captures the sequential dependencies in time series data through its bidirectional structure and learns the contextual features of time series data.

[0075] (4) Transposed Convolutional Layer: The transposed convolutional layer upsamples the feature sequence through transposed convolution to restore its length. Residual convolutional blocks and channel attention mechanisms are used to further extract features, enhancing expressive power and highlighting key information. 1×1 convolutions fuse features and reduce dimensionality. Finally, the fully connected layer maps the features to the prediction space, generating decomposition results and outputting the decomposed power value of the appliance.

[0076] The model of this invention uses three layers of convolution to process the original input sequence, with the number of channels gradually increasing from 1 to 64, 128, and 256, as shown below. Figure 3 As shown.

[0077] The first convolutional layer extracts features from the input temporal data, mapping the single-channel input temporal data to a 64-channel feature representation. The parameters of the 7×1 convolutional block are configured as follows: input channel 1, output channel 64, kernel size 7, stride 1, padding 3, and padding mode is to copy edge values ​​for padding. Convolutional operations can capture local temporal dependencies in temporal data, as shown in Equation (2).

[0078] (2)

[0079] in, It is the output feature sequence. These are the convolution kernel weights. It is the input timing data. This represents the convolution operation. It is a bias term. This represents the activation function; this model uses the LeakyReLU activation function.

[0080] Subsequently, the model introduces residual convolutional blocks to enhance feature learning capabilities. A residual convolutional block consists of two convolutional layers and one residual connection. Each convolutional layer has 64 channels and a kernel size of 5. The residual convolutional block is shown in Equation (3).

[0081] (3)

[0082] in, It is the output feature sequence. and These are the weights of the convolutional layer. This indicates a batch normalization operation. This represents the LeakyReLU activation function.

[0083] Finally, the model applies a channel attention mechanism to highlight important feature channels. The channel attention mechanism obtains channel-level features through global average pooling and global max pooling, and then generates attention weights through a fully connected layer and a sigmoid activation function, as shown in Equation (4).

[0084] (4)

[0085] in, It is an output feature sequence with attention weights. It is the input feature sequence. It is a global average pooling operation. It is a max pooling operation. It is a fully connected layer. It is the sigmoid activation function.

[0086] After initial feature extraction and residual learning, the second convolutional layer further deepens the feature representation by increasing the number of feature channels from 64 to 128 to capture more complex temporal patterns. The parameters of the 5×1 convolutional block are configured as follows: 64 input channels, 128 output channels, kernel size of 5, stride of 1, padding of 2, and padding mode of copying edge values. After the 5×1 convolutional block, the model uses dilated convolution to expand the receptive field and capture dependencies over a longer time range. The parameters of the dilated convolution are 128 channels, kernel size of 3, and dilation rate of 2, as shown in Equation (5).

[0087] (5)

[0088] in, It is the output feature sequence. These are the convolution kernel weights. It is the input feature sequence. This indicates the dilated convolution operation. It is the expansion rate. It is a bias term. It uses the LeakyReLU activation function. After the dilated convolution, this convolutional layer applies channel attention again to emphasize important feature channels.

[0089] The third convolutional layer maps the feature channel count from 128 to the hidden layer dimension of the model, which is 256 channels, preparing for subsequent BERT and BiLSTM processing. The parameters of this convolutional layer are: 128 input channels, 256 output channels, kernel size of 3, stride of 1, padding of 1, and padding mode of copying edge values. Convolutional operations can further extract high-level features while maintaining the sequence length.

[0090] After convolutional feature extraction, the model uses an average pooling layer to downsample the sequence, reducing the sequence length while retaining key features. The pooling layer parameters are a pooling window size of 2 and a stride of 2. The average pooling is shown in formula (6).

[0091] (6)

[0092] in, It is the first of the output sequences One element, It is the pooling window size. It is the input feature sequence.

[0093] Pooling halves the sequence length while preserving important temporal features. To further process the pooled features, the model uses a one-dimensional convolutional block for feature transformation. The parameters of this convolutional block are: 256 input channels, 256 output channels, kernel size of 3, stride of 1, padding of 1, and padding mode of copying edge values.

[0094] Through the processing of the above convolution and pooling layers, the model can extract multi-level feature representations from the original time series data, providing rich feature information for the subsequent BERT and BiLSTM layers, thereby effectively capturing complex patterns and long-term dependencies in the time series data.

[0095] The Transformer block in BERT captures long-distance dependencies between different positions in a feature sequence through a multi-head self-attention mechanism, enhancing the model's ability to perceive global features, such as... Figure 4 As shown.

[0096] from Figure 3 As can be seen, the input layer converts the data into a vector representation, which mainly consists of three parts: word embedding, sentence embedding, and position embedding. The input data processed by the model of this invention is a single sequence of data, and there are no multiple sentences or data segments that need to be distinguished. Therefore, the sentence embedding does not play its due role, so the sentence embedding part is removed. The formula of the input layer is shown in formula (7).

[0097] (7)

[0098] in, Indicates word embedding, This represents the positional encoding. In the model of this invention, This refers to the output of the convolutional layer.

[0099] Multi-head self-attention allows the model to focus on other words in the entire sequence while processing a word, thereby capturing global dependencies, automatically learning meaningful features from historical load data, and effectively coping with the diversity and complexity of load changes.

[0100] This model uses a two-layer bidirectional Transformer encoder, with each layer containing two attention heads. To accelerate training and stabilize gradients, BERT uses residual connections and layer normalization techniques, as shown in Equation (8).

[0101] (8)

[0102] in, The output of each layer is normalized, maintaining a mean of 0 and a variance of 1, thereby accelerating model training and improving stability. BERT's powerful contextual understanding capabilities enable it to capture long-term dependencies, accurately decomposing the states of different electrical devices when dealing with electrical loads.

[0103] The BiLSTM layer, as a temporal feature learning module, effectively captures the sequential dependencies in time series data through its bidirectional structure, learning the contextual features of the time series data, such as... Figure 5 As shown.

[0104] The input to BiLSTM is the power load feature sequence processed by BERT. BERT extracts and transforms power load sequence features through a multi-head self-attention mechanism and a feedforward network, but its ability to model the temporal dependencies of the sequence is limited. BiLSTM overcomes this deficiency by modeling the sequence from both positive and negative perspectives, simultaneously capturing past and future information. When the feature sequence... After being input into BiLSTM, the forward LSTM processes the features of each time step in chronological order and updates the hidden state according to formula (9).

[0105] (9)

[0106] The reverse LSTM starts from the end of the sequence and processes it in reverse order, and calculates the hidden state according to formula (10).

[0107] (10)

[0108] Finally, the hidden states in the two directions are spliced ​​together, as shown in formula (11).

[0109] (11)

[0110] in, This indicates a connection operation, and dropout is enabled to prevent overfitting. The output of BiLSTM is concatenated with the output of the BERT module to form the final feature representation. This output incorporates richer temporal features, providing a higher-quality feature representation for subsequent deconvolution operations and further processing by fully connected layers.

[0111] In this model, the transposed convolutional layer is responsible for transforming the feature representation after multiple layers of feature extraction and sequence modeling into the final decomposition result. The transposed convolutional layer progressively upsamples, fuses, reduces dimensionality, and maps the features, such as... Figure 6 As shown.

[0112] The transposed convolutional layer of the model first uses transposed convolution for upsampling to restore the length of the feature sequence. The parameters of the transposed convolutional layer are 256 input channels, 256 output channels, a kernel size of 4, a stride of 2, and padding of 1. The transposed convolution operation can double the length of the feature sequence while keeping the number of channels unchanged, as shown in Equation (12).

[0113] (12)

[0114] in, It is the output feature sequence. These are the transposed convolution kernel weights. It is the input feature sequence. This indicates the transpose convolution operation. It is a bias term. It's the LeakyReLU activation function. The transposed convolution restores the downsampled feature sequence to its original length for subsequent feature fusion and decomposition. Upsampling allows the model to capture dependencies over a longer time span and provides a more detailed feature representation for the final decomposition.

[0115] After transposed convolution, the model uses a residual convolution block to further enhance the feature representation. The model applies a channel attention mechanism to highlight important feature channels. After the residual convolution block and the channel attention mechanism, the model uses a 1×1 convolutional layer for feature fusion and channel reduction. The parameters of this convolution block are 256 input channels, 128 output channels, and a kernel size of 1.

[0116] The purpose of this convolutional layer is to reduce the number of feature channels while weighted fusion of features from different channels, highlighting important features and suppressing unimportant ones. This not only reduces the computational complexity of the model but also enhances its feature representation capabilities. After feature fusion and channel reduction, the model uses two fully connected layers for final feature mapping and decomposition.

[0117] Step 3: Training the model

[0118] The Conv-BERT-BiLSTM model proposed in this invention aims to achieve efficient decomposition of total power load data, and its training process is as follows: Figure 7 As shown.

[0119] (1) Input. The input is the preprocessed total power timing data.

[0120] (2) Conv-BERT-BiLSTM model processing. First, convolutional layers are used to extract features from the input sequence, and pooling layers are used for downsampling. Then, BERT is used to capture long-term dependencies in the sequence, and BiLSTM is used to further process the contextual features of the time series data. Finally, the output layer is upsampled through transposed convolution and feature fusion is performed. The normalized target electrical power prediction value is output through a fully connected layer.

[0121] (3) Denormalization of output power. The output of the model is the electrical power decomposition value presented in normalized form. Although this value can reflect the relative situation of electrical power decomposition, in order to make its output meet the needs of actual application, it is necessary to perform denormalization operation to restore it to the actual measurable power value, so as to provide accurate quantitative basis for subsequent electrical energy consumption assessment, power allocation, etc.

[0122] The specific method for output power inverse normalization and correction is to multiply the predicted value by a preset maximum power threshold for the appliance to obtain the energy consumption value in its original dimension, and then perform physical constraint correction:

[0123] (a) Zero out the predicted energy consumption of less than 5W to eliminate low-power noise interference.

[0124] (b) Truncate predicted values ​​that exceed the maximum power threshold to ensure that the maximum power of the appliance is not exceeded.

[0125] (4) State judgment algorithm. In order to accurately judge the operating state of electrical appliances, the decomposed power data is binarized according to the preset power threshold. Through this processing flow, the complex and ever-changing power data can be simplified into two clear state signals, providing concise and effective information support for subsequent intelligent control and decision-making, and ensuring the efficient operation and precise control of the entire system.

[0126] (5) Loss Calculation and Weight Update. During training, the model calculates the predicted value through forward propagation, uses the predicted value to determine the on / off state of the appliance, and compares the predicted decomposition value and state with the true value and true state to calculate the loss function. The gradient of the loss with respect to each model parameter is calculated using the backpropagation algorithm. Finally, the Adam optimization algorithm is used to update the model weights based on the gradient, thereby optimizing the model performance.

[0127] This invention uses a joint loss function with a flexible factor to flexibly consider multiple loss terms, aiming to simultaneously optimize the performance of power decomposition and state classification. The joint total loss is a weighted sum of four types of losses, namely KL divergence loss, mean squared error loss (MSE loss), margin loss, and L1 activation loss (L1 on loss), as shown in Equation (13):

[0128] (13)

[0129] in, This represents the total combined loss value; This represents the KL divergence loss value; This represents the mean squared error loss value; Indicates the boundary loss value; This represents the L1 activation loss value. This represents a hyperparameter used to adjust the weights of the L1 term. , , , .

[0130] (a) KL divergence loss: used to constrain the alignment of the probability distribution between the model output and the true label, as shown in Equation (14).

[0131] (14)

[0132] in, Represents the sequence of real-state labels. Represents the predicted state sequence. This represents a temperature parameter, which measures the difference between the predicted state distribution and the actual state distribution. By introducing the temperature parameter... This is used to adjust the softmax output, forcing the model to learn the probabilistic characteristics of device power changes. It is especially suitable for power superposition distribution modeling in multi-device concurrent scenarios, enhancing the model's ability to express uncertainty.

[0133] (b) Mean square error loss: directly optimize the reconstruction accuracy of energy consumption value, as shown in formula (15).

[0134] (15)

[0135] in, Represents the total number of time steps. This represents the predicted power value at time step t. Represents the time step The model accurately fits the power amplitude by calculating the point-by-point squared error between the predicted energy consumption and the actual value, with particular attention to the steady-state energy consumption characteristics of high-power equipment.

[0136] (c) Boundary loss: By maximizing the classification boundary, the determination threshold of the device start-up and shutdown status is calibrated as shown in formula (16).

[0137] (16)

[0138] in, Represents the time step The predicted state value, Represents the time step The actual state value is determined. The binarized device state prediction value and the actual state are mapped to a numerical space of -1 and 1, and the consistency between the two is calculated. This loss function can effectively distinguish the critical points of discrete state switching, especially for weak signals of low-power devices, avoiding false triggering caused by noise.

[0139] (d) L1 activation loss: specifically optimized for power surges during device start-up and shutdown, as shown in formula (17).

[0140] (17)

[0141] in, The set of time steps representing whether an appliance is turned on or its status is incorrectly classified; Represents the time step The predicted power value; Represents the time step The true power value is obtained. Data points within the state transition window are filtered by masking, and the L1 norm error between the predicted and true values ​​is calculated to enhance the model's ability to capture transient features. This function uses hyperparameters... The weights are dynamically adjusted to balance the emphasis on steady-state and transient modeling.

[0142] Step 4: Analyze non-invasive loads

[0143] After the model training is complete, a Conv-BERT-BiLSTM model can be obtained. The input is the preprocessed total power time series data, and the output is the power decomposition value of each appliance. The state of the appliance can be determined based on these predicted power values.

[0144] This invention uses an appliance starting power threshold combined with a state judgment algorithm to discretize continuous energy consumption values ​​into binary states, namely, on and off states. The state judgment process is as follows: Figure 8 As shown, state 1 represents on and state 0 represents off.

[0145] According to the flowchart, it can be concluded that if the power of the device is briefly lower than the lower limit and the time below is less than the minimum shutdown time, it is not judged as a shutdown. Similarly, if the power of the device exceeds the set value for a shorter time than the minimum start time, it is not judged as a start, as shown in formula (18). This formula takes into account both power and time factors, and can effectively filter out power fluctuations caused by the instantaneous start or stop of electrical appliances, thereby improving the accuracy of status judgment.

[0146] (18)

[0147] Wherein, represents the operating status value of the target appliance at time t; represents the power value of the target appliance at time t; represents the starting power threshold of the target appliance; represents the duration of the target appliance being turned on; represents the minimum on time of the target appliance; represents the actual off time of the target appliance; and represents the minimum off time of the target appliance.

[0148] Based on the above method, this invention uses the Adam optimizer with a learning rate of 0.0010, a batch size of 96, and 100 iterations. = = When the value is 0.2500, the evaluation index of the model and the various indicators of each electrical appliance are shown in Table 1.

[0149] Table 1. Results of classification metrics under optimal parameters for the Conv-BERT-BiLSTM model.

[0150] Appliance Name Accuracy (%) Accuracy (%) Recall rate (%) F1 score (%) refrigerator 95.24 94.76 92.94 93.84 washing machine 99.81 87.50 99.09 92.93 Micro-wave oven 99.72 79.07 83.10 81.04 dishwasher 99.06 97.60 79.32 87.52 average value 98.45 89.73 88.61 88.83

[0151] The foregoing has shown and described the basic principles, main features, and advantages of the present invention. Those skilled in the art should understand that the above embodiments do not limit the scope of protection of the present invention in any way, and all technical solutions obtained by equivalent substitution or other means fall within the scope of protection of the present invention. Parts not covered in this invention are the same as or can be implemented using existing technology.

Claims

1. A non-intrusive intelligent load analysis method based on Conv-BERT-BiLSTM, characterized in that, Includes the following steps: Step 1: Data preprocessing, converting the raw electricity consumption data into a format that the model can recognize; Step 2: Construct the Conv-BERT-BiLSTM model: This model includes convolutional layers, BERT layers, BiLSTM layers, and transposed convolutional layers. The model receives time-series data of total power load and power data of each device, and outputs the energy consumption results of individual electrical appliances decomposed from the total power. At the same time, the model determines the operating status of electrical appliances based on the comparison and analysis of preset thresholds and real-time data. Step 3: Train the model; Step 4: Analyze non-intrusive loads: After the model training is completed, input the preprocessed total power time series data, and output the power decomposition value of each appliance. Determine the appliance status based on the predicted power value.

2. The non-invasive intelligent load analysis method based on Conv-BERT-BiLSTM according to claim 1, characterized in that, In step 1, data preprocessing includes: (1) Data merging and cleaning: Merge the main power data and the power data of each device by performing an inner join through time index to ensure that data is retained only at the time points where both the main power data and the device data exist, and ensure the time alignment of the data. After data merging, remove the missing values ​​that still exist after merging, and remove rows with a total power of zero or negative values. (2) Data normalization: Perform normalization operation on the total power load data, and select mean standardization to normalize the data; (3) Determine the window size: Use a window to combine continuous measurements into a sequence as the input to the model; (4) Generate tags: Generate power value tags for electrical appliances based on their power data.

3. The non-invasive intelligent load analysis method based on Conv-BERT-BiLSTM according to claim 1, characterized in that, In step 2, the model structure is as follows: Convolutional Layers: The convolutional layers extract temporal data features and transform dimensions through three layers of convolution and pooling operations. The first convolutional layer maps single-channel data to 64 channels, capturing local temporal dependencies, and residual convolution enhances expressive power and alleviates gradient vanishing. The second convolutional layer increases the number of channels to 128, capturing multi-scale features, and dilated convolution expands the receptive field. The third convolutional layer further increases the number of channels to 256, providing high-dimensional features for subsequent layers. The pooling layer downsamples the feature sequence, halving its length to reduce computational complexity, and the subsequent one-dimensional convolutional blocks perform feature transformations, preserving key feature information. BERT layer: The Transformer block in BERT captures long-distance dependencies between different positions in the feature sequence through a multi-head self-attention mechanism, enhancing the model's ability to perceive global features; BiLSTM layer: As a time series feature learning module, the BiLSTM layer captures the sequential dependencies in time series data through its bidirectional structure and learns the contextual features of time series data. Transposed convolutional layer: The transposed convolutional layer upsamples the feature sequence through transposed convolution to restore its length; it further extracts features using residual convolutional blocks and channel attention mechanisms to enhance expressive power and highlight key information; 1×1 convolution fuses features and reduces dimensionality; Finally, the fully connected layer maps the features to the prediction space, generates the decomposition results, and outputs the decomposed power values ​​of the electrical appliances.

4. The non-invasive intelligent load analysis method based on Conv-BERT-BiLSTM according to claim 3, characterized in that, In the convolutional layer of the aforementioned steps: In the first convolutional layer, the parameters of the 7×1 convolutional block are configured as follows: 1 input channel, 64 output channels, 7 kernels, 1 stride, 3 padding, and the padding mode is to copy edge values ​​for padding. The convolution operation captures local temporal dependencies in the temporal data. Subsequently, residual convolutional blocks are introduced to enhance feature learning capabilities. The residual convolutional block contains two convolutional layers and one residual connection. The parameters of each convolutional layer are 64 channels and the kernel size is 5. Finally, a channel attention mechanism is applied to highlight important feature channels. The channel attention mechanism obtains channel-level features through global average pooling and global max pooling, and then generates attention weights through a fully connected layer and a sigmoid activation function. In the second convolutional layer, the parameters of the 5×1 convolutional block are configured as follows: 64 input channels, 128 output channels, kernel size of 5, stride of 1, padding of 2, and padding mode of copying edge values. After the 5×1 convolutional block, dilated convolution is used to expand the receptive field and capture dependencies over a longer time range. The parameters of the dilated convolution are 128 channels, kernel size of 3, and dilation rate of 2. In the third convolutional layer, the parameters of the convolutional layer are 128 input channels, 256 output channels, 3 kernels, 1 stride, 1 padding, and the padding mode is to copy edge values ​​for padding. The convolution operation can further extract high-level features while keeping the sequence length unchanged. After convolutional feature extraction, an average pooling layer is used to downsample the sequence, reducing the sequence length and retaining key features. The parameters of the pooling layer are a pooling window size of 2 and a stride of 2. Pooling halves the sequence length while preserving important temporal features. To process the pooled features, a one-dimensional convolutional block is used for feature transformation. The parameters of this convolution are 256 input channels, 256 output channels, a kernel size of 3, a stride of 1, padding of 1, and padding mode of copying edge values.

5. The non-invasive intelligent load analysis method based on Conv-BERT-BiLSTM according to claim 3, characterized in that, The BERT layer uses a two-layer bidirectional Transformer encoder, with each layer containing two attention heads; the multi-head self-attention mechanism allows the model to pay attention to other words in the entire sequence while processing a word, thereby capturing global dependencies; In the BiLSTM layer, the input is the power load feature sequence processed by BERT. When the feature sequence After being input into BiLSTM, the forward LSTM processes the features of each time step in chronological order, while the backward LSTM processes them in reverse order starting from the end of the sequence. Finally, the hidden states of the two directions are concatenated together. The output of BiLSTM is concatenated with the output of the BERT module to form the final feature representation.

6. The non-invasive intelligent load analysis method based on Conv-BERT-BiLSTM according to claim 3, characterized in that, In the transposed convolutional layer, upsampling is first performed using transposed convolution to restore the length of the feature sequence. The parameters of the transposed convolutional layer are 256 input channels, 256 output channels, a kernel size of 4, a stride of 2, and padding of 1. After the transposed convolution, a residual convolutional block is used to further enhance the feature representation. A channel attention mechanism is applied to highlight important feature channels. After the residual convolutional block and the channel attention mechanism, a 1×1 convolutional layer is used for feature fusion and channel reduction. The parameters of this convolutional layer are 256 input channels, 128 output channels, and a kernel size of 1.

7. The non-invasive intelligent load analysis method based on Conv-BERT-BiLSTM according to claim 1, characterized in that, In step 3, the training process is as follows: (1) Input: Input the preprocessed total power timing data; (2) Conv-BERT-BiLSTM model processing: First, the convolutional layer is used to extract features from the input sequence and the pooling layer is used for downsampling. Then, BERT is used to capture long-term dependencies in the sequence and BiLSTM is used to further process the context features of the time series data. Finally, the output layer is upsampled through transposed convolution and feature fusion is performed. The normalized target electrical power prediction value is output through the fully connected layer. (3) Denormalization of output power: The output result is the electrical power decomposition value presented in normalized form. Denormalization is performed on it to restore it to the actual measurable power value. (4) Status judgment: Based on the preset power threshold, the decomposed power data is binarized to simplify the complex and ever-changing power data into two distinct status signals, providing information support for subsequent intelligent control and decision-making, and ensuring the efficient operation and precise regulation of the entire system. (5) Loss calculation and weight update: During the training process, the predicted value is calculated through forward propagation. The on / off state of the electrical appliance is determined by the predicted value. The predicted decomposition value and state are compared with the true value and true state to calculate the loss function. The gradient of the loss with respect to each model parameter is calculated using the backpropagation algorithm. Finally, the Adam optimization algorithm is used to update the model weights according to the gradient, thereby optimizing the model performance.

8. The non-invasive intelligent load analysis method based on Conv-BERT-BiLSTM according to claim 7, characterized in that, The joint loss function using a flexible factor simultaneously optimizes the performance of power decomposition and state classification. The joint total loss is a weighted sum of four types of losses: KL divergence loss, mean squared error loss, boundary loss, and L1 activation loss.

9. The non-invasive intelligent load analysis method based on Conv-BERT-BiLSTM according to claim 1, characterized in that, In step 4, the continuous energy consumption value is discretized into binary states, namely on and off states, based on the electrical appliance starting power threshold and combined with the state judgment algorithm. In the state judgment, if the device power is briefly lower than the lower power limit and the time below the limit does not exceed the minimum off time, it is not judged as being off. If the device power exceeds the set value for a shorter time than the minimum on time, it is not judged as being on.

10. A non-invasive intelligent load analysis device based on Conv-BERT-BiLSTM, characterized in that, include: The data preprocessing module is used to convert the raw electricity consumption data into a format that the model can recognize; The model training module is used to train the constructed Conv-BERT-BiLSTM model, which includes convolutional layers, BERT layers, BiLSTM layers and transposed convolutional layers. The model receives time-series data of total power load and power data of each device, and outputs the energy consumption results of individual appliances decomposed from the total power. At the same time, the model determines the operating status of appliances based on the comparison analysis between preset thresholds and real-time data. A non-intrusive load analysis module is used to determine the status of electrical appliances based on predicted power values.