Non-intrusive load identification method and system based on threshold adjustment

By using CNN to extract features in non-invasive load identification technology and combining threshold adjustment algorithms, the problem of unknown electrical appliance recognition is solved, which significantly improves the accuracy and robustness of load identification.

CN120070982APending Publication Date: 2025-05-30STATE GRID JIANGSU ELECTRIC POWER CO LTD MARKETING SERVICE CENT +2
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510143669.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-10
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

Existing non-invasive load identification technology is difficult to accurately identify and classify unknown electrical appliances, resulting in a decrease in recognition rate and possible misidentification or misreport.

Method used

A method based on convolutional neural network (CNN) is used to extract high-dimensional features, and a load identification open set recognition algorithm based on threshold adjustment is introduced to build a load identification model with both accuracy and generalization to identify and classify unknown electrical appliances.

Benefits of technology

It effectively improves the accuracy and robustness of load identification, improves the accuracy of identification of unknown electrical appliances, from 71.1% to 89.1%, while maintaining high identification accuracy of known electrical appliances.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120070982A_ABST
    Figure CN120070982A_ABST
Patent Text Reader

Abstract

The invention discloses a non-intrusive load identification method and system based on threshold adjustment, and the method comprises the steps: dividing a single electric appliance power image sample set, constructing a first training sample set which only contains known electric appliance samples, and constructing a second training sample set which retains the known electric appliance samples and unknown electric appliance samples at the same time; training through the first training sample set to obtain a CNN load pre-classification model based on a convolutional neural network; based on the second training sample set, training by combining a load identification open set identification algorithm of threshold adjustment to obtain a load identification calibration model; the method comprises the following steps: preprocessing real-time total power utilization data collected in a target user residence to obtain single electric appliance power image data, and performing classification and identification on the single electric appliance power image data through a CNN load pre-classification model and a load identification calibration model in sequence to obtain an electric appliance type identification result of a current access load of the target user residence. According to the method, known electric appliances can be effectively identified, unknown electric appliances can be classified, and the accuracy and robustness of load identification are effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of power load identification, and in particular, to a non-invasive load identification method and system based on threshold adjustment. Background Art

[0002] Load identification is a key research area in the power system, aiming to identify and classify different types of electrical loads by analyzing power consumption data. Load identification methods can be divided into non-invasive load identification (NILM) and invasive load identification (ILM). Among them, non-invasive load identification has broad application prospects in fields such as home energy management, smart grid construction, and carbon emission monitoring due to its advantages of low cost, high efficiency, and convenience. Non-invasive load identification usually relies on pre-collected and labeled training data for model training. These training data cover the power usage characteristics of known appliances, and the model learns and identifies these characteristics during the training process. By analyzing the power usage data, the operating states of different appliances are identified, thereby optimizing energy consumption and improving the efficiency of power management.

[0003] Currently, there have been many studies on non-invasive load identification technology. For example, some technologies propose a non-invasive residential load monitoring method based on fine recognition of U-I trajectory curves, which uses a goodness-of-fit test to capture appliance switching events and extracts three types of features, namely active / reactive power changes and U-I trajectories, to achieve rough load identification and identify blind-zone loads with fine granularity. Other technologies develop a steady-state monitoring model based on AOGA, convert the power parameter model into active and reactive power components, and establish a dual-objective function to solve the monitoring error problem, and further improve the accuracy of transient load monitoring by using harmonic characteristics. The prior art also proposes a NILM method based on the LSTM network to obtain user load characteristic data, uses PCA to reduce the number of load characteristics, improves the calculation efficiency, and can accurately identify household appliances including low-power appliances.

[0004] However, from the existing research, the main goal of current non-intrusive load identification technology is still to improve the accuracy of load identification, and there is a lack of effective methods for identifying unknown types of electrical appliances. However, although the existing non-intrusive load identification technology can identify the samples in the known electrical appliance training set, in real application scenarios, the system may encounter new electrical appliances that do not appear in the training set. The existence of these unknown electrical appliances makes it difficult for traditional load identification technology to accurately identify and classify, resulting in a decrease in the recognition rate, and even may be misidentified as known electrical appliances, which is likely to cause false alarms or missed alarms. Therefore, for the existing load identification methods that often focus on how to improve the identification accuracy and perform well on closed data sets, but do not consider the problem of unknown electrical appliances, how to improve the adaptability of the algorithm to new electrical appliances and the robustness in a changing environment has become an urgent problem to be solved. Summary of the Invention

[0005] To solve the deficiencies in the prior art, the present invention provides a non-intrusive load identification method and system based on threshold adjustment, which uses a convolutional neural network CNN to extract high-dimensional features, introduces a load identification open-set recognition algorithm based on threshold adjustment, constructs a load identification model with both accuracy and generalization, and on the premise of effectively identifying known electrical appliances, detects and classifies unknown electrical appliances, thereby overcoming the limitations of traditional load identification algorithms in dealing with unknown electrical appliances and effectively improving the accuracy and robustness of load identification.

[0006] The present invention adopts the following technical solutions.

[0007] In the first aspect, the present invention provides a non-intrusive load identification method based on threshold adjustment, and the method includes:

[0008] Step 1: Collect the total power consumption data of the user's residence for training, and perform preprocessing of time series image conversion to obtain a single-electrical-appliance power image sample set containing various electrical appliances;

[0009] Step 2: Divide the single-electrical-appliance power image sample set to construct a first training sample set containing only known-class electrical appliance samples and a second training sample set that retains both known-class and unknown-class electrical appliance samples;

[0010] Step 3: Train a CNN load pre-classification model based on the convolutional neural network through the first training sample set;

[0011] Step 4: Based on the second training sample set, train a load identification calibration model in combination with a load identification open-set recognition algorithm with threshold adjustment;

[0012] Step 5: The single - appliance power image data obtained after pre - processing the real - time total power consumption data collected in the target user's residence is successively classified and recognized by the CNN load pre - classification model and the load identification and calibration model to obtain the identification result of the types of appliances of the currently connected load in the target user's residence.

[0013] Optionally, in step 1, the steps of performing pre - processing of time - series image formation include:

[0014] S1.1: Label the switch states of various appliances corresponding to the collected total power consumption data to obtain an original data set;

[0015] S1.2: Clean the abnormal data in the original data set and select stable data immediately following the abnormal data for replacement to obtain a complete and available initial sample set;

[0016] S1.3: Calculate the effective power of each cycle according to the voltage and current information of each total power consumption data in the initial sample set to obtain the one - dimensional time series of the total power corresponding to each sample;

[0017] S1.4: Referring to the labeled switch states of the appliances, intercept the jump windows of each one - dimensional time series of the total power and perform normalization processing;

[0018] S1.5: Use the Gram - Angle - Field (GASF) algorithm to convert the sequence data of each normalized jump window to the polar coordinate system respectively to obtain a single - appliance power image sample set containing various appliances.

[0019] Optionally, the known types of appliances include washing machines, televisions, refrigerators, computers, hair dryers, air conditioners, rice cookers, printers, electric fans, floor sweepers, water heaters, humidifiers, microwave ovens, wall - breaking machines, and / or air purifiers;

[0020] The unknown types of appliances include electric garage doors, air compressors, and / or DVR hard disk recorders.

[0021] Optionally, the structure of the CNN load pre - classification model includes a first convolutional layer group, a first pooling layer, a second convolutional layer group, a second pooling layer, a third convolutional layer group, a third pooling layer, a dropout layer, a flattening layer, a fully - connected layer, and an output layer connected in sequence;

[0022] After the input image sample is extracted with primary features by the first convolutional layer group, it is subjected to primary feature dimensionality reduction processing by the first pooling layer. After the intermediate features are extracted by the second convolutional layer group, it is subjected to intermediate feature dimensionality reduction processing by the second pooling layer. After the high-level features are extracted by the third convolutional layer group, it is subjected to high-level feature dimensionality reduction processing by the third pooling layer. Subsequently, the high-level features after dimensionality reduction are sequentially processed by a dropout layer and a flattening layer, and then input to a fully connected layer to learn global features to obtain an activation vector. Finally, after the activation vector is processed by an output layer with Softmax activation, the pre-classification category of the input image sample is finally obtained.

[0023] Optionally, the objective function adopted by the CNN load pre-classification model during training is: cross-entropy loss function.

[0024] Optionally, in step 4, the steps of training the load identification calibration model by the load identification open set recognition algorithm combined with threshold adjustment include:

[0025] S4.1: Calculate the mean activation vector of each category according to the activation vector output by the fully connected layer in the CNN load pre-classification model for the correctly classified samples, and use the Weibull distribution fitting to obtain the cumulative distribution function of each category;

[0026] S4.2: Based on the mean activation vector and cumulative distribution function of each category, calculate the sum of the corrected probability vector elements of the samples in each category in the second training sample set, and sort them in descending order to obtain the element sum sequence of each category;

[0027] S4.3: Based on the current value of n, respectively take the element sum S corresponding to the first n% boundary in the element sum sequence of each category * y as the threshold α of its category i , and obtain the threshold vector β corresponding to the value of n;

[0028] S4.4: For any second training sample y, compare the sum S of the corrected probability vector elements of the sample y y with the threshold α corresponding to its pre-classified category i i . If S y is greater than α i , then classify the sample y into category i, otherwise classify it into the unknown category, so as to obtain the corrected classification result of the second training sample set;

[0029] S4.5: Based on the corrected classification result, calculate the classification accuracy of the known and unknown samples in the second training sample set;

[0030] S4.6: Repeat steps S4.3 to S4.5 to calculate the average classification accuracy of the known classes and the unknown classes for different values of n, and use the threshold vector corresponding to the value of n with the highest average accuracy as the finally determined threshold vector, so as to obtain the finally trained load identification and calibration model.

[0031] Optionally, the specific steps of S4.1 include:

[0032] S411: Input the first training sample set into the trained CNN load pre-classification model for pre-classification, and compare the pre-classification result with its corresponding true value to screen out the correctly classified sample sets of various electrical appliances;

[0033] S412: Calculate the mean activation vector of each category according to the activation vectors output by the fully connected layer in the CNN load pre-classification model for the correctly classified sample sets of various electrical appliances;

[0034] S413: Calculate the Euclidean distance between each sample in each correctly classified sample set and the mean activation vector of its corresponding category;

[0035] S414: Fit the Weibull distribution to the Euclidean distances of each sample in each correctly classified sample set to obtain the cumulative distribution function of each category.

[0036] Optionally, the calculation formula of the mean activation vector of each category is as follows:

[0037]

[0038] where, MV i is the mean activation vector of category i, N i represents the number of correctly classified samples in category i, V x is the activation vector output by the xth correctly classified sample in the fully connected layer, i = 1, 2... c, and c is the total number of known class electrical appliance categories.

[0039] Optionally, the expression of the cumulative distribution function of each category is as follows

[0040]

[0041] where, F(d xi ; k; λ) represents the cumulative distribution function of category i, d xi represents the Euclidean distance between the activation vector V x of the correctly classified sample x and the mean activation vector MV i of its belonging category i, k is the shape parameter of the Weibull distribution, and λ is the scale parameter of the Weibull distribution.

[0042] Optionally, the Euclidean distance d xi is calculated as follows:

[0043]

[0044] wherein, V x is the activation vector output by the x-th correctly classified sample in the fully connected layer, and MV i is the mean activation vector of class i.

[0045] Optionally, in S4.2, the step of separately calculating the sum of the corrected probability vector elements of each class of samples in the second training sample set includes:

[0046] S421: For any second training sample y, input it into the trained CNN-based pre-classification model to obtain its corresponding activation vector V y and the pre-classified class i;

[0047] S422: Separately calculate the Euclidean distance between the activation vector V y and the mean activation vectors of each class to obtain the Euclidean distance vector ε y of the sample y;

[0048] S423: Substitute the Euclidean distance vector ε y into the cumulative distribution functions of each class to calculate the corrected probability vector of the sample y, thereby obtaining the sum S y of the corrected probability vector elements of the sample y.

[0049] In a second aspect, the present invention provides a non-intrusive load identification system based on threshold adjustment, which operates according to the steps of any one of the first aspects of the present invention. The system includes:

[0050] A data acquisition module, configured to acquire the total power consumption data of a user's residence for training, and perform preprocessing of time series image formation to obtain a single-appliance power image sample set including various appliances;

[0051] A sample division module, configured to divide the single-appliance power image sample set, and construct a first training sample set containing only known-class appliance samples and a second training sample set that simultaneously retains known-class and unknown-class appliance samples;

[0052] A first training module, configured to train a CNN load pre-classification model based on the convolutional neural network through the first training sample set;

[0053] A second training module, configured to identify and calibrate a load identification model based on the second training sample set and in combination with a load identification open-set recognition algorithm with threshold adjustment;

[0054] The load identification module is used to obtain the identification result of the types of electrical appliances of the currently connected load in the target user's residence. The single-electrical-appliance power image data obtained by preprocessing the real-time total power consumption data collected in the target user's residence is successively classified and identified by the CNN load pre-classification model and the load identification calibration model.

[0055] Thirdly, the present invention provides a terminal, including a processor and a storage medium;

[0056] The storage medium is used to store instructions;

[0057] The processor is used to operate according to the instructions to execute the steps of the method described in any one of the first aspects of the present invention.

[0058] Fourthly, the present invention provides a computer-readable storage medium, on which a computer program is stored. When the program is executed by a processor, the steps of the method described in any one of the first aspects of the present invention are implemented.

[0059] The beneficial effects of the present invention are as follows. Compared with the prior art,

[0060] 1. The present invention integrates the convolutional neural network of deep learning to construct a CNN load pre-classification model to pre-classify load data, and proposes a load identification open-set recognition algorithm based on threshold adjustment to further calibrate the results of pre-classification by the load identification calibration model. While identifying known classes, it can effectively detect and process unknown classes, thus overcoming the limitations of traditional load identification algorithms in dealing with unknown electrical appliances.

[0061] 2. The present invention trains known electrical appliance classes and corrects the probability distribution, combined with a threshold adjustment strategy, to reduce the influence of extreme samples in the training data, so as to make a reasonable judgment when facing unknown electrical appliances, prevent misclassifying new electrical appliances as known classes, and avoid the recognition failure caused by insufficient training data or overly single electrical appliance types in traditional algorithms.

[0062] 3. Experimental results show that while maintaining a relatively high recognition accuracy for known electrical appliances, the present invention improves the recognition accuracy of unknown electrical appliances from 71.1% to 89.1%, successfully balancing the recognition performance of known and unknown classes, and effectively improving the accuracy and robustness of the load identification system; by integrating the feature extraction method of load identification and the unknown detection technology of the load identification open-set recognition algorithm based on threshold adjustment, an intelligent load identification system with both accuracy and generalization can be constructed, thereby enhancing the perception and response capabilities of the smart grid to complex power environments. Description of the Drawings

[0063] Figure 1 It is a schematic flowchart of the non-intrusive load identification method based on threshold adjustment in the embodiment of the present invention;

[0064] Figure 2 It is a logical schematic diagram of training a load identification and calibration model in an embodiment of the present invention;

[0065] Figure 3 It is a logical schematic diagram of load identification for real-time power consumption data in an embodiment of the present invention;

[0066] Figure 4 It is a structural schematic diagram of a CNN load pre-classification model in an embodiment of the present invention;

[0067] Figure 5 It is a schematic diagram showing the influence of the value of threshold setting n on the accuracy of known and unknown classes in an embodiment of the present invention;

[0068] Figure 6 It is a structural principle block diagram of a non-intrusive load identification system based on threshold adjustment in an embodiment of the present invention. Detailed implementation manners

[0069] To make the objectives, technical solutions, and advantages of the present invention clearer, the technical solutions of the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present invention. The embodiments described in the present invention are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the spirit of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the protection scope of the present invention.

[0070] Embodiment 1:

[0071] Refer to Figure 1 , an embodiment of the present invention provides a non-intrusive load identification method based on threshold adjustment, which specifically includes the following steps:

[0072] Step 1: Collect the total power consumption data of the user's residence for training, and perform preprocessing of time series imaging to obtain a single-appliance power image sample set containing various appliances;

[0073] As an embodiment of the present invention, the steps of performing preprocessing of time series imaging in Step 1 include:

[0074] S1.1: Label the switch states of various appliances corresponding to the collected total power consumption data to obtain an original data set;

[0075] S1.2: Clean the abnormal data in the original data set, and select the stable data immediately following the abnormal data for replacement to obtain a complete and usable initial sample set;

[0076] S1.3: Calculate the effective power of each cycle based on the voltage and current information of the total power consumption data in the initial sample set to obtain the one-dimensional time series of the total power corresponding to each sample;

[0077] S1.4: Refer to the marked switch states of the electrical appliances, intercept the jump windows of the one-dimensional time series of the total power respectively, and perform normalization processing;

[0078] S1.5: Use the Gramian Angular Summation Field (GASF) algorithm to convert the sequence data of each normalized jump window to the polar coordinate system respectively to obtain a single-electrical-appliance power image sample set containing various types of electrical appliances.

[0079] Step 2: Divide the single-electrical-appliance power image sample set to construct a first training sample set containing only known-class electrical appliance samples and a second training sample set that retains both known-class and unknown-class electrical appliance samples;

[0080] Among them, known-class electrical appliances are those that are frequently used, cover the basic living needs of most people, have irreplaceable functions, and can be used or contacted in daily life, such as washing machines, televisions, refrigerators, computers, hair dryers, printers, air conditioners, rice cookers, electric fans, floor sweepers, water heaters, humidifiers, microwave ovens, wall breakers, air purifiers, etc. Unknown-class electrical appliances are those with low popularity and only meet the needs of a few families or specific groups, such as electric garage doors, air compressors, DVR hard disk recorders, etc.

[0081] Specifically, in this embodiment, during the experiment, the Blued dataset released by Cornell University in 2012 was used. It is a high-quality dataset widely used in the field of non-intrusive load identification, containing typical household appliances, including lamps, small appliances, and heating equipment, etc. The sampling frequency is 12 kHz, and the voltage and current waveforms of the total household load are recorded, which is suitable for studying the switching transient signals and feature extraction of electrical appliances. In this embodiment, high-frequency current and voltage sampling information of multiple devices from a single residence over 8 days was obtained, and at the same time, an electrical appliance event list was recorded. From this, 406 pieces of total power consumption data containing the switching events of 12 electrical appliances in Table 2 were selected as the total sample set; each sample in the total sample set of step 1 was preprocessed. First, the effective power of each cycle was calculated based on the high-frequency voltage and current information to obtain a one-dimensional time series. Referring to the event list, the switching window was intercepted, and the Gramian angular field algorithm (GASF) was used to convert the normalized sequence data to the polar coordinate system, realizing the conversion of the total power consumption data sample into a single electrical appliance image sample, while labeling the corresponding labels, and expanding the sample quantity through time translation, so as to output a new single electrical appliance power image sample set; ten commonly used electrical appliances were selected as known classes, and two non-commonly used electrical appliances were selected as unknown classes, and the first training sample set and the second training sample set were divided according to the ratio of 8:2; for the divided training set samples and validation set samples, only the known class electrical appliance samples in Table 1 were retained in the first training sample set, a total of 1659 pieces; the second training sample set retained both known class and unknown class electrical appliance samples, a total of 251 pieces.

[0082] Table 1

[0083]

[0084]

[0085] Step 3: Train a CNN load pre-classification model based on a convolutional neural network through the first training sample set;

[0086] As an embodiment of the present invention, the structure of the CNN load pre-classification model includes a first convolutional layer group, a first pooling layer, a second convolutional layer group, a second pooling layer, a third convolutional layer group, a third pooling layer, a dropout layer, a flattening layer, a fully connected layer, and an output layer connected in sequence; after the input image sample is extracted with primary features by the first convolutional layer group, it is subjected to primary feature dimensionality reduction processing by the first pooling layer, and then after the intermediate features are extracted by the second convolutional layer group, it is subjected to intermediate feature dimensionality reduction processing by the second pooling layer, and then after the high-level features are extracted by the third convolutional layer group, it is subjected to high-level feature dimensionality reduction processing by the third pooling layer. Subsequently, the high-level features after dimensionality reduction are sequentially processed by the dropout layer and the flattening layer, and then input to the fully connected layer to learn global features to obtain an activation vector. Finally, after the activation vector is processed by the output layer with Softmax activation, the pre-classification category of the input image sample is finally obtained.

[0087] Furthermore, it should be noted that the load pre-classification model structure used in the present invention is CNN, and the overall structure is as Figure 4 shown; CNN learns high-dimensional feature representations by continuously applying operations of multiple convolutional, pooling, and fully connected layers to the input data; the convolutional layer is the core of CNN, which performs a sliding window operation on the local area of the image to extract a feature map with high-dimensional features. The pooling layer then reduces the dimensionality of the output of the convolutional layer, and at the same time can also reduce the computational complexity of the network and prevent overfitting; through multiple convolutional and pooling operations, CNN can extract high-level features of the image, and then classify them through the fully connected layer.

[0088] In this embodiment, the CNN load pre-classification model is designed to process grayscale images of 32×32 pixels. The image first goes through two convolutional layers, each using 16 3×3 filters and the ReLU activation function to extract primary features; subsequently, the max pooling layer reduces the spatial dimension of the feature map through a 2×2 window, thereby reducing the computational amount and extracting key features. Next, the model includes two convolutional layers with 32 3×3 filters and the ReLU activation function to extract more complex features, and then goes through the max pooling layer again to reduce the size of the feature map. After that, the model applies a convolutional layer with 64 3×3 filters to extract more high-level features, and then goes through the max pooling layer to further reduce the size. To prevent overfitting, the model uses a dropout layer to randomly discard 20% of the neurons; then, the flattening layer flattens the three-dimensional feature map into one dimension for input to the fully connected layer; the fully connected layer contains 50 neurons and uses the ReLU activation function for feature combination, and at the same time provides an activation vector of the correctly classified sample for the post-processing algorithm; the final output layer contains 10 neurons to provide the pre-classification result.

[0089] In a preferred but non-limiting embodiment, the objective function adopted by the CNN load pre-classification model during training is: the cross-entropy loss function.

[0090] Step 4: Based on the second training sample set, train a load identification calibration model by combining a load identification open-set recognition algorithm with threshold adjustment;

[0091] Refer to Figure 2 , the steps of training a load identification calibration model by combining a load identification open-set recognition algorithm with threshold adjustment (abbreviated as the OpenAppliance algorithm) include:

[0092] S4.1: Calculate the mean activation vector of each category according to the activation vectors output by the fully connected layer in the CNN load pre-classification model for the correctly classified samples, and use the Weibull distribution fitting to obtain the cumulative distribution function of each category; the specific process is as follows:

[0093] S411: Input the first training sample set into the trained CNN load pre-classification model for pre-classification, and compare the pre-classification result with its corresponding true value to screen out the correctly classified sample sets of various electrical appliances;

[0094] S412: Calculate the mean activation vector of each category according to the activation vectors output by the fully connected layer in the CNN load pre-classification model for the correctly classified sample sets of various electrical appliances;

[0095] In the experiment of this embodiment, 1659 known-class first training sample sets are put into the CNN load pre-classification model to screen out 1592 correctly classified samples; calculate the activation vector V output by the fully connected layer for each correctly classified sample x as follows:

[0096] V x =(v x1 ,v x2 ,v x3 ,...,v xc )

[0097] In the formula, V x is the activation vector of the correctly classified sample x, c represents the total number of categories, and v xi represents the average activation value of the sample x in the i-th category.

[0098] According to the activation vectors V of all correctly classified samples in each category i (i = 1, 2... c, c is the total number of categories of known-class electrical appliances) x calculate the mean activation vector MV of each category i i :

[0099]

[0100] In the formula, MV i is the mean activation vector of category i, Ni represents the number of correctly classified samples in class i, V x is the activation vector output by the fully connected layer for the x-th correctly classified sample.

[0101] S413: Calculate the Euclidean distance between each sample in the set of correctly classified samples of each class and the mean activation vector of its corresponding class; the expression is as follows:

[0102]

[0103] In the formula, d xi represents the Euclidean distance between the activation vector V x of the correctly classified sample x and the mean activation vector MV i of its class i, V x is the activation vector output by the fully connected layer for the x-th correctly classified sample, and MV i is the mean activation vector of class i; finally, the set of Euclidean distances for class i is obtained

[0104] S414: Perform Weibull distribution fitting on the Euclidean distances of each sample in the set of correctly classified samples of each class to obtain the cumulative distribution function of each class.

[0105] Specifically, perform Weibull distribution fitting on all the Euclidean distances in D i to obtain the probability density function f(d xi ; λ; k) of class i, which represents the probability density that a random sample in D i takes a certain specific Euclidean distance value.

[0106]

[0107] In the formula, k is the shape parameter of the Weibull distribution, which determines the shape characteristics of the distribution; λ is the scale parameter of the Weibull distribution, which controls the range of expansion of the distribution.

[0108] By integrating the probability density function of the Weibull distribution, the expression of the cumulative distribution function F(d xi ; k; λ) of each class is as follows:

[0109]

[0110] In the formula, k is the shape parameter of the Weibull distribution, and λ is the scale parameter of the Weibull distribution.

[0111] S4.2: Calculate the sum of the elements of the corrected probability vectors of the samples in each category in the second training set respectively based on the mean activation vectors and cumulative distribution functions of each category, and sort them in descending order to obtain the sequence of the sum of elements of each category. The specific process is as follows:

[0112] S421: For any second training sample y, input it into the trained CNN load pre-classification model to obtain its corresponding activation vector V y and the pre-classified category i;

[0113] S422: Calculate the Euclidean distances between the activation vector V y and the mean activation vectors of each category respectively to obtain the Euclidean distance vector ε y of the sample y; where, ε y =(d y1 , d y2 , d y3 ,..., d yc ),

[0114] S423: Substitute the Euclidean distance vector ε y into the cumulative distribution functions of each category respectively to calculate the corrected probability vector of the sample y, so as to obtain the sum S y of the elements of the corrected probability vector of the sample y.

[0115] Furthermore, substitute each d y in ε y1 =(d y2 , d y3 , d yc ) into the cumulative distribution functions of each category to obtain the corrected probability cv yi of the sample belonging to each category, that is yi , and finally obtain the corrected probability vector CV of the sample x y =(cv y1 , cv y2 , cv y3 ,..., cv yc ), so as to calculate the sum y of the elements of CV

[0116] S4.3: Based on the current value of n, respectively take the sum S * y corresponding to the first n% demarcation in the sequence of the sum of elements of each category as the threshold α i of its category (that is: for the i-th type of electrical appliance, the sum of the elements of the corrected probability vectors of n% of the samples is greater than α i ), and obtain the threshold vector β corresponding to the value of n;

[0117] S4.4: For any second training sample y, compare the sum of the elements of the corrected probability vector of sample y with S y the threshold α corresponding to its pre-classified category i i If S y is greater than α i , then classify sample y into category i, otherwise classify it into the unknown category, so as to obtain the corrected classification result of the second training sample set;

[0118] S4.5: Based on the corrected classification result, calculate the classification accuracy rates of the known and unknown samples in the second training sample set;

[0119] S4.6: Repeat steps S4.3 to S4.5 to calculate the average values of the classification accuracy rates of the known and unknown categories for different values of n, and use the threshold vector corresponding to the value of n with the highest average accuracy rate as the finally determined threshold vector β * =(α 1 , α 2 , α 3 ,..., α c ), so as to obtain the finally trained load identification calibration model.

[0120] Furthermore, the value-taking method of n in this embodiment is: first set a minimum value of 80% for the accuracy rates of the known and unknown categories, then compare the average accuracy rates of the known and unknown categories of the overall verification samples for different values of n when this condition is met, and the threshold vector corresponding to the value of n with the highest average accuracy rate is the finally determined threshold vector. In this experiment, five groups of threshold vectors are obtained by taking n as 94, 95, 96, 97, and 98 respectively, and the accuracy rate changes are as Figure 5 shown. The accuracy rate of the unknown category shows an increasing relationship with the value of n; while the accuracy rate of the known category shows a decreasing relationship with the value of n. When n is greater than 96, the accuracy rate of the known category drops significantly; the value of n is determined by calculating the average values of the accuracy rates of the known and unknown categories for different values of n, and finally the value of n is obtained as 96.

[0121] Step 5: After preprocessing the real-time total power consumption data collected in the target user's residence to obtain single-appliance power image data, and passing it through the CNN load pre-classification model and the load identification calibration model for classification and recognition in sequence, obtain the identification result of the appliance types of the currently connected loads in the target user's residence.

[0122] As Figure 3As shown, when performing load identification on real-time total power consumption data in step 5, first, calculate its one-dimensional time series of active power, intercept the power jump window, and use the Gramian angular field algorithm to convert the normalized sequence data to the polar coordinate system to obtain the corresponding single-appliance power image data z; then, input the obtained single-appliance power image data z into the trained CNN load pre-classification model to obtain its activation vector V z and the pre-classified category i; then, input the activation vector V z and the pre-classified category i into the trained load identification calibration model. The load identification calibration model calculates the correction probability vector element sum S z of the single-appliance power image data z and compares it with the threshold α i corresponding to its pre-classified category i to obtain the final classification result, and its expression is as follows:

[0123]

[0124] In the formula, Class(z) is the final classification category of sample z; i is the category to which the maximum probability score of sample z belongs; Unkown is the unknown category; S z is the correction probability element sum of sample z in category i; α i is the threshold of category i.

[0125] Furthermore, traditional load identification methods usually rely on the closed-set assumption, that is, it is considered that the training data covers all possible electrical appliance types and operating modes. However, in actual applications, new electrical appliances are constantly emerging, and users' electricity consumption behaviors are diverse, resulting in the load identification system often encountering unknown load types that have not been seen before; traditional SoftMax classifiers perform poorly in this case because they tend to assign all inputs to known categories; the present invention trains a load identification calibration model through the OpenAppliance algorithm to recalculate the output probability of the neural network, enabling it to better distinguish between known and unknown categories.

[0126] In addition, existing open-set recognition methods reclassify by calculating the difference between the probability output by the original SoftMax classifier of known classes and the correction probability and assigning it to the unknown category. However, due to the highly complex load data and the easy overlap of features, existing open-set recognition algorithms perform poorly when directly applied to the field of load identification, and existing open-set recognition algorithms based on image classification assign to the unknown category by reducing the probability of known classes and cannot be directly applied to load identification, so improvements are needed. The present invention combines the load identification open-set recognition algorithm (OpenAppliance) with threshold adjustment to statistically calculate the threshold α i, that is, for the i-th type of electrical appliance, n% of the sample values are greater than α i , a threshold vector β corresponding to the value of n is obtained i =(α 1 ,α 2 ,α 3 ,...,α c ). By comparing the average accuracy rates of the known and unknown classes of the entire second training sample set under different values of n, the value of n in the case of the highest average accuracy rate is taken, and the corresponding threshold vector is the finally determined threshold vector. By introducing a threshold adjustment strategy and integrating deep learning feature extraction and probability model technology, unknown electrical appliances are detected and classified on the premise of effectively identifying known electrical appliances, thus overcoming the limitations of traditional load identification algorithms in dealing with unknown electrical appliances.

[0127] In the comparative experiment of this embodiment, the load identification method based only on the traditional SoftMax classifier and the experiment based on the traditional open-set OpenMax recognition method were respectively carried out and compared with the OpenAppliance algorithm for load identification with threshold adjustment of the present invention for calibration recognition; the experimental results are shown in Table 2.

[0128] Table 2

[0129]

[0130] The analysis of the experimental results is as follows:

[0131] I. Comparison with the SoftMax classifier

[0132] Softmax is the initial classifier of the CNN model in this paper, used to calculate the prediction probability of each class. When using the SoftMax classification method, the accuracy rate of the known class is 93.75%, but it completely loses the recognition ability for the unknown class, and the accuracy rate for the unknown class is only 0.2%. While through the OpenAppliance algorithm, although the recognition accuracy rate of the known class decreases slightly, the recognition accuracy rate of the unknown class can reach more than 80%, far exceeding the commonly used SoftMax classification method in the CNN model.

[0133] II. Comparison with the OpenMax algorithm

[0134] The OpenMax algorithm is an open-set recognition algorithm mainly used to handle the "unknown class" problem that traditional pattern recognition methods cannot effectively cope with. For unknown classes, OpenMax can judge whether an input sample belongs to a known class based on the output distribution of the classification model. By calculating the difference between the probability output by the original softmax classifier of the known class and the corrected probability, it assigns it to the unknown class for reclassification. When using the OpenMax algorithm in the load identification scenario, the recognition accuracy of the unknown class can only reach 71.1% at most, but the accuracy of the known class drops to 81.2%. After using the OpenAppliance algorithm, not only is the recognition accuracy of the unknown class improved, but also the impact on the recognition accuracy of the known class is reduced. The accuracy of the known class is 6.8% higher than that of the OpenMax algorithm, and the accuracy of the unknown class is 18% higher than that of the OpenMax algorithm.

[0135] Summary of Comparative Experiments

[0136] Through comparative experiments, it is found that the OpenAppliance, an open-set recognition algorithm for load identification based on threshold adjustment in the present invention, can enhance the recognition ability of the load identification algorithm for unknown electrical appliances, and at the same time has less impact on the performance degradation of the algorithm for identifying the load of known electrical appliances, thereby improving the accuracy, robustness and practicality of the load identification algorithm.

[0137] In summary, for the existing non-intrusive load identification algorithms provided in the embodiments of the present invention, although they can identify the samples in the training set, in real application scenarios, there will be new samples that have not appeared in the training set. The existence of these unknown samples will make it difficult for these algorithms to accurately identify, and the algorithm recognition rate will thus drop significantly, which is called the open-set recognition problem. To solve the impact of unknown electrical appliances on the load identification algorithm in practical applications, the present invention proposes an open-set recognition algorithm for load identification based on threshold adjustment, namely OpenAppliance. It converts the load data into an image form suitable for deep learning, uses a convolutional neural network (CNN) to extract high-dimensional features, introduces a threshold adjustment strategy, and by integrating deep learning feature extraction and probability model technology, it can detect and classify unknown electrical appliances on the premise of effectively identifying known electrical appliances, thereby overcoming the limitations of traditional load identification algorithms in dealing with unknown electrical appliances. Verified by the BLUED load data set, the OpenAppliance algorithm can not only maintain a relatively high recognition accuracy for known electrical appliances, but also increase the recognition accuracy of unknown electrical appliances from 71.1% to 89.1%, successfully balancing the recognition performance of known and unknown classes, and effectively improving the accuracy and robustness of the load identification system.

[0138] Embodiment 2:

[0139] As Figure 6As shown in the figure, the present invention provides a non-intrusive load identification system based on threshold adjustment. The system is used to implement the steps of the method in the first embodiment above. Specifically, the system includes:

[0140] A data acquisition module, which is used to acquire the total power consumption data of the user's residence for training and perform preprocessing of time series imaging to obtain a single-appliance power image sample set containing various appliances;

[0141] A sample division module, which is used to divide the single-appliance power image sample set, construct a first training sample set containing only known-class appliance samples, and a second training sample set that retains both known-class and unknown-class appliance samples;

[0142] A first training module, which is used to train a CNN load pre-classification model based on a convolutional neural network through the first training sample set;

[0143] A second training module, which is used to calibrate the load identification model based on the second training sample set and the load identification open-set recognition algorithm with threshold adjustment;

[0144] A load identification module, which is used to perform classification and recognition on the single-appliance power image data obtained after preprocessing the real-time total power consumption data collected in the target user's residence through the CNN load pre-classification model and the load identification calibration model in sequence, so as to obtain the identification result of the appliance types of the current connected load in the target user's residence.

[0145] The non-intrusive load identification system based on threshold adjustment provided by the embodiment of the present invention and the non-intrusive load identification method based on threshold adjustment provided by the first embodiment are based on the same technical concept, and can produce the beneficial effects described in the first embodiment. For the content not described in detail in this embodiment, reference can be made to the first embodiment.

[0146] Embodiment Three:

[0147] An end device provided by an embodiment of the present invention includes a processor and a storage medium;

[0148] The storage medium is used to store instructions;

[0149] The processor is used to operate according to the instructions to execute the steps of the method according to any one of the first embodiments.

[0150] Embodiment Four:

[0151] A computer-readable storage medium provided by an embodiment of the present invention stores a computer program, and when the program is executed by a processor, it implements the steps of the method according to any one of the first embodiments.

[0152] The present invention may be a system, a method, and / or a computer program product. The computer program product may include a computer-readable storage medium having computer-readable program instructions thereon for causing a processor to implement aspects of the present invention.

[0153] A computer-readable storage medium may be a tangible device that can retain and store instructions for use by an instruction execution device. A computer-readable storage medium may be, for example—but not limited to—an electrical storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the foregoing. More specific examples (a non-exhaustive list) of the computer-readable storage medium include: a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), a static random access memory (SRAM), a portable compact disc read-only memory (CD-ROM), a digital versatile disc (DVD), a memory stick, a floppy disk, a mechanically encoded device such as a punch card or raised structures in grooves having instructions stored thereon, and any suitable combination of the foregoing. The computer-readable storage medium as used herein is not construed to be a transitory signal per se, such as a radio wave or other freely propagating electromagnetic wave, an electromagnetic wave propagating through a waveguide or other transmission medium (e.g., an optical pulse through an optical fiber cable), or an electrical signal transmitted through a wire.

[0154] The computer-readable program instructions described herein may be downloaded to various computing / processing devices from a computer-readable storage medium or may be downloaded to an external computer or external storage device through a network, such as the Internet, a local area network, a wide area network, and / or a wireless network. The network may include copper transmission cables, optical fiber transmission, wireless transmission, routers, firewalls, switches, gateway computers, and / or edge servers. A network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards the computer-readable program instructions for storage in a computer-readable storage medium in each computing / processing device.

[0155] The computer program instructions for carrying out the operations of the present invention may be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-related instructions, microcode, firmware instructions, state-setting data, or source code or object code written in any combination of one or more programming languages, including object-oriented programming languages such as Smalltalk, C++, etc., and conventional procedural programming languages such as the "C" language or similar programming languages. The computer-readable program instructions may be executed entirely on the user's computer, partially on the user's computer, executed as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on the remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., through the Internet using an Internet service provider). In some embodiments, by using the state information of the computer-readable program instructions to customize an electronic circuit, such as a programmable logic circuit, a field-programmable gate array (FPGA), or a programmable logic array (PLA), the electronic circuit can execute the computer-readable program instructions to implement various aspects of the present invention.

[0156] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to the above embodiments, those of ordinary skill in the art should understand that: modifications or equivalent substitutions can still be made to the specific embodiments of the present invention, and any modifications or equivalent substitutions that do not depart from the spirit and scope of the present invention should be covered by the protection scope of the claims of the present invention.

Claims

1. A non-intrusive load identification method based on threshold adjustment, characterized in that: Methods include: Step 1: Collect the total electricity consumption data of the user's residence for training, and perform preprocessing of time series imaging to obtain a single appliance power image sample set containing various types of appliances; Step 2: Divide the single electrical appliance power image sample set to construct a first training sample set containing only known class electrical appliance samples and a second training sample set retaining both known class and unknown class electrical appliance samples; Step 3: obtaining a CNN load pre-classification model based on a convolutional neural network through training the first training sample set; Step 4: Based on the second training sample set, a load identification open set recognition algorithm combined with threshold adjustment is trained to obtain a load identification calibration model; Step 5: The single appliance power image data obtained after preprocessing the real-time total power consumption data collected in the target user's residence is classified and identified by the CNN load pre-classification model and the load identification calibration model in turn to obtain the identification result of the appliance type currently connected to the load in the target user's residence.

2. The non-intrusive load identification method based on threshold adjustment according to claim 1, characterized in that: In step 1, the step of preprocessing the time series image includes: S1.1: Label the switch status of various electrical appliances corresponding to the collected total electricity consumption data to obtain the original data set; S1.2: cleaning the original data set of abnormal data, and replacing the data with stable data immediately following the abnormal data, so as to obtain a complete and usable initial sample set; S1.3: Calculate the effective power of each cycle according to the voltage and current information of each total power consumption data in the initial sample set to obtain a one-dimensional time series of the total power corresponding to each sample; S1.4: referring to the switch states marked by the electrical appliances, respectively intercepting the transition windows of the one-dimensional time series of the total power, and performing normalization processing; S1.5: The normalized sequence data of each transition window is converted into a polar coordinate system using the Grami Angular Field (GASF) algorithm to obtain a single electrical appliance power image sample set containing various electrical appliances.

3. The non-intrusive load identification method based on threshold adjustment according to claim 1, characterized in that: The known electrical appliances include washing machines, televisions, refrigerators, computers, hair dryers, air conditioners, rice cookers, printers, electric fans, sweepers, water heaters, humidifiers, microwave ovens, wall breakers and / or air purifiers; The unknown electrical appliances include electric garage doors, air compressors and / or DVRs.

4. The non-intrusive load identification method based on threshold adjustment according to claim 1, characterized in that: The structure of the CNN load pre-classification model includes a first convolutional layer group, a first pooling layer, a second convolutional layer group, a second pooling layer, a third convolutional layer group, a third pooling layer, a random inactivation layer, a flattening layer, a fully connected layer and an output layer connected in sequence; After the input image sample has primary features extracted by the first convolution layer group, the primary features are reduced in dimension by the first pooling layer, and then the intermediate features are extracted by the second convolution layer group. The intermediate features are reduced in dimension by the second pooling layer, and then the high-level features are extracted by the third convolution layer group. The high-level features are reduced in dimension by the third pooling layer. Subsequently, the high-level features after dimensionality reduction are processed by the random inactivation layer and the flattening layer in turn, and then input into the fully connected layer to learn the global features to obtain the activation vector. Finally, the activation vector is processed by the output layer with Softmax activation to finally obtain the pre-classification category of the input image sample.

5. The non-intrusive load identification method based on threshold adjustment according to claim 4, characterized in that: The objective function used by the CNN load pre-classification model during training is: cross entropy loss function.

6. The non-intrusive load identification method based on threshold adjustment according to claim 1, characterized in that: In step 4, the step of training the load identification open set recognition algorithm combined with threshold adjustment to obtain the load identification calibration model includes: S4.1: Calculate the mean activation vector of each category according to the activation vector output by the fully connected layer of the CNN load pre-classification model of the correctly classified samples, and use Weibull distribution fitting to obtain the cumulative distribution function of each category; S4.2: Based on the mean activation vector and cumulative distribution function of each category, respectively calculate the sum of the corrected probability vector elements of each category of samples in the second training sample set, and sort them in descending order to obtain the sum sequence of elements of each category; S4.3: Based on the current value of n, separate the elements of each category and the elements corresponding to the first n% of the sequence and S * y As the threshold α of its category i , get the threshold vector β corresponding to the value of n; S4.4: For any second training sample y, add the corrected probability vector element of sample y and S y The threshold α corresponding to its pre-classification category i i Compare, if S y Greater than α i , then classify sample y into category i, otherwise it is classified into an unknown class, thereby obtaining a corrected classification result of the second training sample set; S4.5: Based on the corrected classification result, calculate the classification accuracy of the known class and the unknown class samples in the second training sample set; S4.6: Repeat steps S4..3 to S4.5 to calculate the average classification accuracy of known classes and unknown classes under different values ​​of n, and use the threshold vector corresponding to the value of n with the highest average accuracy as the final threshold vector to obtain the final trained load identification calibration model.

7. The non-intrusive load identification method based on threshold adjustment according to claim 6, characterized in that: The specific steps of S4.1 include: 411: inputting the first training sample set into the trained CNN load pre-classification model for pre-classification, and comparing the pre-classification result with its corresponding true value to screen out the correctly classified sample sets of various electrical appliances; S412: Calculate the mean activation vector of each category according to the activation vectors output by the fully connected layer in the CNN load pre-classification model according to the correctly classified sample sets of each category of electrical appliances; S413: Calculate the Euclidean distance between each sample in each correct sample set and the mean activation vector of its corresponding category respectively; S414: Perform Weibull distribution fitting on the Euclidean distance of each sample in each category of correct sample set to obtain the cumulative distribution function of each category.

8. The non-intrusive load identification method based on threshold adjustment according to claim 6 or 7, characterized in that: The calculation formula for the mean activation vector of each category is as follows: In the formula, MV i is the mean activation vector of category i, N i represents the number of correctly classified samples in category i, V x is the activation vector output by the fully connected layer of the xth correctly classified sample, i = 1, 2…c, c is the total number of known electrical appliance categories.

9. The non-intrusive load identification method based on threshold adjustment according to claim 6 or 7, characterized in that: The expression of the cumulative distribution function for each category is as follows In the formula, F(d xi ; k; λ) represents the cumulative distribution function of category i, d xi The activation vector V represents the correctly classified sample x x The mean activation vector MV of its category i i The Euclidean distance of , k is the shape parameter of the Weibull distribution, and λ is the scale parameter of the Weibull distribution.

10. The non-intrusive load identification method based on threshold adjustment according to claim 9, characterized in that: The Euclidean distance d xi The calculation formula is as follows: Where V x is the activation vector output by the fully connected layer for the xth correctly classified sample, MV i is the mean activation vector for class i.

11. The non-intrusive load identification method based on threshold adjustment according to claim 6, characterized in that: In S4.2, the step of respectively calculating the sum of the elements of the correction probability vector of each category of samples in the second training sample set includes: S421: For any second training sample y, input it into the trained CNN load pre-classification model to obtain its corresponding activation vector V y and pre-classified category i; S422: Calculate activation vectors V respectively y The Euclidean distance from the mean activation vector of each category is obtained to obtain the Euclidean distance vector ε of sample y y ; S423: The Euclidean distance vector ε y Substitute them into the cumulative distribution function of each category to calculate the corrected probability vector of sample y, thereby obtaining the corrected probability vector elements and S of sample y y .

12. A non-intrusive load identification system based on threshold adjustment, running the non-intrusive load identification method based on threshold adjustment as claimed in any one of claims 1 to 11, characterized in that: The system includes: A data collection module is used to collect the total power consumption data of user residences for training, and perform preprocessing of time series imaging to obtain a single appliance power image sample set containing various types of appliances; A sample division module is used to divide the single electrical appliance power image sample set to construct a first training sample set containing only known class electrical appliance samples and a second training sample set retaining both known class and unknown class electrical appliance samples; A first training module, used for training a CNN load pre-classification model based on a convolutional neural network through the first training sample set; A second training module is used for calibrating a load identification model based on the second training sample set and a load identification open set recognition algorithm combined with a threshold adjustment; The load identification module is used to obtain the single appliance power image data after preprocessing the real-time total power consumption data collected in the target user's residence, and then classify and identify it through the CNN load pre-classification model and the load identification calibration model in turn, so as to obtain the identification result of the type of appliance currently connected to the load in the target user's residence.

13. A terminal comprising a processor and a storage medium; characterized in that: The storage medium is used to store instructions; The processor is configured to operate according to the instructions to execute the steps of the method according to any one of claims 1-11.

14. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the steps of the method according to any one of claims 1 to 11 are implemented.

Citation Information

Cited By

  • Unknown load identification and incremental learning method based on feature space multi-objective optimization

    CN122346734A

  • Unknown load identification and incremental learning method based on feature space multi-objective optimization

    CN122346734B