A High-Precision Low-Frequency NILM Neural Network Method Based on a Synthetic Training Set
Through the high-precision low-frequency NILM neural network method based on the synthetic training set, the overfitting and misjudgment problems of small samples and low-frequency data sets are solved, and the training results are optimized by feature extraction and error analysis, which improves the recognition accuracy and anti-interference of the NILM neural network.
Patent Information
- Application Number
- CN202310383900.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-04-12
- Publication Date
- 2025-07-01
- Estimated Expiration
- 2043-04-12
AI Technical Summary
The existing NILM neural network schemes have problems such as overfitting, frequent noise, and frequent misjudgment when processing domestic small samples and low-frequency data sets, and the evaluation standards are different, resulting in poor recognition results.
The high-precision low-frequency NILM neural network method based on the synthetic training set is adopted to reconstruct the data set through feature extraction and error analysis, and train with lightweight CNN-Seq2point neural network, and optimize the training results through error analysis to improve device density and prediction accuracy.
The feature density of small samples and low-density devices is improved, overfitting is reduced, and the accuracy and anti-interference of prediction results are enhanced, thus achieving higher recognition accuracy.
Smart Images

Figure CN116451744B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a high-precision low-frequency NILM neural network method based on a synthetic training set, belonging to the technical field of big data analysis of power systems. Background Art
[0002] The operation of the power system load usually includes two monitoring methods, intrusive load monitoring (ILM) and non-intrusive load monitoring (NILM). Traditional intrusive load monitoring (ILM) requires installing sensors at each power load, while non-intrusive load monitoring (NILM) only needs to perform real-time measurement at the entrance of the power load, analyze the collected information such as voltage, current, and power, and then the total load, the categories of each decomposed electrical load, and the real-time energy consumption of different electrical loads can be obtained. It has low maintenance costs and is convenient to implement, and has great application prospects in aspects such as power grid demand-side management, energy conservation and emission reduction, and building intelligent power grids.
[0003] Existing NILM implementation schemes are mainly divided into two categories: optimization schemes and pattern recognition schemes. The former does not have a training process. Only a set of evaluation methods for quantifying the matching degree of solutions needs to be designed, and then the optimal solution is found through repeated iteration and verification. However, this method has a large amount of computation, low efficiency, and poor flexibility, and cannot process a large amount of user data; while the pattern recognition scheme mainly based on neural networks can effectively solve this bottleneck problem and is the current mainstream NILM scheme.
[0004] The pattern recognition scheme of the NILM neural network mainly includes steps such as the selection and processing of the data set, the selection and design of the neural network, and evaluation and optimization. Currently, almost all the data sets used in the research in this field are classic foreign data sets. Firstly, the acquisition time is relatively early. Secondly, there are significant differences in load models between domestic and foreign. Thirdly, there are also significant differences in power grid power supply forms between domestic and foreign. In short, the training results are not good for identifying domestic loads. The data sets self-collected by domestic small and medium-sized companies generally have characteristics such as low frequency, data loss, low device working density, and non-obvious and non-concentrated load characteristics due to the different precisions of the acquisition devices, which bring certain difficulties to the recognition work of the neural network.
[0005] Currently, there is an urgent need to propose corresponding solutions for the processing of the above-mentioned small-sample low-frequency data sets and low-density device data. Summary of the Invention
[0006] In view of the problems that small-sample data sets are prone to overfitting, the fitting curves of low-density devices have a lot of spike noises and frequent misjudgments, etc., the present invention provides a high-precision low-frequency NILM neural network method based on a synthetic training set, and designs feature extraction schemes including "pulse-interval" power sequences, device dependency distribution matrices, and device dependency distribution probabilities, etc., and uses these pulses to regenerate the training set, validation set, and test set, improving the device density based on the original device dependencies and retaining more device features. The test results prove that the data processed by this scheme can help the S2P neural network to more accurately (with high precision) fit the power curves of small-sample and low-density devices.
[0007] In view of the problems that the NILM evaluation criteria for neural network-containing are inconsistent and it is difficult to improve the network, the present invention provides a general error analysis method, which can be used to judge the network quality (initial training results), and accordingly adjust the feature ratio and composition of the devices in the input data set to obtain a more satisfactory final training result.
[0008] The present invention adopts the following technical solutions:
[0009] A high-precision low-frequency NILM neural network method based on a synthetic training set, comprising the following steps:
[0010] Step 1, data preprocessing: collect the bus power and each electrical appliance power data, and perform data filling and label making;
[0011] Step 2, feature data extraction: apply the threshold determination method and statistical calculation formulas to the data obtained after the preprocessing in Step 1 to extract and calculate the corresponding feature data;
[0012] Step 3, synthetic data set: reconstruct the data set according to the statistical method for the feature data obtained in Step 2 to prevent overfitting;
[0013] Step 4, lightweight CNN-Seq2point neural network training: use the data set synthesized in Step 3, with the bus power as the input and the power of a certain electrical appliance as the true value to calculate the loss function, and use the lightweight CNN-Seq2point model structure to train to obtain the network model parameters model of this electrical appliance;
[0014] Step 5, decomposition output: load the model obtained in Step 4 into the lightweight CNN-Seq2point network, send the bus power data in the test set into the lightweight CNN-Seq2point network, and use the data of the electrical appliance to which the network belongs as the true value, then output the power prediction data and prediction accuracy of the electrical appliance to which the network model belongs, which is the preliminary prediction result;
[0015] Step 6, Error Analysis: Calculate the error between the prediction result obtained in Step 5 and the true value data of the same electrical appliance at the same timestamp. At the same time, perform an overlap analysis on this electrical appliance and electrical appliances with similar working states. According to the error calculation results and overlap analysis results, count several four-fold tables, and judge the influence degree and influence position of similar electrical appliances on this electrical appliance based on the four-fold tables;
[0016] Step 7, Optimize the Preliminary Prediction Result: According to the judgment of the error analysis in Step 6, adjust the composition of the synthetic dataset in Step 3, and then repeat Steps 4 and 5 until satisfactory training and prediction results (lower loss value or higher accuracy) are obtained.
[0017] Preferably, in Step 1, the bus power and the power data of each electrical appliance (i.e., the power sequence), mainly referring to small sample data, are obtained by collecting the dataset domestic, the public dataset refit, the public dataset REDD (low frequency), etc. For problems such as sampling breakpoints in some devices in the sampled power sequence and inconsistent sampling intervals for different electrical appliances, the following measures are taken:
[0018] First, cut and splice at the breakpoint to obtain a set of power sequences without breakpoints; then, specify a unified sampling interval T. For power sequences with an original sampling interval greater than T, interpolation processing is used; for power sequences with an original sampling interval less than T, decimation processing is used; a series of null values NULL will be generated during this process. Finally, the "mean compensation method" is used to fill in the missing values; where the sampling breakpoint refers to a situation where the sampling interval of a device at a certain moment or within a certain small time period is much larger than the normal sampling interval of the device; the mean compensation method refers to calculating the average value of n points near the null value NULL and setting the value at NULL to
[0019] Label making refers to establishing a correspondence between the device name and the device labels involved in the training process to facilitate the subsequent call of the neural network. The device labels include but are not limited to the mean, standard deviation, the house where the device is located, and the channel where the device is located; among them, the mean and standard deviation are set using the trial-and-error method, and the trial-and-error values are given based on the device power curve, device switch state curve, and device power distribution curve, etc.
[0020] Preferably, in Step 2, the feature data includes the pulse eigen-distribution, the "pulse-interval" sequence, the device dependency distribution matrix, and the device dependency distribution probability;
[0021] Among them, the "pulse" and "interval" refer to scanning the specified power sequence under a specified threshold: the power segments higher than the threshold are recorded as the pulse sequence {Pul n} = [Pul1, Pul2,..., Pul n, where n represents the total number of extracted pulse or interval segments; the power segments below the threshold are recorded as the interval sequence {Int n} = [Int1, Int2,..., Int n , and the threshold is determined by trial and error;
[0022] The pulse eigen-distribution, i.e., the device density, characterizes the density at which the device itself operates, referring to the total pulse time T of the device pul occupying the total time T of its power sequence total ratio
[0023] The "pulse-interval" sequence includes the noise "pulse-interval" sequence, the device "pulse-interval" sequence, and the bus "pulse-interval" sequence. Among them,
[0024] a) The noise "pulse-interval" sequence refers to the "pulse-interval" sequence extracted from the difference sequence P t between the bus power sequence P a and the sum of the power sequences of the electrical appliances ∑P Δ = |P t - ∑P a |;
[0025] b) The device "pulse-interval" sequence refers to the "pulse-interval" sequence extracted from the power sequences {P ax} of each electrical appliance, where x refers to the xth electrical appliance; the following definitions are made for type A devices and type B devices:
[0026] Define α > 30% as high-density devices, i.e., type A devices, and α ≤ 30% as low-density devices, i.e., type B devices, where 30% is given according to the Monte Carlo algorithm;
[0027] For type A devices, the pulse sequence {Pul An} and the interval sequence {Int An} need to be extracted separately; for type B devices, only the pulse sequence {Pul Bn} needs to be extracted, and its interval sequence is regarded as a zero-value sequence;
[0028] c) The bus "pulse-interval" sequence refers to the sum of the noise "pulse-interval" sequence in a) and the device "pulse-interval" sequence in b) respectively to obtain the bus "pulse-interval" sequence;
[0029] The device dependence distribution matrix characterizes the dependence degree of the "on" states of two devices. Denote the total pulse time of type A devices as T pulA , and the number of times type B devices enter the pulse state as C on , then this dependence degree β can be calculated by the following formula: Let the total number of type A devices be N A , and the total number of type B devices be N B , then the matrix size is N A ×N B ;
[0030] The distribution probability γ of the device dependency represents the density exponent of the "on" state of the type B devices being controlled. Based on the device dependency distribution matrix, find the maximum dependency β of the type B devices max (Take the maximum value of β for each column of the matrix (i.e., each type of type B device), that is, the β values of different type B devices max are different), and at the same time calculate the dependency standard deviation β sd as the pessimistic error estimate value, then γ = β max ±β sd ;
[0031] Among them, p is the p-th type B device, and q is the total number of type A devices.
[0032] The extraction of the characteristic data in step 2 includes dividing the data into "pulses" and "intervals" through threshold determination, counting the proportion of the total time of the pulse segments of the devices in the total time of their power sequences as the device density, counting the proportion of the number of times the low-density devices enter the pulse state in the total pulse time of the high-density devices as the device dependency distribution matrix, and counting the pessimistic error estimate value of the matrix as the device dependency distribution probability.
[0033] Preferably, the process of synthesizing the data set in step 3 is as follows:
[0034] Determine the length L of the synthesized sequence according to the general length of the preprocessed power sequence, and generate a set of zero-value sequences Z with length L;
[0035] For type A devices, insert device A into Z with equal probability x of {Pul An} and {Int An}, that is, Pul Ak and Int Ak are inserted alternately, where Pul Ak and Int Ak are both randomly selected from {Pul An} and {Int An}; if the total length of the synthesized sequence exceeds L after the last inserted sequence, then truncate the last sequence before inserting;
[0036] For type B devices, reconstruct sequence Z according to the device dependency distribution probability γ, specifically including: randomly generating a probability number P x , if P x > γ, then insert 1 zero value; P xIf ≤γ, then insert a pulse sequence Pul Bk , where Pul Bk is also randomly selected from {Pul Bn}; similarly, if the total length of the synthesized sequence exceeds L after the last inserted sequence, then truncate the last sequence before inserting it;
[0037] Divide the data set into a training set and a test set at a certain ratio;
[0038] In addition, to solve the overfitting problem in small-sample data, during the process of inserting Pul Bk , expand or truncate Pul Bn at a certain ratio, so that the later neural network can learn richer features.
[0039] Step 3: Synthesize the training set, including inserting pulse or interval feature sequences into the initial value sequence according to different methods according to different device densities, that is, randomly insert device pulses or interval segments with equal probability for high-density devices, and randomly generate a probability value for low-density devices and compare it with the device dependency distribution probability. If the probability value is greater than the distribution probability, insert an initial value, otherwise randomly insert a pulse segment of the device. Finally, if the total length exceeds the length of the initial value sequence, truncate the last inserted segment. The length of the initial value sequence can be set to the length of the regular training set in Step 4.
[0040] Preferably, in Step 4, the lightweight CNN-Seq2point network refers to using the Seq2point end-to-end network architecture. The input data is the total bus power of the synthesized training set, intercepted by a sliding window of 599 data points, and the output is the midpoint of the window, which records the power prediction value of the current electrical appliance being trained; the network model uses CNN, effectively utilizing the characteristic that the CNN neural network is not restricted by time, facilitating any shuffling operation on the data set;
[0041] The lightweight CNN-Seq2point neural network used in the present invention is an existing neural network, and its structure can refer to the prior art;
[0042] The training process adopts the method of cross-validation, including dividing the limited data set into n parts. Each time during training, n - 1 of them are used as the training set, and the remaining 1 part is used as the validation set. In this way, all validation sets are traversed, and the total number of times is n. This operation can make full use of the source data and achieve a sufficient training volume in the case of small samples.
[0043] Preferably, in step 5, predicting the test set mainly refers to fitting the power curves of each electrical appliance in the test set. The model records the parameters of the neural network. After being trained in step 4, several models will be obtained (each electrical appliance corresponds to a model file). Loading the model of a certain device into a neural network with the same structure as that in step 4 and inputting the bus power curve of the test set, the predicted output value of this device can be decomposed. By comparing the output value with the true power value of the corresponding electrical appliance in the test set to calculate the acc, the quality of the model trained previously can be judged according to the size of the acc.
[0044] Preferably, in step 6, the loss calculation for error calculation includes, but is not limited to, analyzing the interference situation of power prediction between different electrical appliances by error functions such as mean square error MSE, mean error ME, mean absolute error MAE, accuracy acc, Huber-Loss, etc., calculating the error between the segmented pulse or interval segment and the true value, and the error between the entire pulse or interval segment and the true value.
[0045] Preferably, the overlap analysis in step 6 specifically refers to the overlap degree of the ON state of a certain segmented pulse or interval segment of the measured device with the corresponding timestamp segment of a similar device. The threshold is set to 70%, that is, if it is greater than 70%, it is considered to have overlap, otherwise there is no overlap.
[0046] Preferably, the contingency table in step 6 specifically includes:
[0047] Taking all the segmented pulses and interval segments of the power data of a certain device predicted in step 5 as the population, counting the number of segmented pulses and interval segments that satisfy both "the error value of this segmented pulse or interval segment is greater than the error value of the entire pulse or interval segment" and "there is overlap between this device and one of its similar devices", as well as the number of segmented pulses and interval segments that only satisfy one of them and those that satisfy neither, and making a contingency table based on this, and calculating the chi-square value of the four tables to characterize the possibility that this similar device has an impact on the prediction result of a certain segmented pulse or interval segment of this device.
[0048] Preferably, when the chi-square value is greater than 90%, it is considered that the influence degree is relatively large, and then return to step 3 to optimize the composition of the synthetic data set, that is, increase the segmented pulses or interval segments of this device affected by the similar device, and resynthesize the "pulse-interval" sequence.
[0049] For details not elaborated in the present invention, reference can be made to the prior art.
[0050] The beneficial effects of the present invention are:
[0051] 1. The present invention extracts feature data from the input measurement data. By introducing statistical analysis methods and reconstructing the data set using the extracted features, the feature density of small-sample and low-density devices is increased. At the same time, an error analysis method is introduced to quickly perform secondary optimization on the prediction results, making the prediction results more accurate. It has significant performance in the field of non-intrusive load decomposition, with strong anti-interference ability and high precision.
[0052] 2. The present invention can make full use of the self-collected small-sample data set for neural network model training. Through two major steps of feature extraction and data set regeneration, the device density is increased on the original device dependence relationship, and more device features are retained, making the neural network more accurate in identifying small-sample data sets. And because the data is randomly generated each time, it ensures that the neural network training can have an infinite data set, reducing overfitting to a certain extent.
[0053] 3. Aiming at the problems of inconsistent evaluation criteria for NILM containing neural networks and difficulty in improving the network, the present invention provides a general error analysis method, which can be used to evaluate the network quality (initial training results), and accordingly adjust the feature ratio and composition of the devices in the input data set to obtain a more satisfactory final training result. BRIEF DESCRIPTION OF THE DRAWINGS
[0054] The specification drawings forming a part of this application are used to provide a further understanding of this application. The schematic embodiments of this application and their descriptions are used to explain this application and do not constitute an improper limitation of this application.
[0055] Figure 1 is a flowchart of the present invention;
[0056] Figure 2 is a dependency relationship distribution matrix of a certain embodiment, where A1 to A3 represent A-type devices 1, 2, and 3, and B1 to B3 represent B-type devices 1, 2, and 3;
[0057] Figure 3 is a schematic diagram of "pulse-interval" sequence extraction of a certain embodiment, {Pul n} is the pulse sequence, {Int n} is the interval sequence;
[0058] Figure 4 is the chi-square calculation result of the four-grid table of the pulse sequence of a certain embodiment, including a short pulse and a long pulse. The error criteria selected are me, mae, and acc (accuracy rate). The research device is PC, the similar device 1 is LightBulb, and the similar device 2 is waterdispenser;
[0059] Figure 5The chi-square calculation results of the interval sequence four-grid table for a certain embodiment, including a short pulse and a long pulse. The error criteria selected are me, mae, and acc (accuracy). The research device is a PC, the similar device 1 is a LightBulb, and the similar device 2 is a water dispenser;
[0060] Figure 6 The prediction result one of the power curve of the water dispenser for a certain embodiment;
[0061] Figure 7 The prediction result two of the power curve of the water dispenser for a certain embodiment;
[0062] Figure 8 The prediction result of the power curve of the PC for a certain embodiment. Detailed implementation manners
[0063] To enable those skilled in the art to better understand the technical solutions in this specification, the technical solutions in the embodiments of this invention will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of this specification. However, this is not limited to this. For those parts that are not elaborated in this invention, they are all in accordance with the conventional technologies in the art.
[0064] Embodiment 1
[0065] A high-precision low-frequency NILM neural network method based on a synthetic training set, as Figure 1 shown, includes the following steps:
[0066] Step 1, data preprocessing: Collect the bus power and the power data of each electrical appliance, and perform data filling and label making;
[0067] Step 2, feature data extraction: Apply the threshold determination method and statistical calculation formulas to the data obtained after preprocessing in Step 1 to extract and calculate the corresponding feature data;
[0068] Step 3, synthetic data set: Reconstruct the data set based on statistical methods for the feature data obtained in Step 2 to prevent overfitting;
[0069] Step 4, lightweight CNN-Seq2point neural network training: Use the data set synthesized in Step 3, with the bus power as the input and the power of a certain electrical appliance as the true value to calculate the loss function, and use the lightweight CNN-Seq2point model structure to train the network model parameters model of this electrical appliance;
[0070] Step 5, Decomposition and Output: Load the model obtained in Step 4 into the lightweight CNN-Seq2point network. Feed the bus power data in the test set into the lightweight CNN-Seq2point network. Using the data of the electrical appliance to which the network belongs as the ground truth, the power prediction data and prediction accuracy of the electrical appliance to which the network model belongs are output, which is the preliminary prediction result.
[0071] Step 6, Error Analysis: Calculate the error between the prediction result obtained in Step 5 and the ground truth data of the same electrical appliance at the same time stamp. Meanwhile, conduct an overlap analysis on the same electrical appliance and electrical appliances with similar working states. According to the error calculation results and overlap analysis results, count several four-fold tables, and judge the influence degree and influence position of similar electrical appliances on this electrical appliance based on the four-fold tables.
[0072] Step 7, Optimize the Preliminary Prediction Result: According to the judgment in the error analysis of Step 6, adjust the composition of the synthetic dataset in Step 3, and then repeat Steps 4 and 5 until satisfactory training and prediction results (lower loss value or higher accuracy) are obtained.
[0073] Embodiment 2
[0074] A high-precision low-frequency NILM neural network method based on a synthetic training set, as described in Embodiment 1. The difference is that in Step 1, the bus power and each electrical appliance power data (i.e., power sequence), mainly referring to small sample data, are obtained by collecting the dataset domestic, the public dataset refit, the public dataset REDD (low frequency), etc. For problems such as sampling breakpoints in some devices in the sampled power sequence and inconsistent sampling intervals for different electrical appliances, the following measures are taken:
[0075] First, cut and splice at the breakpoint to obtain a group of power sequences without breakpoints. Then, specify a unified sampling interval T. For power sequences with an original sampling interval greater than T, interpolation processing is used; for power sequences with an original sampling interval less than T, decimation processing is used. A series of null values NULL will be generated during this process. Finally, the "mean compensation method" is used to fill in the missing values. Here, the sampling breakpoint refers to the situation where the sampling interval of a certain device at a certain moment or within a certain small time period is much larger than the normal sampling interval of the device. The mean compensation method refers to calculating the average value of n points near the null value NULL and making the value at the NULL be
[0076] Label making refers to establishing the correspondence between the device name and the device labels involved in the training process to facilitate the subsequent call of the neural network. The device labels include, but are not limited to, the mean, standard deviation, the house where the device is located, and the channel where the device is located. Among them, the mean and standard deviation are set using the trial-and-error method, and the trial values are given based on the device power curve, device switching state curve, device power distribution curve, etc.
[0077] Embodiment 3
[0078] A high-precision low-frequency NILM neural network method based on a synthetic training set, as described in Embodiment 2. The difference is that in Step 1 and Step 2, the feature data includes the pulse intrinsic distribution, "pulse-interval" sequence, device dependency distribution matrix, and device dependency distribution probability.
[0079] Among them, the "pulse" and "interval" refer to scanning a specified power sequence at a specified threshold: the power segment above the threshold is recorded as the pulse sequence {Pul n} = [Pul1, Pul2,..., Pul n , where n represents the total number of extracted pulse or interval segments; the power segment below the threshold is recorded as the interval sequence {Int n} = [Int1, Int2,..., Int n , as Figure 3 shown, the threshold is determined using the trial-and-error method.
[0080] The pulse intrinsic distribution, that is, the device density, characterizes the density of the device's own operation, referring to the ratio of the total pulse time T pul of the device to the total time T total of its power sequence.
[0081] The "pulse-interval" sequence includes the noise "pulse-interval" sequence, the device "pulse-interval" sequence, and the bus "pulse-interval" sequence. Among them,
[0082] a) The noise "pulse-interval" sequence refers to the "pulse-interval" sequence extracted from the difference sequence P t between the bus power sequence P a and the sum of the appliance power sequences ∑P Δ = |P t - ∑P a |.
[0083] b) The device "pulse-interval" sequence refers to the "pulse-interval" sequence extracted from the power sequence {P ax} of each appliance, where x refers to the xth appliance; the following definitions are made for Class A and Class B devices:
[0084] Define α > 30% as high-density devices, i.e., Class A devices, and α ≤ 30% as low-density devices, i.e., Class B devices, where 30% is given according to the Monte Carlo algorithm;
[0085] For Class A devices, the pulse sequence {Pul An} and the interval sequence {Int An} need to be extracted separately; for Class B devices, only the pulse sequence {Pul Bn} needs to be extracted, and its interval sequence is regarded as a zero-value sequence;
[0086] c) The bus "pulse-interval" sequence refers to summing the noise "pulse-interval" sequence in a) and the device "pulse-interval" sequence in b) respectively to obtain the bus "pulse-interval" sequence;
[0087] The device dependency distribution matrix characterizes the degree of dependence of the "on" states of two devices. Denote the total pulse time of Class A devices as T pulA , and the number of times Class B devices enter the pulse state as C on , then this degree of dependence β can be calculated by the following formula: Let the total number of Class A devices be N A , and the total number of Class B devices be N B , then the size of the matrix is N A ×N B ;
[0088] The device dependency distribution probability γ characterizes the density index of controlling the "on" state of Class B devices. Based on the device dependency distribution matrix, find the maximum dependence β of Class B devices max (take the maximum value of β for each column of the matrix (i.e., each type of Class B device), that is, the β values of different Class B devices are different), and at the same time calculate the dependence standard deviation β max as the pessimistic error estimate value, then γ = β sd ±β max ; as shown in sd ; Figure 2 shown;
[0089] Among them, p is the p-th Class B device, and q is the total number of Class A devices.
[0090] The extraction of characteristic data in Step 2 includes dividing the data "pulse" and data "interval" through threshold determination, statistically calculating the ratio of the total time of the pulse segments of the device to the total time of its power sequence as the device density, statistically calculating the ratio of the number of times the low-density device enters the pulse state to the total pulse time of the high-density device as the device dependency distribution matrix, and statistically calculating the pessimistic error estimate value of the matrix as the device dependency distribution probability.
[0091] Example 4
[0092] A high-precision low-frequency NILM neural network method based on a synthetic training set. As described in Embodiment 3, the difference is that in Step 1, the process of synthesizing the data set in Step 3 is as follows:
[0093] Determine the length L of the synthetic sequence according to the general length of the preprocessed power sequence, and generate a set of zero-value sequences Z with length L;
[0094] For type A devices, insert device A into Z with equal probability x of {Pul An} and {Int An}, that is, Pul Ak and Int Ak are inserted alternately, where Pul Ak and Int Ak are both randomly selected from {Pul An} and {Int An}; if the last inserted sequence makes the total length of the synthetic sequence exceed L, then truncate the last sequence and insert it again;
[0095] For type B devices, reconstruct sequence Z according to the device dependency distribution probability γ, specifically including: randomly generate a probability number P x , if P x >γ, then insert 1 zero value; if P x ≤γ, then insert a pulse sequence Pul Bk , where Pul Bk is also randomly selected from *Pul Bn}; similarly, if the last inserted sequence makes the total length of the synthetic sequence exceed L, then truncate the last sequence and insert it again;
[0096] Divide the data set into a training set and a test set at a certain ratio;
[0097] In addition, to solve the overfitting problem in small-sample data, during the process of inserting Pul Bk , Pul Bn is expanded or truncated at a certain ratio so that the later neural network can learn richer features.
[0098] Step 3: Synthesize the training set, which includes inserting pulse or interval feature sequences into the initial value sequence according to different methods based on device density. That is, for high-density devices, device pulses or interval segments are randomly inserted with equal probability, while for low-density devices, a probability value is randomly generated and compared with the distribution probability of device dependencies. If the probability value is greater than the distribution probability, an initial value is inserted; otherwise, a pulse segment of the device is randomly inserted. Finally, if the total length exceeds the length of the initial value sequence, truncation is performed on the last inserted segment, and the length of the initial value sequence can be set to the length of the conventional training set in Step 4.
[0099] Example 5
[0100] A high-precision low-frequency NILM neural network method based on a synthesized training set, as described in Example 4. The difference is that in Step 4, the lightweight CNN-Seq2point network refers to using a Seq2point end-to-end network architecture. The input data is the total bus power of the synthesized training set, intercepted by a sliding window of 599 data points, and the output is the midpoint of the window, which records the power prediction value of the electrical appliance being trained. The network model uses CNN, effectively utilizing the characteristic of the CNN neural network being unrestricted by time, facilitating arbitrary shuffling operations on the dataset.
[0101] The lightweight CNN-Seq2point neural network used in the present invention is an existing neural network, and its structure can refer to the prior art.
[0102] The training process adopts a cross-validation method, which includes dividing the limited dataset into n parts. Each time during training, n - 1 parts are used as the training set, and the remaining 1 part is used as the validation set. In this way, all validation sets are traversed, and the total number of times is n. This operation can make full use of the source data and achieve a sufficient training volume in the case of small samples.
[0103] Example 6
[0104] A high-precision low-frequency NILM neural network method based on a synthesized training set, as described in Example 5. The difference is that in Step 5, predicting the test set mainly refers to fitting the power curves of each electrical appliance in the test set. The model records the parameters of the neural network, and several models (each electrical appliance corresponds to a model file) are obtained through training in Step 4. Loading the model of a certain device into a neural network with the same structure as in Step 4 and inputting the test set total bus power curve, the predicted output value of the device can be decomposed and obtained. By comparing the output value with the true power value of the corresponding electrical appliance in the test set to calculate acc, the quality of the model trained previously can be judged according to the size of acc.
[0105] Example 7
[0106] A high-precision low-frequency NILM neural network method based on a synthetic training set, as described in Embodiment 6. The difference is that in step 6, the loss calculation for error calculation includes, but is not limited to, error function analysis of different electrical appliances for power prediction using mean square error MSE, mean error ME, mean absolute error MAE, accuracy acc, Huber-Loss, etc., calculating the error between the segmented pulse or interval segment and the true value, and the error between the entire pulse or interval segment and the true value.
[0107] The overlap analysis in step 6 specifically refers to the overlap degree of the ON state of a certain segmented pulse or interval segment of the measured device with the corresponding timestamp segment of a similar device, and the threshold is set to 70%, that is, if it is greater than 70%, it is considered to have overlap, otherwise there is no overlap.
[0108] The contingency table statistics in step 6 specifically include:
[0109] Taking all the segmented pulses and interval segments of the power data of a certain device predicted in step 5 as the population, counting the number of segmented pulses and interval segments that satisfy both "the error value of this segmented pulse or interval segment is greater than the error value of the entire pulse or interval segment" and "there is overlap between this device and one of its similar devices", as well as the number of segmented pulses and interval segments that only satisfy one of them and those that satisfy neither, and making a contingency table based on this, and calculating the chi-square value of the four tables to characterize the possibility that this similar device has an impact on the prediction result of a certain segmented pulse or interval segment of this device.
[0110] When the chi-square value is greater than 90%, it is considered that the degree of influence is relatively large, and then return to step 3 to optimize the composition of the synthetic data set, that is, increase the segmented pulse or interval segment of this device affected by the similar device, and resynthesize the "pulse-interval" sequence.
[0111] For a single electrical appliance, the following processing is done: The pulse is divided into two or three segments according to the pulse length, and the error of each segment is calculated separately. If the pulse length is greater than 2 * 599, the pulse is divided into three segments, where the first segment and the last segment both contain 599 points, and the middle segment contains the remaining (L - 2×599) data points; otherwise, the pulse is evenly divided into two segments. The interval sequence is segmented in the same way.
[0112] For studying the mutual influence between electrical appliances (this is especially applicable to electrical appliances with unsatisfactory prediction results, such as light bulbs), the following calculations are done:
[0113] First, two scenarios are defined. Scenario 1: Whether the error of a certain segment of Q alone is larger than the overall error value (i.e., the training error), and this will generate 2 or 3 contingency tables, specifically depending on whether Q is a long-pulse electrical appliance or a short one;
[0114] Scenario 2: Electrical appliances *A1,A2,...,A nWhether the pulse of the electrical appliance Q overlaps with this section, the overlap standard is set to 70%, that is, if the overlap degree ≥ 70%, it is recorded as overlap, otherwise it is not recorded as overlap. The so-called overlap refers to the time-domain overlap (aligned by timestamp), that is, multiple devices work or are in the ON state simultaneously in the same time period, and thus n fourfold tables will be generated.
[0115] A number of long pulses {Pul ln} can be extracted from a relatively long power data of the device Q, a number of short pulses {Pul sn} and a number of intervals {Int n}. After segmenting the pulses and intervals according to the above criteria, assuming that only the first segment of the long pulse Pul k is concerned, only the influence of the electrical appliance A1 on the electrical appliance Q is concerned, and only the MSE error is calculated. In this way, 1 fourfold table will be generated, as shown in Table 1 below:
[0116] Table 1: A certain fourfold table in the embodiment
[0117] Scenario 1T Scenario 1F Scenario 2T a b Scenario 2F c d
[0118] Among them, T represents true "yes", F represents False "no", and a represents the number of pulses that satisfy that the MSE error of the first segment of Pul k is greater than the overall MSE error of Pul k and the first segment of the pulse overlaps with A1. The χ k value finally calculated in this way characterizes the possibility that the device A1 has a greater influence on the prediction result of the first segment of the long pulse of the electrical appliance Q. 2
[0119] Calculate the chi-square value χ 2 according to the fourfold table, which characterizes the possibility that the training error of the device Q is affected by the pulses of other devices. The larger the χ 2 value, the greater this possibility, or in other words, the higher the correlation degree between the corresponding electrical appliances A x and Q.
[0120] According to the calculated χ 2 value table, feedback to step (3) to synthesize the training set, and find the segmented pulses and intervals of the device Q and the device A x (there may be more than 1, and the device A( x is one of all devices other than the device Q, and this A( x can be some devices that are most likely to affect the prediction result of Q according to the prediction result), and re-synthesize the "pulse-interval" sequence of the device Q. During the synthesis process, increase the number of this difference feature (that is, return to step (3) and randomly insert some more difference features). Empirical evidence shows that the prediction result in this way is better.
[0121] As Figure 4 , 5 shown, they are respectively a schematic diagram of the chi-square calculation result of a pulse sequence four-grid table and a schematic diagram of the chi-square calculation result of an interval sequence four-grid table.
[0122] Figure 6 In , "Aggregate" represents the bus power curve, "Ground Truth" represents the true power curve of the device (water dispenser), and "Predicted" represents the power prediction output curve of the device by the neural network after secondary optimization. In particular, both the bus and the true value of the device are from the test set part of the domestic dataset processed in Example 4. The picture shows the prediction effect of a test window of the Seq2point lightweight neural network. The effects of the remaining windows are basically the same as this window. It can be seen that the predicted value (Predicted) of the device is basically the same as the true value (Ground Truth) of the device, indicating the feasibility of this method.
[0123] Figure 7 Different from Figure 6 , in order to more clearly see the prediction effect, Figure 7 there are only two curves of the predicted value (Predicted) of the device and the true value (Ground Truth) of the device.
[0124] Figure 8 In , "Ground Truth" represents the true power curve of the device (PC), and "Predicted" represents the power prediction output curve of the device by the neural network after secondary optimization. In particular, the true value of the device is from the test set part of the domestic dataset processed in Example 4. The picture shows the prediction effect of a test window of the Seq2point lightweight neural network. The effects of the remaining windows are basically the same as this window. It can be seen that the predicted value (Predicted) of the device is basically the same as the true value (Ground Truth) of the device, indicating the feasibility of this method.
[0125] The above is the preferred implementation manner of the present invention. It should be noted that for those of ordinary skill in the art of this technology, without departing from the principle described in the present invention, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present invention.
Claims
1. A high-precision low-frequency NILM neural network method based on a synthetic training set, characterized in that It includes the following steps: Step 1, data preprocessing: Collect bus power and individual electrical appliance power data, and perform data completion and label making; Step 2, feature data extraction: Apply the threshold determination method and statistical calculation formula to the data obtained after preprocessing in Step 1 to extract and calculate feature data; Step 3, synthetic dataset: Reconstruct the dataset based on the statistical method for the feature data obtained in Step 2; Step 4, lightweight CNN-Seq2point neural network training: Use the dataset synthesized in Step 3, with the bus power as the input and the power of a certain electrical appliance as the true value to calculate the loss function, and use the lightweight CNN-Seq2point model structure to train the network model parameters model of this electrical appliance; Step 5, decomposition output: Load the model obtained by training in Step 4 into the lightweight CNN-Seq2point network, send the bus power data in the test set into the lightweight CNN-Seq2point network, and use the data of the electrical appliance to which the network belongs as the true value, then output the power prediction data and prediction accuracy of the electrical appliance to which the network model belongs, which is the preliminary prediction result; Step 6, error analysis: Calculate the error between the prediction result obtained in Step 5 and the true value data of the same electrical appliance at the same timestamp, and at the same time perform an overlap analysis on this electrical appliance and electrical appliances with similar working states. According to the error calculation result and overlap analysis result, count several contingency tables, and judge the influence degree and influence position of similar electrical appliances on this electrical appliance based on the contingency tables; Step 7, optimize the preliminary prediction result: According to the judgment of the error analysis in Step 6, adjust the composition of the synthetic dataset in Step 3, and then repeat Steps 4 and 5 until satisfactory training and prediction results are obtained; In Step 2, the feature data includes the pulse eigen distribution and the device dependence distribution probability; Among them, the "pulse" and "interval" refer to scanning a specified power sequence under a specified threshold: power segments higher than the threshold are recorded as a pulse sequence {Pul n} = [Pul1, Pul2,..., Pul n , where n represents the total number of extracted pulse or interval segments; power segments lower than the threshold are recorded as an interval sequence {Int n} = [Int1, Int2,..., Int n ; The pulse eigen-distribution, i.e., the device density, characterizes the density at which the device itself operates, referring to the total pulse time T of the device pul occupying the total time T of its power sequence total ratio Make the following definitions for type A devices and type B devices: Define α>30% as high-density devices, that is, type A devices, and α≤30% as low-density devices, that is, type B devices; The device dependency distribution probability γ characterizes the density exponent of the "on" state of the control Class B devices. Based on the device dependency distribution matrix, the maximum dependency β of the Class B devices is found max , and at the same time, the dependency standard deviation β is calculated sd As the pessimistic error estimate value, then γ = β max ±β sd ; Among them, p is the p-th type-B device, and q is the total number of type-A devices; The process of synthesizing the dataset in Step 3 is as follows: Determine the length L of the synthetic sequence according to the general length of the power sequence after preprocessing, and generate a set of zero-value sequences Z with length L; For type A devices, insert device A into Z with equal probability x of {Pul An} and {Int An}, that is, Pil Ak and Int Ak are inserted alternately, where Pil Ak and Int Ak are both randomly selected from {Pil An} and {Int An}; if the total length of the synthesized sequence exceeds L after the last insertion, truncate the last sequence and then insert it; For Class B devices, reconstruct the sequence Z according to the device dependency distribution probability γ, specifically including: randomly generating a probability number P x , if P x >γ, then insert a 0 value; if P x ≤γ, then insert a pulse sequence Pul Bk , where Pul Bk is also randomly selected from {Pul Bn}; similarly, if the total length of the synthesized sequence exceeds L after the last insertion, truncate the last sequence and then insert it.
2. The high-precision low-frequency NILM neural network method based on a synthetic training set according to claim 1, wherein In Step 1, for the problems of sampling breakpoints in some devices in the sampled power sequence and inconsistent sampling intervals for different electrical appliances, take the following measures: First, cut and splice at the breakpoint to obtain a set of power sequences without breakpoints; then, specify a unified sampling interval T. For power sequences with an original sampling interval greater than T, interpolation processing is used; for power sequences with an original sampling interval less than T, decimation processing is used; this process will generate a series of null values NULL. Finally, the "mean compensation method" is used to fill in the missing values; among them, the sampling breakpoint refers to a situation where the sampling interval is much larger than the conventional sampling interval of a certain device at a certain moment or within a certain small time period; the mean compensation method refers to calculating the average value of n points near the null value NULL and set the value at NULL to Label making refers to establishing a corresponding relationship between the device name and the device labels involved in the training process. The device labels include but are not limited to the mean, standard deviation, the house where the device is located, and the channel where the device is located.
3. The high-precision low-frequency NILM neural network method based on a synthetic training set according to claim 2, wherein In Step 2, the feature data also includes the "pulse-interval" sequence and the device dependence distribution matrix; The "pulse-interval" sequence includes the noise "pulse-interval" sequence, the device "pulse-interval" sequence, and the bus "pulse-interval" sequence, where, a) The noise "pulse-interval" sequence refers to the sequence P t obtained from the difference between the bus power sequence P a and the sum of the appliance power sequences ∑P Δ = |P t - ∑P a |, which is the "pulse-interval" sequence extracted; b) The "pulse-interval" sequence of the device refers to the "pulse-interval" sequence extracted from the power sequences {P ax} of each electrical appliance, where x refers to the x-th electrical appliance; For type A devices, the pulse sequence {Pul An} and the interval sequence {Int An} need to be extracted separately; for type B devices, Just extract the pulse sequence {Pul Bn}, and regard its interval sequence as a sequence of 0 values; c) The bus "pulse-interval" sequence refers to summing the noise "pulse-interval" sequence in a) and the device "pulse-interval" sequence in b) respectively to obtain the "pulse-interval" sequence of the bus; The device-dependency distribution matrix characterizes the degree of dependence between the "on" states of two devices. Denote the total pulse time of type A devices as T pulA , and the number of times type B devices enter the pulse state as C on . Then this degree of dependence β can be calculated using the following formula: Let the total number of type A devices be N A , and the total number of type B devices be N B . Then the size of the matrix is N A ×N B .
4. The high-precision low-frequency NILM neural network method based on a synthetic training set according to claim 3, wherein In Step 3, divide the dataset into a training set and a test set at a certain ratio.
5. The high-precision low-frequency NILM neural network method based on a synthetic training set according to claim 4, wherein In step 4, the lightweight CNN-Seq2point network refers to a Seq2point end-to-end network architecture. The input data is the bus power of the synthesized training set, and the output is the midpoint of the window, which records the power prediction value of the electrical appliance being trained. The training process uses cross-validation, which includes dividing the finite data set into n parts. Each time during training, n - 1 of these parts are used as the training set, and the remaining 1 part is used as the validation set. By traversing all the validation sets in this way, the total number of times is n.
6. The high-precision low-frequency NILM neural network method based on a synthetic training set according to claim 5, wherein In step 6, the loss calculation for error calculation includes, but is not limited to, mean squared error (MSE), mean error (ME), mean absolute error (MAE), accuracy (acc), and Huber-Loss error function. Calculate the error between the segmented pulse or interval segment and the true value, as well as the error between the entire pulse or interval segment and the true value.
7. The high-precision low-frequency NILM neural network method based on the synthetic training set according to claim 6, characterized in that The overlap analysis in step 6 specifically refers to the on-state overlap of a certain segmented pulse or interval segment of the measured device with the corresponding timestamp segment of a similar device. The threshold is set at 70%, that is, if it is greater than 70%, it is considered to have overlap, otherwise there is no overlap.
8. The high-precision low-frequency NILM neural network method based on a synthetic training set according to claim 7, wherein The contingency table in step 6 specifically includes: Taking all the segmented pulses and interval segments of the power data of a certain device predicted in step 5 as the population, count the number of segmented pulses and interval segments that satisfy both "the error value of this segmented pulse or interval segment is greater than the error value of the entire pulse or interval segment" and "there is overlap between this device and one of its similar devices", as well as the number of segmented pulses and interval segments that only satisfy one of them and those that satisfy neither. Based on this, a contingency table is made, and the chi-square value of the four tables is calculated to characterize the possibility that this similar device has an impact on the prediction result of a certain segmented pulse or interval segment of this device.
9. The high-precision low-frequency NILM neural network method based on a synthetic training set according to claim 8, wherein When the chi-square value is greater than 90%, it is considered that the degree of influence is relatively large, and then return to step 3 to optimize the composition of the synthesized data set, that is, increase the segmented pulse or interval segment of this device affected by the similar device, and resynthesize the "pulse-interval" sequence.
Citation Information
Patent Citations
Non-intrusive industrial load decomposition method based on deep learning
CN115482123A
Data dual-drive method, apparatus, and device for predicting power grid failure during typhoon
WO2023045278A1