High-frequency-oriented non-intrusive household load decomposition system based on multi-source data
Through the multi-source data fusion and dynamic convolution analysis module, a non-invasive household load decomposition system for high-frequency is built, which solves the problems of environmental factors affecting and insufficient high-frequency signal capture in the existing technology, and achieves more efficient and accurate household load analysis.
Patent Information
- Application Number
- CN202510210473.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-25
- Publication Date
- 2025-06-27
AI Technical Summary
The existing non-invasive household load decomposition methods fail to effectively consider the impact of the environment on users' power usage habits, and when processing high-frequency data, fixed-length patch schemes are difficult to capture the complete high-frequency signal cycle, resulting in insufficient comprehensive analysis.
Using a multi-source data fusion method, combining variable patches and dynamic convolution analysis modules, a non-invasive home load decomposition system based on multi-source data is constructed for high-frequency. The system dynamically adjusts the sampling grid of convolution operations through adaptive cutting modules and multi-source analysis architecture models to better capture high-frequency signals.
This method not only takes into account the impact of environmental factors on users' power usage habits more comprehensively, but also dynamically analyzes high-frequency data to avoid cutting into discontinuous patches, thereby improving the analysis ability and accuracy of high-frequency electrical data.
Smart Images

Figure CN120218471A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of non-intrusive household load analysis in electrical data analysis algorithms, and specifically to an algorithm that uses multi-source data fusion means, faces the characteristics of high-frequency data, and uses variable patches and dynamic convolution analysis modules to improve the analysis ability. Background Technique
[0002] Non-intrusive household load decomposition can, first, predict the electricity demand on the user side, thereby providing quantitative indicators for the power generation plan on the power generation side and reasonably arranging the power generation tasks of each power generation source; second, it can understand the electrical appliance ownership of each household, thereby indirectly understanding the living standards of household users and providing auxiliary data for the government's decision-making.
[0003] The existing non-intrusive household load decomposition methods mainly analyze the data directly from parameters such as current on the electricity meter. They mainly have two disadvantages. One is that they do not consider the influence of the environment on the user's electricity consumption habits. For example, although the overall increase in temperature in summer will cause users to use air conditioners, if the environment in the area where they are located has a high altitude and strong wind, then the willingness to use air conditioners will be reduced, and the economic level will also affect the electricity consumption habits. The other is that high-frequency data is not only the characteristics of some devices such as microwave ovens and kettles, but sometimes it is also some noise.
[0004] Through analysis, for the first disadvantage, obviously, auxiliary data sources such as meteorology, geography, and economy are needed for auxiliary analysis, but which data sources (i.e., channels such as temperature, humidity, and wind direction) are effective is uncertain. For the second disadvantage, in the existing fixed patch solution in multi-source fusion prediction methods, it is not easy to capture the complete change range near the high-frequency signal. It is very likely to cut the effective period of the high-frequency signal into multiple patches, thus unable to effectively capture a complete high-frequency signal period, and these are generally the usage periods of electrical appliances similar to a few minutes.
[0005] Therefore, if we want to better solve the two main disadvantages faced by the existing non-intrusive household load decomposition methods, then we need to improve the existing model, not only need to add auxiliary data sources, but also need to add new solutions to capture the complete high-frequency signal period. Summary of the Invention
[0006] Aiming at the above-mentioned technical deficiencies, the purpose of the present invention is to provide a high-frequency-oriented non-intrusive household load decomposition system based on multi-source data, which can not only take into account the influence of the environment on users' electricity consumption habits, but also dynamically analyze high-frequency data to prevent high-frequency data from being cut into multiple discontinuous patches due to the fixed length; thus, more comprehensively analyze the data channel data, better focus on high-frequency information, and strengthen the analysis ability of high-frequency electrical appliance data.
[0007] To solve the above technical problems, the present invention adopts the following technical solutions: A high-frequency-oriented non-intrusive household load decomposition system based on multi-source data, comprising: a dataset construction module, an adaptive cutting module, a multi-source analysis architecture, and a variable convolution module; The dataset construction module is used to construct a multi-source time-series dataset with time as the coordinate axis, which takes the electricity consumption on the electricity meter as the main data and uses historical real meteorological data, weather forecast data, geographical information data, local economic data, and real user household electrical appliance data willing to participate in the test as auxiliary data to construct an auxiliary dataset; The adaptive cutting module is used to cut the time-series data of the dataset before inputting it into the model; in order to obtain patches with more comprehensive semantic information, it can adaptively expand the length of the patches and adopt bilinear interpolation to uniformly sample a fixed number of time steps within each patch, thus a more complete semantic patch division strategy; The multi-source analysis architecture model is used to perform data fusion on various multi-source auxiliary data; The variable convolution module is used to dynamically adjust the sampling grid of each convolution operation; it calculates the offset for each sampling point to adjust its receptive field, so that it can better match the unique features of the input data to accurately capture the specific information of the target, thereby improving the performance of data analysis.
[0008] Furthermore, the multi-source analysis architecture model adopts the ST-SAN architecture model as the basic framework model, which includes three parts: Conv-LSTM, soft attention mechanism, and SVM-RFE module; among them, Conv-LSTM is used to perform convolution operations on the input data to extract spatial features, and then apply the LSTM unit to capture time-dependent relationships, and retain the spatial structure information through state transfer; The soft attention mechanism is used to selectively ignore some information to perform re-weighted aggregation calculation on the remaining information. All information will be re-weighted in an adaptive manner before being aggregated, and then important information is separated; SVM-RFE is used to select the most influential features on the model from the original multi-source data, thereby improving the running efficiency and generalization ability of the model. It can automatically sort and eliminate according to the weights of the features, effectively reducing the number of features, reducing the risk of overfitting, and enhancing the model interpretability and computational efficiency.
[0009] Furthermore, in the specific operation of the deformable convolution module, for deformable convolution, each element of the convolution kernel has an independent offset parameter; these parameters allow the convolution kernel to deform during the training process, so as to be able to more flexibly adapt to the shape of the input features; by setting the weights to a uniform distribution in the initialization stage and initializing the bias to zero, it helps to optimize the network in the early stage of training; Furthermore, the deformable convolution module changes the Conv part in the Conv-LSTM of the ST-SAN module to deformable convolution and adds weights, as shown in the following formula: ; In the formula, represents in a convolution window, is the position of the center point; p k is the 1-norm distance of the k-th point in the convolution window from the p0 point, is the deformation offset, is the weight of the k-th point, is the scaling factor, which reflects the multiple by which the current data image is scaled or enlarged.
[0010] The beneficial effects of the present invention are as follows: On the basis of adopting the means of multi-source data fusion, aiming at data misalignment and the characteristics of high-frequency data, the present invention uses variable patches and dynamic convolution analysis modules, so as to improve the analysis ability of the algorithm, which has the characteristics of high analysis accuracy for high-frequency data (household appliances with high usage frequency), and at the same time the computational amount is not large.
[0011] The present invention can not only take into account the influence of the environment on the user's electricity consumption habits, but also dynamically analyze high-frequency data to prevent high-frequency data from being cut into multiple discontinuous patches due to the reason of fixed length; thus, it can more comprehensively analyze the data channel data, better focus on high-frequency information, and thereby enhance the analysis ability of high-frequency electrical appliance data. It can decompose the load of the household target and provide a data basis for user-side demand prediction. BRIEF DESCRIPTION OF THE DRAWINGS
[0012] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained according to these drawings.
[0013] Figure 1 Schematic structural diagram of the learning stage of a high-frequency-oriented non-intrusive household load decomposition system based on multi-source data provided by an embodiment of the present invention; Figure 2 Schematic structural diagram of the working stage of a high-frequency-oriented non-intrusive household load decomposition system based on multi-source data provided by an embodiment of the present invention. Specific implementation manners
[0014] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0015] Embodiment A high-frequency-oriented non-intrusive household load decomposition system based on multi-source data, the main structure of which includes a dataset construction module, an adaptive cutting module, a multi-source analysis architecture model, and a variable convolution module; 1. Dataset construction module The function of the dataset construction module is to use the electricity consumption on the electricity meter as the main data, and use historical real meteorological data, weather forecast data, geographical information data, local economic data, and real household appliance data of real users willing to participate in the test as auxiliary data to construct an auxiliary dataset, thereby constituting a multi-source time series dataset with time as the coordinate axis. Thus, the influence of the environment on the user's electricity consumption habits is considered from the source of the dataset.
[0016] Among them, the electricity consumption on the electricity meter, historical real meteorological data, and weather forecast data are the three major time series features, and the real geographical information data, local economic data, and real household appliance data of real users willing to participate in the test are the three major non-time series features.
[0017] Among them, the electricity meter is the ammeter of the real user's home willing to participate in the test, and common electricity consumption data is collected in units of time. There are a total of 20,000 households, with a time resolution of 1 data per minute and 1 attribute value.
[0018] Among them, the historical real meteorological data is the spatio-temporal data of the ammeter where the real user's home willing to participate in the test is located, and it comes from the data of the European Meteorological Agency, including 6 types of attribute values such as wind speed, wind direction, and rainfall. The time series resolution is 1 data per 5 minutes.
[0019] Among them, the meteorological forecast data is the spatio-temporal data of the ammeter where the real users who are willing to participate in the test live. The meteorological forecast data can be purchased from China Meteorological Science and Technology Group, including 9 types of attribute values such as wind speed, wind direction, and rainfall level. It is the forecast data within 1 day, and the time series resolution is unified to 1 data every 5 minutes.
[0020] Among them, the real geographical information data is the spatio-temporal data of the families where the real users who are willing to participate in the test live, including 4 types of attribute values such as geographical coordinates (2), geographical height, and geographical type (city, rural).
[0021] Among them, the local economic data includes 2 attribute values such as the average housing price of the user's community in the previous year and the per capita disposable income of the city.
[0022] Among them, the electrical appliance data of 20,000 real user families who are willing to participate in the test includes the quantities of 11 types of common electrical appliances, including water heaters, induction cookers, microwave ovens, refrigerators, air conditioners, electric blankets, oil radiators, televisions, computers, hair dryers, and others.
[0023] In addition, there is also time data (including 4 attribute values such as year, month, day, and time).
[0024] In the learning stage, the input of the learning dataset includes ammeter data, historical real meteorological data, meteorological forecast data, time data, real geographical information data, and local economic data, which is coded as X(train), and its output is the electrical appliance data of the real user families who are willing to participate in the test, which is coded as Y(train). In this way, a trained model Model(trained) can be trained.
[0025] Among them, each piece of data in X(train) includes 20 time series data of 1 ammeter data, 6 historical real meteorological data, 9 meteorological forecast data, and 4 time data, as well as 6 non-time series data of 4 real geographical information data and 2 local economic data. And the resolution of the time data is interpolated to 1 per minute using the ARIMA model. Each time of prediction, the time length is 24 hours. Then, in the input data of each sample, the time series data is 20 * 24 * 60 = 28,800, and the non-time series data is 6, that is, 28,806 parameters. In the output data Y(train), its output is the quantities of 11 types of common electrical appliance data, that is, 11 values.
[0026] In the use phase, the input of the learning data set includes meter data, historical real meteorological data, weather forecast data, time data, real geographic information data, and local economic data, coded as X(work). It is input into the trained model Model(trained), and its output is the household appliance data of real users who are willing to participate in the test, coded as Y(predict). The data setting of X(work) is the same as that of X(train) in the learning phase.
[0027] like Figure 1 , 2 As shown in the figure, when learning the data set, the X data needs to be integrated, with rows representing channels (corresponding to parameters) and columns representing time (the column is the data sampled at 1 minute), and the data is combined in parallel into a data matrix.
[0028] 2. Adaptive cutting module Since the data faced by the present invention is time series data, the data set needs to be cut before inputting into the model. The traditional solution is a fixed cutting solution, which cuts the data into patches of fixed length. There is a problem with this cutting at present, that is, the result of the fixed-length cutting is likely to be the center of a whole cycle of high-frequency data, which cuts a cycle into multiple patches, which is very unfavorable for later learning. Therefore, in the existing research plan, there is also a length-expandable patch solution on how to cut the data. After the data is adaptively cut, it is input to the feature extraction end of the model, so that it is processed at the source of data cutting.
[0029] The basic model of the data adaptive patch cutting module of the present invention is derived from the solution in the paper "HDMixer: Hierarchical Dependency with Extendable Patch for Multivariate Time Series Forecasting" (a conference paper of AAAI-24). It designs a length-extendable patcher customized for multivariate time series, enriches the boundary information of the patches, and reduces the semantic inconsistency in concatenation. Then, the paper designs a hierarchical dependency explorer based on a multi-layer perceptron. This solution simulates the short-term dependencies within the patches, the long-term dependencies between the patches, and the complex interactions between variables. Its advantage is that in order to obtain patches with more comprehensive semantic information, the adaptive patch cutting module adaptively expands the length of the patches and uses bilinear interpolation to uniformly sample a fixed number of time steps within each patch, thus having a more complete semantic patch division strategy. Its disadvantage is that it overly relies on multi-layer information entropy to learn the window length it needs to expand, with a large computational amount. When the scenario applied by the present invention is 11 known common electrical appliance targets, through manual analysis, the high-frequency characteristics of the 11 common electrical appliance targets can be obtained, and then the computational amount of this single household can be reduced to achieve the task of adapting to more households. Therefore, it is necessary to crop the multivariate time series based on the high-frequency characteristics obtained through manual analysis to adapt to the application scenario of the present invention.
[0030] Here, the present invention proposes that considering the main cycle length of the high-frequency features is 1 - 5 minutes, when the time resolution of the data is 1 minute, the main cycle length occupies 5 time series data points. According to the Nyquist sampling theorem, the length of this cycle should be controlled within 2 times of 5 minutes, that is, 10 consecutive time series data, which means the maximum of this high-frequency cycle is 10 * 1 (10 refers to the width, and 1 is because the convolution object is 1D data). For the same reason, since the maximum window is 10, this solution does not require multi-layer convolution, so there is no need to use multi-layer information entropy for feature extraction, and directly using one layer is sufficient. Therefore, there is no need for a hierarchical dependency explorer based on a multi-layer perceptron in the original text.
[0031] Therefore, by using these data features and making such modifications, the computational amount can be greatly reduced, and the volume of the model can be reduced. The computational amount of the data adaptive cutting module of the present invention is only 2.3% of the original paper, thus greatly reducing the computational amount of this module and adapting to the scenario of the present invention facing more object tasks.
[0032] 3. Multi-source analysis architecture model The multi-source analysis architecture model is used for data fusion of various auxiliary data (geography, meteorology, economy, etc.) greatly increased in this solution.
[0033] In the multi-source analysis architecture model, in this embodiment, the ST-SAN model in the reference paper "Research on High Spatiotemporal Resolution XCO2 Concentration Estimation for Carbon Satellite Data" (Journal of Atmospheric Sciences, 2024, 47(6), 976-992) is referred to. This model mainly consists of three parts: Conv-LSTM, soft attention mechanism, and SVM-RFE module. Among them, Conv-LSTM is used to perform convolution operations on the input data to extract spatial features, and then the LSTM unit is applied to capture temporal dependencies, and the spatial structure information is retained through state transfer.
[0034] Among them, the soft attention mechanism performs reweighted aggregation calculation on the remaining information (such as the effective low-frequency and high-frequency information in the present invention) by selectively ignoring some information (such as uncertain information such as noise in the present invention). All information will be reweighted in an adaptive manner before being aggregated. This can isolate important information and avoid interference from unimportant information, thereby improving accuracy; SVM-RFE is used to effectively select the most influential features on the model from the original multi-source data, thereby improving the running efficiency and generalization ability of the model, and can automatically sort and eliminate according to the weights of the features, effectively reducing the number of features, reducing the risk of overfitting, and enhancing the model interpretability and computational efficiency.
[0035] 4. Deformable Convolution Module After the data is adaptively cut, it is input into the non-intrusive household load decomposition method. The periods of high-frequency data often show the characteristics of instability and small period range. For example, the usage time of a microwave oven is often within 1-5 minutes, an ordinary water heater is about 5 minutes, and a hair dryer is within 1-5 minutes. Therefore, due to its fixed size (commonly 1, 3, 5), the traditional convolution kernel design has a constant receptive field range when processing each region of the data, which makes its adaptability to high-frequency data targets with different period ranges insufficient, and further affects the accuracy of high-frequency data semantic segmentation. For example, in the present invention, it may not be suitable for high-frequency data with a width of 2 and a length of 4.
[0036] Here, to solve this engineering problem, this paper introduces a deformable convolution scheme and makes improvements. Different from traditional convolution, deformable convolution has the ability to dynamically adjust the sampling grid for each convolution operation. This flexibility allows for more precise feature extraction, enabling the neural network to capture complex details in the input data, that is, it can capture the dynamic and relatively small period range features of high-frequency data. By combining offsets for each sampling point, deformable convolution can adjust its receptive field to better match the unique features of the input data, accurately capture the deformation information of the target, and thus improve the performance of data analysis.
[0037] In specific operations, in deformable convolution, each element of the convolution kernel has an independent offset parameter. These parameters allow the convolution kernel to deform during the training process, enabling it to more flexibly adapt to the shape of the input features. In a neural network, the initial values of the weights are crucial for the convergence speed and performance of the model. If the weights are not properly initialized, it may lead to overly large or small activation values, thereby causing problems such as vanishing gradients or exploding gradients. By setting the weights to a uniform distribution during the initialization stage and initializing the biases to zero, it helps to optimize the network in the early stage of training. In the present invention, the Conv part in Conv-LSTM in the overall scheme ST-SAN is changed to deformable convolution, and weights are added, as shown in the following formula: ; In the formula, represents within a convolution window, is the position of the center point. p k is the k-th point within the convolution window (a total of K, for example, in a 5*5 window, excluding the center point after that, K is 24) the 1-norm distance from point p0, is the deformation offset, is the weight of the k-th point, is the scaling factor, reflecting the multiple by which the current data image is scaled or magnified.
[0038] This method enables the model to learn the offset of the sampling points and enhances the model's ability to adaptively adjust the spatial sampling position.
[0039] Here, in order to reduce the computational amount, the present invention proposes that considering the main period length of high-frequency features is 1 to 5 minutes, when the time resolution of the data is 1 minute, that is, the main period length accounts for 5 time-series data points, considering that 6 to 7 minutes are used in some cases. Therefore, the maximum window of the deformable convolution in this article is 7*1 (7 refers to the width, and 1 is because the object of convolution is 1D data).
[0040] As Figure 1 shown, in the training stage of this embodiment, electricity meter data, historical real meteorological data, weather forecast data, geographical information data, and local economic data are used as sample inputs, and user household appliance data is used as the sample output to jointly train a model of a multi-source analysis architecture.
[0041] As Figure 2 shown, in the usage stage of this embodiment, the trained multi-source analysis architecture model in Figure 1 is used. Electricity meter data, historical real meteorological data, weather forecast data, geographical information data, and local economic data are input, and then the unknown user household appliance data to be analyzed by this analysis is obtained.
[0042] In this way, by implementing this embodiment, the user can implement a non-intrusive household load analysis algorithm in an electrical data analysis algorithm. Based on the means of multi-source data fusion, this algorithm is oriented to data misalignment and the characteristics of high-frequency data, and uses variable patches and a dynamic convolution analysis module, so as to improve the analysis ability of the algorithm. It has the characteristics of high analysis accuracy for high-frequency data (household appliances with high usage frequency), and at the same time, the computational complexity is not large.
[0043] Based on the existing artificial intelligence technology and power big data technology, the present invention proposes a non-intrusive household load decomposition method, which can not only take into account the influence of the environment on the user's electricity consumption habits, but also dynamically analyze high-frequency data to prevent high-frequency data from being cut into multiple discontinuous patches due to the fixed length. Thus, it can analyze the data channel data more comprehensively, pay better attention to high-frequency information, and strengthen the analysis ability of high-frequency electrical appliance data.
[0044] Obviously, those skilled in the art can make various changes and modifications to the present invention without departing from the spirit and scope of the present invention. Thus, if these modifications and variations of the present invention fall within the scope of the claims of the present invention and its equivalent technologies, the present invention also intends to include these changes and modifications.
Claims
1. A high-frequency non-intrusive household load decomposition system based on multi-source data, characterized in that: include: Dataset construction module, adaptive cutting module, multi-source analysis architecture, variable convolution module; The data set construction module is used to construct a multi-source time series data set with time as the coordinate axis, which takes the electricity consumption on the electric meter as the main data, and takes the historical real meteorological data, weather forecast data, geographic information data, local economic data, and the household appliance data of real users willing to participate in the test as auxiliary data to construct an auxiliary data set; The adaptive cutting module is used to cut the time series data of the data set before inputting into the model; In order to obtain patches with more comprehensive semantic information, it can adaptively expand the length of the patch and use bilinear interpolation to uniformly sample a fixed number of time steps within each patch, thus achieving a more complete semantic patch partitioning strategy; The multi-source analysis architecture model is used to perform data fusion on various multi-source auxiliary data; The variable convolution module is used to dynamically adjust the sampling grid of each convolution operation; it adjusts its receptive field by calculating an offset for each sampling point so that it is better consistent with the unique characteristics of the input data, so as to accurately capture the unique information of the target, thereby improving the performance of data analysis.
2. According to claim 1, a high-frequency-oriented non-intrusive household load decomposition system based on multi-source data is characterized in that: The multi-source analysis architecture model adopts the ST-SAN architecture model as the basic framework model, which includes three parts: Conv-LSTM, soft attention mechanism, and SVM-RFE module; Conv-LSTM is used to perform convolution operations on input data to extract spatial features, and then LSTM units are applied to capture temporal dependencies, and spatial structure information is retained through state transfer; The soft attention mechanism is used to selectively ignore part of the information to perform reweighted aggregation calculation on the remaining information. All information is reweighted in an adaptive manner before being aggregated, thereby separating important information. SVM-RFE is used to select the most influential features on the model from the original multi-source data, and automatically sort and eliminate them according to the weights of the features.
3. A high-frequency non-intrusive household load decomposition system based on multi-source data according to claim 1 or 2, characterized in that: In the specific operation of the variable convolution module, the deformable convolution is that each element of the convolution kernel has an independent offset parameter; these parameters allow the convolution kernel to deform during the training process, so that it can more flexibly adapt to the shape of the input feature; By setting the weights to a uniform distribution and initializing the biases to zero during the initialization phase, it helps to optimize the network in the early stages of training.
4. The high-frequency non-intrusive household load decomposition system based on multi-source data according to claim 3 is characterized in that: The variable convolution module changes the Conv part in the Conv-LSTM in the ST-SAN module into a variable convolution and adds weights, as shown in the following formula: ; In the formula, it means in a convolution window, is the position of the center point; p k is the 1-norm distance between the kth point in the convolution window and the point p0, is the deformation offset, is the weight of the kth point, It is the scaling factor, which reflects the multiple by which the current data image is scaled or magnified.