Load Monitoring Using Machine Learning
By employing a trained CNN to analyze source location energy usage data, the method effectively predicts the energy usage of specific devices within households, addressing the challenges of NILM and enhancing energy management.
Patent Information
- Application Number
- JP2021555324
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2019-11-27
- Filing Date
- 2020-09-16
- Publication Date
- 2025-06-18
- Estimated Expiration
- 2040-09-16
AI Technical Summary
Non-intrusive load monitoring (NILM) and disaggregation of energy usage in households are challenging due to the diversity of devices and limited availability of labeled datasets, making it difficult to accurately predict the energy usage of target devices from overall source location energy usage.
The use of a trained convolutional neural network (CNN) to predict disaggregated target device energy usage data from source location energy usage data, based on labeled energy usage data from multiple source locations, allowing for the estimation of energy usage of specific devices like electrical appliances and electric vehicles.
This approach enables accurate prediction of energy usage for target devices with high granularity, improving energy management and efficiency by overcoming the limitations of existing NILM techniques.
Smart Images

Figure 0007695195000007 
Figure 0007695195000008 
Figure 0007695195000009
Abstract
Description
Technical Field
[0001] Field Embodiments of the present disclosure generally relate to utility metering devices, and more specifically to non-intrusive load monitoring using utility metering devices.
Background Art
[0002] Background Non-intrusive load monitoring ("NILM") and disaggregation of various energy-consuming devices at a specific source location are known to be very difficult. For example, disaggregating the energy usage of a household's devices and / or electric vehicles from the monitored total household energy usage has been difficult due in part to the diversity of the household's devices and / or electric vehicles (e.g., manufacturing method, model, year of manufacture, etc.). Although advancements in metering devices have provided some opportunities, it remains difficult to successfully achieve disaggregation. The limited availability of labeled datasets, or the availability of source location energy usage values including labeled device energy usage values (e.g., values of household energy usage labeled with energy usage values of electrical appliances 1, electric vehicles 1, electrical appliances 2, etc.) has further hindered progress. Therefore, NILM and disaggregation techniques that can successfully predict the energy usage of target devices from the overall energy usage of the source location by learning from these limited datasets would significantly improve this technical field and benefit users who implement these techniques.
Summary of the Invention
Means for Solving the Problems
[0003] Summary Embodiments of the present disclosure are generally directed to systems and methods for non-intrusive load monitoring using machine learning. A trained convolutional neural network (CNN) can be stored, the CNN including a plurality of layers, the CNN being trained to predict disaggregated target device energy usage data from source location energy usage data based on training data including labeled energy usage data from a plurality of source locations. Input data including energy usage data for a source location over a period of time can be received. Using the trained CNN, disaggregated target device energy usage can be predicted based on the input data.
[0004] The features and advantages of the embodiments will be described in the following description, or will be apparent from this description, or can be known by implementing the present disclosure.
[0005] Still other embodiments, details, advantages, and modifications will become apparent by considering the following detailed description of the preferred embodiments in conjunction with the accompanying drawings.
Brief Description of the Drawings
[0006]
Figure 1
Figure 2
Figure 3
Figure 4A
Figure 4B
Figure 5A
Figure 5B
Figure 5C
Figure 5D
Figure 5E
Figure 5F
Figure 5G
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
Figure 11
[0007] Detailed Description Embodiments perform non-intrusive load monitoring using a novel learning method. NILM and disaggregation take as input the total energy usage at a source location (e.g., the energy usage of a household provided by an advanced metering infrastructure), and estimate the energy usage of one or more electrical appliances, electric vehicles, and other devices that use energy at this source location. Embodiments utilize trained machine learning models to predict the energy usage of a target device based on the overall energy usage at the source location. For example, the target device may be a large electrical appliance or an electric vehicle, the source location may be a household, and the trained machine learning model may receive the energy usage of this household as input and predict the energy usage of the target device (e.g., the energy usage of the target device included in the overall energy usage of this household).
[0008] Embodiments train a machine learning model using labeled energy usage data. For example, the machine learning model may be a designed / selected neural network or the like. Energy usage data can be obtained from multiple source locations (e.g., households), and the energy usage data can be labeled with device-specific energy usage. For example, the value of household energy usage can cover a period, and the values of energy usage of individual devices (e.g., electrical appliance 1, electric vehicle 1, electrical appliance 2, etc.) within this period can be labeled. Then, in some embodiments, by processing this household and device-specific energy usage, training data for the machine learning model can be generated.
[0009] In some embodiments, by training the machine learning model, the energy usage of the target device can be predicted (e.g., disaggregated). For example, the training data may include the energy usage specific to the target device at multiple different source locations (e.g., households), and thus, by training the machine learning model, the trends in the training data can be identified and the energy usage of the target device can be predicted. In some embodiments, while predicting the energy usage of the target device by training the machine learning model, the training data may include the energy usage prediction / loss calculation / gradient update of one or more other devices. For example, when implementing embodiments of training techniques (e.g., prediction generation, loss calculation, gradient propagation, accuracy improvement, etc.) for the machine learning model, a set of other devices can be included together with the target device.
[0010] In some embodiments, a set of other devices can be based on training data and / or device-specific labeled data values available in the training data. For example, the availability of energy usage data at a source location labeled with device-specific energy usage may be limited. Embodiments include a correspondence between a set of other devices used in techniques for training a machine learning model and the values of labeled device-specific energy usage data available in the training data. In other words, the values of labeled device-specific energy usage data available in the training data can include labels for multiple different devices, there may be a number of different combinations of devices that appear within a particular source location of the training data, and the frequency with which different devices appear together at the same source location can vary. The set of other devices used in the training technique can be based on the diversity of devices within the training data, the different combinations of devices at a particular source location, and / or the frequency of appearance of different combinations of devices.
[0011] In some embodiments, when using the training data of a set of other devices in combination with the training data of a target device, the training technique can include the target device and the set of other devices. This implementation enables the trained machine learning model to more accurately predict the energy usage / disaggregation of the target device using features learned based on the set of other devices. In some embodiments, the correspondence between the set of other devices and the available training data further enhances the training / prediction / accuracy benefits achieved by including the set of other devices.
[0012] Embodiments use the total energy within a household provided by an advanced metering infrastructure (AMI) to accurately estimate or predict the corresponding device-specific energy usage. The domain of non-intrusive load monitoring (NILM) has received significant research attention, as described in the insights of Hart, George W., "Nonintrusive appliance load monitoring," Proceedings of the IEEE, vol. 80, no. 12, pp. 1870-1891, 1992. Accurate disaggregation by NILM provides numerous benefits, including energy savings opportunities, personalization, and improved power grid planning.
[0013] Embodiments utilize a deep learning approach that can accurately disaggregate the power loads of a large number of energy-consuming devices, such as large household appliances and electric vehicles, based on a limited training set. Accurate disaggregation can be challenging due to the diversity of energy-consuming devices (such as large household appliances and electric vehicles) in a typical household and their corresponding usage patterns. Additionally, in the NILM domain, the available training data may be limited. Therefore, a learning approach that can maximize the benefits of the training dataset can be particularly effective.
[0014] In embodiments, the training data can be used to train a learning model designed to effectively learn in these difficult situations. The input to the learning model can be provided from the AMI, along with other types of inputs. Embodiments can accurately predict the energy usage of electrical devices at high granularity / resolution and low granularity / resolution (e.g., 1 minute, 5 minutes, 15 minutes, 30 minutes, 1 hour, or longer).
[0015] The embodiments utilize a learning method of NILM designed for the disaggregation of target devices, but one or more other non-target devices can be used in this learning method. For example, a subset of the training dataset can include labeled device energy usage, which creates a larger training set. In the NILM domain, there may be limitations in the measured training data, and there may also be limitations in the number of labeled devices at any particular location within the dataset.
[0016] The realization of conventional NILM by existing learning methods has its own drawbacks. Some of the proposed strategies considered in the past are built on combinatorial optimization, Bayesian methods, hidden Markov models, or deep learning. However, many of these models have various drawbacks and are not useful in real-world scenarios. For example, some of these solutions have high computational costs and are thus not practical. Others often require high-resolution / granularity inputs (such as AMI data or training data) that are not available or practical with certain deployed metering capabilities.
[0017] For example, one proposed strategy focuses on multiple energy-consuming devices, but this scenario does not utilize the limited training dataset due to the constraint that multiple labeled energy-consuming devices at the same source location are required (e.g., for effective training). As a result, the training dataset is not fully utilized. Another proposed strategy selects the target device from the training dataset but does not use other devices in the learning method. In this case, the effectiveness of the system is limited because the number of devices participating in the learning is restricted. The embodiments properly solve the NILM problem within a practical time under real-world constraints.
[0018] Embodiments utilize a limited training dataset within a domain by using a target device and also using non-target devices in a learning approach. For example, learning can be performed based on non-target devices (e.g., labeled energy usage, loss calculation, and gradient propagation), and invalid entries in the training data, such as the lack of labeled data, can be replaced with zero values instead of discarded. Data curation in this form realizes an accurate disaggregation prediction of the target device by also utilizing data from non-target devices. For example, training can be performed on the curated / processed dataset. Embodiments can flexibly make accurate predictions of the disaggregation of the target device. The model can learn from other non-target devices, from a subset of the training data, or from any combination thereof during training for the target device. This flexibility enables better utilization of the training data and results in a higher level of accuracy in the disaggregation of the target device.
[0019] Embodiments can use data (e.g., training and / or input) from any suitable meter (such as AMI, etc.), and the data used (e.g., training and / or input) can have a low granularity, such as 15 minutes, 30 minutes, or 1 hour. Predicting the disaggregated energy usage of the target device can be useful for many reasons, namely, providing energy conservation opportunities to utilities and their customers, providing opportunities for personalization, and enabling better power grid planning, including peak-time demand management. For example, an electric utility can invest in technologies for aggregating energy usage from large electrical appliances or devices. The motivations for these investments include advancements in AMI and smart grid technologies, increased interest in energy efficiency, and customers' interest in more appropriate information.
[0020] Some embodiments implement an architecture on a deep learning framework that includes a convolutional neural network (“CNN”). This architecture is also extensible and can be adjusted to match the sizes of the input and output. The functions of the deep learning framework, such as layer initialization, implemented optimizer, value normalization, dropout, etc., can be utilized, removed, or adjusted. In fact, many applications of CNNs are designed to recognize visual patterns (e.g., directly from images for classification). Examples thereof include LeNet, AlexNet, ZFNet, GoogleNet / Inception, VGGNet, and ResNet. On the other hand, embodiments use the CNN architecture to predict disaggregation of the energy usage of a target device. For example, the CNN can be designed to have multiple convolutional layers that are run in parallel and have various kernel sizes and shapes. This design can be used to learn trends and other aspects of energy usage data measured (e.g., over granularities of 1 minute, 5 minutes, 15 minutes, 30 minutes, 1 hour, or longer).
[0021] Some embodiments utilize multiple trained learning models to achieve higher prediction accuracy. For example, an ensemble strategy can combine the outputs from multiple trained models. Embodiments implementing this ensemble strategy can achieve higher accuracy by combining multiple deep learning models designed to solve disaggregation and detection / identification problems. For example, the optimal accuracy can be achieved by combining the outputs of these models in multiple possible ways. Embodiments solve two different but related problems regarding the same input, disaggregation, and detection / identification. Separate models can be used to more effectively solve each problem. The results obtained from the disaggregation and detection / identification models can be combined in several ways, namely: a) weighting for detection / identification, b) weighting for disaggregation, c) equal (substantially equal) weighting for each model. For example, a particular way of combining two models into a final output can be based on multiple factors, namely, a threshold, and the distance between the predicted output of each model and the final output.
[0022] Embodiments train and construct disaggregation and detection / identification models (e.g., for each target device). Depending on the mutual distance between each model (measured, e.g., by a distance metric) and the distance of each model from the labeled / known values, an ensemble / combination strategy can be selected, namely: a) weighting for detection / identification, b) weighting for disaggregation, c) equally (or substantially equally) weighting each model. Some embodiments may augment the output using data values (e.g., a threshold), for example, based on the disagreement between models. The implementation forms and results show improved disaggregation predictions for multiple energy-consuming devices (e.g., large household appliances and / or electric vehicles) when the models are combined into a final output.
[0023] Next, refer in detail to the embodiments of the present disclosure in which the examples are shown in the accompanying drawings. In the following detailed description, numerous specific details are set forth in order to provide a thorough understanding of the present disclosure. However, it will be apparent to those skilled in the art that the present disclosure may be practiced without these specific details. In other instances, well-known methods, procedures, components, and circuits have not been described in detail so as not to unnecessarily obscure aspects of the embodiments. As much as possible, like reference numerals are used for like elements.
[0024] Figure 1 shows a system for disaggregating energy usage associated with a target device according to an example of an embodiment. System 100 includes a source location 102, a meter 104, a source location 106, a meter 108, devices 110, 112, and 114, and a network node 116. Source location 102 can be any suitable location that includes, or is otherwise associated with, a device that consumes or generates energy, such as a household that includes devices 110, 112, and 114. In some embodiments, devices 110, 112, and 114 can be electrical appliances that use energy and / or electric vehicles, such as washing machines, dryers, air conditioners, heaters, refrigerators, televisions, computing devices, and the like. For example, source location 102 can receive a supply of power (e.g., electricity), and devices 110, 112, and 114 can obtain the power supplied from source location 102. In some embodiments, source location 102 is a household, and the power to the household is supplied from a power grid, a local power source (e.g., solar panels), a combination thereof, or any other suitable source.
[0025] In some embodiments, the meter 104 can be used to monitor the energy usage (e.g., electricity usage) at the source location 102. For example, the meter 104 can be a smart meter, an advanced metering infrastructure ( "AMI") meter, an automatic meter reading ( "AMR") meter, a simple energy usage meter, or the like. In some embodiments, the meter 104 can transmit information regarding the energy usage at the source location 102 to the central power grid, a provider, a third party, or any other suitable entity. For example, the meter 104 can enable two-way communication with an entity to convey the energy usage at the source location 102. In some embodiments, the meter 104 can enable one-way communication with an entity, in which case the measured reading of the meter is transmitted to the entity.
[0026] In some embodiments, the meter 104 can communicate via a wired communication link and / or a wireless communication link, and can utilize wireless communication protocols (e.g., cellular technology), Wi-Fi®, wireless ad hoc networks via Wi-Fi, wireless mesh networks, low-power wide-area wireless ( "LoRa"), Zigbee®, Wi-SUN, wireless local area networks, wired local area networks, and the like. Devices 110, 112, and 114 (and other devices not shown) can use energy at the source location 102, and the meter 104 can monitor the energy usage at the source location and report corresponding data (e.g., to network node 116).
[0027] In some embodiments, source location 106 and meter 108 can be the same as source location 102 and meter 104. For example, networking node 116 can receive energy usage information regarding source location 102 and source location 106 from meter 104 and meter 106. In some embodiments, networking node 116 can be part of a central power grid, a provider, a transmission network, an analytics service provider, a third - party entity, or any other suitable entity.
[0028] The following description includes the description of one criterion or multiple criteria. These terms are used interchangeably throughout the present disclosure, and the scope of multiple criteria is intended to include the scope of one criterion, and the scope of one criterion is intended to include the scope of multiple criteria.
[0029] FIG. 2 is a block diagram of a computer server / system 200 according to an embodiment. All or part of system 200 may be used to implement any of the elements shown in FIG. 1. As shown in FIG. 2, system 200 may include a bus device 212 and / or other communication mechanisms configured to transfer information among the various components of system 200, such as processor 222 and memory 214. Additionally, communication device 220 may enable connectivity between processor 222 and other devices by encoding data to be transmitted from processor 222 to another device via a network (not shown) and decoding data received from another system via the network for processor 222.
[0030] For example, communication device 220 may include a network interface card configured to provide wireless network communication. Various wireless communication technologies can be used, including infrared, wireless, Bluetooth®, Wi-Fi, and / or cellular communication. Alternatively, communication device 220 may be configured to provide a wired network connection such as an Ethernet® connection.
[0031] Processor 222 may include one or more general-purpose or dedicated processors for performing the computing and control functions of system 200. Processor 222 may include a single integrated circuit such as a microprocessing device, or may include multiple integrated circuit devices and / or circuit boards that operate in cooperation to implement the functions of processor 222. Additionally, processor 222 may execute computer programs such as operating system 215, prediction tool 216, and other applications 218 stored in memory 214.
[0032] System 200 may include a memory 214 for storing instructions and information executed by a processor 222. The memory 214 may include various components for retrieving, presenting, modifying, and storing data. For example, the memory 214 may store software modules that, when executed by the processor 222, provide functionality. The modules may include an operating system 215 that provides operating system functionality to the system 200. The modules may include an operating system 215, a prediction tool 216 that implements the NILM and disaggregation functions disclosed herein, and other application modules 218. The operating system 215 provides operating system functionality to the system 200. In some examples, the prediction tool 216 may be implemented as an in-memory configuration. In some implementations, when the system 200 executes the functions of the prediction tool 216, it implements a non-conventional dedicated computer system that executes the functions disclosed herein.
[0033] The non-transitory memory 214 may include various computer-readable media accessible by the processor 222. For example, the memory 214 may include any combination of random access memory ("RAM"), dynamic RAM ("DRAM"), static RAM ("SRAM"), read-only memory ("ROM"), flash memory, cache memory, and / or any other type of non-transitory computer-readable media. The processor 222 is further coupled via a bus 212 to a display 224, such as a liquid crystal display ("LCD"). A keyboard 226 and a cursor control device 228, such as a computer mouse, are further coupled to the communication device 212 to enable a user to interface with the system 200.
[0034] In some embodiments, system 200 can be part of a larger system. Thus, system 200 can include one or more additional functional modules 218 to include additional functionality. Other application modules 218 can include various modules of, for example, Oracle® utility customer cloud services, Oracle® cloud infrastructure, Oracle® cloud platform, Oracle® cloud applications. Prediction tool 216, other application modules 218, and any other suitable components of system 200 can include various modules of Oracle® data science cloud services, Oracle® data integration services, or other suitable Oracle® products or services.
[0035] Database 217 is coupled to bus 212 to provide centralized storage for modules 216 and 218 and store data received, for example, by computer vision tool 216 or other data sources. Database 217 can store data in an integrated collection of logically related records or files. Database 217 can be an operational database, an analytical database, a data warehouse, a distributed database, an end-user database, an external database, a navigation database, an in-memory database, a document-oriented database, a real-time database, a relational database, an object-oriented database, a non-relational database, a NoSQL database, a Hadoop® distributed file system (“HFDS”), or any other database well known in the art.
[0036] Although shown as a single system, the functions of system 200 may be implemented as a distributed system. For example, memory 214 and processor 222 may be distributed among a plurality of different computers collectively representing system 200. In one embodiment, system 200 may be part of a device (such as a smartphone, tablet, computer, etc.). In certain embodiments, system 200 may be independent of the device and may provide the disclosed functions remotely to this device. Further, one or more components of system 200 may be omitted. For example, in the case of functions as a user or consumer device, system 200 may be a smartphone or other wireless device, and this smartphone or other wireless device includes a processor, a memory, and a display, does not include one or more of the other components shown in FIG. 2, and includes additional components not shown in FIG. 2 such as components of an antenna, a transceiver, or any other suitable wireless device.
[0037] FIG. 3 shows a system for disaggregating energy usage associated with a target device using a machine learning model, according to an example of an embodiment. System 300 includes input data 302, a processing module 304, a prediction module 306, training data 308, and output data 310. In some embodiments, input data 302 may include energy usage from a source location, and processing module 304 can process this data. For example, processing module 304 can generate features based on the input data by processing input data 302.
[0038] In some embodiments, the prediction module 306 can be a machine learning module (e.g., a neural network) trained by the training data 308. For example, the training data 308 can include labeled data such as energy usage data values from a plurality of source locations (e.g., source locations 102 and 106 in FIG. 1), including labeled device-specific energy usage data values. In some embodiments, the output from the processing module 304, such as the processed input, can be provided to the prediction module 306 as an input. The prediction model 306 can generate output data 310, such as disaggregated energy usage data, for the input data 302. In some embodiments, the input data 302 can be source location energy usage data, and the output data 310 can be disaggregated energy usage data for the target device (or devices).
[0039] Embodiments use a machine learning model, such as a neural network, to predict the energy usage of a target device. A neural network can include a plurality of nodes called neurons, which are connected to other neurons via links or synapses. Some implementations of neural networks can be for classification tasks and / or can be trained under supervised learning techniques. Often, the labeled data can include features useful for the realization of a prediction task (e.g., classification / prediction of energy usage). In some embodiments, neurons within a trained neural network can perform small mathematical operations on specific input data, in which case, using their corresponding weights (or relevance), they can generate operands (e.g., partially generated by applying non-linearity) that are further sent within the network or provided as an output. Synapses can connect two neurons with corresponding weights / relevance. The prediction model 306 of FIG. 3 can be a neural network.
[0040] In some embodiments, a neural network can be used to learn trends in labeled energy usage data values (e.g., household energy usage data values labeled with device-specific energy usage over a period of time). For example, the training data can include features, and a neural network (or other learning model) can use these features to identify trends and predict energy usage (e.g., identify energy usage by a target device by disaggregating the overall energy usage of a household) from the overall source location energy usage associated with the target device. In some embodiments, once the model is trained / prepared, it can be deployed. Embodiments can be implemented using multiple products or services (e.g., products or services of Oracle®).
[0041] In some embodiments, the design of the prediction module 306 can include components of any suitable machine learning model (e.g., neural network, support vector machine, specialized regression model, etc.). For example, a neural network can be implemented with a specific cost function (e.g., for training / gradient calculation). A neural network can include any number (e.g., 0, 1, 2, 3, or more) of hidden layers and can include feedforward neural networks, recurrent neural networks, convolutional neural networks, modular neural networks, and any other suitable types.
[0042] Figures 4A - 4B show a convolutional neural network according to an example of an embodiment. CNN 400 in Figure 4A includes layers 402, 404, 406, 408, and 410, as well as kernels 412, 414, and 416. For example, in a particular layer of a convolutional neural network, one or more filters or kernels can be applied to the input data of the layer. In the illustrated embodiment, layers 402, 404, and 406 are convolutional layers, kernel 412 is applied to layer 402, kernel 414 is applied to layer 404, and kernel 416 is applied to layer 406. The shape of the data and the underlying data values can be changed from input to output depending on the shape of the filter or kernel being applied (e.g., 1×1, 1×2, 1×3, 1×4, etc.), the method of applying the filter or kernel (e.g., mathematical application), and other parameters (e.g., stride). Kernels 412, 414, and 416 are shown as 1 - dimensional kernels, but any other suitable shape can be realized. In an embodiment, kernels 412, 414, and 416 can have one shape that matches among them, two different shapes, or three different shapes (e.g., all kernels are of different sizes).
[0043] In some examples, the layers of a convolutional neural network can be heterogeneous and can include various mixes / sequences such as convolutional layers, pooling layers, fully - connected layers (similar to the application of a 1×1 filter), etc. In the illustrated embodiment, layers 408 and 410 can be fully - connected layers. Thus, CNN 400 shows an embodiment of a feed - forward convolutional neural network where fully - connected layers follow a plurality of convolutional layers (realizing, for example, 1 - dimensional filters or kernels). The embodiment can realize any other suitable convolutional neural network.
[0044] The CNN 420 in FIG. 4B includes layers 422, 424, 426, 428, 430, and 432, as well as kernels 434, 436, and 438. The CNN 420 may be similar to the CNN 400 in FIG. 4A. However, layers 422, 424, and 426 may be convolutional layers in a parallel orientation, and layer 428 may be a concatenation layer that concatenates the outputs of layers 422, 424, and 426. For example, the input from the input layer can be provided to each of layers 422, 424, and 426, and the outputs from these layers are concatenated at layer 428. In some embodiments, the concatenated output from layer 428 can be provided to layer 430, which may be a fully connected layer. For example, each of layers 430 and 432 may be a fully connected layer, and the output from layer 432 can be used as the prediction generated by the CNN 420.
[0045] In some embodiments, kernels 434, 436, and 438 can be the same as kernels 412, 414, and 416 in FIG. 4A. For example, although kernels 434, 436, and 438 are shown as 1 - dimensional kernels, they can implement any other suitable shape. In an embodiment, kernels 434, 436, and 438 can have one shape that matches between them, two different shapes, or three different shapes (e.g., all kernels have different sizes).
[0046] In some examples, the layers of a convolutional neural network may be heterogeneous and can include various mixes / sequences such as convolutional layers, pooling layers, fully connected layers (similar to the application of 1×1 filters), parallel layers, concatenation layers, etc. For example, layers 422, 424, and 426 can represent three parallel layers, although more or fewer parallel layers can be implemented. Similarly, the output of each of layers 422, 424, and 426 is shown as an input to layer 428, which is a concatenation layer in some embodiments, although one or more of layers 422, 424, and 426 can include additional convolutional or other layers before the concatenation layer. For example, one or more convolutional or other layers may be present between layer 422 (e.g., a convolutional layer) and layer 428 (e.g., a concatenation layer). In some embodiments, another convolutional layer (having a different kernel) can be implemented between layer 422 and layer 428, although such an intervening layer is not implemented for layer 424. In other words, in this example, the input to layer 422 can pass through another convolutional layer before being input to layer 428 (e.g., a concatenation layer), while the input to layer 424 is output directly to layer 428 (without another convolutional layer).
[0047] In some embodiments, layers 422, 424, 426, and 428 (e.g., three parallel convolutional layers and one concatenation layer) can represent a block within CNN420, and one or more additional blocks can be implemented before or after the shown block. For example, a block can be characterized by at least two parallel convolutional layers followed by a concatenation layer. In some embodiments, multiple (e.g., three or more) additional convolutional layers having various parallel structures can be implemented as blocks. CNN420 shows an embodiment of a feedforward convolutional neural network where a fully connected layer follows a plurality of convolutional layers in a parallel orientation (e.g., implementing 1 - dimensional filters or kernels). Embodiments can implement any other suitable convolutional neural network.
[0048] In some embodiments, the neural network can be configured for deep learning, for example, based on the number of hidden layers realized. In some examples, a Bayesian network can be similarly realized, or other types of supervised learning models can be similarly realized. For example, a support vector machine can be realized with one or more kernels (such as a Gaussian kernel, a linear kernel, etc.) in some examples. In some embodiments, the prediction module 306 of FIG. 3 can be, for example, a plurality of stacked models where the output of the first model is provided as the input to the second model. Some implementations may include multiple layers of the prediction model.
[0049] In some embodiments, a test instance can be given to the model to calculate its accuracy. For example, a portion of the training data 308 / labeled energy usage data can be reserved (for example, not for training the model) to test the trained model. By using the accuracy measurement value, the prediction module 306 can be adjusted. In some embodiments, the accuracy evaluation can be based on a subset of the training data / processed data. For example, by using a subset of the data, the accuracy of the trained model can be evaluated (for example, the training - to - test ratio is 75% to 25%, etc.). In some embodiments, data can be randomly selected for the test and training segments over different repetitions of the test.
[0050] In some embodiments, during testing, the trained model can output a predicted data value of the energy usage of the target device based on a specific input (e.g., an instance of test data). For example, an instance of test data can be energy usage data for a specific source location (e.g., a household) over a period of time that includes the unique energy usage of the target device over that period as a known labeled data value. Since the energy usage data value of the target device is known for a specific input / test instance, an accuracy metric can be generated by comparing the predicted value to the known value. Based on testing of the trained model using multiple instances of test data, the accuracy of the trained model can be evaluated.
[0051] In some embodiments, the design of the prediction module 306 can be adjusted based on training, retraining, and / or accuracy calculations during updated training. For example, the adjustment can include adjusting the number of hidden layers within a neural network, adjusting kernel calculations (e.g., those used to implement a support vector machine or neural network), and the like. This adjustment can also include adjustment / selection of features used by the machine learning model, adjustment to the processing of input data, and the like. Embodiments include implementing various adjustment configurations (e.g., different versions of machine learning models and features) while training / calculating accuracy to reach a configuration for the prediction module 306 that achieves a desired performance (e.g., performing predictions at a desired accuracy level, performing according to a desired resource utilization / time metric, etc.) when trained. In some embodiments, the trained model can be saved or stored for further use and to maintain its state. For example, the training of the prediction module 306 can be performed "offline," in which case the trained model can be stored and used as needed to achieve time- and resource-efficient data predictions.
[0052] Embodiments of the prediction module 306 are trained to disaggregate energy usage data within overall source location (e.g., household) energy usage data. NILM and / or disaggregation takes as input the total energy usage at a source location (e.g., the energy usage of a household provided by an advanced metering infrastructure) and estimates the energy usage of one or more electrical appliances, electric vehicles, and other devices that use energy at this source location. FIGS. 5A-5G show sample graphs depicting disaggregated energy usage data according to an example embodiment. The data shown in the sample graphs are tested embodiments of the disaggregation techniques disclosed herein, e.g., representing a trained machine learning model that disaggregates energy usage at an unknown source location (e.g., a household).
[0053] FIG. 5A shows a graph representing total energy usage data, labeled energy usage data of a target device, i.e., an air conditioner, and predicted energy usage data of this target device (e.g., predicted by a trained embodiment of the prediction module 306). In this graph, time is shown on the x-axis and energy usage (in kWh) is shown on the y-axis. The data shown includes a one-hour granularity of measured data (e.g., total energy usage data and labeled energy usage data of the target device). Any other suitable granularity can be similarly achieved.
[0054] Referring to FIG. 5A, the comparison between the labeled energy usage data values (actual / measured data values) of the target device and the predicted energy usage data values of the target device indicates the accuracy of the trained prediction model. For example, the trained prediction model can receive the total energy usage data values (or processed ones) as input and generate a prediction represented graphically. In some embodiments, the total energy usage data values may include the energy usage by the target device and a plurality of other devices. As shown in FIG. 5A, the predicted, disaggregated energy usage values of the target device achieve a high degree of accuracy over multiple days. Other suitable data granularity, period, or other suitable parameters can be achieved.
[0055] Figures 5B to 5G show graphs of total energy usage data, labeled energy usage data of the target device, and predicted energy usage data of the target device. For example, the graphs of Figures 5B to 5G can be the same as the graph of Figure 5A except for the target device. The graph of Figure 5B shows the predicted disaggregation for the target device that is a furnace (for example, an electric furnace). The graph of Figure 5C shows the predicted disaggregation for the target device that is a dryer (for example, an electrical appliance). The graph of Figure 5D shows the predicted disaggregation for the target device that is a pool pump. The graph of Figure 5E shows the predicted disaggregation for the target device that is a water heater. The graph of Figure 5F shows the predicted disaggregation for the target device that is a refrigerator. The graph of Figure 5G shows the predicted disaggregation for the target device that is an electric vehicle. The predictions shown in the graphs of Figures 5A to 5G can utilize the source location energy usage from a household that includes a plurality of electrical appliances / devices that consume energy in addition to the corresponding target device. The input (for example, the input to the learning module 306) used to generate the predicted disaggregation can include energy usage data that has not been used for training. In other words, the trained learning module 306 generates predictions for input data that it has not seen before.
[0056] In some embodiments, the input 302 and / or the training data 308 can include information other than energy usage information. For example, weather information related to the energy usage data (such as the weather like precipitation, temperature, etc. at the time of energy usage measurement), calendar information related to the energy usage data (such as calendar information like month, date, day of the week, etc. at the time of energy usage measurement), time stamps related to the energy usage data, and other related information can be included in the input 302 and / or the training data 308.
[0057] The embodiments generate training data 308 for use in training the prediction module 306 by processing energy usage data from a source location (e.g., a household). For example, the overall source location energy usage data values shown in FIGS. 5A - 5G can be combined with the labeled energy usage data values of one or more devices, and by processing this resulting combination, training data 308 can be reached. In some embodiments, the energy usage data of the source location can be obtained through measurement (e.g., metering). Additionally, by implementing measurements, metering, or some other technique for receiving / monitoring the energy usage of specific devices within the source location, device - specific labeled energy usage data for training can be generated. In other examples, energy usage data including source location energy usage and disaggregated device - specific energy within the source location can be obtained from a third party. For example, the training data can be obtained in any suitable way, such as monitoring a source location (e.g., a household) in known situations, obtaining publicly (or otherwise) available datasets, developing joint ventures or partnerships that result in training data, and using any other suitable means.
[0058] An example of energy usage data that can be processed to generate training data 308 includes the following.
[0059] [Table 1]
[0060] The columns included in the sample rows of data in this example are an identifier, a timestamp, a total (energy usage), and labeled appliance-specific energy usage (such as air conditioners, electric vehicles, washing machines, dryers, dishwashers, refrigerators, etc.). This example includes a granularity of 15 minutes, although other suitable granularities can be achieved as well (e.g., 1 minute, 5 minutes, 15 minutes, 30 minutes, 1 hour, several hours, etc.). In some embodiments, processing the energy usage data (e.g., to generate training data 308) may include reducing the granularity of the data such that, for example, a training corpus can be generated with a consistent granularity (e.g., 1 hour). Such a reduction in granularity can be achieved by summing the data usage values across the elements that make up a unit of time (e.g., summing the data usage values across four 15-minute intervals that make up 1 hour).
[0061] Embodiments include a determined set of devices to be included within training data 308. For example, training of the prediction module 306 can be configured to generate disaggregation predictions for target devices, although the training can utilize labeled data usage for a set of other devices in addition to the target device. In some embodiments, the set of other devices can be based on energy usage data and / or device-specific labeled data values available for training. Training data is often limited, and thus training techniques that make use of available training data are often beneficial.
[0062] An embodiment includes a correspondence between a set of other devices used in a technique for training a machine learning model and values of labeled device-specific energy usage data available in training data. In other words, the values of the labeled device-specific energy usage data available in the training data can include labels of a plurality of different devices, there may be a number of different combinations of devices that appear within a particular source location of the training data, and the frequency with which different devices appear together at the same source location can vary. The set of other devices used in the training technique can be based on the diversity of devices in the training data, different combinations of devices at a particular source location, and / or the frequency of occurrence of different combinations of devices.
[0063] Accordingly, a plurality of different variations of the training data 308 can be generated by processing the preprocessed energy usage data of Table 1 above. Table 2 shows an example of generating the training data 308 that includes a timestamp, a total source location energy usage, and a labeled energy usage of an electric vehicle as a single target device by processing the energy usage data.
[0064]
Table 2
[0065] In some embodiments, the preprocessing can include selecting from Table 1 a subset of columns, a subset of rows, an aggregation of data (or some other mathematical / combinatorial function), and other appropriate processing. For example, data cleaning, normalization, scaling, or other processing used to render data suitable for machine learning can be performed.
[0066] Table 3 represents an example of generating training data 308 that includes a timestamp, the total source location energy usage, the labeled energy usage of the target device (electric vehicle), and the labeled energy usage of other devices (air conditioner) by processing energy usage data.
[0067]
Table 3
[0068] Table 4 shows an example of generating training data 308 that includes a timestamp, the total source location energy usage, the labeled energy usage of the target device (electric vehicle), and the labeled energy usage of three other devices (air conditioner, washing machine, and dryer) by processing energy usage data.
[0069]
Table 4
[0070] In this example shown, the devices with null values in Table 1 are included in Table 4, but the null values have been replaced with zero values. The null values for the washing machine and dryer in Table 1 indicate that the source location energy usage data in Table 1 does not include the labeled energy usage data for the labeled washing machine and dryer. In other words, for example, due to limitations in the training corpus, measurement / metrology devices, or other situations, the labeled energy usage data for the washing machine and dryer within the overall source location energy usage data (shown in Table 1) cannot be utilized.
[0071] In some embodiments, processing the device-specific labeled energy usage data at the source location may include replacing null values (or any other placeholder value) with zero values. For example, if it is determined to use a particular device in the training technique of a given implementation of the prediction module 306, and it is determined that a portion of the energy usage data is missing the labeled device-specific energy usage of the particular device (at one or more source locations), the value of the labeled energy usage of this particular missing device can be replaced with a zero value. As further described herein, it can be determined that a device participates in the training technique of a given implementation, even if the training corpus does not include a comprehensive set of the labeled energy usage data for these devices. Embodiments replace null values with zero values in order to utilize available training data, leverage one or more devices other than the target device for learning purposes, and comprehensively improve machine learning performance.
[0072] Table 5 shows an example of generating training data 308 that includes a timestamp, the total source location energy usage, the labeled energy usage of the target device (electric vehicle), and the labeled energy usage of four other devices (air conditioner, washing machine, dryer, and refrigerator) by processing the energy usage data.
[0073] [Table 5]
[0074] In the example shown for the source locations in Table 5, the target device is accompanied by two devices with labeled energy usage data (e.g., an air conditioner and a refrigerator) and two devices without labeled energy usage data (e.g., a washing machine and a dryer), where in this case the two devices without labeled energy usage data are processed to reflect zero energy usage data labels. Table 6 shows an example of generating training data 308 that includes a timestamp, the total source location energy usage, the labeled energy usage of the target device (an electric vehicle), and the labeled energy usage of five other devices (an air conditioner, a washing machine, a dryer, a dishwasher, and a refrigerator) by processing the energy usage data.
[0075]
Table 6
[0076] In the example shown for the source locations in Table 6, the target device is accompanied by two devices with labeled energy usage data (e.g., an air conditioner and a refrigerator) and three devices without labeled energy usage data (e.g., a dishwasher, a washing machine, and a dryer), where in this case the three devices without labeled energy usage data are processed to reflect zero energy usage data labels. As demonstrated by Tables 2 - 6, different variations of training data can be generated by processing the energy usage data of a particular source location that includes the labeled energy usage data of some devices. Embodiments can utilize one or more of these variations to train a machine learning model and achieve beneficial results.
[0077] For example, a given implementation for training module 306 may include several factors. While one implementation may aim to disaggregate the energy usage of a single target device, several other factors related to the available training data may be problematic, which include the availability of overall energy usage data including labeled energy usage data of the target device at different source locations, the number and diversity of other devices having available labeled energy usage data arranged with the target device at different source locations, the granularity of the available energy usage data, as well as other related factors and the like. Thus, an implementation that achieves the desired prediction results may include the use of labeled energy usage data of the target device, labeled energy usage data of a set of other devices, and energy usage data set to zero for a particular device among the set of other devices in a portion of the training data. Therefore, a particular transformation of the processed training data shown in Tables 2 to 6 that achieves the desired results is based on the available training data and its related factors. In other words, the set of devices to be used within the training data can be determined based on, for example, available energy usage data having labeled device-specific energy usage values and the like.
[0078] In some embodiments, a set of other devices participating in training can meet certain criteria regarding the available training data. For example, the available training data can include energy usage values at a plurality of source locations (e.g., households), and the energy usage of most of the source locations can include the labeled energy usage from the target device and the labeled energy usage from at least one of the set of other devices. In another example, a set of other devices may be determined such that at least a threshold number (e.g., a minimum number) of other devices are used in the training technique. In some embodiments, a set of other devices may be determined such that any given instance of the training data (e.g., a row of the training data) includes a threshold number or fewer other devices whose energy usage data value is set to zero (e.g., 0 or less, 1 or less, 2 or less, 3 or less, etc.). In some embodiments, a set of other devices may be determined to be a null set, for example, based on the limitations indicated by the training data.
[0079] In some embodiments, a set of other devices is determined such that the amount of training data used to train a machine learning model based on this set of other devices meets a criterion. For example, a set of other devices may be determined such that at least a threshold number (e.g., a minimum number) of training instances (e.g., rows of the training data) are useful for training. In another example, a set of other devices may be determined such that the sum of the labeled device-specific energy usage data values set to zero meets a criterion (e.g., less than a threshold percentage of the training data, e.g., less than 5%, 10%, 15%, 20%, 30%, 40%, 50%, etc.).
[0080] Embodiments may also include training and implementation techniques that realize other correspondence relationships between the training data and a set of other devices. For example, most of the instances of the training data (e.g., rows of the training data) may include energy usage values labeled with added zeros (based on a set of other devices that do not exist in the household or dataset). In another example, most of the instances of the training data may include non-zero labeled energy usage values of the target device. In another example, most of the instances of the training data may include non-zero labeled energy usage values of at least one of the target device and a set of other devices. In another example, at least some of the instances of the training data may include energy usage values labeled with added zeros of the target device. In another example, each of the set of other devices may have non-zero labeled energy usage values at least at a threshold amount (e.g., 10%, 20%, 30%, etc.) of the instances of the training data. In another example, one or more instances of the training data may include non-zero labeled energy usage values of at least two of the target device and a set of other devices. By implementing the training techniques and a set of other devices, any of these correspondence relationships, most of them, one of them, or a combination can be obtained, or any other suitable correspondence relationship can be obtained.
[0081] In some embodiments, when using the training data of a set of other devices in combination with the training data of the target device, training techniques (such as prediction generation, loss calculation, gradient propagation, accuracy improvement, etc.) can be implemented using the target device and the set of other devices. For example, by training a machine learning model using the post-processed training data, the energy usage of the target device can be predicted from the energy usage of the source location (e.g., household). In some embodiments, by receiving, processing, and providing input data such as energy usage data of the source location over a period of time to the trained machine learning model, a prediction can be generated as to how much of the energy used at the source location was used by the target device.
[0082] In some embodiments, when training a machine learning model using a determined set of other devices, the predictions generated by the trained model can include a disaggregation prediction for the target device and predictions generated for a set of other devices (e.g., non-target devices). For example, the predictions generated for non-target devices can be useful when calculating accuracy metrics (e.g., based on the labeled device-specific energy usage of the set of other devices) and training the model. In some embodiments that focus on the disaggregation prediction for the target device, the predictions generated for the set of other devices can be discarded. For example, these embodiments utilize the advantages of training using other devices, and these advantages improve the disaggregation prediction for the target device.
[0083] In some embodiments, by training a convolutional neural network using a processed version of the training data, the energy usage of the target device can be disaggregated. Referring to FIG. 4A, CNN 400 includes layers 402, 404, and 406 that can be convolutional layers and layers 408 and 410 that can be fully connected layers. Kernel 412, shown as having a 1×a shape, can be applied to layer 402, kernel 414, shown as having a 1×b shape, can be applied to layer 404, and kernel 416, shown as having a 1×c shape, can be applied to layer 406. In some embodiments, the shapes of kernels 412, 414, and 416 can be adjusted during the training / configuration of CNN 400, so they can be any suitable shapes that achieve effective performance of the disaggregation task.
[0084] Similarly, referring to FIG. 4B, CNN 420 includes layers 422, 424, and 426 that can be convolutional layers, layer 428 that can be a concatenation layer, and layers 430 and 432 that can be fully connected layers. Kernel 434, shown as having a 1×a shape, can be applied to layer 422, kernel 436, shown as having a 1×b shape, can be applied to layer 434, and kernel 438, shown as having a 1×c shape, can be applied to layer 426. In some embodiments, the shapes of kernels 434, 436, and 438 can be adjusted during the training / configuration of CNN 420, so they can be any suitable shapes that achieve effective performance of the disaggregation task.
[0085] In some embodiments, the shapes of kernels 412, 414, and 416 can each change the shape of the data as this data passes through layers 402, 404, and 406. For example, the shape of the input data can be changed by applying kernels 412, 414, and 416 as the input data passes through layers 402, 404, and 406. In some embodiments, the shape of the data passing through layers 402, 404, and 406 is based on the shapes of kernels 412, 414, and 416, the stride of each kernel, and the padding that is implemented. For example, the padding can include adding 0s to the data (e.g., above, below, to the left, and / or to the right of the data). In some embodiments, for one or more of layers 412, 414, and 416, a convolution of the same / original size that does not change the shape of the data can be obtained by a combination of kernel shape, stride, and padding.
[0086] Similarly, the shapes of kernels 434, 436, and 438 can each change the shape of the data as this data passes through layers 422, 424, and 426. For example, the shape of the input data can be changed by applying kernels 434, 436, and 438 as the input data passes through layers 422, 424, and 426. In some embodiments, the shape of the data passing through layers 422, 424, and 426 is based on the shapes of kernels 434, 436, and 438, the stride of each kernel, and the padding that is implemented. In some embodiments, for one or more of layers 434, 436, and 438, a convolution of the same / original size that does not change the shape of the data can be obtained by a combination of kernel shape, stride, and padding.
[0087] As described above, embodiments predict the energy usage disaggregation of a target device, but a set of other devices (e.g., non-target devices) can also participate in the learning techniques (e.g., prediction, loss calculation, gradient propagation). Embodiments that implement a CNN can use the training data, loss calculation, and gradient propagation of non-target devices to configure the weights / values of the kernels implemented in different layers. This training / configuration of the CNN results in neurons trained based on non-target devices that can be effective in improving the accuracy of predictions for the target device.
[0088] In some embodiments, multiple machine learning models can be trained, and by combining the outputs of these models, a predicted energy usage disaggregation of the target device can be realized. FIG. 6 shows a flowchart for disaggregating the energy usage associated with a target device by using multiple machine learning models according to an example of an embodiment.
[0089] System 600 includes input data 602, a processing module 604, prediction modules 606 and 610, training data 608 and 612, a combination module 614, and an output 616. In some embodiments, the input data 602 can include energy usage from a source location, and the processing module 604 can process the data. For example, the processing module 604 can process the input data 602 and generate features based on the input data. In some embodiments, the input data 602 and the processing module 604 can be similar to the input data 302 and the processing module 304 of FIG. 3.
[0090] In some embodiments, prediction modules 606 and 610 can be machine learning modules (e.g., neural networks) trained by training data 608 and 612, respectively. For example, training data 608 can include labeled data such as energy usage data values from a plurality of source locations (e.g., source locations 102 and 106 in FIG. 1), including labeled device-specific energy usage data values. In some embodiments, the output from processing module 604, such as processed input, can be provided as input to prediction modules 606 and 610. In some embodiments, prediction modules 606 and 610 can be similar to prediction module 306 in FIG. 3.
[0091] In some embodiments, training data 608 can predict the disaggregated energy usage of the target device by training prediction module 606, and training data 612 can predict the energy usage of the target device above a threshold by training prediction module 610. For example, training data 608 that generates a disaggregation prediction by training prediction module 606 can include labeled energy usage data having the total energy used (e.g., over a certain time span). In other words, the labeled device-specific energy usage data of training data 608 reflects the total energy used, such as that shown in Tables 1-6 above.
[0092] In some embodiments, the training data 612 that generates the detection prediction by training the prediction module 610 may include the detected energy usage (e.g., over a certain time span). In other words, the labeled device-specific energy data in the training data 612 can represent whether energy exceeding a threshold was used (e.g., a binary value representing on or off). In some embodiments, the training data 612 can be generated by setting any labeled device-specific energy usage value above the threshold to 1 and any labeled device-specific energy usage value below the threshold to 0. For example, the labeled device-specific energy usage in the training data 612 may not include the total of the energy usage of the labeled devices (e.g., instead include binary 1 or 0).
[0093] The prediction module 606 can generate a predicted disaggregated energy usage of the target device based on the input data 602, and the prediction module 610 can generate a target device detection prediction based on the input data 602. These predictions from the prediction modules 606 and 610 can be input into the combination module 614, which can generate a combined disaggregation prediction of the target device as the output 616. In some embodiments, the combination module 614 combines the disaggregation prediction from the prediction module 606 and the detection prediction from the prediction module 610 by adding a value to the predicted disaggregation of the target device based on the detection prediction, such as when the detection prediction does not match the disaggregation prediction.
[0094] For example, if prediction module 606 generates little or no predicted energy usage of the target device during a specific period (e.g., 1 hour), and prediction module 610 generates a prediction that the target device was using energy during this period (e.g., a prediction indicating that the target device was on and using energy), combination module 614 can enhance the disaggregation prediction by adding an energy usage value (e.g., a threshold of the energy usage value or a predetermined total) to the predicted disaggregation for that period. Similarly, if prediction module 606 generates a substantial predicted energy usage value (e.g., an energy usage value exceeding a threshold) of the target device during a specific period (e.g., 1 hour), and prediction module 610 generates a prediction that the target device was not using energy during this period (e.g., a prediction indicating that the target device was off and not using energy), combination module 614 can enhance the disaggregation prediction by subtracting an energy usage value (e.g., a threshold of the energy usage value or a predetermined total) from the predicted disaggregation for that period.
[0095] In some embodiments, combination module 614 can use a weighting algorithm to combine the disaggregation prediction from prediction module 606 and the detection prediction from prediction module 610. For example, prediction module 610 can generate detection predictions (e.g., by indicating on / off) during the usage time of the target device at granularities such as 1 minute, 5 minutes, 15 minutes, 30 minutes, 45 minutes, 1 hour, or other similar granularities. Prediction module 606 can generate disaggregation predictions that estimate how much energy the target device uses at granularities such as 1 minute, 5 minutes, 15 minutes, 30 minutes, 45 minutes, 1 hour, or other similar granularities.
[0096] The combination module 614 can implement a general weighting scheme that uses parameters to prioritize one prediction over another. For example, when combining the generated predictions, one or more parameters can be used to construct weights (e.g., α and / or β weights, as well as first and second thresholds). In this example, based on the values of these parameters, the degree to which the disaggregation prediction is enhanced by the detection prediction is set. In some embodiments, a match between the disaggregation prediction and the detection prediction may be sought. For example, sometimes the disaggregation prediction may predict the energy usage of the target device while the detection prediction indicates that the device was off. Similarly, sometimes the disaggregation prediction may not predict the energy usage of the target device while the detection prediction indicates that the device was on.
[0097] In some embodiments, one or more thresholds may be set to enhance the disaggregation prediction if there is a mismatch with the detection prediction. For example, if the detection prediction indicates that the target device was on but the disaggregation prediction does not include the predicted energy usage over a relevant period (e.g., a relevant 15-minute, 30-minute, 45-minute, or 1-hour time window), a first threshold amount of energy usage can be added to the disaggregation prediction. In this example, if the disaggregation prediction includes a predicted energy usage that is below the first threshold, this enhancement can include increasing the predicted energy usage to the first threshold. Similarly, if the detection prediction indicates that the target device was off but the disaggregation prediction includes a predicted energy usage that is greater than a second threshold amount over the relevant period, the predicted energy usage from the disaggregation prediction can be reduced to the second threshold amount of energy usage.
[0098] In some embodiments, by using one or more weighting parameters (e.g., α and β weights), the degree to which the energy usage value of the disaggregation prediction is enhanced by the detection prediction can be adjusted. For example, the α weight may be related to a first threshold, and the weight may control the energy usage added to the disaggregation prediction. In a sample implementation, when the α weight is set to "1", the energy usage over the relevant time window within the disaggregation prediction can be increased to the first threshold, and when the weight is set to "0", no energy usage is added. Intermediate values of the α weight between "1" and "0" can add an energy usage proportional to the weight. For example, a weight of "0.5" can increase the energy usage to half of the first threshold, or the delta between the energy usage within the disaggregation prediction over the relevant period and the first threshold can be determined, and an energy usage equal to "0.5" of the delta can be added.
[0099] Similarly, the β weight may be related to a second threshold, and the weight may control the reduction in energy usage from the disaggregation prediction. In a sample implementation, when the β weight is set to "1", the energy usage over the relevant time window within the disaggregation prediction can be reduced to the second threshold, and when the weight is set to "0", the energy usage is not reduced. Intermediate values of the β weight between "1" and "0" can reduce the energy usage proportional to the weight. For example, a weight of "0.5" can reduce the energy usage to half of the second threshold, or the delta between the energy usage within the disaggregation prediction over the relevant period and the second threshold can be determined, and an energy usage equal to "0.5" of the delta can be subtracted from the disaggregation prediction. In some embodiments, one or any combination of these parameters (e.g., α, β, the first threshold, and / or the second threshold) can be implemented, or any other suitable weighting scheme can be implemented.
[0100] In some embodiments, the combination module 614 may include a third machine learning model trained / configured to combine the disaggregation prediction and the detection prediction. For example, by training the third trained machine learning model, the total energy usage of the target device can be predicted using the disaggregation prediction and the detection prediction. In some embodiments, the disaggregation prediction and the detection prediction are combined using the prediction generated by the third trained machine learning model. For example, the training data for the third machine learning model may include the disaggregation prediction, the detection prediction, and the labeled (known) energy usage of the target device. In this example, the input to the third machine learning model includes the disaggregation prediction and the detection prediction outputs, and thus the training data includes these predictions along with the labeled (known) energy usage data, and loss and gradient calculations can be performed during training.
[0101] In some embodiments, the training data / input to the third machine learning model may also include the overall source location (e.g., household) energy usage. For example, this overall source location energy usage is part of the training data used to train prediction modules 606 and 610 and also functions as input to prediction modules 606 and 610 to generate disaggregation predictions and detection predictions. The third trained machine learning model may find that trends from the overall source location energy usage affect accuracy when combining disaggregation predictions and detection predictions. Thus, when training the third machine learning model, by using the overall source location energy usage, one can learn how to combine these predictions. Similarly, when generating a combined prediction to combine the disaggregation prediction and the detection prediction, both these predictions and the overall source location energy usage can be used as input.
[0102] In some embodiments, the third machine learning model may be a deep learning model, for example, based on the number of hidden layers implemented. In some embodiments, the disaggregation prediction and the detection prediction are combined by combination module 614 using a decision tree, a random forest algorithm, Bayesian learning, or other suitable combination techniques.
[0103] FIG. 7 shows a flowchart for training a machine learning model to disaggregate energy usage associated with a target device according to an example of an embodiment. In some embodiments, the functions of FIGS. 7-11 may be implemented by software stored in a memory or other computer-readable or tangible medium and executed by a processor. In other embodiments, each function may be executed by hardware (e.g., using an application specific integrated circuit (“ASIC”), programmable gate array (“PGA”), field programmable gate array (“FPGA”), etc.), or may be executed by any combination of hardware and software. In an embodiment, the functions of FIGS. 7-11 can be executed by one or more elements of the system 200 of FIG. 2.
[0104] At 702, energy usage data including energy usage by a target device and one or more non-target devices at a plurality of source locations can be received. For example, the energy usage data can be similar to the data shown in Table 1 above. In some embodiments, the received data can include a timestamp, the overall energy usage (including the energy used by multiple devices) at a source location (e.g., a household), and the labeled device-specific energy usage of one or more of the target and non-target devices. The energy usage data can be received by monitoring energy usage, or from a third party, or based on a joint venture, or through any other suitable channel or entity.
[0105] At 704, a machine learning model can be configured. For example, a machine learning model such as a neural network, CNN, RNN, Bayesian network, support vector machine, or any other suitable machine learning model can be configured. Parameters such as the number of layers (e.g., the number of hidden layers), input shape, output shape, width, depth, direction (e.g., feedforward or bidirectional), activation function, type of layer or unit (e.g., gated recurrent unit, long short-term memory, etc.), or other suitable parameters of the machine learning model can be selected. In some embodiments, these configured parameters can be adjusted (e.g., tuned, globally changed, added, or removed) when training the model.
[0106] In some embodiments, the machine learning model may include a CNN. In this case, parameters such as the type of layer (e.g., convolutional, pooling, fully connected, etc.), kernel size and type, stride, and other parameters can also be configured. These configured parameters can also be adjusted when training the model.
[0107] At 706, energy usage data can be processed to generate training data. For example, training data can be generated by processing the energy usage data based on a target device and a set of other devices (e.g., one or more non-target devices). The processing can be based on a correspondence relationship between the energy usage data (e.g., the availability of labeled device-specific energy usage of various devices within the energy usage data) and the set of other devices. In some embodiments, the set of other devices can be selected based on the available energy usage data and / or the target device, and the energy usage data used to generate the training data can be selected based on the set of other devices and / or the target device, or the set of other devices and the energy usage data can be considered in combination, and both can be selected based on the correspondence relationship between them (e.g., beneficial for training / performance).
[0108] In some embodiments, the set of other devices is determined based on the energy usage of multiple households in the training data. A plurality of the other devices among the set of other devices can also be based on the energy usage of multiple households in the training data. In some embodiments, the set of other devices is determined based on known energy usage values of the set of other devices included in the energy usage of multiple households in the training data. The set of other devices can be determined such that the amount of training data configured to train a machine learning model for the set of other devices meets a criterion (e.g., a threshold amount).
[0109] In some embodiments, by processing energy usage data based on a target device and a set of other devices participating in training, the data can be enhanced with zero-valued energy usage. For example, an instance of data (e.g., a row) can include a timestamp, the overall energy usage at a source location, and various labeled device-specific energy usages. If either the target device or any of the set of other devices is not included in the labeled device-specific energy usage for an instance, these missing (or otherwise invalid) entries can be filled with zero values. In some embodiments, for a particular source location (e.g., a household) where the energy usage does not include labeled energy usage from a subset of the set of other devices, the training data can be processed such that the labeled energy usage of the subset of other devices is set to zero. This processing allows the learning mechanism to still learn from many of the training data in the training data, so that the available training data can be used efficiently. Additionally, the correspondence between the energy usage data and the set of other devices selected to participate in learning mitigates any potential learning problems that may arise due to the insertion of zero values.
[0110] In 708, by training a machine learning model using the generated training data, the disaggregated energy usage of the target device can be predicted. This training can include prediction generation, loss calculation (e.g., based on a loss function), and gradient propagation (e.g., through the layers / neurons of the machine learning model). As described herein, both the labeled energy usage of the target device and the labeled energy usage of the set of other devices are used to train the machine learning model.
[0111] In some embodiments, a trained machine learning model is trained using energy usage values of a plurality of households, labeled energy usage values of a target device, and labeled energy usage values of a set of other devices. Training data used to train the machine learning model may include energy usage from a plurality of households, labeled energy usage values of the target device among the energy usage of these households, and labeled energy usage values of a set of other devices among the energy usage of these households. In some embodiments, training the machine learning model can optimize the accuracy of predicting the energy usage value of the target device.
[0112] In some embodiments, training data including energy usage values of a plurality of households, labeled energy usage values of a target device, and labeled energy usage values of a set of other devices has a granularity of substantially one hour. Other suitable granularities (e.g., 1 minute, 15 minutes, 30 minutes, 45 minutes, etc.) can be similarly achieved. In some embodiments, the machine learning model, the training data utilized, and / or a set of other devices can be adjusted based on the results of training. For example, testing the trained model can show the accuracy of the trained model, and various adjustment and tuning can be performed based on the test accuracy.
[0113] At 710, a trained machine learning model can be stored. For example, the storage of a trained learning model that generates predictions meeting a criterion (e.g., an accuracy criterion or a threshold) can be performed so that disaggregation predictions can be executed using the stored model.
[0114] FIG. 8 shows a flowchart for predicting disaggregated energy usage associated with a target device using a trained machine learning model according to an example of an embodiment. For example, the functions of FIG. 8 can be executed by using a machine learning model trained based on the functions of FIG. 7.
[0115] At 802, household energy usage data over a period of time can be received, where the household energy usage includes the energy consumed by the target device and the energy consumed by a plurality of other devices. For example, the household energy usage data can be divided into time intervals (e.g., with a granularity of substantially one hour) based on timestamps over a period of time such as one day. Other suitable granularities can be achieved.
[0116] At 804, the received energy usage data can be processed. For example, this processing can be similar to the processing of training data. In such an example, this processing can transform the household energy usage input data to be similar to the training data, and by doing so, the trained machine learning model can achieve improved prediction results. This processing can include achieving a granularity of one hour for the energy usage data, normalization, other forms of scaling, and any other suitable processing.
[0117] At 806, the processed data can be provided as input data to the trained machine learning model. For example, a model trained according to the functions of FIG. 7 can be stored, and the processed data can be provided as input to the trained model. At 808, a prediction can be generated by the trained machine learning model. For example, the disaggregated energy usage of the target device based on the overall received energy usage can be predicted by the trained machine learning model.
[0118] In some embodiments, the prediction can have the same granularity as the input provided to the trained model. For example, the predicted energy disaggregation for the target device can have a granularity of substantially one hour. In some embodiments, the predicted energy usage includes the predicted energy usage of the target device over at least one day with a granularity of at least substantially one hour.
[0119] FIG. 9 shows a flowchart for predicting disaggregated energy usage associated with a target device using a trained convolutional neural network according to an example of an embodiment. For example, the function of FIG. 9 can be executed by using a convolutional neural network trained based on the function of FIG. 7.
[0120] In some embodiments, by using the function of FIG. 7, a CNN having a mix of convolutional layers and fully connected layers can be trained. For example, the CNN can include a plurality of convolutional layers followed by one or more fully connected layers. The CNN can be similar to the network shown in FIGS. 4A and / or 4B. For example, one of the CNN layers can be a convolutional layer having a one-dimensional kernel of a first size, and another one of the CNN layers can be a convolutional layer having a one-dimensional kernel of a second size. In some embodiments, the first size is smaller than the second size. In some embodiments, at least two of the CNN layers can be parallel convolutional layers, and as shown in FIG. 4B, a concatenation layer can be used to concatenate the parallel branches within the CNN. In such embodiments, the parallel branches can be configured to learn different features of the input / training data. For example, the kernel size, stride, and padding realized for the first branch among the parallel branches can be made different from the kernel size, stride, and padding realized for the second branch among the parallel branches.
[0121] At 902, input data including energy usage data of a source location over a period of time can be received. For example, the source location can be a household, and the household energy usage can include the energy consumed by the target device and the energy consumed by a plurality of other devices. For example, the energy usage of the source location can be divided into time intervals (e.g., with a granularity of substantially one hour) based on timestamps over a period of time such as one day.
[0122] At 904, the received energy usage data can be processed. For example, this processing can be similar to the processing of training data. In such an example, this processing can change the energy usage input data at the source location to be similar to the training data, and by doing so, the trained machine learning model can achieve improved prediction results. This processing can include achieving a granularity of the energy usage data in hourly units, normalization, other forms of scaling, and any other appropriate processing.
[0123] At 906, the processed data can be provided as input data to the trained convolutional neural network. For example, a convolutional neural network trained according to the functions of FIG. 7 can be stored, and the processed data can be provided as input to the trained network. At 908, a prediction can be generated by the trained convolutional neural network. For example, the disaggregated energy usage of the target device based on the overall received energy usage can be predicted by the trained convolutional neural network. In some embodiments, predicting the disaggregated energy usage of the target device includes, at least, advancing the input data in a feed-forward direction through the trained CNN such that the shape of the input data changes between a first layer (having a first 1D kernel) and a second layer (having a second 1D kernel). In some embodiments, the first layer and the second layer have a parallel orientation within the CNN.
[0124] In some embodiments, the predictions can have the same granularity as the inputs provided to the trained network. For example, the predicted energy disaggregation for a target device can have a granularity of substantially one hour. In some embodiments, the predicted energy usage includes the predicted energy usage of the target device over at least one day at a granularity of at least substantially one hour.
[0125] FIG. 10 shows a flowchart for disaggregating energy usage associated with a target device by training a plurality of machine learning models, according to an example of an embodiment. At 1002, energy usage data including the energy usage by the target device and one or more other non-target devices at a plurality of source locations can be received. For example, the energy usage data can be similar to the data shown in Table 1 above. In some embodiments, the received data can include a timestamp, the overall energy usage (including the energy used by multiple devices) at a certain source location (e.g., a household), and the labeled device-specific energy usage of one or more of the target and non-target devices. The energy usage data can be received by monitoring energy usage, or from a third party, or based on a joint venture, or through any other suitable channel or entity.
[0126] In 1004, a first machine learning model and a second machine learning model can be configured. For example, a machine learning model such as a neural network, CNN, RNN, Bayesian network, support vector machine, or any other suitable machine learning model can be configured. Parameters can be selected, such as the number of layers (e.g., the number of hidden layers), input shape, output shape, width, depth, direction (e.g., feedforward or bidirectional), activation function, type of layer or unit (e.g., gated regression unit, long short-term memory, etc.), or other suitable parameters of the machine learning model. In some embodiments, these configured parameters can be adjusted (e.g., tuned, globally changed, added, or deleted) when training the model.
[0127] In some embodiments, the first machine learning model can be designed / configured to disaggregate the energy usage of the target device from the energy usage at the source location. For example, the machine learning model can be similar to those designed / trained and / or implemented in FIGS. 7, 8, and 9. In some embodiments, the second machine learning model can be designed / configured to detect the energy usage of the target device from the energy usage at the source location. For example, as a point where detection can differ from disaggregation, disaggregation is intended to obtain the total energy usage by the target device, whereas detection is intended to detect the energy usage exceeding a threshold by the target device. The realization of detection can be intended to distinguish between a target device using energy while ON and a target device in standby mode (which can draw low-level energy) by detecting the energy usage exceeding the threshold. The threshold used to distinguish between standby mode (which can be interpreted as off) and on can depend on the target device. In some embodiments, the predicted disaggregation can take numerical values (e.g., within a certain range of values), whereas the predicted detection is binary (e.g., on or off).
[0128] In 1006, energy usage data can be processed to generate a training data set. For example, a training data set can be generated by processing the energy usage data based on a target device and a set of other devices (such as one or more non-target devices). The processing can be based on the correspondence between the energy usage data (such as the availability of labeled device-specific energy usage of various devices within the energy usage data) and a set of other devices. In some embodiments, the set of other devices can be selected based on the available energy usage data and / or the target device, and the energy usage data used to generate the training data can be selected based on the set of other devices and / or the target device, or the set of other devices and the energy usage data can be considered in combination, and both of these can be selected based on the correspondence between them (such as beneficial for training / performance).
[0129] In some embodiments, the set of other devices is determined based on the energy usage of multiple households in the training data. A plurality of the other devices among the set of other devices can also be based on the energy usage of multiple households in the training data. In some embodiments, the set of other devices is determined based on the known energy usage values of the set of other devices included in the energy usage of multiple households in the training data. The set of other devices can be determined such that the amount of training data configured to train a machine learning model for the set of other devices meets a criterion (such as a threshold amount).
[0130] In some embodiments, by processing energy usage data based on a target device and a set of other devices participating in training, the data can be enhanced with zero-valued energy usage. For example, an instance of data (e.g., a row) can include a timestamp, the overall energy usage at a source location, and various labeled device-specific energy usages. If either the target device or any of the set of other devices is not included in the labeled device-specific energy usage of an instance, these missing (or otherwise invalid) entries can be filled with zero values. In some embodiments, for a particular source location (e.g., a household) whose energy usage does not include the labeled energy usages of a subset of the set of other devices, the training data can be processed such that the labeled energy usages of the subset of other devices are set to zero. This processing allows the learning mechanism to still learn from many of the training data in the training data, so that the available training data can be used efficiently. Additionally, the correspondence between the energy usage data and the set of other devices selected to participate in learning alleviates any potential learning problems that may arise due to the insertion of zero values.
[0131] In some embodiments, processing the energy usage data may include generating a first training data set for a first machine learning model and a second training data set for a second machine learning model. For example, a first machine learning model that generates disaggregation predictions is trained using labeled energy usage data that includes the total energy used (e.g., over a certain time span). Thus, the labeled device-specific energy data in the first training data set may include labeled device-specific energy usage that reflects the total energy used, such as that shown in Tables 1-6 above. In some embodiments, a second machine learning model that generates detection predictions is trained using labeled energy usage data that includes the detected energy usage (e.g., over a certain time span). Thus, the labeled device-specific energy data in the second training data set may include labeled device-specific energy usage that indicates whether energy above a threshold was used (e.g., a binary value representing on or off).
[0132] In some embodiments, the second training data set can be generated by setting any labeled device-specific energy usage value above a threshold to 1 and any labeled device-specific energy usage value below the threshold to 0. For example, the labeled device-specific energy usage in the second training data set may not include the total energy usage of the labeled device (and instead includes binary 1 or 0).
[0133] In 1008, by training a first machine learning model and a second machine learning model using the generated first training dataset and second training dataset, it is possible to predict the disaggregated energy usage (e.g., the total energy usage) of the target device and the detected energy usage of the target device (e.g., an on or off detection such as energy usage exceeding a threshold). The training may include prediction generation, loss calculation (e.g., based on a loss function), and gradient propagation (e.g., through the layers / neurons of the machine learning model). As described herein, both the labeled energy usage of the target device and the labeled energy usage of a set of other devices are used to train the machine learning model.
[0134] In some embodiments, the trained machine learning model is trained using the energy usage values of multiple households, the labeled energy usage value of the target device, and the labeled energy usage values of a set of other devices. The first training dataset and the second training dataset used to train the first machine learning model and the second machine learning model may include the energy usage from multiple households, the labeled energy usage value of the target device within the energy usage of these households, and the labeled energy usage values of a set of other devices within the energy usage of these households. In some embodiments, the training of the first machine learning model can optimize the accuracy of predicting the energy usage value (e.g., the total energy usage) of the target device, and the training of the second machine learning model can optimize the accuracy of predicting the detected energy usage of the target device (e.g., an on or off detection such as energy usage exceeding a threshold).
[0135] In some embodiments, the first training data set and the second training data set, which include the values of the energy usage of a plurality of households, the labeled energy usage values of the target device, and the labeled energy usage values of a set of other devices, have a granularity of substantially one hour. Other suitable granularities can be realized as well. In some embodiments, the machine learning model, the training data set utilized, and / or a set of other devices can be adjusted based on the results of the training. For example, testing the trained model can show the accuracy of the trained model, and various adjustment and tuning can be performed based on the test accuracy.
[0136] At 1010, the trained first machine learning model and the trained second machine learning model can be stored. For example, the storage of the trained learning model that generates predictions meeting a criterion (such as an accuracy criterion or a threshold) can be performed so that disaggregation and detection predictions can be executed using the stored model.
[0137] FIG. 11 shows a flowchart for predicting the disaggregated energy usage associated with a target device using a plurality of trained machine learning models according to an example of an embodiment. For example, the functions of FIG. 11 can be executed using a plurality of machine learning models trained based on the functions of FIG. 10.
[0138] At 1102, household energy usage data with a granularity of substantially one hour over a period of time can be received, and the energy usage of this household includes the energy consumed by the target device and the energy consumed by a plurality of other devices. For example, the household energy usage data can be divided into time intervals (such as a granularity of substantially one hour or other suitable granularity) based on timestamps over a period of time such as one day.
[0139] At 1104, the received energy usage data can be processed. For example, this processing can be similar to the processing of the training data. In such an example, by this processing, the household energy usage input data can be changed to be similar to the training data, and by doing so, the trained machine learning model can achieve improved prediction results. This processing may include achieving a granularity of the energy usage data in one-hour units, normalization, other forms of scaling, and any other appropriate processing.
[0140] In some embodiments, this processing may include generating first input data for a first trained machine learning model and second input data for a second trained machine learning model. For example, the first machine learning model can be trained / configured to disaggregate the energy usage of a target device from the input data, and the second machine learning model can be trained / configured to detect the energy usage of the target device from the input data.
[0141] At 1106, the first input data can be provided to the first trained machine learning model, and a prediction can be generated for the disaggregated energy usage from among the energy usage data at the source location. For example, the total of the disaggregated energy usage of the target device over the above period based on the received household energy usage can be predicted by the first trained machine learning model.
[0142] In some embodiments, the first trained machine learning model can be trained according to the functionality of FIG. 10, and the first input data can be provided as input to the trained model. In some embodiments, the prediction can have the same granularity as the first input / energy usage data provided to the trained model. For example, the predicted energy disaggregation for the target device can have a granularity of substantially one hour. In some embodiments, the predicted energy usage includes the predicted energy usage for the target device with a granularity of at least substantially one hour over at least one day.
[0143] At 1108, the second input data can be provided to the second trained machine learning model, and a prediction can be generated for the energy usage detected from the energy usage data at the source location. For example, the detected energy usage of the target device over the above period based on the received household energy usage can be predicted by the second trained machine learning model.
[0144] In some embodiments, the second trained machine learning model can be trained according to the functionality of FIG. 10, and the second input data can be provided as input to the trained model. In some embodiments, the prediction can have the same granularity as the second input / energy usage data provided to the trained model. For example, the prediction of the detected energy usage for the target device can have a granularity of substantially one hour. In some embodiments, the prediction of the detected energy usage includes the detected energy usage for the target device with a granularity of at least substantially one hour over at least one day.
[0145] At 1110, the predicted outputs from the first machine learning model and the second machine learning model can be combined. For example, by combining a disaggregation prediction and a detection prediction, a combined prediction can be formed for the disaggregation of the energy usage of the target device.
[0146] In some embodiments, combining the disaggregation prediction and the detection prediction includes adding a value to the predicted disaggregation of the target device based on the detection prediction, such as when the detection prediction does not match the disaggregation prediction. In some embodiments, the disaggregation prediction and the detection prediction are combined using a weighting scheme that resolves the disagreement between the disaggregation prediction and the detection prediction.
[0147] In some embodiments, a third machine learning model can be trained, and the third machine learning model is trained / configured to combine the disaggregation prediction and the detection prediction. For example, the third trained machine learning model can be trained to predict the total energy usage of the target device using the disaggregation prediction and the detection prediction. In some embodiments, the disaggregation prediction and the detection prediction are combined using the prediction generated by the third trained machine learning model.
[0148] Embodiments perform non-intrusive load monitoring using a novel learning method. NILM and disaggregation take as input the total energy usage at a source location (e.g., the energy usage in a household provided by an advanced metering infrastructure), and estimate the energy usage of one or more electrical appliances, electric vehicles, and other devices that use energy at this source location. Embodiments predict the energy usage of a target device based on the overall energy usage at a source location by leveraging a trained machine learning model. For example, the target device may be a large electrical appliance or an electric vehicle, the source location may be a household, and the trained machine learning model can receive the energy usage of this household as input and predict the energy usage of the target device (e.g., the energy usage of the target device included in the overall energy usage of this household).
[0149] Embodiments train a machine learning model using labeled energy usage data. For example, the machine learning model may be a designed / selected neural network or the like. Energy usage data can be obtained from multiple source locations (e.g., households), and the energy usage data can be labeled with device-specific energy usage. For example, the value of household energy usage can cover a certain period, and the values of the energy usage of individual devices (e.g., electrical appliance 1, electric vehicle 1, electrical appliance 2, etc.) within this period can be labeled. Then, in some embodiments, by processing this household and device-specific energy usage, training data for the machine learning model can be generated.
[0150] In some embodiments, by training a machine learning model, the energy usage of a target device can be predicted (e.g., disaggregated). For example, the training data can include the energy usage specific to the target device at a plurality of different source locations (e.g., households), and thus, by training the machine learning model, the trends in the training data can be identified and the energy usage of the target device can be predicted. In some embodiments, while predicting the energy usage of the target device by training a machine learning model, the training data can include the energy usage prediction / loss calculation / gradient update of one or more other devices. For example, when implementing embodiments of training techniques (e.g., prediction generation, loss calculation, gradient propagation, accuracy improvement, etc.) for a machine learning model, a set of other devices can be included together with the target device.
[0151] The features, structures, or characteristics of the present disclosure described throughout this specification may be combined in any suitable manner in one or more embodiments. For example, throughout this specification, when expressions such as "one embodiment", "some embodiments", "certain embodiments", "certain multiple embodiments", or other similar expressions are used, it means that the specific features, structures, or characteristics described in relation to the embodiments may be included in at least one embodiment of the present disclosure. Thus, throughout this specification, when there are expressions such as "one embodiment", "some embodiments", "certain embodiments", "certain multiple embodiments", or other similar expressions, it does not necessarily mean that they all refer to the same group of embodiments, and the described features, structures, or characteristics may be combined in any suitable manner in one or more embodiments.
[0152] Those skilled in the art will readily understand that the above embodiments can be implemented with steps in a different order and / or with elements in a configuration different from the disclosed configuration. Accordingly, although the present disclosure has considered the embodiments outlined, it will be apparent to those skilled in the art that certain modifications, variations, and alternative configurations are still within the spirit and scope of the present disclosure. Therefore, reference should be made to the appended claims to determine the scope of the present disclosure.
Claims
1. A method for disaggregating energy usage associated with a target device using a convolutional neural network, the method comprising: storing a trained convolutional neural network (CNN), the CNN including a plurality of layers, one or more of the plurality of layers including a convolutional layer having a one-dimensional kernel, the CNN being trained to predict disaggregated target device energy usage data from source location energy usage data based on training data including labeled energy usage data from a plurality of source locations, the method further comprising: receiving input data including energy usage data for a source location over a period of time; and using the trained CNN to predict disaggregated target device energy usage based on the input data, predicting disaggregated target device energy usage including at least advancing the input data through the trained CNN in a feed-forward direction such that the shape of the input data changes between a first layer of the trained CNN and a second layer of the CNN, a first layer of the one or more layers including a convolutional layer having a one-dimensional kernel of a first size, a second layer of the one or more layers including a convolutional layer having a one-dimensional kernel of a second size, the first size being smaller than the second size, the first layer and the second layer having a parallel orientation within the CNN.
2. The method of claim 1, wherein the source location includes a household and the input data includes a granularity of one hour.
3. The method of claim 2, wherein the energy usage data for the source location includes energy consumed by the target device and energy consumed by a set of other devices.
4. The training data includes energy usage values of a plurality of households, and for most of the plurality of households, the energy usage includes the labeled energy usage from the target device and the labeled energy usage from at least one device among a set of other devices. The method according to claim 1.
5. The training data including the energy usage values from the plurality of households, the labeled energy usage value of the target device, and the labeled energy usage values of the set of other devices includes a granularity of one hour. The method according to claim 4.
6. For a particular household whose energy usage does not include the labeled energy usage from a subset of the set of other devices, the method according to claim 5 further includes the step of processing the training data such that the labeled energy usage of the subset of the other devices is set to zero.
7. The predicted energy usage includes the predicted energy usage of the target device with a granularity of at least one hour over at least one day. The method according to claim 6.
8. The set of other devices is determined based on the known energy usage values of the set of other devices among the energy usage of the plurality of households including the training data. The method according to claim 7.
9. The set of other devices is determined such that the amount of training data configured to train a machine learning model for the set of other devices meets a criterion. The method according to claim 8.
10. A system for disaggregating energy usage associated with a target device, the system comprising: a processor; a memory storing instructions executed by the processor, and by the instructions, the processor configured to store a trained convolutional neural network (CNN), the CNN including a plurality of layers, one or more of the plurality of layers including a convolutional layer having a one-dimensional kernel, the CNN being trained to predict disaggregated target device energy usage data from source location energy usage data based on training data including labeled energy usage data from a plurality of source locations, the processor further configured to receive input data including energy usage data for a source location over a period of time, configured to use the trained CNN to predict disaggregated target device energy usage based on the input data, predicting disaggregated target device energy usage including at least advancing the input data through the trained CNN in a feed-forward direction such that the shape of the input data changes between a first layer of the trained CNN and a second layer of the CNN, a first layer of the one or more layers including a convolutional layer having a one-dimensional kernel of a first size, a second layer of the one or more layers including a convolutional layer having a one-dimensional kernel of a second size, the first size being smaller than the second size, the first layer and the second layer having a parallel orientation within the CNN, a system. Claims 11 The system according to claim 10, wherein the source location includes a household and the input data includes a granularity of one hour. Claims 12 The system according to claim 11, wherein the energy usage data for the source location includes energy consumed by the target device and energy consumed by a set of other devices. Claims 13 A program for causing a processor to execute the method according to any one of claims 1 to 9.
Citation Information
Patent Citations
CNN based non-intrusive power consumption load decomposition method
CN108899892A
Power demand prediction apparatus and program
JP2016001951A
Object recognition device
JP2019095339A
Electrical meter system for engery desegregation
US20180328967A1