Non-intrusive load identification method based on reinforcement learning dynamic simulation
By using reinforcement learning dynamic simulation methods and leveraging electrical coupling data conversion models and scene generation strategy networks, high-quality simulation samples are generated. This solves the problems of recognition accuracy and robustness caused by the infinite combinations of equipment start-stop and electrical coupling effects in existing technologies, and achieves high-precision load decomposition and energy efficiency analysis.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- YANTAI DONGFANG WISDOM ELECTRIC
- Filing Date
- 2025-12-16
- Publication Date
- 2026-04-24
AI Technical Summary
Existing non-intrusive load identification methods are difficult to train models with sufficient experimental data in practical applications due to the infinite combinations of device start-up and shutdown timing. Furthermore, the electrical coupling effect causes differences between the characteristics of a single device and the characteristics of multiple devices operating simultaneously, which limits the identification accuracy and robustness.
A reinforcement learning-based dynamic simulation method is adopted. Through an electrical coupling data conversion model and a scene generation strategy network, a closed-loop process of "data acquisition - conversion model training - simulation optimization - classification model training" is constructed to generate high-quality multi-device superimposed simulation samples and train the classification model to improve recognition accuracy and generalization ability.
It significantly improves the model's recognition accuracy and generalization ability in real-world complex power environments, enabling accurate identification of operating equipment and its start-up and shutdown times, supporting power safety early warning and smart energy management.
Smart Images

Figure QLYQS_10 
Figure QLYQS_25 
Figure QLYQS_28
Abstract
Description
Technical Field
[0001] This invention relates to the field of power user-side energy consumption monitoring technology, specifically to a non-intrusive load identification method based on reinforcement learning dynamic simulation. Background Technology
[0002] Non-intrusive load identification is an intelligent sensing technology for end-user electricity consumption monitoring. Its core advantage lies in its "non-intrusiveness," meaning it only requires deploying data acquisition equipment at the main inlet of the user's power distribution system. By collecting electrical quantities such as total load power, voltage, and current at high frequencies, and further calculating its transient, steady-state, and harmonic characteristics, it can decompose the overall electricity load into the independent operating state of individual electrical devices, including the energy consumption contribution, precise start-stop times, and operating duration of each device. Taking low-voltage residential or commercial users as an example, this technology can finely break down the total load curve for a certain period into the specific operating trajectories of different electrical devices, achieving a perception leap from macroscopic total electricity consumption to microscopic electricity consumption classification. This refined decomposition capability is of great value for providing high-quality underlying data support for smart energy management scenarios such as electricity safety early warning, personalized energy-saving strategy formulation, and grid load forecasting and demand response. It is one of the key technologies driving the transformation of power distribution networks towards intelligent and refined operation.
[0003] Current mainstream non-intrusive load identification methods rely on building data-driven classification models. The solution typically involves: first, collecting operational data of one or more electrical devices under specific start-stop time combinations in a controlled experimental environment to build a training sample library; then, constructing machine learning models such as neural networks or recurrent neural networks, and training them based on the above samples to learn the mapping relationship from overall electrical characteristics to the operating state of specific devices; finally, deploying the trained model to the actual application scenario for load decomposition.
[0004] However, existing methods for training models based on experimental data still face technical challenges in practical engineering applications. Firstly, the startup sequence and time intervals of different devices in actual operation have virtually infinite possibilities. Manually exhaustively exploring all possible combinations of start-up and shutdown times to obtain corresponding training data would be exponentially more labor-intensive and costly with the increase in the number of device types, directly limiting the model's ability to generalize to unseen time-series scenarios. Secondly, the electrical characteristics exhibited by a single device operating independently differ from those exhibited by multiple devices operating simultaneously due to coupling factors such as line impedance and voltage fluctuations. Most existing methods rely on single-device independent operation data or simple synthetic data to construct training sets, resulting in a mismatch between the features learned by the model and those in real-world scenarios. This limits the model's accuracy and robustness in complex multi-device concurrent scenarios. The data scarcity problem caused by "infinite time-series combinations" and the feature mismatch problem caused by "electrical coupling" together make training a high-precision, high-generalization-capability load identification model in an experimental environment extremely difficult. Summary of the Invention
[0005] This invention proposes a non-intrusive load identification method based on reinforcement learning dynamic simulation. Its purpose is to solve the problems in the existing technology, such as the difficulty in fully training the model with experimental data due to the infinite combination of equipment start-up and shutdown timing, and the limitation of model generalization ability and identification accuracy caused by the difference between the characteristics of a single device and the characteristics of multiple devices under superposition due to electrical coupling effect. Thus, it can build a load identification model that can accurately adapt to real complex power consumption scenarios at low cost and high efficiency without exhaustive experiments.
[0006] The technical solution of this invention is as follows:
[0007] A non-intrusive load identification method based on reinforcement learning dynamic simulation includes:
[0008] Step S1: Prepare the training sample set for the classification model;
[0009] Step S2: Train the classification model based on the training sample set of the classification model;
[0010] The input to the classification model is the total composite time-series data under the superimposed operation scenario of multiple electrical devices, and the output is the set of devices consisting of the currently operating electrical devices and the start and stop times of each electrical device.
[0011] Step S3: Collect total composite time-series data at the user bus in the actual operation scenario, and input the collected total composite time-series data into the trained classification model to obtain the recognition result;
[0012] In step S1, based on the electrical coupling data conversion model and the scene generation strategy network, multi-device superimposed simulation sample data is generated to form a training sample set for the classification model.
[0013] As a further improvement to the non-intrusive load identification method based on reinforcement learning dynamic simulation, step S1 specifically includes the following sub-steps:
[0014] Step S1-1: Record the operating data of the electrical equipment under different operating states, including the data of each electrical equipment operating independently, as well as the data of the electrical equipment operating under multiple different superimposed operating scenarios;
[0015] Step S1-2: Construct an electrical coupling data conversion model The electrical coupling data conversion model is trained based on the data recorded in step S1-1. ;
[0016] Steps S1-3: Construct the training scene generation policy network The scenario generation policy network is trained based on the data recorded in step S1-1. ;
[0017] Steps S1-4: Generate a training sample set for the classification model based on the data transformation model and the scene generation strategy network.
[0018] As a further improvement to the non-intrusive load identification method based on reinforcement learning dynamic simulation: in step S1-1, recording... Independent state composite time-series data of a typical electrical device operating independently;
[0019] For the Individual electrical equipment, It controls its independent start-up, operation, and shutdown, and collects data through a waveform recording device during this process. Its independent state composite timing data is recorded as follows. , The maximum sampled sequence length is preset, where any runtime data The structure is as follows: In the above formula, Indicates the first Each electrical device at the sampling time power, Indicates the first Each electrical device at the sampling time Voltage data, Indicates the first Each electrical device at the sampling time Current data, Indicates the first Each electrical device at the sampling time Current harmonic amplitude data.
[0020] As a further improvement to the non-intrusive load identification method based on reinforcement learning dynamic simulation: in step S1-1, construct... Two different overlay scenarios were run, and data was recorded separately in each scenario;
[0021] For the One scenario, First, build a device set. and start / stop time configuration set ;in, Indicates the first Startup time of the equipment in this scenario. Indicates its closing time;
[0022] Configure a control system based on the equipment set and start / stop time to operate multiple electrical devices in combination, while collecting two types of data: one is for each participating electrical device. Superposition state composite time series data acquired by the branch. , for The structure and operation data collected in real time The same applies; secondly, the total composite timing data of all participating electrical devices is collected at the bus. , for The structure and operation data collected in real time same;
[0023] Based on the total composite timing data collected on the bus Calculate the first Statistical data vectors for each scenario ;
[0024] The first The complete information of each superimposed running scenario is denoted as a vector. .
[0025] As a further improvement to the aforementioned non-intrusive load identification method based on reinforcement learning dynamic simulation, according to the first Calculate the statistical data vector for a given time series data point in a given scenario. The method is as follows:
[0026] ;
[0027] In the above formula, This represents the mean of the power component in the time series data. This represents the standard deviation of the power component in the time series data. This represents the local variance of the power component in the time series data, calculated using a sliding window. and These represent the maximum and minimum values of the power component in the time series data, respectively. This represents the power factor calculated based on the power, voltage, and current in the time series data.
[0028] As a further improvement to the aforementioned non-intrusive load identification method based on reinforcement learning dynamic simulation: a data transformation model. The input consists of two parts: the first part is the independent state dataset. The second part is the overlay of vectors corresponding to the running scenario. ;
[0029] Data transformation model The architecture employs a "dual encoder-single decoder" structure, specifically including:
[0030] First encoder It employs a bidirectional long short-term memory network, with independent state datasets as input and device latent features as output. ;
[0031] Second encoder A multilayer perceptron is used, with vector input. The output is the scene's latent features. ;
[0032] decoder The decoder is a multilayer perceptron, and the input is the concatenated features. The output is the superposition state composite time series prediction data of all electrical devices in the corresponding scenario. And the total composite timing prediction data at the bus. ; Indicates the first The first superimposed running scenario Individual electrical equipment Real-time runtime data, For the first In a superimposed operating scenario, the bus is at The runtime data at each moment, both have the same structure as same.
[0033] As a further improvement to the aforementioned non-intrusive load identification method based on reinforcement learning dynamic simulation: training an electrical coupling data conversion model. Based on the dataset obtained in step S1-1: independent state dataset, vectors corresponding to each superimposed running scenario, and total composite time series data corresponding to each superimposed running scenario;
[0034] When training a data transformation model, the objective function Defined as:
[0035] ;
[0036] In the above formula, It is the current input vector Statistical data vectors in Replace with total composite time series prediction data Calculated statistical data vector The resulting vector The weighting coefficients are used to balance the two losses.
[0037] As a further improvement to the aforementioned non-intrusive load identification method based on reinforcement learning dynamic simulation: Scene Generation Policy Network Based on the reinforcement learning training method, the environment interaction process for reinforcement learning is designed as follows:
[0038] State space: A state is a vector of currently constructed, superimposed running scenarios. Its structure and definition are similar to those of a vector. Consistent;
[0039] Action space: Actions are defined in two types; type one is a triple. , indicating that the first Each electrical device is added to the device set of the current scene, and a startup time is assigned to it. and stop time Type Two is Special Actions This indicates that the construction of the current scene will be terminated;
[0040] State transition: when an action is performed After that, the state changed from Updated to The specific update method is as follows: [The text abruptly ends here, likely due to an incomplete sentence or a formatting error.] Individual electrical devices are added to the status. The collection of devices, and Join the state The set of start and stop time configurations is then used to calculate a new statistical data vector. : Retrieve the independent-state composite time-series data corresponding to each electrical device in the current device set. Then, based on the corresponding startup time and closing time ,Will The sequence is shifted and truncated on the time axis to form a set of expected independent running data. Then, all the expected independent running data corresponding to the current set of devices are summed and superimposed at the same time to obtain the expected total time series data. Then based on the expected total time series data Calculate the statistical data vector ;
[0041] Scene generation strategy network A multilayer perceptron is used, with the current state as the input, and the output layer simultaneously outputs the following three sub-probability distributions: 1) Adding electrical equipment and... The probability distribution is such that if the highest probability corresponds to the index of the electrical device, then that electrical device will be added in the generation action. The generated action is ;2) Startup time and The probability distribution is such that if the highest probability corresponds to a certain candidate time, then that candidate time is used as the start time in the generation action; if the highest probability corresponds to... The generated action is 3) Stop time and The probability distribution is such that if the maximum probability corresponds to a certain candidate time, then that candidate time is used as the stop time in the generated action; if the maximum probability corresponds to... The generated action is ;
[0042] Reward function: Each time with When constructing the overlay scene after the action ends, based on the final state... Calculate the reward function During training, the parameters of the scene generation policy network are updated based on the reward function.
[0043] As a further improvement to the non-intrusive load identification method based on reinforcement learning dynamic simulation, the reward function The calculation method is as follows:
[0044] ;
[0045] In the above formula, It is a preset target scene vector. This indicates that the final state has been obtained. The first in the process An intermediate state These are the weighting coefficients;
[0046] When training the scene generation policy network, the policy gradient method is used and the parameters of the scene generation policy network are updated based on the reward function. Suppose the current round has experienced Given a series of actions and intermediate states, a trajectory is generated. The gradient of the policy in the current round The calculation is as follows:
[0047] ;
[0048] In the above formula, Is the policy network in state Down Output Action The probability is obtained by multiplying the maximum probability values of the three sub-probability distributions; It is the baseline value, which is the average of the reward functions from the most recent rounds; Parameters for generating a policy network for a scene The gradient;
[0049] Update the network parameters of the scenario generation strategy using gradient ascent. ,in This is the learning rate.
[0050] As a further improvement to the non-intrusive load identification method based on reinforcement learning dynamic simulation, steps S1-4 are specifically implemented as follows: utilizing a trained scene generation policy network generate There are three different superimposed running scenarios, and the corresponding vectors are denoted as follows: For each generated overlay running scenario, its corresponding vector is... The independent state dataset is input into the pre-trained electrical coupling data transformation model. In this process, the total composite time series prediction data for this scenario is obtained. , as training samples corresponding to this superimposed running scenario Meanwhile, the set of devices and the set of start / stop time configurations for the superimposed running scenario are used as training samples. tags Thus, the training sample set is obtained. .
[0051] Compared with the prior art, the present invention has the following beneficial effects:
[0052] 1. This invention constructs a complete closed-loop process of "data acquisition - conversion model training - simulation optimization - classification model training - accurate recognition" by introducing an electrical coupling data conversion model and a scene generation strategy network. The method first collects and learns the electrical coupling rules between the independent operation of a single device and the superimposed operation of multiple devices, realizing the equivalent conversion from single-device data to multi-device superimposed scenarios. Then, through a reinforcement learning-driven dynamic simulation process, it can generate superimposed operation scenarios without manually exhaustively enumerating all start-stop time combinations, thereby obtaining a large number of near-realistic and diverse training samples. Finally, these high-quality simulation samples are used to train the classification model, thereby significantly improving the model's recognition accuracy and generalization ability in actual complex power environments.
[0053] 2. This invention trains an electrical coupling data conversion model with a "dual encoder-single decoder" structure to accurately learn and simulate changes in electrical characteristics caused by factors such as line impedance and voltage fluctuations when multiple devices operate simultaneously. The model takes as input the timing characteristics of a single device operating independently and a scene vector containing device combinations, start / stop times, and statistical data, and outputs the predicted total timing data of the superimposed operation. This design enables the model to effectively capture electrical coupling effects, solving the problem of decreased recognition performance caused by feature mismatch between training data (single-device data) and real-world application scenarios (multi-device superimposed data) in traditional methods, thus ensuring the input quality for subsequent classification model training from the data source.
[0054] 3. This invention also includes a scene generation policy network based on reinforcement learning, used to automatically generate diverse multi-device overlay operation scenarios. This policy network uses scene vectors as states and adds devices and specifies their start / stop times or terminates scene construction as actions. By defining a reward function that balances the "realism" and "novelty" of generated samples, the network is guided to explore and generate a large number of simulation scenarios with high coverage and high diversity in start / stop time sequence combinations. This mechanism fundamentally overcomes the bottleneck of traditional methods, which cannot exhaustively collect training data through experiments due to the near-infinite combinations of device start / stop time sequences. It expands the training sample library of the classification model and enhances the model's adaptability to unseen patterns.
[0055] 4. This invention utilizes a massive sample dataset generated by the aforementioned collaborative mechanism, containing high-quality simulation time-series data and its corresponding scene labels (equipment set and start / stop times of each device), to train the classification model. The trained classification model can more accurately learn the mapping relationship from complex superimposed electrical features to the operating state of specific equipment. In practical deployment, this method can achieve high-precision and robust identification of the composition of operating equipment and its precise start / stop times based solely on electrical monitoring data from the user's main access point, effectively supporting applications such as power safety early warning, refined energy consumption statistics, and smart energy management. Detailed Implementation
[0056] The technical solution of the present invention will be described in detail below. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments.
[0057] A non-intrusive load identification method based on reinforcement learning dynamic simulation is proposed. This method learns the data difference patterns of equipment in independent and superimposed operation states by constructing an electrical coupling data conversion model. It also uses a reinforcement learning-driven scene generation strategy network to generate massive, diverse, and realistic superimposed operation simulation samples at low cost and high efficiency. This solves the problems of insufficient training samples, distortion, and poor generalization ability caused by the infinite combination of equipment start-stop and electrical coupling effects in traditional methods.
[0058] The method specifically includes:
[0059] Step S1: Prepare the training sample set for the classification model. Unlike the conventional training sample preparation work, this step is based on the electrical coupling data conversion model and the scene generation strategy network to collaboratively generate multi-device superimposed simulation sample data with high diversity and high fidelity, so as to form the training sample set for the classification model.
[0060] Step S1 specifically includes the following sub-steps.
[0061] Step S1-1: Record the operating data of the electrical equipment under different operating conditions, including data of each electrical device operating independently, as well as data of the electrical equipment operating under multiple different superimposed operating scenarios. This step aims to obtain the basic real data for training subsequent models.
[0062] First, recording Independent state composite time-series data of a typical electrical device operating independently. For the first... Personal electrical equipment ( The system controls its independent startup, operation, and shutdown, and collects data through a waveform recording device during this process. Its independent-state composite timing data is recorded as follows: , The maximum sampled sequence length is preset, where any runtime data The structure is as follows:
[0063] ;
[0064] In the above formula, Indicates the first Each electrical device at the sampling time power, Represents voltage data. Represents current data. This represents the amplitude data of current harmonics.
[0065] This is used to standardize the data format, for example, by setting it to cover 1000 sampling points to cover a typical device start-up and shutdown process. This is for cases where the actual recorded data length is insufficient. The part is filled with zero values.
[0066] Then, build The system runs through several different overlapping scenarios, and records data for each scenario separately.
[0067] For the One scenario ( First, construct a device set. and start / stop time configuration set .in, Indicates the first Startup time of the electrical equipment in this scenario (sampling point index). Indicates its closing time (sampling point index).
[0068] In the experimental environment, multiple electrical devices were configured to operate in combination according to the device set and start / stop time, while two types of data were collected: one was for each participating electrical device. ( Superposition-state composite time-series data acquired by the branch where it is located , for The structure and operation data collected in real time The two methods are the same, but the length is determined by the actual start and stop times; secondly, the total composite timing data of all participating electrical devices is collected at the bus. , for The structure and operation data collected in real time The same, but the length is determined by the start and stop time of the entire scene. and The difference reflects the electrical coupling effect.
[0069] Furthermore, based on the total composite timing data collected on the bus Calculate the first Statistical data vectors for each scenario :
[0070] ;
[0071] In the above formula, Represents total composite time series data The average value of the medium power portion, Represents total composite time series data The standard deviation of the medium power component, Represents total composite time series data The local variance of the medium-power component is calculated using a sliding window. and These represent the total composite time series data. The maximum and minimum values of the medium power portion, Indicates based on total composite time series data The power factor is calculated from the power, voltage, and current in the data. These statistics describe the overall electrical characteristics of the superimposed operation. This calculation process is also applicable to other timing data with the same structure, which will not be elaborated further below.
[0072] The first The complete information of each superimposed running scenario is denoted as a vector. .
[0073] Step S1-2: Construct and train the electrical coupling data conversion model .
[0074] Electrically Coupled Data Conversion Model The goal is to learn the mapping pattern from independent device operation data and scenario definition to the superimposed operation data of various electrical devices in that scenario, that is, to approximate the real physical coupling function that is difficult to model directly with a parameterized model. Model training is based on the dataset obtained in step S1-1: the independent state dataset. Complete information on each overlay operation scenario and the total composite time series data corresponding to each overlay operation scenario.
[0075] Data transformation model The input consists of two parts: the first part is the independent state dataset, which has been padded with zeros to make it uniform. The first part is a two-dimensional array; the second part is a vector superimposed on the running scenario. ,vector This is obtained by vectorizing the complete information of the superimposed running scenario: First, construct a vector of length... Binary vector representation of device set If a certain electrical device is in the device set, the corresponding position is 1; otherwise, it is 0. Secondly, for each electrical device... Set its start and stop times Arranged in order, for those not in The start and stop times of the devices in the vector are represented by specific placeholders (such as -1). Finally, the device set vector, start and stop time vector, and statistics vector are combined. Concatenate them into a one-dimensional vector as The final representation.
[0076] Data transformation model The architecture employs a "dual encoder-single decoder" structure, specifically including:
[0077] First encoder The system employs a Bidirectional Long Short-Term Memory (Bi-LSTM) network. The input is an independent-state dataset, which passes through two hidden layers (e.g., dimensions 1000->512, 512->512) and one linear layer, outputting the device's latent features. .
[0078] Second encoder The system employs a multilayer perceptron (MLP), with vector input. (length is) After passing through hidden layers and linear layers, the output scene latent features are determined. .
[0079] decoder The decoder is a multilayer perceptron, and the input is the concatenated features. The output is the superposition state composite time series prediction data of all electrical devices in the corresponding scenario. And the total composite timing prediction data at the bus. . Indicates the first The first superimposed running scenario Individual electrical equipment Real-time runtime data, For the first In a superimposed operating scenario, the bus is at The runtime data at each moment, both have the same structure as The same applies. This method focuses more on whether the total composite time series prediction data closely approximates the actual collected data; the superposition state composite time series prediction data of electrical equipment is not included in the subsequent calculations.
[0080] When training a data transformation model, the objective function Defined as:
[0081]
[0082] In the above formula, the first term is the mean squared error (MSE) between the predicted total data and the actual total data. The second term is the constraint term, where... It is the current input vector Statistical data vectors in Replace with total composite time series prediction data Calculated statistical data vector The resulting vector, when generated using the L1 norm constraint, exhibits statistical properties consistent with the real data. To balance the weighting coefficients of the two losses, the parameters of the data transformation model are optimized using gradient descent (e.g., the Adam optimizer). To minimize To enable data transformation models It can accurately simulate electrical coupling effects.
[0083] Steps S1-3: Construct and generate policy networks based on reinforcement learning training scenarios .
[0084] Scene generation strategy network It is a reinforcement learning agent whose goal is to learn to automatically generate diverse and reasonable superimposed operating scenarios for subsequent simulation sample generation. The environment interaction process of reinforcement learning is designed as follows:
[0085] State space: A state is a vector of currently constructed, superimposed running scenarios. Its structure and definition are similar to those of a vector. Consistent.
[0086] Action Space: Actions are defined in two types. Type 1 is a triple. , indicating that the first Each electrical device is added to the device set of the current scene, and a startup time is assigned to it. and stop time Type Two is special actions. This indicates that the construction of the current scene will be terminated.
[0087] State transition: when an action is performed After that, the state changed from Updated to The specific update method is as follows: [The text abruptly ends here, likely due to an incomplete sentence or a formatting error Individual electrical devices are added to the status. The collection of devices, and Join the state The set of start and stop time configurations is then used to calculate a new statistical data vector. : Retrieve the independent-state composite time-series data corresponding to each electrical device in the current device set. Then, based on the corresponding startup time and closing time ,Will The sequence is shifted and truncated on the timeline (to match start and stop times) to form a set of expected independent running data. Then, all the expected independent running data corresponding to the current set of devices are summed and superimposed at the same time to obtain the expected total time series data. Then based on the expected total time series data Calculate the statistical data vector The above operation is based on the assumption that, under the ideal linear superposition approximation, the total electrical signal when multiple electrical devices are running simultaneously is approximately equal to the algebraic sum of the signals of each device operating independently. Although this differs from the actual physical coupling (the electrical coupling effect mentioned earlier), it is efficient and feasible for rapid and continuous evaluation of the "reasonableness" of a scenario during the training of reinforcement learning policy networks.
[0088] Scene generation strategy network A three-layer multilayer perceptron is used for implementation. The network input is the current state, and the output layer is designed to simultaneously output the following three sub-probability distributions: 1) Adding electrical equipment and... The probability distribution is such that if the highest probability corresponds to the index of the electrical device, then that electrical device will be added in the generation action. The generated action is ;2) Startup time and The probability distribution is such that if the highest probability corresponds to a certain candidate time, then that candidate time is used as the start time in the generation action; if the highest probability corresponds to... The generated action is 3) Stop time and The probability distribution is such that if the maximum probability corresponds to a certain candidate time, then that candidate time is used as the stop time in the generated action; if the maximum probability corresponds to... The generated action is In actual sampling, samples are taken sequentially based on these distributions to determine the specific actions.
[0089] Reward function: Each time with When constructing the overlay scene after the action ends, based on the final state... Calculate the reward function :
[0090] ;
[0091] In the above formula, It is a preset target scene vector. This indicates that the final state has been obtained. The first in the process An intermediate state These are the weighting coefficients. The first term represents the degree of completion of the current training round, and the second term represents the minimum distance between the new state and the historical state.
[0092] When training the scene generation policy network, the parameters of the scene generation policy network are updated using a policy gradient method (such as the REINFORCE algorithm). Suppose the current round has experienced... Given a series of actions and intermediate states, a trajectory is generated. The gradient of the policy in the current round The calculation is as follows:
[0093] ;
[0094] In the above formula, Is the policy network in state Down Output Action The probability is obtained by multiplying the maximum probability values of the three sub-probability distributions. This is the baseline value, used to reduce training variance. It is usually the average of the rewards obtained in the most recent rounds (e.g., 5 rounds). Parameters for generating a policy network for a scene The gradient is used to update the parameters of the scenario generation policy network through gradient ascent. ,in This is the learning rate.
[0095] After sufficient training, the scene generation strategy network can efficiently generate a large number of superimposed operating scene definitions that both conform to the target electrical characteristics (realistic) and differ significantly from each other (diverse).
[0096] Steps S1-4: Generate a training sample set for the classification model based on the data transformation model and the scene generation strategy network.
[0097] The specific method is as follows: Utilizing a pre-trained scene generation policy network. generate There are three different superimposed running scenarios, and the corresponding vectors are denoted as follows: ,in For each generated overlay running scenario, its corresponding vector is... The independent state dataset is input into the pre-trained electrical coupling data transformation model. In this process, the total composite time series prediction data for this scenario is obtained. , as training samples corresponding to this superimposed running scenario Meanwhile, the set of devices and the set of start / stop time configurations for the superimposed running scenario are used as training samples. tags .
[0098] Therefore, a large-scale, high-quality training sample set can be automatically constructed for training the final load decomposition classification model. .
[0099] Step S2: Train the classification model based on the training sample set of the classification model.
[0100] This step employs supervised learning to train a load decomposition and classification model (e.g., based on a deep neural network). The model's input is the total composite time-series data under the scenario of multiple electrical devices operating simultaneously, and its output is the set of currently operating electrical devices and the start-up and shutdown times of each device.
[0101] The specific process is as follows: The large-scale sample set generated in steps S1-4 is... The dataset is divided into training and validation sets. A classification model is constructed (such as a sequence-to-sequence model using an encoder-decoder framework, or a model based on an attention mechanism; however, classification models are existing technologies and not the innovation of this invention, so they will not be elaborated upon here). Let be the input, and let the predicted output of the classification model be denoted as . Define the loss function. (e.g., cross-entropy loss or mean squared error loss, by) and (Calculated), the parameters of the classification model are iteratively optimized using backpropagation and gradient descent algorithms (such as Adam) to minimize This enables the model to accurately extract the operating information of each independent device from the superimposed signal. Since the training data comes from a simulation model that has learned the laws of real electrical coupling and covers a massive and diverse range of device combinations and start-up / shutdown sequences, the trained classification model has extremely strong generalization ability and adaptability to real-world scenarios.
[0102] Step S3: Collect total composite timing data at the user bus in the actual operation scenario, input the collected total composite timing data into the trained classification model, and obtain the recognition result, namely: the set of devices consisting of the currently running electrical equipment and the start and stop times of each electrical equipment.
[0103] In practical deployment, a non-intrusive load monitoring device is installed at the main inlet of the power distribution system of the user to be monitored (such as a home or commercial building). This device acquires the raw voltage and current waveforms of the bus at a high frequency (e.g., thousands of points per second). The acquired waveform data is processed to calculate the total power. ,Voltage Current Harmony Combined time-series data are used to form the input to be identified. .Will The data is input into the classification model that has been fully trained and solidified in step S2. The classification model automatically performs analysis and calculation and outputs the recognition results, thereby achieving accurate and non-intrusive decomposition from the total power consumption "black box" to the power consumption "white box" of each device, providing a reliable data foundation for applications such as energy efficiency analysis, abnormal power consumption detection, and demand response.
[0104] It should be noted that, as will be apparent to those skilled in the art, the present invention is not limited to the details of the exemplary embodiments described above, and that the present invention can be implemented in other specific forms without departing from the spirit or essential characteristics thereof. The scope of the present invention is defined by the claims rather than the foregoing description.
Claims
1. A non-intrusive load identification method based on reinforcement learning dynamic simulation, comprising: Step S1: Prepare the training sample set for the classification model; Step S2: Train the classification model based on the training sample set of the classification model; The input to the classification model is the total composite time-series data under the superimposed operation scenario of multiple electrical devices, and the output is the set of devices consisting of the currently operating electrical devices and the start and stop times of each electrical device. Step S3: Collect total composite time-series data at the user bus in the actual operation scenario, and input the collected total composite time-series data into the trained classification model to obtain the recognition result; Its features are: In step S1, based on the electrical coupling data conversion model and the scene generation strategy network, multi-device superimposed simulation sample data is generated to form a training sample set for the classification model. Step S1 specifically includes the following sub-steps: Step S1-1: Record the operating data of the electrical equipment under different operating states, including the data of each electrical equipment operating independently, as well as the data of the electrical equipment operating under multiple different superimposed operating scenarios; In step S1-1, record Independent state composite time-series data of a typical electrical device operating independently; For the Individual electrical equipment, It controls its independent start-up, operation, and shutdown, and collects data through a waveform recording device during this process. Its independent state composite timing data is recorded as follows. , The maximum sampled sequence length is preset, where any runtime data The structure is as follows: In the above formula, Indicates the first Each electrical device at the sampling time power, Indicates the first Each electrical device at the sampling time Voltage data, Indicates the first Each electrical device at the sampling time Current data, Indicates the first Each electrical device at the sampling time Current harmonic amplitude data; In step S1-1, construct Two different overlay scenarios were run, and data was recorded separately in each scenario; For the One scenario, First, build a device set. and start / stop time configuration set ;in, Indicates the first Startup time of the equipment in this scenario. Indicates its closing time; Configure a control system based on the equipment set and start / stop time to operate multiple electrical devices in combination, while collecting two types of data: one is for each participating electrical device. Superposition state composite time series data acquired by the branch. , for The structure and operation data collected in real time The same applies; secondly, the total composite timing data of all participating electrical devices is collected at the bus. , for The structure and operation data collected in real time same; Based on the total composite timing data collected on the bus Calculate the first Statistical data vectors for each scenario ; The first The complete information of each superimposed running scenario is denoted as a vector. ; According to the Calculate the statistical data vector for a given time series data point in a given scenario. The method is as follows: ; In the above formula, This represents the mean of the power component in the time series data. This represents the standard deviation of the power component in the time series data. This represents the local variance of the power component in the time series data, calculated using a sliding window. and These represent the maximum and minimum values of the power component in the time series data, respectively. This represents the power factor calculated based on the power, voltage, and current in the time series data. Step S1-2: Construct an electrical coupling data conversion model The electrical coupling data conversion model is trained based on the data recorded in step S1-1. ; Data transformation model The input consists of two parts: the first part is the independent state dataset. The second part is the overlay of vectors corresponding to the running scenario. ; Data transformation model The architecture employs a "dual encoder-single decoder" structure, specifically including: First encoder It employs a bidirectional long short-term memory network, with independent state datasets as input and device latent features as output. ; Second encoder A multilayer perceptron is used, with vector input. The output is the scene's latent features. ; decoder The decoder is a multilayer perceptron, and the input is the concatenated features. The output is the superposition state composite time series prediction data of all electrical devices in the corresponding scenario. And the total composite timing prediction data at the bus. ; Indicates the first The first superimposed running scenario Individual electrical equipment Real-time runtime data, For the first In a superimposed operating scenario, the bus is at The runtime data at each moment, both have the same structure as same; Train the electrical coupling data conversion model based on the dataset obtained in step S1-1. Independent state datasets, vectors corresponding to each superimposed running scenario, and total composite time series data corresponding to each superimposed running scenario; When training a data transformation model, the objective function Defined as: ; In the above formula, It is the current input vector Statistical data vectors in Replace with total composite time series prediction data Calculated statistical data vector The resulting vector The weighting coefficients are used to balance the two losses; Steps S1-3: Construct the training scene generation policy network The scenario generation policy network is trained based on the data recorded in step S1-1. ; Scene generation strategy network Based on the reinforcement learning training method, the environment interaction process for reinforcement learning is designed as follows: State space: A state is a vector of currently constructed, superimposed running scenarios. Its structure and definition are similar to those of a vector. Consistent; Action space: Actions are defined in two types; type one is a triple. , indicating that the first Each electrical device is added to the device set of the current scene, and a startup time is assigned to it. and stop time Type Two is Special Actions This indicates that the construction of the current scene will be terminated; State transition: when an action is performed After that, the state changed from Updated to The specific update method is as follows: [The text abruptly ends here, likely due to an incomplete sentence or a formatting error.] Individual electrical devices are added to the status. The collection of devices, and Join the state The set of start and stop time configurations is then used to calculate a new statistical data vector. : Retrieve the independent-state composite time-series data corresponding to each electrical device in the current device set. Then, based on the corresponding startup time and closing time ,Will The sequence is shifted and truncated on the time axis to form a set of expected independent running data. Then, all the expected independent running data corresponding to the current set of devices are summed and superimposed at the same time to obtain the expected total time series data. Then based on the expected total time series data Calculate the statistical data vector ; Scene generation strategy network A multilayer perceptron is used, with the current state as the input, and the output layer simultaneously outputs the following three sub-probability distributions: 1) Adding electrical equipment and... The probability distribution is such that if the highest probability corresponds to the index of the electrical device, then that electrical device will be added in the generation action. The generated action is ;2) Startup time and The probability distribution is such that if the highest probability corresponds to a certain candidate time, then that candidate time is used as the start time in the generation action; if the highest probability corresponds to... The generated action is 3) Stop time and The probability distribution is such that if the maximum probability corresponds to a certain candidate time, then that candidate time is used as the stop time in the generated action; if the maximum probability corresponds to... The generated action is ; Reward function: Each time with When constructing the overlay scene after the action ends, based on the final state... Calculate the reward function During training, the parameters of the scene generation policy network are updated based on the reward function. The reward function The calculation method is as follows: ; In the above formula, It is a preset target scene vector. This indicates that the final state has been obtained. The first in the process An intermediate state These are the weighting coefficients; When training the scene generation policy network, the policy gradient method is used and the parameters of the scene generation policy network are updated based on the reward function. Suppose the current round has experienced Given a series of actions and intermediate states, a trajectory is generated. The gradient of the policy in the current round The calculation is as follows: ; In the above formula, Is the policy network in state Down Output Action The probability is obtained by multiplying the maximum probability values of the three sub-probability distributions; It is the baseline value, which is the average of the reward functions from the most recent rounds; Parameters for generating a policy network for a scene The gradient; Update the network parameters of the scenario generation strategy using gradient ascent. ,in The learning rate; Steps S1-4: Generate a training sample set for the classification model based on the data transformation model and the scene generation strategy network; The specific method for steps S1-4 is as follows: Utilize the trained scene generation policy network generate There are three different superimposed running scenarios, and the corresponding vectors are denoted as follows: , For each generated overlay running scenario, its corresponding vector is... The independent state dataset is input into the pre-trained electrical coupling data transformation model. In this process, the total composite time series prediction data for this scenario is obtained. , as training samples corresponding to this superimposed running scenario Meanwhile, the set of devices and the set of start / stop time configurations for the superimposed running scenario are used as training samples. tags Thus, the training sample set is obtained. .
Citation Information
Patent Citations
Non-invasive load monitoring method based on mixed probability label time-varying constraint distribution
CN111598145A
Residential load identification method and device for coupling neural network and dynamic time planning
CN113987910A