Non-intrusive load identification method based on reinforcement learning dynamic simulation

By using reinforcement learning dynamic simulation methods and utilizing electrical coupling data conversion models and scene generation strategy networks, diverse simulation samples of multi-device superimposed operation are generated. This solves the problem of limited recognition accuracy and robustness caused by the infinite combination of device start-up and shutdown timing and electrical coupling effects in existing technologies, and achieves efficient and accurate load identification.

CN121350801AActive Publication Date: 2026-01-16YANTAI DONGFANG WISDOM ELECTRIC

Patent Information

Application Number
CN202511891738.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-16
Publication Date
2026-01-16
Estimated Expiration
2045-12-16

AI Technical Summary

Technical Problem

Existing non-intrusive load identification methods are difficult to train models with sufficient experimental data in practical applications due to the infinite combinations of device start-up and shutdown timing. Furthermore, the electrical coupling effect causes differences between the characteristics of a single device and the characteristics of multiple devices operating simultaneously, which limits the identification accuracy and robustness.

Method used

A dynamic simulation method based on reinforcement learning is adopted. By constructing an electrical coupling data conversion model and a scene generation strategy network, diverse multi-device superimposed simulation samples are generated collaboratively to train a classification model to improve recognition accuracy and generalization ability.

Benefits of technology

It enables the low-cost and high-efficiency construction of load identification models that can accurately adapt to real and complex power consumption scenarios without exhaustive experiments, significantly improving the model's identification accuracy and generalization ability in real complex environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure QLYQS_20
    Figure QLYQS_20
  • Figure QLYQS_29
    Figure QLYQS_29
  • Figure QLYQS_32
    Figure QLYQS_32
Patent Text Reader

Abstract

The invention discloses a non-intrusive load identification method based on reinforcement learning dynamic simulation, and relates to the technical field of power consumer side energy consumption monitoring. The method comprises the following steps: firstly, collecting time sequence data of various types of electric equipment in independent operation and a plurality of superimposed operation scenes; then constructing and training an electrical coupling data conversion model used for learning an electrical characteristic mapping rule between single-device independent operation and multi-device superposition operation; meanwhile, a scene generation strategy network is constructed and trained to automatically generate a diversified multi-device superposition operation scene; then, a massive simulation training sample training classification model close to a real scene is generated through cooperation of the trained conversion model and the strategy network; according to the method, the problems of training data shortage and feature mismatch caused by infinity of equipment start-stop time sequence combination and an electrical coupling effect are effectively solved, and the precision, robustness and generalization ability of a non-intrusive load identification model in an actual complex power utilization environment are improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of power user side energy consumption monitoring, and particularly relates to a non-intrusive load identification method based on reinforcement learning dynamic simulation. BACKGROUND

[0002] Non-intrusive load identification is an intelligent sensing technology for terminal power consumption monitoring, and its core advantage lies in "non-intrusiveness", that is, only data acquisition equipment needs to be deployed at the total entrance side of the user power distribution system, and through high-frequency acquisition of total load power, voltage, current and other electrical quantities, and further calculation of transient, steady-state and harmonic characteristics and other high-order electrical characteristics, the overall power consumption load can be decomposed into the independent running state of a single power consumption device, including the energy consumption contribution, accurate start-stop time and running time of each device. Taking a low-voltage household or commercial user as an example, this technology can finely disassemble the total load curve in a certain period into the specific running track of different power consumption devices, realizing the sensing transition from macro power consumption total amount to micro power consumption classification. This fine disassembly capability has important value for providing high-quality bottom layer data support for power consumption safety early warning, personalized energy-saving strategy formulation, and smart energy management scenarios such as power grid load prediction and demand response, and is one of the key technologies for promoting the transformation of power distribution networks to intelligent and fine operation.

[0003] The current mainstream non-intrusive load identification method relies on the construction of a data-driven classification model, and its solution is usually as follows: first, in a controlled experimental environment, pre-acquire the running data of a single or multiple power consumption devices under a specific start-stop time combination to construct a training sample library; then, construct machine learning models such as neural networks, recurrent neural networks and the like, and train them based on the above samples to learn the mapping relationship from the overall electrical characteristics to the specific device running state; finally, deploy the trained model to the actual application scenario for load decomposition.

[0004] However, the existing model training method based on experimental data still has technical problems in actual engineering application. On the one hand, in actual operation, there are almost infinite permutations of the starting sequence and time interval of different devices, and the workload will increase exponentially with the increase of the number of devices if all possible start-stop time combinations are exhausted by manual experiments to obtain the corresponding training data, which is costly and difficult to achieve, directly restricting the generalization recognition ability of the model for unseen time sequence scenarios. Secondly, the electrical characteristics of a single device running independently are different from those of multiple devices running simultaneously due to coupling factors such as line impedance and voltage fluctuations. Most existing methods rely on single-device independent running data or simple synthetic data to construct the training set, which causes the features learned by the model to mismatch the features in the real superposition scenario, resulting in limited recognition accuracy and robustness of the model in actual complex multi-device concurrent scenarios. The above-mentioned data scarcity problem caused by "infinite time sequence combination" and the feature mismatch problem caused by "electrical coupling" make it extremely difficult to train a high-precision and high-generalization-capability load identification model in an experimental environment. SUMMARY

[0005] The present application provides a non-intrusive load identification method based on reinforcement learning dynamic simulation, which aims to solve the problem of limited model generalization ability and recognition accuracy caused by the infinite combination of device start-stop time sequences and the difference between single-device characteristics and multi-device superposition running characteristics due to electrical coupling effects, so as to construct a load identification model that can accurately adapt to real complex power consumption scenarios without exhaustive experiments, at low cost and high efficiency.

[0006] The technical scheme of the present application is as follows:

[0007] A non-intrusive load identification method based on reinforcement learning dynamic simulation, comprising:

[0008] Step S1: preparing a classification model training sample set;

[0009] Step S2: training a classification model based on the classification model training sample set;

[0010] The input of the classification model is the total composite time sequence data in the superposition running scenario of multiple power consumption devices, and the output is the device set composed of the power consumption devices currently running and the start-stop time of each power consumption device;

[0011] Step S3: collecting total composite time sequence data at the user bus in the actual running scenario, inputting the collected total composite time sequence data into the trained classification model, and obtaining the recognition result;

[0012] In step S1, the electrical coupling data conversion model and the scene generation strategy network are generated based on the electrical coupling data conversion model and the scene generation strategy network, and the simulation sample data of the multiple devices are stacked and run cooperatively to constitute the training sample set of the classification model.

[0013] As a further improvement of the non-intrusive load identification method based on reinforcement learning dynamic simulation, step S1 specifically includes the following sub-steps:

[0014] Step S1-1, record the running data of the electrical equipment in different running states, including the data of each electrical equipment running independently, and the data of the electrical equipment running in multiple different stacked running scenes;

[0015] Step S1-2, construct an electrical coupling data conversion model , and train the electrical coupling data conversion model based on the data recorded in step S1-1 ;

[0016] Step S1-3, construct a training scene generation strategy network , and train the scene generation strategy network based on the data recorded in step S1-1 ;

[0017] Step S1-4, generate a classification model training sample set based on the data conversion model and the scene generation strategy network.

[0018] As a further improvement of the non-intrusive load identification method based on reinforcement learning dynamic simulation: in step S1-1, record independent state composite time sequence data of a typical electrical equipment running independently;

[0019] For the th electrical equipment, , control it to start alone, run to stop, and collect data through a recording device during this process, and its independent state composite time sequence data is denoted as , is the preset maximum sampling sequence length, wherein the running data at any moment has the following structure: ; in the above formula, denotes the power of the th electrical equipment at the sampling moment , denotes the voltage data of the th electrical equipment at the sampling moment , denotes the current data of the th electrical equipment at the sampling moment , denotes the power of the The current harmonic amplitude data of the electric device at the sampling moment .

[0020] As a further improvement of the non-intrusive load identification method based on reinforcement learning dynamic simulation: in step S1-1, a plurality of different superimposed operation scenarios are constructed, and data is recorded under each scenario;

[0021] For the first scenario, , first construct a device set and a start-stop time configuration set ; wherein, represents the start time of the first electric device in this scenario, represents the closing time;

[0022] According to the device set and the start-stop time configuration set, a plurality of electric devices are combined to operate, and two types of data are collected: one is the superimposed state composite timing data collected at each participating electric device , is the operation data collected at the moment , the structure is the same as ; the second is the total composite timing data of all participating electric devices superimposed at the bus , is the operation data collected at the moment , the structure is the same as ;

[0023] According to the total composite timing data collected at the bus, the statistical data vector of the first scenario is calculated;

[0024] The complete information of the first superimposed operation scenario is recorded as vector .

[0025] As a further improvement of the non-intrusive load identification method based on reinforcement learning dynamic simulation, the way to calculate the statistical data vector of a scenario according to a timing data in the scenario is:

[0026]

[0027] In the above formula, represents the mean of the power part in the timing data, represents the standard deviation of the power part in the timing data,​​ This represents the local variance of the power component in the time series data, calculated using a sliding window. and These represent the maximum and minimum values ​​of the power component in the time series data, respectively. This represents the power factor calculated based on the power, voltage, and current in the time series data.

[0028] As a further improvement to the aforementioned non-intrusive load identification method based on reinforcement learning dynamic simulation: a data transformation model. The input consists of two parts: the first part is the independent state dataset. The second part is the overlay of vectors corresponding to the running scenario. ;

[0029] Data transformation model The architecture employs a "dual encoder-single decoder" structure, specifically including:

[0030] First encoder It employs a bidirectional long short-term memory network, with independent state datasets as input and device latent features as output. ;

[0031] Second encoder A multilayer perceptron is used, with vector input. The output is the scene's latent features. ;

[0032] decoder The decoder is a multilayer perceptron, and the input is the concatenated features. The output is the superposition state composite time series prediction data of all electrical devices in the corresponding scenario. And the total composite timing prediction data at the bus. ; Indicates the first The first superimposed running scenario Individual electrical equipment Real-time runtime data, For the first In a superimposed operating scenario, the bus is at The runtime data at each moment, both have the same structure as same.

[0033] As a further improvement to the aforementioned non-intrusive load identification method based on reinforcement learning dynamic simulation: training an electrical coupling data conversion model. Based on the dataset obtained in step S1-1: independent state dataset, vectors corresponding to each superimposed running scenario, and total composite time series data corresponding to each superimposed running scenario;

[0034] When training the data conversion model, the objective function is defined as:

[0035] ;

[0036] In the above formula, is a vector obtained by replacing the statistical data vector in the current input vector with a statistical data vector calculated based on the total composite timing prediction data , is a weight coefficient for balancing the two losses.

[0037] As a further improvement of the non-intrusive load identification method based on reinforcement learning dynamic simulation, a scene generation strategy network is provided obtained based on reinforcement learning training, and the environment interaction process of reinforcement learning is designed as follows:

[0038] State space: the state is a vector of the constructed superimposed running scene, and its structure and definition are consistent with the vector ;

[0039] Action space: the action is defined as two types; type one is a triple , indicating that the first power equipment is added to the device set of the current scene, and the start time and stop time are assigned to it; type two is a special action , indicating the termination of the construction of the current scene;

[0040] State transition: after executing the action , the state is updated from to ; the specific update method is: the first power equipment is added to the device set of the state , and is added to the start-stop time configuration set of the state , and then a new statistical data vector is calculated: the independent state composite timing data corresponding to each power equipment in the current device set is obtained , and then the sequence is shifted and truncated on the time axis according to the corresponding start time and stop time to form an expected independent running data, and then all the expected independent running data corresponding to the current device set are summed at the same time to complete superposition, and the expected total timing data Then based on the expected total time series data Calculate the statistical data vector ;

[0041] Scene generation strategy network A multilayer perceptron is used, with the current state as the input, and the output layer simultaneously outputs the following three sub-probability distributions: 1) Adding electrical equipment and... The probability distribution is such that if the highest probability corresponds to the index of the electrical device, then that electrical device will be added in the generation action. The generated action is ;2) Startup time and The probability distribution is such that if the highest probability corresponds to a certain candidate time, then that candidate time is used as the start time in the generation action; if the highest probability corresponds to... The generated action is 3) Stop time and The probability distribution is such that if the maximum probability corresponds to a certain candidate time, then that candidate time is used as the stop time in the generated action; if the maximum probability corresponds to... The generated action is ;

[0042] Reward function: Each time with When constructing the overlay scene after the action ends, based on the final state... Calculate the reward function During training, the parameters of the scene generation policy network are updated based on the reward function.

[0043] As a further improvement to the non-intrusive load identification method based on reinforcement learning dynamic simulation, the reward function The calculation method is as follows:

[0044] ;

[0045] In the above formula, It is a preset target scene vector. This indicates that the final state has been obtained. The first in the process An intermediate state These are the weighting coefficients;

[0046] When training the scene generation policy network, the policy gradient method is used and the parameters of the scene generation policy network are updated based on the reward function. Suppose the current round has experienced Given a series of actions and intermediate states, a trajectory is generated. The gradient of the policy in the current round The calculation is as follows:

[0047] ;

[0048] In the above formula, is the probability of the policy network outputting the action in state , which is obtained by multiplying the maximum probability value in the three sub-probability distributions; is a baseline value, which takes the average value of the reward function in the last several rounds; is the gradient of the parameters of the scene generation policy network for the scene;

[0049] The scene generation policy network parameters are updated by gradient ascent , wherein is the learning rate.

[0050] As a further improvement of the non-intrusive load identification method based on reinforcement learning dynamic simulation, the specific manner of step S1-4 is: using the trained scene generation policy network to generate different superimposed running scenes, and the corresponding vectors are denoted as ; for each generated superimposed running scene, input the corresponding vector and the independent state data set into the trained electrical coupling data conversion model to obtain the total composite time series prediction data under the scene, which is taken as the training sample corresponding to the superimposed running scene, and meanwhile, the device set and the start-stop time configuration set of the superimposed running scene are taken as the label of the training sample , so as to obtain the training sample set .

[0051] Compared with the prior art, the present application has the following beneficial effects:

[0052] 1. The present application introduces an electrical coupling data conversion model and a scene generation policy network, and constructs a complete closed-loop process of "data acquisition-conversion model training-simulation optimization-classification model training-accurate identification". This method first acquires and learns the electrical coupling law between single-device independent running and multi-device superimposed running states, realizes the equivalent conversion from single-device data to multi-device superimposed scenes; then through the reinforcement learning driven dynamic simulation process, the superimposed running scenes can be generated without manually enumerating all start-stop time combinations, and then a large number of training samples close to reality and rich in diversity are obtained; finally, the classification model is trained using these high-quality simulation samples, thereby significantly improving the recognition accuracy and generalization ability of the model in the actual complex power consumption environment.

[0053] 2. The application trains an electrically coupled data conversion model of a "double-encoder-single-decoder" structure to accurately learn and simulate the changes in electrical characteristics caused by factors such as line impedance and voltage fluctuations when multiple devices are running simultaneously. The model takes the time series characteristics of a single device running independently and the scenario vector containing device combinations, start-stop times, and statistical data as input, and outputs the predicted superimposed running total time series data. This design enables the model to effectively capture the electrical coupling effect, solving the problem of reduced recognition performance caused by the mismatch between training data (single device data) and real application scenarios (multiple device superimposed data) in traditional methods, ensuring the input quality of subsequent classification model training from the data source.

[0054] 3. The application also includes a scenario generation strategy network based on reinforcement learning, which is used to automatically generate diverse multiple device superimposed running scenarios. The strategy network takes the scenario vector as the state and adds devices and specifies their start-stop times or terminates scenario construction as the action. By defining a reward function that takes into account the "realism" and "novelty" of generated samples, the network is guided to explore and generate a large number of simulation scenarios with high coverage and high diversity in start-stop time series combinations. This mechanism fundamentally overcomes the bottleneck of traditional methods that cannot collect training data through exhaustive experiments due to the almost infinite combinations of device start-stop times, expanding the training sample library of the classification model and enhancing the model's ability to adapt to unseen patterns.

[0055] 4. The application uses the vast amount of samples generated by the above-mentioned collaborative mechanism, including high-quality simulation total time series data and their corresponding scenario labels (device set and start-stop time of each device), to train the classification model. The trained classification model can more accurately learn the mapping relationship from complex superimposed electrical characteristics to specific device operating states. In actual deployment, this method can achieve high-precision and high-robustness recognition of the composition of running devices and their accurate start-stop times based only on electrical monitoring data at the user's total entry side, effectively supporting applications such as electricity safety warning, fine energy consumption statistics, and smart energy management. DETAILED DESCRIPTION

[0056] The technical solutions of the application will be described in detail below. Obviously, the described embodiments are only a part of the embodiments of the application, not all embodiments.

[0057] A non-intrusive load identification method based on reinforcement learning dynamic simulation. This method learns the data difference rules of devices in independent and superimposed running states by constructing an electrically coupled data conversion model, and uses a reinforcement learning driven scenario generation strategy network to generate a large number of, diverse, and realistic physical law superimposed running simulation samples at low cost and high efficiency, thereby solving the problems of insufficient, distorted, and poor generalization of training samples caused by the infinite combinations of device start-stop times and electrical coupling effects in traditional methods.

[0058] The method specifically comprises:

[0059] Step S1: preparing a classification model training sample set, which is different from the conventional training sample preparation work in that this step cooperates with the electrical coupling data conversion model and the scene generation strategy network to jointly generate multi-device superimposed operation simulation sample data with high diversity and high fidelity to constitute the training sample set of the classification model.

[0060] Step S1 specifically comprises the following sub-steps.

[0061] Step S1-1, recording the running data of the electrical equipment in different running states, including the data of each electrical equipment running independently, and the data of the electrical equipment running in multiple different superimposed running scenes. This step aims to obtain the basic real data for training the subsequent model.

[0062] First, record the independent state composite time sequence data of a typical electrical equipment running independently. For the first electrical equipment (e.g. ), control it to start alone, run to stop, and collect data through a recording device during the process, and the independent state composite time sequence data is denoted as , is a preset maximum sampling sequence length, and the running data at any time has the following structure:

[0063] ;

[0064] In the above formula, represents the power of the first electrical equipment at the sampling time , represents the voltage data, represents the current data, represents the current harmonic amplitude data.

[0065] is used to unify the data format, for example, set to cover 1000 sampling points of the typical start-stop process of the equipment. For the part with an effective length less than , fill it with zero values.

[0066] Then, construct different superimposed running scenes and record data in each scene.

[0067] For the first scene (e.g. ), first construct a device set​ and start / stop time configuration set .in, Indicates the first Startup time of the electrical equipment in this scenario (sampling point index). Indicates its closing time (sampling point index).

[0068] In the experimental environment, multiple electrical devices were configured to operate in combination according to the device set and start / stop time, while two types of data were collected: one was for each participating electrical device. ( Superposition-state composite time-series data acquired by the branch where it is located , for The structure and operation data collected in real time The two methods are the same, but the length is determined by the actual start and stop times; secondly, the total composite timing data of all participating electrical devices is collected at the bus. , for The structure and operation data collected in real time The same, but the length is determined by the start and stop time of the entire scene. and The difference reflects the electrical coupling effect.

[0069] Furthermore, based on the total composite timing data collected on the bus Calculate the first Statistical data vectors for each scenario :

[0070] ;

[0071] In the above formula, Represents total composite time series data The average value of the medium power portion, Represents total composite time series data The standard deviation of the medium power component, Represents total composite time series data The local variance of the medium-power component is calculated using a sliding window. and These represent the total composite time series data. The maximum and minimum values ​​of the medium power portion, Indicates based on total composite time series data The power factor is calculated from the power, voltage, and current in the data. These statistics describe the overall electrical characteristics of the superimposed operation. This calculation process is also applicable to other timing data with the same structure, which will not be elaborated further below.

[0072] The first The complete information of each superimposed running scenario is denoted as a vector. .

[0073] Step S1-2: Construct and train the electrical coupling data conversion model .

[0074] Electrically Coupled Data Conversion Model The goal is to learn the mapping pattern from independent device operation data and scenario definition to the superimposed operation data of various electrical devices in that scenario, that is, to approximate the real physical coupling function that is difficult to model directly with a parameterized model. Model training is based on the dataset obtained in step S1-1: the independent state dataset. Complete information on each overlay operation scenario and the total composite time series data corresponding to each overlay operation scenario.

[0075] Data transformation model The input consists of two parts: the first part is the independent state dataset, which has been padded with zeros to make it uniform. The first part is a two-dimensional array; the second part is a vector superimposed on the running scenario. ,vector This is obtained by vectorizing the complete information of the superimposed running scenario: First, construct a vector of length... Binary vector representation of device set If a certain electrical device is in the device set, the corresponding position is 1; otherwise, it is 0. Secondly, for each electrical device... Set its start and stop times Arranged in order, for those not in The start and stop times of the devices in the vector are represented by specific placeholders (such as -1). Finally, the device set vector, start and stop time vector, and statistics vector are combined. Concatenate them into a one-dimensional vector as The final representation.

[0076] Data transformation model The architecture employs a "dual encoder-single decoder" structure, specifically including:

[0077] First encoder The system employs a Bidirectional Long Short-Term Memory (Bi-LSTM) network. The input is an independent-state dataset, which passes through two hidden layers (e.g., dimensions 1000->512, 512->512) and one linear layer, outputting the device's latent features. .

[0078] Second encoder The system employs a multilayer perceptron (MLP), with vector input. (length is) ), the output scene hidden features after the hidden layer and linear layer .

[0079] Decoder : the decoder is a multi-layer perceptron, the input is the spliced features , and the output is the superposition state composite time series prediction data of all electrical equipment under the corresponding scene and the total composite time series prediction data at the bus . represents the running data of the th electrical equipment under the th superposition running scene at the th moment, is the running data of the bus at the th moment under the th superposition running scene, and the structures of both are the same as . This method pays more attention to whether the total composite time series prediction data approximates the actually collected data, and the superposition state composite time series prediction data of the electrical equipment does not participate in the subsequent calculation.

[0080] When training the data conversion model, the objective function is defined as:

[0081]

[0082] In the above formula, the first term is the mean square error (MSE) between the predicted total data and the real total data. The second term is the constraint term, where is the vector obtained by replacing the statistical data vector in the current input vector with the statistical data vector calculated based on the total composite time series prediction data , and this term uses the L1 norm to constrain the statistical characteristics of the generated data to be consistent with the real data. is the weight coefficient for balancing the two loss terms. The parameters of the data conversion model are optimized by the gradient descent method (such as the Adam optimizer) to minimize , so that the data conversion model can accurately simulate the electrical coupling effect.

[0083] Step S1-3, construct and train the scene generation strategy network based on reinforcement learning .

[0084] The scene generation strategy network is a reinforcement learning agent, and its goal is to learn to automatically generate diverse and reasonable superposition running scenes for subsequent simulation sample generation. The environment interaction process of reinforcement learning is designed as follows:

[0085] State space: State is the vector of the current constructed superposition running scenario, its structure and definition are consistent with the vector .

[0086] Action space: Action is defined as two types. Type one is triple , which represents adding the th electrical equipment to the device set of the current scenario and assigning it start time and stop time . Type two is special action , which represents terminating the construction of the current scenario.

[0087] State transition: After executing action , state is updated from to . The specific update method is: add the th electrical equipment to the device set of state , and add to the start-stop time configuration set of state , then calculate the new statistical data vector : get the independent state composite time series data corresponding to each electrical equipment in the current device set , then according to the corresponding start time and close time , shift and truncate the sequence on the time axis (to match the start-stop time) to form an expected independent running data, then sum all the expected independent running data corresponding to the current device set at the same time to complete the superposition, get the expected total time series data , then calculate the statistical data vector based on the expected total time series data . The above operation is based on an assumption: under the ideal linear superposition approximation, the total electrical signal when multiple electrical equipment runs simultaneously is approximately equal to the algebraic sum of the independent running signals of each electrical equipment, although this is different from the real physical coupling (the aforementioned electrical coupling effect), but it is efficient and feasible for the rapid and continuous evaluation of the "reasonableness" of the scenario in the reinforcement learning strategy network training process.

[0088] Scenario generation strategy network : implemented by a three-layer multilayer perceptron. The network input is the current state, and the output layer is designed to output the following three sub-probability distributions simultaneously: 1) the probability distribution of adding electrical equipment and , if the maximum probability corresponds to the index of the electrical equipment, then the electrical equipment will be added as the electrical equipment to be added in the generated action, if the maximum probability corresponds to , then the generated action is ; 2) the start time is sampled from the probability distribution of , if the maximum probability corresponds to a candidate time, then the candidate time is taken as the start time in the generated action, if the maximum probability corresponds to , then the generated action is ; 3) the stop time is sampled from the probability distribution of , if the maximum probability corresponds to a candidate time, then the candidate time is taken as the stop time in the generated action, if the maximum probability corresponds to , then the generated action is . In actual sampling, sampling is performed in turn according to these distributions to determine the specific action.

[0089] Reward function: each time the construction of the running scene is ended with action, the reward function is calculated according to the final state :

[0090] ;

[0091] In the above formula, is a preset target scene vector, represents the intermediate state in the process of obtaining the final state , is a weight coefficient. The first term represents the completion degree of the current round of training, and the second term represents the minimum distance between the new state and the historical state.

[0092] When training the scene generation policy network, the policy gradient method (such as the REINFORCE algorithm) is used to update the parameters of the scene generation policy network. Assuming that actions and intermediate states are experienced in the current round, and the generated trajectory is , then the current round policy gradient is calculated as follows:

[0093] ;

[0094] In the above formula, is the probability of the policy network outputting the action under the state , which is obtained by multiplying the maximum probability values in the three sub-probability distributions. is a baseline value used to reduce training variance, and is usually the average of the rewards obtained in the last several rounds (for example, 5 rounds). is the gradient of the parameters of the scene generation policy network. The scene generation policy network parameters are updated by gradient ascent wherein is a learning rate.

[0095] After being fully trained, the scene generation strategy network can efficiently generate a large number of superimposed operation scene definitions that are both consistent with the target electrical characteristics (realistic) and significantly different from each other (diverse).

[0096] Step S1-4, based on the data conversion model and the scene generation strategy network, generate a set of classification model training samples.

[0097] The specific way is: using the trained scene generation strategy network to generate different superimposed operation scenes, and the corresponding vector is denoted as wherein . For each generated superimposed operation scene, input its corresponding vector and the independent state data set into the trained electrical coupling data conversion model , to obtain the total composite time sequence prediction data under this scene as the training sample corresponding to this superimposed operation scene, and at the same time, the device set and the start-stop time configuration set of this superimposed operation scene are taken as the label of the training sample .

[0098] Thus, a large-scale and high-quality training sample set for training the final load decomposition classification model can be automatically constructed .

[0099] Step S2: train a classification model based on the set of classification model training samples.

[0100] This step trains a load decomposition classification model (for example, based on a deep neural network) in a supervised learning manner. The input of the model is the total composite time sequence data under the superimposed operation scene of multiple electrical equipment, and the output is the device set composed of the currently running electrical equipment and the start-stop time of each electrical equipment.

[0101] The specific process is: divide the large-scale sample set generated in step S1-4 into a training set and a validation set. Construct a classification model (such as a sequence-to-sequence model using an encoder-decoder framework, or a model based on an attention mechanism, the classification model belongs to the prior art and is not the innovation point of the present application, and will not be described here) with as the input, and the predicted output of the classification model is denoted as . Define the loss function (such as cross-entropy loss or mean square error loss, which is determined by and (Computed), the parameters of the classification model are iteratively optimized by back propagation and gradient descent algorithm (such as Adam) to minimize , so that the model can accurately analyze the running information of each independent device from the superimposed signal. Since the training data comes from the simulation model that has learned the real electrical coupling law and covers a large number of diversified device combinations and start-stop time sequences, the trained classification model has strong generalization ability and actual scene adaptability.

[0102] Step S3: Collecting total composite time sequence data at the user bus in the actual running scene, inputting the collected total composite time sequence data into the trained classification model to obtain the recognition result, i.e., the device set composed of the currently running electrical equipment and the start-stop time of each electrical equipment.

[0103] In the actual application deployment stage, a non-intrusive load monitoring device is installed at the total entrance of the power distribution system of the user to be monitored (such as a family, a commercial building). The device collects the voltage and current raw waveforms of the bus at a high frequency (for example, thousands of points per second). The collected waveform data is processed to calculate the total power , voltage , current and harmonic composite time sequence data, forming the input to be identified. The is input into the classification model which has been fully trained and solidified in step S2. The classification model automatically analyzes and calculates and outputs the recognition result, thereby realizing the accurate and non-intrusive decomposition from the total power "black box" to the device power "white box", providing a reliable data basis for energy efficiency analysis, abnormal power consumption detection, demand response and other applications.

[0104] It should be noted that for those skilled in the art, it is obvious that the present application is not limited to the details of the above exemplary embodiments, and the present application can be implemented in other specific forms without departing from the spirit or essential characteristics of the present application. The scope of the present application is defined by the claims rather than the above description.

Claims

1. A non-intrusive load identification method based on reinforcement learning dynamic simulation, comprising: Step S1: preparing a classification model training sample set; Step S2: training a classification model based on the classification model training sample set; The input of the classification model is the total composite time series data under the superimposed operation scene of multiple electrical equipment, and the output is the equipment set composed of the currently operated electrical equipment and the start-stop time of each electrical equipment; Step S3: collecting total composite time series data at the user bus in the actual operation scene, inputting the collected total composite time series data into the trained classification model, and obtaining the identification result; Characterized in that: In step S1, based on the electrical coupling data conversion model and the scene generation strategy network, multiple device superimposed operation simulation sample data are cooperated to constitute the training sample set of the classification model.

2. The non-intrusive load identification method based on reinforcement learning dynamic simulation of claim 1, wherein, Step S1 specifically includes the following sub-steps: Step S1-1, record the operation data of electrical equipment in different operation states, including the data of each electrical equipment operating independently, and the data of electrical equipment operating in multiple different superimposed operation scenes; Step S1-2, constructing an electrical coupling data conversion model and training the electrical coupling data conversion model based on the data recorded in step S1-1 ; Step S1-3, constructing a training scene generation policy network and training the scene generation policy network based on the data recorded in step S1-1 ; Step S1-4, based on the data conversion model and the scene generation strategy network, generate a classification model training sample set.

3. The non-intrusive load identification method based on reinforcement learning dynamic simulation of claim 2, wherein: In step S1-1, the recording of the individual state composite time series data of the individual typical electrical equipment running independently; For the first power equipment, , control its single start, run to stop, and in the process through the recording device to collect data, its independent state composite timing data is recorded as , is the preset maximum sampling sequence length, wherein any moment of running data structure is: ; in the formula, indicates the power of the first power equipment at the sampling time , indicates the voltage data of the first power equipment at the sampling time , indicates the current data of the first power equipment at the sampling time , indicates the current harmonic amplitude data of the first power equipment at the sampling time .

4. The non-intrusive load identification method based on reinforcement learning dynamic simulation of claim 3, wherein: In step S1-1, a plurality of different superposition operation scenarios are constructed, and data is recorded under each scenario respectively. In step S1-1, a plurality of different superposition operation scenarios are constructed, and data is recorded under each scenario respectively. For the first scenario, , first construct a device set and a start-stop time configuration set ; wherein, represents the start time of the th electric device in this scenario, represents its shutdown time; According to the device set and the start-stop time configuration set, a plurality of electric device groups are controlled to operate in combination, and two types of data are collected: one is superimposed state complex timing data collected at each branch of the participating electric device , , ​​​​​​​ According to the total composite timing data collected on the bus Computing the statistical data vector of the nth scenario ; Let the complete information of the nth superposition running scenario be recorded as vector . .

5. The non-intrusive load identification method based on reinforcement learning dynamic simulation of claim 4, wherein, According to the first way, the statistical data vector of the scene is calculated according to the time sequence data of the scene . ; In the above formulae, denotes the mean of the power part in the time series data, denotes the standard deviation of the power part in the time series data, denotes the local variance of the power part in the time series data calculated by a sliding window, and denote the maximum and minimum of the power part in the time series data, respectively, denotes the power factor calculated based on the power, voltage and current in the time series data.

6. The non-intrusive load identification method based on reinforcement learning dynamic simulation of claim 4, wherein: Data conversion model The input includes two parts: the first part is an independent state data set , and the second part is a vector corresponding to a superimposed operation scene ; Data conversion model Adopting a "double-encoder-single-decoder" structure, specifically comprising: First encoder : bidirectional long short-term memory network, input is independent state data set, output is device hidden feature ; Second encoder : using a multi-layer perceptron with input being a vector , output being a scene latent representation ; Decoder : the decoder is a multi-layer perception, the input is the spliced features , the output is the superposition state composite time series prediction data of all electrical equipment in the corresponding scene and the total composite time series prediction data at the bus ; represents the running data of the th electrical equipment in the th superposition running scene at the th moment, is the running data of the bus at the th moment in the th superposition running scene, both structures are the same as .

7. The non-intrusive load identification method based on reinforcement learning dynamic simulation of claim 6, wherein: Training an electrically coupled data conversion model Based on the data set obtained in step S1-1: the independent state data set, the vector corresponding to each superposition running scene, and the total composite timing data corresponding to each superposition running scene; When training the data conversion model, the objective function is defined as: ; In the above formula, It is the current input vector Statistical data vectors in Replace with total composite time series prediction data Calculated statistical data vector The resulting vector The weighting coefficients are used to balance the two losses.

8. The non-intrusive load identification method based on reinforcement learning dynamic simulation of claim 4, wherein: Scene generation policy network The training method based on reinforcement learning is obtained, and the environment interaction process of the reinforcement learning is designed as follows: State space: The state is a vector of the currently constructed superposition of running scenarios, whose structure and definition are consistent with the vector of the current state. Action space: Actions are defined as two types; Type one is a triple , which represents adding the th electrical device to the device set of the current scene and assigning it a start time and a stop time ; Type two is a special action , which represents terminating the construction of the current scene; State transition: when action is executed After, state is updated from to ; The specific updating method is: adding the first power equipment to the state equipment set, adding to the start-stop time configuration set in the state , and then calculating a new statistical data vector : obtaining the independent state composite time sequence data corresponding to each power equipment in the current equipment set , and then shifting and truncating the sequence on the time axis according to the corresponding start time and the closing time to form an expected independent running data, superimposing all the expected independent running data corresponding to the current equipment set at the same time to obtain expected total time sequence data , and then calculating the statistical data vector based on the expected total time sequence data ; Scenario generation policy network : adopt multi-layer perception, input is current state, output layer simultaneously outputs following three sub-probability distribution: 1) probability distribution of adding electrical equipment and , if the maximum probability corresponds to the index of electrical equipment, the electrical equipment is taken as the electrical equipment to be added in the generated action, if the maximum probability corresponds to , the generated action is ; 2) probability distribution of start time and , if the maximum probability corresponds to a candidate time, the candidate time is taken as the start time in the generated action, if the maximum probability corresponds to , the generated action is ; 3) probability distribution of stop time and , if the maximum probability corresponds to a candidate time, the candidate time is taken as the stop time in the generated action, if the maximum probability corresponds to , the generated action is ; Reward function: each time with At the end of the action, the construction of the superimposed running scene is based on the final state Calculate the reward function ; Update the parameters of the scene generation policy network based on the reward function during training.

9. The non-intrusive load identification method based on reinforcement learning dynamic simulation of claim 8, wherein, The reward function is calculated as follows: ; In the above formula, is a preset target scene vector, represents the first intermediate state in the process of obtaining the final state , and is a weight coefficient. When training the scene generation policy network, a policy gradient method is adopted and parameters of the scene generation policy network are updated based on a reward function : Suppose that the current round experiences actions and intermediate states, and the generated trajectory is , then the current round policy gradient is calculated as follows: ; In the above formula, Is the policy network in state Down Output Action The probability is obtained by multiplying the maximum probability values ​​of the three sub-probability distributions; It is the baseline value, which is the average of the reward functions from the most recent rounds; Parameters for generating a policy network for a scene The gradient; updating the scene generation policy network parameters by gradient ascent wherein is a learning rate.

10. The non-intrusive load identification method based on reinforcement learning dynamic simulation of claim 4, wherein, The specific method for steps S1-4 is as follows: Utilize the trained scene generation policy network generate There are three different superimposed running scenarios, and the corresponding vectors are denoted as follows: For each generated overlay running scenario, its corresponding vector is... The independent state dataset is input into the pre-trained electrical coupling data transformation model. In this process, the total composite time series prediction data for this scenario is obtained. , as training samples corresponding to this superimposed running scenario Meanwhile, the set of devices and the set of start / stop time configurations for the superimposed running scenario are used as training samples. tags Thus, the training sample set is obtained. .

Citation Information

Patent Citations

  • Non-invasive load monitoring method based on mixed probability label time-varying constraint distribution

    CN111598145A

  • Residential load identification method and device for coupling neural network and dynamic time planning

    CN113987910A

  • Non-intrusive load monitoring method based on zero sample learning

    CN114113773A

  • Non-intrusive electrical load monitoring method based on FedProx federated learning and GRU

    CN115952907A

  • Sequence-to-subsequence non-intrusive load identification method and device and storage medium

    CN117493929A

Cited By

  • Method for calculating dynamic load power of electric energy meter at autonomous controllable gateway

    CN122218301A