A User-Distributed Resource Energy Management Method Based on Intelligent Decision Control
By constructing a user-distributed resource energy management method based on intelligent decision control, and utilizing deep reinforcement learning and the Internet of Things platform, the problems of low resource allocation efficiency and untimely regulation in the power system are solved, achieving efficient and low-cost energy management.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- SOUTHEAST UNIV
- Filing Date
- 2022-10-17
- Publication Date
- 2026-04-24
AI Technical Summary
The current power system's resource allocation and pricing are still based on the traditional hierarchical top-down approach, which leads to problems such as high energy interaction costs, low efficiency, and untimely regulation response.
A user-distributed resource energy management method based on intelligent decision control is constructed. Deep reinforcement learning and Markov decision process (MDP) models are adopted, combined with an Internet of Things platform. The model is trained by reinforcement learning algorithm, and Raspberry Pi 4B is used as the hardware basic unit for decision control.
It enables intelligent decision-making for distributed resources, improves the efficiency and response speed of energy management, reduces energy interaction costs, and enhances the precision of resource regulation.
Smart Images

Figure CN115587539B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of power system demand response technology, specifically relating to a user-distributed resource energy management method based on intelligent decision control. Background Technology
[0002] Greenhouse gas emissions are causing environmental pollution and climate change worldwide, and fossil fuel resources are experiencing a persistent shortage. These issues are driving the power system to seek new energy management strategies. Currently, although the power system is moving towards more decentralized management, the allocation and pricing of electricity resources in the electricity market are still based on the traditional hierarchical, top-down approach of power system management. This makes producers and consumers passive recipients, resulting in high energy interaction costs, low efficiency, and untimely regulatory responses. Summary of the Invention
[0003] In view of the shortcomings of the prior art, the purpose of this invention is to provide a user-distributed resource energy management method based on intelligent decision control, so as to solve the problems mentioned in the background art.
[0004] The objective of this invention can be achieved through the following technical solutions:
[0005] A user-distributed resource energy management method based on intelligent decision control includes:
[0006] Step 1: Construct an intelligent decision-making framework for the distributed system;
[0007] Step 2: Based on the intelligent decision-making framework obtained in Step 1, construct a distributed system intelligent decision-making model based on deep reinforcement learning and MDP, and train it using reinforcement learning algorithms. After training, save the trained model.
[0008] Step 3: Build an IoT-based hardware and software application platform and call the trained model from Step 2 to make decisions.
[0009] Preferably, the intelligent decision-making framework of the distributed system constructed in step one includes a power grid company, an EMS, an IoT control terminal, and N producers and consumers.
[0010] Preferably, the EMS interacts with the grid side to exchange electricity price and power data, and interacts with producers and consumers as well as with each other to exchange local market electricity price and electricity price data. The EMS generates local market electricity price based on the grid company's electricity price data and the energy demand data of producers and consumers. The Internet of Things control terminal transmits control signals to producers and consumers through artificial intelligence technology to regulate the relevant electrical equipment.
[0011] The producer-consumer unit contains electrical equipment, including fixed loads, distributed photovoltaics, electric vehicles in distributed energy storage, air conditioners and electric water heaters in adjustable loads, and washing machines in deferred loads.
[0012] Preferably, the MDP components in step two consist of the agent's state space S, action space A, reward function R, and state transition functions P for air conditioners, water heaters, electric vehicles, and washing machines.
[0013] Preferably, the state transition function P is as follows:
[0014]
[0015]
[0016]
[0017]
[0018] The following constraints must also be met:
[0019]
[0020]
[0021] SOC min ≤SOC t ≤SOC max (7)
[0022] SOC α ≥SOC H (8)
[0023] α wash ≤t start ≤t≤t end ≤β wash (9)
[0024] t end -t start +1=T set (10)
[0025] β wash -α wash ≥T set (11)
[0026]
[0027] Preferably, the state space S and action space A are as follows:
[0028]
[0029]
[0030] Preferably, in addition to the state quantities of the aforementioned air conditioners, electric water heaters, electric vehicles, and washing machines, the state space S also includes time, local market electricity prices, photovoltaic output, and fixed load status.
[0031] Preferably, the reward function R is as follows:
[0032]
[0033]
[0034]
[0035]
[0036]
[0037]
[0038]
[0039] Preferably, in step three, the basic unit of the IoT hardware platform is a Raspberry Pi 4B. Based on the Respbain system, a deep reinforcement learning environment is built using Tensorflow and Gym libraries. One Raspberry Pi 4B serves as the overall energy management center, and each of the remaining Raspberry Pi 4B corresponds to an energy management terminal for a prosumer, performing decision control on the corresponding prosumer. The Raspberry Pis establish an information transmission path through a Wi-Fi module to exchange various data and decision information. Through the built Raspberry Pi platform, the intelligent decision model framework is not repeatedly constructed, and the saved training model is directly run.
[0040] Preferably, a user-distributed resource energy manager based on intelligent decision control stores a program for running a user-distributed resource energy management method based on intelligent decision control.
[0041] The beneficial effects of this invention are:
[0042] 1. This invention proposes a user-distributed resource energy management method based on reinforcement learning intelligent decision control. This method can better handle the intelligent decision-making problem of local consumption of user-side distributed resources, effectively making up for the shortcomings of low intelligence level of distribution network demand-side resource decision-making and insufficient precision of control methods. At the same time, the built software and hardware Internet of Things platform can better match the training model for decision-making, and has practicality and reliability. Attached Figure Description
[0043] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0044] Figure 1 This is a flowchart of the solution method of the present invention;
[0045] Figure 2 This is a diagram of the intelligent decision-making framework of the distributed transaction system in this invention;
[0046] Figure 3 This is a framework diagram of the Internet of Things platform in this invention. Detailed Implementation
[0047] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0048] Please see Figures 1 to 3 As shown, a user-distributed resource energy management method based on intelligent decision control includes the following steps:
[0049] Step 1: Construct an intelligent decision-making framework for the distributed system;
[0050] The intelligent decision-making framework of the distributed system constructed in this invention includes a power grid company, an energy management system (EMS), an Internet of Things (IoT) control terminal, and N prosumers. The EMS interacts with the power grid side to exchange electricity prices and power consumption data. The EMS also interacts with prosumers and among prosumers to exchange local market electricity prices and electricity consumption data. The EMS generates local market electricity prices based on the power grid company's electricity price data and the energy demand data of prosumers. The IoT control terminal transmits control signals to prosumers through artificial intelligence technology to regulate the corresponding electrical equipment. The electrical equipment within the prosumers includes fixed loads, distributed photovoltaics, electric vehicles in distributed energy storage, air conditioners and electric water heaters in adjustable loads, and washing machines in deferred loads.
[0051] Step 2: Construct a distributed system intelligent decision-making model based on deep reinforcement learning and Markov decision process (MDP), and train it using reinforcement learning algorithms;
[0052] The air conditioners, electric water heaters, electric vehicles, and washing machines in Step 1 are typical loads in a distributed system. These types of loads are the main focus of the intelligent decision-making model for distributed systems. The Markov property states that the state in the next time period depends only on the state and control signals in the current time period, and is independent of the state and actions in previous time periods. The intelligent decision-making model for distributed systems is constructed based on MDP.
[0053] The MDP consists of the agent's state space S, action space A, reward function R, and state transition function P.
[0054] The state transition equations for air conditioners (both heating and cooling), electric water heaters, electric vehicles, and washing machines are as follows:
[0055]
[0056]
[0057]
[0058]
[0059] In the formula, T t in T t water SOC t and Let represent the indoor temperature, electric water heater water temperature, electric vehicle SOC state, and washing machine operating state at time t, respectively. These represent the state variables in the MDP model and must satisfy the following constraints:
[0060]
[0061]
[0062] SOC min ≤SOC t ≤SOC max (7)
[0063] SOC α ≥SOC H (8)
[0064] α wash ≤t start ≤t≤t end ≤β wash (9)
[0065] t end -t start +1=T set (10)
[0066] β wash -α wash ≥T set (11)
[0067]
[0068] Constraints (5), (6), and (7) limit the upper and lower limits of the state quantities of air conditioners, electric water heaters, and electric vehicles; constraint (8) ensures that the electric vehicle's power needs to meet basic commuting requirements; constraint (9) limits the working hours of the washing machine; constraints (10) and (11) ensure that the washing machine's operation is uninterrupted; constraint (12) limits the time when the washing machine finishes its work.
[0069] The state space and action space of the constructed distributed system intelligent decision-making model are as follows:
[0070]
[0071]
[0072] In addition to the state variables of the four types of electrical equipment mentioned above, the state space also includes time, local market electricity prices, photovoltaic output, and fixed load status.
[0073] The reward function of the constructed distributed system intelligent decision-making model is as follows:
[0074] P t s =P t fix -P t pv +P t ac +P t water +P t ch -P t dis +P t wash (15)
[0075]
[0076]
[0077]
[0078]
[0079]
[0080]
[0081] Equations (15) and (16) are the reward functions for total electricity cost; Equations (17) and (18) are the reward functions for comfort of air conditioning and electric water heater, respectively; Equation (19) is the reward function for electric vehicle, taking into account power anxiety and SOC limitation; Equation (20) is the reward function for washing machine, taking into account the difference between the end time of work and the optimal end time.
[0082] Based on the above-mentioned intelligent decision-making model for distributed systems, this invention uses the DQN algorithm in reinforcement learning for training to optimize intelligent decision-making actions, thereby enabling effective energy management of distributed resources.
[0083] Step 3: Build an IoT-based hardware and software application platform and call the model to make decisions;
[0084] This invention uses a Raspberry Pi 4B as the basic unit of an IoT hardware platform. Based on the Respbain system, it utilizes Tensorflow and the Gym library to build a deep reinforcement learning environment. One Raspberry Pi 4B serves as the central energy management hub, while the remaining Raspberry Pi 4Bs each act as an energy management terminal for a prosumer, performing decision-making and control over their respective prosumers. The Raspberry Pis communicate with each other via a Wi-Fi module to exchange various data and decision information. By using the pre-built Raspberry Pi platform, the intelligent decision-making model framework is not repeatedly constructed; instead, the saved model is run directly.
[0085] This invention saves the trained reinforcement learning model to Raspberry Pi hardware and transmits various data information of the corresponding prosumers via the Wi-Fi module, including weather temperature, water consumption at any time, commuting electricity consumption, local market electricity prices, etc. These data information are substituted as part of the state variables into the reinforcement learning environment to make decisions and controls. The decision information made by each prosumer interacts with the energy management center, thereby affecting some data information such as local market electricity prices. Then, in the next decision process, the updated data information is substituted for dynamic decision control. Each Raspberry Pi 4B hardware is configured with a visualization window to display the current status of various electrical devices and the current control decision information in real time.
[0086] In the description of this specification, references to terms such as "an embodiment," "example," "specific example," etc., indicate that a specific feature, structure, material, or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of the invention. In this specification, illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples.
[0087] Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application can take the form of a computer program product embodied on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0088] This application is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this application. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart... Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0089] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.
[0090] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0091] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit it. Although the present invention has been described in detail with reference to the above embodiments, those skilled in the art should understand that modifications or equivalent substitutions can still be made to the specific implementation of the present invention. Any modifications or equivalent substitutions that do not depart from the spirit and scope of the present invention should be covered within the scope of protection of the claims of the present invention.
Claims
1. A user-distributed resource energy management method based on intelligent decision control, characterized in that, include: Step 1: Construct an intelligent decision-making framework for the distributed system; Step 2: Based on the intelligent decision-making framework obtained in Step 1, construct a distributed system intelligent decision-making model based on deep reinforcement learning and MDP, and train it using reinforcement learning algorithms. After training, save the trained model. Step 3: Build an IoT-based hardware and software application platform and call the trained model from Step 2 to make decisions; The MDP components in step two include the agent's state space S, action space A, reward function R, and state transition functions P for air conditioners, water heaters, electric vehicles, and washing machines. The state transition function P is as follows: (1) (2) (3) (4) In the formula, , , and Let represent the indoor temperature, electric water heater water temperature, electric vehicle SOC state, and washing machine operating state at time t, respectively. These represent the state variables in the MDP model and must satisfy the following constraints: (5) (6) (7) (8) (9) (10) (11) (12) Constraints (5), (6), and (7) limit the upper and lower limits of the state quantities of air conditioners, electric water heaters, and electric vehicles; constraint (8) ensures that the electric vehicle's power needs to meet basic commuting requirements; constraint (9) limits the working hours of the washing machine; constraints (10) and (11) ensure that the washing machine is uninterrupted; and constraint (12) limits the time when the washing machine finishes its work.
2. The user-distributed resource energy management method based on intelligent decision control according to claim 1, characterized in that, The intelligent decision-making framework for the distributed system constructed in step one includes a power grid company, an EMS (Electronic Management System), an IoT control terminal, and N producers and consumers.
3. The user-distributed resource energy management method based on intelligent decision control according to claim 2, characterized in that, The EMS interacts with the grid side to exchange electricity prices and power data, and interacts with producers and consumers as well as with each other to exchange local market electricity prices and electricity prices. The EMS generates local market electricity prices based on the grid company's electricity price data and the energy demand data of producers and consumers. The Internet of Things control terminal transmits control signals to producers and consumers through artificial intelligence technology to regulate the relevant electrical equipment. The producer-consumer unit contains electrical equipment, including fixed loads, distributed photovoltaics, electric vehicles in distributed energy storage, air conditioners and electric water heaters in adjustable loads, and washing machines in deferred loads.
4. The user-distributed resource energy management method based on intelligent decision control according to claim 1, characterized in that, The state space S and action space A are as follows: (13) (14)。 5. A user-distributed resource energy management method based on intelligent decision control according to claim 4, characterized in that, In addition to the state variables of air conditioners, electric water heaters, electric vehicles and washing machines mentioned above, the state space S also includes time, local market electricity prices, photovoltaic output and fixed load status.
6. The user-distributed resource energy management method based on intelligent decision control according to claim 1, characterized in that, The reward function R is as follows: (15) (16) (17) (18) (19) (20) (21) Equations (15) and (16) are the reward functions for total electricity cost; Equations (17) and (18) are the reward functions for comfort of air conditioning and electric water heater, respectively; Equation (19) is the reward function for electric vehicle, taking into account power anxiety and SOC limitation; Equation (20) is the reward function for washing machine, taking into account the difference between the end time of work and the optimal end time.
7. The user-distributed resource energy management method based on intelligent decision control according to claim 1, characterized in that, In step three, the basic unit of the IoT hardware platform is a Raspberry Pi 4B. Based on the Respbain system, a deep reinforcement learning environment is built using Tensorflow and the Gym library. One Raspberry Pi 4B serves as the central energy management hub, while the remaining Raspberry Pi 4Bs act as energy management terminals for each prosumer, enabling decision-making and control over their respective prosumers. The Raspberry Pis communicate with each other via a Wi-Fi module to exchange various data and decision information. By utilizing the pre-built Raspberry Pi platform, the intelligent decision-making model framework is not repeatedly constructed, and the saved training model is run directly.
8. A user-distributed resource energy manager based on intelligent decision control, storing a program for running the user-distributed resource energy management method based on intelligent decision control as described in any one of claims 1-7.
Citation Information
Patent Citations
Household energy management method combining LSTM and deep reinforcement learning and medium
CN114841409A