Energy internet cloud-edge collaborative optimization scheduling method based on deep reinforcement learning
Patent Information
- Application Number
- CN202610878757.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-17
- Publication Date
- 2026-09-11
AI Technical Summary
[0004]针对现有技术的不足,本发明拟解决的技术问题是,提出了一种基于深度强化学习的能源互联网云边协同优化调度方法,旨在解决标准SAC算法无法平衡经济性与安全性,Actor网络梯度退化、关键信息丢失,固定折扣因子无法适配电价动态波动、长期效益分配不合理的问题,实现云边协同能源互联网快速、精准调度
1. 传统奖励函数仅以运行成本为导向,易忽略各种约束条件,导致策略生成无效动作;同时固定权重难以适配动态工况,无法平衡经济性与安全性。引入安全系数与硬约束惩罚项后,既能通过负成本项保证经济优化目标,又能通过惩罚信号引导智能体规避约束,还可通过动态权重自适应调整成本与安全的优先级,有效提升了策略的可行性与训练效率。
Smart Images

Figure CN122736184A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the interdisciplinary field of intelligent scheduling of energy internet and artificial intelligence technology, specifically involving a cloud-edge collaborative optimization scheduling method for energy internet based on deep reinforcement learning algorithm. Background Technology
[0002] With the advancement of energy transition and carbon neutrality goals, the energy internet, characterized by the coupling of multiple energy sources such as electricity, heat, and gas, has become an important form of future energy systems. The energy internet integrates fluctuating renewable energy sources such as wind and solar power, multi-energy loads (electricity, heat, and gas), energy storage systems, and various coupled devices. Complex coupling constraints exist among the devices in the energy internet, and the optimization scheduling problem exhibits strong nonlinearity, multiple constraints, and high dimensionality. Traditional mathematical programming methods (such as mixed-integer linear programming) are computationally inefficient when facing large-scale time-series optimization and struggle to adapt to the real-time random fluctuations in solar power output and load demand.
[0003] Deep Reinforcement Learning (DRL) offers a novel solution to the aforementioned problems. Among them, the Soft Actor-Critic (SAC) algorithm is an offline policy-based deep reinforcement learning algorithm based on the maximum entropy framework. By introducing an entropy regularization term into the objective function, it balances policy optimality with exploration diversity, demonstrating excellent performance in continuous action space optimization problems. With the continuous expansion of the energy internet, the requirements for computing power, latency, and data privacy in scheduling systems are becoming increasingly stringent. Against this backdrop, the cloud-edge collaborative architecture, as a core technology for overcoming the challenges of "high latency in centralized decision-making, weak autonomy of edge nodes, and low resource scheduling efficiency" in the energy internet, has become a crucial support for its efficient operation, providing a new technical path for multi-energy collaborative scheduling. In this cloud-edge collaborative architecture, the cloud is responsible for handling computationally intensive model training tasks, aggregating global historical data, and completing offline training and model version management for the SAC algorithm; the edge is responsible for handling latency-sensitive online inference and real-time control tasks, generating scheduling commands within a short window using a pre-trained Actor network, eliminating the need for round-trip communication with the cloud for each decision. However, the standard SAC algorithm still has the following shortcomings when applied to energy internet scheduling: the traditional reward function design ignores system operation constraints and cannot balance economy and security; the Actor network is prone to feature aliasing and gradient vanishing, and the policy is not aware of the constraint boundary, resulting in a decline in decision quality; the fixed discount factor keeps the agent's weight of long-term returns unchanged, ignoring the timeliness impact of energy price signals on current decisions. Summary of the Invention
[0004] To address the shortcomings of existing technologies, this invention proposes a cloud-edge collaborative optimization scheduling method for the energy internet based on deep reinforcement learning. This method aims to solve problems such as the inability of the standard SAC algorithm to balance economy and security, gradient degradation in the Actor network, loss of key information, and the inability of fixed discount factors to adapt to dynamic fluctuations in electricity prices and unreasonable long-term benefit distribution. The goal is to achieve rapid and accurate scheduling of the cloud-edge collaborative energy internet.
[0005] The present invention solves the aforementioned technical problem by adopting the following technical solution: A cloud-edge collaborative optimization scheduling method for the energy internet based on deep reinforcement learning, characterized by the following steps: Step 1: Construct an energy internet operation model; Step 2: Construct an energy internet optimization scheduling model based on the energy internet operation model; The objective function of the energy internet optimization scheduling model is: (5) In the formula, For intraday system operating costs, The cost of purchasing electricity from the power grid, To the cost of purchasing natural gas, For equipment operating costs, and These are the electricity purchase price and the natural gas price, for Electricity purchase during specific time periods for The amount of natural gas consumed by the combined heat and power unit during a given period. for The amount of natural gas consumed by the gas-fired boiler during a given period. The size of the scheduling time scale, , They are respectively Time-of-use equipment Operating power and operating cost, The scheduling period; The constraints of the energy internet optimization scheduling model include power balance constraints, equipment operation constraints, equipment ramping constraints, and energy storage unit state of charge constraints. Step 3: Construct a reinforcement learning framework based on the energy internet optimization scheduling model; The status includes electrical load, thermal load, photovoltaic power generation, state of charge of energy storage units, electricity purchase price, and time period. Status of the time period Represented as: (11) In the formula, , They are respectively Electrical and thermal loads during different time periods for Photovoltaic power generation during the period for State of charge of the time-limited energy storage unit; Actions during a period of time Represented as: (12) In the formula, for The power of the time-of-use energy storage unit, for The electrical power generated by the cogeneration unit during the period for The electrical power consumed by the electric boiler during a given period. for The amount of natural gas consumed by the gas-fired boiler during a given period; The reward function is: (13) (14) (15) (16) In the formula, for Instant rewards for specific time periods For safety reasons, and These are the coefficients for penalties for violating constraints. It is the number of constraints that are satisfied. It is the total number of constraints. for The heat power generated by the cogeneration unit during the period, and They are respectively The thermal power output of electric boilers and gas boilers during certain periods for The thermal power output of the electric boiler during the period for The heat output of the gas-fired boiler during the period for The electrical power output of the combined heat and power unit during the specified time period. , and These are the maximum ramp power limits for electric boilers, gas-fired boilers, and combined heat and power units, respectively. To take the absolute value; Step 4: Construct a reinforcement learning model; Step 5: Train the reinforcement learning model offline and execute it online; the cloud layer uses historical data to train the model, sends the trained parameters to the edge layer, the edge layer receives the parameters, generates scheduling commands and executes them, and feeds back the scheduling execution results to the cloud layer, forming a closed loop.
[0006] Compared with the prior art, the beneficial effects of the present invention are: 1. Traditional reward functions, driven solely by operating costs, tend to overlook various constraints, leading to ineffective policy actions. Furthermore, fixed weights struggle to adapt to dynamic conditions, failing to balance economy and safety. Introducing a safety factor and hard constraint penalty terms ensures economic optimization through negative cost terms, guides the agent to avoid constraints through penalty signals, and adaptively adjusts the priority of cost and safety through dynamic weights, effectively improving policy feasibility and training efficiency.
[0007] 2. To address the physical differences between external environmental variables and internal state variables in energy systems, a "dual-branch + feature residual learning" Actor network structure is proposed. This structure divides the state vector into an external environmental information group and an internal state information group. Through independent feature extraction branches and residual connections, it alleviates the problems of information aliasing and gradient vanishing in traditional single-branch networks, and improves the network's ability to perceive constraint boundaries.
[0008] 3. Traditional reinforcement learning uses a fixed discount factor to measure the importance of future rewards, which cannot adapt to the time-varying characteristics of dynamic electricity prices in energy internet dispatch scenarios. The proposed dynamic discount factor mechanism can adjust the agent's weight for future rewards according to dynamic electricity prices, enabling the strategy to focus on reducing immediate electricity purchase costs during peak electricity price periods and prioritize long-term energy reserves during off-peak periods, significantly improving the economic efficiency and scenario adaptability of the dispatch strategy. Attached Figure Description
[0009] Figure 1 This is a structural diagram of the energy internet system; Figure 2 This is the overall flowchart. Detailed Implementation
[0010] Specific embodiments are given below with reference to the accompanying drawings. These specific embodiments are only used to describe the technical solution of the present invention in detail, and are not intended to limit the scope of protection of this application.
[0011] like Figure 1-2 As shown, this invention proposes a cloud-edge collaborative optimization scheduling method for the energy internet based on deep reinforcement learning, comprising the following steps: Step 1: Construct an energy internet operation model to deploy various energy supply and coupling devices and load terminals at the edge layer; mainly including electric boilers (EB), gas boilers (GB), combined heat and power (CHP) units, electrical energy storage (BES), and photovoltaic (PV) equipment; Electric boilers and gas-fired boilers consume electricity and natural gas to generate heat, respectively. Their operating models are shown below: (1) (2) In the formula, and They are respectively The thermal power output of electric boilers and gas boilers during specific time periods; for The electrical power consumed by the electric boiler during a given time period; for The amount of natural gas consumed by the gas-fired boiler during a given period; and The conversion efficiencies are for electric boilers and gas-fired boilers, respectively.
[0012] A combined heat and power (CHP) unit consumes natural gas to generate electricity and heat, and its operating model is as follows: (3) In the formula, and They are respectively The electrical and thermal power generated by the cogeneration unit during the specified time period; for The amount of natural gas consumed by the combined heat and power unit during a given period; and These are the electrical energy and thermal energy conversion efficiencies, respectively.
[0013] The formula for calculating the State of Charge (SOC) of an energy storage unit is as follows: (4) In the formula, , They are respectively , State of charge of the time-limited energy storage unit for The charging / discharging power of the time-limited energy storage unit is positive when it indicates discharging and negative when it indicates charging. The capacity of the energy storage unit; The size of the scheduling time scale; This represents the charge / discharge coefficient of electrical energy storage.
[0014] Step 2: Construct an energy internet optimization scheduling model, including the objective function and constraints; 1) Objective function The goal of energy internet-based optimized scheduling is to minimize intraday system operating costs. C This includes the cost of purchasing electricity from the power grid. Cost of purchasing natural gas and equipment operating costs , represented as: (5) In the formula, and These are the electricity purchase price and the natural gas price, respectively. for Electricity purchase during a specific time period; , They are respectively Time-of-use equipment Operating power and operating costs; The scheduling period is [number].
[0015] 2) Constraints (1) Power balance constraint (6) (7) In the formula, for Photovoltaic power generation during the period , They are respectively Electrical and thermal loads during different time periods; (2) Equipment operation constraints Each device in the energy internet has upper and lower operating limits, as shown below: (8) In the formula, and These are the lower and upper limits of the output thermal power of the electric boiler, respectively. and These are the lower and upper limits of the output thermal power of the gas-fired boiler, respectively. and These are the lower and upper limits of the output electrical power of a combined heat and power unit, respectively. and These represent the lower and upper limits of the charging / discharging power of the energy storage unit, respectively.
[0016] (3) Equipment ramping constraints All devices in the energy internet must meet ramp constraints, as shown below: (9) In the formula, , They are respectively The heat output of electric boilers and gas boilers during certain periods. for The electrical power output of the combined heat and power unit during the specified time period. , and These are the maximum ramp power limits for electric boilers, gas-fired boilers, and combined heat and power units, respectively.
[0017] (4) State of charge constraints of energy storage units To avoid damage to the energy storage unit from deep charge / discharge, its State of Charge (SOC) is also limited to a certain range, which means: (10) In the formula, and These represent the lower and upper limits of the State of Charge (SOC) for the energy storage unit, respectively.
[0018] Step 3: Construct a reinforcement learning framework based on the energy internet optimization scheduling model; 1) State Space For each agent, the state space includes electrical and thermal loads, photovoltaic power generation, the state of charge of the energy storage unit, the electricity purchase price, and the current time period. Status of the time period Represented as: (11) 2) Action Space The action space of each agent is represented by the output of each device. Actions during a period of time Represented as: (12) 3) Reward function The optimization objective of reinforcement learning is to maximize cumulative reward, while the objective of this application is to minimize system operating cost. C Therefore, the reward function is defined as a negative operating cost. Furthermore, the system operation process needs to satisfy four types of constraints: power balance constraints and equipment operation constraints are mandatory hard constraints, while equipment ramp-up constraints and energy storage unit state-of-charge constraints are soft constraints that can be appropriately relaxed. To effectively address these constraints, a safety factor is introduced. Furthermore, by embedding the penalty for violating hard constraints into the reward function, the reward function is expressed as: (13) (14) (15) (16) in, for Instant rewards for specific time periods For safety reasons, It is the number of constraints that are satisfied. This is the total number of constraints. When all constraints are satisfied... The value is 1. At this time, Maximum. Conversely, the more constraints are not satisfied, the greater. The smaller the value. When all constraints are not met... , To set a very small value to ensure that the denominator is not zero. and These are the coefficients for penalties for violating the constraints.
[0019] Step 4: Construct a reinforcement learning model to enable the network to focus on learning the differences between different states and retain the underlying information that is important for decision-making; introduce a dynamic discount factor mechanism to balance immediate rewards and long-term benefits. The reinforcement learning model based on the SAC algorithm consists of 5 core networks: 1 action network. 2 evaluation networks and and two target networks and Action networks output the probability distribution of actions based on the state. Traditional Actor networks employ fully connected layers, with each layer performing feature transformation using a nonlinear activation function (such as ReLU). However, this often ignores the differences in physical characteristics of the system state and is prone to problems such as feature aliasing, gradient vanishing, and insufficient awareness of boundary constraints by the generation strategy. To address these issues, a "two-branch + feature residual learning" design is adopted, classifying electrical load, thermal load, and photovoltaic power generation as environmental states. The remaining system operation information is classified into the internal status group. The two sets of states are mapped to the same dimension through independent fully connected layers and activation functions, and then added to the original input state dimension by dimension to preserve the original features, enhance the environmental state and enhance the internal state. (17) (18) In the formula, , They are respectively The environmental and internal conditions after the time period change; , They are respectively Enhanced environmental and internal conditions during different time periods. It is the ReLU activation function. , This is the weight matrix. , For bias; The enhanced environmental state and the enhanced internal state are spliced together to obtain the fusion feature. ; to integrate features After extracting high-order features through two fully connected layers, the mean and standard deviation of the action distribution are output. This structure can effectively distinguish the feature patterns between external disturbances and the internal state of the system, alleviate the gradient vanishing and feature degradation problems of deep networks, improve the adaptability of the policy to system operation constraints and the decision accuracy, and enhance training stability and algorithm convergence performance.
[0020] The optimization objective of SAC adds an entropy regularization term to the traditional reward maximization, as shown below: (19) In the formula, As expected, It is the policy entropy, which measures the uncertainty of action distribution. The higher the entropy, the stronger the exploratory nature. It is a temperature coefficient used to control the weight of entropy; It is a dynamic discount factor; Traditional reinforcement learning uses a fixed discount factor While emphasizing the importance of future rewards, dynamic electricity prices are crucial in energy internet dispatch scenarios. Therefore, a dynamic discount factor mechanism is proposed. During peak electricity price periods, when electricity purchase costs are high and the economic impact of immediate operational decisions is significant, the discount factor should be reduced to encourage agents to focus on minimizing current costs, prioritizing the use of photovoltaics, CHP, and BES to reduce high-priced electricity purchases. During off-peak electricity price periods, when electricity purchase costs are low, the discount factor should be increased to encourage agents to focus on long-term benefits and formulate energy storage dispatch strategies to prepare for low-cost operation during peak periods. During flat electricity price periods, the discount factor is set to a median value to balance immediate and long-term benefits. The specific calculation method is as follows: (20) (twenty one) in, and Do not specify the maximum and minimum values for dynamic electricity prices; and The maximum and minimum values of the discount factor can be set to 0.99 and 0.9, respectively. During network updates, the loss function is minimized. The network parameters are optimized using the following formula: (twenty two) (twenty three) (twenty four) The action network parameter update rule is to minimize the policy loss. : (25) Temperature coefficient The update needs to minimize the loss, as shown below: (26) in, The dimension set as the action space; Target network parameters The following soft update method is used: (27) in This is the soft update coefficient.
[0021] Step 5: Perform offline training and online execution. The cloud layer uses historical data to train the network and sends the trained parameters to the edge layer. The edge layer receives the parameters, generates scheduling commands and executes them, and at the same time feeds back the scheduling execution results to the cloud layer, forming a closed loop.
[0022] In the initial training phase, the cloud layer first initializes the parameters of the action and evaluation networks of the SAC algorithm and constructs an experience buffer to store historical interaction data from the edge layers, providing data support for subsequent network updates. During continuous training, the cloud layer continuously and randomly extracts training data of a fixed batch size (Batch_size) from the experience buffer. This batch of data consists of valid interaction samples from the historical feedback of the edge layers. After data extraction, the cloud layer strictly follows the preset training steps (i.e., the update rules for the evaluation and action networks in step four), sequentially updating the soft Q-values of the dual evaluation networks, updating the policy parameters of the action networks, adaptively adjusting the temperature coefficient α, and softly updating the target evaluation network. The entire update process iterates continuously until the reward value reaches the preset convergence condition. At this point, the parameters of the mature and optimally performing Actor network are obtained, completing this round of cloud training.
[0023] After the cloud layer completes the action network training, it packages and distributes the optimal trained parameters to the edge layer nodes, ensuring that the edge layer can execute real-time decision-making tasks based on the latest optimization strategies. As the core carrier for interaction with the real-world scenario, the edge layer collects multi-dimensional state information of the current environment in real time, including multi-energy load data, the operating status of various devices, dynamic electricity prices, and other key data. Through data preprocessing and feature fusion, it forms the current environmental state. Subsequently, the edge layer invokes the locally deployed action network, loads the optimal parameters sent from the cloud, and inputs the current state. The action network outputs actions that conform to the maximum entropy strategy. The edge layer performs the action. Afterwards, the system will monitor the action execution effect in real time and calculate the reward value for this round of interaction based on the preset reward function. After an action is executed, the environmental state will change accordingly. The edge layer further collects the updated environmental information to form a new state. and a complete data sample of this interaction. Standardize the records and store them in a local data list.
[0024] To ensure the validity of feedback data and the efficiency of batch processing, the edge layer monitors its local data list in real time. When the number of interactive data samples stored in the list reaches a preset threshold, the entire data list is packaged and encrypted, and then fed back to the cloud layer through a reliable communication link, completing an edge-to-cloud data interaction. After receiving the data from the edge layer, the cloud layer first verifies the data's validity (removing outliers and missing values), and then stores the verified high-quality data in an experience buffer pool. At this time, the cloud layer checks the total amount of existing data in the experience buffer pool in real time. If the amount of data has reached the buffer pool's preset maximum capacity, to avoid data redundancy and ensure the timeliness and high quality of the data in the buffer pool, a "first-in, first-out" iterative mechanism is adopted. The oldest data initially stored in the buffer pool is deleted, and replaced with the latest data samples transmitted by the edge layer, ensuring that the experience buffer pool always stores the most valuable interactive data, providing reliable support for subsequent network updates.
[0025] As the edge layer continuously interacts with the environment and feeds back new data samples to the cloud, high-quality data in the experience buffer pool accumulates and updates dynamically. After a preset time period, the cloud layer restarts a new round of network training. Based on the latest data in the buffer pool, it updates the parameters of the evaluation network and action network, further optimizing policy performance and obtaining action network parameters that are more suitable for the current real-world scenario. Subsequently, the cloud layer sends the updated optimal parameters back to the edge layer, which then makes real-time decisions based on the new parameters. This forms a continuous optimization loop of "cloud training - edge execution - data feedback - cloud iteration," ensuring that the SAC algorithm's policy can always adapt to the dynamically changing environment and achieve long-term optimal decision-making performance.
[0026] Any aspects not covered in this invention are applicable to existing technologies.
Claims
1. A cloud-edge collaborative optimization scheduling method for the energy internet based on deep reinforcement learning, characterized in that, Includes the following steps: Step 1: Construct an energy internet operation model; Step 2: Construct an energy internet optimization scheduling model based on the energy internet operation model; The objective function of the energy internet optimization scheduling model is: (5) In the formula, For intraday system operating costs, The cost of purchasing electricity from the power grid, To the cost of purchasing natural gas, For equipment operating costs, and These are the electricity purchase price and the natural gas price, for Electricity purchase during specific time periods for The amount of natural gas consumed by the combined heat and power unit during the period. for The amount of natural gas consumed by the gas-fired boiler during a given period. The scheduling time scale, , They are respectively Time-of-use equipment Operating power and operating cost, The scheduling period; The constraints of the energy internet optimization scheduling model include power balance constraints, equipment operation constraints, equipment ramping constraints, and energy storage unit state of charge constraints. Step 3: Construct a reinforcement learning framework based on the energy internet optimization scheduling model; The status includes electrical load, thermal load, photovoltaic power generation, state of charge of energy storage units, electricity purchase price, and time period. Status of the time period Represented as: (11) In the formula, , They are respectively Electrical and thermal loads during different time periods for Photovoltaic power generation during the period for State of charge of the time-limited energy storage unit; Actions during a period of time Represented as: (12) In the formula, for The power of the time-of-use energy storage unit, for The electrical power generated by the cogeneration unit during the period for The electrical power consumed by the electric boiler during a given period. for The amount of natural gas consumed by the gas-fired boiler during a given period; The reward function is: (13) (14) (15) (16) In the formula, for Instant rewards for specific time periods For safety reasons, and These are the coefficients for penalties for violating constraints. It is the number of constraints that are satisfied. It is the total number of constraints. for The heat power generated by the cogeneration unit during the period, and They are respectively The thermal power output of electric boilers and gas boilers during certain periods for The thermal power output of the electric boiler during the period for The heat output of the gas-fired boiler during the period for The electrical power output of the combined heat and power unit during the specified time period. , and These are the maximum ramp power limits for electric boilers, gas-fired boilers, and combined heat and power units, respectively. To take the absolute value; Step 4: Construct a reinforcement learning model; Step 5: Train the reinforcement learning model offline and execute it online; the cloud layer uses historical data to train the model, sends the trained parameters to the edge layer, the edge layer receives the parameters, generates scheduling commands and executes them, and at the same time feeds back the scheduling execution results to the cloud layer, forming a closed loop.