A robust reinforcement learning-based day-and-night dual-mode intelligent scheduling method for energy flow of an oceanic island group
Patent Information
- Application Number
- CN202610427280.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-04-02
- Publication Date
- 2026-07-10
Smart Images

Figure CN122371320A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of energy optimization scheduling technology for island microgrids, specifically to a day-night dual-mode intelligent scheduling method for energy flow in offshore island clusters based on robust reinforcement learning. Background Technology
[0002] Because of their distance from the mainland, remote island groups cannot be connected to the main power grid, and their energy supply is highly dependent on local renewable energy and controllable power sources. Therefore, these groups face three major technical challenges: First, uneven resource endowment and geographical isolation mean that resource islands have abundant renewable energy but low load demand, while load islands have concentrated loads but scarce energy, and natural isolation between islands limits direct energy transmission. Second, strong environmental uncertainty means that renewable energy output fluctuates due to weather conditions, and the marine environment can easily disrupt battery swapping vessel transportation; traditional dispatching methods struggle to address these uncertainties. Third, significant differences in daytime and nighttime supply and demand mean that renewable energy sources such as photovoltaics have high output during the day, but production loads are concentrated; existing solutions lack differentiated dispatching mechanisms, easily leading to energy stagnation, supply shortages, or excessively high operating costs.
[0003] Existing island energy dispatching technologies mostly focus on single islands, single energy types, or deterministic scenarios, resulting in the following shortcomings: 1) They do not fully utilize the mobile energy storage characteristics of battery-swapping vessels, leading to low cross-island energy transmission efficiency; 2) They lack differentiated optimization objectives for day and night phases, making it difficult to adapt to dynamic changes in supply and demand conditions; 3) They lack robustness to uncertainties such as renewable energy fluctuations and transportation interruptions, resulting in poor system reliability; 4) The algorithms mostly employ traditional optimization methods or basic reinforcement learning, which suffer from insufficient learning efficiency and stability, making it difficult to meet the complex dispatching needs of distant island clusters. Therefore, a method for optimizing energy flow dispatching that balances cross-island collaboration, day-night adaptation, and robust disturbance resistance is needed. Summary of the Invention
[0004] To address the shortcomings and deficiencies of existing technologies, this invention provides a day-night dual-mode intelligent scheduling method for energy flow in ocean-going island clusters based on robust reinforcement learning. The core objective is to solve problems such as limited energy transmission due to geographical isolation in ocean-going island clusters, insufficient adaptation to day-night supply and demand differences, and unstable operation under uncertain scenarios, thereby achieving energy self-sufficiency and efficient and economical operation of the island clusters.
[0005] To achieve the above-mentioned technical objectives, the present invention provides the following technical solution: A robust reinforcement learning-based intelligent scheduling method for energy flow in ocean-going island groups, operating in both day and night modes, comprises the following steps: Construct a comprehensive energy system for the island clusters in the open ocean, forming an energy flow transmission framework for the island clusters, encompassing energy production, energy conversion, inter-island transportation, and energy consumption; Based on the energy flow transmission architecture of island clusters, and considering the differences in energy supply and demand between day and night in the ocean island clusters, a two-stage scheduling mechanism including daytime scheduling mode and nighttime scheduling mode is established, and daytime and nighttime scheduling optimization objectives are formulated. The problem of energy flow scheduling in remote island groups is modeled as a robust Markov decision process, and a dual-agent collaborative decision-making system adapted to day and night is constructed to achieve the optimization objective under the two-stage scheduling mechanism. A robust adversarial reinforcement learning algorithm is used to solve the robust Markov decision process, and outputs a day-night adapted robust real-time scheduling strategy to drive the energy flow transmission and coordinated operation of each unit in the integrated energy system architecture of the ocean island cluster.
[0006] Furthermore, the construction of a comprehensive energy system for a group of distant-water islands, forming an energy flow transmission architecture for energy production, energy conversion, inter-island transportation, and energy consumption, specifically involves: Based on the geographically separated distribution characteristics of resource islands and load islands in the ocean island cluster, the islands are divided into resource islands that undertake energy production and energy conversion functions and load islands that consume energy. Battery swapping vessels are configured as mobile energy storage and transportation carriers connecting resource islands and load islands for cross-island transportation. The resource island is equipped with renewable energy generating units, energy conversion devices, and energy storage systems. The renewable energy generating units include wind turbines and photovoltaic units, and the energy conversion devices include electrolyzers, nitrogen production modules, ammonia synthesis reactors, and seawater desalination devices. Under the premise that the resource island can meet its own load, the energy transported by battery-swapping vessels across the island, and the power supply needs of the load island, the surplus electrical energy will be converted and stored through the electro-ammonia conversion device and the seawater desalination device. The load island is equipped with a controllable power supply and an energy storage system; the battery swapping vessel is equipped with an energy storage system; the energy storage system of the battery swapping vessel includes storing the electrical energy required for transporting the load island and the energy converted from the energy of the resource island for cross-island transportation; the energy storage systems of the resource island and the load island are used to store and allocate electrical energy or energy converted from energy. For renewable energy generating units, we will construct a power output prediction model that considers prediction errors; for battery-swapping vessels, we will construct a navigation energy consumption model that considers real-time changes in sea state and load; for controllable power sources, we will construct a dynamic power source model that considers efficiency adjustment and response lag; and for energy storage systems, we will construct a dynamic energy storage model that considers charging and discharging efficiency. The above model is used to characterize the uncertainties in the energy flow transmission architecture of island clusters, providing a decision-making basis for subsequent robust scheduling.
[0007] Furthermore, based on the energy flow transmission architecture of island clusters, considering the differences in energy supply and demand between day and night in the offshore island clusters, a two-stage scheduling mechanism including daytime and nighttime scheduling modes is established, and the specific daytime and nighttime scheduling optimization objectives are formulated as follows: Based on the differences in energy supply and demand between day and night, the scheduling cycle is divided into daytime and nighttime periods, and corresponding daytime and nighttime scheduling modes are established respectively. The two scheduling modes take minimizing operating costs as the optimization objective. The daytime dispatch mode takes into account high renewable energy output, concentrated production and public service loads, and cross-island transportation by battery swapping vessels. The daytime operating costs include: the navigation power consumption and charging and discharging losses incurred by battery swapping vessels in performing cross-island transportation tasks, the cost of renewable energy abandonment caused by excess photovoltaic / wind power output on resource islands, and the cost of controllable loads cut off during daytime periods. The nighttime dispatch mode takes into account the low output of renewable energy, the main load of residential life and the high dependence on energy storage. Its optimization goal is to minimize the nighttime operating cost. The nighttime operating cost includes: the cost of controllable load cut off during the nighttime period, the power supply cost on the load island, and the cost of renewable energy abandonment on the resource island due to the continued wind power output. A unified day and night scheduling optimization target is constructed, and a dynamic weighting mechanism and scene identification factor are introduced to weight and fuse the daytime optimization target and the nighttime optimization target. The dynamic weight is adaptively adjusted according to the real-time energy supply and demand tension. When the energy supply is tight, the weight of the nighttime optimization target is increased, and when the energy supply is sufficient, the weight of the daytime optimization target is increased. At the same time, cross-stage collaborative constraints are set, including: at the end of daytime scheduling, the remaining amount of electrical energy in the energy storage systems of the resource island and the load island is not lower than a preset threshold; the ammonia and freshwater reserves formed by energy conversion in the resource island can meet the consumption needs of the load island.
[0008] Furthermore, the problem of day-night dual-mode scheduling of energy flow in offshore island groups is modeled as a robust Markov decision process, and a day-night adapted dual-agent collaborative decision-making system is constructed to achieve the optimization objective under the dual-stage scheduling mechanism. Specifically, this is as follows: The day-night dual-mode scheduling problem of energy flow in a remote island group is modeled as a robust Markov decision process (RMDP), defining its state space S, action space A, reward function R, and uncertainty set P; the reward function is the negative of the day-night scheduling optimization objective; the uncertainty set P is contained in the state space S. Execute action Then transition to each possible state All possible transition probabilities; Based on the uncertain set P, a dual-agent collaborative decision-making system consisting of a robust agent and an adversarial agent is constructed. The adversarial agent is used to simulate the worst-case scenario by minimizing the reward function, while the robust agent is used to learn the energy scheduling strategy under the worst-case scenario and output a robust scheduling scheme that adapts to the switching between day and night scenarios.
[0009] More specifically, the state space S includes: Resource-side state variables used to characterize the energy output and energy storage status of the resource island; transportation state variables used to characterize the operating status of the battery swapping vessel; and load-side state variables used to characterize the load demand and local power output of the load island.
[0010] More specifically, the action space A includes: Energy storage scheduling variables are used to control the charging and discharging of the electrical energy storage section in the resource island energy storage system; ship scheduling variables are used to schedule the navigation and berthing of battery swapping vessels; and load control variables are used to adjust the output of controllable power sources and controllable loads in the load island.
[0011] Furthermore, the adversarial agent is used to simulate the worst-case scenario by minimizing the reward function, and the robust agent is used to learn an energy scheduling strategy under the worst-case scenario to output a robust scheduling scheme that adapts to day-night scene switching. Specifically: Both adversarial agents and robust agents consist of a policy network and a value network; The adversarial agent generates interference actions through a policy network to represent the adverse conditions that may occur during the state transition process, and fits the value of the interference actions through a value network to output the interference action with the minimum value, i.e. the interference action representing the worst condition, as the interference scenario. The robust agent generates robust scheduling actions based on the current state and the interference scenario output by the adversarial agent through its policy network, and fits the robust action value that can be obtained by executing the robust scheduling action under the interference scenario through its value network. Based on the robust action value, the policy network parameters of the robust agent are optimized, so that the robust agent learns the scheduling strategy that maximizes the robust value function under environmental interference. The robust value function is defined as the expected cumulative discount reward function under the worst case.
[0012] Furthermore, in the robust adversarial reinforcement learning algorithm: Introducing a risk distortion mechanism into the optimization objective of adversarial agents: when calculating the value of perturbation optimization, the standard expectation of the reward function R is replaced with the risk-distorted expectation, which amplifies the weight of low-reward scenarios based on quantile transformation; A regularized baseline mechanism is introduced in the policy optimization of robust agents: the regularized baseline is obtained by weighting the value of robust scheduling actions generated under the current policy; the advantage function is obtained by subtracting the regularized baseline from the value of robust actions, and the policy network parameters of robust agents are optimized based on the advantage function.
[0013] Based on the above technical solution, the present invention has at least the following beneficial effects: This invention combines the spatial layout characteristics of "load islands + resource islands + battery swapping vessel transportation" in ocean-going island clusters, renewable energy endowments, and the mobile energy storage characteristics of battery swapping vessels. It constructs operational models for the navigation energy consumption and transportation capacity of battery swapping vessels, as well as charging and discharging models. This solves the problem of limited direct energy transmission caused by geographical isolation of island clusters, and achieves adaptive matching to changes in load demand on load islands. By leveraging the constructed integrated energy system architecture and energy management model of island clusters, it achieves orderly scheduling of cross-island energy flow in scenarios with limited energy transmission. Furthermore, through a robust reinforcement learning optimization framework and a model-based robust adversarial reinforcement learning algorithm (MB-RAC), it solves the energy management challenges of multiple energy sources and multiple uncertainties among island clusters. Ultimately, it achieves energy self-sufficiency for ocean-going island clusters, providing technical support for their sustainable development, and also providing a technical path for the implementation of the energy internet concept in ocean-going island scenarios. Attached Figure Description
[0014] The accompanying drawings, which are included to provide a further understanding of this application and form part of this application, illustrate exemplary embodiments and are used to explain this application, but do not constitute an undue limitation of this application. In the drawings: Figure 1 This is a schematic diagram of the energy flow scheduling model for an island group in an embodiment of the present invention; Figure 2 This is a flowchart of a day-night dual-mode intelligent scheduling method for energy flow in ocean-going island groups based on robust reinforcement learning, as proposed in this invention. Detailed Implementation
[0015] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the invention.
[0016] Although the steps in this invention are arranged by reference numerals, this is not intended to limit the order of the steps. Unless the order of the steps is explicitly stated or the execution of a step requires other steps as a basis, the relative order of the steps can be adjusted. It is understood that the term "and / or" as used herein refers to and covers any and all possible combinations of one or more of the associated listed items.
[0017] like Figure 2 As shown, this invention proposes a day-night dual-mode intelligent scheduling method for energy flow in offshore island groups based on robust reinforcement learning, which specifically includes the following steps: S1. Construct a comprehensive energy system for the island clusters in the open ocean, forming an energy flow transmission framework for the island clusters, encompassing energy production, energy conversion, cross-island transportation, and energy consumption; In a preferred embodiment, step S1 specifically comprises: Based on the geographically separated distribution characteristics of resource islands and load islands in the ocean island cluster, the islands are divided into resource islands that undertake energy production and energy conversion functions and load islands that consume energy. Battery swapping vessels are configured as mobile energy storage and transportation carriers connecting resource islands and load islands for cross-island transportation. The resource island is equipped with renewable energy generating units, energy conversion devices, and energy storage systems. The renewable energy generating units include wind turbines and photovoltaic units, and the energy conversion devices include electrolyzers, nitrogen production modules, ammonia synthesis reactors, and seawater desalination devices. Under the premise that the resource island can meet its own load, the energy transported by battery-swapping vessels across the island, and the power supply needs of the load island, the surplus electrical energy will be converted and stored through the electro-ammonia conversion device and the seawater desalination device. The load island is equipped with a controllable power supply and an energy storage system; the battery swapping vessel is equipped with an energy storage system; the energy storage system of the battery swapping vessel includes storing the electrical energy required for transporting the load island and the energy converted from the energy of the resource island for cross-island transportation; the energy storage systems of the resource island and the load island are used to store and allocate electrical energy or energy converted from energy. In this embodiment, as Figure 1 As shown, a certain group of distant ocean islands is defined as comprising one load island (equipped with a 10MWh energy storage system, two 500kW gas turbines, power poles, etc.) and three resource islands (located 5km, 20km, and 25km from the load island, respectively); the resource islands are configured as follows: Resource Island 1 is equipped with a 500kW photovoltaic unit, an 800kW wind turbine, a 100Nm³ / d seawater desalination unit, and a 50kg / d electrolytic ammonia production unit. Resource Island 2 is equipped with a 600kW photovoltaic unit, a 600kW wind turbine unit, a 130Nm³ / d seawater desalination unit, and an 80Nm³ / d electrolytic ammonia production unit. Resource Island 3 is equipped with a 400kW photovoltaic unit, a 1000kW wind turbine unit, a 120Nm³ / d seawater desalination unit, and a 60kg / d electrolytic ammonia production unit. There are 3 battery-swapping vessels, each with a rated load capacity of 5t, an energy storage capacity of 800kWh, a rated navigation power of 100kW, and a charge / discharge efficiency of 0.92.
[0018] Considering uncertainties such as renewable energy fluctuations and transportation disruptions, this application constructs a generator output prediction model that takes into account prediction errors for renewable energy generator units, a navigation energy consumption model that takes into account real-time changes in sea state and load for battery-swapping vessels, a dynamic power supply model that takes into account efficiency regulation and response lag for controllable power sources (gas turbines), and a dynamic energy storage model that takes into account charging and discharging efficiency for the energy storage function in energy storage systems.
[0019] In this embodiment, the formula for the generator output prediction model of renewable energy units is expressed as follows: ; ; in, Resource islands Photovoltaic units and wind turbines on Robust predictive output at any given moment , Resource islands Photovoltaic units and wind turbines on Predicting output at any given moment; , Resource islands Photovoltaic units and wind turbines on The allowable range of prediction error at any given time; , This is the corresponding error adjustment coefficient; Resource islands The state factors for the positive and negative errors of the photovoltaic and wind turbine units are defined, and the constraints of the state factors are as follows: ; ; in, For the sampling time window, Resource islands The maximum cumulative error coefficient of photovoltaic and wind turbine units within the sampling time window; These are the change rates of the adjustment coefficients for photovoltaic units and wind turbine units, respectively.
[0020] This embodiment also describes the relevant process and energy consumption for converting and storing surplus electrical energy through an electro-ammonia and seawater desalination device, as detailed below: First, an electrolyzer is used to decompose seawater to produce hydrogen using surplus electrical energy, represented as follows: ; in, This refers to the electrical power of the electrolytic cell; To improve the power distribution efficiency of the resource island; To reduce system power consumption; This refers to the effective input power of the electrolytic cell; Electrolysis efficiency correction term; The enthalpy of water decomposition; For hydrogen production; This is the electrolysis temperature.
[0021] Next, pure nitrogen is separated from the air using a nitrogen generation module, as shown in the formula: ; in, The energy consumption coefficient per unit of nitrogen production; Power consumption of the nitrogen generator module; This represents nitrogen production.
[0022] Finally, hydrogen and nitrogen are combined in the ammonia synthesis reactor to synthesize ammonia, as shown in the following formula: ; in, , These are the molar masses of ammonia and hydrogen, respectively. For ammonia synthesis efficiency, ammonia synthesis temperature, This is the pressure required for ammonia synthesis. The ammonia synthesis reaction is exothermic. The synthesis reaction is exothermic.
[0023] The seawater desalination process is also represented by the power consumption of the desalination unit, as shown in the formula: ; in, For seawater desalination equipment Power consumed during a period of time For the specific energy consumption of the device, The amount of fresh water produced, To downplay efficiency; This refers to the seawater salinity coefficient. This is a correction term for salinity loss.
[0024] In this embodiment, the formula for the navigation energy consumption model of the battery-swapping vessel is expressed as follows: ; ; ; ; ; in, Rated navigation electrical power; For island wind speed, For island high waves, The water flow velocity acting on the battery-swapping vessel in the marine environment; , , This is the sea state influence coefficient; The weight of the ship itself; For ships Sea state influence coefficient under sea state conditions; For ships Load influence coefficient under load conditions; Ships Total mass of the energy storage; For the storage quality of ammonia and water; This refers to the number of shipboard energy storage units. Mass of a single energy storage unit.
[0025] In this embodiment, the formula for the dynamic power supply model of the gas turbine is expressed as follows: ; ; in, The actual output power of the gas turbine at time t of load island j reflects the current power supply capacity. Let be the power adjustment amount at time t. Power regulation efficiency factor; and They are respectively The lower limit and upper limit.
[0026] In this embodiment, a dynamic energy storage model is established for the part of the energy storage system that performs the function of storing electrical energy. The formula is as follows: For the electrical energy storage unit (ESU), including the resource island electrical energy storage unit and the load island electrical energy storage unit, its state of charge update is expressed as: ; in, Indicates the rated capacity of the energy storage; , These are the battery charging and discharging outputs, respectively. , This refers to the charge / discharge efficiency factor. The scheduling time step; , The states of charge at time t+1 and time t are respectively. The above model is used to characterize the uncertainties in the energy flow transmission architecture of island clusters, providing a decision-making basis for subsequent robust scheduling.
[0027] S2. Based on the energy flow transmission architecture of the island cluster, considering the differences in energy supply and demand between day and night in the ocean island cluster, a two-stage scheduling mechanism including daytime scheduling mode and nighttime scheduling mode is established, and daytime and nighttime scheduling optimization objectives are formulated. In a preferred embodiment, step S2 specifically includes: Based on the differences in energy supply and demand between day and night, the scheduling cycle is divided into daytime and nighttime periods, and corresponding daytime and nighttime scheduling modes are established respectively. The two scheduling modes take minimizing operating costs as the optimization objective. The daytime dispatch mode considers high renewable energy output, concentrated production and public service loads, and inter-island transportation by battery-swapping vessels. Therefore, the daytime operating costs include: the navigation power consumption and charging / discharging losses incurred by the battery-swapping vessels during inter-island transportation tasks; the cost of renewable energy abandonment due to excess photovoltaic / wind power output on resource islands; and the cost of controllable load cut-off during daytime periods. In this embodiment, the daytime optimization objective function... With the goal of minimizing overall cost, its formula is expressed as: ; in, For ships Self-consumption of energy and charging / discharging losses of the transport carrier, They are respectively Energy waste on the resource island at all times As a penalty factor; A controllable workload for day-operated resections; This refers to the daytime scheduling period; This is the load shedding penalty factor; This refers to resource islands, load islands, and the number of ships.
[0028] The nighttime dispatch mode considers low renewable energy output, primarily residential load, and high reliance on energy storage. Its optimization objective is to minimize nighttime operating costs, which include: the cost of controllable load shedding during nighttime hours, the power supply cost on the load island, and the cost of renewable energy curtailment on the resource island due to ongoing wind turbine output. During nighttime hours, since the photovoltaic output on the resource island is zero, battery swapping vessels do not conduct inter-island transportation. The resource island stores surplus electricity through ammonia conversion or seawater desalination for use in subsequent periods. At this time, the load island mainly relies on gas turbines and energy storage systems for power supply. Considering the cost of shedding controllable load, the mode aims to ensure the stability and reliability of the island cluster's power system; therefore, the target design nighttime operating cost is... The formula is expressed as: ; in, This is the nighttime dispatch period. A controllable workload for nighttime resection; This is the load shedding penalty factor; , This represents the cost coefficient for gas turbines. For the load island gas turbine in Output power at any given moment; The use of renewable energy on resource islands has been abandoned. This is a penalty factor for abandoning energy use.
[0029] To comprehensively consider the scheduling objectives during the day and night, and to achieve a smooth transition and adaptive switching between the two modes, this embodiment constructs a unified day-night scheduling optimization objective. A dynamic weighting mechanism and scene identification factors are introduced to weight and fuse the daytime and nighttime optimization objectives, ensuring the continuity and robustness of the scheduling strategy over time. The formula for the day-night scheduling optimization objective F is expressed as: ; in, , This is a dynamic weighting function used to balance the optimization priorities between day and night, and it is dynamically adjusted according to the energy supply and demand tension of the island group. Indicates the current state. ; The scenario identification factor is used to distinguish between daytime and nighttime scheduling scenarios, ensuring that the weighting function is applied correctly to the corresponding scheduling stage and that the daytime and nighttime optimization goals are smoothly connected. At the same time, combined with the real-time energy supply and demand tension index, the scenario identification factor enables the scheduling strategy to adaptively adjust the daytime and nighttime weights to achieve the optimization of load matching and energy utilization.
[0030] The dynamic weights are adaptively adjusted based on real-time energy supply and demand tension. When energy supply is tight, the weight of the nighttime optimization target is increased; when energy supply is ample, the weight of the daytime optimization target is increased. In this embodiment, the weights are based on the energy supply and demand tension index. The dynamic weights are calculated using the following formula: ; in, ; This is the supply and demand balance threshold; This is the adjustment coefficient; when the system is in a state of energy shortage, Increase. Prioritize nighttime targets; conversely, during periods of ample energy supply, Increased costs, with daytime operating costs dominating.
[0031] Furthermore, to ensure the consistency of the day-night scheduling strategy, this application also sets cross-stage collaborative constraints, including: at the end of daytime scheduling, the remaining electrical energy in the energy storage systems of the resource island and the load island is not lower than a preset threshold; the ammonia and freshwater reserves of the resource island are sufficient to meet the consumption needs of the load island; the formula for the above constraints is: ; ; ; ; in, , This represents the percentage of remaining electricity stored at the end of the day; , For ammonia and freshwater production on resource islands; , This is the initial storage amount; The amount of freshwater consumed to support the ammonia load island.
[0032] S3. The energy flow day-night dual-mode scheduling problem of the ocean island group is modeled as a robust Markov decision process, and a day-night adapted dual-agent collaborative decision system is constructed to achieve the optimization objective under the dual-stage scheduling mechanism. In a preferred embodiment, step S3 specifically includes: The energy flow day-night dual-mode scheduling problem of the distant island group is modeled as a robust Markov decision process (RMDP), and its state space S, action space A, reward function R and uncertainty set P are defined. In this embodiment, the state space S includes: Resource-side state variables characterizing the energy output and storage status of resource islands (including battery energy storage, ammonia conversion, and freshwater conversion); transportation-side state variables characterizing the operating status and energy storage status of battery-swapping vessels (including the quality of converted energy obtained from resource islands); load-side state variables characterizing the load demand and local power output of load islands; the formula is expressed as: ; in, For resource island The remaining energy storage capacity at time t; For resource island ammonia at time t Freshwater storage capacity; The scheduling status of the k-th battery-swapping vessel; This refers to the output power of the gas turbine. Indicates the ship swapping batteries at time t. The remaining power of the energy storage unit This represents the total load demand of the load island at time t; It indicates the output power of photovoltaics and wind turbines.
[0033] The action space A includes: energy storage scheduling variables for controlling the charging and discharging of the electrical energy storage portion of the resource island energy storage system; ship scheduling variables for scheduling the navigation and berthing of battery swapping vessels; and load control variables for adjusting the controllable power output and controllable load of the load island; expressed by the formula: ; in, The energy storage charging and discharging power of resource island i at time t; Dispatch instructions for ships undergoing battery swapping; Let be the controllable load shedding amount of the load island at time t; Let J be the power adjustment of the gas turbine at time t.
[0034] The reward function is the negative of the day-night scheduling optimization objective, i.e. The uncertain set Included in state Execute action Then transition to each possible state All possible transition probabilities ,Right now ; Based on the uncertain set P, a dual-agent collaborative decision-making system consisting of a robust agent and an adversarial agent is constructed. The adversarial agent is used to simulate the worst-case scenario by minimizing the reward function, and the robust agent is used to learn the energy scheduling strategy under the worst-case scenario and output a robust scheduling scheme that adapts to the switching between day and night scenarios. In a preferred embodiment, the adversarial agent is used to simulate the worst-case scenario by minimizing the reward function, and the robust agent is used to learn an energy scheduling strategy under the worst-case scenario to output a robust scheduling scheme that adapts to day-night scene switching. Both adversarial agents and robust agents consist of a policy network and a value network; Adversarial agents generate disruptive actions through a policy network. The adverse working conditions that may occur during the state transition are represented by the following countermeasures: ;in, To determine the policy network parameters for the adversarial agent, the state is fitted using a value network. Lower interference action Corresponding interference action value , To counteract the agent's value network parameters, the value of the perturbation action is used to quantify the negative impact of the generated perturbation action on the robust scheduling strategy. The smaller the value, the greater the negative impact, and the worse the operating condition. The perturbation action with the smallest value, representing the worst operating condition, is output as the perturbation scenario. ; The robust agent generates a robust scheduling action a based on its current state s and the interference scenario d output by the adversarial agent through its policy network. Its robust policy is represented as follows: ;in, The parameters of the robust agent's policy network are used to evaluate its performance in interference scenarios through its value network. Robust action value that can be obtained by performing robust scheduling actions ,That The parameters of the robust agent's value network are defined as follows: The policy network parameters of the robust agent are optimized based on the robust action value, enabling the robust agent to learn a scheduling policy that maximizes the robust value function under environmental disturbances; the robust value function is defined as the expected cumulative discounted reward function under the worst-case scenario, i.e., the robust Markov decision process in the policy... The following optimization objectives The formula is expressed as: ; in, Represents an indeterminate set. Indicates the state transition probability and strategy The expected operator below, Indicates the first The discount weight of the reward for each time step. Representing state Next action Instant rewards obtained at that time Indicates the initial state is From the above formulas, it can be seen that the adversarial agent is responsible for minimizing the inner layer reward, while the robust agent is responsible for maximizing its own reward under uncertainty in the outer layer; thus, the solution can be obtained. It can gradually learn scheduling strategies that are highly adaptable to environmental disturbances.
[0035] S4. The robust adversarial reinforcement learning algorithm MB-RAC is used to solve the robust Markov decision process RMDP, and the day and night adapted robust real-time scheduling strategy is output to drive the energy flow transmission and coordinated operation of each unit in the integrated energy system architecture of the ocean island group. In a preferred embodiment, the robust adversarial reinforcement learning algorithm is employed as follows: Introducing a risk distortion mechanism into the optimization objective of adversarial agents: when calculating the value of interference optimization, the standard expectation of the reward function R is replaced with the risk distortion expectation. The risk distortion expectation amplifies the weight of low-reward scenarios based on quantile transformation to generate interference scenarios that are closer to the actual risk characteristics. A regularized baseline mechanism is introduced in the policy optimization of robust agents: the regularized baseline is obtained by weighting the value of robust scheduling actions generated under the current policy; the advantage function is obtained by subtracting the regularized baseline from the value of robust actions, and the policy network parameters of robust agents are optimized based on the advantage function to increase the randomness of generating scheduling policies.
[0036] In this embodiment, the specific process of the MB-RAC algorithm is as follows: P1. Initialize the core parameters of the dynamic environment model of the integrated energy system of the island cluster: Based on step S1, we construct the integrated energy system for the offshore island cluster and the state space defined in step S3. Action space The reward function R, and all states. neighboring state set Establish an initial system dynamic model; Initialize a robust agent policy network Value network parameters And copy the parameters to obtain the target value network parameters of the robust agent. Initialize the policy network parameters of the adversarial agent. Value network parameters And copy the parameters to obtain the target value network parameters of the adversarial agent. ; Set hyperparameters, including: model learning rate Robust agent policy / value network learning rate Adversarial agent policy / value network learning rate Discount factor Uncertainty level Target network update coefficients Model update cycle Batch sampling size: 64.
[0037] P2. Model-assisted dynamic environmental modeling: Robust agents and adversarial agents interact in a real island environment, collecting tuple data such as state, action, disturbance, reward, and next state, and storing them in the experience replay pool D; Based on the data in the experience pool, the environment model is trained using the temporal difference learning method, and the model parameters are updated every T steps to reduce the interaction cost in the real environment. By using error feedback between real-world environmental sampling data and model prediction data, the boundary of the uncertainty set P is dynamically adjusted to ensure that the model output conforms to the actual fluctuation characteristics of the island environment.
[0038] P3, Adversarial Agent Training: From the experience replay pool Randomly selected sample data x is the sample number; Calculate the risk distortion expectation, which is the reward in the sample. The expected risk distortion is calculated using quantile transformation, and the formula is as follows: ; in, , for right The first derivative, Let R be the quantile function of the reward function R. Using quantile variables, this calculation amplifies the weight of low-reward scenarios, making the adversarial agent more focused on low-probability, high-risk situations. Based on risk-distorted expectations, the value network parameters of the adversarial agent are updated, enabling it to more accurately assess the destructive extent of interference actions.
[0039] P4. Robust Agent Training: Regarding the current state and interference scenarios Calculate the regularized baseline , represented as: ; in, For action space; Candidate actions; Regularization coefficient; regularization term This is used to enhance the exploratory nature of strategies; The action value assessment results are adjusted by constructing a dominance function, which is: Based on the aforementioned advantage function, optimize the policy network parameters of the robust agent. To update, the formula is expressed as: ; in, Let the policy objective function of the robust agent be denoted as . Indicates the policy parameters The gradient is used to optimize the policy network; As the expectation operator, the gradient of the state-action distribution is averaged to obtain a stable update direction; By introducing a regularized baseline, the robust agent can avoid its policy from becoming fixed to a local optimum too early during learning, thus maintaining its exploratory capabilities. By solving the robust Markov decision process RMDP using the MB-RAC described above, a day-night adaptive robust real-time scheduling strategy can be output to drive the energy flow transmission and coordinated operation of each unit in the integrated energy system architecture of the ocean island cluster.
[0040] In summary, this invention proposes a day-night dual-mode intelligent scheduling method for energy flow in ocean-going island clusters based on robust reinforcement learning. It fully considers the layout characteristics of island clusters, renewable energy endowments, and mobile energy storage characteristics of battery-swapping vessels, effectively addressing uncertainties such as renewable energy fluctuations and transportation interruptions. It achieves adaptive matching of load demand on load islands, solving the problems of energy retention, supply shortages, and high operating costs associated with traditional methods, and providing technical support for energy self-sufficiency and sustainable development of ocean-going island clusters.
[0041] In the description of this specification, the references to terms such as "one embodiment," "some embodiments," "example," "specific example," or "some examples," etc., indicate that a specific feature, structure, material, or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of this application. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples. Moreover, without contradiction, those skilled in the art can combine and integrate the different embodiments or examples described in this specification, as well as the features of those different embodiments or examples.
[0042] The logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as a sequenced list of executable instructions for implementing logical functions, and can be embodied in any computer-readable medium for use by, or in conjunction with, an instruction execution system, apparatus or device (such as a computer-based system, a processor-included system or other system that can fetch and execute instructions from, an instruction execution system, apparatus or device).
[0043] The above embodiments provide a detailed description of the present invention. Specific examples have been used to illustrate the principles and implementation methods of the present invention. The descriptions of the above embodiments are only for the purpose of helping to understand the method and core ideas of the present invention. At the same time, for those skilled in the art, there will be changes in the specific implementation methods and application scope based on the ideas of the present invention. Therefore, the content of this specification should not be construed as a limitation of the present invention.
Claims
1. A day-night dual-mode intelligent scheduling method for energy flow in offshore island groups based on robust reinforcement learning, characterized in that, Specifically, the following steps are included: Construct a comprehensive energy system for the island clusters in the open ocean, forming an energy flow transmission framework for the island clusters, encompassing energy production, energy conversion, inter-island transportation, and energy consumption; Based on the energy flow transmission architecture of island clusters, and considering the differences in energy supply and demand between day and night in the ocean island clusters, a two-stage scheduling mechanism including daytime scheduling mode and nighttime scheduling mode is established, and daytime and nighttime scheduling optimization objectives are formulated. The problem of energy flow scheduling in remote island groups is modeled as a robust Markov decision process, and a dual-agent collaborative decision-making system adapted to day and night is constructed to achieve the optimization objective under the two-stage scheduling mechanism. A robust adversarial reinforcement learning algorithm is used to solve the robust Markov decision process, and outputs a day-night adapted robust real-time scheduling strategy to drive the energy flow transmission and coordinated operation of each unit in the integrated energy system architecture of the ocean island cluster.
2. The day-night dual-mode intelligent scheduling method for energy flow in offshore island groups based on robust reinforcement learning as described in claim 1, characterized in that, The construction of a comprehensive energy system for a group of distant-water islands, forming an energy flow transmission architecture for energy production, energy conversion, inter-island transportation, and energy consumption, specifically involves: Based on the geographically separated distribution characteristics of resource islands and load islands in the ocean island cluster, the islands are divided into resource islands that undertake energy production and energy conversion functions and load islands that consume energy. Battery swapping vessels are configured as mobile energy storage and transportation carriers connecting resource islands and load islands for cross-island transportation. The resource island is equipped with renewable energy units, energy conversion devices, and energy storage systems. The renewable energy units include wind power and photovoltaic units, and the energy conversion devices include electrolyzers, nitrogen production modules, ammonia synthesis reactors, and seawater desalination devices. Under the premise that the resource island can meet its own load, the energy transported by battery-swapping vessels across the island, and the power supply needs of the load island, the surplus electrical energy will be converted into energy through the electro-ammonia conversion device and the seawater desalination device. The load island is equipped with a controllable power supply and an energy storage system; the battery swapping vessel is equipped with an energy storage system; the energy storage system of the battery swapping vessel includes storing the electrical energy required for transporting the load island and the energy converted from the energy of the resource island for cross-island transportation; the energy storage systems of the resource island and the load island are used to store and allocate electrical energy or energy converted from energy. For renewable energy generating units, we will construct a power output prediction model that considers prediction errors; for battery-swapping vessels, we will construct a navigation energy consumption model that considers real-time changes in sea state and load; for controllable power sources, we will construct a dynamic power source model that considers efficiency adjustment and response lag; and for energy storage systems, we will construct a dynamic energy storage model that considers charging and discharging efficiency. The above model is used to characterize the uncertainties in the energy flow transmission architecture of island clusters, providing a decision-making basis for subsequent robust scheduling.
3. The day-night dual-mode intelligent scheduling method for energy flow in offshore island groups based on robust reinforcement learning according to claim 2, characterized in that, The energy flow transmission architecture based on island clusters considers the differences in energy supply and demand between day and night in the offshore island clusters, establishes a two-stage scheduling mechanism including daytime and nighttime scheduling modes, and formulates specific daytime and nighttime scheduling optimization objectives as follows: Based on the differences in energy supply and demand between day and night, the scheduling cycle is divided into daytime and nighttime periods, and corresponding daytime and nighttime scheduling modes are established respectively. The two scheduling modes take minimizing operating costs as the optimization objective. The daytime dispatch mode takes into account high renewable energy output, concentrated production and public service loads, and cross-island transportation by battery swapping vessels. The daytime operating costs include: the navigation power consumption and charging and discharging losses incurred by battery swapping vessels in performing cross-island transportation tasks, the cost of renewable energy abandonment caused by the overcapacity of photovoltaic and wind power output on resource islands, and the cost of controllable loads cut off during the daytime period. The nighttime dispatch mode takes into account the low output of renewable energy, the main load of residential life and the high dependence on energy storage. Its optimization goal is to minimize the nighttime operating cost. The nighttime operating cost includes: the cost of controllable load cut off during the nighttime period, the power supply cost on the load island, and the cost of renewable energy abandonment on the resource island due to the continued wind output. A unified day and night scheduling optimization target is constructed, and a dynamic weighting mechanism and scene identification factor are introduced to weight and fuse the daytime optimization target and the nighttime optimization target. The dynamic weight is adaptively adjusted according to the real-time energy supply and demand tension. When the energy supply is tight, the weight of the nighttime optimization target is increased, and when the energy supply is sufficient, the weight of the daytime optimization target is increased. At the same time, cross-stage collaborative constraints are set, including: at the end of daytime scheduling, the remaining amount of electrical energy in the energy storage systems of the resource island and the load island is not lower than a preset threshold; the ammonia and fresh water reserves formed by energy conversion in the resource island can meet the consumption needs of the load island.
4. The day-night dual-mode intelligent scheduling method for energy flow in ocean-going island groups based on robust reinforcement learning according to claim 2, characterized in that, The proposed modeling of the day-night dual-mode scheduling problem of energy flow in offshore island groups as a robust Markov decision process, and the construction of a day-night adapted dual-agent collaborative decision-making system to achieve the optimization objective under the two-stage scheduling mechanism, specifically includes: The day-night dual-mode scheduling problem of energy flow in a remote island group is modeled as a robust Markov decision process (RMDP), defining its state space S, action space A, reward function R, and uncertainty set P; the reward function is the negative of the day-night scheduling optimization objective; the uncertainty set P is contained in the state space S. Execute action Then transition to each possible state All possible transition probabilities; Based on the uncertain set P, a dual-agent collaborative decision-making system consisting of a robust agent and an adversarial agent is constructed. The adversarial agent is used to simulate the worst-case scenario by minimizing the reward function, while the robust agent is used to learn the energy scheduling strategy under the worst-case scenario and output a robust scheduling scheme that adapts to the switching between day and night scenarios.
5. A day-night dual-mode intelligent scheduling method for energy flow in ocean-going island groups based on robust reinforcement learning, as described in claim 4, is characterized in that... The state space S includes: Resource-side state variables used to characterize the energy output and energy storage status of the resource island; transportation-side state variables used to characterize the operating status and energy storage status of the battery swapping vessel; and load-side state variables used to characterize the load demand and local power output of the load island.
6. The day-night dual-mode intelligent scheduling method for energy flow in ocean-going island groups based on robust reinforcement learning according to claim 4, characterized in that, The action space A includes: Energy storage scheduling variables are used to control the charging and discharging of the electrical energy storage section in the resource island energy storage system; ship scheduling variables are used to schedule the navigation and berthing of battery swapping vessels; and load control variables are used to adjust the output of controllable power sources and controllable loads in the load island.
7. A day-night dual-mode intelligent scheduling method for energy flow in ocean-going island groups based on robust reinforcement learning, as described in claim 4, is characterized in that... The adversarial agent is used to simulate the worst-case scenario by minimizing the reward function, and the robust agent is used to learn an energy scheduling strategy under the worst-case scenario to output a robust scheduling scheme that adapts to day-night scene switching. Adversarial agents and robust agents consist of a policy network and a value network; The adversarial agent generates interference actions through a policy network to represent the adverse conditions that may occur during the state transition process, and fits the value of the interference actions through a value network to output the interference action with the minimum value, i.e. the interference action representing the worst condition, as the interference scenario. The robust agent generates robust scheduling actions based on the current state and the interference scenario output by the adversarial agent through its policy network, and fits the robust action value that can be obtained by executing the robust scheduling action under the interference scenario through its value network. Based on the robust action value, the policy network parameters of the robust agent are optimized, so that the robust agent learns the scheduling strategy that maximizes the robust value function under environmental interference. The robust value function is defined as the expected cumulative discount reward function under the worst case.
8. A day-night dual-mode intelligent scheduling method for energy flow in ocean-going island groups based on robust reinforcement learning, as described in claim 7, is characterized in that... The robust adversarial reinforcement learning algorithm employed in this context: Introducing a risk distortion mechanism into the optimization objective of adversarial agents: when calculating the value of perturbation optimization, the standard expectation of the reward function R is replaced with the risk-distorted expectation, which amplifies the weight of low-reward scenarios based on quantile transformation; A regularized baseline mechanism is introduced in the policy optimization of robust agents: the regularized baseline is obtained by weighting the value of robust scheduling actions generated under the current policy; the advantage function is obtained by subtracting the regularized baseline from the value of robust actions, and the policy network parameters of robust agents are optimized based on the advantage function.