Island energy system based on deep reinforcement learning and scheduling method thereof
By constructing a balanced constraint system with coupled electricity, hydrogen, oxygen, ammonia, and water multi-energy flow in an isolated energy system based on deep reinforcement learning, the problems of high carbonization, inefficient multi-energy synergy, and marginalization of ecological protection in isolated energy systems are solved, achieving low-carbon energy supply and ecological restoration, and optimizing the global relationship of multi-energy co-generation.
Patent Information
- Application Number
- CN202511192147.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-25
- Publication Date
- 2025-12-12
AI Technical Summary
Existing isolated energy systems suffer from problems such as high carbonization of energy supply, inefficiency of multi-energy coordination, and marginalization of ecological protection. Furthermore, the use of deterministic optimization algorithms in scheduling strategies leads to insufficient balancing capacity and difficulty in real-time scheduling.
An islanded energy system based on deep reinforcement learning is adopted. Water provided by the seawater desalination unit is decomposed into hydrogen and oxygen through an electrolysis device. The hydrogen and oxygen are stored and then supplied to the ammonia synthesis unit to produce ammonia. A multi-dimensional state vector is constructed in combination with a scheduling device to form a balanced constraint system with coupled electricity, hydrogen, oxygen, ammonia and water multi-energy flow. The near-end policy optimization algorithm is used for decision-making and outputs the optimal policy parameters.
It has achieved a three-dimensional closed loop of low-carbon energy supply, ecological restoration and load demand, optimized electrolytic hydrogen production and multi-energy cogeneration, improved the efficient conversion and ecological value of the energy system, and met the real-time dispatch requirements.
Smart Images

Figure CN121124178A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of energy system optimization, in particular to an island energy system based on deep reinforcement learning and a scheduling method thereof. BACKGROUND
[0002] For off-grid islands or islands, the existing energy system generally takes diesel generator set as the core energy supply unit, wind energy, solar energy and other renewable energy as supplementary power supply, and electrolytic hydrogen production device as passive energy storage unit to store surplus electricity to form a multi-energy complementary system.
[0003] However, the above energy system has many shortcomings: first, the energy supply is high-carbon, the intermittent output of wind and solar power generation is outstanding due to the influence of natural condition fluctuation, and the unique tidal energy and other marine resources of off-grid islands or islands are not effectively integrated and not included in the system scheduling strategy, resulting in low overall utilization rate of renewable energy. This single energy supply mode leads to long-term dependence on high-carbon power to maintain base load, and the carbon emission intensity far exceeds the ecological carrying threshold.
[0004] Second, the multi-energy collaboration is inefficient, the hydrogen production link is not dynamically coupled with seawater desalination, chemical synthesis and other load demands, and a cross-medium energy conversion network cannot be built. The hydrogen energy re-generation efficiency is significantly limited, and the nitrogen required for ammonia synthesis needs to be prepared with additional energy consumption, resulting in an inefficient cycle of the energy system. The isolated operation mode cannot smooth the fluctuations of renewable energy, nor can it respond to the spatial and temporal differences in demand for multiple types of loads.
[0005] Finally, the ecological protection is marginalized, a large amount of oxygen produced in the electrolytic hydrogen production process is usually directly discharged, not only wasting the associated resources of clean energy, but also causing the continuous deterioration of the near-shore ecological environment. At the same time, the pollutants emitted by traditional diesel power generation conflict with the demand for marine ecological protection.
[0006] For the above energy system, related scheduling strategies often use deterministic optimization algorithms, such as simplifying dynamic parameters such as wind and light output, load demand, etc. as fixed time sequence curves, which is difficult to describe the nonlinear characteristics of multi-energy flow coupling; or the step-by-step optimization framework separates the internal relationship between hydrogen production, energy storage scheduling and load response, leading to decision bias in multi-objective optimization. Such methods have insufficient balancing capacity in power fluctuation scenarios, and the calculation efficiency decreases sharply with the expansion of system scale, making it difficult to meet the real-time scheduling demand. SUMMARY
[0007] The present application provides an island energy system based on deep reinforcement learning and a scheduling method thereof, which aims to solve the technical problems of high-carbon energy supply, inefficient multi-energy collaboration, marginalized ecological protection, and insufficient balancing capacity and difficulty in real-time scheduling caused by the use of deterministic optimization algorithms in related technologies.
[0008] This application provides an isolated energy system based on deep reinforcement learning, which includes: an energy device, an electrolysis device, an energy storage battery, a hydrogen storage device, an oxygen storage device, a pressure swing adsorption device, an ammonia synthesis device, a seawater desalination device, and a scheduling device connected to each device respectively. The energy device serves as an energy input source, inputting electrical energy into the electrolysis device, energy storage battery, pressure swing adsorption device, and seawater desalination device. The energy device includes a diesel generator and a renewable energy device, which includes a photovoltaic generator, a wind generator, and a tidal generator. Based on the electrical energy input from the energy device, the electrolysis device decomposes the water provided by the seawater desalination device into hydrogen and oxygen; the pressure swing adsorption device provides nitrogen; the hydrogen storage device stores the hydrogen, which is then fed together with the nitrogen into the ammonia synthesis device for ammonia production; and the oxygen storage device stores the oxygen. The scheduling device constructs the real-time operating status of the isolated energy system into a multi-dimensional state vector, forming a balanced constraint system with coupled electricity-hydrogen-oxygen-ammonia-water multi-energy flow. Based on the near-end policy optimization algorithm in deep reinforcement learning, the device makes decisions on the multi-dimensional state vector and outputs the optimal policy parameters.
[0009] In one embodiment, the island energy system also includes a marine ranch that receives oxygen from the oxygen storage device.
[0010] This application also provides a scheduling method for an isolated energy system based on deep reinforcement learning, which applies the isolated energy system based on deep reinforcement learning as described in any of the above claims, and includes the following steps: The real-time operating state of an isolated energy system is constructed as a multi-dimensional state vector, represented as follows: ; in, for The output power of the photovoltaic power generation device at any time for The output power of wind power generation devices at any given time for The output power of the tidal power generation device at any time For load exist The demand at any given moment for Real-time energy storage battery capacity, for Capacity of hydrogen storage device at any time; Input the multidimensional state vector, make a decision based on the proximal policy optimization algorithm in deep reinforcement learning, and output the control action vector; By embedding physical constraints, the control action vector is constrained and filtered to construct a balanced constraint system with multi-energy flow coupling of electricity, hydrogen, oxygen, ammonia, and water. Calculate multidimensional real-time rewards by defining a reward function that includes operation and maintenance cost items, environmental protection cost items, economic benefit items, and penalty items to generate reinforcement learning reward signals; Store the multidimensional state vector, control action vector, and reinforcement learning reward signal, and update the policy parameters in batches; Determine whether the training rounds have reached the preset threshold; If so, stop learning and output the optimal policy parameters; If not, the closed-loop feedback executes the updated strategy parameters as the multi-dimensional state vector, and outputs the control action vector for the next round.
[0011] In one implementation, the embedded physical constraints, which constrain and filter the control action vector to construct a multi-energy flow coupled equilibrium constraint system of electricity-hydrogen-oxygen-ammonia-water, include: Set physical constraints, including the multi-energy flow balance constraint of electricity-hydrogen-oxygen-ammonia-water. The physical constraints are integrated into the multidimensional state vector; Based on the construction of a hybrid action space for each device, the control action vector is constrained and mapped to construct a balanced constraint system with multi-energy flow coupling of electricity, hydrogen, oxygen, ammonia, and water.
[0012] In one embodiment, setting the multi-energy flow balance constraint conditions for the electricity-hydrogen-oxygen-ammonia-water system includes: The energy balance constraint is defined as follows: ; in, for The charging and discharging power of the energy storage battery at all times. for The output power of the diesel generator at any given time. for Constant load power, for The output power of the electrolysis unit at any given time. for The output power of the seawater desalination unit at all times. for The charging and discharging power of the pressure swing adsorption device at constant time. for The output power of the ammonia synthesis unit at any given time; The hydrogen balance constraint is defined as follows: ; in, for The hydrogen production rate of the electrolysis unit at any given time. for The amount of hydrogen absorbed or released by the hydrogen storage device at any given time. for Hydrogen consumption of the ammonia synthesis unit at any given time; The oxygen balance constraint is defined as follows: ; in, for The oxygen production rate of the electrolysis unit at all times. The daily oxygen load of a marine ranch that receives oxygen from an oxygen storage device; The ammonia equilibrium constraint is defined as follows: ; in, for The ammonia production rate of the ammonia synthesis unit at any given time. This represents the daily ammonia load. The water balance constraints are defined as follows: ; in, for The water production rate of the seawater desalination unit at any given time. for Water consumption of the electrolysis unit at any given time; For water storage devices The amount of water stored or released at any given time.
[0013] In one implementation, the setting of physical constraints further includes: Set operational constraints for each device and solve them in conjunction with the multi-energy flow balance constraints of the electricity-hydrogen-oxygen-ammonia-water system to achieve constraint coupling.
[0014] In one implementation, when constraining the control action vector, the near-end policy optimization algorithm uses a shearing loss function for constraint protection.
[0015] In one implementation, the calculation of the multidimensional real-time reward, defining a reward function including a power generation cost term, an environmental cost term, an economic benefit term, and a penalty term, to generate a reinforcement learning reward signal includes: The reward function for reinforcement learning is defined as follows: ; in, For power generation costs, As an environmental cost item, For economic benefits, D For the penalty function term; Weighted generation of reinforcement learning reward signals.
[0016] In one implementation, the reward function includes the power generation cost item. Represented as: ; in, For maintenance costs, For fuel costs; The operation and maintenance costs Represented as: ; in, For the first Type of equipment Constant effort For the first The unit operation and maintenance cost of this type of equipment; The environmental cost item Represented as: ; in, For the types of polluting gases, To process the first The unit cost of the pollutants, for This device generates The emission coefficient of the polluting gas; The economic benefits item Represented as: ; in, Price per unit of oxygen; This refers to the unit price of ammonia. The penalty function term D Represented as: ; in, Penalty for unit difference in electricity consumption; , They are respectively The imbalance of total energy at any given moment and the over-discharge or over-charge of energy storage devices.
[0017] In one implementation, the calculation of the multidimensional real-time reward, defining the reward function as including a power generation cost item, an environmental cost item, an economic benefit item, and a penalty item, to generate a reinforcement learning reward signal further includes: The reward function is constrained and driven by hard constraint penalties.
[0018] The beneficial effects of the technical solutions provided in this application include: This application provides an isolated energy system based on deep reinforcement learning. Using an electrolysis unit as a hub, it utilizes electricity generated from fluctuating renewable energy sources such as wind, solar, and tidal power to convert water into hydrogen and oxygen. The hydrogen is then supplied to an ammonia synthesis unit via a hydrogen storage device to produce green ammonia, which in turn powers a seawater desalination unit to meet various load demands. The oxygen is stored in an oxygen storage device, forming an energy-ecological dual cycle of "fluctuating power source → hydrogen and oxygen dual products → chemical and ecological dual value output." Furthermore, a near-end policy optimization algorithm based on deep reinforcement learning is used to perceive the real-time dynamics of the isolated energy system. Operating status (such as electricity / hydrogen / water / oxygen / ammonia load demand, energy storage battery status) is constructed as a multi-dimensional state vector, forming a balanced constraint system of multi-energy flow coupling of electricity-hydrogen-oxygen-ammonia-water. Dynamic decision-making is used for real-time energy management and device collaborative control to output optimal strategy parameters (such as the output power of the electrolysis unit, the power of the ammonia synthesis unit and the adjustment of oxygen and ammonia loads, etc.), optimize the global relationship between electrolytic hydrogen production and multi-energy supply, realize a three-dimensional closed loop of low-carbon energy supply, ecological restoration and load demand, achieve efficient energy conversion and ecological value enhancement, and provide a systematic solution for the low-carbon transformation of island energy systems. Attached Figure Description
[0019] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0020] Figure 1 This is a schematic diagram of an isolated energy system based on deep reinforcement learning in one embodiment of the present invention.
[0021] Figure 2 This is a flowchart of a scheduling method for an isolated energy system based on deep reinforcement learning in one embodiment of the present invention.
[0022] Figure 3 This is a power balance analysis diagram of an islanded energy system in one embodiment of the present invention.
[0023] Figure 4 This is a hydrogen balance analysis diagram of an islanded energy system in one embodiment of the present invention.
[0024] Figure 5 This is an oxygen balance analysis diagram of an isolated energy system in one embodiment of the present invention.
[0025] Figure 6 This is a water balance analysis diagram of an islanded energy system in one embodiment of the present invention.
[0026] Figure 7This is a reward training curve for an islanded energy system in one embodiment of the present invention.
[0027] In the diagram: 1. Energy device; 11. Diesel power generation device; 12. Photovoltaic power generation device; 13. Wind power generation device; 14. Tidal power generation device; 2. Electrolysis device; 3. Energy storage battery; 4. Hydrogen storage device; 5. Oxygen storage device; 6. Pressure swing adsorption device; 7. Ammonia synthesis device; 8. Seawater desalination device. Detailed Implementation
[0028] To enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present application, and not all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of the present application.
[0029] This application provides an isolated energy system and its scheduling method based on deep reinforcement learning. The purpose is to solve the technical problems of existing isolated energy systems in related technologies, such as high carbonization of energy supply, low efficiency of multi-energy coordination, marginalization of ecological protection, and insufficient balancing capacity and difficulty in real-time scheduling due to the use of deterministic optimization algorithms in scheduling strategies.
[0030] like Figure 1 As shown, where, Figure 1 This is a schematic diagram of an isolated energy system based on deep reinforcement learning in one embodiment of the present invention.
[0031] This embodiment provides an isolated energy system based on deep reinforcement learning, which includes: an energy device 1, an electrolysis device 2, an energy storage battery 3, a hydrogen storage device 4, an oxygen storage device 5, a pressure swing adsorption device 6, an ammonia synthesis device 7, a seawater desalination device 8, and a scheduling device connected to each device respectively. Energy device 1 serves as an energy input source, inputting electrical energy into electrolysis device 2, energy storage battery 3, pressure swing adsorption device 4, and seawater desalination device 8. Energy device 1 includes diesel generator 11 and renewable energy device. Renewable energy device includes photovoltaic generator 12, wind power generator 13, and tidal power generator 14. Based on the electrical energy input from energy device 1, electrolysis device 2 decomposes the water provided by the seawater desalination device into hydrogen and oxygen; pressure swing adsorption device 6 provides nitrogen; hydrogen storage device 4 stores hydrogen, which is then fed together with nitrogen into ammonia synthesis device 7 for ammonia production; and oxygen storage device 5 stores oxygen. The scheduling device constructs the real-time operating status of the isolated energy system into a multi-dimensional state vector, forming a balanced constraint system with coupled electricity-hydrogen-oxygen-ammonia-water multi-energy flow. Based on the near-end policy optimization algorithm in deep reinforcement learning, the device makes decisions on the multi-dimensional state vector and outputs the optimal policy parameters.
[0032] This embodiment provides an isolated energy system based on deep reinforcement learning. Using an electrolysis unit as a hub, it utilizes electricity generated from fluctuating renewable energy sources such as wind, solar, and tidal power to convert water into hydrogen and oxygen. The hydrogen is then supplied to an ammonia synthesis unit via a hydrogen storage device to produce green ammonia, which in turn powers a seawater desalination unit to meet various load demands. The oxygen is stored in an oxygen storage device, forming an energy-ecological dual cycle of "fluctuating power source → hydrogen and oxygen dual products → chemical and ecological dual value output." Furthermore, a near-end policy optimization algorithm based on deep reinforcement learning is used to perceive the real-time status of the isolated energy system. The system constructs a multi-dimensional state vector based on the real-time operating status (such as electricity / hydrogen / water / oxygen / ammonia load demand and energy storage battery status), forming a balanced constraint system with multi-energy flow coupling of electricity, hydrogen, oxygen, ammonia, and water. Dynamic decision-making enables real-time energy management and coordinated control of devices to output optimal strategy parameters (such as the output power of the electrolysis unit, the power of the ammonia synthesis unit, and the adjustment of oxygen and ammonia loads). This optimizes the global relationship between electrolytic hydrogen production and multi-energy supply, achieving a three-dimensional closed loop of low-carbon energy supply, ecological restoration, and load demand. It also promotes efficient energy conversion and enhances ecological value, providing a systematic solution for the low-carbon transformation of island energy systems.
[0033] In an isolated energy system, the mathematical models of each device are as follows: The fuel cost of the diesel generator set 11 is its consumption characteristic function, expressed as: ; in, Fuel costs for diesel generator sets; This refers to the output power of the diesel generator set; , , This is a coefficient representing the fuel cost of diesel generator sets.
[0034] Photovoltaic power generation device 12 is represented as: ; in, This is the reduction factor; , Photovoltaic power generation devices t The actual output power at any given time and the maximum output power under standard test conditions; , They are respectively t The actual light intensity at any given time and the light intensity under standard test conditions; The power temperature coefficient; , They are respectively t The actual temperature at any given time and the temperature under standard test conditions.
[0035] Wind power generation device 13 is represented as: ; in: , Wind power generation devices t Output power and rated power at any given time; , , , Wind power generation devices t Real-time wind speed, cut-in wind speed, cut-out wind speed, and rated wind speed.
[0036] Tidal power generation device 14 is represented as: ; in, Tidal power generation device Output power at any given moment; The energy capture coefficient of a tidal power generation device; The density of seawater; The area swept by the blades of the tidal power generation device; , , These are the tidal current velocity, the cut-in velocity of the tidal power generation device, and the rated velocity, respectively. This refers to the rated output power of the tidal power generation device.
[0037] The aforementioned energy device 1 serves as an energy input source, transmitting electrical energy to the electrolysis device 2 in the middle via electrical energy flow.
[0038] Electrolysis device 2 is represented as: ; in, Electrolysis device t Hydrogen production at any given time, in kg; The efficiency of the electrolysis unit; Electrolysis device t Output power at any given moment; The electro-hydrogen conversion coefficient is given.
[0039] Energy storage battery 3 smooths out power fluctuations, as shown below: ; in, , For energy storage batteries Time and The battery status at any given time; This represents the self-discharge coefficient of the energy storage battery. , These represent the charging and discharging power of the energy storage battery; , These represent the charge and discharge efficiencies of the energy storage battery.
[0040] Hydrogen storage device 4 is represented as: ; In the formula: , Hydrogen storage devices Time and The amount of hydrogen stored at any given time; , They are respectively The amount of hydrogen absorbed and released by the hydrogen storage device at any given time; This represents the hydrogen loss rate.
[0041] Pressure swing adsorption device 6 is represented as: ; in, The operating energy consumption of the pressure swing adsorption (PSA) unit. Energy consumption for producing one unit volume of nitrogen for PSA The volume of nitrogen produced by PSA.
[0042] In ammonia synthesis unit 7, the volume of ammonia synthesized is determined using a hydrogen fixation ammonia method, expressed as: ; in, , , These are the operating energy consumption, ammonia volume produced, and hydrogen volume consumed for P2A, respectively. The energy consumption required to produce one unit volume of ammonia for P2A.
[0043] Seawater desalination unit 8 is represented as: ; In the formula: for t The water production rate of a seawater desalination plant at any given time is the volume of fresh water produced. ; for t The electrical power consumed by the seawater desalination unit at all times. ; The water production rate of the seawater desalination unit.
[0044] A water storage device 81 is provided between the seawater desalination device 8 and the electrolysis device 2.
[0045] Water storage device 81 is represented as: ; in, , The water storage device is located at Time and The amount of water stored at any given time is the volume of fresh water stored. The self-loss rate of the water storage device; and The water storage device is located at The volume of freshwater stored and released at any given time; and These refer to the storage and release efficiencies of the water storage device, respectively.
[0046] In one embodiment, the island energy system also includes a marine ranch that receives oxygen from the oxygen storage device 5.
[0047] The specific functions of the dispatching device will be explained in detail below in conjunction with the dispatching method of the isolated energy system.
[0048] like Figure 2 As shown, Figure 2 This is a flowchart of a scheduling method for an isolated energy system based on deep reinforcement learning in one embodiment of the present invention.
[0049] This embodiment also provides a scheduling method for an isolated energy system based on deep reinforcement learning. The method, applied to the aforementioned isolated energy system, includes the following steps: Step S1: The real-time operating state of the isolated energy system is constructed as a multi-dimensional state vector, represented as follows: ; in, for The output power of the photovoltaic power generation device at any time for The output power of wind power generation devices at any given time for The output power of the tidal power generation device at any time For load exist The demand at any given moment for Real-time energy storage battery capacity, for Capacity of hydrogen storage device at any time; Step S2: Input a multi-dimensional state vector, make a decision based on the proximal policy optimization algorithm in deep reinforcement learning, and output a control action vector; Step S3: Embed physical constraints, perform constraint filtering on the control action vector, and construct a balanced constraint system with multi-energy flow coupling of electricity, hydrogen, oxygen, ammonia, and water. Step S4: Calculate multidimensional real-time rewards and define a reward function that includes operation and maintenance cost items, environmental protection cost items, economic benefit items, and penalty items to generate reinforcement learning reward signals; Step S5: Store the multi-dimensional state vector, control action vector, and reinforcement learning reward signal, and update the policy parameters in batches; Step S6: Determine whether the number of training rounds has reached the preset threshold; If so, stop learning and output the optimal policy parameters; If not, the closed-loop feedback executes the updated strategy parameters as a multi-dimensional state vector, and outputs the control action vector for the next round.
[0050] Through the above scheme, the island energy system dynamically coordinates the fluctuating power generation of wind and solar power, the timing of hydrogen storage and release, the elasticity of ammonia load, and the ecological release strategy of oxygen based on the deep reinforcement learning strategy. This achieves precise matching of the five-dimensional energy flow of electricity, hydrogen, oxygen, fresh water, and green ammonia at the electricity load, hydrogen load, water load, oxygen load, and ammonia load ends, forming a closed loop of the entire chain of renewable energy consumption, multi-energy supply, and marine ecological restoration. This overcomes the dual challenges of high carbon dependence and fragmented ecological governance in island systems.
[0051] In response to the multi-objective and highly dynamic characteristics of the operation and scheduling of isolated energy systems, this application adopts a deep reinforcement learning algorithm to construct an optimized architecture and realizes the coordinated control of all system equipment through state-action space mapping.
[0052] Furthermore, the Proximal Policy Optimization (PPO) algorithm is employed. PPO is a deep reinforcement learning algorithm proposed by OpenAI in 2017. It is an improved version of the Policy Gradient method, aiming to optimize the policy through a more stable training method while avoiding the training instability problem common in traditional Policy Gradient methods.
[0053] The PPO algorithm structure is as follows: (1) State Space: Defines the environmental observation vector, including device state parameters, dynamic input signals and external environmental variables; (2) Action Space: Outputs continuous control commands and supports multi-dimensional action decision-making; (3) Policy network (Actor): Generates the probability distribution of actions and parameterizes it as a deep neural network; (4) Value Network (Critic): Evaluates the state value function and provides a benchmark for policy updates.
[0054] Its core mechanism is as follows: (1) Proximal policy constraint: Limit the policy update magnitude by cutting the objective function to ensure training stability.
[0055] Represented as: ; in, Here is the shear loss function; For expectation operators; The probability ratio of the strategy; The shear threshold; It is the shearing function; This is the dominant function.
[0056] (2) Generalized advantage estimation (GAE): balance the bias and variance, and calculate the advantage value.
[0057] Represented as: ; in, This is the estimate of the generalized advantage. It is an exponential decay factor; and These are the current and next time-series differential errors (TD errors), respectively. For instant rewards; Discount factor; and The current state is respectively Value function and next state The value function of .
[0058] (3) Experience Replay: Storing historical data Random sampling updates network parameters.
[0059] In step S1, for an isolated energy system, the information provided by the environment to the agent is the output of renewable energy, load, and the state of charge of energy storage. Therefore, the real-time operating state of the isolated energy system is defined as a multi-dimensional state vector representation. Specifically, the power generation of photovoltaic power generation device 12, wind power generation device 13, and tidal power generation device 14 is collected in real time, and the state of charge of energy storage battery 3, as well as the load requirements for electricity, hydrogen, oxygen, ammonia, and water, are integrated to construct a standardized state vector, which covers wind / solar / tidal fluctuations, energy storage status, and ecological load requirements (such as oxygen release).
[0060] In step S2, the PPO algorithm is used to process the state vector and output the control action vector, such as the reference power command of electrolysis device 2, hydrogen storage release rate, and control actions such as electric cooling and heating load adjustment.
[0061] In one embodiment, step S3, embedding physical constraints and filtering the control action vector to construct a multi-energy flow coupling equilibrium constraint system of electricity-hydrogen-oxygen-ammonia-water, includes: Step S31: Set physical constraints, including the balance constraints of multiple energy flows of electricity-hydrogen-oxygen-ammonia-water. Step S32: Integrate the physical constraints into a multi-dimensional state vector; Step S33: Construct a hybrid action space based on each device, and perform constraint mapping on the control action vector to construct a balanced constraint system with multi-energy flow coupling of electricity, hydrogen, oxygen, ammonia, and water.
[0062] Here, we define the action space. After observing the state information of the environment, the agent determines its own policy set. In action space Choose one action from the options. Therefore, based on the physical characteristics of the device, a hybrid action space is constructed as follows: ; in, for The output power of the energy conversion device at any given time; for The energy input or output of the energy storage device at any given moment. The action space mapping control variables cover energy conversion (electrolysis unit), energy storage scheduling (hydrogen storage) and ecological load (oxygen regulation), and chemical production (ammonia regulation), to achieve global closed-loop optimization.
[0063] Furthermore, constraint mapping is applied to the action space to achieve dynamic coordination. Taking action masking as an example, illegal actions are disabled based on the current state. Specifically, for example, when the hydrogen storage device is full ( =100%), shielding the hydrogen absorption action to prevent overflow. For example, when the SOC of energy storage battery 3 is below 20%, the discharge action is prohibited to avoid over-discharge. Taking projection as another example, the original actions output by the strategy network are post-processed. Specifically, if the power command of electrolysis device 2... Out of range, forcibly truncate to or For example, if the release rate of the hydrogen storage device... If the flow rate exceeds the valve's maximum capacity, the upper limit of the device shall apply.
[0064] In one embodiment, step S31, setting the multi-energy flow balance constraint condition for the electricity-hydrogen-oxygen-ammonia-water includes: The energy balance constraint is defined as follows: ; in, for The charging and discharging power of the energy storage battery at all times. for The output power of the diesel generator at any given time. for Constant load power, for The output power of the electrolysis unit at any given time. for The output power of the seawater desalination unit at all times. for The charging and discharging power of the pressure swing adsorption device at constant time. for The output power of the ammonia synthesis unit at any given time; The hydrogen balance constraint is defined as follows: ; in, for The hydrogen production rate of the electrolysis unit at any given time. for The amount of hydrogen absorbed or released by the hydrogen storage device at any given time. for Hydrogen consumption of the ammonia synthesis unit at any given time; The oxygen balance constraint is defined as follows: ; in, for The oxygen production rate of the electrolysis unit at all times. The daily oxygen load of a marine ranch that receives oxygen from an oxygen storage device; The ammonia equilibrium constraint is defined as follows: ; in, for The ammonia production rate of the ammonia synthesis unit at any given time. This represents the daily ammonia load. The water balance constraints are defined as follows: ; in, for The water production rate of the seawater desalination unit at any given time. for Water consumption of the electrolysis unit at any given time; For water storage devices The amount of water stored or released at any given time.
[0065] Through the above scheme, the power balance constraint achieves real-time supply and demand equilibrium, ensuring that the power generation side and the power consumption side are consistent, forcing instantaneous power balance between power generation, energy storage, and load, and avoiding voltage / frequency instability. The hydrogen balance constraint, through the buffering capacity of the hydrogen storage device, mitigates the impact of renewable energy fluctuations on the hydrogen energy chain and prevents interruptions in fuel cell hydrogen supply. The hydrogen-ammonia-water multi-energy flow coupling buffer, and the hydrogen-ammonia-water synergistic balance (flexible control of material circulation) decouples rigid loads (such as electrical loads) from flexible production units (such as seawater desalination units). When there is a power surplus, the power of the electrolysis unit is increased to increase green hydrogen production, driving the green ammonia synthesis unit to store it. Chemical raw materials; when there is a power shortage, green ammonia fuel is used for power generation and the priority of seawater desalination is reduced, forming a cross-media buffer mechanism of "hydrogen storage for electricity and ammonia regulation for water"; oxygen balance constraints bind oxygen production to ecological restoration needs, avoiding excessive oxygen production and waste or insufficient oxygen supply leading to ecological degradation of marine ranches, improving the sustainable stability of the system from the perspective of resource cycle, and achieving ecological energy synergy; therefore, a balance constraint system of multi-energy flow coupling of electricity-hydrogen-oxygen-ammonia-water is constructed. By matching the supply and demand relationship of multi-energy flow in real time, the dynamic coupling and mutual support of each energy conversion link are ensured, thereby maintaining the dynamic stability of multi-energy flow coordinated operation of the island energy system.
[0066] In one embodiment, step S31, setting physical constraints further includes: Set the operating constraints for each device and solve them together with the multi-energy flow balance constraints of electricity-hydrogen-oxygen-ammonia-water to achieve constraint coupling.
[0067] The operating constraints for each device are as follows: The operating constraints of the diesel generator set 11 are expressed as follows: ; in, , These are the minimum and maximum output power of the diesel generator set, respectively.
[0068] The operating constraints of the photovoltaic power generation device 12 are expressed as follows: ; in, , These are the minimum and maximum output power of the photovoltaic power generation device, respectively.
[0069] The equipment operation constraints of wind power generation device 13 are expressed as follows: ; in, , These are the minimum and maximum output power of the wind power generation device, respectively.
[0070] The equipment operation constraints of the tidal power generation device 14 are expressed as follows: ; in, , These represent the minimum and maximum output power of the tidal power generation device, respectively.
[0071] The operating constraints of electrolysis unit 2 are expressed as follows: ; in, , These are the minimum and maximum output power of the electrolysis unit, respectively.
[0072] The operating constraints of energy storage battery 3 are expressed as follows: ; in, , These are the minimum and maximum capacities of the energy storage battery, respectively. , These represent the minimum and maximum charge / discharge power of the energy storage battery per unit time, respectively.
[0073] The operating constraints of hydrogen storage device 4 are expressed as follows: ; in, , These are the minimum and maximum capacities of the hydrogen storage device, respectively. for The amount of hydrogen absorbed or released by the hydrogen storage device at any given time; , These represent the minimum and maximum hydrogen absorption / desorption rates of the hydrogen storage device per unit time, respectively.
[0074] The operating constraints of the pressure swing adsorption device 6 are expressed as follows: ; in, , These are the minimum and maximum output power of the pressure swing adsorption device, respectively.
[0075] The equipment operation constraints of ammonia synthesis unit 7 are expressed as follows: ; in, , These are the minimum and maximum output power of the pressure swing adsorption device, respectively.
[0076] The operating constraints of the seawater desalination unit 8 are expressed as follows: ; in, , These are the minimum and maximum output power of the seawater desalination unit, respectively.
[0077] The operating constraints of the water storage device 81 are expressed as follows: ; in, , These are the minimum and maximum capacities of the water storage device, respectively. for The volume of fresh water stored in the water storage device at any given time; , These represent the minimum and maximum freshwater storage volumes of the water storage device per unit time, respectively.
[0078] Through the above scheme, the device operation constraints, by limiting the physical boundaries of the power and capacity of each device (such as maximum / minimum output limits and energy storage state of charge range), force each device to operate within its optimal efficiency range, avoiding ineffective energy consumption under low load and the risk of losses under over-limit conditions. Simultaneously, power change rate constraints suppress dynamic response overshoot, ensuring the safe operation and lifespan of each device while locking the overall system energy efficiency within the peak superposition region of the efficiency curves of each unit, forming a stable and efficient multi-energy conversion link. The PPO strategy senses the device's safety margin in real time. Furthermore, the multi-energy flow balance constraints (electricity-hydrogen-oxygen-water-ammonia coupling) are integrated into the state vector to reflect system-level supply and demand gaps. These two constraints are jointly optimized and coupled, such as correcting hydrogen storage deviations, verifying load power balance, and applying gradient penalties for overcharging and discharging of energy storage. The physical constraints serve as the environment model for reinforcement learning, ensuring that the strategy output naturally conforms to physical laws.
[0079] In one embodiment, when performing constraint filtering on the control action vector, the near-end policy optimization algorithm uses a shearing loss function for constraint protection.
[0080] Specifically, the shearing loss function limits the policy update magnitude (shearing threshold ϵ=0.1) to prevent policy mutations during training from exceeding the limit. For example, if the new policy attempts to update the policy... The power output suddenly increased from 80kW to 120kW (exceeding the limit). =100kW), the shearing mechanism will reduce the updated weights, forcing the strategy to converge to the compliance range.
[0081] In one embodiment, step S4, calculating a multidimensional real-time reward, defines a reward function including a power generation cost term, an environmental cost term, an economic benefit term, and a penalty term, to generate a reinforcement learning reward signal, including: The reward function for reinforcement learning is defined as follows: ; in, For power generation costs, As an environmental cost item, For economic benefits, D For the penalty function term; Weighted generation of reinforcement learning reward signals.
[0082] In one embodiment, the reward function includes the power generation cost item. Represented as: ; in, For maintenance costs, For fuel costs; Operation and maintenance costs Represented as: ; in, For the first Type of equipment Constant effort For the first The unit operation and maintenance cost of this type of equipment; Environmental cost item Represented as: ; in, For the types of polluting gases, To process the first The unit cost of the pollutants, for This device generates The emission coefficient of the polluting gas; Economic benefits Represented as: ; in, Price per unit of oxygen; This refers to the unit price of ammonia. penalty function term D Represented as: ; in, Penalty for unit difference in electricity consumption; , They are respectively The imbalance of total energy at any given moment and the over-discharge or over-charge of energy storage devices.
[0083] The above scheme primarily considers operation and maintenance costs and fuel costs for power generation, while the environmental protection cost primarily considers system costs. , as well as The emission treatment cost is considered. The economic benefit term mainly considers the ecological benefits (quantified as the oxygenation revenue of marine ranches) and the sales revenue of ammonia. The penalty term ensures that each variable meets the energy balance constraint and energy storage constraint during the iteration process. The constructed reward function comprehensively considers multiple dimensions such as power generation economics, environmental governance costs and system constraint satisfaction, and optimizes the ecological benefits (oxygenation), chemical product economic benefits (ammonia production) in conjunction with the economic goal.
[0084] In one embodiment, step S4, calculating the multidimensional real-time reward, and defining the reward function including a power generation cost term, an environmental cost term, an economic benefit term, and a penalty term, to generate a reinforcement learning reward signal further includes: The reward function is driven by hard constraint penalties.
[0085] Specifically, this could involve imposing penalties for transgressions. For example, setting... When the SOC exceeds the limit, the penalty increases sharply, forcing the strategy to be quickly corrected.
[0086] Step S5 involves performing experience replay and distributed training; Step S6 involves termination determination and outputting optimization results.
[0087] Through the above scheme, the PPO algorithm is deeply coupled with the various devices of the isolated energy system, achieving dynamic coordination and precise matching. First, it has a natural adaptability to constraint handling. The isolated energy system needs to meet strict physical constraints. PPO's Clipped Surrogate Objective (CSO) ensures that the action space remains within the compliance range during training by limiting the policy update step size, avoiding the constraint violation risk caused by policy mutations in other algorithms (such as DDPG). Second, it efficiently explores the high-dimensional continuous action space. The control variables of the isolated energy system include continuous parameters such as electrolysis unit power, hydrogen storage rate, and flexible load adjustment. PPO's Stochastic Policy can simultaneously output multi-dimensional continuous actions and explore the optimal parameter combination through probability distribution. Compared with related algorithms: DQN / A3C: only supports discrete actions and cannot precisely control power regulation; DDPG: although it supports continuous actions, its deterministic strategy is prone to getting trapped in local optima and cannot cover the complex solution space of multi-energy flow coordination; finally, PPP has training stability and robustness. PPO's experience replay and multi-round small-batch update mechanism can effectively deal with the strong dynamics (wind and solar fluctuations, load abrupt changes) and sparse reward problem (such as the need for long-term accumulation of ecological benefits) of isolated energy systems. Compared with related algorithms: TRPO: although it also guarantees policy stability, it requires the calculation of complex second derivatives (Fisher information matrix), resulting in excessive computational cost; SAC: relies on entropy regularization to adjust the exploration intensity, and is prone to triggering over-limit penalties due to over-exploration in hard-constraint scenarios.
[0088] Specific embodiments are provided below to further illustrate this application.
[0089] Taking an island in the East China Sea as the research object, typical daily data were extracted based on the island's meteorological and load data for a certain year, and the energy system operation and scheduling method provided in this application was simulated and verified.
[0090] like Figures 3 to 7 As shown, where,Figure 3 This is a power balance analysis diagram of an islanded energy system in one embodiment of the present invention. Figure 4 This is a hydrogen balance analysis diagram of an islanded energy system in one embodiment of the present invention. Figure 5 This is an oxygen balance analysis diagram of an isolated energy system in one embodiment of the present invention. Figure 6 This is a water balance analysis diagram of an islanded energy system in one embodiment of the present invention. Figure 7 This is a reward training curve for an islanded energy system in one embodiment of the present invention.
[0091] like Figure 3 As shown, the power system integrates photovoltaic power generation device 12 (peak value > 150kW), tidal power generation device 14 (stable 50kW at night), and energy storage battery 3 with intelligent charging and discharging (±80-120kW), achieving full coverage of high-energy-consuming loads such as electrolysis device 2 (specifically, electrolyzer) and seawater desalination device 8 for 86% of the time. Only during the short period from 04:00 to 08:00 does it rely on diesel power generation device 11 (specifically, diesel engine) to make up for the gap, significantly improving the utilization rate of renewable energy.
[0092] like Figure 4 As shown, in the hydrogen balance, relying on the continuous hydrogen production (8-15 kg / h) of the electrolysis unit 2 (specifically the electrolyzer) and the precise peak regulation (±5-8 kg / h) of the hydrogen storage unit (specifically the hydrogen storage tank), the production and consumption are completely self-consistent (15 kg / h) during the period from 12:00 to 18:00, with a balance error of less than 5%.
[0093] like Figure 5 As shown, in the oxygen balance, the stable oxygen supply (5-20kW) through the electrolysis device 2 (specifically the electrolysis cell) is seamlessly connected with the demand of the marine ranch (5-15kW). Zero deviation between supply and demand is achieved from 00:00 to 08:00. During the other periods, excess capacity is efficiently absorbed through compression storage.
[0094] like Figure 6 As shown, in the water balance, the seawater desalination unit 8 provides a stable supply throughout the day (0-40kW), and its water supply capacity significantly covers the water demand (0-20kW). In particular, during the period from 03:00 to 07:00, the excess capacity provides flexible space for energy storage scheduling.
[0095] As can be seen from the above, the islanded energy system provided in this application embodiment takes multi-energy coupling and real-time regulation as its core, and achieves efficient dynamic balance through the synergy of water, electricity, hydrogen and oxygen multi-energy, achieving industrial-grade low-carbon operation efficiency with a renewable energy utilization rate of over 75% and a diesel engine dependence rate of less than 7%.
[0096] like Figure 7As shown, the convergence and stability of the PPO algorithm provided in this application embodiment are clearly demonstrated: the horizontal axis covers 0 to 200,000 training rounds, and the vertical axis quantifies the original reward and the moving average reward (range -2.25 to 0.50). The dark-colored original reward curve initially exhibits dramatic fluctuations (dominated by negative values), gradually converging to near 0 as training progresses, indicating that the strategy exploration is gradually optimizing in an effective direction. The light-colored moving average reward curve simultaneously verifies the improved stability, smoothing towards the 0.50 threshold in the later stages, reflecting the repeatability and anti-interference capability of the strategy improvement. Dual-axis data coupling shows that the system enters a stable learning phase after approximately 75,000 rounds, with reward fluctuations reduced by more than 90%, verifying the training algorithm's adaptability and convergence reliability for complex tasks, providing crucial data support for the engineering deployment of intelligent control strategies.
[0097] In the description of this application, it should be noted that the terms "upper," "lower," etc., indicating the orientation or positional relationship are based on the orientation or positional relationship shown in the accompanying drawings, and are only for the convenience of describing this application and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation, and therefore should not be construed as a limitation of this application. Unless otherwise expressly specified and limited, the terms "installed," "connected," and "linked" should be interpreted broadly. For example, they can refer to a fixed connection, a detachable connection, or an integral connection; they can refer to a mechanical connection or an electrical connection; they can refer to a direct connection or an indirect connection through an intermediate medium; they can refer to the internal communication between two elements. For those skilled in the art, the specific meaning of the above terms in this application can be understood according to the specific circumstances.
[0098] It should be noted that in this application, relational terms such as "first" and "second" are used merely to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.
[0099] In the description of the embodiments of this application, unless otherwise stated, " / " means "or". For example, A / B can mean A or B. The "and / or" in the text is merely a description of the relationship between related objects, indicating that there can be three relationships. For example, A and / or B can mean: A exists alone, A and B exist simultaneously, and B exists alone. In addition, in the description of the embodiments of this application, "multiple" means two or more.
[0100] The above are merely specific embodiments of this application, enabling those skilled in the art to understand or implement this application. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of this application. Therefore, this application is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features claimed herein.
Claims
1. An islanded energy system based on deep reinforcement learning, characterized in that, It includes an energy device (1), an electrolysis device (2), an energy storage battery (3), a hydrogen storage device (4), an oxygen storage device (5), a pressure swing adsorption device (6), an ammonia synthesis device (7), a seawater desalination device (8), and a dispatching device connected to each device respectively. The energy device (1) serves as an energy input source, inputting electrical energy into the electrolysis device (2), energy storage battery (3), pressure swing adsorption device (6), and seawater desalination device (8). The energy device (1) includes a diesel generator (11) and a renewable energy device, which includes a photovoltaic generator (12), a wind power generator (13), and a tidal power generator (14). Based on the electrical energy input from the energy device (1), the electrolysis device (2) decomposes the water provided by the seawater desalination device (8) into hydrogen and oxygen; the pressure swing adsorption device (6) provides nitrogen; the hydrogen storage device (4) stores the hydrogen and then inputs it together with the nitrogen into the ammonia synthesis device (7) for ammonia production; the oxygen storage device (5) stores the oxygen. The scheduling device constructs the real-time operating status of the isolated energy system into a multi-dimensional state vector, forming a balanced constraint system with coupled electricity-hydrogen-oxygen-ammonia-water multi-energy flow. Based on the near-end policy optimization algorithm in deep reinforcement learning, the device makes decisions on the multi-dimensional state vector and outputs the optimal policy parameters.
2. The isolated energy system based on deep reinforcement learning as described in claim 1, characterized in that, The isolated energy system also includes a marine ranch that receives oxygen from the oxygen storage device (5) in a directed manner.
3. A scheduling method for an isolated energy system based on deep reinforcement learning, applying the isolated energy system based on deep reinforcement learning as described in any one of claims 1 to 2, characterized in that, It includes the following steps: The real-time operating state of an isolated energy system is constructed as a multi-dimensional state vector, represented as follows: ; in, for The output power of the photovoltaic power generation device at any time for The output power of wind power generation devices at any given time for The output power of the tidal power generation device at any time For load exist The demand at any given moment for Real-time energy storage battery capacity, for Capacity of hydrogen storage device at any time; Input the multidimensional state vector, make a decision based on the proximal policy optimization algorithm in deep reinforcement learning, and output the control action vector; By embedding physical constraints, the control action vector is constrained and filtered to construct a balanced constraint system with multi-energy flow coupling of electricity, hydrogen, oxygen, ammonia, and water. Calculate multidimensional real-time rewards by defining a reward function that includes operation and maintenance cost items, environmental protection cost items, economic benefit items, and penalty items to generate reinforcement learning reward signals; Store the multidimensional state vector, control action vector, and reinforcement learning reward signal, and update the policy parameters in batches; Determine whether the training rounds have reached the preset threshold; If so, stop learning and output the optimal policy parameters; If not, the closed-loop feedback executes the updated strategy parameters as the multi-dimensional state vector, and outputs the control action vector for the next round.
4. The scheduling method for an isolated energy system based on deep reinforcement learning as described in claim 3, characterized in that, The embedded physical constraints, which filter the control action vector to construct a multi-energy flow coupled equilibrium constraint system of electricity, hydrogen, oxygen, ammonia, and water, include: Set physical constraints, including the multi-energy flow balance constraint of electricity-hydrogen-oxygen-ammonia-water. The physical constraints are integrated into the multidimensional state vector; Based on the construction of a hybrid action space for each device, the control action vector is constrained and mapped to construct a balanced constraint system with multi-energy flow coupling of electricity, hydrogen, oxygen, ammonia, and water.
5. The scheduling method for an isolated energy system based on deep reinforcement learning as described in claim 4, characterized in that, The equilibrium constraints for the multi-energy flow of electricity, hydrogen, oxygen, ammonia, and water include: The energy balance constraint is defined as follows: ; in, for The charging and discharging power of the energy storage battery at all times. for The output power of the diesel generator at any given time. for Constant load power, for The output power of the electrolysis unit at any given time. for The output power of the seawater desalination unit at all times. for The charging and discharging power of the pressure swing adsorption device at constant time. for The output power of the ammonia synthesis unit at any given time; The hydrogen balance constraint is defined as follows: ; in, for The hydrogen production rate of the electrolysis unit at any given time. for The amount of hydrogen absorbed or released by the hydrogen storage device at any given time. for Hydrogen consumption of the ammonia synthesis unit at any given time; The oxygen balance constraint is defined as follows: ; in, for The oxygen production rate of the electrolysis unit at all times. The daily oxygen load of a marine ranch that receives oxygen from an oxygen storage device; The ammonia equilibrium constraint is defined as follows: ; in, for The ammonia production rate of the ammonia synthesis unit at any given time. This represents the daily ammonia load. The water balance constraints are defined as follows: ; in, for The water production rate of the seawater desalination unit at any given time. for Water consumption of the electrolysis unit at any given time; For water storage devices The amount of water stored or released at any given time.
6. The scheduling method for an isolated energy system based on deep reinforcement learning as described in claim 4, characterized in that, The set physical constraints also include: Set operational constraints for each device and solve them in conjunction with the multi-energy flow balance constraints of the electricity-hydrogen-oxygen-ammonia-water system to achieve constraint coupling.
7. The scheduling method for an isolated energy system based on deep reinforcement learning as described in claim 4, characterized in that, When performing constraint filtering on the control action vector, the near-end policy optimization algorithm uses a shearing loss function for constraint protection.
8. The scheduling method for an isolated energy system based on deep reinforcement learning as described in claim 3, characterized in that, The calculation of multidimensional real-time rewards defines a reward function that includes a power generation cost term, an environmental protection cost term, an economic benefit term, and a penalty term, to generate reinforcement learning reward signals, including: The reward function for reinforcement learning is defined as follows: ; in, For power generation costs, As an environmental cost item, For economic benefits, D For the penalty function term; Weighted generation of reinforcement learning reward signals.
9. The scheduling method for an isolated energy system based on deep reinforcement learning as described in claim 8, characterized in that, In the reward function, the power generation cost item Represented as: ; in, For maintenance costs, For fuel costs; The operation and maintenance costs Represented as: ; in, For the first Type of equipment Constant effort For the first The unit operation and maintenance cost of this type of equipment; The environmental cost item Represented as: ; in, For the types of polluting gases, To process the first The unit cost of the pollutants, for This device generates The emission coefficient of the polluting gas; The economic benefits item Represented as: ; in, Price per unit of oxygen; This refers to the unit price of ammonia. penalty function term D Represented as: ; in, Penalty for unit difference in electricity consumption; , They are respectively The imbalance of total energy at any given moment and the over-discharge or over-charge of energy storage devices.
10. The scheduling method for an isolated energy system based on deep reinforcement learning as described in claim 8, characterized in that, The calculation of multidimensional real-time rewards, defining a reward function that includes a power generation cost term, an environmental protection cost term, an economic benefit term, and a penalty term, and generating reinforcement learning reward signals, also includes: The reward function is constrained and driven by hard constraint penalties.
Citation Information
Cited By
Stable operation method of seawater hydrogen production system under variable load working condition
CN122147449A