Real-time energy management method and system for mobile energy network based on deep reinforcement learning

By designing the state space, action space, and reward function using deep reinforcement learning methods and training the Q-network model, the problem of real-time dynamic changes in energy management of mobile energy networks was solved, realizing intelligent energy management of all-electric ships and improving energy regulation efficiency and decision-making level.

CN116523228BActive Publication Date: 2026-08-04SHANGHAI JIAOTONG UNIV
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
SHANGHAI JIAOTONG UNIV
Filing Date
2023-04-24
Publication Date
2026-08-04

AI Technical Summary

Technical Problem

Existing energy management methods for mobile energy networks have failed to effectively address real-time dynamic changes during navigation, especially the dynamic changes in ship speed and load, making it difficult to accurately predict and optimize energy regulation.

Method used

By employing a deep reinforcement learning-based approach, a state space, action space, and reward function are designed. A Q-network model is trained using the DQN algorithm to achieve intelligent decision-making for real-time energy management of ships and optimize the power allocation of diesel generator sets and energy storage systems.

Benefits of technology

It improves the adaptability of mobile energy networks to dynamic changes during navigation, enables online decision-making for all-electric ships, enhances energy management and intelligent decision-making capabilities, and reduces fuel consumption and carbon emissions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116523228B_ABST
    Figure CN116523228B_ABST
Patent Text Reader

Abstract

The application provides a mobile energy network real-time energy management method and system based on deep reinforcement learning, comprising the following steps: S1, representing the real-time energy regulation process of a full-electric ship based on a Markov decision process, including a state space, an action space and a reward function; S2, constructing a Q network model representing an action value function, and training the Q network model by using the state space, the action space and the reward function through a DQN algorithm; and S3, selecting a decision action based on the current state space through the trained Q network model, so as to realize intelligent decision of real-time energy management of the ship. The Q network model is fitted to the behavior process of the ship expected to make an optimal energy management intelligent decision through input and output of a neural network, realizes mapping from the state space to the action space, and achieves the purpose of optimal energy management according to the real-time state of the ship operation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of electrical engineering and computer science, and more specifically, to a real-time energy management method for mobile energy networks based on deep reinforcement learning. Background Technology

[0002] With increasingly stringent emission reduction policies, mobile energy networks, represented by electric vehicles, electrified ships, and mobile energy storage vehicles, have become an irreversible trend in transportation electrification. Thanks to the continuous development of electric propulsion technology and integrated power systems, the penetration rate of electrified ships and electric vehicles is gradually increasing. In traditional transportation operation modes, human control plays a crucial role; however, as the complexity of mobile energy networks continues to rise, intelligentization has become an inevitable trend in their development.

[0003] Currently, energy management in mobile energy networks largely relies on accurate forecasts of energy and load, focusing on establishing mathematical optimization models for the entire journey, but failing to consider real-time dynamic changes during flight. However, during real-time flight, due to the complexity and uncertainty of the environment, both the mobile energy network's own energy system and the surrounding environment are in a dynamic process of change, making accurate forecasting difficult to achieve in real-world scenarios. Therefore, real-time energy management systems for mobile energy networks need to enhance their adaptability to load changes and the flexibility of energy regulation.

[0004] Patent document CN114498753A (application number: 202210160754.8) discloses a data-driven real-time energy management method for low-carbon ship microgrids. First, a ship net load scenario set considering the temporal correlation of prediction errors is established through prediction error fitting, equal-probability inverse transformation scenario set generation, and synchronous back-substitution scenario set reduction. Second, combining scenario set information and rolling optimization and feedback correction mechanisms, a stochastic model predictive control energy management model is established that minimizes the expected sum of control action operating costs and state-of-charge deviation penalty costs under each scenario. Subsequently, a large number of training data samples are generated based on stochastic model predictive control, and a random forest algorithm is trained to perform multivariate regression on the data samples. Data-driven stochastic model predictive control real-time energy management strategies for low, medium, and high power load levels are obtained. This patent proposes a data-driven real-time energy management method for ship microgrids. This method focuses on accurate prediction of ship load and obtains control variables through mathematical optimization models; however, accurate prediction is often difficult to achieve. Furthermore, this patent does not consider the impact of dynamic changes in ship speed on real-time energy regulation. The proposed solution does not require accurate load prediction in advance. The trained ship agent can optimize the power allocation of the diesel generator set and energy storage in real time based on the dynamic changes in ship speed and load.

[0005] Y. Hu, W. Li, K. Xu, T. Zahid, F. Qin, and C. Li, “Energy Management Strategy for a Hybrid Electric Vehicle Based on Deep Reinforcement Learning,” Applied Sciences, vol. 8, no. 2, p. 187, Jan. 2018. This paper studies the real-time energy management strategy for hybrid electric vehicles using deep reinforcement learning. This method can autonomously learn the optimal strategy based on data input; however, the design of the state space, action space, and reward function in this paper is not applicable to all-electric ships. This invention designs a corresponding state space, action space, and reward function based on the characteristics of the energy system of all-electric ships, which can effectively solve the real-time energy management problem of all-electric ships.

[0006] Kumar S. Deep Reinforcement learning based energy management in marine hybrid vehicle [D]. NTNU, 2021. This paper studies the real-time energy management strategy of hybrid-powered ships based on deep reinforcement learning. However, this paper does not comprehensively consider the dynamic changes in the ship's navigation process, only considering the uncertainty of the load, and does not consider the impact of ship speed changes on the intelligent decision-making system for real-time energy management. This invention takes into account state variables such as ship speed and acceleration, which can better identify the dynamic changes in ship navigation and further improve the intelligent decision-making level of ship navigation.

[0007] To achieve real-time energy optimization and control of mobile energy networks, improve the intelligent decision-making level of mobile energy networks during navigation, and reduce fuel consumption, this invention proposes a real-time energy management method for mobile energy networks based on deep reinforcement learning, which can significantly improve the operating efficiency of mobile energy networks. Summary of the Invention

[0008] To address the shortcomings of existing technologies, the purpose of this invention is to provide a real-time energy management method and system for mobile energy networks based on deep reinforcement learning.

[0009] A real-time energy management method for mobile energy networks based on deep reinforcement learning, provided by the present invention, includes:

[0010] Step S1: Characterize the real-time energy regulation process of the all-electric ship based on the Markov decision process, including: state space, action space and reward function;

[0011] Step S2: Construct a Q-network model representing the action value function, and train the Q-network model using the DQN algorithm with the state space, action space, and reward function;

[0012] Step S3: Based on the current state space, select decision actions through the trained Q-network model to realize real-time intelligent energy management decision-making for the ship;

[0013] The Q-network model describes the process by which a ship makes optimal energy management decisions by fitting the input and output of a neural network to the ship's desired behavior. This process achieves a mapping from the state space to the action space, thus enabling optimal energy management based on the real-time operating status of the ship.

[0014] Preferably, the state space adopts:

[0015]

[0016] in, This represents the ship's speed during time interval t; This represents the ship's acceleration during time interval t; State of Charge (SOC) represents the power demand of residential services during time period t. t This represents the state of charge of the energy storage system during time period t;

[0017] The action space adopts:

[0018] A t ={ratio t}

[0019] in, P represents the output power of the diesel generator set during time period t; GN Indicates the rated power of the diesel generator set; ratio t Discretize the values ​​within the range of 0 to 1; when ratio t A ratio equal to 1 indicates that the diesel generator set is operating at maximum power; t When the value is 0, it means that the diesel generator set is running under no-load, and the energy storage system provides full load support.

[0020] The reward function adopts:

[0021] The intelligent decision-making system makes a decision action A t The state of charge of the energy storage system is determined by SOC. t Transform into SOC t+1 If SOC t+1 If the state of charge exceeds the specified upper and lower limits, then SOC is obtained. t+1 and SOC tIf the trend of change is opposite to the expectation, then a penalty is imposed through the reward function;

[0022] When SOC t+1 <0 or SOC t+1 >1 hour:

[0023] r t =-C

[0024] Where, r t Indicates that the intelligent decision-making system makes a decision action A. t The reward value obtained later; C represents a positive number;

[0025] When 0≤SOC t+1 <SOC region_L hour:

[0026]

[0027] Among them, SOC region_L This represents the lower limit of the safe state of charge range for an energy storage system; |ΔSOC max | represents the absolute value of the maximum change in the state of charge of an energy storage system over a given time period;

[0028] When SOC region_H ≤SOC t+1 <1 hour:

[0029]

[0030] Among them, SOC region_H This indicates the upper limit of the safe range of the state of charge of the energy storage system;

[0031] When SOC region_L ≤SOC t+1 <SOC region_H When the state of charge is within the safe range, the reward function is designed based on the optimal operating point of the diesel generator set's fuel efficiency.

[0032]

[0033] in, β and are fitting parameters, such that r at this time t The value of varies approximately within the range of [-1, 1]. opt This is the optimal operating point for the fuel efficiency of the diesel generator set.

[0034] Preferably, the Q-network model adopts:

[0035]

[0036] Where t represents time, Gt S represents the return during time period t. t A represents the state during time period t. t R represents the decision-making action made by the intelligent decision-making system during time period t. t+k E represents the reward for the t+k time period. π This represents the expectation under strategy π, and γ represents the discount factor.

[0037] Preferably, step S2 employs:

[0038] Step S2.1: Initialize the current Q-network model Q ω (s,a), and initialize the target network with the same parameters.

[0039] Step S2.2: Initialize the experience replay pool R;

[0040] Step S2.3: Obtain the initial state s1 based on the Markov decision sequence;

[0041] Step S2.4: Based on the current network Q ω (s,a) selects the current state s using an ε-greedy strategy. t The following action a t Execute action a t Receive reward r t The environmental state changes as s t+1 ; will (s t ,a t ,r t ,s t+1 Store the data in the experience replay pool; repeat step S2.4. When the data in the experience replay pool meets the preset requirements, sample N data points {(s)}. i ,a i ,r i ,s′ i )} i=1,…,N For each data point, the loss function is calculated using the target network, and the loss is minimized using the stochastic gradient descent algorithm to update the current network Q. ω The parameters of (s,a) are used; the parameters of the current network are synchronized to the target network at regular intervals; step S2.4 is triggered repeatedly until the current Markov decision sequence is in a terminated state; a new Markov decision sequence is obtained, and steps S2.3 to S2.4 are triggered repeatedly until training is completed.

[0042] Preferably, step S2.4 employs the following:

[0043] Utilizing target network computation

[0044] Where γ represents the discount factor;

[0045] Calculate the loss function:

[0046]

[0047] A real-time energy management system for mobile energy networks based on deep reinforcement learning, according to the present invention, includes:

[0048] Module M1: Based on Markov decision process, it represents the real-time energy regulation process of all-electric ships, including: state space, action space and reward function;

[0049] Module M2: Construct a Q-network model representing the action value function, and train the Q-network model using the DQN algorithm with the state space, action space, and reward function;

[0050] Module M3: Based on the current state space, the trained Q-network model selects decision actions to realize real-time intelligent energy management decision-making for ships;

[0051] The Q-network model describes the process by which a ship makes optimal energy management decisions by fitting the input and output of a neural network to the ship's desired behavior. This process achieves a mapping from the state space to the action space, thus enabling optimal energy management based on the real-time operating status of the ship.

[0052] Preferably, the state space adopts:

[0053]

[0054] in, This represents the ship's speed during time interval t; This represents the ship's acceleration during time interval t; State of Charge (SOC) represents the power demand of residential services during time period t. t This represents the state of charge of the energy storage system during time period t;

[0055] The action space adopts:

[0056] A t ={ratio t}

[0057] in, P represents the output power of the diesel generator set during time period t; GN Indicates the rated power of the diesel generator set; ratio t Discretize the values ​​within the range of 0 to 1; when ratio t A ratio equal to 1 indicates that the diesel generator set is operating at maximum power; t When the value is 0, it means that the diesel generator set is running under no-load, and the energy storage system provides full load support.

[0058] The reward function adopts:

[0059] The intelligent decision-making system makes a decision action A t The state of charge of the energy storage system is determined by SOC. t Transform into SOC t+1 If SOC t+1 If the state of charge exceeds the specified upper and lower limits, then SOC is obtained. t+1 and SOC t If the trend of change is opposite to the expectation, then a penalty is imposed through the reward function;

[0060] When SOC t+1 <0 or SOC t+1 >1 hour:

[0061] r t =-C

[0062] Where, r t Indicates that the intelligent decision-making system makes a decision action A. t The reward value obtained later; C represents a positive number;

[0063] When 0≤SOC t+1 <SOC region_L hour:

[0064]

[0065] Among them, SOC region_L This represents the lower limit of the safe state of charge range for an energy storage system; |ΔSOC max | represents the absolute value of the maximum change in the state of charge of an energy storage system over a given time period;

[0066] When SOC region_H ≤SOC t+1 <1 hour:

[0067]

[0068] Among them, SOC region_H This indicates the upper limit of the safe range of the state of charge of the energy storage system;

[0069] When SOC region_L ≤SOC t+1 <SOC region_H When the state of charge is within the safe range, the reward function is designed based on the optimal operating point of the diesel generator set's fuel efficiency.

[0070]

[0071] in, β and are fitting parameters, such that r at this time t The value of varies approximately within the range of [-1, 1]. opt This is the optimal operating point for the fuel efficiency of the diesel generator set.

[0072] Preferably, the Q-network model adopts:

[0073]

[0074] Where t represents time, G t S represents the return during time period t. t A represents the state during time period t. t R represents the decision-making action made by the intelligent decision-making system during time period t. t+k E represents the reward for the t+k time period. π This represents the expectation under strategy π, and γ represents the discount factor.

[0075] Preferably, the module M2 adopts:

[0076] Module M2.1: Initialize the current Q-network model Q ω (s,a), and initialize the target network with the same parameters.

[0077] Module M2.2: Initialize the experience replay pool R;

[0078] Module M2.3: Obtain the initial state s1 based on the Markov decision sequence;

[0079] Module M2.4: Based on the current network Q ω (s,a) selects the current state s using an ε-greedy strategy. t The following action a t Execute action a t Receive reward r t The environmental state changes as s t+1 ; will (s t ,a t ,r t ,s t+1 The data is stored in the experience replay pool; the repeated trigger module M2.4 samples N data points when the data in the experience replay pool meets the preset requirements. i ,a i ,r i ,s′ i )} i=1,...,N For each data point, the loss function is calculated using the target network, and the loss is minimized using the stochastic gradient descent algorithm to update the current network Q. ωThe parameters of (s,a) are used; the parameters of the current network are synchronized to the target network at regular intervals; module M2.4 is triggered repeatedly until the current Markov decision sequence is in a terminated state; a new Markov decision sequence is obtained, and modules M2.3 to M2.4 are triggered repeatedly until training is completed.

[0080] Preferably, module M2.4 adopts:

[0081] Utilizing target network computation

[0082] Where γ represents the discount factor;

[0083] Calculate the loss function:

[0084]

[0085] Compared with the prior art, the present invention has the following beneficial effects:

[0086] 1. The real-time energy management method proposed in this invention can improve the adaptability of mobile energy networks to dynamic changes during navigation, and enable all-electric ships to make online decisions during navigation;

[0087] 2. This invention uses ship speed and living service load as state variables, which can take into account the dynamic changes in ship speed and the uncertainty of living service load to realize real-time energy management of all-electric ships, thereby achieving a better level of energy management.

[0088] 3. This invention proposes to deploy a real-time energy management intelligent decision-making system for mobile energy networks using deep reinforcement learning methods. This eliminates the need for precise mathematical modeling of the ship's energy system, has wide applicability and good scalability, and can effectively improve the intelligent decision-making level of all-electric ships.

[0089] 4. This invention is not limited to real-time energy management of all-electric ships, but has a certain degree of universality for real-time energy management of mobile energy networks;

[0090] 5. This invention does not require accurate prediction of the ship's own energy system and marine environment during real-time energy regulation of ships. It can make optimal action decisions based on real-time status, which has significant advantages over traditional mathematical optimization methods. Attached Figure Description

[0091] Other features, objects, and advantages of the present invention will become more apparent from the following detailed description of non-limiting embodiments with reference to the accompanying drawings:

[0092] Figure 1 This is a schematic diagram of the real-time energy regulation process for an all-electric ship.

[0093] Figure 2 This is a flowchart of a real-time energy management method for all-electric ships based on deep reinforcement learning.

[0094] Figure 3 This is a schematic diagram of the training process.

[0095] Figure 4 This is a schematic diagram showing the state of charge change curve and charging / discharging power of an energy storage system.

[0096] Figure 5 This is a schematic diagram showing the power distribution of the diesel generator set and energy storage system.

[0097] Figure 6 The flowchart shows the process of training a Q-network model using the DQN algorithm. Detailed Implementation

[0098] The present invention will now be described in detail with reference to specific embodiments. These embodiments will help those skilled in the art to further understand the present invention, but do not limit the invention in any way. It should be noted that those skilled in the art can make several changes and improvements without departing from the concept of the present invention. These all fall within the protection scope of the present invention.

[0099] This invention aims to address the shortcomings of existing real-time energy management methods for mobile power networks. To this end, this invention provides a real-time energy management method and system for mobile power networks based on deep reinforcement learning, aiming to achieve a superior level of real-time energy management and effectively cope with dynamic changes during real-time navigation in mobile power networks.

[0100] This invention relates to real-time energy management of mobile energy networks based on deep reinforcement learning. It utilizes deep reinforcement learning to deploy a real-time energy management intelligent decision-making system for all-electric ships. Focusing on all-electric ships, the invention represents the real-time energy regulation process of all-electric ships as a Markov decision process and designs corresponding state space, action space, and reward function. The ship's intelligent agent is trained based on the DQN algorithm to obtain the optimal real-time energy management strategy, thereby realizing intelligent navigation of all-electric ships.

[0101] Example 1

[0102] A real-time energy management method for mobile energy networks based on deep reinforcement learning, provided by the present invention, includes:

[0103] Step S1: Characterize the real-time energy regulation process of the all-electric ship based on the Markov decision process, including: state space, action space and reward function;

[0104] Step S2: Construct a Q-network model representing the action value function, and train the Q-network model using the DQN algorithm with the state space, action space, and reward function;

[0105] Step S3: Based on the current state space, select decision actions through the trained Q-network model to realize real-time intelligent energy management decision-making for the ship;

[0106] The Q-network model describes the process by which a ship makes optimal energy management decisions by fitting the input and output of a neural network to the ship's desired behavior. This process achieves a mapping from the state space to the action space, thus enabling optimal energy management based on the real-time operating status of the ship.

[0107] Specifically, the real-time energy regulation process of mobile energy networks is characterized based on Markov decision processes, including:

[0108] The key to real-time energy regulation in mobile energy networks lies in tailoring an intelligent decision-making system for real-time energy management that is both adaptive and cost-effective. This system needs to be able to coordinate and control the power flow of the ship's energy system at each time period based on the ship's operating status, in order to minimize operating costs and improve economic and environmental benefits.

[0109] The real-time energy management problem of all-electric ships can be viewed as a sequential decision-making problem within a finite time domain, which can be characterized by a Markov decision process. Its fundamental property is that the current state of the system depends only on the state and actions of the previous moment. In this invention, the ship's real-time energy management intelligent decision-making system can be considered as an agent within a Markov decision process, interacting with the entire ship's energy system and the surrounding marine environment. At each time interval, the ship's real-time energy management intelligent decision-making system determines the action decision for that time interval based on the environmental feedback and the reward signal from the action decision of the previous time interval. Its basic interaction process is as follows: Figure 1 As shown.

[0110] In Markov decision-making, the agent is responsible for performing actions and interacting with its environment. The process of environmental state changes is influenced by both the random process of spontaneous changes in the environment and the actions and behaviors of the agent.

[0111] A Markov decision process can be represented by a quintuple (S, A, T, r, γ), where S represents the state space, A represents the action space, T represents the state transition function, T(s′|s,a) represents the probability of reaching state s′ after performing action a in state s, r is the reward function, which depends on state s and action a, and γ is a discount factor with a value in the range [0,1). When γ is close to 1, the agent pays more attention to long-term cumulative rewards, and when it is close to 0, it considers short-term rewards more.

[0112] Given a state, the probability distribution of an agent's actions in the action space is called the policy, usually denoted by π. Suppose the agent takes action a in state s, and subsequent actions still follow the policy π. The expected reward obtained in this case is defined as the action value function, denoted by Q. π (s,a) represents, mathematically, as:

[0113]

[0114] Where t represents time, G t S represents the return during time period t. t A represents the state during time period t. t R represents the decision-making action made by the intelligent decision-making system during time period t. t+k This represents the reward for the t+k time period.

[0115] The action-value function under the optimal policy is called the optimal action-value function, denoted as Q. * (s,a), its mathematical expression is:

[0116]

[0117] Deploying real-time energy management strategies for all-electric ships using deep reinforcement learning does not rely on precise mathematical modeling of the ship's energy system. The key lies in clearly defining the objectives of the real-time energy management intelligent decision-making system and the state feedback that can be obtained from the ship's energy system and environment. Therefore, it is necessary to design the state space, action space, and reward function for real-time energy regulation of the ship based on Markov decision processes.

[0118] The state space adopts:

[0119] The selection of state variables needs to characterize the state of the all-electric ship during navigation. During navigation, ship speed changes dynamically, and this speed change further determines the propulsion power demand. The ship's real-time energy management intelligent decision-making system needs to track this dynamic change. Therefore, ship speed and acceleration are selected as state variables. Furthermore, since the demand for shipboard living services is inherently random and highly uncertain, making accurate prediction difficult, the ship's real-time energy management intelligent decision-making system can only accurately determine the living service load demand for the current period. Therefore, the power demand for living services is selected as a state variable. In this invention, the energy source for the all-electric ship is considered to consist of a diesel generator set and an energy storage system, where the energy storage system can both generate power and absorb power from the diesel generator set. To maintain the state of charge of the energy storage system within a safe range, the state of charge is also selected as a state variable. In summary, the state space in this invention is defined as follows:

[0120]

[0121] in, This represents the ship's speed during time period t. This represents the ship's acceleration during time interval t. State of Charge (SOC) represents the power demand of residential services during time period t. t This represents the state of charge of the energy storage system during time period t. The first three state variables... and For ship real-time energy management intelligent decision-making systems, the state variable (SOC) is an uncertain exogenous variable that cannot be controlled. t These are endogenous variables influenced by the decisions made by intelligent decision-making systems.

[0122] The action space adopts:

[0123] The action space represents the actions that an intelligent decision-making system can control. To determine the power allocation among energy sources in an all-electric ship, this invention proposes using a diesel generator set load ratio factor, denoted as ratio, as the decision variable. t Its definition is:

[0124]

[0125] in, P represents the output power of the diesel generator set during time period t. GN This indicates the rated power of the diesel generator set. When ratio t A ratio of 1 indicates that the diesel generator set is operating at maximum power; t A value of 0 indicates that the diesel generator set is operating under no-load conditions, with the energy storage system providing full load support. Therefore, the operational space of this invention is defined as follows:

[0126] A t ={ratio t} (5)

[0127] Where, ratio t Discretize the values ​​within the range of 0 to 1, and set them as {0, 0.2, 0.4, 0.6, 0.8, 1.0}.

[0128] The reward function adopts:

[0129] The reward function determines the performance metrics the algorithm focuses on and serves as a benchmark guiding the agent to achieve optimal decisions. To improve fuel efficiency, reduce operating costs, and decrease carbon emissions, the ship's real-time energy management intelligent decision-making system needs to operate the diesel generator set as close to its economic operating point as possible and maintain the state of charge (SOC) of the energy storage system within a safe range. This invention sets the reward function in segments based on the SOC of the energy storage system. The intelligent decision-making system then makes a decision action A. tThe state of charge of the energy storage system is determined by SOC. t Transform into SOC t+1 If SOC t+1 If the state of charge exceeds the specified upper and lower limits, then the state of charge (SOC) must be considered. t+1 and SOC t If the trend of change is opposite to the expectation, then a penalty needs to be imposed through the reward function, as shown below.

[0130] 1) SOC t+1 <0 or SOC t+1 >1 o'clock:

[0131] r t =-C (6)

[0132] Where C represents a positive number, and in this invention, C is 3, which yields the best results; r t Indicates that the intelligent decision-making system makes a decision action A. t The reward value obtained afterward. The formula represents SOC. t+1 When the value of has no actual physical meaning, the reward value is -3.

[0133] 2) 0 ≤ SOC t+1 <SOC region_L hour:

[0134]

[0135] Among them, SOC region_L This represents the lower limit of the safe state of charge range for an energy storage system. |ΔSOC max | represents the absolute value of the maximum change in the state of charge of an energy storage system over a given time period.

[0136] 3) SOC region_H ≤SOC t+1 <1 o'clock:

[0137]

[0138] Among them, SOC region_H This indicates the upper limit of the safe range of the state of charge of the energy storage system.

[0139] 4) SOC region_L ≤SOC t+1 <SOC region_H That is, when the state of charge is within the safe range, the reward function is designed based on the optimal operating point of the diesel generator set's fuel efficiency:

[0140]

[0141] in, β and are fitting parameters, such that r at this timet The value of varies approximately within the range of [-1, 1]. opt This is the optimal operating point for the fuel efficiency of the diesel generator set.

[0142] A real-time intelligent decision-making system for ship energy management trained based on the DQN algorithm.

[0143] The core of deploying a real-time energy management intelligent decision-making system for ships lies in continuously improving and optimizing real-time energy management strategies. If the optimal action value function can be solved, the optimal strategy can be easily determined. However, directly solving for the optimal action value function is usually not feasible in real-world situations. Deep reinforcement learning methods can approximate the optimal action value function through deep neural networks. This invention uses the Deep Q Network (DQN) algorithm to train the ship's intelligent agent, specifically training a real-time energy management intelligent decision-making system for all-electric ships.

[0144] The DQN algorithm introduces two Q-networks to fit the optimal action value function, namely the current network Q... ω (s,a) and target network Where ω and This section describes the parameters of the current network and the target network, respectively. The target network can be considered a copy of the current network, and its network parameters are synchronized with the current network periodically. For example... Figure 6 As shown, the basic flow of the DQN algorithm is illustrated in steps a to o:

[0145] Step a: Initialize the current network Q ω (s,a) and initialize the target network using the same parameters.

[0146] Step b: Initialize the experience replay pool.

[0147] Step c: Begin a new Markov decision, executing steps d through n:

[0148] Step d: Obtain the initial state s1 of the Markov decision sequence.

[0149] Step e: For each time step t = 1 → T, execute steps f to m:

[0150] Step f: Based on the current network Q ω (s,a) selects the current state s using an ε-greedy strategy. t The following action a t .

[0151] Step g: Perform action a t Receive reward r t The environmental state becomes s t+1 .

[0152] Step h: (s) t ,a t ,r t ,s t+1 Stored in the experience replay pool.

[0153] Step i: If there is enough data in the experience replay pool, then sample N data points {(s)} from it. i ,a i ,r i ,s′ i )} i=1,…,N .

[0154] Step j: For each data point, compute using the target network.

[0155] Step k: Calculate the loss

[0156] Step 1: Minimize the loss using the stochastic gradient descent algorithm and update the current network Q. ω The parameters of (s,a).

[0157] Step m: Synchronize the parameters of the current network to the target network at regular intervals.

[0158] Step n: Return to step e until the termination state of the Markov decision sequence is reached.

[0159] Step o: Return to step c until all Markov decision sequences have been trained.

[0160] After training the ship agent using the DQN algorithm, the Q network can approximate the optimal action value function. In any given state, the agent can identify the action with the highest action value based on the Q network.

[0161] A Real-Time Energy Management Method for Mobile Energy Networks Based on Deep Reinforcement Learning

[0162] First, the real-time energy regulation process of an all-electric ship is formulated as a Markov decision process considering environmental uncertainties, designing the state space, action space, and reward function for real-time energy regulation. Next, the DQN algorithm, combining reinforcement learning and neural networks, is used to train the ship's real-time energy management intelligent decision system, learning from historical data to capture the uncertainty characteristics of ship speed and load consumption. Finally, based on the trained real-time energy management intelligent decision system, real-time energy management of the all-electric ship is implemented. The basic process is as follows: Figure 2 As shown.

[0163] The real-time energy management system for mobile energy networks based on deep reinforcement learning provided by this invention can be implemented through the steps and flow of the real-time energy management method for mobile energy networks based on deep reinforcement learning provided by this invention. Those skilled in the art can understand the aforementioned real-time energy management method for mobile energy networks based on deep reinforcement learning as a preferred example of the real-time energy management system for mobile energy networks based on deep reinforcement learning.

[0164] This invention is not limited to traditional mathematical optimization methods. It proposes a real-time energy management method for mobile energy networks that takes into account dynamic changes in ship speed and uncertainties in the load of living services. This method helps mobile energy networks to further improve navigation economy and reduce carbon emissions.

[0165] Meanwhile, this invention differs from existing mobile energy network energy management methods by characterizing the real-time energy regulation process of the mobile energy network as a Markov decision process, enriching and refining the state space, action space, and reward function. The design of the reward function takes into account the economic operating point of the diesel generator set and the dynamic change trend of the state of charge of the energy storage system.

[0166] This invention presents an algorithmic framework for deploying a real-time energy management intelligent decision-making system based on deep reinforcement learning. It employs a neural network to fit the optimal action value function to obtain the optimal real-time energy management strategy, which can effectively improve the intelligent decision-making level of mobile energy networks during navigation.

[0167] Example 2

[0168] The present invention will now be described in further detail with reference to the accompanying drawings and specific embodiments.

[0169] This invention uses actual navigation data from a fully electric ship as an example, constructs a DQN algorithm according to the parameters in Table 1, and trains the ship's intelligent agent using 500 Markov decision sequences. The training process is as follows: Figure 3 As shown. The trained intelligent decision-making system for real-time ship energy management was tested in a test scenario, and the test results are as follows. Figure 4 and Figure 5 As shown.

[0170] Table 1. Q-Network and Hyperparameter Design

[0171]

[0172]

[0173] It can be seen that during the online decision-making process, although the total load demand fluctuates between 159.2kW and 715.9kW, with a peak-to-valley difference of approximately 557kW, the well-trained ship real-time energy management intelligent decision-making system can consistently maintain the energy storage system's state of charge within a safe range of 0.3 to 0.8. Furthermore, during peak load periods, the energy storage system has sufficient power to support the load, and during off-peak periods, it can promptly replenish energy. In addition, the diesel generator sets mostly operate at 400kW and 500kW under load, rarely operating in the low-efficiency range.

[0174] This embodiment illustrates that the intelligent decision-making system for real-time energy management of ships trained based on the DQN algorithm can realize online decision-making during ship navigation and can cope with the uncertainty of dynamic load changes. It can rationally allocate the power flow of diesel generator sets and energy storage according to the current load conditions and the charge status of the energy storage system.

[0175] Those skilled in the art will understand that, in addition to implementing the system, apparatus, and their modules provided by this invention in purely computer-readable program code, the same program can be implemented in the form of logic gates, switches, application-specific integrated circuits, programmable logic controllers, and embedded microcontrollers by logically programming the method steps. Therefore, the system, apparatus, and their modules provided by this invention can be considered a hardware component, and the modules included therein for implementing various programs can also be considered structures within the hardware component; alternatively, modules for implementing various functions can be considered both software programs implementing the method and structures within the hardware component.

[0176] Specific embodiments of the present invention have been described above. It should be understood that the present invention is not limited to the specific embodiments described above, and those skilled in the art can make various changes or modifications within the scope of the claims, which do not affect the essence of the present invention. Unless otherwise specified, the embodiments and features described in this application can be arbitrarily combined with each other.

Claims

1. A method for real-time energy management of mobile energy networks based on deep reinforcement learning, characterized in that, include: Step S1: Characterize the real-time energy regulation process of the all-electric ship based on the Markov decision process, including: state space, action space and reward function; Step S2: Construct a Q-network model representing the action value function, and train the Q-network model using the DQN algorithm with the state space, action space, and reward function; Step S3: Based on the current state space, select decision actions through the trained Q-network model to realize real-time intelligent energy management decision-making for the ship; The Q-network model is a process of making optimal energy management intelligent decisions by fitting the ship's expectations through the input and output of a neural network. It realizes the mapping from the state space to the action space and achieves the goal of optimal energy management based on the real-time state of the ship's operation. The state space adopts: in, express t Ship speed during the time period; express t Ship speed acceleration over a given period of time; express t Power demand of residential service load during a given time period; express t State of charge of the time-phase energy storage system; The action space adopts: in, , express t The output power of the diesel generator set during the specified time period; Indicates the rated power of the diesel generator set; Discretize the values ​​within the range of 0 to 1; when When the value is 1, it means that the diesel generator set is operating at maximum power. When the value is 0, it means that the diesel generator set is running under no-load, and the energy storage system provides full load support. The reward function adopts: Intelligent decision-making system makes decisions The state of charge of the energy storage system is changed from Become ,like If the charge exceeds the specified upper and lower limits, then obtain and If the trend of change is opposite to the expectation, then a penalty is imposed through the reward function; When time: wherein, represents the decision action made by the intelligent decision system the reward value obtained later; C represents a positive number; When Time: wherein, represents a lower limit of a state of charge safety interval of the energy storage system; represents an absolute value of a maximum variation of the state of charge of the energy storage system for a time period. When Time: wherein, represents the upper limit of the energy storage system state of charge safety interval; when When the state of charge is within the safe range, the reward function is designed based on the optimal operating point of the diesel generator set's fuel efficiency. wherein , and are fitting parameters such that the value of varies in the interval [-1, 1] at this time, is the diesel genset fuel efficiency optimum operating point.

2. The deep reinforcement learning based mobile energy network real-time energy management method according to claim 1, characterized in that, The Q-network model adopts: in, Indicates time, express t Time period rewards express t The state of the time period express t The decision-making actions made by the time-based intelligent decision-making system express Rewards for specific time periods; Indicating in strategy Seeking expectations This represents the discount factor. 3.The deep-reinforcement learning based mobile energy network real-time energy management method of claim 1, wherein, Step S2 employs the following: Step S2.1 : initialize the current Q-network model and initialize the target network with the same parameters ; Step S2.2: Initialize the experience replay pool R; Step S2.3: Obtain initial state based on Markov decision sequence ; Step S2.4: Based on the current network by - Greedy strategy selects the current state The following action Execute actions Receive rewards The environmental state changes as ;Will Stored in the experience replay pool; Repeat step S2.4, when the data in the experience replay pool meets the preset requirement, then sample N data ; for each data, calculate the loss function using the target network, and minimize the loss by a stochastic gradient descent algorithm to update the parameters of the current network ; synchronize the parameters of the current network to the target network at certain time intervals; Repeat step S2.4 until the current Markov decision sequence is in a terminated state; obtain a new Markov decision sequence, and repeat steps S2.3 to S2.4 until training is complete.

4. The deep reinforcement learning based mobile energy network real-time energy management method according to claim 3, characterized in that, Step S2.4 adopts the following: Computing with target networks ; wherein represents a discount factor; Calculate the loss function: 。 5. A deep reinforcement learning based mobile energy network real-time energy management system, characterized in that, include: Module M1: Based on Markov decision process, it represents the real-time energy regulation process of all-electric ships, including: state space, action space and reward function; Module M2: Construct a Q-network model representing the action value function, and train the Q-network model using the DQN algorithm with the state space, action space, and reward function; Module M3: Based on the current state space, the trained Q-network model selects decision actions to realize real-time intelligent energy management decision-making for ships; The Q-network model is a process of making optimal energy management intelligent decisions by fitting the ship's expectations through the input and output of a neural network. It realizes the mapping from the state space to the action space and achieves the goal of optimal energy management based on the real-time state of the ship's operation. The state space adopts: in, express t Ship speed during the time period; express t Ship speed acceleration over a given period of time; express t Power demand of residential service load during a given time period; express t State of charge of the time-phase energy storage system; The action space adopts: in, , express t The output power of the diesel generator set during the specified time period; Indicates the rated power of the diesel generator set; Discretize the values ​​within the range of 0 to 1; when When the value is 1, it means that the diesel generator set is operating at maximum power. When the value is 0, it means that the diesel generator set is running under no-load, and the energy storage system provides full load support. The reward function adopts: Intelligent decision-making system makes decisions The state of charge of the energy storage system is changed from Become ,like If the charge exceeds the specified upper and lower limits, then obtain and If the trend of change is opposite to the expectation, then a penalty is imposed through the reward function; When Time: wherein, represents the decision action made by the intelligent decision system the reward value obtained later; C represents a positive number; When Time: wherein, represents a lower limit of a state of charge safety interval of the energy storage system; represents an absolute value of a maximum variation of the state of charge of the energy storage system for a time period. When Time: wherein, represents the upper limit of the energy storage system state of charge safety interval; When the state of charge is in the safe interval, the reward function is designed according to the diesel generator set fuel efficiency optimal operating point: in, , and For the fitting parameters, such that at this time The value of varies within the interval [-1, 1]. This is the optimal operating point for the fuel efficiency of the diesel generator set.

6. The deep reinforcement learning based mobile energy network real-time energy management system of claim 5, wherein, The Q-network model adopts: in, Indicates time, express t Time period rewards, express t The state of the time period express t The decision-making actions made by the time-based intelligent decision-making system express Rewards for specific time periods; Indicating in strategy Seeking expectations This represents the discount factor.

7. The deep reinforcement learning based mobile energy network real-time energy management system of claim 5, wherein, The module M2 adopts: Module M2.1: Initialize the current Q-network model The target network was initialized using the same parameters. ; Module M2.2: Initialize the experience replay pool R; Module M2.3: Obtaining initial state based on Markov decision sequence ; Module M2.4: Based on the current network by - Greedy strategy selects the current state The following action Execute actions Receive rewards The environmental state changes as ;Will Stored in the experience replay pool; The repeated trigger module M2.4 samples N data points when the data in the experience playback pool meets preset requirements. For each data point, the loss function is calculated using the target network, and the loss is minimized using the stochastic gradient descent algorithm to update the current network. The parameters are set; the parameters of the current network are synchronized to the target network at regular intervals; module M2.4 is triggered repeatedly until the current Markov decision sequence is terminated; a new Markov decision sequence is obtained, and modules M2.3 to M2.4 are triggered repeatedly until training is complete.

8. The deep reinforcement learning based mobile energy network real-time energy management system of claim 7, wherein, The module M2.4 adopts: Computing with target networks ; wherein represents a discount factor; Calculate the loss function: 。