Layered coordination control method for urban rail hybrid energy storage system based on deep reinforcement learning
By employing a hierarchical coordinated control method for urban rail hybrid energy storage systems based on deep reinforcement learning, and utilizing the MATD3 algorithm for online optimization of the urban rail hybrid energy storage system, the problem of how to fully utilize the regenerative braking energy of trains was solved, achieving stable energy saving of the traction power supply network and extending the lifespan of energy storage components.
Patent Information
- Application Number
- CN202510960173.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-11
- Publication Date
- 2025-10-31
AI Technical Summary
How to fully utilize the regenerative braking energy of trains, leverage the complementary advantages of HESS characteristics, achieve stable energy saving of the traction power supply network, and extend the service life of energy storage components.
A hierarchical coordinated control method for urban rail hybrid energy storage systems based on deep reinforcement learning is adopted. The multi-agent dual-delay deep deterministic policy gradient algorithm (MATD3) is used to optimize the upper-level power allocation strategy and the lower-level current control online, respectively, to construct a hierarchical coordinated control architecture. The hierarchical optimization and coordinated control of the hybrid energy storage system are realized through an online training-online decision-making method.
It improves the energy-saving and voltage-stabilizing characteristics of DC traction networks, avoids overcharging and over-discharging phenomena in hybrid energy storage systems, extends the service life of energy storage components, and enhances the energy-saving and economic benefits of urban rail transit systems.
Smart Images

Figure CN120879700A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of urban rail transit energy management technology, specifically to a hierarchical coordinated control method for urban rail hybrid energy storage systems based on deep reinforcement learning. Background Technology
[0002] Due to the high operating speed and short distances between stations in urban rail transit, trains need to start and brake frequently, generating a large amount of regenerative braking energy during braking. How to fully utilize this regenerative braking energy is one of the key issues in the field of energy conservation in urban rail transit. Using a hybrid energy storage system (HESS) composed of onboard supercapacitors and ground-based lithium batteries can achieve the recycling of regenerative braking energy. By designing a reasonable energy management method, efficient hierarchical coordinated control can be achieved to fully leverage the complementary advantages of HESS characteristics and realize peak shaving and valley filling of traction network voltage fluctuations.
[0003] In addition, frequent train starts and stops can affect the lifespan of energy storage components. Therefore, how to improve the lifespan of energy storage components while effectively achieving the energy-saving and voltage-stabilizing characteristics of the energy storage system is a technical problem that technicians need to consider and solve. Summary of the Invention
[0004] Technical problem: To fully utilize the regenerative braking energy of trains, leverage the complementary advantages of HESS characteristics, and efficiently achieve hierarchical coordinated control, thereby ensuring the stability and energy saving of the traction power supply network while extending the service life of energy storage components, is a problem that urgently needs to be solved.
[0005] Technical Solution: To address the aforementioned issues, this invention proposes a hierarchical coordinated control method for urban rail hybrid energy storage systems based on deep reinforcement learning. This method constructs a hierarchical coordinated control architecture and utilizes the Multi-agent Twin Delayed Deep Deterministic Policy Gradient (MATD3) algorithm to optimize the upper-layer power allocation strategy and lower-layer current control online. Furthermore, it employs an "online training-online decision-making" method to achieve hierarchical optimization and coordinated control of the hybrid energy storage system.
[0006] The method is implemented in the following steps:
[0007] First, a hierarchical coordinated control architecture for the urban rail hybrid energy storage system is established. This includes an upper-level energy management system and a lower-level converter control system. The upper-level energy management system generates charging and discharging current commands for the ground-based lithium battery and the on-board supercapacitor through voltage control, traction power feedforward control, and power distribution strategies. The lower-level converter control system compares the current commands given by the upper level with the actual current to control the DC / DC converter, thereby controlling the charging and discharging of the energy storage system. The lower-level converter control system also includes a dynamic current compensation module for dynamically compensating the current control portion of the on-board supercapacitor.
[0008] The current dynamic compensation module will calculate the actual power P of the ground-mounted lithium battery. bat With a given low-frequency power P bat_ref After differential calculation, the voltage U across the vehicle-mounted supercapacitor is... sc By combining these methods, the real-time current compensation amount ΔI of the vehicle-mounted supercapacitor is obtained. sc Then, the current I of the vehicle-mounted supercapacitor. sc Given the charging and discharging current combined with I sc_ref0 By combining these methods, the actual charging and discharging current I of the vehicle-mounted supercapacitor can be obtained. sc_ref The current compensation formula is as follows (1):
[0009]
[0010] Wherein, ΔP bat This represents the difference in power absorbed by the ground-based lithium battery.
[0011] Secondly, based on the hierarchical coordinated control architecture, MATD3 is used to optimize the hierarchical coordinated control architecture online. Based on the control application of multi-agent reinforcement learning algorithms in urban rail hybrid energy storage, the training framework, environmental state, agent actions, and reward function of the agents are designed respectively. The specific methods are as follows:
[0012] 1) Design of the training framework for the agent:
[0013] The method employs a "centralized training-distributed learning" training framework for controlling urban rail hybrid energy storage systems. MATD3's agents consist of Agent A and Agent B, which dynamically adjust the power distribution of the hybrid energy storage system and the charging and discharging of the ground-based lithium batteries, respectively. Each agent has a decentralized actuator and two centralized critics. The two agents have both a cooperative relationship (sharing the urban rail system's state information and sharing the common goal of improving the system's energy-saving and voltage-stabilizing characteristics) and a competitive relationship (having different goals, requiring a reward mechanism for differentiation).
[0014] 2) Design of power supply environment and state space for urban rail permanent magnet traction:
[0015] Taking the urban rail permanent magnet traction power supply system as the environment in which the two agents exist, and selecting the DC traction voltage U... dc Hybrid energy storage power given P hess_ref Vehicle-mounted supercapacitors' State of Charge (SOC) sc State of charge (SOC) of ground-mounted lithium batteries bat The voltage U across the vehicle-mounted supercapacitor sc Ground-based lithium battery terminal voltage U bat Train speed ω m and acceleration a c The state space S, representing the state of the environment as observed by the two agent agents, is expressed by equation (2):
[0016]
[0017] Among them, U dcn P hess_refn P hess_refn SOC scn SOC batn U scn U batn a cn , representing the DC traction grid voltage at the nth sampling time, the total power demand of the hybrid energy storage system, the state of charge of the on-board supercapacitor, the state of charge of the ground-mounted lithium battery, the terminal voltage of the on-board supercapacitor, the terminal voltage of the ground-mounted lithium battery, the train speed, and the train acceleration, respectively.
[0018] 3) Selection of continuous motion space and execution of motion:
[0019] After observing the state information of the environment, the agent selects an action from the action space based on its own policy π, i.e., the mapping relationship between state and action. Therefore, the action selected by Agent A is the power adjustment amount ΔP of the on-board supercapacitor in the upper-level control. sc The continuous action space of Action A is A; the selected action of Agent B, Action B, is the adjustment amount ΔI of the charging and discharging current of the ground-based lithium battery controlled by the lower layer. bat The continuous action space of Action B is B, as shown in equation (3) below:
[0020]
[0021] Where, ΔP scn ΔI batnThese represent the power adjustment of the vehicle-mounted supercapacitor and the charging / discharging current adjustment of the ground-based lithium battery at the nth sampling time, respectively.
[0022] 4) Reward function design:
[0023] The reward function determines that Agent A and Agent B have both cooperative and competitive relationships. The common rewards for the two agents are the voltage regulation rate v% and the energy saving rate e%, but the two agents also have different rewards.
[0024] Agent A's reward r1 is selected as the voltage regulation rate v% and energy saving rate e% within a time step ΔT, as well as the SOC of the on-board supercapacitor. sc The weighted sum of safety variation ranges aims to achieve optimal voltage regulation and energy saving while minimizing the SOC of the vehicle-mounted supercapacitor. sc Able to remain within a safe range [SOC] scmin SOC scmax Within ], as shown in equation (4):
[0025]
[0026] The Agent B reward r2 is selected as the voltage regulation efficiency v% and energy saving efficiency e% within the time step ΔT, as well as the SOC of the ground-mounted lithium battery. bat The weighted sum of the monitoring optimization index d%, and the monitoring optimization index d% can be kept within [d]. min ,d max Within ], as shown in the following formula (5):
[0027] r2=max[λ2·v%+μ2·e%+σ2(d min ≤d%≤d max )-η2(d% <d min or d%>d max (5)
[0028] In equations (4)-(5), λ1,μ1,σ1,η1,λ2,μ2,σ2,η2 are weighting coefficients; SOC scmin SOC scmax These are vehicle-mounted supercapacitor SOCs sc The minimum and maximum safety ranges; d min d max The minimum and maximum monitoring optimization index ranges are respectively defined; the energy saving rate e% is defined as the percentage change in the total output energy of the substation before and after the installation of the hybrid energy storage system relative to the total output energy of the substation without the energy storage system, as shown in the following formula (6):
[0029]
[0030] in, These represent the DC traction grid voltage with and without hybrid energy storage installed, respectively. These represent the DC traction network currents with and without mixed energy storage, respectively, and T represents the train's running time from start-up to braking.
[0031] The voltage regulation rate v% is evaluated by integrating the portion of the DC traction voltage that exceeds / below the limit, as shown in equation (7) below:
[0032]
[0033] in, These represent the set upper and lower safety limits for the DC traction network voltage, respectively, and Δh / Δl represent the time during which the DC traction voltage exceeds the upper and lower safety limits during train operation.
[0034] Ground-based lithium battery SOC bat The monitoring optimization index d% is defined as the current SOC of the ground-based lithium battery compared to the initial SOC of the ground-based lithium battery without the MATD3 algorithm, relative to the maximum capacity SOC of the ground-based lithium battery. batmax The percentage deviation is shown in equation (8) below:
[0035]
[0036] Among them, SOC bat The current state of charge (SOC) of the ground-based lithium battery. bat0 The initial SOC of a ground-based lithium battery without using the MATD3 algorithm.
[0037] Finally, an online training-online decision-making method was designed:
[0038] 1) In the online training module, the urban rail permanent magnet traction power supply system platform is used as the intelligent agent environment, enabling each intelligent agent to interact with the platform; under the condition of random initialization of train running speed, each intelligent agent shares training data and centrally updates the strategy until a stable control strategy is obtained.
[0039] 2) In the online decision-making module, based on real-time operating data, the intelligent agent quickly makes the optimal control decision according to the current state, so as to achieve energy saving and voltage stabilization and improve the service life of the energy storage system.
[0040] The system is running, and the specific steps are as follows:
[0041] Step 1: Two agents learn online from the urban rail permanent magnet traction power supply system environment. By collecting the system's status, actions and reward information, they interact with the environment repeatedly N times until the accumulated reward stabilizes and converges.
[0042] Step 2: Deploy the trained Agent strategy to the actual urban rail permanent magnet traction power supply system platform;
[0043] Step 3: After deployment, the system outputs the real-time power compensation amount ΔP of the on-board supercapacitor according to the train's operating conditions. sc and the real-time current monitoring value ΔI of the ground-mounted lithium battery. bat The high-frequency power P of the vehicle-mounted supercapacitor is respectively related to the power of the supercapacitor. sc_ref0 The given reference current I for ground-mounted lithium batteries bat_ref2 By combining these methods, the real-time power P of the vehicle-mounted supercapacitor can be obtained. sc_ref and the actual charge and discharge current I of ground-mounted lithium batteries bat_ref After passing through the lower-level current control, the drive pulse signal for controlling the switching transistor of the bidirectional DC / DC converter is obtained, which controls the charging and discharging of the hybrid energy storage system.
[0044] Step 4: During online training, monitor the system's energy efficiency (e%), voltage regulation (v%), and the SOC of the on-board supercapacitor in real time. sc The system monitors and optimizes the performance index d%. If each performance index remains within the set range, the system continues to use the existing strategy without retraining. If each performance index is below the set range, the system returns to Step 1.
[0045] Step 5: When the system detects that the train has come to a complete stop, the system ends the control task and terminates the operation.
[0046] Beneficial effects: This invention establishes a hierarchical coordinated control architecture and adopts the MATD3 multi-agent reinforcement learning algorithm to perform online learning and optimization of the hierarchical control of the urban rail hybrid energy storage system, thereby improving the energy-saving and voltage-stabilizing characteristics of the DC traction network, avoiding overshoot and over-discharge phenomena in the hybrid energy storage system, and thus extending the life of energy storage components; through online training-online decision-making process, the system achieves online optimization control in the complex and ever-changing urban rail operation environment, comprehensively improving the energy-saving and economic benefits of the urban rail transit system. Attached Figure Description
[0047] Figure 1 This is a diagram of the hierarchical coordinated control structure of the urban rail hybrid energy storage system of the present invention;
[0048] Figure 2 This is a control structure diagram of the urban rail hybrid energy storage system based on the MATD3 algorithm of the present invention;
[0049] Figure 3 This is a schematic diagram of the training framework for the intelligent agent of the present invention;
[0050] Figure 4 This is a structural diagram of the online learning-online decision optimization of the present invention. Detailed Implementation
[0051] To further explain the technical solution of the present invention, the following will be combined with the appendix. Figure 1-4 The present invention will be described in further detail below.
[0052] First, establish a hierarchical coordinated control architecture for the urban rail hybrid energy storage system, such as... Figure 1 As shown, the upper-level energy management system, based on train operation status, energy storage system status, and the energy-saving and voltage-stabilizing requirements of the DC traction network, generates charging and discharging current commands for the ground-based lithium battery and the on-board supercapacitor through voltage control, traction power feedforward control, and power distribution strategies. The lower-level converter control system compares the current commands given by the upper level with the actual current through current control, and outputs PWM signals to control the DC / DC converter for controlling the charging and discharging of the energy storage system. The lower-level converter control system also includes a current dynamic compensation module for dynamically compensating the current control portion of the on-board supercapacitor.
[0053] The current dynamic compensation module will calculate the actual power P of the ground-mounted lithium battery. bat With a given low-frequency power P bat_ref After differential calculation, the voltage U across the vehicle-mounted supercapacitor is... sc By combining these methods, the real-time current compensation amount ΔI of the vehicle-mounted supercapacitor is obtained. sc Then, the current I of the vehicle-mounted supercapacitor. sc Given the charging and discharging current combined with I sc_ref0 By combining these methods, the actual charging and discharging current I of the vehicle-mounted supercapacitor can be obtained. sc_ref The current compensation formula is as follows (1):
[0054]
[0055] Wherein, ΔP bat This represents the difference in power absorbed by the ground-based lithium battery.
[0056] Based on the hierarchical coordination control architecture, MATD3 is used to optimize the hierarchical coordination control architecture online. The control structure is as follows: Figure 2 As shown. In the upper-level energy management system, the angular velocity ω and electromagnetic torque T of the permanent magnet synchronous motor (PMSM) are detected in real time through traction feedforward control. e The traction power requirement P is obtained. load Where η is the energy transfer efficiency; in the ground-based lithium battery voltage control section, the ground-based lithium battery is given a charge / discharge voltage U. bat_char / U bat_dis (i.e., 1520 / 1480) and traction grid voltage U dc The difference is processed by a PI controller, and then by a current limiting module to obtain the given current I of the ground-based lithium battery. bat_ref0In the voltage control section of the vehicle-mounted supercapacitor, the vehicle-mounted supercapacitor is given a charge / discharge voltage U. sc_char / U sc_dis (i.e., 1520 / 1480) and traction grid voltage U dc The difference is processed by a PI controller to obtain the power correction amount ΔP, which is then compared with the traction power demand P. load Combined, the given power P of the hybrid energy storage is obtained hess_ref The high-frequency power P of the vehicle-mounted supercapacitor is then obtained through a low-pass filter (LPF). sc_ref0 and ground-mounted battery low-frequency power P bat_ref0 Based on this, the MATD3 algorithm is used for optimization, and after passing through the PI controller in the lower-level converter control, the drive pulse signal for controlling the switching transistors of the bidirectional DC / DC converter is obtained, thereby realizing the control of charging and discharging of the hybrid energy storage system.
[0057] Based on the control application of multi-agent reinforcement learning algorithm in urban rail hybrid energy storage, the training framework, environmental state, agent actions, and reward function of the agents are designed respectively. The specific method is as follows:
[0058] 1) Design of the training framework for the agent:
[0059] The method employs a "centralized training-distributed learning" training framework for controlling urban rail hybrid energy storage systems. The training framework is as follows: Figure 3 As shown.
[0060] Training Phase (Centralized): Agent A and Agent B share training data, namely, the observations and actions of all agents. In this phase, Agent A and Agent B each have two Critic networks: CriticA1 and CriticA2, and CriticB1 and CriticB2. Each Critic network can access global information, including the states and actions of all agents, to more accurately evaluate the value of their respective actions, thereby improving training effectiveness.
[0061] Each Agent also contains an Actor network, where Actor A and Actor B control the actions of Agent A and Agent B respectively. The Actor network only accesses its own local observations. During training, the Actor network optimizes its action selection by receiving feedback from the Critic network.
[0062] Execution phase (distributed): The trained agents run independently. Agent A decides its actions based only on the partial states it observes, and Agent B decides its actions based only on the partial states it observes.
[0063] Objective: Both agents share the common goal of optimizing the energy-saving and voltage-stabilizing performance of the urban rail system, but each also has its own objective: Agent A focuses on the State of Charge (SOC) of onboard supercapacitors. sc For safety, Agent B focuses on ground-mounted lithium battery SOCs. bat Monitoring and optimization are achieved through the design of reward functions.
[0064] 2) Design of power supply environment and state space for urban rail permanent magnet traction:
[0065] Taking the urban rail permanent magnet traction power supply system as the environment in which the two agents exist, and selecting the DC traction voltage U... dc Hybrid energy storage power given P hess_ref Vehicle-mounted supercapacitors' State of Charge (SOC) sc State of charge (SOC) of ground-mounted lithium batteries bat The voltage U across the vehicle-mounted supercapacitor sc Ground-based lithium battery terminal voltage U bat Train speed ω m and acceleration a c The state space S, representing the state of the environment as observed by the two agent agents, is expressed by equation (2):
[0066]
[0067] Where n represents the nth sampling time of the interaction between MATD3 and the permanent magnet traction power supply system, and n is 700; U dcn P hess_refn P hess_refn SOC scn SOC batn U scn U batn a cn , representing the DC traction grid voltage at the nth sampling time, the given values of the hybrid energy storage system, the state of charge of the on-board supercapacitor, the state of charge of the ground-mounted lithium battery, the terminal voltage of the on-board supercapacitor, the terminal voltage of the ground-mounted lithium battery, the train speed, and the train acceleration, respectively.
[0068] 3) Selection of continuous motion space and execution of motion:
[0069] After observing the state information of the environment, the agent selects an action from the action space according to its own policy π. Therefore, the selected action Action A of Agent A is the power adjustment amount ΔP of the on-board supercapacitor in the upper-level control. sc The continuous action space of Action A is A; the selected action of Agent B, Action B, is the charging and discharging current adjustment amount ΔI of the ground-based lithium battery controlled by the lower layer. bat The continuous action space of Action B is B.
[0070] The continuous action space of the two agents is shown in equation (3) below:
[0071]
[0072] Wherein, ΔP scn ΔI batn These represent the power adjustment of the vehicle-mounted supercapacitor and the charging / discharging current adjustment of the ground-based lithium battery at the nth sampling time, respectively.
[0073] 4) Reward function design:
[0074] The reward function determines that Agent A and Agent B have both cooperative and competitive relationships. The common rewards for the two agents are the voltage regulation rate v% and the energy saving rate e%, but the two agents also have different rewards.
[0075] Agent A's reward r1 is selected as the voltage regulation rate v% and energy saving rate e% within a time step ΔT, as well as the SOC of the on-board supercapacitor. sc The weighted sum of safety variation ranges aims to achieve optimal voltage regulation and energy saving while minimizing the SOC of the vehicle-mounted supercapacitor. sc It can be kept within a safe range of [0.15, 0.85], as shown in equation (4) below:
[0076] r1=max[λ1·v%+μ1·e%+σ1(0.15≤SOC sc ≤0.85)-η1(SOC sc ≤0.15orSOC sc ≥0.85)] (4)
[0077] Its expression is as follows:
[0078] λ1·v%: Voltage regulation rate bonus; the higher the v%, the greater the bonus.
[0079] μ1·e%: Energy saving rate bonus; the higher the e% is, the greater the bonus.
[0080] σ1(0.15≤SOC sc≤0.85): State of Charge (SOC) of vehicle-mounted supercapacitors sc Security Coverage Bonus, i.e., SOC sc Within the safe range [0.15, 0.85], a positive reward σ1 is given;
[0081] η1(SOC sc ≤0.15orSOC sc ≥0.85): State of Charge (SOC) of vehicle-mounted supercapacitors sc Penalty for exceeding limits, i.e., SOC sc If the value is below 0.15 or above 0.85, a negative penalty η1 is applied.
[0082] The Agent B reward r2 is selected as the voltage regulation efficiency v% and energy saving efficiency e% within the time step ΔT, as well as the SOC of the ground-mounted lithium battery. bat The weighted sum of the monitoring optimization index d% is such that the monitoring optimization index d% can be kept within [0,100], as shown in the following formula (5):
[0083] r2=max[λ2·v%+μ2·e%+σ2(0≤d%≤100)-η2(d%<0 or d%>100)] (5)
[0084] Its expression is as follows:
[0085] λ2·v%: Voltage regulation rate bonus; the higher the v%, the greater the bonus.
[0086] μ2·e%: Energy saving rate bonus; the higher the e% is, the greater the bonus.
[0087] σ2 (0≤d%≤100): State of Charge (SOC) of terrestrial lithium batteries bat The monitoring optimization index is rewarded within a safe range, i.e., if d% is within [0,100], a positive reward σ2 is given;
[0088] η2 (d% < 0 or d% > 100): Ground-based lithium battery SOC bat If the monitoring and optimization indicators exceed the limit, i.e., d% is lower than 0 or higher than 100, a negative penalty η2 will be given.
[0089] In equations (4)-(5), λ1,μ1,σ1,η1,λ2,μ2,σ2,η2 are weighting coefficients, with values of 2, 2, 10, 50, 2, 2, 10, and 50, respectively; the energy saving rate e% is defined as the percentage change in the total output energy of the substation before and after the installation of the hybrid energy storage system relative to the total output energy of the substation without the energy storage system, as shown in equation (6) below:
[0090]
[0091] in, These represent the DC traction grid voltages with and without hybrid energy storage, respectively. The DC traction grid voltage without hybrid energy storage is shown below. The value is 1350V; These represent the DC traction network current with and without mixed energy storage, respectively; T represents the train's running time from start-up to braking.
[0092] The voltage regulation rate v% is evaluated by integrating the portion of the DC traction voltage that exceeds / below the limit, as shown in equation (7) below:
[0093]
[0094] in, These represent the set upper and lower safety limits for the DC traction network voltage, i.e., 1520V and 1480V, respectively; Δh and Δl represent the time during which the DC traction voltage exceeds the upper and lower safety limits during train operation.
[0095] Ground-based lithium battery SOC bat The monitoring optimization index d% is defined as the current SOC of the ground-based lithium battery compared to the initial SOC of the ground-based lithium battery without using the MATD3 algorithm, relative to the maximum capacity SOC of the ground-based lithium battery. batmax The percentage deviation is shown in equation (8) below:
[0096]
[0097] Among them, SOC bat The current state of charge (SOC) of the ground-based lithium battery. bat0 The initial SOC of a ground-based lithium battery without using the MATD3 algorithm.
[0098] Finally, an online training-online decision-making method was designed, such as... Figure 4 As shown;
[0099] 1) In the online training module, the urban rail permanent magnet traction power supply system platform is established as the intelligent agent environment, enabling each intelligent agent to interact with the platform; under the condition of random initialization of train running speed, each intelligent agent shares training data and centrally updates the strategy until a stable control strategy is obtained.
[0100] 2) In the online decision-making module, based on real-time operating data, the intelligent agent quickly makes the optimal control decision according to the current state, so as to realize the energy saving and voltage stabilization of the system and improve the service life of the energy storage system.
[0101] The system is running, and the specific steps are as follows:
[0102] Step 1: Two agents learn online from the urban rail permanent magnet traction power supply system environment. By collecting the system's status, actions, and reward information, they interact with the environment repeatedly N times (N≥100) until the cumulative reward stabilizes and converges.
[0103] Step 2: Deploy the trained Agent strategy to the actual urban rail permanent magnet traction power supply system platform;
[0104] Step 3: After deployment, the system outputs the real-time power compensation amount ΔP of the on-board supercapacitor according to the train's operating conditions. sc and the real-time current monitoring value ΔI of the ground-mounted lithium battery. bat The high-frequency power P of the vehicle-mounted supercapacitor is respectively related to the power of the supercapacitor. sc_ref0 The given reference current I for ground-mounted lithium batteries bat_ref2 By combining these methods, the real-time power P of the vehicle-mounted supercapacitor can be obtained. sc_ref and the actual charge and discharge current I of ground-mounted lithium batteries bat_ref After passing through the lower-level current control, the drive pulse signal for controlling the switching transistor of the bidirectional DC / DC converter is obtained, which controls the charging and discharging of the hybrid energy storage system.
[0105] Step 4: During online training, monitor the system's energy efficiency (e%), voltage regulation (v%), and the SOC of the on-board supercapacitor in real time. sc And monitoring and optimization indicators d%, if each performance indicator is maintained within the set range, i.e. e% ≥ 29.86%, v% ≥ 81.43%, SOC sc If the values are within [0.15, 0.85] and d% are within [0, 100], the system will continue to use the existing strategy and will not need to be retrained; if the performance indicators are below the set range, the system will return to Step 1.
[0106] Step 5: When the system detects that the train has come to a complete stop, the system ends the control task and terminates the operation.
[0107] This invention employs a hierarchical coordinated control method for urban rail hybrid energy storage systems based on deep reinforcement learning. Utilizing the MATD3 multi-agent reinforcement learning algorithm, the hierarchical control of the urban rail hybrid energy storage system is learned and optimized online, thereby improving the energy-saving and voltage-stabilizing characteristics of the DC traction network, avoiding overshoot and over-discharge phenomena in the hybrid energy storage system, and extending the lifespan of energy storage components. Through an online training-online decision-making method, online optimized control of the system is achieved under complex and ever-changing urban rail operating environments, comprehensively improving the energy-saving and economic benefits of the urban rail transit system.
[0108] The present invention and its embodiments have been described above. This description is not restrictive, and the accompanying drawings are only one embodiment of the present invention; the actual structure is not limited thereto. In conclusion, if those skilled in the art are inspired by this description and design similar structures and embodiments without departing from the spirit of the invention, such designs should fall within the protection scope of the present invention.
Claims
1. A hierarchical coordinated control method for urban rail hybrid energy storage systems based on deep reinforcement learning, characterized in that... This method constructs a hierarchical coordinated control architecture and utilizes the Multi-agent Twin Delayed Deep Deterministic Policy Gradient (MATD3) algorithm to perform online optimization of the upper-layer power allocation strategy and the lower-layer current control, respectively. Furthermore, it achieves hierarchical optimization and coordinated control of the hybrid energy storage system through an "online training-online decision-making" method. The invention is implemented as follows: First, a hierarchical coordinated control architecture for the urban rail hybrid energy storage system is established, including an upper-level energy management system and a lower-level converter control system. The upper-level energy management system generates charging and discharging current commands for the ground-mounted lithium battery and the vehicle-mounted supercapacitor through voltage control, traction power feedforward control, and power distribution strategies. The lower-level converter control system compares the current commands given by the upper level with the actual current through current control and controls the DC / DC converter to control the charging and discharging of the energy storage system. The lower-level converter control system also includes a current dynamic compensation module, which is used to dynamically compensate the current control part of the vehicle-mounted supercapacitor. The current dynamic compensation module will calculate the actual power P of the ground-mounted lithium battery. bat With a given low-frequency power P bat_ref After differential calculation, the voltage U across the vehicle-mounted supercapacitor is... sc By combining these methods, the real-time current compensation amount ΔI of the vehicle-mounted supercapacitor is obtained. sc Then, the current I of the vehicle-mounted supercapacitor... sc Given the charging and discharging current combined with I sc_ref0 By combining these methods, the actual charging and discharging current I of the vehicle-mounted supercapacitor can be obtained. sc_ref The current compensation formula is as follows (1): Wherein, ΔP bat This represents the difference in power absorbed by the ground-based lithium battery. Secondly, based on the hierarchical coordinated control architecture, MATD3 is used to optimize the control system online. Based on the control application of multi-agent reinforcement learning algorithms in urban rail hybrid energy storage, the training framework, environmental state, agent actions, and reward function of the agents are designed. The specific methods are as follows: 1) Design of the training framework for the agent: The method employs a "centralized training-distributed learning" training framework for controlling urban rail hybrid energy storage systems. The MATD3 agent comprises Agent A and Agent B, which are used to dynamically adjust the power distribution of the hybrid energy storage system and the charging and discharging of the ground-based lithium batteries, respectively. Each agent has a decentralized actuator and two centralized critics. The two agents have both a cooperative relationship (sharing the state information of the urban rail system and sharing the common goal of improving the energy-saving and voltage-stabilizing characteristics of the urban rail system) and a competitive relationship (having different goals, requiring a reward mechanism for differentiation). 2) Design of power supply environment and state space for urban rail permanent magnet traction: Taking the urban rail permanent magnet traction power supply system as the environment in which the two agents exist, and selecting the DC traction voltage U... dc Hybrid energy storage power given P hess_ref Vehicle-mounted supercapacitors' State of Charge (SOC) sc State of charge (SOC) of ground-mounted lithium batteries bat The voltage U across the vehicle-mounted supercapacitor sc Ground-based lithium battery terminal voltage U bat Train speed ω m and acceleration a c The state space S, representing the state of the environment as observed by the two agent agents, is expressed by equation (2): Among them, U dcn P hess_refn P hess_refn SOC scn SOC batn U scn U batn a cn These represent the DC traction network voltage, hybrid energy storage power setpoint, on-board supercapacitor state of charge, ground-mounted lithium battery state of charge, on-board supercapacitor terminal voltage, ground-mounted lithium battery terminal voltage, train speed, and train acceleration, respectively, at the nth sampling time. 3) Selection of continuous motion space and execution of motion: After observing the state information of the environment, the agent selects an action from the action space according to its own policy π, i.e., the mapping relationship from state to action; therefore, the selected action Action A of Agent A is the power adjustment amount ΔP of the on-board supercapacitor in the upper-level control. sc The continuous action space of Action A is A; the selected action of Agent B, Action B, is the charging and discharging current adjustment amount ΔI of the ground-based lithium battery controlled by the lower layer. bat The continuous action space of Action B is B, as shown in equation (3) below: Wherein, ΔP scn ΔI batn These represent the power adjustment of the vehicle-mounted supercapacitor and the charging / discharging current adjustment of the ground-based lithium battery at the nth sampling time, respectively. 4) Reward function design: The reward function determines that Agent A and Agent B have both cooperative and competitive relationships. The common rewards for the two agents are the voltage regulation rate v% and the energy saving rate e%, but the two agents also have different rewards. Agent A's reward r1 is selected as the voltage regulation rate v% and energy saving rate e% within a time step ΔT, as well as the SOC of the on-board supercapacitor. sc The weighted sum of safety variation ranges aims to achieve optimal voltage regulation and energy saving while minimizing the SOC of the vehicle-mounted supercapacitor. sc Able to remain within a safe range [SOC] scmin SOC scmax Within ], as shown in equation (4): The Agent B reward r2 is selected as the voltage regulation efficiency v% and energy saving efficiency e% within the time step ΔT, as well as the SOC of the ground-mounted lithium battery. bat The weighted sum of the monitoring optimization index d%, and the monitoring optimization index d% can be kept within [d]. min ,d max Within ], as shown in the following formula (5): r2=max[λ2·v%+μ2·e%+σ2(d min ≤d%≤d max )-η2(d%<d min or d%>d max )] (5) In equations (4)-(5), λ1,μ1,σ1,η1,λ2,μ2,σ2,η2 are weighting coefficients; SOC scmin SOC scmax These are vehicle-mounted supercapacitor SOCs sc The minimum and maximum safety ranges; d min d max The minimum and maximum monitoring optimization index ranges are respectively defined; the energy saving rate e% is defined as the percentage change in the total output energy of the substation before and after the installation of the hybrid energy storage system relative to the total output energy of the substation without the energy storage system, as shown in the following formula (6): in, These represent the DC traction grid voltage with and without hybrid energy storage installed, respectively. These represent the DC traction network currents with and without hybrid energy storage, respectively, and T represents the train's running time from start-up to braking. The voltage regulation rate v% is evaluated by integrating the portion of the DC traction voltage that exceeds / below the limit, as shown in equation (7) below: in, These represent the set upper and lower safety limits for the DC traction network voltage, respectively, and Δh / Δl represent the time during which the DC traction voltage exceeds the upper and lower safety limits during train operation. Ground-based lithium battery SOC bat The monitoring optimization index d% is defined as the current SOC of the ground-based lithium battery compared to the initial SOC of the ground-based lithium battery without using the MATD3 algorithm, relative to the maximum capacity SOC of the ground-based lithium battery. batmax The percentage deviation is shown in equation (8) below: Among them, SOC bat The current state of charge (SOC) of the ground-based lithium battery. bat0 The initial SOC of a ground-based lithium battery without using the MATD3 algorithm; Finally, an online training-online decision-making method was designed: 1) In the online training module, the urban rail permanent magnet traction power supply system platform is established as the intelligent agent environment, enabling each intelligent agent to interact with the platform; under the condition of random initialization of train running speed, each intelligent agent shares training data and centrally updates the strategy until a stable control strategy is obtained. 2) In the online decision-making module, based on real-time operating data, the intelligent agent quickly makes the optimal control decision according to the current state, so as to achieve energy saving and voltage stabilization and improve the service life of the energy storage system; The system is running, and the specific steps are as follows: Step 1: Two agents learn online from the urban rail permanent magnet traction power supply system environment. By collecting the system's status, actions and reward information, they interact with the environment repeatedly N times until the accumulated reward stabilizes and converges. Step 2: Deploy the trained Agent strategy to the actual urban rail permanent magnet traction power supply system platform; Step 3: After deployment, the system outputs the real-time power compensation amount ΔP of the on-board supercapacitor according to the train's operating conditions. sc and the real-time current monitoring value ΔI of the ground-mounted lithium battery. bat The high-frequency power P of the vehicle-mounted supercapacitor is respectively related to the power of the supercapacitor. sc_ref0 The given reference current I for ground-mounted lithium batteries bat_ref2 By combining these methods, the real-time power P of the vehicle-mounted supercapacitor can be obtained. sc_ref and the actual charge and discharge current I of ground-mounted lithium batteries bat_ref After passing through the lower-level current control, the drive pulse signal for controlling the switching transistor of the bidirectional DC / DC converter is obtained, which controls the charging and discharging of the hybrid energy storage system. Step 5: During online training, monitor the system's energy efficiency, voltage regulation rate, and the SOC of the on-board supercapacitor in real time. sc And monitoring and optimization index d%, if each performance index is maintained within the set range, i.e. e%≥E, v%≥V, SOC sc In [SOC] scmin SOC scmax Within ], d% is within [d min ,d max If the performance indicators are within the set range, the system will continue to use the existing strategy and no retraining is required; if the performance indicators are below the set range, return to Step 1. Step 6: When the system detects that the train has come to a complete stop, the system ends the control task and terminates the operation.