A flexible DC voltage regulator intelligent control system and method based on reinforcement learning

By adopting an intelligent control system based on reinforcement learning in a flexible DC voltage regulator, optimizing the control strategy and adjusting the output parameters in real time, the problem that traditional control strategies are difficult to adapt to when load changes and grid fluctuations is solved, and a voltage control effect with high stability and low energy loss is achieved.

CN119628039BActive Publication Date: 2025-05-09STATE GRID SHANGHAI MUNICIPAL ELECTRIC POWER CO
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510161693.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-02-14
Publication Date
2025-05-09
Estimated Expiration
2045-02-14

AI Technical Summary

Technical Problem

When the load changes violently or the power grid fluctuates greatly, traditional control strategies are difficult to adapt, resulting in system response lag and voltage fluctuations, making it difficult to effectively control the voltage under complex power grid conditions.

Method used

An intelligent control system based on reinforcement learning is adopted, through the reinforcement learning controller, the operating status of the source and user-side equipment is monitored, and the control strategy is optimized using the DQN algorithm to adjust the output voltage, power and duty cycle of the source side rectifier module and the user-side inverter module in real time.

Benefits of technology

It realizes rapid response and adaptive optimization under load changes and grid fluctuations, improves the stability and control accuracy of the system, improves the sensitivity and response speed to grid and load changes, and minimizes energy loss while maintaining voltage stability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119628039B_ABST
    Figure CN119628039B_ABST
Patent Text Reader

Abstract

The present invention relates to a flexible DC voltage regulator intelligent control system and method based on reinforcement learning, the system includes a source-side rectifier module, a user-side inverter module and a reinforcement learning controller, the source-side rectifier module and the user-side inverter module are connected through a flexible DC line, the source-side rectifier module is used to convert the AC power of the source side into DC power, and transmit the electric energy to the user side through the flexible DC line; the user-side inverter module is used to reconvert the DC power into AC power for use by the load; the reinforcement learning controller is used to monitor the operating status of the source-side and user-side devices, analyze real-time data and interact with the environment, and obtain the optimized control strategy of the source-side rectifier module and the user-side inverter module through the reinforcement learning intelligent control algorithm, and adjust the output voltage of the source-side rectifier module and the user-side inverter module according to the optimized control strategy. Compared with the prior art, the present invention realizes real-time voltage regulation through the optimized control strategy under load changes and power grid fluctuations.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of intelligent control technology, and in particular to a flexible direct current regulator intelligent control system and method based on reinforcement learning. Background Art

[0002] With the access of renewable energy sources such as photovoltaic and wind power to the power grid, the load volatility of the power system has increased, especially under dynamically changing load conditions, which puts higher requirements on voltage stability. Existing flexible DC voltage regulators generally use traditional control algorithms based on PID control, fuzzy control, etc. to adjust the voltage. Although the structure is simple and easy to implement, traditional control strategies are often difficult to adapt to when the load changes drastically or the power grid fluctuates greatly, resulting in delayed system response and large voltage fluctuations. It is difficult to effectively control the voltage under complex power grid conditions.

[0003] Existing flexible DC voltage regulators usually include a source-side rectifier module and a user-side inverter module, which are connected through a ±375VDC bipolar DC power supply. On the source side, the rectifier module converts AC power into DC power and transmits it to the user-side inverter module through a DC line to complete the conversion from DC to AC. These modules rely on traditional control algorithms with fixed parameters for adjustment. When faced with dynamic load changes, traditional control strategies are difficult to achieve rapid response and adaptive optimization under drastic load changes and grid fluctuations, and control performance will also decline. Summary of the invention

[0004] The purpose of the present invention is to overcome the defects of the above-mentioned prior art and to provide a flexible DC voltage regulator intelligent control system and method based on reinforcement learning, which realizes real-time voltage regulation by optimizing control strategy under load changes and power grid fluctuations.

[0005] The purpose of the present invention can be achieved by the following technical solutions:

[0006] A flexible DC voltage regulator intelligent control system based on reinforcement learning comprises: a source-side rectifier module, a user-side inverter module and a reinforcement learning controller, wherein the source-side rectifier module and the user-side inverter module are connected through a flexible DC line, the source-side rectifier module is used to convert the AC power of the source side into DC power, and transmit the electric energy to the user side through the flexible DC line; the user-side inverter module is used to reconvert the DC power into AC power for use by the load; the reinforcement learning controller is used to monitor the operating status of the source-side and user-side devices, and obtain an optimization control strategy based on the operating status through a reinforcement learning intelligent control algorithm, and jointly adjust the output voltage, power and duty cycle of the source-side rectifier module and the user-side inverter module according to the optimization control strategy, take the operating status of the system as the state space of the reinforcement learning intelligent control algorithm, take the control actions of the source-side rectifier module and the user-side inverter module as the action space of the reinforcement learning intelligent control algorithm, set a reward function according to the voltage regulation effect and energy loss, and the reinforcement learning intelligent control algorithm obtains the optimization control strategy by updating the Q value through a DQN algorithm, and the Q value is related to the state space and the action space.

[0007] Furthermore, the state space is:

[0008] ,

[0009] In the formula, is the state space, is the AC voltage input on the source side, is the voltage of the flexible DC line, is the current of the user side load, is the power of the user side load, is the deviation between the actual load current and the target current, is the deviation between the actual output voltage and the target voltage.

[0010] Furthermore, the action space is:

[0011] ,

[0012] In the formula, is the action space, is the output voltage of the source side rectifier module, is the output voltage of the inverter module on the user side, is the duty cycle.

[0013] Furthermore, the source-side rectifier module and the user-side inverter module are respectively controlled by the reinforcement learning controller, and the output voltage of the source-side rectifier module is:

[0014] ,

[0015] In the formula, is the output voltage of the source side rectifier module, is the voltage of the flexible DC line, is the duty cycle;

[0016] The output voltage of the user-side inverter module is:

[0017] ,

[0018] In the formula, is the output voltage of the inverter module on the user side, is the load output voltage.

[0019] Furthermore, the reward function is:

[0020] ,

[0021] In the formula, is the load output voltage, is the target voltage, To balance the voltage stability and energy loss, is the energy loss in the system.

[0022] Furthermore, the energy loss is:

[0023] ,

[0024] In the formula, is the energy loss, is the required power of the user side load, is the output power of the inverter.

[0025] Furthermore, the update formula of the Q value is:

[0026] ,

[0027] In the formula, Status Take action The value function of is the learning rate, is the reward at the current moment, is the discount factor, For the next state The maximum Q value, For the status An action from the set of all possible actions.

[0028] Further, according to the Q value, The greedy strategy selects the action.

[0029] Furthermore, the reinforcement learning intelligent control algorithm stabilizes the reinforcement learning process through an experience replay mechanism, and the experience replay mechanism updates the network parameters by minimizing the loss function, and the loss function is:

[0030] ,

[0031] ,

[0032] In the formula, is the loss function, is the average value of all sampled data, For interactive experience, is the target Q value, is the prediction value of the main Q network, is the current state, For the actions taken, are the parameters of the main Q network, is the reward at the current moment, For the next state, is the discount factor, To use the next state of the target network The maximum Q value, For the status An action in the set of all actions below, are the parameters of the target network.

[0033] According to another aspect of the present invention, a method for intelligent control of a flexible DC regulator based on reinforcement learning is provided, comprising the following steps:

[0034] The operation status of the source-side and user-side devices is monitored by a reinforcement learning controller, and the optimization control strategy of the source-side rectifier module and the user-side inverter module is obtained according to the operation status by a reinforcement learning intelligent control algorithm;

[0035] The output voltage, power and duty cycle of the source-side rectifier module and the user-side inverter module are adjusted according to the optimization control strategy.

[0036] Compared with the prior art, the present invention has the following beneficial effects:

[0037] 1. The present invention monitors the operating status of the source side and the user side through a reinforcement learning controller, continuously optimizes and adjusts the control strategy through the DQN algorithm according to the operating status, obtains the optimized control strategy of the source side rectifier module and the user side inverter module, and adjusts the output voltage and duty cycle of the source side rectifier module and the user side inverter module according to the optimized control strategy to ensure that the voltage regulator can quickly respond to load fluctuations, improve the stability and control accuracy of the system, and enhance the sensitivity and response speed of the system to changes in the power grid and load.

[0038] 2. The present invention takes into account the intelligent control effect and energy consumption through the design of the reinforcement learning reward function, balances the two in real time, avoids unnecessary adjustment actions, and minimizes energy loss while maintaining voltage stability. Through real-time adjustment, the voltage stabilizer can achieve optimal energy efficiency under various load conditions and reduce long-term operating costs. BRIEF DESCRIPTION OF THE DRAWINGS

[0039] Figure 1 A structural schematic diagram of a flexible DC voltage regulator intelligent control system based on reinforcement learning proposed by the present invention;

[0040] Figure 2 A schematic diagram of a flow chart of a flexible DC voltage regulator intelligent control method based on reinforcement learning proposed by the present invention;

[0041] Figure 3 Flowchart of the reinforcement learning intelligent control algorithm. DETAILED DESCRIPTION

[0042] The present invention is described in detail below in conjunction with the accompanying drawings and specific embodiments. This embodiment is implemented based on the technical solution of the present invention, and provides a detailed implementation method and specific operation process, but the protection scope of the present invention is not limited to the following embodiments.

[0043] Abbreviations involved:

[0044] Deep Q Network: Deep Q Network, DQN Example 1

[0045] This embodiment provides a flexible DC regulator intelligent control system based on reinforcement learning, such as Figure 1As shown, it includes: a source-side rectifier module 1, a user-side inverter module 2 and a reinforcement learning controller 3, wherein the source-side rectifier module 1 and the user-side inverter module 2 are connected via a flexible DC line, the source-side rectifier module 1 is used to convert the AC power on the source side into DC power, and transmit the electric energy to the user side via the flexible DC line; the user-side inverter module 2 is used to reconvert the DC power into AC power for use by the load; the reinforcement learning controller 3 is used to monitor the operating status of the source-side and user-side devices, and obtain the optimized control strategy of the source-side rectifier module 1 and the user-side inverter module 2 according to the operating status through the reinforcement learning intelligent control algorithm, and adjust the output voltage and duty cycle of the source-side rectifier module 1 and the user-side inverter module 2 according to the optimized control strategy.

[0046] The source-side rectifier module 1 is responsible for converting the AC power on the source side into ±375V DC power, outputting it to the DC line, and providing power transmission to the user side. The output voltage of the source-side rectifier module 1 is adjustable to adapt to different loads and grid conditions. The DC voltage is transmitted to the user side through a flexible DC line, which can reduce energy loss during long-distance power transmission.

[0047] The user-side inverter module 2 converts ±375V DC power back into AC power for the load. The output voltage of the user-side inverter module 2 is also adjustable to ensure voltage stability at the load end. In the user-side inverter module 2, the system adjusts the duty cycle 𝐷 to control the output voltage and power, thereby optimizing the energy efficiency of the system.

[0048] The reinforcement learning controller 3 is responsible for monitoring the operating status of the source-side and user-side devices, analyzing real-time data and interacting with the environment, and obtaining the optimal control strategy of the source-side rectifier module 1 and the user-side inverter module 2 through the reinforcement learning algorithm. The reinforcement learning controller 3 can adaptively adjust the system parameters to ensure voltage stability while minimizing energy loss.

[0049] The design elements of reinforcement learning intelligent control algorithm include state space, action space and reward function.

[0050] The reinforcement learning intelligent control algorithm takes the operating status of the source-side and user-side equipment as the state space of the reinforcement learning intelligent control algorithm, considers the global optimal control of voltage and power in the flexible DC system, and optimizes the duty cycle of the source-side rectifier module and the user-side inverter module to ensure the coordinated control of the double-end voltage, optimize voltage stability and energy loss, and takes the control actions of the source-side rectifier module 1 and the user-side inverter module 2 as the action space of the reinforcement learning intelligent control algorithm, and sets the reward function according to the voltage regulation effect and energy loss.

[0051] The elements in the state space are the operating states of the system, and the state space is:

[0052] ,

[0053] In the formula, is the state space, is the AC voltage input on the source side, is the voltage of the flexible DC line, is the current of the user side load, is the power of the user side load, is the deviation between the actual load current and the target current, is the deviation between the actual output voltage and the target voltage.

[0054] The action space is the control action of the source-side rectifier module 1 and the user-side inverter module 2. The action space is:

[0055] ,

[0056] In the formula, is the action space, is the output voltage of the source side rectifier module, is the output voltage of the inverter module on the user side, is the duty cycle.

[0057] The voltage and current changes of the system follow the dynamic equation of the power electronic converter. The source-side rectifier module 1 and the user-side inverter module 2 are controlled by the reinforcement learning controller 3 respectively. The output voltage of the source-side rectifier module 1 is:

[0058] ,

[0059] In the formula, is the output voltage of the source side rectifier module 1, is the voltage of the flexible DC line, is the duty cycle;

[0060] The output voltage of the user-side inverter module 2 is:

[0061] ,

[0062] In the formula, is the output voltage of the user-side inverter module 2, is the load output voltage.

[0063] The reward function is:

[0064] ,

[0065] In the formula, is the load output voltage, is the target voltage, To balance the voltage stability and energy loss, is the energy loss in the system.

[0066] The energy loss is:

[0067] ,

[0068] In the formula, is the energy loss, is the required power of the user side load, is the output power of the inverter.

[0069] The reinforcement learning intelligent control algorithm updates the Q value through the DQN algorithm to obtain the optimal control strategy. The reinforcement learning intelligent control algorithm updates the Q value through the DQN algorithm to obtain the optimal control strategy. The Q value is related to the state space and action space. The update formula of the Q value is:

[0070] ,

[0071] In the formula, Status Take action The value function of is the learning rate, is the reward at the current moment, is the discount factor, For the next state The maximum Q value, For the status An action from the set of all possible actions.

[0072] The reinforcement learning intelligent control algorithm stabilizes the reinforcement learning process through the experience replay mechanism. The experience replay mechanism updates the network parameters by minimizing the loss function. The loss function is:

[0073] ,

[0074] ,

[0075] In the formula, is the loss function, is the average value of all sampled data, For interactive experience, is the target Q value, is the prediction value of the main Q network, is the current state, For the actions taken, are the parameters of the main Q network, is the reward at the current moment, For the next state, is the discount factor, To use the next state of the target network The maximum Q value, For the status An action in the set of all actions below, are the parameters of the target network.

[0076] The reinforcement learning intelligent control algorithm adopts the DQN algorithm, which avoids the control oscillation problem caused by too fast strategy update in the DRL method through experience replay and target Q network optimization strategy, and improves voltage stability and convergence speed. The specific steps to obtain the optimized control strategy are as follows: Figure 3 As shown, the following steps are included.

[0077] Step 1: Initialization

[0078] Initialize main Q network parameters , and initialize the target network parameters . Initialize the experience replay pool.

[0079] Step 2: State Observation

[0080] At time t, the system observes the current operating status , including grid voltage, load current, voltage error and other information.

[0081] Step 3: Action Selection

[0082] According to the current Q value and Greedy strategy, choosing actions , including adjusting the output voltage and duty cycle of the rectifier and inverter.

[0083] Step 4: Perform actions and calculate rewards

[0084] Execute an action , adjust the output of the rectifier and inverter, and calculate the current reward , based on load voltage stability and energy loss.

[0085] Step 5: Store experience and update Q network

[0086] The current state ,action ,award and the next state Stored in the experience replay pool.

[0087] Sample from the experience replay pool and update the main Q network parameters via backpropagation .

[0088] Step 6: Update the target network

[0089] Periodically copy the Q network parameters to the target network.

[0090] Step 7: Iterate Learning

[0091] By repeating steps 2 to 6, the system gradually learns the optimal control strategy to achieve voltage stability and maximize energy efficiency. Example 2

[0092] This embodiment provides a flexible DC regulator intelligent control method based on reinforcement learning, such as Figure 2 As shown, the following steps are included:

[0093] S1, monitor the operating status of the source-side and user-side devices through the reinforcement learning controller 3, and obtain the optimization control strategy of the source-side rectifier module 1 and the user-side inverter module 2 according to the operating status through the reinforcement learning intelligent control algorithm;

[0094] S2. Adjust the output voltage and duty cycle of the source-side rectifier module 1 and the user-side inverter module 2 according to the optimization control strategy.

[0095] The rest is the same as in Example 1.

[0096] The preferred specific embodiments of the present invention are described in detail above. It should be understood that a person skilled in the art can make many modifications and changes based on the concept of the present invention without creative work. Therefore, any technical solution that can be obtained by a person skilled in the art through logical analysis, reasoning or limited experiments based on the concept of the present invention on the basis of the prior art should be within the scope of protection determined by the claims.

Claims

1. A flexible DC voltage regulator intelligent control system based on reinforcement learning, characterized in that: include: A source-side rectifier module (1), a user-side inverter module (2) and a reinforcement learning controller (3), wherein the source-side rectifier module (1) and the user-side inverter module (2) are connected via a flexible DC line, the source-side rectifier module (1) is used to convert the alternating current (AC) on the source side into DC, and transmit the electric energy to the user side via the flexible DC line; the user-side inverter module (2) is used to convert the DC into AC for use by the load; the reinforcement learning controller (3) is used to monitor the operating status of the source-side and user-side devices, and obtain the optimal control algorithm based on the operating status through the reinforcement learning intelligent control algorithm. The invention discloses a method for optimizing a control strategy, wherein the output voltage, power and duty cycle of the source-side rectifier module (1) and the user-side inverter module (2) are jointly adjusted according to the optimized control strategy, the operating state of the system is used as the state space of the reinforcement learning intelligent control algorithm, the control actions of the source-side rectifier module (1) and the user-side inverter module (2) are used as the action space of the reinforcement learning intelligent control algorithm, and a reward function is set according to the voltage regulation effect and the energy loss. The reinforcement learning intelligent control algorithm updates the Q value through the DQN algorithm to obtain the optimized control strategy, and the Q value is related to the state space and the action space.

2. According to the reinforcement learning-based flexible DC voltage regulator intelligent control system of claim 1, it is characterized in that: The state space is: S t =[V grid ,V dc ,I load ,P load ,I err ,V err ], In the formula, S t is the state space, V grid is the AC voltage input on the source side, V dc is the voltage of the flexible DC line, I load is the current of the user side load, P load is the power of the user side load, I err is the deviation between the actual load current and the target current, V err is the deviation between the actual output voltage and the target voltage.

3. The flexible DC voltage regulator intelligent control system based on reinforcement learning according to claim 1 is characterized in that: The action space is: And t =[V rect ,In inv ,D], In the formula, A t is the action space, V rect is the output voltage of the source side rectifier module (1), V inv is the output voltage of the user-side inverter module (2), and D is the duty cycle.

4. The flexible DC voltage regulator intelligent control system based on reinforcement learning according to claim 1 is characterized in that: The source-side rectifier module (1) and the user-side inverter module (2) are respectively controlled by the reinforcement learning controller (3), and the output voltage of the source-side rectifier module (1) is: Where V rect is the output voltage of the source side rectifier module (1), V dc (t) is the voltage of the flexible DC line, D(t) is the duty cycle; The output voltage of the user-side inverter module (2) is: V inv =V out (t)·D(t), Where V inv is the output voltage of the user-side inverter module (2), V out (t) is the load output voltage.

5. The flexible DC voltage regulator intelligent control system based on reinforcement learning according to claim 1 is characterized in that: The reward function is: Where V out is the load output voltage, V target is the target voltage, λ is the weight coefficient for weighing voltage stability and energy loss, P loss is the energy loss in the system.

6. The flexible DC voltage regulator intelligent control system based on reinforcement learning according to claim 5 is characterized in that: The energy loss is: P loss (t)=P load (t)-P output (t), Where P loss (t) is energy loss, P load (t) is the power of the user side load, P output (t) is the output power of the inverter.

7. The flexible DC voltage regulator intelligent control system based on reinforcement learning according to claim 1 is characterized in that: The update formula of the Q value is: In the formula, Q(s t , a t ) is state s t Take action a t The value function, α is the learning rate, r t is the reward at the current moment, γ is the discount factor, For the next state s t+1 The maximum Q value of a′ is the maximum Q value of the state s t+1 An action from the set of all possible actions.

8. The flexible DC voltage regulator intelligent control system based on reinforcement learning according to claim 7 is characterized in that: The action is selected according to the Q value through an ε-greedy strategy.

9. The intelligent control system for flexible DC voltage regulator based on reinforcement learning according to claim 1 is characterized in that: The reinforcement learning intelligent control algorithm stabilizes the reinforcement learning process through the experience replay mechanism, and the experience replay mechanism updates the network parameters by minimizing the loss function, and the loss function is: Where L(θ) is the loss function, is the average value of all sampled data, (s t , a t , r t ,s t+1 ) is the interaction experience, y t is the target Q value, Q(s t , a t ; θ) is the predicted value of the main Q network, s t is the current state, a t is the action taken, θ is the parameter of the main Q network, r t is the reward at the current moment, s t+1 is the next state, γ is the discount factor, To use the next state s of the target network t+1 The maximum Q value of a′ is the maximum Q value of the state s t+1 An action in the set of all actions, θ - are the parameters of the target network.

Citation Information

Patent Citations

  • Flexible DC system DC bus voltage control method and device based on deep reinforcement learning

    CN113113928A