Direct current microgrid voltage dynamic trajectory optimization method and system based on reinforcement learning

CN122823364APending Publication Date: 2026-09-25SHANGHAI JIAOTONG UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611317704.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-08-28
Publication Date
2026-09-25

AI Technical Summary

Technical Problem

[0007]本发明提供了基于强化学习的直流微网电压动态轨迹优化方法及系统,解决了现有技术中无法使强化学习智能体在稳态期自动退出不干扰下垂控制的功率分配功能及无法在不改变稳态工作点的前提下改善暂态电压响应等的技术问题

Benefits of technology

现有基于强化学习的补偿方法多针对恒压控制设计,不适用于稳态电压随负载电流变化的下垂控制系统;且现有方法的强化学习智能体在稳态期仍输出非零控制量,偏移稳态工作点,在多变换器并联系统中导致功率分配误差。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122823364A_ABST
    Figure CN122823364A_ABST
Patent Text Reader

Abstract

The application provides a direct-current microgrid voltage dynamic trajectory optimization method and system based on reinforcement learning. It is applied to the technical field of direct-current microgrid voltage control, comprising real-time acquisition of bus voltage and output current of the direct-current microgrid, and calculation of voltage tracking error in combination with rated voltage of the direct-current microgrid and droop coefficient of a droop controller; the rate of change of the voltage tracking error is calculated, a three-dimensional observation signal is constructed based on the voltage tracking error, the rate of change and the output current, and after normalized processing, the three-dimensional observation signal is input into a reinforcement learning intelligent agent, and the reinforcement learning intelligent agent is used for outputting compensation voltage; a gating weight is calculated according to the voltage tracking error and a soft gating function, and a gated compensation voltage is calculated according to the compensation voltage and the gating weight; the gated compensation voltage is superimposed with the output voltage of the droop controller and then input into the controller, and the final control instruction is output to an inner loop controller for execution of control. The application fundamentally solves the problem of interference with the droop control of the reinforcement learning intelligent agent in the steady state period.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of DC microgrid voltage control technology, and in particular to a method and system for optimizing the dynamic trajectory of DC microgrid voltage based on reinforcement learning. Background Technology

[0002] DC microgrids have been widely used in distributed generation systems due to their advantages such as simple structure, high conversion efficiency, and convenient access to renewable energy. In a DC microgrid, the DC-DC converter is the core power electronic device connecting the power source and the DC bus, and its control performance directly affects the operational stability of the microgrid.

[0003] Droop control is a commonly used distributed control strategy in DC microgrids, enabling autonomous power allocation among different converters. However, droop control suffers from insufficient transient response performance during load surges: Sudden load changes can cause significant overshoot and long recovery time in the bus voltage, especially under the negative impedance characteristics of constant power loads (CPL). The transient process may be accompanied by severe voltage oscillations or even instability.

[0004] In recent years, reinforcement learning (RL) technology has been introduced into the field of power electronic control to improve the transient response performance of traditional controllers. However, existing reinforcement learning-based power electronic control methods mainly fall into the following three paradigms: Paradigm A uses reinforcement learning to replace traditional control structures, including replacing specific control loops, replacing the entire regulation loop to directly output the duty cycle, or bypassing PWM to directly perform gate-level switching control. Such methods are simple but sacrifice the stability guarantee and physical protection of traditional controllers. Paradigm B utilizes reinforcement learning to adaptively adjust controller parameters such as PI gain or droop coefficient online. This approach preserves the underlying controller but requires access to internal parameters, which may not be available in commercially available converters with closed firmware. Paradigm C uses reinforcement learning agents as supplementary compensators, adding correction signals to the existing control structure. This approach requires minimal modification and can be deployed as an add-on module, fully preserving the inherent protection mechanisms.

[0005] Of the three paradigms mentioned above, paradigm C is the most attractive for practical deployment. However, applying it to droop-controlled DC microgrids still presents two key challenges: First, existing compensation methods are mostly designed for constant voltage control and are not suitable for droop control systems where steady-state voltage changes with load current. Second, the non-zero output of the reinforcement learning agent during the steady-state period inevitably deviates from the steady-state operating point. In a multi-converter parallel system, the steady-state voltage set by the droop control is crucial for proportional power distribution. Even a small compensation offset will reduce the power distribution accuracy.

[0006] Therefore, there is an urgent need for a method to optimize the dynamic trajectory of droop-controlled DC microgrid voltage in order to simultaneously solve the following two key problems: (1) How to enable the RL agent to automatically exit the power distribution function without interfering with droop control during the steady state period; (2) How to design a compensation architecture suitable for droop control systems to improve transient voltage response without changing the steady-state operating point. Summary of the Invention

[0007] This invention provides a method and system for optimizing the dynamic voltage trajectory of a DC microgrid based on reinforcement learning, which solves the technical problems in the prior art, such as the inability to enable the reinforcement learning agent to automatically exit the power allocation function without interfering with the droop control during the steady state period and the inability to improve the transient voltage response without changing the steady-state operating point.

[0008] According to a first aspect of the present invention, a method for optimizing the dynamic voltage trajectory of a DC microgrid based on reinforcement learning is provided. The method includes: The bus voltage and output current of the DC microgrid are collected in real time, and the voltage tracking error is calculated by combining the rated voltage of the DC microgrid and the droop coefficient of the droop controller. The rate of change of the voltage tracking error is calculated, and a three-dimensional observation signal is constructed based on the voltage tracking error, the rate of change, and the output current. After normalization, the signal is input into the reinforcement learning agent, which is used to output the compensation voltage. The gate weight is calculated based on the voltage tracking error and the soft gate function, and the post-gated compensation voltage is calculated based on the compensation voltage and the gate weight. The gated compensation voltage is superimposed with the output voltage of the droop controller and then input into the inner loop PI controller. The final control command is then output to the inner loop controller for execution. Specifically, based on the voltage tracking error, the gating compensation voltage, and the historical values ​​of the gating compensation voltage, a comprehensive reward signal is calculated using a constructed four-component reward function, and the comprehensive reward signal is used to train the reinforcement learning agent.

[0009] Furthermore, the voltage tracking error is obtained by subtracting the droop coefficient and the output current from the rated voltage and then comparing it with the bus voltage; In steady state, the voltage tracking error approaches zero.

[0010] Furthermore, the three-dimensional observation signal uses the voltage tracking error as the first dimension, the rate of change of the voltage tracking error as the second dimension, and the output current as the third dimension, and maps the data signals of each dimension to a preset normalization boundary.

[0011] Furthermore, the gating weights are obtained using the Sigmoid soft-gating function and with the magnitude of the voltage tracking error as input; When the DC microgrid is in a steady state and the voltage tracking error amplitude is less than a first preset threshold, the gating weight tends to 0. When the DC microgrid is in a transient period and the voltage tracking error amplitude is greater than the second preset threshold, the gating weight tends to 1, and the second preset threshold is greater than the first preset threshold.

[0012] Furthermore, the four-component reward function includes a transient error penalty term, a steady-state zero output mandatory term, a directional guidance term, and a rate of change penalty term, all of which are controlled by gating weights; The transient error penalty term is used to drive the agent to minimize the voltage tracking error during the transient period; The steady-state zero-output coercion term is used to guide the reinforcement learning agent to approach zero in the steady-state period; The directional guidance term is used to guide the agent to output a compensation voltage that is consistent with the direction of the voltage tracking error during the transient period; The action change rate penalty term is used to suppress oscillations in the agent's output.

[0013] Furthermore, the training process of the reinforcement learning agent includes: Only when the DC microgrid reaches steady state is the state of the integrator in the DC microgrid saved as a steady-state snapshot. Based on the steady-state snapshot, normalized three-dimensional observation signal samples are obtained under random load disturbances; The three-dimensional observation signal samples are input into the reinforcement learning agent for training. It is determined whether the four-component reward function has converged. If it has, the trained reinforcement learning agent is obtained. Otherwise, the steady-state snapshot is loaded as the initial state, and then the process proceeds to step 2 until the four-component reward function converges.

[0014] According to a second aspect of the present invention, the present invention provides a DC microgrid voltage dynamic trajectory optimization system based on reinforcement learning, comprising: The voltage tracking error calculation module is used to collect the bus voltage and output current of the DC microgrid in real time, and calculate the voltage tracking error by combining the rated voltage of the DC microgrid and the droop coefficient of the droop controller. The reinforcement learning agent module is connected to the voltage tracking error calculation module. It is used to calculate the rate of change of the voltage tracking error, construct a three-dimensional observation signal based on the voltage tracking error, the rate of change and the output current, and input the signal into the reinforcement learning agent after normalization. The reinforcement learning agent is used to output the compensation voltage. A soft gating module, connected to the reinforcement learning agent module, is used to calculate the gating weights based on the voltage tracking error and the soft gating function, and to calculate the post-gating compensation voltage based on the compensation voltage and the gating weights. The output voltage compensation module is connected to the soft gating module and is used to superimpose the gating compensation voltage with the output voltage of the droop controller and input it into the inner loop PI controller, and output the final control command to the inner loop controller for control execution. The reward calculation and training module is connected to the voltage tracking error calculation module, the reinforcement learning agent module, the soft gating module, and the output voltage compensation module. It is used to calculate the comprehensive reward signal based on the voltage tracking error, the gating compensation voltage, and the historical value of the gating compensation voltage through a constructed four-component reward function, and to train the reinforcement learning agent using the comprehensive reward signal.

[0015] According to a third aspect of the present invention, an electronic device is provided, including a memory, a processor, and a computer program; wherein the computer program is stored in the memory and configured to be executed by the processor to implement the method as described in the first aspect.

[0016] According to a fourth aspect of the present invention, a computer-readable storage medium is provided, characterized in that a computer program is stored thereon; the computer program is executed by a processor to implement the method as described in the first aspect.

[0017] The beneficial effects of this invention are: Existing reinforcement learning-based compensation methods are mostly designed for constant voltage control and are not suitable for droop control systems where steady-state voltage varies with load current. Furthermore, the reinforcement learning agent in existing methods still outputs non-zero control values ​​during the steady-state period, deviating from the steady-state operating point and causing power distribution errors in multi-converter parallel systems.

[0018] The method of this invention fundamentally solves the problem of droop control of reinforcement learning agents in the steady state by designing voltage tracking error (naturally returning to zero in steady state) and using a dual guarantee mechanism (soft gating hard constraints and reward function soft constraints). At the same time, it achieves voltage dynamic trajectory optimization without changing the steady-state operating point through an output voltage compensation architecture.

[0019] Specifically, this application incorporates the characteristics of a droop controller into the voltage tracking error, so that the error naturally approaches zero in steady state. By utilizing the characteristic that the voltage tracking error naturally returns to zero in steady state and reflects the direction and magnitude of the deviation in transient state, the automatic distinction between transient and steady state is realized, providing a signal basis for soft gating and reward function design. Furthermore, voltage tracking error provides deviation direction and amplitude information for the core dimension of the observation space, and also controls the on / off state of the reinforcement learning agent's output using the gate signal of soft gating. At the same time, the sign of voltage tracking error naturally indicates the correct compensation direction, thus achieving simplicity of system architecture and consistency of signal links. Soft gating is used as a physical hard constraint to force the RL output to decay to near zero during the steady state; the steady-state zero output forced term in the four-component reward function is used as a complementary soft constraint to guide the reinforcement learning agent to learn the zero output policy; the two together ensure that the droop control power allocation function is not disturbed. By using a four-component reward function, the transient error penalty and the steady-state zero output forced by gating weights are coordinated. The directional guiding term accelerates the convergence of the strategy, and the action change rate penalty term suppresses bang-bang oscillations. Through the voltage compensation architecture, only the output voltage of the droop control is compensated, rather than the control output. The complete inner loop controller structure is retained, and it can be deployed as an add-on module without modifying the controller parameters. In summary, this invention can achieve the complementary advantages of reinforcement learning agents and droop control. Droop control ensures steady-state power distribution, while reinforcement learning agents optimize the dynamic trajectory of transient voltage. The two work together through a soft gating mechanism without changing the steady-state operating point.

[0020] It should be understood that the description in the Summary of the Invention is not intended to limit the key or essential features of the embodiments of the present invention, nor is it intended to restrict the scope of the invention. Other features of the invention will become readily apparent from the following description. Attached Figure Description

[0021] The above and other features, advantages, and aspects of the various embodiments of the present invention will become more apparent from the accompanying drawings and the following detailed description. The drawings are provided for a better understanding of the invention and are not intended to limit the invention. In the drawings, the same or similar reference numerals denote the same or similar elements, wherein: Figure 1 A flowchart of the DC microgrid voltage dynamic trajectory optimization method based on reinforcement learning provided in an embodiment of the present invention is shown. Figure 2 This diagram illustrates the architecture of the DC microgrid voltage dynamic trajectory optimization method based on reinforcement learning provided in an embodiment of the present invention. Figure 3The diagram shows the operating characteristic curves of the soft gating mechanism provided in the embodiments of the present invention. Figure 4 A flowchart illustrating the calculation of the four-component reward function provided in an embodiment of the present invention is shown. Figure 5 A schematic diagram of the network structure of the reinforcement learning agent provided in an embodiment of the present invention is shown; Figure 6 A flowchart illustrating the training process of the method provided in an embodiment of the present invention is shown; Figure 7 The figure shows a comparison of the improvement effect of the method of the present invention under the first load condition provided by the embodiment of the present invention; Figure 8 The figure shows a comparison of the improvement effect of the method of the present invention under the second load condition provided by the embodiment of the present invention; Figure 9 The figure shows a comparison of the improvement effect of the method of the present invention under the third load condition provided by the embodiment of the present invention; Figure 10 The diagram shows a block diagram of a DC microgrid voltage dynamic trajectory optimization system based on reinforcement learning provided in an embodiment of the present invention. Figure 11 A block diagram of an electronic device according to an embodiment of the present invention is shown. Detailed Implementation

[0022] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0023] This invention provides a method for optimizing the dynamic voltage trajectory of a DC microgrid based on reinforcement learning. (See also...) Figures 1 to 9 This includes the following steps: S1. Real-time acquisition of the bus voltage of the DC microgrid. and output current And in conjunction with the rated voltage of the DC microgrid The droop coefficient k of the droop controller is used to calculate the voltage tracking error. ; For details on calculating voltage tracking error, please refer to the following formula: ; The voltage tracking error e is obtained by subtracting the droop factor and the output current from the rated voltage and then comparing it with the bus voltage. The voltage tracking error e has the following key characteristics: (i) In steady state, the voltage tracking error e approaches zero because droop control ensures... Approaching the output voltage of the droop controller, i.e. ; (ii) When loading, the voltage tracking error e>0 (voltage is lower than the reference value), indicating that positive compensation is required; (iii) When the load is reduced, the voltage tracking error e < 0 (voltage is higher than the reference value), indicating that negative compensation is required; (iv) The sign of the voltage tracking error e naturally corresponds to the correct compensation direction, that is, e and the compensation voltage. They should have the same symbols; (v) The absolute value of voltage tracking error |e| directly reflects the amplitude of transient deviation.

[0024] The voltage tracking error e serves as the core dimension of the three-dimensional observation signal, providing information on the direction and magnitude of the deviation. On the other hand, it acts as the gating signal for subsequent soft gating (based on |e|), controlling the on / off state of the reinforcement learning agent's output. Furthermore, the sign of the voltage tracking error e naturally indicates the correct compensation direction. The aforementioned design ensures the consistency of the signal chain and the simplicity of the system architecture.

[0025] S2. Calculate the rate of change of the voltage tracking error. Based on the voltage tracking error rate of change and output current Constructing three-dimensional observation signals After normalization, the voltage is input into the reinforcement learning agent, which is used to output a compensation voltage. ; In this invention, the normalization process uses the rescale-symmetric method to map the three-dimensional observation signal to the interval [-1, 1]; the normalization boundary is specifically calibrated by multiplying the P99.5 percentile by 1.2. In the three-dimensional observation signal, the first dimension is the voltage tracking error e, which provides information on the direction and magnitude of the deviation; the second dimension is the rate of change of the voltage tracking error. When the load changes abruptly, |e| initially changes, but Mutations will occur accordingly, allowing for advance prediction and action execution, given a fixed value for e. If it is negative, it means that the voltage is dropping rapidly and positive compensation is needed; A positive value indicates a rapid voltage recovery, requiring reduced compensation or reverse braking; the third dimension is the output current. It reflects the load status and helps the agent distinguish different working points.

[0026] The reinforcement learning agent uses the TD3 algorithm, specifically including an Actor network and a dual Critic network. The action space is one-dimensional, and the output range is... ∈[-10V, +10V].

[0027] Among them, see Figure 5 The Actor network specifically includes an input layer, a fully connected layer, and an output layer. The fully connected layer specifically includes two fully connected network layers with an input dimension of 3, using the rescale-symmetric normalization algorithm. Each layer is followed by a ReLU activation function. The output layer uses a fully connected network and a tanh activation function, with an output range of [-1V, +1V]. The output is then mapped to [-10V, +10V] by a scaling layer. For example, the scaling layer is scalingLayer(Scale=10, Bias=0). Both Critic networks employ a dual-path structure, namely a state path and an action path. The two paths pass through fully connected layers and are then merged in hidden layers. The minimum value of the two outputs is taken as the final output Q-value estimate. The Q-value estimate is used to update the parameters of the actor network, as shown in the following formula. ; in, Indicates the discount factor. For the subsequent comprehensive reward signal, y represents the target label, and the parameters of the actor network are updated to make the Q-value estimate close to the target label.

[0028] The two Critic networks have the same structure but independent parameters, which solves the problem of Q-value overestimation and greatly improves the stability of network training in DC microgrid control scenarios.

[0029] S3, based on the voltage tracking error The gate weight w is calculated using the soft gate function, and then adjusted according to the compensation voltage. The post-gating compensation voltage is calculated using the gating weight w. ; Specifically, based on the voltage tracking error The gating weight w is calculated using the Sigmoid soft-gating function; see the following formula for the Sigmoid soft-gating function: ; in, Indicates the center point threshold. This represents the switching parameters. The center point threshold is the center point of the voltage error boundary between transient and steady state. It needs to be selected in conjunction with the system voltage noise level and the allowable steady-state error, and this value should not be less than the voltage noise level to avoid fluctuations in the gating weights under steady state. If this value is too large, the RL agent will only compensate when the voltage deviation is large. The switching parameters are used to control the steepness of the transition interval of the Sigmoid soft-gating function. The larger the value, the smoother the transition between steady state and transient state; the smaller the value, the steeper the transition.

[0030] Gated compensation voltage For the detailed calculation process, please refer to the following formula: ; Referring to Table 1, for a rated bus voltage of 375V, the soft gating mechanism can take the center point threshold. =2, switch parameters =0.5.

[0031] The gating weights are obtained using the Sigmoid soft-gating function and with the magnitude of the voltage tracking error as input; When the DC microgrid is in a steady state and the voltage tracking error amplitude is less than the first preset threshold, the gating weight tends to be 0. When the DC microgrid is in a transient state and the voltage tracking error amplitude is greater than the second preset threshold, the gating weight tends to 1, and the second preset threshold is greater than the first preset threshold. In practical applications, the first and second preset thresholds can be selected based on factors such as voltage noise level, rated bus voltage, and soft gating parameters. For example, the first preset threshold needs to be greater than the voltage noise level to avoid noise disturbances causing gating weight jitter, and the second preset threshold needs to be greater than the first preset threshold to ensure the distinction between steady-state and transient states.

[0032] Based on the aforementioned parameter settings, when the DC microgrid is in steady state, the voltage tracking error e approaches zero, i.e., |e| is less than the first preset threshold, such as 0.5V. The gating weight w approaches zero, such as w≈0-0.12, and the compensation voltage after gating... Approaching zero, this ensures that the reinforcement learning agent does not interfere with the power allocation function of droop control during the steady-state period; for example, see... Figure 3 In the aforementioned =2, When |e|=2V, w=0.5; When the DC microgrid is in a transient state, the voltage tracking error e increases, that is, |e| is greater than the second preset threshold. For example, when it is 3V, the gating weight approaches 1, such as w≈0.88-1.0. After gating, the compensation voltage is approximately equal to the compensation voltage, so that the reinforcement learning agent provides voltage compensation during the transient period.

[0033] Table 1. Characteristics of Soft Gating

[0034] The gating mechanism and the subsequent four-component reward function constitute a dual guarantee mechanism: At the control signal level, soft gating, as a physical hard constraint, ensures that even if the reinforcement learning agent outputs a non-zero value during the steady state, it will be attenuated to near zero after passing through soft gating, thus fundamentally ensuring that the droop control power distribution function is not disturbed. At the training level, the steady-state zero-output coercion in the reward function serves as a complementary soft constraint, guiding the agent to actively learn the zero-output policy and reducing unnecessary non-zero exploration. This dual guarantee (the combination of physical hard constraints and training soft constraints) ensures the robustness of steady-state behavior, while the smoothness of the Sigmoid function avoids control quantity jumps caused by hard switching.

[0035] S4, the gated compensation voltage With the output voltage of the droop controller After superposition The input is to the inner-loop PI controller, and the output is to the DC-DC converter to execute the control, thereby realizing dynamic control of the voltage; The gated compensation voltage is only superimposed on the output voltage of the droop controller, rather than being used for direct control output. This preserves the complete inner-loop PI control structure and enables plug-and-play modular deployment. In this invention, the inner-loop controller can be a DC-DC converter.

[0036] Among them, based on the voltage tracking error e and the compensation voltage The historical value of the gated compensation voltage is used to calculate the comprehensive reward signal through the constructed four-component reward function, and the comprehensive reward signal is used to train the reinforcement learning agent.

[0037] The four-component reward function includes a transient error penalty term, a steady-state zero output mandatory term, a directional guidance term, and a rate of change penalty term, which are controlled by a gating weight w.

[0038] Among them, transient error penalty term This can be expressed by the following formula: ; in, The transient improvement weights, modulated by the gate weights w, are fully effective in the transient state, driving the reinforcement learning agent to minimize |e|, and are suppressed in the steady state; if If the value is too large, it can easily cause oscillations in the output of the reinforcement learning agent; if the value is too small, the improvement effect on transient voltage will be weak.

[0039] Steady-state zero output forced term This can be expressed by the following formula: , in, As a second punishment, This represents the steady-state quadratic penalty coefficient. The steady-state penalty coefficient, which forces the reinforcement learning agent's output to approach zero, is also modulated by the gating weight w, resulting in full activation in steady state. If... and If the value is too small, the reinforcement learning agent will still output a large compensation voltage in steady state after training; if the value is too large, it will suppress the compensation voltage output by the reinforcement learning agent in transient state, weakening the transient optimization capability.

[0040] Directional guiding items This can be expressed by the following formula: ; in, Indicates directional guiding weight. The sign function is indicated by *, which indicates multiplication. It only takes effect during the transient period and guides the reinforcement learning agent to align its output direction with the voltage tracking error direction. If the directional guidance weight is too small, the directional guidance effect will be weak. If the weight is too large, it will suppress the effect of the transient error penalty term, and the voltage dynamic optimization effect will decrease.

[0041] Action change rate penalty item This can be expressed by the following formula: ; in, This represents the penalty coefficient for the rate of change of action. Delay steps The pre-step compensation voltage is the historical value of the post-gating compensation voltage. The action change rate penalty term is used to suppress potential bang-bang oscillations in the output of the reinforcement learning agent, ensuring the smoothness of the control quantity. The larger the action change rate penalty coefficient, the smoother the control output, but too large a value will limit the transient response speed of the agent, while too small a value will fail to suppress the oscillations in the output.

[0042] Final reward, i.e., comprehensive reward signal It is calculated using the following formula: .

[0043] See Figure 6 The training process includes: Step 1: Only when the DC microgrid has reached a steady state, save the integrator state in the DC microgrid as a steady-state snapshot.

[0044] In the droop controller of the DC microgrid, the inner-loop PI controller includes an integrator. A steady-state snapshot is retained as the initial state for subsequent iterations, preventing invalid data generated during the initial training phase due to voltage ramp-ups from contaminating the experience replay pool.

[0045] Step 2: Based on the steady-state snapshot, obtain normalized three-dimensional observation signal samples under random load disturbance; Step 3: Input the three-dimensional observation signal sample into the reinforcement learning agent for training, and determine whether the four-component reward function has converged. If it has, the trained reinforcement learning agent is obtained; otherwise, load the steady-state snapshot as the initial state and proceed to step 2 until convergence.

[0046] For example, the parameter settings are shown in Table 2. In actual engineering, each parameter can be determined through iterative debugging in simulation based on the converter's power level, voltage level, etc.

[0047] Table 2 Parameter Settings

[0048] Through the above steps, the inner-loop PI controller outputs the final control command, which includes the duty cycle, to control the inner-loop PI controller and achieve dynamic adjustment of the DC microgrid bus voltage. Specifically, by adjusting the switching state of the inner-loop PI controller according to the duty cycle, the DC microgrid bus voltage can quickly stabilize on the desired droop characteristic curve when the load changes abruptly, thereby optimizing the voltage dynamic trajectory.

[0049] In one specific embodiment, the present invention has undergone relevant verification: A DC-DC Buck-Boost converter system with a constant power load (CPL) was built on the MATLAB / Simulink platform, with a rated bus voltage of... = 375V, the method of the present invention is applied to the aforementioned converter system.

[0050] The dual-path structure of the dual-Critic network is shown in the following settings: The state path is obs(3)→FC(64) and ReLU function→FC(64); where obs(3) represents the three-dimensional observation signal and is the input layer, and FC(64) represents a fully connected layer with 64 neurons; the action path is act(1)→FC(64), where act(1) represents the compensation voltage output by the Actor network. .

[0051] After the two paths mentioned above are added together in the merging layer, they pass through the ReLU function, FC(64), ReLU function, and FC(1) in sequence, and finally output the Q value estimate. FC(1) represents a fully connected layer of one neuron.

[0052] The following training steps are followed: Steady-state snapshot generation phase: Disable reinforcement learning agent, set load to zero, run simulation for 0.5s to bring voltage to rated value of 375V, and then save integrator state. This serves as a steady-state snapshot and verifies whether the voltage is close to 375V. The steady-state snapshot ensures that 100% of the data in the experience replay pool is valid, avoiding invalid data generated during the voltage establishment process when training from scratch, and significantly improving training efficiency.

[0053] Observation space normalization calibration stage: Based on steady-state snapshots, three-dimensional observation signal samples are acquired under random load disturbances. The 99.5 percentile values ​​are calculated for each sample and multiplied by a safety margin of 1.2 to determine the upper and lower bounds of normalization. The three-dimensional observation signals are mapped to the [-1, 1] interval using the following formula to achieve a coverage greater than 99%. ; in, Represents a three-dimensional observation signal. This represents the normalized three-dimensional observation signal.

[0054] Training round phase: At the beginning of each training round, load... As the initial state, the load conditions were then randomized. The total training time was 0.5s. The base load was applied at t=0s; a sudden load reduction was applied at t≈0.25s, and the reinforcement learning agent compensated for the transient state to offset the negative impact of the sudden load reduction. The range settings of each parameter are shown in Table 3.

[0055] Table 3 Load Randomization Parameters

[0056] Online deployment and control phase: The parameters of the reinforcement learning agent after training convergence are fixed, and it is deployed as an online controller to the DC-DC Buck-Boost converter system; Real-time acquisition of bus voltage and output current Calculate the voltage tracking error e and its rate of change. and along with the output current The input voltage is normalized using calibrated normalization parameters and then fed into the Actor network; the Actor network outputs a compensation voltage after forward inference. The gate weight w is calculated using a soft gate function and then output. The voltage is superimposed on the output voltage of the droop controller and then input to the inner loop controller to output the final control command.

[0057] After training, the training hyperparameter settings are shown in Table 4.

[0058] Table 4 Hyperparameter Settings

[0059] Reference Figures 7 to 9 Three different load step conditions were tested.

[0060] Under all three load conditions, the under-adjustment and over-adjustment of the bus voltage after load reduction were significantly reduced after compensation by the reinforcement learning agent (RL).

[0061] During the transition to steady state after a load disturbance, the bus voltage, after compensation by a reinforcement learning agent, showed slight fluctuations compared to the case with only droop control, because it entered a steady state. Figure 3 The transition zone shown: The compensation output of the reinforcement learning agent gradually decreases, while the gating weight w decreases synchronously, with both decaying together. This indicates that the soft gating mechanism achieves a smooth transition from transient to steady state, rather than a sudden cutoff, thus avoiding the control quantity jump caused by hard switching.

[0062] It is worth noting that the steady-state value of the bus voltage after compensation by the reinforcement learning agent is completely consistent with the steady-state voltage of the droop control alone, which verifies that the method of the present invention optimizes the voltage dynamic trajectory without changing the steady-state operating point, and proves the effectiveness of the dual protection mechanism: the soft gating mechanism forces the reinforcement learning agent to return to zero output during the steady-state period at the physical level, and the reward function guides the reinforcement learning agent to learn the zero-output strategy at the training level. The two together ensure that the droop control power distribution function is not disturbed.

[0063] Clearly, the method of the present invention has the following advantages, as verified by the present invention: (1) It significantly reduced voltage overshoot and undershoot during load changes and optimized the voltage dynamic trajectory; (2) The steady-state voltage is unaffected, the steady-state operating point is not changed, and the power distribution function of the droop control is not interfered with. (3) The transition from transient to steady state is smooth, with no jumps in control quantities; (4) The output voltage compensation architecture retains the complete inner loop control structure and realizes plug-and-play additional modular deployment.

[0064] Based on the above technical solution, the following beneficial effects are achieved: Existing reinforcement learning-based compensation methods are mostly designed for constant voltage control and are not suitable for droop control systems where steady-state voltage varies with load current. Furthermore, the reinforcement learning agent in existing methods still outputs non-zero control values ​​during the steady-state period, deviating from the steady-state operating point and causing power distribution errors in multi-converter parallel systems.

[0065] The method of this invention fundamentally solves the problem of droop control of reinforcement learning agents in the steady state by designing voltage tracking error (naturally returning to zero in steady state) and using a dual guarantee mechanism (soft gating hard constraints and reward function soft constraints). At the same time, it achieves voltage dynamic trajectory optimization without changing the steady-state operating point through an output voltage compensation architecture.

[0066] Specifically, by designing the voltage tracking error, and utilizing its characteristics of naturally returning to zero in steady state and reflecting the direction and magnitude of the deviation in transient state, the automatic distinction between transient and steady state is realized, providing a signal basis for soft gating and reward function design; Furthermore, voltage tracking error provides deviation direction and amplitude information for the core dimension of the observation space, and also controls the on / off state of the reinforcement learning agent's output using the gate signal of soft gating. At the same time, the sign of voltage tracking error naturally indicates the correct compensation direction, thus achieving simplicity of system architecture and consistency of signal links. Soft gating is used as a physical hard constraint to force the RL output to decay to near zero during the steady state; the steady-state zero output forced term in the four-component reward function is used as a complementary soft constraint to guide the reinforcement learning agent to learn the zero output policy; the two together ensure that the droop control power allocation function is not disturbed. By using a four-component reward function, the transient error penalty and the steady-state zero output forced by gating weights are coordinated. The directional guiding term accelerates the convergence of the strategy, and the action change rate penalty term suppresses bang-bang oscillations. Through the output voltage compensation architecture, only the output voltage of the droop control is compensated, rather than the control output. The complete inner loop controller structure is retained, and it can be deployed as an add-on module without modifying the controller parameters. In summary, this invention can achieve the complementary advantages of reinforcement learning agents and droop control. Droop control ensures steady-state power distribution, while reinforcement learning agents optimize the dynamic trajectory of transient voltage. The two work together through a soft gating mechanism without changing the steady-state operating point.

[0067] It should be noted that, for the sake of simplicity, the foregoing method embodiments are all described as a series of actions. However, those skilled in the art should understand that the present invention is not limited to the described order of actions, because according to the present invention, some steps can be performed in other orders or simultaneously. Furthermore, those skilled in the art should also understand that the embodiments described in the specification are all optional embodiments, and the actions and modules involved are not necessarily essential to the present invention.

[0068] The acquisition, storage, and application of user personal information involved in the technical solution of this invention all comply with the provisions of relevant laws and regulations and do not violate public order and good morals.

[0069] This invention also provides a reinforcement learning-based DC microgrid voltage dynamic trajectory optimization system 1000, see [link to relevant documentation]. Figure 10 ,include: The voltage tracking error calculation module 1010 is used to collect the bus voltage and output current of the DC microgrid in real time, and calculate the voltage tracking error by combining the rated voltage of the DC microgrid and the droop coefficient of the droop controller. The reinforcement learning agent module 1020 is connected to the voltage tracking error calculation module 1010. It is used to calculate the rate of change of the voltage tracking error, construct a three-dimensional observation signal based on the voltage tracking error, the rate of change and the output current, and input the signal into the reinforcement learning agent after normalization. The reinforcement learning agent is used to output the compensation voltage. The soft gating module 1030 is connected to the reinforcement learning agent module 1020 and is used to calculate the gating weight based on the voltage tracking error and the soft gating function, and to calculate the post-gating compensation voltage based on the compensation voltage and the gating weight. The output voltage compensation module 1040 is connected to the soft gating module 1030 and is used to superimpose the gating compensation voltage and the output voltage of the droop controller and input them into the inner loop PI controller, and output the final control command to the inner loop controller for control execution. The reward calculation and training module 1050 is connected to the voltage tracking error calculation module 1010, the reinforcement learning agent module 1020, the soft gating module 1030, and the output voltage compensation module 1040. It is used to calculate the comprehensive reward signal based on the voltage tracking error, the gating compensation voltage, and the historical value of the gating compensation voltage through the constructed four-component reward function, and to train the reinforcement learning agent using the comprehensive reward signal.

[0070] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working process of the described module can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here.

[0071] Figure 11 A schematic block diagram of an electronic device 1100 that can be used to implement embodiments of the present invention is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the invention described and / or claimed herein.

[0072] Electronic device 1100 includes a computing unit 1101, which can perform various appropriate actions and processes according to a computer program stored in ROM 1102 or a computer program loaded into RAM 1103 from storage unit 1108. RAM 1103 may also store various programs and data required for the operation of electronic device 1100. The computing unit 1101, ROM 1102, and RAM 1103 are interconnected via bus 1104. I / O interface 1105 is also connected to bus 1104.

[0073] Multiple components in electronic device 1100 are connected to I / O interface 1105, including: input unit 1106, such as keyboard, mouse, etc.; output unit 1107, such as various types of displays, speakers, etc.; storage unit 1108, such as disk, optical disk, etc.; and communication unit 1109, such as network card, modem, wireless transceiver, etc. Communication unit 1109 allows electronic device 1100 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.

[0074] The computing unit 1101 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 1101 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 1101 performs the various methods and processes described above. For example, in some embodiments, the methods may be implemented as computer software programs tangibly contained in a machine-readable medium, such as storage unit 1108. In some embodiments, part or all of the computer program may be loaded and / or installed on the electronic device 1100 via ROM 1102 and / or communication unit 1109. When the computer program is loaded into RAM 1103 and executed by the computing unit 1101, one or more steps of the reinforcement learning-based DC microgrid voltage dynamic trajectory optimization method described above may be performed. Alternatively, in other embodiments, computing unit 1101 may be configured to perform the reinforcement learning-based DC microgrid voltage dynamic trajectory optimization method by any other suitable means (e.g., by means of firmware).

[0075] It should be understood that the various forms of processes shown above can be used to reorder, add, or delete steps. For example, the steps described in this invention can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution disclosed in this invention can be achieved, and this is not limited herein.

[0076] The specific embodiments described above do not constitute a limitation on the scope of protection of this invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this invention should be included within the scope of protection of this invention.

Claims

1. A method for optimizing the dynamic voltage trajectory of a DC microgrid based on reinforcement learning, characterized in that, include: The bus voltage and output current of the DC microgrid are collected in real time, and the voltage tracking error is calculated by combining the rated voltage of the DC microgrid and the droop coefficient of the droop controller. The rate of change of the voltage tracking error is calculated, and a three-dimensional observation signal is constructed based on the voltage tracking error, the rate of change, and the output current. After normalization, the signal is input into the reinforcement learning agent, which is used to output the compensation voltage. The gate weight is calculated based on the voltage tracking error and the soft gate function, and the post-gated compensation voltage is calculated based on the compensation voltage and the gate weight. The gated compensation voltage is superimposed with the output voltage of the droop controller and then input into the inner loop PI controller. The final control command is then output to the inner loop controller for execution. Specifically, based on the voltage tracking error, the gating compensation voltage, and the historical values ​​of the gating compensation voltage, a comprehensive reward signal is calculated using a constructed four-component reward function, and the comprehensive reward signal is used to train the reinforcement learning agent.

2. The method for optimizing the dynamic voltage trajectory of a DC microgrid based on reinforcement learning according to claim 1, characterized in that, The voltage tracking error is obtained by subtracting the droop factor and the output current from the rated voltage and then comparing it with the bus voltage. In steady state, the voltage tracking error approaches zero.

3. The method for optimizing the dynamic voltage trajectory of a DC microgrid based on reinforcement learning according to claim 1, characterized in that, The three-dimensional observation signal uses the voltage tracking error as the first dimension, the rate of change of the voltage tracking error as the second dimension, and the output current as the third dimension, mapping the data signals of each dimension to a preset normalization boundary.

4. The method for optimizing the dynamic voltage trajectory of a DC microgrid based on reinforcement learning according to claim 1, characterized in that, The gating weights are obtained using the Sigmoid soft-gating function and with the magnitude of the voltage tracking error as input; When the DC microgrid is in a steady state and the voltage tracking error amplitude is less than a first preset threshold, the gating weight tends to 0. When the DC microgrid is in a transient period and the voltage tracking error amplitude is greater than the second preset threshold, the gating weight tends to 1, and the second preset threshold is greater than the first preset threshold.

5. The method for optimizing the dynamic voltage trajectory of a DC microgrid based on reinforcement learning according to claim 1, characterized in that, The four-component reward function includes a transient error penalty term, a steady-state zero output mandatory term, a directional guidance term, and a rate of change penalty term, all of which are controlled by gating weights. The transient error penalty term is used to drive the agent to minimize the voltage tracking error during the transient period; The steady-state zero-output coercion term is used to guide the reinforcement learning agent to approach zero in the steady-state period; The directional guidance term is used to guide the agent to output a compensation voltage that is consistent with the direction of the voltage tracking error during the transient period; The action change rate penalty term is used to suppress oscillations in the agent's output.

6. The method for optimizing the dynamic voltage trajectory of a DC microgrid based on reinforcement learning according to claim 1, characterized in that, The training process of the reinforcement learning agent includes: Only when the DC microgrid reaches steady state is the state of the integrator in the DC microgrid saved as a steady-state snapshot. Based on the steady-state snapshot, normalized three-dimensional observation signal samples are obtained under random load disturbances; The three-dimensional observation signal samples are input into the reinforcement learning agent for training. It is determined whether the four-component reward function has converged. If it has, the trained reinforcement learning agent is obtained. Otherwise, the steady-state snapshot is loaded as the initial state, and then the process proceeds to step 2 until the four-component reward function converges.

7. A DC microgrid voltage dynamic trajectory optimization system based on reinforcement learning, used to implement the DC microgrid voltage dynamic trajectory optimization method based on reinforcement learning as described in any one of claims 1 to 6, characterized in that, include: The voltage tracking error calculation module is used to collect the bus voltage and output current of the DC microgrid in real time, and calculate the voltage tracking error by combining the rated voltage of the DC microgrid and the droop coefficient of the droop controller. The reinforcement learning agent module is connected to the voltage tracking error calculation module. It is used to calculate the rate of change of the voltage tracking error, construct a three-dimensional observation signal based on the voltage tracking error, the rate of change and the output current, and input the signal into the reinforcement learning agent after normalization. The reinforcement learning agent is used to output the compensation voltage. A soft gating module, connected to the reinforcement learning agent module, is used to calculate the gating weights based on the voltage tracking error and the soft gating function, and to calculate the post-gating compensation voltage based on the compensation voltage and the gating weights. A voltage compensation module, connected to the soft gating module, is used to superimpose the gating compensation voltage with the output voltage of the droop controller and input it into the inner loop PI controller, and output the final control command to the inner loop controller for control execution. The reward calculation and training module is connected to the voltage tracking error calculation module, the reinforcement learning agent module, the soft gating module, and the output voltage compensation module. It is used to calculate the comprehensive reward signal based on the voltage tracking error, the gating compensation voltage, and the historical value of the gating compensation voltage through a constructed four-component reward function, and to train the reinforcement learning agent using the comprehensive reward signal.

8. An electronic device, characterized in that, It includes a memory, a processor, and a computer program; wherein the computer program is stored in the memory and configured to be executed by the processor to implement the reinforcement learning-based DC microgrid voltage dynamic trajectory optimization method as described in any one of claims 1 to 6.

9. A computer-readable storage medium, characterized in that, It stores a computer program; the computer program is executed by a processor to implement the DC microgrid voltage dynamic trajectory optimization method based on reinforcement learning as described in any one of claims 1 to 6.