Method for enhancing stability of power hardware-in-the-loop system based on DDPG-RP

By adopting a hierarchical collaborative control architecture based on DDPG-RP, combined with a phase-compensated repetitive-voltage feedforward composite controller and a deep reinforcement learning controller, the stability and dynamic response problems of the PHIL simulation system are solved, and efficient and stable control under complex power grid fault conditions is achieved.

CN122136980AActive Publication Date: 2026-06-02TIANJIN UNIV

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
TIANJIN UNIV
Filing Date
2026-04-30
Publication Date
2026-06-02

AI Technical Summary

Technical Problem

Traditional PHIL simulation systems suffer from stability issues, including system oscillations caused by delays and bandwidth limitations from interface devices, power fluctuations caused by impedance mismatch between the digital side and the device under test, and simulation accuracy affected by nonlinear loads and harmonic interference. Furthermore, existing control methods are ill-suited for handling complex operating conditions.

Method used

A hierarchical collaborative control architecture based on DDPG-RP is adopted. By using a phase-compensated repetitive-voltage feedforward composite controller and a deep deterministic policy gradient algorithm, bottom-level and upper-level intelligent controllers are constructed. The deep reinforcement learning controller is combined for offline training and online deployment to optimize control parameters and enhance system stability and dynamic response.

Benefits of technology

It effectively solves the problem that traditional control systems cannot balance stability and performance under complex power grid fault conditions, improves the steady-state accuracy and rapid dynamic response capability of power hardware in-loop systems, and significantly enhances dynamic response and stability performance under three-phase short-circuit faults.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122136980A_ABST
    Figure CN122136980A_ABST
Patent Text Reader

Abstract

This invention proposes a method for enhancing the stability of power hardware-in-the-loop systems based on DDPG-RP. First, an equivalent circuit model of the system containing grid fault disturbances is established, its dynamic interaction characteristics are analyzed, and the impedance matching conditions for stable operation are determined. Based on this, a phase-compensated and enhanced repetitive-voltage feedforward composite controller is designed. An upper-level intelligent controller is constructed by combining this with a deep deterministic policy gradient algorithm. These two controllers are deployed in the bottom and upper-level control units, respectively, forming a hierarchical collaborative control architecture for grid fault simulation. A test platform is built based on this architecture. After offline training and online deployment of the deep reinforcement learning controller, a three-phase short-circuit fault condition is simulated using a high-power grid-connected interface device to verify its dynamic response and stability. This invention, relying on hierarchical collaboration and intelligent optimization, effectively solves the problem of traditional control systems struggling to balance stability and performance under complex faults, significantly improving system adaptability and robustness.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of smart grid simulation technology, specifically to a method for enhancing the stability of power hardware-in-the-loop systems based on DDPG-RP. Background Technology

[0002] With the rapid development of smart grid technology, Power Hardware-in-the-Loop (PHIL) simulation technology has become an important means of testing and verifying power system equipment. PHIL simulation builds a power system model in a real-time digital simulation system and connects it to the physical device under test through interface devices to form a closed-loop simulation system. This system can simulate various power grid operating conditions, effectively reducing testing costs and risks.

[0003] However, traditional PHIL simulation systems suffer from stability issues, primarily manifested in: system oscillations caused by delays and bandwidth limitations from interface devices; power fluctuations due to impedance mismatch between the digital side and the device under test; and simulation accuracy affected by nonlinear loads and harmonic interference. Existing control methods mostly employ linear controllers, but their performance degrades when the system deviates from its equilibrium point, making them unsuitable for complex operating conditions.

[0004] Deep reinforcement learning, as a data-driven algorithm, possesses strong environmental adaptability and optimization control capabilities, and has demonstrated superior performance in the field of microgrid control. However, existing research is mostly limited to simulation environments, lacking practical verification on hardware platforms, and has not fully considered the special stability requirements of PHIL systems. Summary of the Invention

[0005] To address the current technological shortcomings, this invention proposes a method for enhancing the stability of power hardware-in-the-loop systems based on DDPG-RP, which is used to solve the stability, accuracy, and dynamic response problems in PHIL simulation systems.

[0006] To achieve the above objectives, the present invention provides the following technical solution: a method for enhancing the stability of a power hardware-in-the-loop system based on DDPG-RP, comprising: Step S1: Establish an equivalent circuit model of the power hardware-in-the-loop system that includes the characteristics of grid fault disturbances, analyze the dynamic interaction characteristics of the power hardware-in-the-loop system under grid fault conditions, and determine the impedance matching conditions for maintaining the stable operation of the power hardware-in-the-loop system. Step S2: Based on the impedance matching condition for stable operation of the power hardware-in-the-loop system, design a phase-compensated and enhanced repetitive-voltage feedforward composite controller. Step S3: Construct the upper-layer intelligent controller using the deep deterministic strategy gradient algorithm; Step S4: Deploy the phase-compensated enhanced repetitive-voltage feedforward composite controller in the bottom control unit, and deploy the upper-level intelligent controller in the upper control unit. Based on the bottom control unit and the upper control unit, a hierarchical collaborative control architecture for grid fault simulation is formed. Step S5: Based on the hierarchical collaborative control architecture, build a power hardware-in-the-loop test platform, and complete the offline training and online deployment of the deep reinforcement learning controller on the power hardware-in-the-loop test platform to obtain the trained deep reinforcement learning controller. Step S6: Simulate a three-phase short-circuit fault condition using a high-power grid-connected interface device, and verify the dynamic response and stability performance under the simulated three-phase short-circuit fault condition using a trained deep reinforcement learning controller.

[0007] Furthermore, the specific process of step S1 is as follows: An equivalent circuit model of the power hardware-in-the-loop system is established based on an ideal transformer model. The digital simulation side of the power hardware-in-the-loop system includes models simulating symmetrical / asymmetrical short circuits and voltage sags / droops in the power grid. The Thevenin equivalent impedance on the digital simulation side is denoted as... The high-power grid-connected interface device and the power electronic device together constitute the measured side; the equivalent impedance of the measured side is denoted as... Establish a hardware-in-the-loop total delay that takes power into account. The open-loop transfer function model is expressed as: ; In the formula, For the open-loop transfer function model of the power hardware-in-the-loop system; It is a complex frequency variable; It is a natural constant; Using the Pad approximation After rationalization approximation, it can be expressed as: ; The open-loop transfer function model after Pad approximation is expressed as: ; ; In the formula, This is the open-loop transfer function model after Pad approximation; Define the impedance ratio function as The open-loop transfer function model after Pad approximation is written as: ,Will Convert to ;right Perform mathematical transformations from the complex frequency domain to the frequency domain. The open-loop frequency response of the power hardware-in-the-loop system is obtained, which is expressed as: , ; In the formula, This is the impedance ratio function of the power hardware in the frequency domain of the loop system. It is the imaginary unit in the field of engineering. Angular frequency; This refers to the open-loop frequency response of the power hardware-in-the-loop system. Based on the Nyquist stability criterion, for The Nyquist curve is analyzed to determine the stability of the power hardware-in-the-loop system; based on the stability of the power hardware-in-the-loop system, [further analysis is needed]. Processing is performed to make By satisfying both phase stability constraints and amplitude stability constraints, impedance matching conditions that ensure stable operation of the power hardware in-loop system are obtained.

[0008] Furthermore, the specific process for obtaining the impedance matching conditions for the stable operation of the power hardware in-loop system is as follows: Based on the Nyquist stability criterion, for The Nyquist curve was analyzed. when The Nyquist curve does not enclose the complex plane. At this point, the stability of the power hardware in-loop system is considered; when The Nyquist curve encloses the complex plane on At this point, the closed-loop power hardware-in-the-loop system is unstable, which affects the power hardware-in-the-loop system. , Total delay Re-identification and adjustment; Based on the stability of the power hardware in-loop system, Processing is performed to make Both phase stability constraints and magnitude stability constraints are satisfied, specifically: When the power hardware in the loop system is in the low-frequency range, by adjusting... The circuit parameters were corrected. Phase curve, avoid The phase lag is -180°, which satisfies the phase stability constraint. When the power hardware in-loop system is in the high-frequency band, by... The amplitude is limited to prevent the Nyquist curve from approaching the critical point, thus satisfying the amplitude stability constraint. when By satisfying both phase stability constraints and amplitude stability constraints, impedance matching conditions that ensure stable operation of the power hardware in-loop system are obtained.

[0009] Furthermore, the specific process of step S2 is as follows: Based on the total delay of the power hardware in the loop system identified in step S1 Design a phase compensator; introduce the phase compensator into the repetitive controller to form a phase-compensated repetitive controller, which includes a repetitive control gain; A voltage feedforward controller is connected in parallel with a phase-compensated repetitive controller to form a phase-compensated repetitive-voltage feedforward composite controller. At the instant a three-phase short-circuit fault occurs in the power grid, the voltage feedforward coefficient is obtained by real-time adjustment of the repetitive controller; The difference between the actual voltage output by the high-power grid-connected interface device and the reference voltage of the power hardware in-loop system is input to the phase-compensated and enhanced repetitive-voltage feedforward composite controller for processing, generating harmonic suppression components and transient support components. The harmonic suppression components and transient support components are linearly superimposed to obtain the modulation wave command of the high-power grid-connected interface device. Based on the modulation wave command, the repetitive control gain and voltage feedforward coefficient are adaptively optimized online through a deep deterministic strategy gradient algorithm to obtain the optimal combination of control parameters in real time, thus forming intelligent parameter risk avoidance. The phase compensator implements a phase stabilization mechanism through hardware circuitry to offset the phase lag caused by the inherent delay of the power hardware in the loop system. The phase stabilization mechanism of the hardware circuitry is called physical structure phase stabilization. An enhanced control logic is formed by organically integrating a two-layer mechanism of physical structural stability and intelligent parameter risk avoidance. The final design of a phase-compensated repetitive-voltage feedforward composite controller based on enhanced control logic was completed.

[0010] Furthermore, the specific process of step S3 is as follows: A deep deterministic policy gradient algorithm is adopted, combined with state space vectors and reward functions to design an upper-level intelligent controller; The state-space vector includes: the output voltage of the high-power grid-connected interface device. , Shaft components, output voltage tracking error of high-power grid-connected interface devices, frequency deviation of power hardware-in-the-loop systems, and characteristic quantities characterizing fault types.

[0011] Furthermore, the specific process of step S4 is as follows: A phase-compensated enhanced repetitive-voltage feedforward composite controller is deployed on the programmable gate array (PGA) of the underlying control unit. The multiphase current output by the high-power grid-connected interface device is controlled through the interface controller on the PGA. With capacitor voltage Perform synchronous sampling; for Mutually Mutually The collection of phase currents; for Mutually Mutually Phase capacitor voltage; right and After performing anti-aliasing filtering and scaling transformation sequentially, a Clarke transformation is executed to obtain a two-phase stationary coordinate system. , Feedback volume , , , ;in, Two-phase stationary coordinate system for feedback shaft current; Two-phase stationary coordinate system for feedback shaft current; Two-phase stationary coordinate system for feedback Shaft capacitor voltage; Two-phase stationary coordinate system for feedback Shaft capacitor voltage; Within each control cycle, the reference voltage command in the two-phase stationary coordinate system is read. , Subtract the corresponding feedback voltage respectively , ,get and ;in, For the current number Under one control cycle, the two-phase stationary coordinate system Shaft voltage error; For the current number Under one control cycle, the two-phase stationary coordinate system Shaft voltage error; For the current number Under one control cycle, the two-phase stationary coordinate system Reference voltage command for the shaft; For the current number Under one control cycle, the two-phase stationary coordinate system Reference voltage command for the shaft; Will and The data is stored in the on-chip circular memory. During the repetitive control calculation, the on-chip circular memory is accessed to obtain the historical voltage error corresponding to the previous fundamental cycle, and combined with the current voltage error. and The output is calculated after phase compensation filter and repetitive control gain calculation. ;in, For the current number The control output quantity obtained by repeated control calculation under each control cycle; Based on reference voltage and voltage feedforward coefficient calculation ; This refers to the control output obtained by voltage feedforward calculation in the current k-th control cycle; This is the reference voltage command for the current k-th control cycle; Will and Superimpose the results to obtain the current number. k Final control voltage command under each control cycle ; based on A modulated wave is generated and compared with multiple set triangular carrier waves to output the switching decision logic command of the power conversion submodule of the high-power grid-connected interface device. Using the high-frequency clock of the programmable gate array, the switching decision logic instruction of the submodule is calibrated for the switching time, and a dead time is inserted to generate a prototype of the driving pulse; The prototype drive pulse is input to the gate drive circuit of each sub-module for amplification and isolation, and the actual multi-level PWM drive signal is output. The upper-level intelligent controller is deployed within the high-performance microprocessor of the upper-level control unit, which operates with millisecond-level control cycles. The lower-level control unit acquires the state space vector of the power hardware-in-the-loop system and uploads it to the upper-level control unit. The upper-level control unit processes the state space vector and optimizes the repetitive control gain and voltage feedforward coefficient based on a deep deterministic policy gradient algorithm to obtain optimized dynamic control parameters. , ; The optimized repetitive control gain; The optimized voltage feedforward coefficient; Optimized dynamic control parameters , The data is sent to the lower-level control unit, which then uses the optimized dynamic control parameters. , Regenerate the multi-level PWM drive signal to complete the interaction mechanism between the upper-level control unit and the lower-level control unit; Through the collaborative deployment and interaction mechanism between the phase-compensated enhanced repetitive-voltage feedforward composite controller and the upper-level intelligent controller, a hierarchical collaborative control architecture for power grid fault simulation is formed.

[0012] Furthermore, the specific process of step S5 is as follows: Based on the hierarchical collaborative control architecture for power grid fault simulation, a hardware-in-the-loop test platform for the high-power grid-connected interface device was built. Based on the hardware-in-the-loop test platform, the deep reinforcement learning controller is trained offline to obtain the trained deep reinforcement learning controller, specifically as follows: A digital simulation environment consistent with the actual high-power grid-connected interface device is established based on the hardware-in-the-loop test platform. In the digital simulation environment, an interaction mechanism between the deep reinforcement learning controller and the digital simulation environment is constructed; the output voltage of the high-power grid-connected interface device is used. , The axial component, the output voltage tracking error of the high-power grid-connected interface device, the frequency deviation of the power hardware-in-the-loop system, and the characteristic quantities characterizing the fault type are used as status inputs. Using the repetitive control gain and voltage feedforward coefficient in the phase-compensated repetitive-voltage feedforward composite controller as the continuous action output is the control strategy of the deep reinforcement learning controller. Within each simulation time step of the interactive mechanism, the repetitive control gain and voltage feedforward coefficient of the phase-compensated repetitive-voltage feedforward composite controller are updated based on the continuous action output. The updated repetitive control gain and voltage feedforward coefficient drive the state evolution of the simulation system. At the same time, a reward function is constructed based on the state input. The effect of the control strategy on driving the state evolution of the simulation system is quantitatively evaluated to obtain the reward value, and the next state of the power hardware-in-the-loop system is calculated. The current state of the power hardware-in-the-loop system, the continuous action output, the reward value, and the next state constitute the interactive data and are stored in the experience playback pool. Based on the experience replay mechanism, historical interaction data is randomly extracted from the experience replay pool. Based on the historical interaction data, the evaluation network of the deep reinforcement learning controller is optimized by minimizing the value function, and then the action network of the deep reinforcement learning controller is optimized by the policy gradient method. After optimizing the evaluation network and action network of the deep reinforcement learning controller, the iterative training process of the control policy begins: During the iterative training of the control strategy, changes in grid parameters, perturbations of interface device parameters, and measurement noise are introduced to repeatedly train the control strategy. The changes in power grid parameters are set based on the impedance fluctuation range under typical operating conditions of the power system, the perturbation of interface device parameters is set based on the manufacturing tolerance range of power electronic equipment, and the measurement noise is set based on the measurement accuracy level of industrial-grade sensors. When the reward function converges during training and the power hardware-in-the-loop system exhibits good stability and control performance under typical three-phase short-circuit fault conditions, the training of the deep reinforcement learning controller is completed, the trained deep reinforcement learning controller is obtained, and the trained deep reinforcement learning controller is deployed to the high-power grid-connected interface device.

[0013] Furthermore, the specific process of step S6 is as follows: A high-power grid-connected interface device is used to simulate a three-phase short-circuit fault condition. The three-phase short-circuit fault condition is accompanied by voltage drop, instantaneous current surge, and frequency fluctuation of the power hardware in-loop system. During the fault process, the upper-level control unit calls the trained deep reinforcement learning controller to process the voltage drop, instantaneous current surge, and frequency fluctuation of the power hardware-in-the-loop system in real time, and outputs dynamic control parameters. The lower-level control unit generates multi-level PWM drive signals based on the dynamic control parameters, and drives the high-power grid-connected interface device through the multi-level PWM drive signals to verify the effectiveness of the dynamic control parameters and the dynamic stability of the waveform. The effectiveness of the dynamic control parameters and the dynamic stability of the waveform together support the dynamic response and stability performance of the power hardware-in-the-loop system.

[0014] A computer storage medium storing a computer program, which, when run, implements a method for enhancing the stability of a power hardware-in-the-loop system based on DDPG-RP.

[0015] An electronic device includes a processor, a memory, and a bus, wherein the processor and the memory are connected via the bus, wherein the memory is used to store a set of program code, and the processor is used to call the program code stored in the memory to execute a power hardware-in-the-loop system stability enhancement method based on DDPG-RP.

[0016] Compared with existing technologies, the present invention has the following advantages: (1) By constructing a hierarchical collaborative control architecture, the present invention adopts a phase-compensated enhanced repetitive-voltage feedforward composite controller in the bottom control unit to ensure the steady-state accuracy and fast dynamic response of the high-power grid-connected interface device, and adopts an upper-level intelligent controller in the upper control unit to realize the power hardware in the loop system-level adaptive optimization, which effectively solves the problem that traditional control is difficult to balance stability and performance under complex grid fault conditions.

[0017] (2) This invention introduces a phase compensator into the internal model of the repetitive controller to correct the phase lag of the power hardware in-loop system, thereby effectively suppressing unfavorable dynamic characteristics and providing reliable underlying control support for the stable operation of the power hardware in-loop system. Attached Figure Description

[0018] Figure 1 This is a flowchart of the method of the present invention.

[0019] Figure 2 This is a comparison diagram of the system frequency response of the method proposed in this invention and the benchmark control method under a three-phase short-circuit fault.

[0020] Figure 3 The graphs show the convergence curves of the reward function during the training process of the DDPG-RP controller of the present invention in different scenarios. Detailed Implementation

[0021] like Figure 1 As shown, the present invention provides a technical solution: a method for enhancing the stability of a power hardware-in-the-loop system based on DDPG-RP, comprising: Step S1: Establish an equivalent circuit model of the power hardware-in-the-loop system that includes the characteristics of grid fault disturbances, analyze the dynamic interaction characteristics of the power hardware-in-the-loop system under grid fault conditions, and determine the impedance matching conditions for maintaining the stable operation of the power hardware-in-the-loop system. Step S2: Based on the impedance matching condition for stable operation of the power hardware-in-the-loop system, design a phase-compensated and enhanced repetitive-voltage feedforward composite controller. Step S3: Construct the upper-layer intelligent controller using the deep deterministic strategy gradient algorithm; Step S4: Deploy the phase-compensated enhanced repetitive-voltage feedforward composite controller in the bottom control unit, and deploy the upper-level intelligent controller in the upper control unit. Based on the bottom control unit and the upper control unit, a hierarchical collaborative control architecture for grid fault simulation is formed. Step S5: Based on the hierarchical collaborative control architecture, build a power hardware-in-the-loop test platform, and complete the offline training and online deployment of the deep reinforcement learning controller on the power hardware-in-the-loop test platform to obtain the trained deep reinforcement learning controller. Step S6: Simulate a three-phase short-circuit fault condition using a high-power grid-connected interface device, and verify the dynamic response and stability performance under the simulated three-phase short-circuit fault condition using a trained deep reinforcement learning controller.

[0022] The specific process of step S1 is as follows: An equivalent circuit model of a Power Hardware-in-the-Loop (PHIL) system is established based on an ideal transformer model. The digital simulation side of the PHIL system includes a grid fault model capable of simulating symmetrical / asymmetrical short circuits and voltage sags / droops. The Thevenin equivalent impedance on the digital simulation side is denoted as... The high-power grid-connected interface device and the power electronic device together constitute the measured side; the equivalent impedance of the measured side is denoted as... Establish a hardware-in-the-loop total delay that takes power into account. The open-loop transfer function model is expressed as: ; In the formula, For the open-loop transfer function model of the power hardware-in-the-loop system; It is a complex frequency variable; It is a natural constant; Using the Pad approximation After rationalization approximation, it can be expressed as: ; The open-loop transfer function model after Pad approximation is expressed as: ; ; In the formula, This is the open-loop transfer function model after Pad approximation; Define the impedance ratio function as The open-loop transfer function model after Pad approximation is written as: ,Will Convert to ;right Perform mathematical transformations from the complex frequency domain to the frequency domain. The open-loop frequency response of the power hardware-in-the-loop system is obtained, which is expressed as: , ; In the formula, This is the impedance ratio function of the power hardware in the frequency domain of the loop system. It is the imaginary unit in the field of engineering. Angular frequency; This refers to the open-loop frequency response of the power hardware-in-the-loop system. Based on the Nyquist stability criterion, for The Nyquist curve is analyzed to determine the stability of the power hardware-in-the-loop system; based on the stability of the power hardware-in-the-loop system, [further analysis is needed]. Processing is performed to make By satisfying both phase stability constraints and amplitude stability constraints, impedance matching conditions that ensure stable operation of the power hardware in-loop system are obtained.

[0023] The specific process for obtaining the impedance matching conditions for stable operation of the power hardware in the loop system is as follows: Based on the Nyquist stability criterion, for The Nyquist curve was analyzed. when The Nyquist curve does not enclose the complex plane. At this point, the stability of the power hardware in-loop system is considered; when The Nyquist curve encloses the complex plane on At this point, the closed-loop power hardware-in-the-loop system is unstable, which affects the power hardware-in-the-loop system. , Total delay Re-identification and adjustment; Based on the stability of the power hardware in-loop system, Processing is performed to make Both phase stability constraints and magnitude stability constraints are satisfied, specifically: When the power hardware in the loop system is in the low frequency range Internally, through adjustment The circuit parameters were corrected. Phase curve, avoid The phase lag is close to -180°, which satisfies the phase stability constraint. When the power hardware in-loop system is in the high-frequency band, by... The amplitude is limited to prevent the Nyquist curve from approaching the critical point, thus satisfying the amplitude stability constraint. when By satisfying both phase stability constraints and amplitude stability constraints, impedance matching conditions that ensure stable operation of the power hardware in-loop system are obtained.

[0024] The specific process of step S2 is as follows: Based on the total delay of the power hardware in the loop system identified in step S1 Design a phase compensator Phase compensator By introducing a repetitive controller, a phase-compensated repetitive controller is formed. The transfer function of the phase-compensated repetitive controller is expressed as: ; In the formula, The transfer function for a phase-compensated repetitive controller; To control the gain repeatedly; Forgetting factor; A voltage feedforward controller is connected in parallel with a phase-compensated repetitive controller to form a phase-compensated repetitive-voltage feedforward composite controller. At the instant a three-phase short-circuit fault occurs in the power grid, the voltage feedforward coefficient is obtained by real-time tuning of the repetitive controller. ; The difference between the actual voltage output by the high-power grid-connected interface device and the reference voltage of the power hardware in-loop system is input to the phase-compensated and enhanced repetitive-voltage feedforward composite controller for processing, generating harmonic suppression components and transient support components. The harmonic suppression components and transient support components are linearly superimposed to obtain the modulation wave command of the high-power grid-connected interface device. Based on the modulation wave command, the repetitive control gain is adjusted using a deep deterministic policy gradient algorithm. and voltage feedforward coefficient Online adaptive optimization is performed to obtain the optimal combination of control parameters in real time, forming intelligent parameter risk avoidance; Phase compensator C ( s A phase stabilization mechanism is implemented through hardware circuitry to offset the phase lag caused by the inherent delay of the power hardware in the loop system. The phase stabilization mechanism of the hardware circuitry is called physical structure phase stabilization. An enhanced control logic is formed by organically integrating a two-layer mechanism of physical structural stability and intelligent parameter risk avoidance. The final design of a phase-compensated repetitive-voltage feedforward composite controller based on enhanced control logic was completed.

[0025] The specific process of step S3 is as follows: A deep deterministic policy gradient algorithm is adopted, combined with state space vectors and reward functions to design an upper-level intelligent controller; state space vector Including: the output voltage of high-power grid-connected interface devices , Axial components and High-power grid-connected interface device output voltage tracking error and Power hardware-in-the-loop system frequency deviation and the characteristic quantities that characterize the type of fault. F c ; For the output voltage of high-power grid-connected interface devices Axial components; For the output voltage of high-power grid-connected interface devices Axial components; For the output voltage tracking error of high-power grid-connected interface devices Axial components; For the output voltage tracking error of high-power grid-connected interface devices Axial components; reward function ,express: ; In the formula, This represents the weighting coefficient for the voltage tracking error; This is the weighting factor for the frequency deviation of the power hardware in the loop system; This is the weighting factor for total harmonic distortion; This is for total harmonic distortion.

[0026] The specific process of step S4 is as follows: A phase-compensated enhanced repetitive-voltage feedforward composite controller is deployed on the programmable gate array (PGA) of the underlying control unit. The multiphase current output by the high-power grid-connected interface device is controlled through the interface controller on the PGA. With capacitor voltage Perform synchronous sampling; for Mutually Mutually The collection of phase currents; for Mutually Mutually Phase capacitor voltage; right and After performing anti-aliasing filtering and scaling transformation sequentially, a Clarke transformation is executed to obtain a two-phase stationary coordinate system. , Feedback volume , , , ;in, Two-phase stationary coordinate system for feedback shaft current; Two-phase stationary coordinate system for feedback shaft current; Two-phase stationary coordinate system for feedback Shaft capacitor voltage; Two-phase stationary coordinate system for feedback Shaft capacitor voltage; Within each control cycle, the reference voltage command in the two-phase stationary coordinate system is read. , Subtract the corresponding feedback voltage respectively , ,get and ;in, For the current number Under one control cycle, the two-phase stationary coordinate system Shaft voltage error; For the current number Under one control cycle, the two-phase stationary coordinate system Shaft voltage error; For the current number Under one control cycle, the two-phase stationary coordinate system Reference voltage command for the shaft; For the current number Under one control cycle, the two-phase stationary coordinate system Reference voltage command for the shaft; Will and The data is stored in the on-chip circular memory. During the repetitive control calculation, the on-chip circular memory is accessed to obtain the historical voltage error corresponding to the previous fundamental cycle, and combined with the current voltage error. and After phase compensation filter and Calculation output ;in, For the transform domain The phase compensation filter below; For the current number The control output quantity obtained by repeated control calculation under each control cycle; According to the reference voltage and voltage feedforward coefficient ,calculate ; For the current number The control output obtained by calculating the voltage feedforward coefficient under each control cycle; For the current number Reference voltage under each control cycle; Will and Superimpose the results to obtain the current number. k Final control voltage command under each control cycle ; based on A modulated wave is generated and compared with multiple set triangular carrier waves to output the switching decision logic command of the power conversion submodule of the high-power grid-connected interface device. Using the high-frequency clock of the programmable gate array, the switching decision logic instruction of the submodule is precisely calibrated for the switching time, and a dead time is inserted to generate a prototype of the driving pulse. The prototype drive pulse is input to the gate drive circuit of each sub-module for amplification and isolation, and the actual multi-level PWM drive signal is output. The upper-level intelligent controller is deployed within the high-performance microprocessor of the upper-level control unit, which operates with millisecond-level control cycles. The lower-level control unit acquires the state space vector of the power hardware-in-the-loop system and uploads it to the upper-level control unit. The upper-level control unit processes the state space vector and, based on a deep deterministic policy gradient algorithm, adjusts the repetitive control gain. and voltage feedforward coefficient Optimization is performed to obtain the optimized dynamic control parameters. , ; The optimized repetitive control gain; The optimized voltage feedforward coefficient; Optimized dynamic control parameters , The data is sent to the lower-level control unit, which then uses the optimized dynamic control parameters. , Regenerate the multi-level PWM drive signal to complete the interaction mechanism between the upper-level control unit and the lower-level control unit; Through the collaborative deployment and interaction mechanism between the phase-compensated enhanced repetitive-voltage feedforward composite controller and the upper-level intelligent controller, a hierarchical collaborative control architecture for power grid fault simulation is formed.

[0027] The specific process of step S5 is as follows: Based on the hierarchical collaborative control architecture for power grid fault simulation, a hardware-in-the-loop test platform for the high-power grid-connected interface device was built. Based on the hardware-in-the-loop test platform, the deep reinforcement learning controller is trained offline to obtain the trained deep reinforcement learning controller, specifically as follows: A digital simulation environment consistent with the actual high-power grid-connected interface device is established based on the hardware-in-the-loop test platform. In the digital simulation environment, an interaction mechanism between the deep reinforcement learning controller and the digital simulation environment is constructed; the output voltage of the high-power grid-connected interface device is used. , Axial components and High-power grid-connected interface device output voltage tracking error and System frequency deviation and the characteristic quantities that characterize the type of fault. As status input; Repetitive control gain in a phase-compensated repetitive-voltage feedforward composite controller and voltage feedforward coefficient As a continuous action output, it is the control strategy of the deep reinforcement learning controller; Within each simulation time step of the interactive mechanism, the repetitive control gain of the phase-compensated repetitive-voltage feedforward composite controller is adjusted based on the continuous action output. and voltage feedforward coefficient Update, from the updated , The simulation system's state evolution is driven; a reward function is constructed based on the state input; the effect of driving the simulation system's state evolution is quantitatively evaluated through control strategies to obtain reward values, and the next state of the power hardware-in-the-loop system is calculated; the current state of the power hardware-in-the-loop system, continuous action outputs, reward values, and the next state constitute interactive data and are stored in the experience playback pool. The updated repetitive control gain; The updated voltage feedforward coefficients; Based on the experience replay mechanism, historical interaction data is randomly extracted from the experience replay pool. Based on the historical interaction data, the evaluation network of the deep reinforcement learning controller is optimized by minimizing the value function, and then the action network of the deep reinforcement learning controller is optimized by the policy gradient method. After optimizing the evaluation network and action network of the deep reinforcement learning controller, the iterative training process of the control policy begins: During the iterative training of the control strategy, changes in grid parameters, perturbations of interface device parameters, and measurement noise are introduced to repeatedly train the control strategy. The changes in power grid parameters are set based on the impedance fluctuation range under typical operating conditions of the power system, the perturbation of interface device parameters is set based on the manufacturing tolerance range of power electronic equipment, and the measurement noise is set based on the measurement accuracy level of industrial-grade sensors. When the reward function converges during training and the power hardware-in-the-loop system exhibits good stability and control performance under typical three-phase short-circuit fault conditions, the training of the deep reinforcement learning controller is completed, the trained deep reinforcement learning controller is obtained, and the trained deep reinforcement learning controller is deployed to the high-power grid-connected interface device.

[0028] The specific process of step S6 is as follows: A high-power grid-connected interface device is used to simulate a three-phase short-circuit fault condition. The three-phase short-circuit fault condition is accompanied by voltage drop, instantaneous current surge, and frequency fluctuation of the power hardware in-loop system. During the fault process, the upper-level control unit calls the trained deep reinforcement learning controller to process the voltage drop, instantaneous current surge, and frequency fluctuation of the power hardware-in-the-loop system in real time, and outputs dynamic control parameters. The lower-level control unit generates multi-level PWM drive signals based on the dynamic control parameters, and drives the high-power grid-connected interface device through the multi-level PWM drive signals to verify the effectiveness of the dynamic control parameters and the dynamic stability of the waveform. The effectiveness of the dynamic control parameters and the dynamic stability of the waveform together support the dynamic response and stability performance of the power hardware-in-the-loop system.

[0029] A computer storage medium storing a computer program, which, when run, implements a method for enhancing the stability of a power hardware-in-the-loop system based on DDPG-RP.

[0030] An electronic device includes a processor, a memory, and a bus, wherein the processor and the memory are connected via the bus, wherein the memory is used to store a set of program code, and the processor is used to call the program code stored in the memory to execute a power hardware-in-the-loop system stability enhancement method based on DDPG-RP.

[0031] To verify the effectiveness of the proposed trained deep reinforcement learning controller DDPG-RP, a comparative analysis was conducted between the trained DDPG-RP and the benchmark strategy of traditional repetitive control combined with voltage feedforward P control (RC-P). Figure 2 As shown, Figure 2 The dynamic response results of the power hardware-in-the-loop system under three-phase short-circuit fault conditions are shown. The horizontal axis Time is in seconds (s) and ranges from 0 to 50 s, clearly presenting the entire process from the occurrence of the fault to the recovery of the system. Figure 2 The medium curve contains four key electrical quantities: For power hardware-in-the-loop system frequency, For grid-connected side voltage, the first phase voltage is output by the high-power grid-connected interface device. The high-power grid-connected interface device outputs the second phase voltage. The solid red line represents the control result of the deep reinforcement learning controller DDPG-RP after training, and the dashed black line represents the control result of RC-P.

[0032] From the power hardware-in-the-loop system frequency response (corresponding to) Figure 2 middle fAs can be seen from the waveform, when the fault occurs at 5s, the frequency under RC-P control shows obvious overshoot (peak value exceeds 1.1pu) and oscillation, and it takes a long time to recover to steady state (1pu). In contrast, the trained deep reinforcement learning controller DDPG-RP can effectively suppress frequency overshoot and recover the frequency to the rated value in a short time after only a small fluctuation, showing stronger transient frequency support capability.

[0033] At the same time, on the grid-connected side voltage and the output first phase voltage of the high-power grid-connected interface device The high-power grid-connected interface device outputs the second phase voltage. In the dynamic process (corresponding to) Figure 2 (corresponding waveforms in the diagram) The voltage drop amplitude under the control of the trained deep reinforcement learning controller DDPG-RP is smaller, the oscillation decay is faster, and the steady-state recovery is smoother; while the voltage under RC-P control shows obvious oscillation after the fault, indicating that the trained deep reinforcement learning controller DDPG-RP has better voltage regulation performance under fault disturbance.

[0034] Overall, compared to RC-P, the trained deep reinforcement learning controller DDPG-RP achieves smaller frequency and voltage fluctuations under three-phase short-circuit faults and significantly shortens the transient recovery time of the power hardware-in-the-loop system. This indicates that the trained deep reinforcement learning controller DDPG-RP strategy possesses strong adaptive adjustment capabilities and comprehensive stability support effects, effectively improving the dynamic operating performance of high-power grid-connected interface devices under severe grid disturbances.

[0035] also, Figure 3 The convergence curve of the reward function of the trained deep reinforcement learning controller DDPG-RP during training under a three-phase short-circuit fault condition is shown (the horizontal axis represents the training period, and the vertical axis represents the reward value). In this scenario, the trained deep reinforcement learning controller DDPG-RP gradually explores and optimizes the control strategy of the upper-level control unit through real-time interaction with the power hardware-in-the-loop system. In the early stage of training, as the trained deep reinforcement learning controller DDPG-RP is in the exploration phase of the state space and action space, the reward function fluctuates greatly, and the reward value even drops to around -120 at one point. As the training period increases, the trained deep reinforcement learning controller DDPG-RP gradually learns the optimal control mapping to cope with fault disturbances, and the reward value gradually recovers. After about 400 training periods, the reward value curve tends to stabilize and finally converges to a stable optimal range of about -60.

[0036] This convergence process not only verifies the learning ability and stability of the trained deep reinforcement learning controller DDPG-RP in complex control environments, but also shows that the DDPG-RP trained offline can provide accurate intelligent decision support for voltage feedforward control, thereby ensuring that the power hardware-in-the-loop system has extremely high initial stability and response speed when deployed online in the future.

[0037] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.

Claims

1. A method for enhancing the stability of a power hardware-in-the-loop system based on DDPG-RP, characterized in that, include: Step S1: Establish an equivalent circuit model of the power hardware-in-the-loop system that includes the characteristics of grid fault disturbances, analyze the dynamic interaction characteristics of the power hardware-in-the-loop system under grid fault conditions, and determine the impedance matching conditions for maintaining the stable operation of the power hardware-in-the-loop system. Step S2: Based on the impedance matching condition for stable operation of the power hardware-in-the-loop system, design a phase-compensated and enhanced repetitive-voltage feedforward composite controller. Step S3: Construct the upper-layer intelligent controller using the deep deterministic strategy gradient algorithm; Step S4: Deploy the phase-compensated enhanced repetitive-voltage feedforward composite controller in the bottom control unit, and deploy the upper-level intelligent controller in the upper control unit. Based on the bottom control unit and the upper control unit, a hierarchical collaborative control architecture for grid fault simulation is formed. Step S5: Based on the hierarchical collaborative control architecture, build a power hardware-in-the-loop test platform, and complete the offline training and online deployment of the deep reinforcement learning controller on the power hardware-in-the-loop test platform to obtain the trained deep reinforcement learning controller. Step S6: Simulate a three-phase short-circuit fault condition using a high-power grid-connected interface device, and verify the dynamic response and stability performance under the simulated three-phase short-circuit fault condition using a trained deep reinforcement learning controller.

2. The method for enhancing the stability of a power hardware-in-the-loop system based on DDPG-RP according to claim 1, characterized in that: The specific process of step S1 is as follows: An equivalent circuit model of the power hardware-in-the-loop system is established based on an ideal transformer model. The digital simulation side of the power hardware-in-the-loop system includes models simulating symmetrical / asymmetrical short circuits and voltage sags / droops in the power grid. The Thevenin equivalent impedance on the digital simulation side is denoted as... The high-power grid-connected interface device and the power electronic device together constitute the measured side; the equivalent impedance of the measured side is denoted as... Establish a hardware-in-the-loop total delay that takes power into account. The open-loop transfer function model is expressed as: ; In the formula, For the open-loop transfer function model of the power hardware-in-the-loop system; It is a complex frequency variable; It is a natural constant; Using the Pad approximation After rationalization approximation, it can be expressed as: ; The open-loop transfer function model after Pad approximation is expressed as: ; ; In the formula, This is the open-loop transfer function model after Pad approximation; Define the impedance ratio function as The open-loop transfer function model after Pad approximation is written as: ,Will Convert to ;right Perform mathematical transformations from the complex frequency domain to the frequency domain. The open-loop frequency response of the power hardware-in-the-loop system is obtained, which is expressed as: , ; In the formula, This is the impedance ratio function of the power hardware in the frequency domain of the loop system. It is the imaginary unit in the field of engineering. Angular frequency; This refers to the open-loop frequency response of the power hardware-in-the-loop system. Based on the Nyquist stability criterion, for The Nyquist curve is analyzed to determine the stability of the power hardware-in-the-loop system; based on the stability of the power hardware-in-the-loop system, [further analysis is needed]. Processing is performed to make By satisfying both phase stability constraints and amplitude stability constraints, impedance matching conditions that ensure stable operation of the power hardware in-loop system are obtained.

3. The method for enhancing the stability of a power hardware-in-the-loop system based on DDPG-RP according to claim 2, characterized in that: The specific process for obtaining the impedance matching conditions for stable operation of the power hardware in a loop system is as follows: Based on the Nyquist stability criterion, for The Nyquist curve was analyzed. when The Nyquist curve does not enclose the complex plane. At this point, the stability of the power hardware in-loop system is considered; when The Nyquist curve encloses the complex plane on At this point, the closed-loop power hardware-in-the-loop system is unstable, which affects the power hardware-in-the-loop system. , Total delay Re-identification and adjustment; Based on the stability of the power hardware in-loop system, Processing is performed to make Both phase stability constraints and magnitude stability constraints are satisfied, specifically: When the power hardware in the loop system is in the low-frequency range, by adjusting... The circuit parameters were corrected. Phase curve, avoid The phase lag is -180°, which satisfies the phase stability constraint. When the power hardware in-loop system is in the high-frequency band, by... The amplitude is limited to prevent the Nyquist curve from approaching the critical point, thus satisfying the amplitude stability constraint. when By satisfying both phase stability constraints and amplitude stability constraints, impedance matching conditions that ensure stable operation of the power hardware in-loop system are obtained.

4. The method for enhancing the stability of a power hardware-in-the-loop system based on DDPG-RP according to claim 3, characterized in that: The specific process of step S2 is as follows: Based on the total delay of the power hardware in the loop system identified in step S1 Design a phase compensator; introduce the phase compensator into the repetitive controller to form a phase-compensated repetitive controller, which includes a repetitive control gain; A voltage feedforward controller is connected in parallel with a phase-compensated repetitive controller to form a phase-compensated repetitive-voltage feedforward composite controller. At the instant a three-phase short-circuit fault occurs in the power grid, the voltage feedforward coefficient is obtained by real-time adjustment of the repetitive controller; The difference between the actual voltage output by the high-power grid-connected interface device and the reference voltage of the power hardware in-loop system is input to the phase-compensated and enhanced repetitive-voltage feedforward composite controller for processing, generating harmonic suppression components and transient support components. The harmonic suppression components and transient support components are linearly superimposed to obtain the modulation wave command of the high-power grid-connected interface device. Based on the modulation wave command, the repetitive control gain and voltage feedforward coefficient are adaptively optimized online through a deep deterministic strategy gradient algorithm to obtain the optimal combination of control parameters in real time, thus forming intelligent parameter risk avoidance. The phase compensator implements a phase stabilization mechanism through hardware circuitry to offset the phase lag caused by the inherent delay of the power hardware in the loop system. The phase stabilization mechanism of the hardware circuitry is called physical structure phase stabilization. An enhanced control logic is formed by organically integrating a two-layer mechanism of physical structural stability and intelligent parameter risk avoidance. The final design of a phase-compensated repetitive-voltage feedforward composite controller based on enhanced control logic was completed.

5. The method for enhancing the stability of a power hardware-in-the-loop system based on DDPG-RP according to claim 4, characterized in that: The specific process of step S3 is as follows: A deep deterministic policy gradient algorithm is adopted, combined with state space vectors and reward functions to design an upper-level intelligent controller; The state-space vector includes: the output voltage of the high-power grid-connected interface device. , Shaft components, output voltage tracking error of high-power grid-connected interface devices, frequency deviation of power hardware-in-the-loop systems, and characteristic quantities characterizing fault types.

6. The method for enhancing the stability of a power hardware-in-the-loop system based on DDPG-RP according to claim 5, characterized in that: The specific process of step S4 is as follows: A phase-compensated enhanced repetitive-voltage feedforward composite controller is deployed on the programmable gate array (PGA) of the underlying control unit. The multiphase current output by the high-power grid-connected interface device is controlled through the interface controller on the PGA. With capacitor voltage Perform synchronous sampling; for Mutually Mutually The collection of phase currents; for Mutually Mutually Phase capacitor voltage; right and After performing anti-aliasing filtering and scaling transformation sequentially, a Clarke transformation is executed to obtain a two-phase stationary coordinate system. , Feedback volume , , , ;in, Two-phase stationary coordinate system for feedback shaft current; Two-phase stationary coordinate system for feedback shaft current; Two-phase stationary coordinate system for feedback Shaft capacitor voltage; Two-phase stationary coordinate system for feedback Shaft capacitor voltage; Within each control cycle, the reference voltage command in the two-phase stationary coordinate system is read. , Subtract the corresponding feedback voltage respectively , ,get and ;in, For the current number Under one control cycle, the two-phase stationary coordinate system Shaft voltage error; For the current number Under one control cycle, the two-phase stationary coordinate system Shaft voltage error; For the current number Under one control cycle, the two-phase stationary coordinate system Reference voltage command for the shaft; For the current number Under one control cycle, the two-phase stationary coordinate system Reference voltage command for the shaft; Will and The data is stored in the on-chip circular memory. During the repetitive control calculation, the on-chip circular memory is accessed to obtain the historical voltage error corresponding to the previous fundamental cycle, and combined with the current voltage error. and The output is calculated after phase compensation filter and repetitive control gain calculation. ;in, For the current number The control output quantity obtained by repeated control calculation under each control cycle; Based on reference voltage and voltage feedforward coefficient calculation ; This refers to the control output obtained by voltage feedforward calculation in the current k-th control cycle; This is the reference voltage command for the current k-th control cycle; Will and Superimpose the results to obtain the current number. k Final control voltage command under each control cycle ; based on A modulated wave is generated and compared with multiple set triangular carrier waves to output the switching decision logic command of the power conversion submodule of the high-power grid-connected interface device. Using the high-frequency clock of the programmable gate array, the switching decision logic instruction of the submodule is calibrated for the switching time, and a dead time is inserted to generate a prototype of the driving pulse; The prototype drive pulse is input to the gate drive circuit of each sub-module for amplification and isolation, and the actual multi-level PWM drive signal is output. The upper-level intelligent controller is deployed within the high-performance microprocessor of the upper-level control unit, which operates with millisecond-level control cycles. The lower-level control unit acquires the state space vector of the power hardware-in-the-loop system and uploads it to the upper-level control unit. The upper-level control unit processes the state space vector and optimizes the repetitive control gain and voltage feedforward coefficient based on a deep deterministic policy gradient algorithm to obtain optimized dynamic control parameters. , ; The optimized repetitive control gain; The optimized voltage feedforward coefficient; Optimized dynamic control parameters , The data is sent to the lower-level control unit, which then uses the optimized dynamic control parameters. , Regenerate the multi-level PWM drive signal to complete the interaction mechanism between the upper-level control unit and the lower-level control unit; Through the collaborative deployment and interaction mechanism between the phase-compensated enhanced repetitive-voltage feedforward composite controller and the upper-level intelligent controller, a hierarchical collaborative control architecture for power grid fault simulation is formed.

7. The method for enhancing the stability of a power hardware-in-the-loop system based on DDPG-RP according to claim 6, characterized in that: The specific process of step S5 is as follows: Based on the hierarchical collaborative control architecture for power grid fault simulation, a hardware-in-the-loop test platform for the high-power grid-connected interface device was built. Based on the hardware-in-the-loop test platform, the deep reinforcement learning controller is trained offline to obtain the trained deep reinforcement learning controller, specifically as follows: A digital simulation environment consistent with the actual high-power grid-connected interface device is established based on the hardware-in-the-loop test platform. In the digital simulation environment, an interaction mechanism between the deep reinforcement learning controller and the digital simulation environment is constructed; the output voltage of the high-power grid-connected interface device is used. , The axial component, the output voltage tracking error of the high-power grid-connected interface device, the frequency deviation of the power hardware-in-the-loop system, and the characteristic quantities characterizing the fault type are used as status inputs. Using the repetitive control gain and voltage feedforward coefficient in the phase-compensated repetitive-voltage feedforward composite controller as the continuous action output is the control strategy of the deep reinforcement learning controller. Within each simulation time step of the interactive mechanism, the repetitive control gain and voltage feedforward coefficient of the phase-compensated repetitive-voltage feedforward composite controller are updated according to the continuous action output. The updated repetitive control gain and updated voltage feedforward coefficient drive the state evolution of the simulation system. At the same time, a reward function is constructed according to the state input. The effect of the control strategy on driving the state evolution of the simulation system is quantitatively evaluated to obtain the reward value, and the next state of the power hardware in the loop is calculated. The current state of the power hardware in-loop system, continuous action output, reward value and the next state constitute interactive data and are stored in the experience replay pool. Based on the experience replay mechanism, historical interaction data is randomly extracted from the experience replay pool. Based on the historical interaction data, the evaluation network of the deep reinforcement learning controller is optimized by minimizing the value function, and then the action network of the deep reinforcement learning controller is optimized by the policy gradient method. After optimizing the evaluation network and action network of the deep reinforcement learning controller, the iterative training process of the control policy begins: During the iterative training of the control strategy, changes in grid parameters, perturbations of interface device parameters, and measurement noise are introduced to repeatedly train the control strategy. The changes in power grid parameters are set based on the impedance fluctuation range under typical operating conditions of the power system, the perturbation of interface device parameters is set based on the manufacturing tolerance range of power electronic equipment, and the measurement noise is set based on the measurement accuracy level of industrial-grade sensors. When the reward function converges during training and the power hardware-in-the-loop system exhibits good stability and control performance under typical three-phase short-circuit fault conditions, the training of the deep reinforcement learning controller is completed, the trained deep reinforcement learning controller is obtained, and the trained deep reinforcement learning controller is deployed to the high-power grid-connected interface device.

8. The method for enhancing the stability of a power hardware-in-the-loop system based on DDPG-RP according to claim 7, characterized in that: The specific process of step S6 is as follows: A high-power grid-connected interface device is used to simulate a three-phase short-circuit fault condition. The three-phase short-circuit fault condition is accompanied by voltage drop, instantaneous current surge, and frequency fluctuation of the power hardware in-loop system. During the fault process, the upper-level control unit calls the trained deep reinforcement learning controller to process the voltage drop, instantaneous current surge, and frequency fluctuation of the power hardware-in-the-loop system in real time, and outputs dynamic control parameters. The lower-level control unit generates multi-level PWM drive signals based on the dynamic control parameters, and drives the high-power grid-connected interface device through the multi-level PWM drive signals to verify the effectiveness of the dynamic control parameters and the dynamic stability of the waveform. The effectiveness of the dynamic control parameters and the dynamic stability of the waveform together support the dynamic response and stability performance of the power hardware-in-the-loop system.

9. A computer storage medium, characterized in that: The computer storage medium stores a computer program, which, when run, implements a method for enhancing the stability of a power hardware-in-the-loop system based on DDPG-RP as described in any one of claims 1 to 8.

10. An electronic device, characterized in that: The system includes a processor, a memory, and a bus, wherein the processor and the memory are connected via the bus, the memory is used to store a set of program code, and the processor is used to call the program code stored in the memory to execute a power hardware-in-the-loop stability enhancement method based on any one of claims 1 to 8.