Method, device and readable storage medium for adaptively adjusting loop parameters of digital power supply loop
By applying reinforcement learning algorithms in the digital power loop to optimize PID parameters, combined with the protection measurement feedback switching mechanism, the problems of insufficient adjustment speed and poor stability in complex scenarios are solved in the existing technology, and fast response, high-precision control and circuit safety protection are achieved.
Patent Information
- Application Number
- CN202510412671.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-03
- Publication Date
- 2025-06-17
- Estimated Expiration
- 2045-04-03
AI Technical Summary
The existing digital power loops have insufficient adjustment speed and poor stability in complex scenarios, and cannot optimize control performance in real time according to dynamic load changes, and lack deep coordination with the protection measurement alarm mechanism.
The reinforcement learning algorithm is adopted to optimize PID parameters in real time through the deep deterministic strategy gradient algorithm (DDPG), and combined with the protection measurement feedback switching mechanism, adaptive adjustment of loop parameters is achieved.
It realizes fast response, high-precision control and circuit safety protection in complex scenarios, shortening the adjustment time by more than 50%, and reducing the overshoot by more than 60%, improving the stability and safety of the system.
Smart Images

Figure CN119921540B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of power control, and in particular to a method, device and readable storage medium for adaptively adjusting loop parameters of a digital power loop, which are applicable to scenarios such as high-current power systems and automatic test equipment (such as chip testers). Background Art
[0002] Due to the advantages of high precision, flexibility and integration, digital power loops are widely used in industrial power control. In the prior art, digital loops usually adopt a fixed-parameter PID algorithm to adjust the output through voltage / current ADC feedback. However, when the system faces complex scenarios (such as sudden load changes and abnormal contact resistance), fixed parameters are difficult to adapt dynamically, resulting in insufficient adjustment speed, excessive overshoot, and even the risk of circuit burnout caused by invalid feedback. For example, after a traditional scheme triggers a protection measurement alarm (such as a Kelvin alarm), although the feedback source is switched to the protection measurement value (such as the Kelvin voltage drop) to quickly reduce the safety risk, the loop parameters still rely on manual presetting and cannot be dynamically optimized according to real-time load changes. In addition, existing adaptive control methods (such as fuzzy PID) need to preset a rule base, are difficult to handle high-dynamic and multi-variable scenarios, and lack in-depth coordination with the protection measurement alarm mechanism.
[0003] Therefore, there is an urgent need for a new method, device and readable storage medium for adaptively adjusting loop parameters of a digital power loop to solve the problems existing in the prior art. Summary of the Invention
[0004] Embodiments of the present invention provide a method, device and readable storage medium for adaptively adjusting loop parameters of a digital power loop, aiming at the problems existing in the current technology that although the safety risk is reduced through feedback source switching and fixed-parameter adjustment when a protection measurement alarm is triggered, the fixed loop parameters result in insufficient adjustment speed and poor stability in complex scenarios, and the control performance cannot be optimized in real time according to dynamic load changes.
[0005] The core technology of the present invention is mainly an adaptive adjustment technology of loop parameters based on intelligent optimization algorithms, which dynamically optimizes PID parameters (such as proportional, integral and differential coefficients) in real time through reinforcement learning (such as DDPG), genetic algorithms, etc., and combines a protection measurement feedback switching mechanism to achieve fast response, high-precision control and circuit safety protection in complex scenarios.
[0006] In a first aspect, the present invention provides a method for adaptively adjusting loop parameters of a digital power loop, the method comprising the following steps:
[0007] (a)Collect the operation status data of the digital power supply in real time, including output voltage, output current, protection measurement value, load impedance change rate, and loop regulation error, to construct a state vector;
[0008] (b)When the protection measurement value triggers an alarm, switch the feedback source of the digital loop from the conventional power supply output measurement value to the protection measurement value, and set the reference value to a preset safe value;
[0009] (c)Dynamically adjust the control parameters of the digital loop based on the reinforcement learning algorithm, specifically including:
[0010] (c1)Through the Actor network in the deep deterministic policy gradient algorithm, output the incremental adjustment amounts of the proportional coefficient Kp, integral coefficient Ki, and derivative coefficient Kd according to the state vector;
[0011] (c2)When the protection measurement alarm is triggered, forcibly increase the proportional coefficient Kp to a preset multiple of the current value, and limit the adjustment range of the integral coefficient Ki within a preset percentage range or reduce the integral coefficient to a preset percentage of the current value;
[0012] (c3)Update the parameters of the PID controller in real time according to the incremental adjustment amounts;
[0013] (d)When the protection measurement value drops to the safe value, disconnect the digital loop feedback and cut off the physical connection between the power supply and the load.
[0014] Furthermore, the reward function of the reinforcement learning algorithm in step (c) is:
[0015] R = -(αT s + βσ% + γe ss + λP DAC )
[0016] where T s is the regulation time, σ% is the overshoot, e ss is the steady-state error, P DAC is the DAC output power consumption, α is the weight coefficient of the regulation time, β is the weight coefficient of the overshoot, γ is the weight coefficient of the steady-state error, and λ is the weight coefficient of the DAC output power consumption;
[0017] When the protection measurement alarm is triggered, apply an additional penalty term R penalty = -100×T alarm , T alarm is the alarm duration.
[0018] Further, in step (c1), the Actor network is a lightweight fully connected network, including an input layer, a hidden layer, and an output layer. The number of nodes in the input layer is the dimension of the state vector, and the number of nodes in the output layer is 3, corresponding to ΔKp, ΔKi, and ΔKd. The network weights are quantized using 8-bit fixed-point numbers.
[0019] Further, in step (c2), when the protection measurement alarm is triggered, it is forced to switch to a preset combination of safety parameters, including increasing the proportional coefficient to 2 times the current value and restricting the adjustment range of the integral coefficient to ±5%.
[0020] Further, in step (c2), when the protection measurement alarm is triggered, it is forced to switch to a preset combination of safety parameters, including increasing the proportional coefficient to 2 times the current value and reducing the integral coefficient to 0.5 times the current value.
[0021] Further, in step (c1), the state vector further includes:
[0022] The trigger flag bit of the protection measurement alarm;
[0023] The power output mode.
[0024] Further, it further includes:
[0025] Deploy a reinforcement learning algorithm in the embedded controller to ensure that the control cycle is synchronized with the digital loop, and the single inference time ≤ 50 microseconds.
[0026] In a second aspect, the present invention provides a device for adaptively adjusting the loop parameters of a digital power loop, including:
[0027] A data acquisition module that real-time collects the operating state data of the digital power supply, including output voltage, output current, protection measurement value, load impedance change rate, and loop regulation error, to construct a state vector;
[0028] A feedback switching module that, when the protection measurement value triggers an alarm, switches the feedback source of the digital loop from the conventional power output measurement value to the protection measurement value and sets the reference value to a preset safety value;
[0029] A reinforcement learning control module that dynamically adjusts the control parameters of the digital loop based on the reinforcement learning algorithm, specifically including:
[0030] Through the Actor network in the deep deterministic policy gradient algorithm, output the incremental adjustment amounts of the proportional coefficient Kp, the integral coefficient Ki, and the differential coefficient Kd according to the state vector;
[0031] The execution module, when the protection measurement alarm is triggered, forcibly increases the proportional coefficient Kp to a preset multiple of the current value, and restricts the adjustment range of the integral coefficient Ki within a preset percentage range or reduces the integral coefficient to a preset percentage of the current value; updates the parameters of the PID controller in real time according to the incremental adjustment amount; when the protection measurement value drops to the safety value, disconnects the digital loop feedback and cuts off the physical connection between the power supply and the load.
[0032] In a third aspect, the present invention provides an electronic device, including a memory and a processor. A computer program is stored in the memory, and the processor is configured to run the computer program to execute the above-mentioned loop parameter adaptive adjustment method of the digital power loop.
[0033] In a fourth aspect, the present invention provides a readable storage medium. A computer program is stored in the readable storage medium, and the computer program includes program codes for controlling a process to execute the process. The process includes the loop parameter adaptive adjustment method of the digital power loop as described above.
[0034] The main contributions and innovations of the present invention are as follows:
[0035] 1. Dynamic parameter optimization driven by reinforcement learning
[0036] The deep deterministic policy gradient algorithm (DDPG) is used to optimize the PID parameters in real time, breaking through the limitation that traditional fixed parameters cannot adapt to dynamic working conditions such as load mutations and contact resistance fluctuations, and can shorten the adjustment time by more than 50% (compared with traditional PID).
[0037] The reward function introduces multi-objective optimization of adjustment time, overshoot, steady-state error and power consumption to ensure a balance among the response speed, stability and energy efficiency of the system.
[0038] 2. Forced parameter strategy when the alarm is triggered
[0039] When the protection measurement alarm is triggered, the proportional coefficient Kp is forcibly increased (such as 2 times) and the integral coefficient Ki is restricted (such as ±5% or reduced to 0.5 times), significantly accelerating the decline speed of the protection measurement value.
[0040] Avoid the problem of adjustment failure of traditional fixed parameters during alarm, such as ineffective feedback or oscillation during Kelvin alarm.
[0041] 3. Lightweight algorithm and hardware co-optimization
[0042] The Actor network adopts 8-bit fixed-point quantization to reduce the consumption of computing resources, realizes real-time deployment on an embedded controller (such as STM32H7), and the single inference time ≤ 50 μs, meeting the high-speed control requirements of the digital loop.
[0043] The control period is synchronized with the loop to ensure that the parameter adjustment is consistent with the system's dynamic response.
[0044] 4. Multi-dimensional state perception and operating condition adaptation
[0045] The state vector includes information such as protection alarm flag bits and power output modes. The algorithm can dynamically adjust the optimization objectives according to different operating conditions (constant voltage / constant current / short circuit) to improve adaptability in complex scenarios.
[0046] Real-time estimation of the load impedance change rate (based on Kalman filtering) anticipates the parameter adjustment requirements in advance and reduces the overshoot by more than 60%.
[0047] 5. Improvement of safety and reliability
[0048] An additional penalty term (R penalty =-100×T alarm ) during alarm strengthens the constraint on the alarm duration to ensure rapid risk elimination.
[0049] Synchronous control and parameter gradual transition algorithm at the hardware level avoid system oscillations caused by parameter mutations and enhance stability.
[0050] 6. Generalizability and expandability
[0051] The reinforcement learning model can be continuously optimized through online learning to adapt to new operating conditions; the modular design facilitates integration into existing digital power systems without additional hardware costs.
[0052] Details of one or more embodiments of the present invention are set forth in the following drawings and description to make other features, objects, and advantages of the present invention more concise and understandable. BRIEF DESCRIPTION OF THE DRAWINGS
[0053] The drawings described herein are used to provide a further understanding of the present invention and constitute a part of the present invention. The illustrative embodiments and descriptions of the present invention are used to explain the present invention and do not constitute an improper limitation of the present invention. In the drawings:
[0054] Figure 1 is a flowchart of a method for adaptively adjusting loop parameters of a digital power loop according to an embodiment of the present invention;
[0055] Figure 2 is a flowchart of dynamically adjusting control parameters of a digital loop based on a reinforcement learning algorithm according to an embodiment of the present invention;
[0056] Figure 3 is a schematic diagram of a digital power loop in the prior art;
[0057] Figure 4 is a schematic diagram of the hardware structure of an electronic device according to an embodiment of the present invention. Detailed Implementation Modes
[0058] Here, exemplary embodiments will be described in detail, and examples thereof are shown in the accompanying drawings. When the following description refers to the accompanying drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The implementation modes described in the following exemplary embodiments do not represent all implementation modes consistent with one or more embodiments of this specification. On the contrary, they are merely examples of devices and methods consistent with some aspects of one or more embodiments of this specification as detailed in the appended claims.
[0059] It should be noted that: In other embodiments, the steps of the corresponding method are not necessarily executed in the order shown and described in this specification. In some other embodiments, the steps included in the method may be more or less than those described in this specification. In addition, a single step described in this specification may be decomposed into multiple steps for description in other embodiments; and multiple steps described in this specification may also be combined into a single step for description in other embodiments.
[0060] As Figure 3 shown is a schematic diagram of an existing digital power loop based on protection measurement feedback. After triggering an alarm for protection measurement (such as a Kelvin alarm), by using the protection measurement value as the feedback of the digital loop and setting a safety reference, the protection measurement value (Kelvin Voltage Drop measurement value, Kelvin high voltage drop or Kelvin low voltage drop) can be quickly reduced below a safety value (such as 0).
[0061] However, when the system faces complex scenarios (such as sudden load changes, abnormal contact resistance), fixed parameters are difficult to adapt dynamically, resulting in insufficient adjustment speed, excessive overshoot, and even the risk of circuit burnout due to invalid feedback. For example, after triggering an alarm for protection measurement (such as a Kelvin alarm) in the traditional scheme, although the feedback source is switched to the protection measurement value (such as Kelvin voltage drop) to quickly reduce the safety risk, the loop parameters still rely on manual presetting and cannot be dynamically optimized according to real-time load changes. In addition, existing adaptive control methods (such as fuzzy PID) require a preset rule base, are difficult to handle high-dynamic, multi-variable scenarios, and lack in-depth coordination with the protection measurement alarm mechanism.
[0062] Based on this, the present invention uses reinforcement learning to solve the problems existing in the prior art.
[0063] Embodiment 1
[0064] In Figure 3Based on the above, the loop parameter adaptive adjustment system of the digital power loop of the present invention mainly consists of a data acquisition module, a feedback switching module, a reinforcement learning control module, and an execution module. The actual hardware it is applied to includes a controller, a digital-to-analog converter (DAC), an analog-to-digital converter (ADC), a protection measurement module, and a digital isolator.
[0065] Among them, the controller: runs digital loop algorithms (such as PID control) and reinforcement learning algorithms (DDPG), and an embedded processor (such as ARM Cortex-M7 or STM32H7) can be used.
[0066] Digital-to-analog converter (DAC): Converts the digital signal output by the controller into an analog signal to drive the power output, such as Figure 3 the main DAC in
[0067] Analog-to-digital converter (ADC): Includes a conventional voltage / current ADC (feedback power output value) and a protection measurement ADC (such as a Kelvin voltage drop ADC, monitoring the contact resistance between the Force Line and the load, such as Figure 3 the four ADCs in
[0068] Power amplifier: Amplifies the DAC output signal to drive the power output terminal, such as Figure 3 the power amplifier in
[0069] Digital isolator (optional): Isolates the protection measurement signal and the main power signal to prevent interference.
[0070] Physical switch: Disconnects the connection between the power supply and the load after the alarm is lifted, such as Figure 3 the output connection switch in
[0071] The following elaborates on the specific implementation manners of the present invention in combination with each module.
[0072] Specifically, the embodiment of the present invention provides a method for adaptively adjusting the loop parameters of a digital power loop. Specifically, referring to Figure 1 , the method includes:
[0073] (a) Real-time collect the operating state data of the digital power supply, including output voltage, output current, protection measurement value, load impedance change rate, and loop regulation error, to construct a state vector;
[0074] In this embodiment, the (a) step includes an initialization step:
[0075] 1. Set the conventional loop parameters:
[0076] Configure the initial parameters (Kp = 1.0, Ki = 0.05, Kd = 0.1) of the PID controller (integrated in the controller);
[0077] Set the power supply output mode to voltage source (reference voltage 2.5V) or current source (reference current 1A).
[0078] 2. Real-time monitoring parameters:
[0079] Output voltage Vout: Sampled by voltage ADC, with an accuracy of ±0.1%;
[0080] Output current Iout: Sampled by current ADC, with an accuracy of ±0.1%;
[0081] Protection measurement value Vprot: The protection measurement value is the Kelvin voltage drop, collected by Kelvin high / low voltage drop ADCs, such as Kelvin high voltage drop ADC and Kelvin low voltage drop ADC, with a measurement range of 0 - 5V;
[0082] Load impedance change rate: According to Ohm's law R(t)=Iout / Vout, the controller calculates the instantaneous value of the load impedance R(t) from the real-time collected Vout and Iout. The controller performs time series analysis on R(t) and estimates the load impedance change rate dR / dt using the difference method:
[0083]
[0084] where Δt is the sampling interval, synchronized with the ADC sampling rate (if the ADC sampling rate is 1MHz, then Δt = 1μs). Filtering can also be performed on R(t) and dR / dt, such as using a 5-point moving average filter with a window size synchronized with the ADC sampling rate (e.g., a 5μs window when the sampling rate is 1MHz), Kalman filter parameters (e.g., process noise covariance Q = 0.01, observation noise covariance R = 0.1), and the Kalman filter parameters Q and R are dynamically adjusted according to the load mutation rate: when dR / dt ≥ 100Ω / μs, Q = 0.1 (increase the process noise to adapt to the mutation), R = 0.05 (reduce the weight of the observation noise); in normal scenarios, Q = 0.01 and R = 0.1.
[0085] The loop regulation error e(t) is divided into two cases:
[0086] When the alarm is not triggered, it is the difference between the power supply output value (voltage or current) and the measured power supply output value. In voltage source mode, e(t)=set target voltage value - real-time measurement value of voltage ADC; in current source mode, e(t)=set target current value - real-time measurement value of current ADC;
[0087] When the alarm is triggered, it is the difference between the protection measurement value (such as Kelvin voltage drop) and the safety set value. For example, e(t)=preset safety value (such as 0) - real-time measurement value of Kelvin high voltage drop ADC.
[0088] Thus, based on the input parameters: Kelvin voltage drop (protection measurement value Vprot), load impedance change rate (dR / dt), current PID parameters (Kp, Ki, Kd), protection measurement alarm flag bit (0 / 1), power supply output mode (voltage source / current source); the output parameters: incremental adjustment amounts of PID parameters (ΔKp, ΔKi, ΔKd).
[0089] Thus, a state vector State can be constructed, such as:
[0090] State = [Vout, Iout, Vprot, dR / dt, e(t), alarm_flag, mode]
[0091] where, alarm_flag is 0 - no alarm, 1 - alarm occurred; mode is 0 - constant voltage mode (voltage source), 1 - constant current mode (current source).
[0092] 3. Configure Kelvin loop parameters:
[0093] Set the alarm threshold of Kelvin voltage drop (such as 5V) and the safety reference value (0V);
[0094] Adjust the resolution of Kelvin ADC ( Figure 3 where the ADC referred to by Kelvin high voltage and Kelvin low voltage drop) through digital gain to match the resolution of voltage ADC (for example, both are adjusted to 12 - bit precision, or it can also be 16 - bit, 24 - bit, etc.).
[0095] 4. Enable the alarm function:
[0096] Enable the Kelvin voltage drop monitoring interrupt with the highest priority set.
[0097] (b)When the protection measurement value triggers an alarm, switch the feedback source of the digital loop from the conventional power supply output measurement value to the protection measurement value, and set the reference value to the preset safety value;
[0098] In this embodiment, when the Kelvin voltage drop exceeds 5V, the protection measurement module triggers an interrupt, and the controller immediately responds: record the current power supply output status (voltage / current, load impedance, etc.); force - stop the conventional loop operation and save the current PID parameters.
[0099] (c)Dynamically adjust the control parameters of the digital loop based on the reinforcement learning algorithm, as Figure 2 shown, specifically including:
[0100] (c1) Through the Actor network in the Deep Deterministic Policy Gradient (DDPG) algorithm, the incremental adjustment amounts of the proportional coefficient Kp, integral coefficient Ki, and derivative coefficient Kd are output according to the state vector.
[0101] Among them, the Actor network is a lightweight fully connected network (input layer with 10 nodes → hidden layer with 8 nodes → output layer with 3 nodes (corresponding to the incremental adjustment amounts of PID parameters (ΔKp, ΔKi, ΔKd))). The output layer uses the Tanh activation function (such as the output range [-1, 1], corresponding to a parameter adjustment range of ±10%). The hidden layer uses the ReLU activation function, and the weights are quantized with 8-bit fixed-point numbers. In this way, the network scale can be compressed to 21KB (input layer 10×8 + hidden layer 8×3 = 104 parameters), meeting the real-time control requirements. The output layer is mapped to the PID parameter increments, and the proportional coefficient is forcibly adjusted in combination with the alarm flag bit to achieve a fast response in emergency scenarios. For example, ΔKp = Kp × Scale Factor × tanh(Actor output value), where Scale Factor is the preset adjustment range (such as 10%).
[0102] There is also a Critic network in DDPG, which is one of the core components of the Actor-Critic architecture. Its core function is to evaluate the value (Q value) of the state-action pair and guide the Actor network to optimize the policy. In the present invention, the Critic network is used to dynamically adjust the PID parameters (such as the proportional coefficient Kp, integral coefficient Ki, and derivative coefficient Kd) to improve the control performance of the digital power loop in the protection measurement alarm scenario. When the protection measurement alarm is triggered, the Critic network predicts the long-term benefits of this action (for example: the adjustment time is reduced by 50μs, and the overshoot is reduced to less than 5%) according to the real-time state (such as the current value of the Kelvin voltage drop and the load mutation rate) and the actions of the Actor (ΔKp = +10%, ΔKi = -5%, etc.). Therefore, the network structure of the Critic network in the present invention is:
[0103] Input layer: 13 nodes (10-dimensional state vector + 3-dimensional action vector);
[0104] Hidden layer: A 16-node fully connected layer using the LeakyReLU activation function;
[0105] Output layer: 1 node, outputting the Q value (state-action value), using a linear activation function.
[0106] Among them, the learning rate of the Critic network can be 1e-3, and the soft update coefficient τ can be 0.01. The specific parameters can be set according to actual needs.
[0107] The reward function is set as:
[0108] R = -(αT s + βσ% + γe ss + λP DAC )
[0109] Where, T s is the adjustment time, σ% is the overshoot, e ss is the steady-state error, P DAC is the power consumption of the DAC output, α is the weight coefficient of the adjustment time, β is the weight coefficient of the overshoot, γ is the weight coefficient of the steady-state error, λ is the weight coefficient of the DAC output power consumption. When the protection measurement alarm is triggered, an additional penalty term R penalty = -100×T alarm , T alarm is the alarm duration. In this way, the Critic network drives the Actor network to preferentially reduce the alarm duration through the penalty term R penalty to solve the "circuit burnout risk" problem. The weight coefficients α, β, γ, λ are determined by grid search: the value range of α is 0.5 - 1.5 (weight for adjustment time optimization), β is 1.0 - 2.0 (weight for overshoot suppression), γ is 0.1 - 0.5 (weight for steady-state error), λ is 0.01 - 0.1 (weight for power consumption optimization), and the optimal combination is selected through cross-validation (such as α = 1.0, β = 1.5, γ = 0.2, λ = 0.05).
[0110] Among them, the overshoot σ (Overshoot) is a key index used to describe the dynamic response characteristics in the control system, referring to the maximum instantaneous deviation of the output value exceeding the target set value during the adjustment process of the system, usually expressed as a percentage.
[0111] In this embodiment, the Actor-Critic network can build a digital power simulation model in MATLAB / Simulink to simulate ADC sampling, DAC output and protection measurement alarm; generate an initial experience pool (100,000 data) by randomly initializing the Actor-Critic network; use batch gradient descent (Batch Size=64) to optimize the network, with learning rates Actor=1e-4 and Critic=1e-3. The initial training data generates 100,000 state-action pairs under random load mutation scenarios through MATLAB / Simulink simulation. The experience replay buffer has a capacity of 10,000, a sampling interval of 50ms, and 128 data are updated each time, where the alarm trigger event is marked as 3 times the priority coefficient; save the optimal network weights and embed them into the controller firmware. For example, in the STM32H7 controller, the CMSIS-NN library is used to optimize 8-bit quantized reasoning, and the memory usage is compressed to 21KB; the interrupt priority is set to the highest (priority 0), and the task scheduling adopts a time slice rotation strategy. The single reasoning time is measured to be 48μs (±2μs).
[0112] In this embodiment, online learning can also be used to adopt Priority Experience Replay (PER) to focus on learning alarm trigger event data and improve the response speed in emergency scenarios. The state vector, action, reward and next state are recorded in real time, and the alarm trigger event is marked as a high priority and stored in the playback buffer; 128 data are sampled from the buffer every 50ms to update the Actor-Critic network; in the alarm state, the integral coefficient adjustment range is limited (±5% or 10%) to prevent parameter mutations. Among them, during the online learning process, action limits are set: ΔKp adjustment range ≤±20%, ΔKi ≤±10%, ΔKd ≤±5%; the exploration noise uses OU process noise, and the standard deviation decays with the number of training steps (initial σ=0.1, decaying to 0.01 every 1000 steps) to avoid parameter mutations causing system oscillations.
[0113] For example, in the actual system, dR / dt=500Ω / μs is detected, triggering the Kelvin alarm; the Actor network outputs ΔKp=+100% (Kp=2.0), ΔKi=-50% (Ki=0.5), and the alarm is lifted within 50μs; PER (priority experience recovery, a means of online learning) marks the event as high priority, and the adjustment time for similar scenarios in subsequent training is optimized to 30μs.
[0114] In this way, basic performance is ensured through pre-training, and online learning improves the robustness of complex scenarios, which is seamlessly connected with the "dynamic parameter adaptation" technical feature.
[0115] (c2)When the protection measurement alarm is triggered, the proportional coefficient Kp is forced to increase to a preset multiple of the current value, and the adjustment range of the integral coefficient Ki is restricted within a preset percentage range or the integral coefficient is reduced to a preset percentage of the current value;
[0116] For example, in a scenario where a large current mutation causes extreme overshoot (overshoot amount ≥ 10% or load mutation rate ≥ 100 A / μs), when the protection measurement alarm is triggered, it is forced to switch to a preset safe parameter combination, including increasing the proportional coefficient to 2 times the current value and reducing the integral coefficient to 0.5 times the current value, which can achieve rapid suppression of overshoot and urgently reduce the protection measurement value.
[0117] In a conventional alarm scenario, in a scenario that requires smooth adjustment (overshoot amount < 10% and load mutation rate < 100 A / μs), when the protection measurement alarm is triggered, it is forced to switch to a preset safe parameter combination, including increasing the proportional coefficient to 2 times the current value and restricting the adjustment range of the integral coefficient to ±5%.
[0118] Among them, the extreme scenario determination threshold (overshoot amount ≥ 10%, load mutation rate ≥ 100 A / μs) is determined based on the power supply specification: for a power supply system with a rated current ≥ 50 A, the load mutation rate threshold is adjusted to 200 A / μs; when the alarm is triggered, the parameter combination (Kp × 2, Ki × 0.5) is forced to switch, and the stability is verified through hardware-in-the-loop testing.
[0119] In this way, the hierarchical mechanism is only enabled when the alarm is triggered, and in extreme scenarios, preset parameters are forced to switch, and in conventional scenarios, dynamic optimization is performed through reinforcement learning. By introducing a hierarchical adjustment mechanism, the emergency suppression and stable adjustment logics are separated, and the alarm trigger scenario is bound to the feedback source switching.
[0120] (c3)Update the parameters of the PID controller in real time according to the incremental adjustment amount;
[0121] (d)When the protection measurement value drops to the safe value, disconnect the digital loop feedback and cut off the physical connection between the power supply and the load.
[0122] This specific implementation fully discloses the hardware architecture, algorithm implementation, and operation process. Through the cooperation of dynamic switching of the feedback source and parameter optimization of reinforcement learning, the problems of rapid response and stability in the protection measurement alarm scenario are solved. This solution can be directly integrated into the existing digital power supply system without additional hardware costs.
[0123] Embodiment 2
[0124] Based on the same concept, the present invention also proposes an adaptive adjustment device for the loop parameters of a digital power loop, including:
[0125] The data acquisition module collects the operation status data of the digital power supply in real time, including output voltage, output current, protection measurement value, load impedance change rate, and loop regulation error, so as to construct a state vector;
[0126] The feedback switching module, when the protection measurement value triggers an alarm, switches the feedback source of the digital loop from the conventional power supply output measurement value to the protection measurement value, and sets the reference value to a preset safe value;
[0127] The reinforcement learning control module dynamically adjusts the control parameters of the digital loop based on the reinforcement learning algorithm, specifically including:
[0128] Through the Actor network in the deep deterministic policy gradient algorithm, the incremental adjustment amounts of the proportional coefficient Kp, integral coefficient Ki, and derivative coefficient Kd are output according to the state vector;
[0129] The execution module, when the protection measurement alarm is triggered, forcibly increases the proportional coefficient Kp to a preset multiple of the current value, and limits the adjustment range of the integral coefficient Ki within a preset percentage range or reduces the integral coefficient to a preset percentage of the current value; updates the parameters of the PID controller in real time according to the incremental adjustment amount; when the protection measurement value drops to the safe value, disconnects the digital loop feedback and cuts off the physical connection between the power supply and the load.
[0130] Embodiment III
[0131] This embodiment also provides an electronic device, refer to Figure 4 , including a memory 404 and a processor 402. A computer program is stored in the memory 404, and the processor 402 is configured to run the computer program to execute the steps in any one of the above method embodiments.
[0132] Specifically, the above-mentioned processor 402 may include a central processing unit (CPU), or an application specific integrated circuit (ASIC for short), or one or more integrated circuits configured to implement the embodiments of the present invention.
[0133] Among them, the memory 404 may include a mass storage 404 for data or instructions. By way of example and not limitation, the memory 404 may include a hard disk drive (HDD), a floppy disk drive, a solid state drive (SSD), a flash memory, an optical disc, a magneto-optical disc, a magnetic tape, or a universal serial bus (USB) drive, or a combination of two or more of these. Where appropriate, the memory 404 may include removable or non-removable (or fixed) media. Where appropriate, the memory 404 may be internal or external to the data processing device. In a particular embodiment, the memory 404 is non-volatile memory. In a particular embodiment, the memory 404 includes a read-only memory (ROM) and a random access memory (RAM). Where appropriate, the ROM may be a mask-programmed ROM, a programmable ROM (PROM), an erasable PROM (EPROM), an electrically erasable PROM (EEPROM), an electrically alterable ROM (EAROM), or a flash memory, or a combination of two or more of these. Where appropriate, the RAM may be a static random access memory (SRAM) or a dynamic random access memory (DRAM), where the DRAM may be a fast page mode dynamic random access memory (FPMDRAM), an extended data output dynamic random access memory (EDODRAM), a synchronous dynamic random access memory (SDRAM), etc.
[0134] The memory 404 can be used to store or cache various data files required for processing and / or communication, as well as possible computer program instructions executed by the processor 402.
[0135] The processor 402 reads and executes the computer program instructions stored in the memory 404 to implement the loop parameter adaptive adjustment method of any one of the digital power loops in the above embodiments.
[0136] Optionally, the above electronic device may further include a transmission device 406 and an input / output device 408, wherein the transmission device 406 is connected to the above processor 402, and the input / output device 408 is connected to the above processor 402.
[0137] The transmission device 406 can be used to receive or send data via a network. Specific examples of the above network may include wired or wireless networks provided by the communication provider of the electronic device. In one example, the transmission device includes a Network Interface Controller (NIC), which can be connected to other network devices through a base station and thus communicate with the Internet. In one example, the transmission device 406 can be a Radio Frequency (RF) module, which is used to communicate with the Internet wirelessly.
[0138] The input / output device 408 is used to input or output information.
[0139] Embodiment 4
[0140] This embodiment also provides a readable storage medium, in which a computer program is stored. The computer program includes program code for controlling a process to execute the process. The process includes the loop parameter adaptive adjustment method of the digital power loop according to Embodiment 1.
[0141] It should be noted that the specific examples in this embodiment can refer to the examples described in the above embodiments and optional implementation manners, and will not be elaborated herein.
[0142] Generally, various embodiments can be implemented in hardware or dedicated circuits, software, logic, or any combination thereof. Some aspects of the present invention can be implemented in hardware, while other aspects can be implemented by firmware or software executed by a controller, a microprocessor, or other computing devices, but the present invention is not limited thereto. Although various aspects of the present invention can be shown and described as block diagrams, flowcharts, or using some other graphical representations, it should be understood that, as a non-limiting example, the blocks, devices, systems, technologies, or methods described herein can be implemented in hardware, software, firmware, dedicated circuits or logic, general hardware or a controller, or other computing devices, or some combination thereof.
[0143] Embodiments of the present invention can be implemented by computer software, which can be executed by a data processor of a mobile device, such as in a processor entity, or implemented by hardware, or implemented by a combination of software and hardware. A computer software or program (also referred to as a program product), including software routines, applets, and / or macros, can be stored in any device-readable data storage medium, and they include program instructions for performing specific tasks. The computer program product can include one or more computer-executable components that are configured to execute the embodiments when the program runs. One or more computer-executable components can be at least one software code or a part thereof. Additionally, in this regard, it should be noted that any box in the logical flow, as Figure 1 described in [reference], can represent a program step, or interconnected logic circuits, boxes, and functions, or a combination of program steps and logic circuits, boxes, and functions. The software can be stored on physical media such as memory chips or storage blocks implemented within a processor, magnetic media such as hard disks or floppy disks, and optical media such as, for example, DVDs and their data variants, CDs. The physical media is a non-transitory medium.
[0144] Those skilled in the art should understand that the technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as within the scope described in this specification.
[0145] The above embodiments only represent several implementation manners of the present invention, and their descriptions are relatively specific and detailed, but they should not be construed as limiting the scope of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present invention, several modifications and improvements can still be made, and these all belong to the protection scope of the present invention. Therefore, the protection scope of the present invention should be subject to the appended claims.
Claims
1. A method for adaptively adjusting loop parameters of a digital power loop, characterized in that: The following steps are involved: (a) Real-time collection of operating status data of the digital power supply, including output voltage, output current, protection measurement value, load impedance change rate and loop regulation error, to construct a state vector; (b) when the protection measurement value triggers an alarm, switching the feedback source of the digital loop from the conventional power output measurement value to the protection measurement value, and setting the reference value to a preset safety value; (c) Dynamically adjust the control parameters of the digital loop based on the reinforcement learning algorithm, including: (c1) outputting incremental adjustments of the proportional coefficient Kp, the integral coefficient Ki and the differential coefficient Kd according to the state vector through the Actor network in the deep deterministic policy gradient algorithm; (c2) When the protection measurement alarm is triggered, the proportional coefficient Kp is forced to increase to a preset multiple of the current value, and the adjustment range of the integral coefficient Ki is limited to a preset percentage range or the integral coefficient is reduced to a preset percentage of the current value; (c3) updating the parameters of the PID controller in real time according to the incremental adjustment amount; (d) When the protection measurement value drops to the safety value, the digital loop feedback is disconnected and the physical connection between the power supply and the load is cut off.
2. A method for adaptively adjusting loop parameters of a digital power loop as claimed in claim 1, characterized in that: The reward function of the reinforcement learning algorithm described in step (c) is: R=-(αT s +βσ%+γe ss +λP DAC ) Among them, T s is the adjustment time, σ% is the overshoot, e ss is the steady-state error, P DAC is the DAC output power consumption, α is the weight coefficient of the adjustment time, β is the weight coefficient of the overshoot, γ is the weight coefficient of the steady-state error, and λ is the weight coefficient of the DAC output power consumption; When the protection measurement alarm is triggered, an additional penalty term R is imposed penalty = -100 × T alarm , T alarm The alarm duration.
3. The method for adaptively adjusting loop parameters of a digital power loop according to claim 1, characterized in that: The Actor network described in step (c1) is a lightweight fully connected network, including an input layer, a hidden layer and an output layer. The number of nodes in the input layer is the dimension of the state vector, the number of nodes in the output layer is 3, corresponding to the proportional coefficient adjustment ΔKp, the integral coefficient adjustment ΔKi, the differential coefficient adjustment ΔKd, and the network weights are quantized using 8-bit fixed-point numbers.
4. The method for adaptively adjusting loop parameters of a digital power loop according to claim 1, characterized in that: In step (c2), when the protection measurement alarm is triggered, the system is forced to switch to the preset safety parameter combination, including increasing the proportional coefficient to twice the current value and limiting the adjustment range of the integral coefficient to ±5%.
5. The method for adaptively adjusting loop parameters of a digital power loop according to claim 1, characterized in that: In step (c2), when the protection measurement alarm is triggered, the preset safety parameter combination is forced to be switched, including increasing the proportional coefficient to 2 times the current value and reducing the integral coefficient to 0.5 times the current value.
6. The method for adaptively adjusting loop parameters of a digital power loop according to claim 1, characterized in that: In step (c1), the state vector also includes: Protection measurement alarm trigger flag; Power output mode.
7. A method for adaptively adjusting loop parameters of a digital power loop according to any one of claims 1 to 6, characterized in that: Also includes: The reinforcement learning algorithm is deployed in an embedded controller to ensure that the control cycle is synchronized with the digital loop and the single inference time is ≤50 microseconds.
8. A loop parameter adaptive adjustment device for a digital power loop, characterized in that: include: The data acquisition module collects the operating status data of the digital power supply in real time, including output voltage, output current, protection measurement value, load impedance change rate and loop regulation error, so as to construct a state vector; A feedback switching module, when the protection measurement value triggers an alarm, switches the feedback source of the digital loop from the conventional power supply output measurement value to the protection measurement value, and sets the reference value to a preset safety value; The reinforcement learning control module dynamically adjusts the control parameters of the digital loop based on the reinforcement learning algorithm, including: Outputting incremental adjustments of the proportional coefficient Kp, the integral coefficient Ki, and the differential coefficient Kd according to the state vector through the Actor network in the deep deterministic policy gradient algorithm; The execution module, when the protection measurement alarm is triggered, forces the proportional coefficient Kp to be increased to a preset multiple of the current value, and limits the adjustment range of the integral coefficient Ki to a preset percentage range or reduces the integral coefficient to a preset percentage of the current value; updates the parameters of the PID controller in real time according to the incremental adjustment amount; when the protection measurement value drops to the safety value, disconnects the digital loop feedback and cuts off the physical connection between the power supply and the load.
9. An electronic device comprising a memory and a processor, characterized in that: A computer program is stored in the memory, and the processor is configured to run the computer program to execute the loop parameter adaptive adjustment method of a digital power loop according to any one of claims 1 to 7.
10. A readable storage medium, characterized in that: A computer program is stored in the readable storage medium, wherein the computer program includes a program code for controlling a process to execute a process, wherein the process includes the method for adaptively adjusting loop parameters of a digital power loop according to any one of claims 1 to 7.
Citation Information
Patent Citations
Shared integral term PID dual-closed-loop controller for PWM digital power supply
CN107104593A
Adaptive reconfigurable proportional-integral-differential controller based on BP neural network
CN113641096A