Wide-area flexible power supply adaptive control method based on reinforcement learning

By using the GoSafeOpt security exploration mechanism based on reinforcement learning and particle-style updates, voltage targets, current limiting thresholds, ripple compensation parameters, and energy buffer thresholds are generated. This solves the stability and security issues of power supply systems for information and communication equipment in multi-device integrated application scenarios, achieves rapid adaptive control and parameter convergence, and improves the stability and security of the power supply system.

CN122052047APending Publication Date: 2026-05-15TIANJIN JINLI TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
TIANJIN JINLI TECHNOLOGY CO LTD
Filing Date
2026-02-26
Publication Date
2026-05-15

AI Technical Summary

Technical Problem

Existing power supply systems for information and communication equipment struggle to balance performance and safety goals under wide-range input and load conditions in multi-device integrated application scenarios. Furthermore, they lack effective adaptive control methods, resulting in weak parameter mobility, long adaptation cycles, and a high risk of overvoltage, overcurrent, or excessive temperature rise.

Method used

The GoSafeOpt safety exploration mechanism based on reinforcement learning and particle-style updates are adopted to generate voltage targets, current limiting thresholds, ripple compensation parameters and energy buffer thresholds. By correcting the boundary reflection projection, candidate points are prevented from going out of bounds, forming a fixed control parameter vector for different load states, thus achieving control with high safety, fast parameter convergence and strong wide-range adaptability.

Benefits of technology

It achieves integrated constraints on transient and steady-state processes, improves power supply stability and constraint satisfaction consistency, reduces the risk of overvoltage, overcurrent and temperature rise exceeding limits, shortens the adaptation cycle, and improves the reusability and migration efficiency of control parameters in multi-device integrated applications.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122052047A_ABST
    Figure CN122052047A_ABST
Patent Text Reader

Abstract

The invention discloses a reinforcement learning-based wide-area flexible power supply adaptive control method. The method comprises the following steps of 1, collecting power supply observed quantity to generate an observation sequence; 2, generating a working condition context vector according to the observation sequence, reading a historical working condition parameter set, and initializing a security agent model, a security feasible region and a sample set; 3, setting a control parameter vector and establishing control mapping; 4, setting a security constraint index set; 5, on the basis of the working condition context vector, the security agent model, the security feasible region and the sample set, adopting a GoSafeOpt particle form to update and generate a candidate control parameter vector; step 6, executing boundary reflection projection correction on the boundary crossing candidate points to obtain security candidate control parameter vectors; and step 7, issuing execution acquisition output, calculating a security constraint index set, updating the security agent model and the security feasible region, and outputting a solidification control parameter vector. The power supply stability and the adaptation efficiency are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of power supply technology for information and communication equipment, and in particular to a wide-area flexible power supply adaptive control method based on reinforcement learning. Background Technology

[0002] Existing power supply systems for information and communication equipment are typically designed for fixed input voltage ranges and relatively stable load conditions. Control strategies often employ threshold comparison, tiered switching, fixed compensation parameters, and preset current-limiting logic to achieve voltage regulation, current limiting, and protection. However, in multi-device integrated applications, the power supply needs to adapt to the rated voltage, current, and dynamic load characteristics of different devices simultaneously, while the output must maintain stability under constraints such as ripple, transient overshoot, and transient sag. To accommodate these differences, existing solutions often rely on manual experience to repeatedly adjust parameters between voltage targets, current-limiting thresholds, compensation parameters, and buffer thresholds, and increase safety margins to prevent exceeding limits. This results in long adaptation cycles, weak parameter migration, and a difficulty in balancing performance and safety objectives under wide input and load conditions.

[0003] In dynamic load scenarios, abrupt changes in load conditions can cause transient overshoot and sag in the output voltage. The parameter settings of the ripple suppression unit and energy buffer unit significantly affect the transient recovery process. Existing technologies mainly focus on single-point thresholds or steady-state indicators for safety performance constraints, lacking process constraint expressions oriented towards preset time windows. This makes it difficult to simultaneously cover the joint constraint requirements of voltage overshoot peak, voltage sag peak, ripple peak, recovery time, and temperature rise peak. Some solutions rely on offline simulation or static margin configuration, failing to form a unified observation sequence of output voltage, output current, temperature, ripple sequence, transient response sequence, and load state for use in operating condition characterization. This results in the inability to stably identify constraint triggering mechanisms under complex operating conditions, leading to significant changes in the safety boundary of the same set of parameters under different load states.

[0004] Furthermore, the application of learning-based methods for adaptive parameter tuning in power supply systems remains limited by safety and sample efficiency. During the exploration process, if control parameters exceed limits, it may lead to risks of overvoltage, overcurrent, or excessive temperature rise. Existing technologies lack a safety exploration mechanism that explicitly incorporates the safety agent model and the safe feasible region into the search process. They also lack a safety determination and candidate point selection process based on confidence boundaries, and a mechanism for boundary projection and reflection correction when candidate points exceed limits. This results in a need to conservatively narrow the search range for parameter searches, making it difficult to balance global optimization with safety constraints. For repeated adaptation processes involving multiple operating conditions and multiple devices, existing technologies lack a closed-loop process for constructing a sample set based on historical operating condition parameter sets and iteratively updating the safety agent model and safe feasible region. This means that each adaptation still requires starting from scratch, making it difficult to quickly converge to a stable, fixed control parameter vector while meeting the threshold of the safety constraint index set.

[0005] Therefore, how to provide a wide-range flexible power supply adaptive control method based on reinforcement learning is a problem that urgently needs to be solved by those skilled in the art. Summary of the Invention

[0006] One objective of this invention is to propose a wide-domain flexible power supply adaptive control method based on reinforcement learning. This invention utilizes the GoSafeOpt safe exploration mechanism and particle-style updates to adaptively generate voltage targets, current limiting thresholds, ripple compensation parameters, and energy buffer thresholds under the constraint of the safe feasible domain. It also avoids candidate points from going out of bounds through boundary reflection projection correction, forming a fixed control parameter vector for different load states. This method has the advantages of high safety, fast parameter convergence, and strong wide-domain adaptability.

[0007] A wide-area flexible power supply adaptive control method based on reinforcement learning according to an embodiment of the present invention includes the following steps: Step 1: Collect input voltage, output voltage, output current, temperature, ripple sequence, transient response sequence, and load status to generate an observation sequence; Step 2: Generate operating condition context vectors from the observation sequence, read the historical operating condition parameter set, and initialize the safety agent model, safe feasible region, and sample set; Step 3: Set the control parameter vector, which includes the voltage target, current limiting threshold, ripple compensation parameters, and energy buffer threshold. Establish the control parameter vector and the control mapping between the voltage regulation unit, current control unit, ripple suppression unit, and buffer unit. Step 4: Set the safety constraint index set, which includes the voltage overshoot peak value, voltage drop peak value, ripple peak value, recovery time, and temperature rise peak value within a preset time window; Step 5: Based on the operating context vector, safety agent model, safe feasible region and sample set, use GoSafeOpt to generate candidate control parameter vectors, and use particle-style update to generate candidate points; Step 6: Perform boundary reflection projection correction on candidate points that do not meet the safety feasible region to obtain the safety candidate control parameter vector; Step 7: Send the safety candidate control parameter vector to the wide-domain adaptive power supply module for execution, collect the output response, calculate the safety constraint index set, write it into the sample set and update the safety proxy model and the safety feasible region, and output the fixed control parameter vector.

[0008] Optionally, step one specifically includes: Collect sampling sequences of input voltage, output voltage, output current, and temperature, and set sampling timestamps for the input voltage sampling sequence, output voltage sampling sequence, output current sampling sequence, and temperature sampling sequence; Ripple extraction is performed on the output voltage sampling sequence to generate a ripple sequence; transient response extraction is performed on the output voltage sampling sequence and the output current sampling sequence to generate a transient response sequence; and load state is calculated based on the output voltage sampling sequence, the output current sampling sequence, and the temperature sampling sequence to generate a load state sequence. The input voltage sampling sequence, output voltage sampling sequence, output current sampling sequence, temperature sampling sequence, ripple sequence, transient response sequence, and load state sequence are aligned and spliced ​​according to the sampling timestamps to generate the observation sequence.

[0009] Optionally, step two specifically includes: The operating condition feature vector is calculated based on the observation sequence. The operating condition feature vector includes input voltage features, output voltage features, output current features, temperature features, ripple features, transient response features, and load state features. The operating condition feature vector is normalized and encoded to obtain the operating condition context vector. Read the historical operating condition parameter set, which includes historical control parameter vectors, historical safety constraint index sets, and historical target indexes, and write the historical control parameter vectors, historical safety constraint index sets, and historical target indexes into the sample set; The safety agent model is trained based on the sample set. The safety agent model takes the control parameter vector and the operating condition context vector as input and outputs the estimated value of the safety constraint index set and the confidence bound of the safety constraint index set. Based on the security agent model, the lower confidence bound of the security constraint index set is calculated for the control parameter vector. The control parameter vector that satisfies the lower confidence bound of the security constraint index set threshold is defined as the safe feasible region.

[0010] Optionally, step three specifically includes: Set a control parameter vector, which includes voltage target, current limiting threshold, ripple compensation parameter and energy buffer threshold, and set the value range for voltage target, current limiting threshold, ripple compensation parameter and energy buffer threshold; Establish a control mapping, which maps the control parameter vector to voltage regulation unit parameters, current control unit parameters, ripple suppression unit parameters, and buffer unit parameters. The voltage target corresponds to the voltage regulation unit parameters, the current limiting threshold corresponds to the current control unit parameters, the ripple compensation parameters correspond to the ripple suppression unit parameters, and the energy buffer threshold corresponds to the buffer unit parameters. Based on the control mapping to generate the power supply module parameter set, the wide-domain adaptive power supply module drives the voltage regulation unit, current control unit, ripple suppression unit and buffer unit to perform power supply output according to the power supply module parameter set.

[0011] Optionally, step four specifically involves: Set a preset time window. The start timestamp of the preset time window is taken from the transition trigger point timestamp of the transient response sequence. The preset time window length is the preset time length. Within a preset time window, calculate the voltage overshoot peak, voltage drop peak, ripple peak peak, and temperature rise peak. The voltage overshoot peak is the maximum positive deviation of the output voltage sampling sequence relative to the voltage target. The voltage drop peak is the maximum negative deviation of the output voltage sampling sequence relative to the voltage target. The ripple peak peak is the maximum value of the ripple sequence minus the minimum value of the ripple sequence. The temperature rise peak is the maximum value of the temperature sampling sequence minus the temperature sampling value corresponding to the start timestamp of the preset time window. The recovery time is calculated within a preset time window. The recovery time is the shortest time from the start timestamp of the preset time window until the output voltage sampling sequence enters the deviation threshold range and remains continuously within the deviation threshold range. The deviation threshold range is determined by the threshold of the safety constraint index set.

[0012] Optionally, step five specifically includes: The target indicator proxy model is trained based on the sample set. The target indicator proxy model takes the control parameter vector and the operating condition context vector as input, and the target indicator recorded in the sample set as the supervised output. It outputs the target indicator estimate and the target indicator uncertainty, and calculates the upper confidence bound of the target indicator based on the target indicator estimate and the target indicator uncertainty. The safety agent model is updated based on the sample set. The safety agent model takes the control parameter vector and the operating condition context vector as inputs and outputs the estimated value of the safety constraint index set and the uncertainty of the safety agent model. The lower confidence bound of the safety constraint index set is calculated based on the estimated value of the safety constraint index set and the uncertainty of the safety agent model. Initialize the candidate point set within the range of control parameter vector values. The candidate point set contains particle position vector and particle velocity vector. The particle position vector components correspond one-to-one with voltage target, current limiting threshold, ripple compensation parameter, and energy buffer threshold. Establish the historical optimal position vector and global optimal position vector for the candidate point set. For the candidate point set, calculate the upper confidence bound of the target index and the lower confidence bound of the safety constraint index set. Based on the lower confidence bound of the safety constraint index set satisfying the threshold of the safety constraint index set, select the safe candidate point set and update the global optimal position vector by selecting the particle position vector corresponding to the largest upper confidence bound of the target index in the safe candidate point set. GoSafeOpt is used to perform particle-based updates on the safety candidate point set. The particle-based update includes updating the particle position vector based on the particle velocity vector, the particle's historical best position vector, and the global best position vector. The particle position vector is subjected to interval constraint processing according to the value range of the control parameter vector. The upper confidence bound of the target index and the lower confidence bound of the safety constraint index set corresponding to the updated particle position vector are recalculated. The historical best position vector of the particle is updated according to the upper confidence bound of the target index, and the global best position vector is updated according to the safety candidate point set. When the iteration reaches the iteration termination condition, the control parameter vector corresponding to the global best position vector is determined as the candidate control parameter vector.

[0013] Optionally, step six specifically includes: Perform a safe and feasible region determination on the control parameter vectors corresponding to the candidate points in the candidate point set. Control parameter vectors that are determined not to belong to the safe and feasible region are marked as out-of-bounds control parameter vectors. For the out-of-bounds control parameter vector, the boundary projection point is determined. The boundary projection point belongs to the safe and feasible region. The distance between the boundary projection point and the out-of-bounds control parameter vector is the minimum distance, which is determined by the sum of the squares of the differences between the components of the control parameter vector. The boundary projection point is used as the center of symmetry to reflect the out-of-bounds control parameter vector to obtain the reflected control parameter vector. The value range of the reflected control parameter vector is clipped to obtain the corrected control parameter vector. The safe and feasible region is determined on the corrected control parameter vector. If it is determined that it does not belong to the safe and feasible region, the boundary projection point determination and reflection are re-executed using the corrected control parameter vector as the out-of-bounds control parameter vector to obtain the safe candidate control parameter vector.

[0014] Optionally, step seven specifically includes: The safety candidate control parameter vector is input into the wide-domain adaptive power supply module. The wide-domain adaptive power supply module generates a power supply module parameter set based on the control mapping. The power supply module parameter set includes voltage regulation unit parameters, current control unit parameters, ripple suppression unit parameters, and buffer unit parameters. The wide-domain adaptive power supply module performs power supply output based on the power supply module parameter set. The transition trigger point timestamp of the transient response sequence is used as the start timestamp of the preset time window, and the preset time window is determined based on the preset time length. The output voltage sampling sequence, output current sampling sequence, and temperature sampling sequence within the preset time window are collected to form the output response. The ripple sequence is extracted based on the output voltage sampling sequence, and the transient response sequence is extracted based on the output voltage sampling sequence and the output current sampling sequence. The safety constraint index set is calculated based on the output response, and the target index is calculated. The safety constraint index set includes voltage overshoot peak value, voltage drop peak value, ripple peak value, recovery time, and temperature rise peak value. The candidate control parameter vector, operating condition context vector, target index, and safety constraint index set are written into the sample set to form sample entries. The safety agent model is updated based on the sample set, and the safety feasible region is updated based on the safety agent model. The fixed control parameter vector is determined from the sample set based on the target index.

[0015] The beneficial effects of this invention are: This invention constructs a working condition context vector through observation sequences and introduces a closed-loop update mechanism for a safety proxy model, a safe feasible domain, and a sample set. This transforms wide-range adaptive power supply control from empirical parameter tuning to an iterative safety exploration process. In each power supply output, the safety constraint index set covers voltage overshoot peak, voltage drop peak, ripple peak, recovery time, and temperature rise peak within a preset time window. This achieves integrated constraints on transient and steady-state processes, enabling the parameter co-optimization of the voltage regulation unit, current control unit, ripple suppression unit, and buffer unit to be performed within a clear safety boundary. This improves power supply stability and constraint satisfaction consistency under wide input and wide load ranges.

[0016] This invention employs GoSafeOpt combined with particle-based updates to generate candidate control parameter vectors. Boundary projection and reflection correction are performed on out-of-bounds candidate points, automatically mapping control parameter vectors that do not meet the safe feasible region back to the safe region. This reduces the risk of overvoltage, overcurrent, and temperature rise exceeding limits during the exploration process, while maintaining the ability to search for the globally optimal parameter region. Continuous updates to the sample set and the safety proxy model gradually converge the safe feasible region. The fixed control parameter vectors can adaptively select based on changes in the operating context vector, shortening the adaptation cycle under different load conditions and improving the reusability and migration efficiency of control parameters in multi-device integrated applications. Attached Figure Description

[0017] The accompanying drawings are provided to further illustrate the invention and form part of the specification. They are used in conjunction with embodiments of the invention to explain the invention and do not constitute a limitation thereof. In the drawings: Figure 1 This is a flowchart of a wide-area flexible power supply adaptive control method based on reinforcement learning proposed in this invention; Figure 2 This is a schematic diagram illustrating the GoSafeOpt security exploration of a wide-area flexible power supply adaptive control method based on reinforcement learning proposed in this invention. Detailed Implementation

[0018] The present invention will now be described in further detail with reference to the accompanying drawings. These drawings are simplified schematic diagrams, illustrating only the basic structure of the invention, and therefore only show the components relevant to the invention.

[0019] refer to Figures 1-2 A wide-area flexible power supply adaptive control method based on reinforcement learning includes the following steps: Step 1: Collect input voltage, output voltage, output current, temperature, ripple sequence, transient response sequence, and load status to generate an observation sequence; Step 2: Generate operating condition context vectors from the observation sequence, read the historical operating condition parameter set, and initialize the safety agent model, safe feasible region, and sample set; Step 3: Set the control parameter vector, which includes the voltage target, current limiting threshold, ripple compensation parameters, and energy buffer threshold. Establish the control parameter vector and the control mapping between the voltage regulation unit, current control unit, ripple suppression unit, and buffer unit. Step 4: Set the safety constraint index set, which includes the voltage overshoot peak value, voltage drop peak value, ripple peak value, recovery time, and temperature rise peak value within a preset time window; Step 5: Based on the operating context vector, safety agent model, safe feasible region and sample set, use GoSafeOpt to generate candidate control parameter vectors, and use particle-style update to generate candidate points; Step 6: Perform boundary reflection projection correction on candidate points that do not meet the safety feasible region to obtain the safety candidate control parameter vector; Step 7: Send the safety candidate control parameter vector to the wide-domain adaptive power supply module for execution, collect the output response, calculate the safety constraint index set, write it into the sample set and update the safety proxy model and the safety feasible region, and output the fixed control parameter vector.

[0020] Step one is as follows: The system collects sampling sequences of input voltage, output voltage, output current, and temperature. Each sampling sequence consists of a sampled value and a sampling timestamp. The sampling timestamp is taken from a unified clock source. The input voltage sampling sequence, output voltage sampling sequence, output current sampling sequence, and temperature sampling sequence use the same time base to record the sampling timestamp. Ripple extraction is performed on the output voltage sampling sequence. Ripple extraction includes estimating the baseline component of the output voltage sampling sequence and subtracting the baseline component from the output voltage sampling sequence. The baseline component is obtained by smoothing filtering. The subtraction result is arranged according to the sampling timestamp to form a ripple sequence. Transient response extraction is performed on the output voltage sampling sequence and the output current sampling sequence. Transient response extraction includes calculating the difference between adjacent sampling points and locating the transition trigger point based on the difference threshold. At each transition trigger point, the output voltage sampling value and the output current sampling value within a preset length time window are extracted and spliced ​​in the order of timestamps to form a transient response sequence. The load state is calculated based on the output voltage sampling sequence, the output current sampling sequence, and the temperature sampling sequence. The load state calculation includes calculating the power index and the equivalent load index and introducing the temperature index to form the load state value. The load state value is arranged according to the sampling timestamp to form a load state sequence. The input voltage sampling sequence, output voltage sampling sequence, output current sampling sequence, temperature sampling sequence, ripple sequence, transient response sequence, and load state sequence are aligned according to the sampling timestamps. Alignment includes sorting the sampling timestamps and resampling each sequence to a unified set of sampling timestamps. The resampling uses interpolation to fill in the sampling values ​​corresponding to missing sampling timestamps. The aligned input voltage sampling values, output voltage sampling values, output current sampling values, temperature sampling values, ripple values, transient response values, and load state values ​​are then concatenated according to the sampling timestamp order to generate the observation sequence.

[0021] Step two is as follows: The operating condition feature vector is calculated based on the observation sequence. The operating condition feature vector is calculated from the input voltage sampling sequence, output voltage sampling sequence, output current sampling sequence, temperature sampling sequence, ripple sequence, transient response sequence, and load state sequence. The input voltage feature, output voltage feature, output current feature, temperature feature, ripple feature, transient response feature, and load state feature are respectively composed of the mean, standard deviation, extreme value amplitude, and rate of change index of the corresponding sequence within a preset statistical window. Normalization and encoding are performed on the operating condition feature vector. Normalization uses the feature mean and feature variance provided by the historical operating condition parameter set to linearly scale each component of the operating condition feature vector. Encoding uses a fixed-dimensional mapping to map the normalized operating condition feature vector into an operating condition context vector. Read the historical operating condition parameter set, which includes historical control parameter vectors, historical safety constraint index sets, and historical target indexes. After associating the historical control parameter vectors, historical safety constraint index sets, and historical target indexes according to the record identifier, write them into the sample set. The sample entries in the sample set include control parameter vectors, operating condition context vectors, safety constraint index sets, and target indexes. The safety agent model is trained based on the sample set. The safety agent model takes the control parameter vector and the operating condition context vector as input and the safety constraint index set as the supervision output. The training process uses regression learning to obtain the estimation function of the safety constraint index set, and gives the confidence bound of the safety constraint index set in the regression output. The confidence bound of the safety constraint index set is jointly determined by the sample set residual statistics and the model uncertainty. Based on the security agent model, the lower confidence bound of the security constraint index set is calculated for the control parameter vector. The lower confidence bound of the security constraint index set is obtained by subtracting the model uncertainty from the estimated value of the security constraint index set according to the confidence level. The control parameter vector that satisfies the lower confidence bound of the security constraint index set threshold is defined as the safe feasible region.

[0022] Step three specifically involves: Set a control parameter vector, which includes voltage target, current limiting threshold, ripple compensation parameter and energy buffer threshold. Set the value range for voltage target, current limiting threshold, ripple compensation parameter and energy buffer threshold. The value range is determined by the rated input range, rated output range, rated temperature range of device, adjustable range of ripple suppression unit and adjustable range of buffer unit of wide-range adaptive power supply module. A control mapping is established, which uses parameter correspondence to map the control parameter vector to voltage regulation unit parameters, current control unit parameters, ripple suppression unit parameters, and buffer unit parameters. The voltage target is mapped to the target voltage setting parameter of the voltage regulation unit, the current limiting threshold is mapped to the current limiting setting parameter of the current control unit, the ripple compensation parameter is mapped to the compensation setting parameter of the ripple suppression unit, and the energy buffer threshold is mapped to the charge and discharge threshold setting parameter of the buffer unit. The unit parameters output by the control mapping are composed of parameter records in a fixed field order. The power supply module parameter set is generated based on the control mapping. The power supply module parameter set is composed of unit parameter records and establishes a one-to-one correspondence with the control parameter vector. The wide-domain adaptive power supply module drives the voltage regulation unit, current control unit, ripple suppression unit and buffer unit to perform power supply output according to the power supply module parameter set. During the power supply output process, the unit working status is updated by using the unit parameter records in the power supply module parameter set.

[0023] Step four is as follows: Set a preset time window. The start timestamp of the preset time window is taken from the transition trigger point timestamp of the transient response sequence. The transition trigger point timestamp is output by the transient response extraction process and used as the event marker of the transient response sequence. The preset time window length is a preset time length, which is given by the configuration parameters and recorded in association with the operating condition context vector. Within a preset time window, the peak voltage overshoot, peak voltage drop, peak ripple, and peak temperature rise are calculated. The output voltage sampling sequence within the preset time window is taken from the set of sampling points of the output voltage sampling sequence from the start time stamp to the end time stamp of the preset time window. The voltage target is taken from the voltage target in the control parameter vector. The peak voltage overshoot is calculated by subtracting the output voltage sampling value within the preset time window from the voltage target to obtain the deviation sequence and taking the maximum positive value of the deviation sequence. The peak voltage drop is calculated by taking the minimum negative value of the deviation sequence and taking the absolute value. The ripple sequence is taken from the set of sampling points of the ripple sequence obtained by ripple extraction within the preset time window. The peak ripple is obtained by calculating the difference between the maximum value and the minimum value of the ripple sequence within the preset time window. The temperature sampling sequence is taken from the set of sampling points of the temperature sampling sequence within the preset time window. The peak temperature rise is obtained by calculating the difference between the maximum value of the temperature sampling value within the preset time window and the temperature sampling value corresponding to the start time stamp of the preset time window. The recovery time is calculated within a preset time window. The calculation of the recovery time is based on the deviation sequence. The deviation threshold range is determined by the threshold of the safety constraint index set. The deviation threshold range corresponds to the upper and lower limits of the output voltage deviation. The recovery time is obtained by searching for the timestamp position in the deviation sequence where the first time the deviation threshold range is entered and the timetamp position is maintained continuously within the deviation threshold range for a preset duration. The recovery time is the time difference between the timestamp of the entry time and the timestamp of the preset time window start time.

[0024] Step five is as follows: The target indicator proxy model is trained based on the sample set. The target indicator recorded in the sample set is associated with the control parameter vector and the operating condition context vector. The target indicator proxy model uses regression learning to establish a mapping function from the control parameter vector and the operating condition context vector to the target indicator. The target indicator estimate is taken from the prediction output of the target indicator proxy model on the control parameter vector and the operating condition context vector. The target indicator uncertainty is jointly determined by the sample set residual distribution and the model variance. The upper confidence bound of the target indicator is obtained by synthesizing the target indicator estimate and the target indicator uncertainty at a preset confidence level. The safety agent model is updated based on the sample set. The safety constraint index set recorded in the sample set is associated with the control parameter vector and the operating condition context vector. The safety agent model uses regression learning to establish a mapping function from the control parameter vector, the operating condition context vector to the safety constraint index set. The estimated value of the safety constraint index set is taken from the predicted output of the safety agent model on the control parameter vector and the operating condition context vector. The uncertainty of the safety agent model is jointly determined by the residual distribution of the sample set and the model variance. The lower confidence bound of the safety constraint index set is synthesized by the estimated value of the safety constraint index set and the uncertainty of the safety agent model under a preset confidence level. Initialize a candidate point set within the range of control parameter vector values. The candidate point set contains particle position vector and particle velocity vector. The particle position vector components correspond one-to-one with voltage target, current limiting threshold, ripple compensation parameter, and energy buffer threshold. The particle velocity vector components correspond one-to-one with particle position vector components. The historical best position vector of the particle is assigned by the position vector initialized from the candidate point set. The global best position vector is assigned by the particle position vector corresponding to the largest confidence boundary of the target index within the candidate point set. For the candidate point set, calculate the upper confidence bound of the target index and the lower confidence bound of the safety constraint index set. The safety judgment result is obtained by comparing the lower confidence bound of the safety constraint index set with the threshold of the safety constraint index set component by component. The particle position vectors that meet the conditions of the safety judgment result constitute the safety candidate point set. The global optimal position vector is updated by searching for the particle position vector corresponding to the maximum upper confidence bound of the target index in the safety candidate point set. GoSafeOpt is used to perform particle-based updates on the safety candidate point set. The particle-based update includes calculating the velocity increment based on the particle's historical best position vector and the global best position vector and updating the particle velocity vector; updating the particle position vector based on the updated particle velocity vector; interval constraint processing is performed by pruning the particle position vector components to the range of control parameter vector values ​​to obtain constrained particle position vectors; recalculating the upper confidence bound of the target index and the lower confidence bound of the safety constraint index set for the constrained particle position vectors; updating the historical best position vector by comparing the upper confidence bound of the target index corresponding to the constrained particle position vector with the upper confidence bound of the target index corresponding to the historical best position vector; updating the global best position vector by searching for the constrained particle position vector corresponding to the maximum upper confidence bound of the target index in the safety candidate point set; and determining the control parameter vector corresponding to the global best position vector as a candidate control parameter vector when the iteration termination condition is met.

[0025] This invention introduces contextual condition modeling and a dual-agent confidence bound collaboration mechanism oriented towards power supply conditions into the GoSafeOpt safety exploration framework. It adopts upper confidence bounds of target indicators to drive the search and lower confidence bounds of safety constraint indicator sets to perform safety judgments, achieving the same evaluation caliber for target optimization and safety constraints. At the same time, it establishes a one-to-one correspondence between particle position vectors and voltage targets, current limiting thresholds, ripple compensation parameters, and energy buffer thresholds, and completes constraint processing through interval pruning, so that the generation and updating of candidate points always remain within the executable parameter space. Combined with the update rules of the particle's historical optimal position vector and global optimal position vector, it improves search stability and reduces sensitivity to the initial point, thereby accelerating convergence speed, reducing invalid exploration, and reducing the risks of trigger voltage overshoot, voltage drop, ripple over-limit, and temperature rise over-limit under wide-domain input and load change conditions.

[0026] Step six specifically involves: For the control parameter vectors corresponding to candidate points in the candidate point set, a safe and feasible region determination is performed. The safe and feasible region determination is based on the calculation of the lower confidence bound of the safety constraint index set by the safety proxy model and compared with the threshold of the safety constraint index set component by component. The control parameter vectors that meet the conditions of the comparison result are determined to belong to the safe and feasible region. The control parameter vectors that do not meet the conditions of the comparison result are determined to not belong to the safe and feasible region and are marked as out-of-bounds control parameter vectors. For out-of-bounds control parameter vectors, boundary projection points are determined. The boundary projection points are within the value range of the control parameter vectors and satisfy the safe and feasible region determination. The boundary projection points are obtained by searching within the value range of the control parameter vectors. The search is based on the out-of-bounds control parameter vectors and uses the safe and feasible region determination as a constraint. The objective function is to minimize the distance, which is determined by the sum of the squares of the differences between the components of the control parameter vectors. The reflected control parameter vector is obtained by reflecting the out-of-bounds control parameter vector with the boundary projection point as the center of symmetry. The reflected control parameter vector is obtained by taking the inverse of the component difference between the out-of-bounds control parameter vector and the boundary projection point and adding it back to the boundary projection point. The value range of the reflected control parameter vector is clipped to obtain the corrected control parameter vector. The clipping restricts each component of the corrected control parameter vector to within the value range of the control parameter vector. The safe and feasible region is determined for the corrected control parameter vector. If it is determined that it does not belong to the safe and feasible region, the boundary projection point determination and reflection are re-executed using the corrected control parameter vector as the out-of-bounds control parameter vector to obtain the safe candidate control parameter vector.

[0027] Step seven is as follows: The safety candidate control parameter vector is input into the wide-domain adaptive power supply module. The wide-domain adaptive power supply module generates a power supply module parameter set based on the control mapping. The power supply module parameter set includes voltage regulation unit parameters, current control unit parameters, ripple suppression unit parameters, and buffer unit parameters in the order of fields. The wide-domain adaptive power supply module updates the power supply module parameter set to the unit operating parameters. The voltage regulation unit updates the target voltage based on the voltage regulation unit parameters and performs regulated output. The current control unit updates the current limiting threshold based on the current control unit parameters and performs current limiting control. The ripple suppression unit updates the compensation intensity based on the ripple suppression unit parameters and performs ripple compensation. The buffer unit updates the charge and discharge threshold based on the buffer unit parameters and performs energy buffering. Using the transition trigger point timestamp of the transient response sequence as the start timestamp of the preset time window and determining the preset time window based on the preset time length, the output voltage sampling sequence, output current sampling sequence, and temperature sampling sequence are acquired within the preset time window and form the output response. The ripple sequence is obtained by performing ripple extraction on the output voltage sampling sequence and intercepting the preset time window ripple sequence within the preset time window. The transient response sequence is obtained by performing transient response extraction on the output voltage sampling sequence and output current sampling sequence and intercepting the preset time window transient response sequence within the preset time window. Safety constraints are calculated based on the output response. The target indicators are calculated by combining the indicator set. The voltage overshoot peak value is the maximum positive deviation of the output voltage sample value relative to the voltage target within the preset time window. The voltage drop peak value is the maximum negative deviation of the output voltage sample value relative to the voltage target within the preset time window. The ripple peak value is the maximum value of the ripple sequence within the preset time window minus the minimum value of the ripple sequence within the preset time window. The temperature rise peak value is the maximum value of the temperature sample value within the preset time window minus the temperature sample value corresponding to the start timestamp of the preset time window. The recovery time is the shortest time from the start timestamp of the preset time window until the output voltage sample value enters the deviation threshold range and continuously remains within the deviation threshold range. The candidate control parameter vector, operating condition context vector, target index, and safety constraint index set are written into the sample set to form sample entries. The sample entries establish association records between the control parameter vector and the safety constraint index set and the target index, and establish association records between the operating condition context vector and the sample entries. The safety agent model is updated by incrementally training the sample set to update the safety agent model parameters and output the estimated value of the safety constraint index set and the lower confidence bound of the safety constraint index set. The safe feasible region is updated by comparing the lower confidence bound of the safety constraint index set with the threshold of the safety constraint index set component by component to obtain the safe feasible region member determination rule. The solidified control parameter vector is obtained by searching within the sample set for sample entries where the target index meets the preset optimal criterion and the safety constraint index set meets the threshold of the safety constraint index set.

[0028] Example 1: To verify the feasibility of this invention in practice, it was applied to a power supply scenario for communication equipment sharing a DC bus among multiple devices. The power supply end uses a wide-range adaptive power supply module to simultaneously supply power to three types of loads: a wireless communication processing unit, an edge computing unit, and an image acquisition unit. These three types of loads experience sudden current changes during startup, transmission, and encoding / decoding switching, leading to transient overshoot and transient drop in output voltage, increased peak-to-peak ripple, and accelerated temperature rise. Traditional tiered current limiting and fixed compensation parameters require repeated manual parameter adjustments and are prone to the problem of "stable under one operating condition, exceeding limits under another" under different load conditions.

[0029] In application, the sampling circuit continuously collects input voltage, output voltage, output current, and temperature. It extracts the ripple sequence from the output voltage and the transient response sequence from the output voltage and current. Simultaneously, it calculates the load state based on the output voltage, current, and temperature, forming an observation sequence. The operating condition feature vector is calculated from the observation sequence and encoded to obtain the operating condition context vector. This vector is then used to initialize the sample set, security proxy model, and safe feasible region, combined with historical operating condition parameter sets. The control parameter vector is set as the voltage target, current limiting threshold, ripple compensation parameter, and energy buffer threshold. The control mapping transforms the control parameter vector into voltage regulation unit parameters, current control unit parameters, ripple suppression unit parameters, and buffer unit parameters. The wide-domain adaptive power supply module outputs power accordingly. The safety constraint index set uses a preset time window to statistically measure peak voltage overshoot, peak voltage drop, peak ripple, recovery time, and peak temperature rise. GoSafeOpt generates candidate points under the prediction of the safety agent model and confidence boundary constraints and performs particle-based updates. When a candidate point does not meet the safety feasible region, boundary projection and reflection correction are performed to obtain a safety candidate control parameter vector, which is then sent out for execution. After each execution, the safety candidate control parameter vector, operating condition context vector, target index, and safety constraint index set are written into the sample set and the safety agent model and safety feasible region are updated until the target index converges. Then, a fixed control parameter vector is output for stable operation under this load condition.

[0030] The comparative verification used the same hardware, the same load sequence, and the same sampling conditions. The baseline scheme was a tiered current limiting and fixed compensation parameter scheme, with a preset time length of 500 milliseconds. The threshold values ​​for the safety constraint index set were set as follows: voltage overshoot peak value not exceeding 0.25V, voltage drop peak value not exceeding 0.35V, ripple peak value not exceeding 60mV, recovery time not exceeding 30ms, and temperature rise peak value not exceeding 18℃. The target index was a comprehensive index composed of the output voltage steady-state error and the ripple peak value weighted average. The load sequence included three typical abrupt changes: step from low load to full load, periodic pulse load, and mixed abrupt change load. The results are shown in Table 1. In the table, "Number of Boundary Exceedances" counts the number of times the safety constraint index set thresholds were exceeded in 10,000 consecutive abrupt change events, and "Number of Convergence Iterations" counts the number of iterations required to obtain the fixed control parameter vector.

[0031] Table 1 Performance Comparison Statistics of Wide-Area Flexible Power Supply Adaptive Control Methods

[0032] The data shows that this invention consistently controls the peak voltage overshoot, peak voltage drop, peak ripple, recovery time, and peak temperature rise within the threshold values ​​under three typical abrupt load conditions, without exceeding the limits during continuous abrupt events. The baseline scheme, however, exhibits significant limit exceedances under mixed abrupt load conditions, and requires repeated manual adjustments to the current limiting threshold and compensation parameters, still struggling to balance ripple and recovery time. This invention, without altering the hardware structure, achieves automatic convergence and solidification of the control parameter vector through a secure proxy model and GoSafeOpt search of the safe feasible domain constraints. The adaptation process is transformed from manual parameter tuning to iterative learning. The solidified control parameter vector maintains consistent performance under the same load conditions, solving the problems of large load differences, strong transient disturbances, low efficiency of manual parameter tuning, and easy exceedances under wide-range power supply conditions.

[0033] The above description is only a preferred embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any equivalent substitutions or modifications made by those skilled in the art within the scope of the technology disclosed in the present invention, based on the technical solution and inventive concept of the present invention, should be covered within the scope of protection of the present invention.

Claims

1. A wide-area flexible power supply adaptive control method based on reinforcement learning, characterized in that, Includes the following steps: Step 1: Collect input voltage, output voltage, output current, temperature, ripple sequence, transient response sequence, and load status to generate an observation sequence; Step 2: Generate operating condition context vectors from the observation sequence, read the historical operating condition parameter set, and initialize the safety agent model, safe feasible region, and sample set; Step 3: Set the control parameter vector, which includes the voltage target, current limiting threshold, ripple compensation parameters, and energy buffer threshold. Establish the control parameter vector and the control mapping between the voltage regulation unit, current control unit, ripple suppression unit, and buffer unit. Step 4: Set the safety constraint index set, which includes the voltage overshoot peak value, voltage drop peak value, ripple peak value, recovery time, and temperature rise peak value within a preset time window; Step 5: Based on the operating context vector, safety agent model, safe feasible region and sample set, use GoSafeOpt to generate candidate control parameter vectors, and use particle-style update to generate candidate points; Step 6: Perform boundary reflection projection correction on candidate points that do not meet the safety feasible region to obtain the safety candidate control parameter vector; Step 7: Send the safety candidate control parameter vector to the wide-domain adaptive power supply module for execution, collect the output response, calculate the safety constraint index set, write it into the sample set and update the safety proxy model and the safety feasible region, and output the fixed control parameter vector.

2. The wide-area flexible power supply adaptive control method based on reinforcement learning according to claim 1, characterized in that, Step one specifically involves: Collect sampling sequences of input voltage, output voltage, output current, and temperature, and set sampling timestamps for the input voltage sampling sequence, output voltage sampling sequence, output current sampling sequence, and temperature sampling sequence; Ripple extraction is performed on the output voltage sampling sequence to generate a ripple sequence, transient response extraction is performed on the output voltage sampling sequence and the output current sampling sequence to generate a transient response sequence, and the load state is calculated based on the output voltage sampling sequence, the output current sampling sequence and the temperature sampling sequence to generate a load state sequence; The input voltage sampling sequence, output voltage sampling sequence, output current sampling sequence, temperature sampling sequence, ripple sequence, transient response sequence, and load state sequence are aligned and spliced ​​according to the sampling timestamps to generate the observation sequence.

3. The wide-area flexible power supply adaptive control method based on reinforcement learning according to claim 1, characterized in that, Step two specifically involves: The operating condition feature vector is calculated based on the observation sequence. The operating condition feature vector includes input voltage features, output voltage features, output current features, temperature features, ripple features, transient response features, and load state features. The operating condition feature vector is normalized and encoded to obtain the operating condition context vector. Read the historical operating condition parameter set, which includes historical control parameter vector, historical safety constraint index set and historical target index, and write the historical control parameter vector, historical safety constraint index set and historical target index into the sample set; The safety agent model is trained based on the sample set. The safety agent model takes the control parameter vector and the operating condition context vector as input and outputs the estimated value of the safety constraint index set and the confidence bound of the safety constraint index set. Based on the security agent model, the lower confidence bound of the security constraint index set is calculated for the control parameter vector. The control parameter vector that satisfies the lower confidence bound of the security constraint index set threshold is defined as the safe feasible region.

4. The wide-area flexible power supply adaptive control method based on reinforcement learning according to claim 1, characterized in that, Step three specifically involves: Set a control parameter vector, which includes voltage target, current limiting threshold, ripple compensation parameter and energy buffer threshold, and set the value range for voltage target, current limiting threshold, ripple compensation parameter and energy buffer threshold; Establish a control mapping, which maps the control parameter vector to voltage regulation unit parameters, current control unit parameters, ripple suppression unit parameters, and buffer unit parameters. The voltage target corresponds to the voltage regulation unit parameters, the current limiting threshold corresponds to the current control unit parameters, the ripple compensation parameters correspond to the ripple suppression unit parameters, and the energy buffer threshold corresponds to the buffer unit parameters. Based on the control mapping to generate the power supply module parameter set, the wide-domain adaptive power supply module drives the voltage regulation unit, current control unit, ripple suppression unit and buffer unit to perform power supply output according to the power supply module parameter set.

5. The wide-area flexible power supply adaptive control method based on reinforcement learning according to claim 1, characterized in that, Step four specifically involves: Set a preset time window. The start timestamp of the preset time window is taken from the transition trigger point timestamp of the transient response sequence. The preset time window length is the preset time length. Within a preset time window, calculate the voltage overshoot peak, voltage drop peak, ripple peak peak, and temperature rise peak. The voltage overshoot peak is the maximum positive deviation of the output voltage sampling sequence relative to the voltage target. The voltage drop peak is the maximum negative deviation of the output voltage sampling sequence relative to the voltage target. The ripple peak peak is the maximum value of the ripple sequence minus the minimum value of the ripple sequence. The temperature rise peak is the maximum value of the temperature sampling sequence minus the temperature sampling value corresponding to the start timestamp of the preset time window. The recovery time is calculated within a preset time window. The recovery time is the shortest time from the start timestamp of the preset time window until the output voltage sampling sequence enters the deviation threshold range and remains continuously within the deviation threshold range. The deviation threshold range is determined by the threshold of the safety constraint index set.

6. The wide-area flexible power supply adaptive control method based on reinforcement learning according to claim 1, characterized in that, Step five specifically involves: The target indicator proxy model is trained based on the sample set. The target indicator proxy model takes the control parameter vector and the operating condition context vector as input, and the target indicator recorded in the sample set as the supervised output. It outputs the target indicator estimate and the target indicator uncertainty, and calculates the upper confidence bound of the target indicator based on the target indicator estimate and the target indicator uncertainty. The safety agent model is updated based on the sample set. The safety agent model takes the control parameter vector and the operating condition context vector as inputs and outputs the estimated value of the safety constraint index set and the uncertainty of the safety agent model. The lower confidence bound of the safety constraint index set is calculated based on the estimated value of the safety constraint index set and the uncertainty of the safety agent model. Initialize the candidate point set within the range of control parameter vector values. The candidate point set contains particle position vector and particle velocity vector. The particle position vector components correspond one-to-one with voltage target, current limiting threshold, ripple compensation parameter, and energy buffer threshold. Establish the historical optimal position vector and global optimal position vector for the candidate point set. For the candidate point set, calculate the upper confidence bound of the target index and the lower confidence bound of the safety constraint index set. Based on the lower confidence bound of the safety constraint index set satisfying the threshold of the safety constraint index set, select the safe candidate point set and update the global optimal position vector by selecting the particle position vector corresponding to the largest upper confidence bound of the target index in the safe candidate point set. GoSafeOpt is used to perform particle-based updates on the safety candidate point set. The particle-based update includes updating the particle position vector based on the particle velocity vector, the particle's historical best position vector, and the global best position vector. The particle position vector is subjected to interval constraint processing according to the value range of the control parameter vector. The upper confidence bound of the target index and the lower confidence bound of the safety constraint index set corresponding to the updated particle position vector are recalculated. The historical best position vector of the particle is updated according to the upper confidence bound of the target index, and the global best position vector is updated according to the safety candidate point set. When the iteration reaches the iteration termination condition, the control parameter vector corresponding to the global best position vector is determined as the candidate control parameter vector.

7. The wide-area flexible power supply adaptive control method based on reinforcement learning according to claim 1, characterized in that, Step six specifically involves: Perform a safe and feasible region determination on the control parameter vectors corresponding to the candidate points in the candidate point set. Control parameter vectors that are determined not to belong to the safe and feasible region are marked as out-of-bounds control parameter vectors. For the out-of-bounds control parameter vector, the boundary projection point is determined. The boundary projection point belongs to the safe and feasible region. The distance between the boundary projection point and the out-of-bounds control parameter vector is the minimum distance, which is determined by the sum of the squares of the differences between the components of the control parameter vector. The boundary projection point is used as the center of symmetry to reflect the out-of-bounds control parameter vector to obtain the reflected control parameter vector. The value range of the reflected control parameter vector is clipped to obtain the corrected control parameter vector. The safe and feasible region is determined on the corrected control parameter vector. If it is determined that it does not belong to the safe and feasible region, the boundary projection point determination and reflection are re-executed using the corrected control parameter vector as the out-of-bounds control parameter vector to obtain the safe candidate control parameter vector.

8. The wide-area flexible power supply adaptive control method based on reinforcement learning according to claim 1, characterized in that, Step seven specifically involves: The safety candidate control parameter vector is input into the wide-domain adaptive power supply module. The wide-domain adaptive power supply module generates a power supply module parameter set based on the control mapping. The power supply module parameter set includes voltage regulation unit parameters, current control unit parameters, ripple suppression unit parameters, and buffer unit parameters. The wide-domain adaptive power supply module performs power supply output based on the power supply module parameter set. The transition trigger point timestamp of the transient response sequence is used as the start timestamp of the preset time window, and the preset time window is determined based on the preset time length. The output voltage sampling sequence, output current sampling sequence, and temperature sampling sequence within the preset time window are collected to form the output response. The ripple sequence is extracted based on the output voltage sampling sequence, and the transient response sequence is extracted based on the output voltage sampling sequence and the output current sampling sequence. The safety constraint index set is calculated based on the output response, and the target index is calculated. The safety constraint index set includes voltage overshoot peak value, voltage drop peak value, ripple peak value, recovery time, and temperature rise peak value. The candidate control parameter vector, operating condition context vector, target index, and safety constraint index set are written into the sample set to form sample entries. The safety agent model is updated based on the sample set, and the safety feasible region is updated based on the safety agent model. The fixed control parameter vector is determined from the sample set based on the target index.