Regulation method based on dynamic modeling of PUE and reinforcement loop optimization

By combining a sparse Gaussian process multi-output dynamic model with reinforcement learning, the problems of multi-loop control interference and uncertainty in data center energy efficiency and thermal management are solved, achieving robust energy efficiency and temperature and humidity control, and reducing frequent equipment switching and safety risks.

CN121276989BActive Publication Date: 2026-04-17CHINA RAILWAY CONSTRUCTION ENGINEERING GROUP
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
CHINA RAILWAY CONSTRUCTION ENGINEERING GROUP
Filing Date
2025-10-22
Publication Date
2026-04-17

AI Technical Summary

Technical Problem

Existing technologies for energy efficiency and thermal management in data centers suffer from problems such as multi-loop control interference, insufficient model and uncertainty handling, and inadequate feasibility and safety constraints. This leads to reversed actions between devices, oscillations, and repeated energy consumption, making it difficult to adapt to non-stationary operating conditions.

Method used

A sparse Gaussian process multi-output dynamic model is used for prediction. The action sensitivity matrix is ​​calculated, an interference index is constructed, and token-based rotation and hierarchical reinforcement learning are introduced. Combined with virtual friction terms and jerk upper limits, multi-indicator fusion and hysteresis threshold closure are performed to achieve robust regulation.

Benefits of technology

It significantly suppresses opposite-direction movements and oscillations, improves the stability and fairness of multi-loop coordination, enhances the accuracy of PUE and cabinet inlet temperature and humidity control, and reduces the risk of model mismatch and misadjustment, as well as the risk of frequent equipment switching.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121276989B_ABST
    Figure CN121276989B_ABST
Patent Text Reader

Abstract

The application discloses a kind of based on PUE dynamic modeling and reinforcement loop optimization regulation and control method, to solve the problem of mutual interference of PID and logic control and outer layer optimization and reinforcement strategy in data center, the present application is through sparse gaussian process multi-output dynamic model prediction, multi-index fusion and uncertainty weighting of interference index, token rotation scheduling and hierarchical reinforcement learning adjustment, in model predictive control Introducing virtual friction and jerk upper limit and combining prediction safety filter, inject reference bias and cost weight to local closed loop, realizes the technical effect that outer layer optimization and inner layer control cooperate stability, suppress reverse action and saturation, improve PUE and cabinet inlet temperature and humidity control performance and safety.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of data center energy efficiency management, and in particular to a control method based on dynamic PUE modeling and enhanced loop optimization. Background Technology

[0002] As data center scale and load volatility continue to increase, the synergy between energy efficiency and thermal management has become crucial. Current engineering practices commonly employ local proportional-integral control and logic control to distribute the regulation of equipment such as chillers, cooling towers, pumps, fans, and valves, with Power Usage Effectiveness (PUE) serving as a key performance indicator.

[0003] To further improve energy efficiency, the industry has introduced rule-based energy-saving strategies, mechanism-based simplified model optimization, data-driven predictive models, and model predictive control. In some scenarios, reinforcement learning or federated optimization is being explored for global coordination. Although these methods have improved automation and energy efficiency, they still face challenges under conditions of multi-device coupling, non-stationary operating conditions, and safety constraints.

[0004] The shortcomings of existing technologies are mainly reflected in:

[0005] Multi-loop control interference: While the upper-level optimization changes the set value or weight, the local closed loop responds quickly, which can easily lead to actions between devices or actions that are opposite to the state response direction, phase mismatch, and hovering near the limit, resulting in oscillation and repeated energy consumption. There is a general lack of quantitative evaluation indicators, hysteresis triggering mechanisms, and fair coordination scheduling for such "fighting" phenomena.

[0006] Inadequate model and uncertainty handling: Common static regression or single-output models are difficult to characterize the multi-output coupling and cross-device linkage of PUE, rack inlet temperature and humidity, lack output covariance and prediction uncertainty, and lack action sensitivity measurement for control allocation, resulting in unstable weighting and arbitration basis, and difficulty in adapting to the non-stationarity of outdoor weather and IT load.

[0007] Insufficient feasibility and safety constraints: Many optimization or learning strategies do not explicitly consider minimum start-stop time, action slope and jerk upper limits, gear dispersion and equipment and environmental safety boundaries, lacking a safety projection based on uncertainty, which can easily lead to frequent switching, overshoot and comfort or equipment safety risks.

[0008] To address the aforementioned issues, there is an urgent need for a control method that can quantify disturbances at each stage of modeling, evaluation, and execution, balance robustness and security, and coordinate the actions of multiple devices in a distributed manner. Summary of the Invention

[0009] One objective of this invention is to propose a control method based on PUE dynamic modeling and reinforcement loop optimization. Addressing the problems of mutual interference, oscillation, and saturation between local PID control, logic closed-loop control, and outer optimization and reinforcement strategies in existing technologies, this invention proposes a technical solution that uses a sparse Gaussian process multi-output dynamic model for prediction and uncertainty measurement; calculates an action sensitivity matrix for weighting and arbitration; constructs a multi-index fusion of interference indices and a hysteresis threshold closed loop; implements cost weights and reference biases using token-based rotation and hierarchical reinforcement learning; and introduces a virtual friction term and jerk upper limit in model predictive control, along with predictive safety filtering and safety projection. This invention achieves the technical effects of suppressing reverse actions and frequent switching, realizing multi-loop collaborative stability and fair arbitration, improving the accuracy and safety of PUE and cabinet inlet temperature and humidity control, and ensuring robust operation under non-stationary conditions.

[0010] A control method based on PUE dynamic modeling and reinforcement loop optimization according to an embodiment of the present invention is characterized by comprising:

[0011] S1. Collect raw data, complete time alignment, outlier handling and feature extraction, and output aligned dataset;

[0012] S2. Input the aligned dataset, construct and train the sparse Gaussian process multi-output PUE dynamic model, and output the model prediction results;

[0013] S3. Input the model prediction results, calculate and fuse multiple indicators to obtain the interference index, and perform confidence correction and hysteresis threshold setting based on the model uncertainty, and output the interference assessment results including the interference index.

[0014] S4. Input the interference assessment results and action sensitivity matrix, execute token-based rotation scheduling, generate the set of devices holding tokens, the cooldown period parameter and the minimum hold time parameter, generate the cost weight set and the reference trajectory bias set for the devices holding tokens, and adaptively shrink the weight adjustment amplitude according to the interference index, and output the weight adjustment result containing the set of devices holding tokens, the cost weight set and the reference trajectory bias set;

[0015] S5. Input the weighting result, adaptively calculate the virtual friction coefficient and jerk upper limit based on the disturbance index, introduce the virtual friction term and jerk upper limit into the objective function and constraints of the model predictive control respectively, and solve the candidate control action. Perform safe projection after predictive safety filtering, output the safe control action, and merge the cost weight set and the reference trajectory bias set with the safe control action to form the execution input set.

[0016] S6. Input the execution input set, inject the reference trajectory bias set and cost weight set into the automatic control closed loop of the device's local controller, execute safety control actions, output the final execution command and complete the regulation.

[0017] Optionally, step S1 specifically includes:

[0018] Raw data is collected from the monitoring system and field sensors. The raw data includes at least the equipment speed, valve position, air volume, pump speed, compressor status, outdoor weather conditions, IT load, energy consumption metering, cabinet inlet temperature, and cabinet inlet humidity.

[0019] The original data is time-aligned on a unified reference time axis, including resampling data from different sampling periods, correcting and filling timestamp offsets, missing data, and inconsistencies in clocks from different data sources, so that each variable corresponds under the same time index.

[0020] Outlier processing is performed on the raw data, including identifying outliers based on the reasonable value range of the device, the rate of change constraint and statistical discrimination, and correcting or removing the identified outliers.

[0021] Based on this, feature extraction is performed to form a feature set for modeling. The feature set includes at least the current value of each original variable, several orders of historical lag values, rate of change and statistical characteristics, and includes a normalized or coded representation of the discrete device status. It may also include combined features of cross-device and environmental variables constructed based on correlation analysis.

[0022] Output an aligned dataset with a uniform time index.

[0023] Optionally, step S2 specifically includes:

[0024] Using the aligned dataset as input, a sparse Gaussian process multi-output dynamic model for discrete-time prediction is constructed. The model uses features in the aligned dataset as independent variables, including at least the current value of the equipment control quantity and several orders of historical lag values, equipment state quantity and environmental quantity, and uses PUE, rack inlet temperature and rack inlet humidity as multiple outputs.

[0025] The parameters and hyperparameters of the model are trained by maximum likelihood training or variational inference, and the computational complexity is reduced by using sparse approximation.

[0026] Given the current state and the device's actions, the trained model is used to infer at least one prediction step, outputting the predicted values ​​and uncertainties of each output variable. The uncertainties include at least the prediction variances of each output variable and may include the covariance between the output variables.

[0027] And calculate the action sensitivity matrix of equipment actions to PUE. The elements of the action sensitivity matrix are defined as the partial derivatives of the predicted PUE value with respect to each equipment control quantity under a given state, or the numerical sensitivity obtained by applying a small perturbation to the equipment control quantity. The action sensitivity matrix can be calculated at the current prediction step or summarized over multiple prediction steps.

[0028] The predicted value, the uncertainty, and the action sensitivity matrix are used as the model prediction results.

[0029] Optionally, step S3 specifically includes:

[0030] Based on the model prediction results, interference assessment is performed on the device action and at least one state response within a set sliding time window. Specifically, three types of indicators are calculated simultaneously and fused to generate an interference index. The first type of indicator is frequency domain coherence and phase margin. The coherence is measured based on the coherence obtained from the power spectral density and cross power spectral density of the device action and state response. The phase margin is measured based on the margin of the phase difference relative to the stable reference at the main peak of coherence or at the main energy frequency band.

[0031] The second type of indicator is the proportion of actions in opposite directions. The proportion is defined as the percentage of samples where the direction of the device control quantity is opposite to that of the corresponding state response within the sliding time window. The selection of the device pairs and the counting weight can be determined based on the action sensitivity matrix or the device coupling relationship.

[0032] The third type of indicator is the proportion of the local controller output of the device that is close to saturation. The proportion is defined as the percentage of time that the controller output or actuator opening falls into a preset range close to the upper limit or the lower limit.

[0033] The above three types of indicators are weighted and fused to obtain the interference index. The weights are adaptively adjusted according to the model uncertainty, and the interference index is subjected to confidence correction to obtain a conservative estimate.

[0034] Based on the model uncertainty, set the trigger threshold and release threshold with hysteresis, and output the interference evaluation result, which includes at least the interference index.

[0035] Optionally, step S4 specifically includes:

[0036] Using the interference assessment results as input, token-based rotation scheduling is performed. Within each scheduling cycle, tokens are issued, renewed, and revoked based on the interference index, device coupling relationship, and current holding status. The set of devices holding tokens is determined, and a cooldown period parameter and a minimum holding time parameter associated with token management are generated for each device. The cooldown period parameter is used to limit that a token cannot be reissued during the cooldown period after it is revoked, and the minimum holding time parameter is used to limit that a token cannot be revoked within the time after it is issued to suppress frequent switching.

[0037] The determination of equipment coupling relationships can be based on correlation analysis of historical data or on motion sensitivity matrices;

[0038] When multiple devices compete for a token, arbitration is conducted based on the interference index and preset priority to ensure rotation and fairness.

[0039] On a given set of devices holding tokens, the hierarchical reinforcement learning weighting module calculates a cost weight set and a reference trajectory bias set based on the interference index and the action sensitivity matrix. The cost weight set is used to adjust the weight of the cost term in the subsequent optimization objective, and the reference trajectory bias set is used to adjust the reference value offset of each device or each state.

[0040] To ensure stability and feasibility, the weighting amplitude is adaptively shrunk according to the magnitude of the interference index. The shrinkage has a monotonic relationship with the interference index and has upper and lower bounds.

[0041] Output the weight adjustment result, which includes at least the set of devices holding tokens, the set of cost weights, and the set of reference trajectory biases.

[0042] Optionally, step S5 specifically includes:

[0043] Using the weighting result as input, the virtual friction coefficient and the upper limit of jerk are adaptively calculated based on the disturbance index through a preset mapping relationship. The virtual friction coefficient is used to measure the additional damping strength on the control increment, and the upper limit of jerk is used to limit the amplitude of the difference between two adjacent control increments in discrete time.

[0044] The virtual friction term is incorporated into the objective function of the model predictive control, so that the control increment is penalized with respect to the virtual friction coefficient.

[0045] The upper limit of jerk is incorporated into the constraint set of model predictive control, and the constraint set explicitly includes at least the following constraints: minimum start-stop time constraint for equipment with start-stop characteristics, action slope limit for continuously adjustable equipment, and gear discreteness constraint for equipment with discrete gears.

[0046] Under the aforementioned objective function and constraint set, the optimization problem is solved in the prediction time domain to obtain candidate control actions;

[0047] The candidate control actions are input into the predictive safety filtering module for safety projection. A safety set is constructed based on the safety boundaries of the equipment and the environment and can be tightened according to the model uncertainty, so that the control actions after safety projection satisfy the safety set and prediction constraints, and a safe control action is output.

[0048] The cost weight set and reference trajectory bias set in the weighting result are merged with the safety control action to form the execution input set.

[0049] Optionally, step S6 specifically includes:

[0050] The automatic control closed loop of the input set input device local controller will be executed, and injection and execution will only be performed on devices holding tokens;

[0051] The reference trajectory bias set is applied to the set value of the corresponding device to form a corrected reference trajectory, and the cost weight set is applied to the target weight, logic threshold or priority of the local controller to adjust the control decision;

[0052] Safety control actions are issued as external control commands and given a higher priority in the output synthesis of the local controller, so that they can cover or superimpose the free response of the local controller within the scope of safety constraints;

[0053] The local controller includes proportional-integral control or logic control, which generates the final execution instruction and completes the regulation in the closed loop based on real-time feedback.

[0054] The beneficial effects of this invention are:

[0055] 1. By integrating multiple indicators of the interference index and using the hysteresis threshold closed loop, combined with token-based rotation scheduling and hierarchical reinforcement learning weight adjustment, the outer layer optimization only applies to devices holding tokens and adaptively shrinks the weight adjustment amplitude, significantly suppressing opposite-direction actions, oscillations, and near-saturation phenomena, thereby improving the stability and fairness of multi-loop collaboration.

[0056] 2. A sparse Gaussian process multi-output dynamic model is adopted to simultaneously predict PUE, rack inlet temperature and inlet humidity, and output uncertainty and action sensitivity matrices to provide a robust basis for weighting and arbitration. This improves energy efficiency and temperature and humidity control accuracy under load and weather non-stationary conditions and reduces the risk of misadjustment caused by model mismatch.

[0057] 3. In model predictive control, a virtual friction term and an upper limit for jerk are introduced, and candidate actions are projected onto the safety set through predictive safety filtering. This explicitly satisfies constraints such as minimum start-stop time, slope limit, and gear dispersion, making the control action smooth, reducing equipment fatigue and frequent switching, and reducing overshoot and safety risks. Attached Figure Description

[0058] The accompanying drawings are provided to further illustrate the invention and form part of the specification. They are used in conjunction with embodiments of the invention to explain the invention and do not constitute a limitation thereof. In the drawings:

[0059] Figure 1 This is a flowchart of a control method based on PUE dynamic modeling and reinforcement loop optimization proposed in this invention. Detailed Implementation

[0060] The present invention will now be described in further detail with reference to the accompanying drawings. These drawings are simplified schematic diagrams, illustrating only the basic structure of the invention, and therefore only show the components relevant to the invention.

[0061] refer to Figure 1 A control method based on PUE dynamic modeling and reinforcement loop optimization, characterized by comprising:

[0062] S1. Collect raw data, complete time alignment, outlier handling and feature extraction, and output aligned dataset;

[0063] S2. Input the aligned dataset, construct and train the sparse Gaussian process multi-output PUE dynamic model, and output the model prediction results;

[0064] S3. Input the model prediction results, calculate and fuse multiple indicators to obtain the interference index, and perform confidence correction and hysteresis threshold setting based on the model uncertainty, and output the interference assessment results including the interference index.

[0065] S4. Input the interference assessment results and action sensitivity matrix, execute token-based rotation scheduling, generate the set of devices holding tokens, the cooldown period parameter and the minimum hold time parameter, generate the cost weight set and the reference trajectory bias set for the devices holding tokens, and adaptively shrink the weight adjustment amplitude according to the interference index, and output the weight adjustment result containing the set of devices holding tokens, the cost weight set and the reference trajectory bias set;

[0066] S5. Input the weighting result, adaptively calculate the virtual friction coefficient and jerk upper limit based on the disturbance index, introduce the virtual friction term and jerk upper limit into the objective function and constraints of the model predictive control respectively, and solve the candidate control action. Perform safe projection after predictive safety filtering, output the safe control action, and merge the cost weight set and the reference trajectory bias set with the safe control action to form the execution input set.

[0067] S6. Input the execution input set, inject the reference trajectory bias set and cost weight set into the automatic control closed loop of the device's local controller, execute safety control actions, output the final execution command and complete the regulation.

[0068] In this specific embodiment, S1 specifically refers to:

[0069] For multi-source raw data from monitoring systems and field sensors (including at least equipment speed, valve position, air volume, pump speed, compressor status, outdoor weather conditions, IT load, energy consumption metering, cabinet inlet temperature, and cabinet inlet humidity), the overall process is to align data with different sampling periods and inconsistent clocks on a unified timeline, complete outlier handling and missing value imputation, and extract features for subsequent modeling. First, a unified baseline timeline is established:

[0070] ;

[0071] in Indicates the first Timestamp of each reference time point Indicates start time, Time index representing a non-negative integer, Indicates the resampling period for positive values;

[0072] After mapping all data sources to this timeline, in each The above performs window resampling and alignment on the variables, and the variable index is denoted as... in Represents a specific variable, The set of variables being processed is represented by the alignment value calculated using a resampling operator matched to the variables:

[0073] ;

[0074] in Representing variables At any moment Alignment value, Indicates the variable Resampling operators (such as linear interpolation or weighted mean) Representing variables Time on the original timeline Observations Represents time variables and intervals on the original timeline. Indicates Centered on and with a width of The resampling window;

[0075] When data is missing within a window, linear interpolation or forward padding strategies are used to maintain time continuity, and clock consistency is achieved through unified correction before alignment when there are timestamp offsets in multi-source data.

[0076] After alignment, outlier handling is performed by first calculating the moving average statistics and then constructing standardized statistics:

[0077] ;

[0078] in Representing variables At any moment Standardized scores Indicates length as The moving average obtained from the sliding window Indicates the same window length The obtained moving standard deviation, Indicates the number of samples in the sliding window. Represents a very small positive number used for numerical stability;

[0079] Anomalies are identified by combining the reasonable range of variable values ​​with the value between the rates of change at adjacent time points. Anomalies are corrected by interpolation or neighborhood statistics, and those that cannot be corrected are removed.

[0080] Based on this, feature extraction is performed, retaining the current alignment value and several orders of historical lag for each variable, calculating the rate of change and moving statistical features, and normalizing or encoding the discrete device status. At the same time, combined features across devices and environmental variables can be constructed based on correlation analysis.

[0081] The final result is an aligned dataset with a unified time index:

[0082] ;

[0083] in Indicates the alignment of datasets, Indicates at time The feature vector composed of the above features, Represents the total number of reference time points and is determined by the length of the sampling interval and A joint decision.

[0084] In this specific embodiment, S2 specifically refers to:

[0085] Using an aligned dataset with a unified time index as input, a multi-output dynamic model of a sparse Gaussian process is constructed and trained to simultaneously predict power efficiency, rack inlet temperature, and rack inlet humidity, and to quantify the uncertainty. The model's multiple outputs at discrete time are represented as follows:

[0086] ;

[0087] in Indicates time Multiple output response vectors Indicates time The response of power use efficiency Indicates time Rack entrance temperature response, Indicates time Humidity response at the rack entrance Indicates the first The timestamp of each reference time point and It is a non-negative integer;

[0088] The input consists of features and device actions, with the device action vector denoted as:

[0089] ;

[0090] in Indicates time Device control vector, Indicates the first Each control channel in time Action value The number of control channels is represented by the feature vector denoted as . And it originates from the feature set extracted in step S1;

[0091] Given the current state and action, construct features for prediction for at least one future prediction step, and denot them as follows:

[0092] ;

[0093] in Indicates time Predicted input features This represents a constructor that maps current features and actions to future features. The prediction step size represents a positive integer;

[0094] The model employs a multi-output covariance structure to characterize the coupling relationship between responses and reduces computational complexity with a sparse approximation of induced points. The covariance function can be expressed using an intrinsic covariance model as follows:

[0095] ;

[0096] in Represents the covariance function between multiple output pairs. and Indicate the output index and correspond to each. Component index, The elements of the covariance coefficient matrix between outputs, This represents a scalar basis kernel function that satisfies the positive definite condition, such as the ARD radial basis function kernel. and This represents two eigenvectors;

[0097] During the training phase, kernel hyperparameters are learned through maximum likelihood or variational inference, and an induced point set is selected to achieve sparse approximation. In the online phase, features are tested... The above gives the multi-output prediction distribution and uncertainty. The prediction distribution is denoted as:

[0098] ;

[0099] in Represents a multivariate normal distribution, Represents the predicted mean vector, This represents a multi-output prediction covariance matrix, where the diagonal elements are the prediction variances of each output and the off-diagonal elements are the covariances between the outputs.

[0100] To reflect the computational form under the sparse approximation, the predicted mean and covariance can be written as:

[0101] and ;

[0102] in The cross-covariance matrix of the test input and the induced point set is represented by... The autocovariance matrix of the induced point set, Represents the variance of observation noise. Indicates and identity matrices of the same dimension This represents the posterior mean vector or equivalent representation of the latent function at the induction point. Represents the autocovariance matrix of the test input. The transpose of the cross-covariance matrix between the inducing points and the test input, and the set of inducing points are denoted as . in Represents the induced input set, Indicates the first A induced input, Indicates the number of induction points;

[0103] To support subsequent weighting and arbitration, it is also necessary to calculate the sensitivity of equipment actions to PUE and express it in gradient form as follows:

[0104] ;

[0105] in Indicates time Action sensitivity row vector, Indicates the prediction step size The mean of the predicted PUE is compared with the first Partial derivatives of each action component, Indicates time The predicted mean of PUE;

[0106] For engineering feasibility, when the analytical gradient is unavailable, a finite difference approximation can be used, denoted as:

[0107] ;

[0108] in Indicates the first Numerical sensitivity of each action component A semicolon indicates a sufficiently small positive disturbance amplitude.

[0109] Indicates except the first The conditional notation that keeps all components except the action component unchanged is used to ultimately obtain the multi-output prediction mean, prediction covariance, and action sensitivity as model prediction results for subsequent disturbance assessment and optimization control.

[0110] In this specific embodiment, S3 specifically refers to:

[0111] Based on the predicted mean, uncertainty, and action sensitivity, disturbance assessment is performed on the device's actions and state responses within a sliding time window, generating triggering results with hysteresis thresholds. First, a sliding window index set is defined:

[0112] ;

[0113] in Indicates the index based on the current reference time. The right endpoint is, and the length is window Indicates the time index within the window, and They represent the first With the The timestamps for each baseline time point are recorded separately for actions and responses. and ,in Indicates the first Each device control channel at time Action value This indicates the status response being evaluated (which can be PUE, rack inlet temperature, or inlet humidity). Indicates the control channel index, Indicates the total number of control channels;

[0114] The first type of metric is frequency domain coherence and phase margin, which is calculated by using spectral estimation within a window and selecting the dominant frequency within the band of interest, followed by weighted coherence-phase metric:

[0115] ;

[0116] in Indicates the first category of indicators, Represents sensitivity-weighted overall coherence. Indicates the first Channel coherence at the main frequency, The cross-power spectral density representing the action and response, and Representing the corresponding self-power spectral density, Representing frequency variables, Indicates focus on frequency bands The main frequency of coherence within, Indicates the weighted main frequency phase, Indicates the first Channel phase, The weights are obtained by normalizing the action sensitivity. This indicates that the PUE obtained in step S2 is related to the first... Sensitivity to motion components Indicates phase penalty weight, Indicates stable reference phase, Represents pi;

[0117] The second type of indicator is the proportion of actions in opposite directions, first defined as the increment between adjacent time moments:

[0118] and ;

[0119] And calculate: ;

[0120] in Indicates the second type of indicator, and Let represent the changes in action and response, respectively, and let 1[.] represent an indicator function that takes 1 if the predicate is true and 0 otherwise;

[0121] The third type of indicator is the proportion of the local controller output that is close to saturation:

[0122] ;

[0123] in Indicates the third type of indicator, Indicates the first Channel at time Local controller output, and These represent the upper and lower limits of the output, respectively. Indicates buffer bandwidth approaching saturation. Indicates the number of samples in the window. Represents a logical OR operation;

[0124] The above three types of indicators are adaptively weighted and fused into an interference index based on uncertainty:

[0125] ;

[0126] in Indicates the interference index, These represent three types of indicators respectively. Represents normalized weights, Indicates base weights, The mapping function that adjusts for uncertainty Indicates the adjustment coefficient, This represents the average of the predicted standard deviations of the three types of outputs. Indicates the first The output at time... The predictive standard deviation These represent the output indices for PUE, rack inlet temperature, and rack inlet humidity, respectively.

[0127] Confidence correction is performed to obtain a conservative estimate:

[0128] ;

[0129] in Indicates the corrected interference index, Indicates the confidence correction coefficient;

[0130] Finally, a threshold with hysteresis is set based on the uncertainty. and ;

[0131] in and These represent the trigger and release thresholds, respectively. Indicates the baseline trigger threshold, The tightening slope represents the rate of change of uncertainty. Indicates the hysteresis width;

[0132] and with Determine the trigger, with The system determines whether to release the interference and outputs an interference evaluation result that includes the interference index and the trigger state.

[0133] In this specific embodiment, S4 specifically refers to:

[0134] Using interference assessment results and action sensitivity as input, a token-based round-robin scheduling method is used for the device set at each scheduling time to generate weight adjustment results, where device indexes are defined. Indicates the first Taiwanese equipment and The total number of devices is a positive integer, and the current scheduling time is [time value missing]. Indicates the first The normalized values ​​of the interference index at each reference time point are denoted as follows: Indicates at time The normalized interference intensity and the action sensitivity to PUE are denoted as Indicates at time No. A scalar measure of the sensitivity of device actions to PUE, and normalized importance weights constructed based on this sensitivity:

[0135] ;

[0136] in Indicates the first The relative importance of the equipment Indicates absolute value operation, given device priority Indicates the first Prioritize or prioritize the prior or online updates of the equipment and weigh them against a coefficient. The arbitration score is obtained by merging the scores. ;

[0137] in Indicates the first The equipment is at all times The overall score and token distribution are determined by a binary variable. Decision and Indicates to the first Taiwan device issues tokens, to This indicates that no tokens will be issued, limited by the maximum number of concurrent tokens. This indicates the maximum number of devices allowed to hold tokens simultaneously and eligibility criteria. Indicates the first The equipment is at all times Whether a pair qualifies for a token (qualification is determined by the minimum hold and cooldown constraints of the state), while also considering the set of coupled mutually exclusive pairs. Given a set of device pairs that cannot simultaneously hold tokens, token allocation is determined through the following optimization:

[0138] st ;

[0139] for The objective is to maximize the overall score, and the constraints represent the concurrency upper limit, eligibility restriction, and coupling mutual exclusion restriction, respectively. The optimal solution is denoted as... And thus obtain the set of devices holding the token. ;

[0140] in Indicates time The token set is used to adaptively shrink the weighting amplitude to match the interference intensity;

[0141] Define a monotonically non-increasing contraction factor:

[0142] ;

[0143] in Represents the contraction function, and Let the lower and upper bounds of the contraction factor be respectively, and satisfy the following conditions: ;

[0144] Based on this, the generation cost weight and reference bias for token-holding devices are:

[0145] ;

[0146] in Indicates the first The cost weight of each device in subsequent optimization objectives, Indicates its benchmark weight, Indicates the first Reference trajectory offset of the device or its associated status Indicates the maximum allowable offset amplitude. This indicates the direction factor determined by the goal of reducing PUE and The symbolic function takes the value of ;

[0147] To suppress frequent switching while ensuring fairness, set the time parameters as follows:

[0148] ;

[0149] in Indicates the first Minimum holding time of the device Indicates its cooling period, and Indicates the corresponding baseline value, and This represents the gain coefficient that increases with increasing interference.

[0150] Final output weighting result .

[0151] In this specific embodiment, S5 specifically includes:

[0152] Using the weighting result and disturbance index as input, the disturbance intensity is mapped to virtual friction and jerk upper limits and embedded into the target and constraints of model predictive control. Simultaneously, a predictive safety filter is used to perform a safety projection on the candidate control. Finally, this is merged with the cost weight and reference bias to form the execution input set. Specifically, the current scheduling time is denoted as... ,in Indicates the first The timestamp of each reference time point and For non-negative integers, the normalized value of the interference exponent is denoted as . ,in Indicates at time The normalized disturbance intensity, virtual friction coefficient, and upper limit of jerk are given by linear mappings:

[0153] and ;

[0154] in The virtual friction coefficient used to penalize the control increment. and These represent the lower and upper bounds of the coefficient, respectively. The upper limit of the jerk represents the difference between adjacent control increments in discrete time. and These represent the lower and upper bounds of the upper limit, respectively.

[0155] When constructing a model for predictive control problems, let the prediction time domain length be denoted as . ,in The number of optimization steps and the control sequence are denoted as . ,in Indicates at time Solving for the future control vector sequence, Indicates at time The control vector and discrete control increment are denoted as:

[0156] ;

[0157] in The control differential and prediction outputs at adjacent time points are denoted as... ,in Indicates at time Multiple output prediction vectors (including PUE, rack inlet temperature and inlet humidity);

[0158] The reference trajectory is denoted as ,in Indicates at time The reference vector (obtained from the reference offset given in step S4 of the benchmark reference superposition step), and the objective function incorporating the virtual friction term are:

[0159] ;

[0160] in Indicates at time Optimization target value, Represented by the weight matrix Defined weighted 2-norm and The elements are set by the cost weight vector given in step S4. Represents the L2 norm;

[0161] Explicitly add a jerk limit to the constraints:

[0162] ;

[0163] in It represents the infinite norm and limits the magnitude of the difference between adjacent control increments in discrete time;

[0164] Other engineering feasibility constraints (such as minimum start-stop time, slope limit, and gear dispersion) are described in words and included in the constraint set during the actual solution. The candidate control sequence obtained under the above objectives and constraints is denoted as... ,in This means satisfying the constraints and minimizing To ensure safety, the candidate solutions are input into the prediction safety filter module to construct a safety margin that tightens with uncertainty:

[0165] ;

[0166] in Indicates at time The tightening of security measures Indicates the tightening coefficient, This represents the average of the standard deviations of the multiple output predictions;

[0167] Define a safety set at the safety boundary between the device and the environment. ,in Indicates time Acceptable control set and as Tightening, safety projection is achieved with minimal deviation:

[0168] ;

[0169] in This represents the control sequence after secure projection;

[0170] Finally, the first control in the sequence is taken as the safety control action and combined with the cost weight and reference bias to form the execution input set:

[0171] ;

[0172] in Indicates at time Execution input set, Indicates safety control actions, Represents the cost weight vector, This represents the reference trajectory bias vector.

[0173] In this specific embodiment, S6 specifically refers to:

[0174] To implement an automatic control closed loop that injects the execution input set into the device's local controller, injection is only performed on devices holding tokens, and security control actions are prioritized within security constraints. The execution input set is first defined as follows:

[0175] ;

[0176] in Indicates at time Execution input set, Represents the security control vector at the next sampling time. Represents the cost weight vector, Represents the reference trajectory bias vector, Indicates the first Each reference time point and It is a non-negative integer;

[0177] Let the token set be The injection range is defined using a binary injection indicator. ,in Indicates the first The device holds the token and at time Participate in injection, This indicates that those who do not hold tokens will not participate in the injection. Represents device index, Indicates the total number of devices;

[0178] To implement reference bias injection, the local reference setting is corrected as follows:

[0179] ;

[0180] in Indicates the first The equipment is at all times Correction reference, Indicates the corresponding reference base. Indicates the reference trajectory offset component;

[0181] Cost weight injection is used to adjust the target weights, logical thresholds, or priorities of the local controller, and updates them additively as follows:

[0182] ;

[0183] in Indicates the first Current weight of the device Indicates the corresponding benchmark weights, Indicates the cost weight components;

[0184] To assign higher priority to safety control actions in output synthesis, device components are first extracted from the safety control vector. and the free response of the local controller. By priority weight Linear superposition yields the synthesis instruction:

[0185] ;

[0186] in Indicates the first Safety control components and symbols of the equipment The vector represents the first Selecting operators for each component Indicates the free output of the local controller. Indicates the first Synthetic control commands for the equipment The closer the value is to 1, the higher the priority of the safety control action.

[0187] To satisfy the constraints of the device actuator, projection saturation is applied to obtain the final command:

[0188] ;

[0189] in Indicates the first The equipment is at all times The final execution instruction Indicates the interval to the closed interval Projection operator, and These represent the minimum and maximum allowable control values, respectively.

[0190] Finally, the commands from each channel are combined. Issued and implemented, among which Indicates at time The final execution instruction vector, superscript Indicates vector transpose;

[0191] The local controller can be a proportional-integral control or logic control and updates its internal state and output based on real-time feedback. Devices without tokens maintain local closed-loop free operation, while devices with tokens execute according to the above-mentioned injection and synthesis mechanism until the next scheduling cycle.

[0192] The above description is only a preferred embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any equivalent substitutions or modifications made by those skilled in the art within the scope of the technology disclosed in the present invention, based on the technical solution and inventive concept of the present invention, should be covered within the scope of protection of the present invention.

[0193] This application employs a collaborative closed loop comprising model, evaluation, scheduling, optimization, safety, and execution, connecting outer-layer optimization with local PID and logic control. A sparse Gaussian process multi-output dynamic model simultaneously provides predictions and uncertainties for PUE, rack inlet temperature, and rack inlet humidity, and calculates an action sensitivity matrix as a robust basis for weighting and arbitration. Within a sliding time window, frequency domain coherence and phase margin, the proportion of actions in opposite directions, and the near-saturation proportion are fused to construct an interference index. This index is then combined with uncertainty weighting, confidence correction, and hysteresis thresholds to achieve stable identification of multi-loop interference. Through token-based rotation and hierarchical reinforcement learning, reference trajectory bias and cost weights are injected only into devices holding tokens, and the weighting amplitude is adaptively reduced according to the interference intensity. Virtual friction and jerk upper limits are introduced into model predictive control, and candidate controls are projected onto a safety set using predictive safety filtering. Finally, in the local closed loop, safety control actions are assigned higher priority for output synthesis. The aforementioned causal chain effectively suppresses opposite-direction actions and oscillations, avoids near-saturation and frequent switching, improves the stability of multi-loop coordination, and enhances the accuracy and safety of PUE and cabinet inlet temperature and humidity control under non-steady operating conditions.

[0194] In terms of algorithm structure, this case addresses the challenges of multi-loop interference, strong device coupling, high prediction uncertainty, and numerous engineering constraints with targeted improvements: At the modeling layer, a multi-output dynamic model of a sparse Gaussian process is adopted, explicitly providing covariance and uncertainty and quantifying the sensitivity to PUE actions, enhancing the reliability of weighting and token arbitration. At the evaluation layer, three types of indicators are fused and combined with uncertainty weighting, confidence correction, and hysteresis thresholds to form a closed loop, reducing trigger jitter and misjudgment. At the coordination layer, token-based rotation with a cooldown period and minimum hold time, combined with hierarchical reinforcement learning, distributes cost weights and reference biases only to token-holding devices, balancing fairness and feasibility. At the optimization and safety layer, virtual friction and jerk upper limits are incorporated into the model's predictive control objectives and constraints, and predictive safety filtering achieves uncertainty-based safety set tightening and safety projection. At the execution layer, high-priority safety control actions are synthesized with local controller free responses to complete the distribution. These structural improvements make identification more robust, intervention more orderly, and implementation safer, further strengthening the technical effects of suppressing interference and improving energy efficiency and temperature and humidity control performance.

Claims

1. A method for regulating based on dynamic modeling of PUE and reinforcement loop optimization, characterized in that, include: S1. Collect raw data, complete time alignment, outlier handling and feature extraction, and output aligned dataset; S2. Input the aligned dataset, construct and train the sparse Gaussian process multi-output PUE dynamic model, and output the model prediction results; S3. Input the model prediction results, calculate and fuse multiple indicators to obtain the interference index, and perform confidence correction and hysteresis threshold setting based on the model uncertainty, and output the interference assessment results including the interference index. The interference index is obtained by simultaneously calculating and fusing three types of indicators. The first type of indicator is frequency domain coherence and phase margin. The coherence is measured based on the power spectral density and cross power spectral density of the device action and state response. The phase margin is measured based on the margin of the phase difference relative to the stable reference at the main peak of coherence or at the main energy frequency band. The second type of indicator is the proportion of actions in opposite directions. The proportion is defined as the percentage of samples where the direction of the device control quantity is opposite to that of the corresponding state response within the sliding time window. The selection of the pairing objects for counting and the counting weight are determined based on the action sensitivity matrix or the device coupling relationship. The third type of indicator is the proportion of the local controller output of the device that is close to saturation. The proportion is defined as the percentage of time that the controller output or actuator opening falls into a preset range close to the upper limit or the lower limit. The above three types of indicators are weighted and fused to obtain the interference index. The weighting weights are adaptively adjusted according to the model uncertainty, and the interference index is subjected to confidence correction to obtain a conservative estimate. S4. Input the interference assessment results and action sensitivity matrix, execute token-based rotation scheduling, generate the set of devices holding tokens, the cooldown period parameter and the minimum hold time parameter, generate the cost weight set and the reference trajectory bias set for the devices holding tokens, and adaptively shrink the weight adjustment amplitude according to the interference index, and output the weight adjustment result containing the set of devices holding tokens, the cost weight set and the reference trajectory bias set; S5. Input the weighting result, adaptively calculate the virtual friction coefficient and jerk upper limit based on the disturbance index, introduce the virtual friction term and jerk upper limit into the objective function and constraints of the model predictive control respectively, and solve the candidate control action. Perform safe projection after predictive safety filtering, output the safe control action, and merge the cost weight set and the reference trajectory bias set with the safe control action to form the execution input set. The virtual friction coefficient is used to measure the additional damping strength on the control increment, and the upper limit of jerk is used to limit the magnitude of the difference between two adjacent control increments in discrete time. S6. Input the execution input set, inject the reference trajectory bias set and cost weight set into the automatic control closed loop of the device's local controller, execute safety control actions, output the final execution command and complete the regulation.

2. The method of claim 1, wherein the method is based on PUE dynamic modeling and reinforcement loop optimization. S1 specifically refers to: Raw data is collected from the monitoring system and field sensors. The raw data includes at least the equipment speed, valve position, air volume, pump speed, compressor status, outdoor weather conditions, IT load, energy consumption metering, cabinet inlet temperature, and cabinet inlet humidity. The original data is time-aligned on a unified reference time axis, including resampling data from different sampling periods, correcting and filling timestamp offsets, missing data, and inconsistencies in clocks from different data sources, so that each variable corresponds under the same time index. Outlier processing is performed on the raw data, including identifying outliers based on the reasonable value range of the device, the rate of change constraint and statistical discrimination, and correcting or removing the identified outliers. Based on this, feature extraction is performed to form a feature set for modeling. The feature set includes at least the current value of each original variable, several orders of historical lag values, rate of change and statistical characteristics, and includes a normalized or coded representation of the discrete device status, as well as combined features of cross-device and environmental variables constructed based on correlation analysis. Output an aligned dataset with a uniform time index.

3. The control method based on PUE dynamic modeling and reinforcement loop optimization according to claim 1, characterized in that, S2 specifically refers to: Using the aligned dataset as input, a sparse Gaussian process multi-output dynamic model for discrete-time prediction is constructed. The model uses features in the aligned dataset as independent variables, including at least the current value of the equipment control quantity and several orders of historical lag values, equipment state quantity and environmental quantity, and uses PUE, rack inlet temperature and rack inlet humidity as multiple outputs. The parameters and hyperparameters of the model are trained by maximum likelihood training or variational inference, and the computational complexity is reduced by using sparse approximation. Given the current state and the device's actions, the trained model is used to infer at least one prediction step, outputting the predicted values ​​and uncertainties of each output variable. The uncertainties include at least the prediction variance of each output variable and the covariance between the output variables. And calculate the action sensitivity matrix of equipment actions to PUE. The elements of the action sensitivity matrix are defined as the partial derivatives of the predicted PUE value with respect to each equipment control quantity under a given state, or the numerical sensitivity obtained by applying a small perturbation to the equipment control quantity. The action sensitivity matrix is ​​calculated at the current prediction step or summarized over multiple prediction steps. The predicted value, the uncertainty, and the action sensitivity matrix are used as the model prediction results.

4. The control method based on PUE dynamic modeling and reinforcement loop optimization according to claim 1, characterized in that, S3 specifically refers to: Based on the model prediction results, interference assessment is performed on the device action and at least one state response within a set sliding time window. Specifically, three types of indicators are calculated simultaneously and fused to generate an interference index. The first type of indicator is frequency domain coherence and phase margin. The second type of indicator is the proportion of actions in the opposite direction; The third type of indicator is the proportion of the local controller output of the device that is close to saturation. The interference index is obtained by weighting and fusing the above three types of indicators; Based on the model uncertainty, set the trigger threshold and release threshold with hysteresis, and output the interference evaluation result, which includes at least the interference index.

5. The control method based on PUE dynamic modeling and reinforcement loop optimization according to claim 1, characterized in that, S4 specifically refers to: Using the interference assessment results as input, token-based rotation scheduling is performed. Within each scheduling cycle, tokens are issued, renewed, and revoked based on the interference index, device coupling relationship, and current holding status. The set of devices holding tokens is determined, and a cooldown period parameter and a minimum holding time parameter associated with token management are generated for each device. The cooldown period parameter is used to limit that a token cannot be reissued during the cooldown period after it is revoked, and the minimum holding time parameter is used to limit that a token cannot be revoked within the time after it is issued to suppress frequent switching. The determination of equipment coupling relationships is based on correlation analysis of historical data or on action sensitivity matrices; When multiple devices compete for a token, arbitration is conducted based on the interference index and preset priority to ensure rotation and fairness. On a given set of devices holding tokens, the hierarchical reinforcement learning weighting module calculates a cost weight set and a reference trajectory bias set based on the interference index and the action sensitivity matrix. The cost weight set is used to adjust the weight of the cost term in the subsequent optimization objective, and the reference trajectory bias set is used to adjust the reference value offset of each device or each state. To ensure stability and feasibility, the weighting amplitude is adaptively shrunk according to the magnitude of the interference index. The shrinkage has a monotonic relationship with the interference index and has upper and lower bounds. Output the weight adjustment result, which includes at least the set of devices holding tokens, the set of cost weights, and the set of reference trajectory biases.

6. The control method based on PUE dynamic modeling and reinforcement loop optimization according to claim 1, characterized in that, S5 specifically refers to: Using the weighting results as input, the virtual friction coefficient and the upper limit of jerk are adaptively calculated based on the disturbance index and a preset mapping relationship. The virtual friction term is incorporated into the objective function of the model predictive control, so that the control increment is penalized with respect to the virtual friction coefficient. The upper limit of jerk is incorporated into the constraint set of model predictive control, and the constraint set explicitly includes at least the following constraints: minimum start-stop time constraint for equipment with start-stop characteristics, action slope limit for continuously adjustable equipment, and gear discreteness constraint for equipment with discrete gears. Under the aforementioned objective function and constraint set, the optimization problem is solved in the prediction time domain to obtain candidate control actions; The candidate control actions are input into the predictive safety filtering module for safety projection. A safety set is constructed based on the safety boundaries of the equipment and the environment and tightened according to the model uncertainty, so that the control actions after safety projection satisfy the safety set and prediction constraints, and the safe control actions are output. The cost weight set and reference trajectory bias set in the weighting result are merged with the safety control action to form the execution input set.

7. The control method based on PUE dynamic modeling and reinforcement loop optimization according to claim 1, characterized in that, S6 specifically refers to: The automatic control closed loop of the input set input device local controller will be executed, and injection and execution will only be performed on devices holding tokens; The reference trajectory bias set is applied to the set value of the corresponding device to form a corrected reference trajectory, and the cost weight set is applied to the target weight, logic threshold or priority of the local controller to adjust the control decision; Safety control actions are issued as external control commands and given a higher priority in the output synthesis of the local controller, so that they can cover or superimpose the free response of the local controller within the scope of safety constraints; The local controller includes proportional-integral control or logic control, which generates the final execution instruction and completes the regulation in the closed loop based on real-time feedback.

Citation Information

Patent Citations

  • Data center energy efficiency optimization method and system based on reinforcement learning

    CN114970358A

  • System, controller, and method for predictive control of energy management for a segmented load centre

    US20230275438A1