A server intelligent scheduling control method
By identifying the dynamic transfer function of the server cluster online and constructing a dynamic safe operating envelope domain, combined with model predictive control and multivariate collaborative tracking control, the problems of temperature overshoot and system instability of the server cluster under high-frequency fluctuations were solved, and the stability and energy efficiency were improved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- 百信信息技术有限公司
- Filing Date
- 2026-02-25
- Publication Date
- 2026-06-30
AI Technical Summary
Existing server cluster control methods cannot effectively cope with high-frequency fluctuations in input stimuli, leading to temperature overshoot and oscillations. Furthermore, they cannot adapt to changes in the physical characteristics of the controlled object, resulting in divergent control accuracy and system instability.
By identifying the dynamic transfer function matrix of the controlled unit online, extracting the thermal response lag time constant and steady-state gain parameters, constructing the dynamic safe operation envelope domain, and combining model predictive control and multivariable collaborative tracking control, closed-loop tracking and adaptive correction are achieved to generate the optimal control command.
It achieves dynamic stability and safety under load fluctuation conditions, suppresses temperature overshoot, improves the robustness and adaptive adjustment accuracy of the control system, and ensures the thermal safety and energy efficiency of the server cluster.
Smart Images

Figure CN122308040A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of industrial process control technology, specifically to a server intelligent scheduling and control method. Background Technology
[0002] With the intensive development of large-scale computing infrastructure, modern server clusters have evolved into thermo-fluidic physical systems with highly nonlinear and strongly coupled characteristics. When dealing with high-frequency fluctuating input stimuli, the controlled units are not only processing nodes for logic data but also physical sources of energy consumption and heat generation. Fluctuations in computing load are instantaneously converted into heat energy through changes in current, while the response of the cooling system is physically limited by the material's heat capacity and the environmental thermal resistance, exhibiting a significant hysteresis effect. Achieving a dynamic balance between input throughput and system response while ensuring the thermal safety and electrical stability of the physical hardware is a severe control engineering challenge in this field.
[0003] Most existing control methods employ open-loop regulation based on static rules or simple negative feedback strategies. This control mode has significant and substantial drawbacks: First, the control logic is disconnected from the physical process. Traditional methods ignore the physical thermal inertia and response hysteresis characteristics of the controlled object, adjusting only based on the current instantaneous state. This easily leads to significant temperature overshoot or oscillations in the system when the input flow fluctuates drastically. Second, the model parameters are rigid and cannot adapt to time-varying environments. Existing strategies typically assume that the physical characteristics of the controlled object are constant, failing to detect system transfer function drift caused by dust accumulation, aging of the heat dissipation medium, or fan wear. This results in a gradual divergence in control accuracy as operating time increases, ultimately leading to increased steady-state error or even system instability. Summary of the Invention
[0004] To address the shortcomings of existing technologies, this invention provides a server intelligent scheduling and control method, which solves the problems mentioned in the background.
[0005] To achieve the above objectives, the present invention provides the following technical solution: a server intelligent scheduling and control method, comprising the following steps: S1. Synchronously differentially sampling the input excitation flow signal and thermal process variable response signal of each controlled unit in the server cluster, identifying the dynamic transfer function matrix of each controlled sub-unit online through a recursive least squares estimation algorithm with a forgetting factor, and extracting the thermal response lag time constant and steady-state gain parameter characterizing the current physical thermal properties of each controlled unit; S2. Performing step response stability analysis based on the thermal response lag time constant, calculating the maximum allowable input rheotropic rate of each controlled sub-unit, and simultaneously mapping the thermal load saturation boundary of each controlled sub-unit through the steady-state gain parameter, combining the maximum input rheotropic rate with the thermal load saturation boundary. S3. Construct a dynamic safety operation envelope domain for the next control cycle adjustment action under constraints; S4. Obtain the current total input throughput setpoint of the system, and within the constraints of the dynamic safety operation envelope domain, establish a model predictive control objective function including energy consumption index and tracking error index, solve for the optimal flow distribution setpoint and energy efficiency operating condition command of each controlled subunit, and send them to the underlying execution control unit; S5. During the execution control cycle, drive the front-end flow distribution gateway and node voltage regulation module in coordination according to the setpoint and command through a multivariable collaborative tracking control algorithm, so as to track and control the service request injection rate and chip core operating voltage amplitude of each controlled subunit in a closed loop, and use the state observer to calculate the observation residual to correct the forgetting factor in S1 online.
[0006] Furthermore, the specific process of identifying the dynamic transfer function matrix of each controlled sub-unit online using the recursive least squares estimation algorithm with a forgetting factor is as follows: A data regression vector containing the increments of historical input excitation flow signals and historical thermal process variable response signals is constructed; the prior prediction error between the actual output of the system at the current moment and the predicted output based on the parameter estimates of the previous moment is calculated; the orthogonal gain correction vector at the current moment is calculated using the forgetting factor and the covariance matrix of the previous moment; the coefficient parameter vector of the discrete-time transfer function matrix is iteratively updated according to the product of the prior prediction error and the orthogonal gain correction vector; the update step size of the covariance matrix is dynamically adjusted according to the excitation intensity of the data regression vector to obtain a parameter-converged discrete-time transfer function matrix.
[0007] Furthermore, the specific process for extracting the thermal response lag time constant and steady-state gain parameter characterizing the current physical thermal properties of each controlled unit is as follows: Perform an inverse bilinear transformation from the Z-domain to the S-domain on the discrete-time transfer function matrix, mapping the discrete coefficients to the system pole distribution in the continuous-time domain; extract the dominant pole closest to the imaginary axis of the S-plane, and calculate the absolute value of the reciprocal of the real part of the dominant pole as the thermal response lag time constant; perform steady-state limit operation on the discrete-time transfer function matrix using the final value theorem, and calculate the output amplitude of the unit step response when the time approaches infinity as the steady-state gain parameter.
[0008] Furthermore, based on the thermal response hysteresis time constant, a step response stability analysis is performed to calculate the maximum allowable input rheological rate of each controlled subunit. Simultaneously, the specific process of mapping the thermal load saturation boundary of each controlled subunit using the steady-state gain parameter is as follows: Using the thermal response hysteresis time constant as the first derivative decay factor, the maximum allowable step response rise slope of each controlled subunit under critical damping state is analytically calculated and defined as the maximum input rheological rate. The difference between the current ambient temperature of each controlled subunit and the upper limit of the chip's physical tolerance temperature is obtained as the thermal capacity margin. The thermal capacity margin is divided by the steady-state gain parameter to calculate the maximum absolute input flux allowed to be loaded by each controlled subunit under steady-state equilibrium conditions, which is then used as the thermal load saturation boundary.
[0009] Furthermore, the specific process of constructing the dynamic safe operating envelope domain for constraining the adjustment action in the next control cycle by combining the maximum input rheological rate with the thermal load saturation boundary is as follows: In the state space of the control variable, with the current input flux value as the reference point, the dynamic reachable cone region of the next control cycle is defined by the maximum input rheological rate; a static absolute amplitude limiting hyperplane is defined in the state space of the control variable by the thermal load saturation boundary, and a multidimensional geometric intersection operation is performed on the dynamic reachable cone region and the static absolute amplitude limiting hyperplane, and the closed convex polyhedral space overlapping the two is extracted as the dynamic safe operating envelope domain.
[0010] Furthermore, to obtain the current total input flux setpoint of the system, and within the constraints of the dynamic safety operation envelope domain, the specific process of establishing a model predictive control objective function including energy consumption and tracking error indices is as follows: Set the prediction time domain length and control time domain length of the controller; extend the total input flux setpoint to a reference trajectory sequence within the future prediction time domain; define the tracking error index as the square of the Euclidean norm between the sum of the predicted output flows of each controlled subunit and the reference trajectory sequence; define the energy consumption index as the quadratic product of the control input vector of each controlled subunit and the energy consumption weighting matrix; transform the dynamic safety operation envelope domain into a set of linear inequality constraints of state variables within the prediction time domain; perform a linear weighted summation of the tracking error index and the energy consumption index; and construct a quadratic programming optimization objective function under the constraints of the set of linear inequality constraints of state variables.
[0011] Furthermore, the specific process of solving for the optimal flow allocation setpoint and energy efficiency operation command of each controlled subunit and sending them to the underlying execution control unit is as follows: The quadratic programming objective function is optimized by a numerical optimization solver through rolling optimization to calculate the optimal control increment sequence in the future control time domain; the first element of the optimal control increment sequence is extracted as the optimal cooperative control vector at the current moment according to the rolling time domain control principle; the optimal cooperative control vector is decomposed into a flow truncation threshold for the front-end flow distribution gateway and a voltage amplitude adjustment code for the node voltage regulation module, which are respectively sent out in parallel through the hardware driver interface as the optimal flow allocation setpoint and energy efficiency operation command.
[0012] Furthermore, the specific process of using a multivariable collaborative tracking control algorithm to collaboratively drive the front-end traffic distribution gateway and node voltage regulation module according to set values and instructions, and to achieve closed-loop tracking control of the service request injection rate and the operating voltage amplitude of the chip core in each controlled sub-unit, is as follows: The optimal traffic allocation set value is converted into the token generation rate of the token bucket algorithm in the front-end traffic distribution gateway, and the service request injection rate entering each controlled sub-unit is physically limited by adjusting the token generation rate; the energy efficiency operating condition command is converted into the pulse width modulation signal of the node voltage regulation module, which drives the power switch to adjust the operating voltage amplitude output to the chip core; the response delay time of the voltage regulation process is monitored, and transient hold logic is applied to the token generation rate before the voltage reaches the target value, so as to achieve time synchronization and coordination between the traffic injection action and the voltage regulation action.
[0013] Furthermore, the specific process of online correction of the forgetting factor in S1 by calculating the observation residuals through the state observer is as follows: the actual output of each controlled subunit is subtracted from the predicted output of the dynamic transfer function matrix to obtain the instantaneous observation residual vector; the modulus of the instantaneous observation residual vector is calculated, and a nonlinear inverse proportional mapping function is established between the forgetting factor and the observation residual modulus; when the observation residual modulus increases, the value of the forgetting factor is reduced through the nonlinear inverse proportional mapping function to enhance the sensitivity of the algorithm to new data; when the observation residual modulus decreases, the value of the forgetting factor is increased to enhance the smoothness of parameter estimation, thus completing the adaptive update of the recursive least squares estimation algorithm.
[0014] The present invention has the following beneficial effects:
[0015] (1) A server intelligent scheduling and control method, which extracts the lag time constant and gain parameter representing the current physical and thermal characteristics of the controlled unit in real time, and analyzes and calculates the maximum allowable input rheological rate and thermal load saturation boundary of the system, thereby constructing a dynamic safe operating envelope domain for constrained control actions. The advantage of this mechanism is that it can dynamically adjust the control constraint range according to the current actual health status of the controlled object, introduce physical limitations at the source of control command generation, effectively suppress transient temperature overshoot and system oscillation caused by thermal inertia, and ensure the dynamic stability and safety of the control system under conditions of severe load fluctuations. Through online identification and dynamic boundary construction, the problem of control instability caused by time-varying physical characteristics and thermal response lag of the controlled object is solved.
[0016] (2) A server intelligent scheduling and control method solves the objective function containing energy consumption and tracking error indices within the dynamic safe operation envelope domain, and performs online adaptive correction of model parameters by combining the residual feedback mechanism of the state observer. This closed-loop control strategy, which combines feedforward optimization decision-making with feedback tracking correction, not only achieves optimal trajectory tracking under physical constraints, but also automatically compensates for model mismatch caused by hardware aging or environmental thermal disturbances by adjusting the forgetting factor online, completely eliminating the steady-state control deviation of the system and significantly improving the robustness and adaptive adjustment accuracy of the control system to the drift of the controlled object parameters. Through model prediction optimization and state observation closed loop, the problems of difficulty in multi-objective cooperative control and poor system anti-disturbance capability are solved.
[0017] Of course, any product implementing this invention does not necessarily need to achieve all of the advantages described above at the same time. Attached Figure Description
[0018] Figure 1 This is a flowchart of a server intelligent scheduling and control method according to the present invention. Detailed Implementation
[0019] This application provides a server intelligent scheduling and control method that solves the problems of poor thermal stability and low energy efficiency caused by unknown model parameters of the controlled object and environmental disturbances in a distributed coupled physical system.
[0020] The overall approach of the scheme in this application is as follows: The server cluster is treated as a multivariable coupled controlled physical system. First, the dynamic transfer function of each unit is extracted online using system identification theory, and its thermal inertia and gain parameters are quantified. Then, a dynamic safe operating envelope domain, including rheological rate and amplitude limits, is defined based on the physical parameters. Within this safe domain, model predictive control is used to solve for the optimal flow and voltage control commands. Finally, a closed-loop feedback is formed through a state observer and parameter adaptive mechanism to eliminate the influence of environmental disturbances and model biases in real time, achieving high-precision coordinated control of the server cluster's thermal process variables and input excitation signals.
[0021] Please see Figure 1 This invention provides a technical solution: a server intelligent scheduling and control method, comprising the following steps: S1. Synchronously differentially sampling the input excitation flow signal and thermal process variable response signal of each controlled unit in the server cluster, identifying the dynamic transfer function matrix of each controlled sub-unit online through a recursive least squares estimation algorithm with a forgetting factor, and extracting the thermal response lag time constant and steady-state gain parameter characterizing the current physical thermal characteristics of each controlled unit; S2. Performing step response stability analysis based on the thermal response lag time constant, calculating the maximum allowable input rheotropic rate of each controlled sub-unit, and mapping the thermal load saturation boundary of each controlled sub-unit through the steady-state gain parameter, combining the maximum input rheotropic rate and the thermal load saturation boundary to construct approximately S3. Obtain the current total input throughput setpoint of the system. Within the constraints of the dynamic safety operation envelope, establish a model predictive control objective function that includes energy consumption indicators and tracking error indicators. Solve for the optimal flow distribution setpoint and energy efficiency operating instructions for each controlled subunit and send them to the underlying execution control unit. S4. During the execution control cycle, use a multivariate collaborative tracking control algorithm to collaboratively drive the front-end flow distribution gateway and node voltage regulation module according to the setpoint and instructions. Use closed-loop tracking control to control the service request injection rate and chip core operating voltage amplitude of each controlled subunit. Use the state observer to calculate the observation residual and correct the forgetting factor in S1 online.
[0022] In this implementation scheme, step S1 is mainly used to construct a real-time mathematical model of the controlled object. By synchronously differentiating and sampling the input excitation flow signal and the thermal process variable response signal, the transient characteristics of the system during dynamic changes can be captured, rather than focusing only on the static amplitude. The dynamic transfer function matrix refers to a mathematical model describing the dynamic characteristics of a multi-input multi-output system in the complex frequency domain. It reflects how changes in the input signal are transformed into the output response through the hysteresis and attenuation of the physical system. The recursive least squares estimation algorithm with a forgetting factor is an online parameter identification method. The forgetting factor is used to assign higher weights to new data and gradually discard old data, enabling the algorithm to track system parameters that change over time. The technical role of this step is to use online identification technology to quantify the physical thermal characteristic drift of the controlled unit caused by hardware aging or environmental changes in real time, extracting accurate thermal response hysteresis time constants (reflecting the magnitude of thermal inertia) and steady-state gain parameters (reflecting input-output conversion efficiency), providing a model benchmark for subsequent precise control. Step S2 is mainly used to define the safe operating range of the controller to prevent system instability caused by control commands exceeding physical limits. Step response stability analysis refers to simulating the transient response of a system under ideal step signal excitation using mathematical methods to assess the system's damping characteristics and overshoot risk. Maximum input rheological rate refers to the maximum rate at which the input signal can change per unit time; exceeding this rate will cause temperature overshoot due to thermal inertia preventing timely heat dissipation. The dynamic safe operating envelope domain refers to a closed geometric region in a multidimensional state space enclosed by a series of linear or nonlinear constraints. The technical role of this step is to transform the abstract physical parameters extracted in S1 into hard constraint boundaries that the controller can execute. By limiting the input rheological rate, dynamic overheating is suppressed; by limiting the thermal load saturation boundary, steady-state overload is prevented, thereby constructing a dynamic feasible domain that maximizes system performance while ensuring thermal safety. Step S3 is mainly used to generate optimal control decisions, finding the optimal operating point of the system while satisfying safety constraints. The model predictive control objective function is a mathematical expression used to quantify the quality of system control performance, typically including a weighted sum of energy consumption indicators (representing operating costs) and tracking error indicators (representing control accuracy). The optimal flow allocation setpoint and energy efficiency operating condition command are ideal reference values output by the control algorithm, used to guide the actions of the underlying actuators. The technical role of this step is to utilize the forward-looking predictive capability of Model Predictive Control (MPC) to perform rolling optimization within the dynamic safe operating envelope domain constructed in S2, automatically balancing the trade-off between energy consumption and performance, and calculating the control law that minimizes the overall system cost at future moments, thus achieving globally optimal scheduling in a multivariable coupled system. Step S4 is mainly used to implement the control command and perform closed-loop adaptive correction, ensuring that the actual operating trajectory of the physical system closely follows the setpoint generated in S3.Among them, the multivariable cooperative tracking control algorithm refers to a low-level control strategy used to drive multiple actuators to cooperate in order to eliminate tracking errors. The state observer is a virtual sensor algorithm based on the system model to reconstruct the unmeasurable states within the system. The observation residual refers to the deviation between the actual output value of the system and the predicted value of the observer. The technical role of this step is twofold: firstly, by driving physical actuators such as the front-end flow distribution gateway and the node voltage regulation module, the logical setpoint is transformed into the physical injection rate and voltage amplitude, achieving closed-loop control; secondly, the observation residual reflects the degree of mismatch between the model and the actual system, and the sensitivity of the identification algorithm is dynamically adjusted by online correction of the forgetting factor, eliminating steady-state errors caused by environmental disturbances and ensuring the robustness of the control system throughout its entire lifecycle.
[0023] Specifically, the process of online identification of the dynamic transfer function matrix of each controlled sub-unit using the recursive least squares estimation algorithm with a forgetting factor is as follows: A data regression vector containing the increments of historical input excitation flow signals and historical thermal process variable response signals is constructed; the prior prediction error between the actual output of the system at the current moment and the predicted output based on the parameter estimates of the previous moment is calculated; the orthogonal gain correction vector at the current moment is calculated using the forgetting factor and the covariance matrix of the previous moment; the coefficient parameter vector of the discrete-time transfer function matrix is iteratively updated according to the product of the prior prediction error and the orthogonal gain correction vector; the update step size of the covariance matrix is dynamically adjusted according to the excitation intensity of the data regression vector to obtain a parameter-converged discrete-time transfer function matrix.
[0024] In this implementation scheme, firstly, a data regression vector is constructed to structure the historical input-output data of the controlled subunit over time, making it conform to the matrix operation format of the recursive algorithm. The data regression vector is essentially a memory carrier of the system's historical behavior, containing the flow rate increment and temperature response increment of the previous moment, reflecting the current dynamic trend of the system. Based on this, the prior prediction error is calculated, that is, the output of the current moment is predicted using the model parameters estimated at the previous moment, and this predicted value is compared with the actual values collected by the sensors. The technical role of this step is to quantify the degree of deviation between the current model and the actual physical system; this deviation is the core driving force for parameter updates. Next, the orthogonal gain correction vector is calculated and the covariance matrix is updated using the core formula of the recursive least squares algorithm with a forgetting factor. This process is the soul of adaptive control, where the orthogonal gain correction vector determines the direction and step size of the model parameters to be adjusted under the current error, while the covariance matrix reflects the uncertainty of parameter estimation. To overcome the defect of traditional algorithms that cannot track time-varying parameters due to data saturation, a forgetting factor is introduced to exponentially decay the weight of historical data. The specific iterative update process can be described by the following set of state equations: ; ; The meanings of each parameter are as follows: : The current discrete sampling time; : The orthogonal gain correction vector calculated at the current time; : The covariance matrix stored in the previous time step; The constructed data regression vector contains the difference values of historical input and output data; Forgetting factor: This is used to adjust the algorithm's sensitivity to new and old data. Its value is usually determined to be between 0.95 and 0.99 based on the system noise level. The vector of coefficient parameters of the discrete-time transfer function matrix updated at the current time. The actual thermal process variable response signal collected by the sensor at the current moment; The identity matrix is used in the above calculation process. The technical significance lies in projecting the prior prediction error onto the parameter space using an orthogonal gain correction vector, thus correcting the coefficient parameter vector. This allows the corrected model to more accurately fit the input-output relationship at the current moment. Simultaneously, the dynamic adjustment of the covariance matrix ensures that the algorithm can converge quickly when the system experiences abrupt changes, thereby obtaining a discrete-time transfer function matrix that accurately describes the dynamic mapping relationship between input and output under the current operating conditions.
[0025] Specifically, the process of extracting the thermal response lag time constant and steady-state gain parameter that characterize the current physical thermal properties of each controlled unit is as follows: Perform an inverse bilinear transformation from the Z-domain to the S-domain on the discrete-time transfer function matrix to map the discrete coefficients to the system pole distribution in the continuous-time domain; extract the dominant pole closest to the imaginary axis of the S-plane and calculate the absolute value of the reciprocal of the real part of the dominant pole as the thermal response lag time constant; perform steady-state limit operation on the discrete-time transfer function matrix through the final value theorem and calculate the output amplitude of the unit step response when the time approaches infinity as the steady-state gain parameter.
[0026] In this implementation scheme, firstly, an inverse bilinear transformation from the Z-domain to the S-domain is performed on the discrete-time transfer function matrix. This is because the recursive least squares method yields a difference equation model based on discrete sampling points (Z-domain), while physical thermal properties (such as thermal inertia) are physical quantities in the continuous-time domain (S-domain). The inverse bilinear transformation (also known as the inverse process of the Tustin transform) is a mathematical method that maps the complex plane of a discrete system back to the complex plane of a continuous system while minimizing frequency distortion. Through this transformation, the discrete coefficients can be mapped to the distribution of system poles in the continuous-time domain. The position of the poles in the complex plane directly determines the dynamic modes of the system, where the real part reflects the decay rate and the imaginary part reflects the oscillation frequency. Next, the dominant poles are extracted and the thermal response lag time constant is calculated. In multi-order systems, the pole closest to the imaginary axis of the S-plane is called the dominant pole, and its corresponding transient response decays the slowest, playing a bottleneck role in determining the overall response speed of the system. For thermodynamic systems, this characteristic directly corresponds to the thermal inertia of the system. The formula for calculating the thermal response lag time constant is as follows: The meanings of each parameter are as follows: The extracted thermal response hysteresis time constant characterizes the degree of hysteresis in the temperature response of the controlled unit; Mathematical operators that take the real part of a complex number; The dominant pole closest to the imaginary axis is obtained through inverse bilinear transformation mapping. The technical significance of this step lies in transforming the abstract mathematical pole into a time parameter with explicit physical meaning. This parameter directly tells the controller how long it takes for the temperature to begin to change significantly when the load changes, thus providing a physical basis for subsequent calculations of the maximum input rheological rate. Finally, the steady-state gain parameter is calculated using the final value theorem. The steady-state gain reflects the change in output caused by a unit change in input after the system reaches thermal equilibrium, i.e., the system's heat-to-fluid conversion efficiency. The final value theorem allows us to calculate the steady-state value directly through the frequency domain limit without performing time-domain integration. The formula for calculating the steady-state gain parameter is as follows: The meanings of each parameter are as follows: The calculated steady-state gain parameters characterize the thermal sensitivity of the controlled unit. : A complex variable in the Z-transform domain, when it approaches 1, represents a frequency approaching 0, i.e., a DC steady-state component; The identified discrete-time transfer function; : The i-th coefficient of the numerator polynomial of the discrete-time transfer function; The j-th coefficient of the denominator polynomial of the discrete-time transfer function; The order of the numerator polynomial; The order of the denominator polynomial. Through the above calculations, the complex transfer function matrix is simplified to two core physical parameters: (Fast and Slow) and (Size) This method completes the dimensionality reduction extraction from the black-box model to the physical feature parameters, providing a direct quantitative indicator for constructing a dynamic safe operation envelope domain.
[0027] Specifically, the following process is used to perform step response stability analysis based on the thermal response hysteresis time constant, calculate the maximum allowable input rheological rate of each controlled subunit, and map the thermal load saturation boundary of each controlled subunit through the steady-state gain parameter: Based on the thermal response hysteresis time constant as the first derivative decay factor, the maximum rise slope of the step response allowed by each controlled subunit under critical damping state is analytically calculated and defined as the maximum input rheological rate; the difference between the current ambient temperature of each controlled subunit and the upper limit of the chip's physical tolerance temperature is obtained as the thermal capacity margin, and the thermal capacity margin is divided by the steady-state gain parameter to calculate the maximum absolute input flux allowed to be loaded by each controlled subunit under steady-state equilibrium conditions as the thermal load saturation boundary.
[0028] In this implementation scheme, firstly, the thermal response hysteresis time constant is used as the first derivative decay factor for analytical calculation. The core of this step lies in transforming the abstract time constant into a specific control rate constraint. In control theory, the critical damped state refers to the ideal state in which the system reaches steady state at the fastest speed without oscillation during a step response. If the rate of increase of the control input exceeds the thermal response speed of the physical system, heat that cannot be dissipated in time will accumulate inside the system, leading to temperature overshoot. Therefore, the maximum input rheological rate is essentially a dynamic constraint on the rate of change of the input signal by the physical system. It is defined by the following inequality equation: The meanings of each parameter are as follows: The calculated maximum input rheological rate represents the upper limit of the load injection rate allowed by the controlled subunit, which is the flow increment per unit time. The upper limit of the physical temperature tolerance of a chip is determined by the technical specifications provided by the hardware manufacturer. The core temperature of the controlled subunit is currently being collected in real time by the sensor. Steady-state gain parameters extracted in the preceding steps; : The thermal response hysteresis time constant extracted in the preceding steps; The preset critical damping safety factor, typically ranging from 0.6 to 0.8, is used to maintain a safety margin beyond the theoretical critical value, preventing noise interference from triggering the boundary. The technical meaning of the above formula is that the numerator represents the remaining temperature space, and the denominator represents the system's physical thermal inertia (the product of gain and time constant). The physical meaning of this ratio is clear: the greater the system's thermal inertia, or the smaller the remaining temperature space, the smaller the allowable rate of change of input must be. Simultaneously, the heat load saturation boundary is calculated. This belongs to the steady-state constraint calculation of the system, and its purpose is to determine the total absolute load that the system can withstand under long-term operation. The heat capacity margin reflects the heat dissipation potential under the current environment. The formula for calculating the heat load saturation boundary is as follows: The meanings of each parameter are as follows: The calculated thermal load saturation boundary characterizes the maximum absolute input flux allowed to be loaded into the controlled sub-unit; The current ambient temperature at the air inlet is collected. Physical temperature limits (e.g., not exceeding 85 degrees Celsius) are mapped back to the domain of control variables (e.g., maximum allowable TPS flow rate or watts of power), allowing subsequent controllers to operate directly within the input space without repeatedly probing temperature boundaries.
[0029] Specifically, the process of constructing the dynamic safe operating envelope domain for constraining the adjustment action in the next control cycle by combining the maximum input rheological rate with the thermal load saturation boundary is as follows: In the state space of the control variable, with the current input flux value as the reference point, the dynamic reachable conical region of the next control cycle is defined by the maximum input rheological rate; a static absolute amplitude limiting hyperplane is defined in the state space of the control variable by the thermal load saturation boundary; a multidimensional geometric intersection operation is performed on the dynamic reachable conical region and the static absolute amplitude limiting hyperplane; and the closed convex polyhedral space overlapping the two is extracted as the dynamic safe operating envelope domain.
[0030] In this implementation plan, firstly, we understand the concepts of state space and dynamically reachable cone region. State space is a multi-dimensional geometric space describing all possible behaviors of the control system. Taking the current input flux value as a reference point, since the maximum input rheological rate calculated in the previous step limits the rate of change of the control quantity, the feasible control quantity in the next step must fall within a cone region with the current point as its vertex and an opening size determined by the rheological rate. This region is called the dynamically reachable cone region, representing the range that the physical system can reach in one step. Secondly, we understand the static absolute amplitude limiting hyperplane. The thermal load saturation boundary is an absolute value; regardless of the current state, the control quantity cannot exceed this absolute limit. In the multi-dimensional state space, this absolute limit is represented by a hyperplane that cuts through the space. The process of constructing the dynamic safe operating envelope domain is essentially solving for the intersection of the two geometric spaces mentioned above. Its mathematical description is as follows: The specific solution to this set is expressed in matrix form using a system of linear inequalities: The meanings of each parameter are as follows: The constructed dynamic safety operation envelope for the next control cycle; The vector of control variables to be solved; A real space of dimension m, where m is the dimension of the controlled variable; The actual control quantity applied at the current moment; The sampling time interval of the control period; Possible control vectors at the next moment; The constraint matrix, composed of the maximum input rheological rate and the identity matrix, defines the normal vector direction of the polyhedron. : Constraint boundary vector, containing the... The calculated dynamic upper bound, by The dynamic lower bound of computation and A static absolute upper bound is defined. The technical significance of this step lies in transforming complex physical constraints into a set of standard linear inequality constraints. This closed convex polyhedral space (dynamically safe operating envelope) precisely defines the control range that is both safe and reachable. For subsequent model predictive control (MPC), this means that the optimization solver only needs to search within this convex polyhedron and will inevitably find a solution that is physically safe, thus avoiding the problem of solution divergence or control command out-of-bounds errors caused by improper constraint settings in traditional methods.
[0031] Specifically, the process of obtaining the current total input flux setpoint of the system and establishing a model predictive control objective function including energy consumption and tracking error indices within the constraints of the dynamic safety operation envelope domain is as follows: The prediction time domain length and control time domain length of the controller are set, and the total input flux setpoint is extended to a reference trajectory sequence within the future prediction time domain; the tracking error index is defined as the square of the Euclidean norm between the sum of the predicted output flows of each controlled subunit and the reference trajectory sequence, and the energy consumption index is defined as the quadratic product of the control input vector of each controlled subunit and the energy consumption weighting matrix; the dynamic safety operation envelope domain is transformed into a set of linear inequality constraints of state variables within the prediction time domain, and the tracking error index and energy consumption index are linearly weighted and summed to construct a quadratic programming optimization objective function under the constraints of the set of linear inequality constraints of state variables.
[0032] In this implementation scheme, firstly, the prediction time domain length and control time domain length of the controller are defined. In model predictive control theory, the prediction time domain refers to how far into the future the controller can look; a longer prediction time domain helps the system perceive future constraints and changes in advance. The control time domain refers to the number of steps the controller plans to take in the future. Extending the total input flux setpoint to a reference trajectory sequence within the future prediction time domain essentially constructs an ideal curve that the system output is expected to follow. The technical role of this step is to provide the controller with a clear navigation target, preventing it from blindly adjusting but instead guiding it to approach the target in a planned manner. Next, tracking error and energy consumption indices are defined. This is to transform the control task into a mathematical optimization problem. The tracking error index reflects whether the system can process business demands on time and in sufficient quantity; the square of the Euclidean norm is used to penalize large deviations. The energy consumption index reflects the electrical costs incurred; a quadratic product is used to impose a stronger penalty when the control input is large, thereby suppressing energy consumption spikes. To mathematically quantify this trade-off, the following quadratic programming optimization objective function is constructed: The meanings of each parameter are as follows: The objective function value of the quadratic programming optimization constructed at the current moment represents the comprehensive cost of system operation; The set prediction time domain length represents the number of steps the controller needs to predict future states; The set control time domain length indicates the number of steps the controller plans for the control increment; The sum of the output flow of each controlled sub-unit at the p-th future time predicted based on the information at the current time; The expected value of the expanded reference trajectory sequence at step p is usually generated by the total input flux setting. The tracking error weighting matrix, the parameters of which are determined by Bryson's rule based on the maximum allowable tracking error of the system; The control input increment vector at the c-th future time step to be solved; The energy consumption weighting matrix, whose parameters are determined based on the gradient of the energy efficiency ratio characteristic curve of each controlled sub-unit, is used. Finally, the dynamic safety operation envelope is transformed into a set of linear inequality constraints for state variables. The envelope obtained in the previous steps is a geometric space; to be recognizable by the optimization solver, it must be transformed into standard mathematical inequality form. This step plays a crucial role in hard-coding the physical safety boundary into the optimization problem, ensuring that no matter how the algorithm optimizes, its result will never exceed the safety limits allowed by the physical system.
[0033] Specifically, the process of solving for the optimal flow allocation setpoint and energy efficiency operation command of each controlled subunit and sending them to the underlying execution control unit is as follows: A numerical optimization solver performs rolling optimization on the quadratic programming objective function to calculate the optimal control increment sequence in the future control time domain; based on the rolling time domain control principle, the first element of the optimal control increment sequence is extracted as the optimal cooperative control vector at the current moment; the optimal cooperative control vector is decomposed into a flow truncation threshold for the front-end flow distribution gateway and a voltage amplitude adjustment code for the node voltage regulation module, which are then sent out in parallel through the hardware driver interface as the optimal flow allocation setpoint and energy efficiency operation command, respectively.
[0034] In this implementation scheme, firstly, a rolling optimization is performed on the quadratic programming objective function using a numerical optimization solver. The numerical optimization solver is an algorithm engine specifically designed for solving convex optimization problems. It can quickly find the sequence of control variables that minimizes the objective function while satisfying all linear inequality constraints. Rolling optimization means that the controller performs a full-process optimization again at each sampling time. This mechanism can promptly correct prediction deviations caused by model errors or environmental disturbances. After calculating the optimal control increment sequence in the future control time domain, the first element is truncated according to the rolling time domain control principle. This is the core strategy of model predictive control: only the most urgent step at the current moment is executed, while subsequent planning is discarded and re-planned based on the new system state at the next moment. This reflects the idea of combining the real-time performance of feedback control with the predictive nature of feedforward control. The optimal cooperative control vector at the current moment is calculated as follows: The meanings of each parameter are as follows: The calculated optimal cooperative control vector at the current moment includes the absolute control quantities of each controlled subunit; The control vector actually applied at the previous moment is used to ensure the continuity of control actions. The first vector element in the optimal control increment sequence output by the numerical optimization solver. Finally, the optimal cooperative control vector is decomposed and distributed. This is because the calculated... These are abstract numerical values in mathematics, which must be translated into specific instructions for physical devices. For the front-end traffic distribution gateway, the corresponding control components are converted into traffic truncation thresholds to physically limit the inbound rate of data packets at the network layer. For the node voltage regulation module, the corresponding control components are converted into voltage amplitude regulation codes (such as VID codes) to drive the power management chip to adjust the output voltage. Parallel distribution through the hardware driver interface ensures that the logical calculation results can be applied synchronously and accurately to the physical actuators, realizing the mapping from the algorithm space to the physical space.
[0035] Specifically, the process of using a multivariable collaborative tracking control algorithm to collaboratively drive the front-end traffic distribution gateway and node voltage regulation module based on setpoints and instructions, and to achieve closed-loop tracking control of the service request injection rate and the operating voltage amplitude of the chip core in each controlled sub-unit, is as follows: The optimal traffic allocation setpoint is converted into the token generation rate of the token bucket algorithm in the front-end traffic distribution gateway, and the service request injection rate entering each controlled sub-unit is physically limited by adjusting the token generation rate; the energy efficiency operating condition command is converted into a pulse width modulation signal of the node voltage regulation module, which drives the power switch to adjust the operating voltage amplitude output to the chip core; the response delay time of the voltage regulation process is monitored, and transient hold logic is applied to the token generation rate before the voltage reaches the target value, so as to achieve time synchronization and coordination between the traffic injection action and the voltage regulation action.
[0036] In this implementation scheme, firstly, the optimal flow allocation setting is transformed into the physical control parameters of the front-end flow distribution gateway. The front-end flow distribution gateway acts as a regulating valve in fluid control, while the token bucket algorithm is the specific execution logic controlling the valve opening. The token generation rate determines the upper limit of the number of data packets allowed to pass through the gateway per unit time, physically limiting the injection rate of service requests into the controlled subunit. The technical function of this step is to forcibly transform the abstract flow value calculated at the upper layer into a hard constraint at the physical network layer, preventing sudden flow surges from breaching the system's defenses. Next, the energy efficiency operating condition command is transformed into an electrical signal for the node voltage regulation module. The node voltage regulation module is a precision power supply actuator used to supply power to the chip core; it regulates the output voltage amplitude by rapidly switching the on and off states of power switching transistors (such as metal-oxide-semiconductor field-effect transistors). The duty cycle of the pulse width modulation signal directly determines the output voltage level. Subsequently, the timing synchronization of flow injection and voltage regulation is implemented. This is to address the physical risks caused by inconsistent response speeds of different actuators. In physical reality, voltage regulation modules require charging and discharging capacitors to adjust output voltage, resulting in a millisecond-level physical response delay; while flow gateways adjusting current limiting thresholds are purely logical operations with an almost instantaneous response. If flow increases before voltage, the chip will operate under high load at low voltage, leading to logic errors or system crashes. Therefore, transient hold logic must be introduced. This cooperative control logic is described by the following state determination formula: The meanings of each parameter are as follows: The actual token generation rate applied to the front-end traffic distribution gateway at any given moment determines the actual traffic injected into the system. The optimal flow allocation setting value obtained from the previous steps; The actual operating voltage amplitude of the chip core is monitored in real time through a voltage sampling circuit. The target voltage value corresponding to the energy efficiency operating condition command calculated in the previous steps; The preset voltage steady-state tracking tolerance band is typically set to one to three percent of the target voltage value according to the chip's power supply specifications. The transient hold rate is typically the token generation rate from the previous moment or the minimum flow rate required for the system to maintain a minimum level of activity. The technical function of this logic is to establish a physical interlocking mechanism, forcing the flow injection action to wait for the voltage regulation action to complete and stabilize within the target range before execution. This achieves strict timing matching between energy supply and load consumption within a microsecond-level control cycle, ensuring the electrical safety of the hardware during varying operating conditions.
[0037] Specifically, the process of online correction of the forgetting factor in S1 by calculating the observation residuals through the state observer is as follows: the actual output of each controlled subunit is subtracted from the predicted output of the dynamic transfer function matrix to obtain the instantaneous observation residual vector; the modulus of the instantaneous observation residual vector is calculated, and a nonlinear inverse proportional mapping function is established between the forgetting factor and the observation residual modulus; when the observation residual modulus increases, the value of the forgetting factor is reduced through the nonlinear inverse proportional mapping function to enhance the sensitivity of the algorithm to new data; when the observation residual modulus decreases, the value of the forgetting factor is increased to enhance the smoothness of parameter estimation, thus completing the adaptive update of the recursive least squares estimation algorithm.
[0038] In this implementation scheme, firstly, the instantaneous observation residual vector is calculated. The state observer acts as a calibrator in the control system, simulating the system's operation in parallel using the dynamic transfer function matrix established by S1. The difference between the actual outputs (such as actual temperature rise) collected by the sensors of each controlled subunit and the predicted output calculated by the observer based on the model is the observation residual. This residual reflects the degree of mismatch between the current mathematical model and the actual physical object. Next, a nonlinear mapping relationship is established between the forgetting factor and the modulus of the observation residual. The forgetting factor is a key parameter determining the memory length in the recursive least squares algorithm. When the system is running smoothly, the model parameters change slowly, requiring a larger forgetting factor to utilize more historical data to smooth out noise; when the system is disturbed by the environment or experiences sudden changes, the model error increases, requiring a smaller forgetting factor to quickly forget old data, making the algorithm more sensitive to new data, thereby quickly tracking the physical changes of the system. This adaptive adjustment mechanism is achieved through the following nonlinear inverse proportional mapping function: The meanings of each parameter are as follows: The calculated adaptive forgetting factor at the current time is used to update the recursive least squares estimation algorithm in S1; The preset lower limit of the forgetting factor is used to prevent the forgetting factor from becoming too small under severe perturbation, which would cause the algorithm value to be unstable. It is usually set between 0.90 and 0.95. Sensitivity adjustment coefficient, used to set the algorithm's response gain to residual changes. This parameter is determined based on the system signal-to-noise ratio through offline simulation experiments. The exponential decay rate parameter determines the steepness of the mapping curve; The Euclidean modulus of the instantaneous observation residual vector calculated at the current moment characterizes the magnitude of the comprehensive model error. A closed-loop feedback control circuit is constructed using the above function, directly feeding back the accuracy of the model prediction to the parameter identification algorithm itself. When the characteristics of the actual physical system drift due to aging, dust accumulation, or external heat source interference, the increased residual automatically triggers a reduction in the forgetting factor, causing the model parameters to converge quickly to the new physical true values. This achieves adaptive correction throughout the entire lifecycle of the control system, eliminating the steady-state control error that inevitably occurs in fixed-parameter models after long-term operation.
[0039] In summary, this application has at least the following effects:
[0040] A server intelligent scheduling and control method identifies the physical thermal characteristic parameters of the controlled object online and constructs a dynamic safe operating envelope domain. This method proactively suppresses transient overshoot caused by thermal response hysteresis from the source. By combining model predictive control and multivariate collaborative tracking strategies, it not only achieves a globally optimal trade-off between energy consumption and tracking error under physical constraints, but also adaptively corrects model parameters through the residual feedback mechanism of the state observer. This effectively overcomes the time-varying nature of the system caused by hardware aging or environmental disturbances, completely eliminates steady-state control deviation, and significantly improves the control accuracy and full life-cycle robustness of complex coupled physical systems.
[0041] Those skilled in the art will understand that embodiments of the present invention can be provided as methods, systems, or computer program products. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0042] This invention is described with reference to flowchart illustrations and / or block diagrams of systems, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations and / or block diagrams. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0043] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.
[0044] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0045] Although preferred embodiments of the invention have been described, those skilled in the art, upon learning the basic inventive concept, can make other changes and modifications to these embodiments. Therefore, the appended claims are intended to be interpreted as including both the preferred embodiments and all changes and modifications falling within the scope of the invention.
[0046] Obviously, those skilled in the art can make various modifications and variations to this invention without departing from its spirit and scope. Therefore, if these modifications and variations fall within the scope of the claims of this invention and their equivalents, this invention also intends to include these modifications and variations.
Claims
1. A server intelligent scheduling and control method, characterized in that, Includes the following steps: S1. The input excitation flow signal and thermal process variable response signal of each controlled unit in the server cluster are synchronously differentially sampled. The dynamic transfer function matrix of each controlled sub-unit is identified online by the recursive least squares estimation algorithm with forgetting factor, and the thermal response lag time constant and steady-state gain parameter characterizing the current physical thermal properties of each controlled unit are extracted. S2. Based on the thermal response hysteresis time constant, step response stability analysis is performed to calculate the maximum allowable input rheological rate of each controlled subunit. At the same time, the thermal load saturation boundary of each controlled subunit is mapped by the steady-state gain parameter. The maximum input rheological rate and the thermal load saturation boundary are combined to construct the dynamic safe operation envelope domain of the control cycle adjustment action under constraint. S3. Obtain the current total input flux setting value of the system, and within the constraints of the dynamic safety operation envelope domain, establish a model predictive control objective function that includes energy consumption index and tracking error index, solve for the optimal flow distribution setting value and energy efficiency operating condition command of each controlled sub-unit, and send them down to the underlying execution control unit. S4. During the execution of the control cycle, the multivariable collaborative tracking control algorithm drives the front-end traffic distribution gateway and node voltage regulation module in coordination with the set value and instructions to track and control the service request injection rate and chip core operating voltage amplitude of each controlled sub-unit in a closed loop, and the forgetting factor in S1 is corrected online by calculating the observation residual through the state observer.
2. The server intelligent scheduling and control method according to claim 1, characterized in that: The specific process of identifying the dynamic transfer function matrix of each controlled sub-unit online using the recursive least squares estimation algorithm with a forgetting factor is as follows: Construct a data regression vector containing the increments of historical input excitation flow signals and the increments of historical thermal process variable response signals, and calculate the prior prediction error between the actual system output at the current moment and the predicted output based on the parameter estimates at the previous moment. The orthogonal gain correction vector at the current time is calculated by using the forgetting factor and the covariance matrix of the previous time step. The coefficient parameter vector of the discrete-time transfer function matrix is then iteratively updated based on the product of the prior prediction error and the orthogonal gain correction vector. The update step size of the covariance matrix is dynamically adjusted based on the excitation intensity of the data regression vector to obtain the discrete-time transfer function matrix with parameter convergence.
3. The server intelligent scheduling and control method according to claim 2, characterized in that: The specific process for extracting the thermal response hysteresis time constant and steady-state gain parameters that characterize the current physical and thermal properties of each controlled unit is as follows: An inverse bilinear transformation from the Z-domain to the S-domain is performed on the discrete-time transfer function matrix to map the discrete coefficients to the system pole distribution in the continuous-time domain. Extract the dominant pole that is closest to the imaginary axis of the S-plane, and calculate the absolute value of the reciprocal of the real part of the dominant pole as the thermal response hysteresis time constant. The steady-state limit operation of the discrete-time transfer function matrix is performed using the final value theorem, and the output amplitude of the unit step response when the time approaches infinity is used as the steady-state gain parameter.
4. The server intelligent scheduling and control method according to claim 1, characterized in that: The stability analysis of the step response based on the thermal response hysteresis time constant is performed to calculate the maximum allowable input rheological rate of each controlled sub-unit. The specific process of mapping the thermal load saturation boundary of each controlled sub-unit using steady-state gain parameters is as follows: Based on the thermal response hysteresis time constant as the first derivative decay factor, the maximum rise slope of the step response allowed by each controlled sub-unit under critical damping state is analytically calculated and defined as the maximum input rheological rate. The difference between the current ambient temperature of each controlled sub-unit and the upper limit of the chip's physical tolerance temperature is obtained as the thermal capacity margin. The thermal capacity margin is divided by the steady-state gain parameter to calculate the maximum absolute input flux that each controlled sub-unit can be loaded under steady-state equilibrium conditions, which is then used as the thermal load saturation boundary.
5. The server intelligent scheduling and control method according to claim 4, characterized in that: The specific process of constructing the dynamic safe operating envelope domain for constrained control cycle regulation by combining the maximum input rheological rate with the thermal load saturation boundary is as follows: In the state space of the control variables, the current input flux value is used as the reference point, and the dynamic reachable cone region of the next control cycle is defined by the maximum input rheological rate. By defining a static absolute amplitude limiting hyperplane in the state space of the control variables through the thermal load saturation boundary, a multidimensional geometric intersection operation is performed on the dynamically reachable conical region and the static absolute amplitude limiting hyperplane, and the overlapping closed convex polyhedral space of the two is extracted as the dynamic safe operation envelope.
6. The server intelligent scheduling and control method according to claim 1, characterized in that: The specific process of obtaining the current total input flux setpoint of the system and establishing the model predictive control objective function, which includes energy consumption and tracking error indices, within the constraints of the dynamic safe operation envelope domain, is as follows: Set the prediction time domain length and control time domain length of the controller, and extend the total input flux setpoint to a reference trajectory sequence in the future prediction time domain; The tracking error index is defined as the square of the Euclidean norm between the sum of the predicted output flow of each controlled subunit and the reference trajectory sequence, and the energy consumption index is defined as the quadratic product of the control input vector of each controlled subunit and the energy consumption weighting matrix. The dynamic safety operation envelope domain is transformed into a set of linear inequality constraints of state variables in the prediction time domain. The tracking error index and energy consumption index are linearly weighted and summed to construct a quadratic programming optimization objective function under the constraints of the set of linear inequality constraints of state variables.
7. The server intelligent scheduling and control method according to claim 6, characterized in that: The specific process of determining the optimal flow allocation setpoint and energy efficiency operating instructions for each controlled subunit and then sending them to the underlying execution control unit is as follows: The optimal control increment sequence in the future control time domain is calculated by using a numerical optimization solver to perform rolling optimization on the quadratic programming objective function. Based on the rolling time-domain control principle, the first element of the optimal control increment sequence is extracted as the optimal cooperative control vector at the current moment; The optimal collaborative control vector is decomposed into a traffic truncation threshold for the front-end traffic distribution gateway and a voltage amplitude adjustment code for the node voltage regulation module. These are then sent out in parallel through the hardware driver interface as the optimal traffic allocation setting value and the energy efficiency operating condition command, respectively.
8. The server intelligent scheduling and control method according to claim 1, characterized in that: The specific process of using a multivariable collaborative tracking control algorithm to collaboratively drive the front-end traffic distribution gateway and node voltage regulation module according to set values and instructions, and to achieve closed-loop tracking control of the service request injection rate and the core operating voltage amplitude of each controlled sub-unit, is as follows: The optimal traffic allocation setting is converted into the token generation rate of the token bucket algorithm in the front-end traffic distribution gateway. The injection rate of business requests into each controlled sub-unit is physically limited by adjusting the token generation rate. The energy efficiency operating condition command is converted into a pulse width modulation signal of the node voltage regulation module, which drives the power switch to adjust the operating voltage amplitude of the chip core. Monitor the response delay time of the voltage regulation process, apply transient hold logic to the token generation rate before the voltage reaches the target value, and achieve timing synchronization and coordination between the flow injection action and the voltage regulation action.
9. The server intelligent scheduling and control method according to claim 1, characterized in that: The specific process of online correction of the forgetting factor in S1 by calculating the observation residuals using the state observer is as follows: The instantaneous observation residual vector is obtained by subtracting the actual output of each controlled subunit from the predicted output of the dynamic transfer function matrix. Calculate the modulus of the instantaneous observation residual vector and establish a nonlinear inverse proportional mapping function between the forgetting factor and the observation residual modulus; When the magnitude of the observation residual increases, the forgetting factor is reduced through a nonlinear inverse proportional mapping function to enhance the algorithm's sensitivity to new data. When the magnitude of the observation residual decreases, the forgetting factor is increased to enhance the smoothness of parameter estimation, thus completing the adaptive update of the recursive least squares estimation algorithm.