Intelligent computing park computing power load prediction-based computing power and electricity collaborative scheduling method

CN122736130APending Publication Date: 2026-09-11HUNAN LUXIN IOT TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610715831.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-05-22
Publication Date
2026-09-11

AI Technical Summary

Technical Problem

算电运行体系割裂:传统算力负载调度仅关注任务排队与服务器资源分配,未联动园区配电网分时电价、碳排特性、微电网惯量支撑能力及储能老化损耗;电网调度也未考虑算力任务可暂停、降频、迁移的柔性调节潜力,无法实现算力与电力资源协同最优配置

Benefits of technology

[0019]综上所述,本发明至少具有以下有益效果:通过分数阶微分预处理兼顾时序全局趋势与局部随机细节,有效挖掘弱关联隐含特征,从底层提升多源数据可用性与后续预测精度。采用时序超图替代传统两两关联图谱,可表征多设备间热-电-算多体耦合关系,适配智算园区复杂拓扑与关联特性,建模维度与适配性显著提升。融合变分模态分解与神经微分方程实现连续时间动态建模,完成负载多尺度分解与概率区间预测,实现不确定性完整量化,为鲁棒调度提供可靠输入支撑。优化多目标智能算法迭代筛选逻辑,消除非支配排序与个体择优的内在矛盾,建立目标权重与拥挤度计算的关联机制,保证帕累托前沿的完备性与合理性。修正分布鲁棒评估最劣判定方向与成本增量符号,使鲁棒决策完全贴合电力与算力效用物理机理,打通预测与调度的对抗闭环,规避惩罚项失效问题。依托智能算法优化卡尔曼滤波噪声协方差,大幅提升微电网惯量与储能老化参数辨识精度;精细化划分算力任务弹性类型,实现算力与电力资源柔性联动调度。引入纳什博弈议价与区间二型模糊决策,实现多目标帕累托最优解集向唯一工程可落地调度方案的平滑过渡,解决多目标优化难以实际决策的行业痛点。搭建日前规划+实时微调的双闭环调度架构,依托信息间隙决策适配源荷与电价未知不确定性,依托李雅普诺夫理论保障系统运行约束稳态,抗扰动与自适应调节能力大幅增强。建立全流程误差反馈、效果评估与模型自演进机制,无需人工干预即可持续适配园区工况、气象、设备老化的动态变化,长期维持系统最优调度性能。打破算力与电力调度壁垒,实现园区运行经济性、电网频率稳定性、储能设备延寿及算力服务品质的协同兼顾,达成智算园区全资源一体化智能协同运行。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122736130A_ABST
    Figure CN122736130A_ABST
Patent Text Reader

Abstract

The application provides a computing power load prediction-based intelligent computing park computing power and electricity collaborative scheduling method, comprising the following steps: S1. Multi-source heterogeneous data preprocessing and spatiotemporal hypergraph dynamic construction based on fractional order chaotic bat algorithm; S2. Computing power load probability prediction and uncertainty quantification based on variational mode decomposition and time series hypergraph neural differential equation; S3. Real-time identification of microgrid inertia and energy storage aging and task flexibility analysis based on improved grey wolf optimization-extended Kalman filter; S4. Multi-objective Pareto front construction based on bacterial foraging-quantum particle swarm fusion optimization and non-dominated sorting; S5. Distribution robust multi-objective scheduling decision based on Wasserstein ball interval second type fuzzy Nash product; S6. Double closed-loop real-time flexible scheduling and model online evolution based on Lyapunov drift plus information gap decision. The application realizes integrated intelligent collaborative scheduling of computing power, electricity, energy storage and cooling resources in an intelligent computing park.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of intelligent computing park resource scheduling, microgrid collaborative optimization, and artificial intelligence time-series prediction, and in particular to a method for intelligent computing park computing-power collaborative scheduling based on computing load prediction. Background Technology

[0002] With the large-scale aggregation of computing infrastructure, large-scale intelligent computing parks gather a large number of servers, racks, cooling equipment, distributed photovoltaics, wind power, and energy storage battery packs, exhibiting operational characteristics of strong fluctuations in computing load, thermal-electrical decoupling of multiple devices, and significant uncertainty in source-load. Currently, the operation and scheduling of intelligent computing parks suffer from numerous technical deficiencies: The computing power operation system is fragmented: Traditional computing power load scheduling only focuses on task queuing and server resource allocation, without linking the time-of-use electricity price of the park's distribution network, carbon emission characteristics, microgrid inertia support capacity, and energy storage aging losses; the power grid scheduling also does not consider the flexible adjustment potential of computing power tasks that can be paused, reduced in frequency, or migrated, and cannot achieve optimal allocation of computing power and power resources.

[0003] Insufficient load forecasting capability: Conventional time series forecasting models can only fit a single load change trend, making it difficult to simultaneously retain time series trend information and random details, and have poor ability to mine weakly correlated implicit features; moreover, most of them are single-point deterministic predictions, lacking probability intervals and uncertainty quantification mechanisms, and cannot support robust scheduling decisions.

[0004] The disconnect between multi-objective optimization and implementation decision-making: Although existing multi-objective intelligent optimization can generate Pareto optimal solution sets, they generally suffer from problems such as conflicting algorithm selection logic and disconnect between objective weights and congestion calculations; moreover, Pareto solution sets lack subsequent implementation selection mechanisms, resulting in a break in data flow with the scheduling decision-making process, remaining only at the theoretical optimization level and unable to be implemented in engineering.

[0005] The robust scheduling evaluation logic is mismatched: the existing sub-bluish scheduling determines the worst-case scenario of the load uncertainty set in a way that contradicts the physical meaning of utility, the definition of cost increment symbols is confusing, and the design of regularization penalty terms is ineffective, which leads to the inability of the adversarial interaction mechanism between prediction and scheduling to close the loop properly.

[0006] Low accuracy of key parameter identification: Traditional extended Kalman filter microgrid inertia and energy storage aging identification methods have problems such as fixed noise covariance and poor adaptability to operating conditions; and lack of fine classification of task elasticity categories, making it impossible to quantify the adjustable potential of computing power.

[0007] The system lacks closed-loop adaptive capability: existing scheduling architectures are mostly open-chain single-time scheduling, without real-time deviation feedback, multi-dimensional effect evaluation and model self-evolution mechanism, and cannot adapt to the dynamic changes in park load, weather, electricity price and equipment aging in the long term. Summary of the Invention

[0008] This invention provides a computing and power collaborative scheduling method for intelligent computing parks based on computing load prediction. It constructs a multi-body coupled time-series hypergraph model of thermal, electrical, and computing resources to achieve fractional-order refined preprocessing of multi-source data; completes multi-scale probabilistic prediction and uncertainty quantification of computing load; establishes a multi-objective Pareto optimization system that considers economy, stability, and energy storage life extension; corrects the sub-Blu-ray evaluation logic to achieve feasible decision-making from Pareto solutions to scheduling schemes; accurately identifies microgrid inertia and energy storage aging parameters, and analyzes the flexibility of computing tasks; and builds a full-link closed-loop architecture of day-ahead planning, real-time adjustment, effect evaluation, and model self-evolution to achieve integrated intelligent collaborative scheduling of computing power, power, energy storage, and cooling resources in intelligent computing parks.

[0009] To achieve the above objectives, the present invention adopts the following technical solution: A computing power-coordinated scheduling method for intelligent computing parks based on computing power load prediction includes: S1. Collect multi-source heterogeneous data from the computing power side, power side, environment side, and cooling system of the intelligent computing park. After imputing missing values ​​and removing outliers from the multi-source heterogeneous data, use the fractional-order chaotic bat algorithm to determine the optimal fractional-order number, and perform fractional-order differential transformation on each data sequence accordingly. At the same time, construct a time-series hypergraph representing the power supply, pipeline, and thermal coupling relationship between devices, and optimize the edge weight matrix of the time-series hypergraph. S2. Based on the data after fractional-order differential transformation and the edge weight matrix, perform multi-scale variational mode decomposition on the computing power load time series, construct a time series hypergraph neural differential equation prediction model, and output a probability prediction surface containing multiple quantiles and the prediction interval width; introduce the robust cost increment fed back by S5 as a regularization term in the training of the prediction model to form an adversarial learning closed loop, wherein the robust cost increment is the difference between the power operation cost under the worst distribution and the power operation cost under the nominal distribution; S3. Based on the transformed data and the edge weight matrix, the improved gray wolf optimization and extended Kalman filter fusion algorithm is used to identify the equivalent inertia and energy storage aging coefficient of the microgrid in real time, and the computing power tasks are flexibly classified to generate a task flexibility contribution matrix. S4. Based on the edge weight matrix, the prediction interval width, the microgrid equivalent inertia, and the energy storage aging coefficient, with the power operation cost, microgrid frequency deviation accumulation, and energy storage aging rate as objectives, and dynamically relaxing the constraint boundary according to the prediction interval width, the bacterial foraging-quantum particle swarm fusion algorithm is adopted to filter according to the non-dominated level and the weighted congestion degree calculated by the multi-objective preference weight, and to generate the Pareto optimal solution set and scheduling feasible domain envelope. S5. Construct a Wasserstein uncertainty set based on the probability prediction surface, perform robustness evaluation on each candidate solution in the Pareto optimal solution set, calculate the generalized Nash bargaining product score weighted by the Nash bargaining coefficient, and use an interval type II fuzzy logic system to select the optimal day-ahead scheduling scheme; at the same time, calculate the robust cost increment corresponding to each candidate solution, and feed back the maximum value of the robust cost increment to S2; S6. Based on the optimal day-ahead scheduling scheme, combined with the ultra-short-term prediction results, the task flexibility contribution matrix, and the scheduling feasible region envelope, minute-level dual-closed-loop real-time elastic scheduling is performed using Lyapunov drift plus information gap decision theory to generate a prediction error matrix and a scheduling deviation matrix, which are fed back to S1, S3, and S4 respectively. The penalty term for Lyapunov drift includes the Lyapunov penalty coefficient. S7. Collect actual execution data throughout the day and evaluate the scheduling effect from five dimensions: economy, stability, environmental protection, computing power service quality and energy storage life. When abnormal operating conditions are triggered, automatically switch to the corresponding emergency scheduling mode. S8. Summarize the entire process operation data and evaluation data, and update the model parameters of each stage through online transfer learning, elastic weight consolidation algorithm and Bayesian optimization method to achieve global closed-loop self-evolution.

[0010] In this specification, in S1, the fractional-order chaotic bat algorithm uses the weighted sum of the maximum autocorrelation coefficient of the data sequence and the stationarity index as the fitness function; the initial value of the edge weight matrix of the time-series hypergraph is determined by the absolute value of the Pearson correlation coefficient between nodes, and is dynamically optimized through another instance of the fractional-order chaotic bat algorithm using the weighted sum of the node degree distribution entropy and the average clustering coefficient as the fitness function; S1 receives the prediction error matrix and the scheduling deviation matrix fed back from S6, and adaptively adjusts the optimal fractional-order order and the maximum inertia weight in the fractional-order chaotic bat algorithm.

[0011] In this specification, in S2, the quadratic penalty factor of the variational mode decomposition is positively correlated with the optimal fractional order obtained in S1; the temporal hypergraph neural differential equation prediction model uses a Dormand-Prince fifth-order Runge-Kutta integrator for forward integration, outputting predicted values ​​with quantiles of 0.05, 0.25, 0.50, 0.75, and 0.95; the robust cost increment is added as a regularization term to the training loss function of the prediction model.

[0012] In this specification, in S3, the improved gray wolf optimization algorithm adopts a nonlinear convergence factor and optimizes the process noise covariance and measurement noise covariance of the extended Kalman filter using the sum of squared prediction errors as the fitness function; the improved gray wolf optimization algorithm is executed by combining offline pre-training with online sliding window rolling updates; the elasticity categories of the computing power tasks are divided into five types: rigid uninterruptible tasks, tasks that can be paused for a short time, tasks that can be paused for a long time, tasks that can run at reduced frequency, and tasks that can migrate across nodes; S3 receives the scheduling deviation matrix fed back from S6 and adaptively adjusts the convergence factor of the improved gray wolf optimization algorithm.

[0013] In this specification, in S4, the method of dynamically relaxing the constraint boundary is as follows: calculate the dynamic adjustment coefficient based on the predicted interval width, and multiply the upper and lower limits of the microgrid inertia constraint, distributed power output constraint, energy storage operation constraint, and computing power task constraint by the dynamic adjustment coefficient respectively; the weighted congestion calculation introduces the multi-objective preference weight, and the population screening is performed in the order of non-dominated level from high to low, and the weighted congestion of the same level from large to small; S4 receives the scheduling deviation matrix fed back by S6, dynamically updates the multi-objective preference weight, and substitutes it into the weighted congestion calculation of the next iteration.

[0014] In this specification, S5, the robustness assessment specifically involves: solving for the minimum value of computing power task completion and the minimum negative cost of power operation for each candidate solution under the worst load distribution within the Wasserstein uncertainty set; the generalized Nash bargaining product takes the minimum value as input and uses the retention utility obtained from the statistics of the lowest power operation cost in the park's history over the past seven days as a benchmark; the robust cost increment is always non-negative.

[0015] In this specification, S6, the dual-closed-loop real-time elastic scheduling specifically refers to: the inner loop fine-tuning the scheduling instructions at the minute level through a proportional-integral-derivative controller to quickly respond to real-time fluctuations; the outer loop generating the prediction error matrix and the scheduling deviation matrix on a daily basis and feeding them back; the information gap decision theory adopts a double-layer nested optimization structure, with the outer layer solving for the maximum allowable uncertainty radius of the electricity price, and the inner layer minimizing the Lyapunov drift plus penalty function under this radius.

[0016] In this specification, the evaluation indicators for the five dimensions in S7 include: economic dimension: actual power operation cost, unit computing power power cost and net demand response revenue; stability dimension: maximum frequency deviation, average frequency deviation, standard deviation and voltage qualification rate; environmental dimension: carbon dioxide emissions and diesel consumption; computing power service quality dimension: task on-time completion rate, average response time and computing power utilization rate; and energy storage life dimension: average health status and average aging rate.

[0017] In this specification, the abnormal operating condition judgment conditions in S7 include: power operation cost exceeding 20% ​​of the threshold, frequency deviation exceeding ±0.2 Hz, task on-time completion rate below 95%, any energy storage health status below 80%, distributed power output dropping by more than 50%, or the park actively disconnecting from the grid; the emergency dispatch mode includes economic priority mode, stability priority mode, computing power priority mode, and energy storage protection mode.

[0018] In this specification, S8, the global closed-loop self-evolution specifically includes: updating the non-intrusive load decomposition model and server power-computing power mapping coefficients through online transfer learning; fine-tuning the parameters of the time-series hypergraph neural differential equation prediction model through the elastic weight consolidation algorithm; automatically adjusting the Lyapunov penalty coefficient, the multi-objective preference weight, and the Nash bargaining coefficient through Bayesian optimization; and seamlessly switching to the scheduling process of the next day after all model updates are completed.

[0019] In summary, this invention offers at least the following advantages: Fractional-order differential preprocessing balances global temporal trends with local stochastic details, effectively uncovering weakly correlated latent features and improving the availability of multi-source data and subsequent prediction accuracy from the ground up. The use of a temporal hypergraph instead of traditional pairwise correlation graphs characterizes the multi-body coupling relationships between multiple devices (thermal, electrical, and computational), adapting to the complex topology and correlation characteristics of intelligent computing parks, significantly enhancing modeling dimensionality and adaptability. The fusion of variational mode decomposition and neural differential equations enables continuous-time dynamic modeling, completing multi-scale load decomposition and probability interval prediction, achieving complete quantification of uncertainty, and providing reliable input support for robust scheduling. The optimized iterative selection logic of the multi-objective intelligent algorithm eliminates the inherent contradiction between non-dominated sorting and individual optimization, establishing a correlation mechanism between objective weights and congestion calculation, ensuring the completeness and rationality of the Pareto front. The correction of the worst-case judgment direction and cost increment sign in the fractional-order robust evaluation ensures that robust decision-making fully aligns with the physical mechanisms of power and computing utility, breaking down the adversarial loop between prediction and scheduling, and avoiding the problem of penalty term failure. By leveraging intelligent algorithms to optimize the noise covariance of Kalman filters, the accuracy of identifying microgrid inertia and energy storage aging parameters is significantly improved. The elasticity of computing tasks is refined to achieve flexible coordinated scheduling of computing power and power resources. Nash game negotiation and interval-based type II fuzzy decision-making are introduced to smoothly transition from multi-objective Pareto optimal solutions to a unique, engineering-feasible scheduling scheme, addressing the industry pain point of difficulty in practical decision-making for multi-objective optimization. A dual-closed-loop scheduling architecture of day-ahead planning and real-time fine-tuning is established. Information gap decision-making adapts to the unknown uncertainties of source load and electricity prices, while Lyapunov theory ensures the system operates in a constrained steady state, significantly enhancing its anti-disturbance and adaptive adjustment capabilities. A full-process error feedback, effect evaluation, and model self-evolution mechanism is established, enabling continuous adaptation to dynamic changes in park operating conditions, weather, and equipment aging without manual intervention, maintaining optimal system scheduling performance in the long term. Breaking down the barriers between computing power and power scheduling, a synergistic balance is achieved between park operation economy, grid frequency stability, energy storage equipment life extension, and computing service quality, realizing integrated intelligent collaborative operation of all resources in the intelligent computing park. Attached Figure Description

[0020] Figure 1 This is a schematic diagram illustrating the steps of a computing power-coordinated scheduling method for intelligent computing parks based on computing power load prediction.

[0021] Figure 2 This is a flowchart illustrating the intelligent computing park's computing power collaborative scheduling method based on computing power load prediction.

[0022] Figure 3 This is a schematic diagram of the S4-S5 optimization decision-making process involved in this invention.

[0023] Figure 4 This is a schematic diagram of the S2-S5 adversarial interaction process involved in this invention. Detailed Implementation

[0024] The embodiments of the present invention will now be described in detail with reference to the accompanying drawings.

[0025] like Figure 1 and Figure 2 As shown, this embodiment provides a computing power collaborative scheduling method for intelligent computing parks based on computing power load prediction, including the following: S1. Preprocessing of multi-source heterogeneous data and dynamic construction of spatiotemporal hypergraph based on fractional-order chaotic bat algorithm: Within the intelligent computing park, deploy server-side resource monitoring agents, non-intrusive power load monitoring devices for rack power distribution units, cooling system sensors, environmental meteorological stations, photovoltaic combiner boxes, energy storage converters, gateway electricity meters, and synchronous phasor measurement units to uniformly collect the following multi-source heterogeneous data and align them all to Coordinated Universal Time, with a sampling interval of fifteen seconds.

[0026] The types of raw data on the computing power side are as follows: CPU utilization, GPU utilization, GPU memory usage, memory bandwidth utilization, disk read / write operations per second, network transmission throughput, network reception throughput, task wait queue length, process context switching rate, task submission time, task execution duration, task type, task priority, rack inlet temperature, rack outlet temperature, and total fan speed. The types of raw power-side data are as follows: substation bus voltage, bus current, grid frequency, total active power at the junction, total reactive power at the junction, power factor, estimated active and reactive power per server obtained from non-intrusive load monitoring devices, total harmonic distortion rate of current, active power on the AC side of distributed photovoltaic inverters, output power of distributed wind turbines, total voltage of lithium iron phosphate battery packs, total current of lithium iron phosphate battery packs, state of charge of lithium iron phosphate battery packs, maximum cell temperature of lithium iron phosphate battery packs, minimum cell temperature of lithium iron phosphate battery packs, charge and discharge power of lithium iron phosphate battery packs, output power of diesel generators, and branch current of each distribution cabinet. The types of raw environmental data are as follows: outdoor dry-bulb temperature, outdoor wet-bulb temperature, relative humidity, total solar irradiance on the horizontal plane, wind speed, wind direction, and atmospheric pressure. The raw data types for the cooling system are as follows: chiller compressor suction pressure, chiller compressor discharge pressure, chilled water supply temperature, chilled water return temperature, chilled water volumetric flow rate, cooling water inlet temperature, cooling water outlet temperature, cooling tower fan speed, cooling water pump frequency, and chilled water valve opening at each air conditioning terminal. In addition, day-ahead time-of-use marginal electricity price forecast sequences and real-time electricity price forecast sequences are obtained through the power trading center interface; time-of-use carbon emission factor forecast sequences are obtained through the power grid carbon emission data service platform; and outdoor dry-bulb temperature, outdoor wet-bulb temperature, relative humidity, and solar irradiance forecasts for the next 72 hours are obtained through the European Centre for Medium-Range Weather Forecasts (ECMWF) high-resolution numerical forecast interface.

[0027] Missing value imputation and outlier removal were performed on all raw data. Missing value imputation employed a weighted average method based on spatiotemporal proximity. For a missing value of a device at a given time, a distance-weighted average was calculated using the valid values ​​of the device at the two times before and after it, as well as the valid values ​​of adjacent devices within the same rack at the same time. The weights were inversely proportional to the time interval or spatial distance. Outlier removal used a three-standard-deviation criterion combined with a box plot method. Data exceeding three standard deviations and located 1.5 times the interquartile range outside the upper and lower quartiles of the box plot were marked and replaced with the mean of the device at the one time before and after it.

[0028] Initialize the parameters of the fractional-order chaotic bat algorithm. Set the bat population size to 100, the maximum number of iterations to 200, the initial loudness to 0.5, the initial pulse emission rate to 0.5, the fractional-order range to the interval 0 to 1, and the initial value of the chaotic mapping to 0.3.

[0029] In the fractional-order chaotic bat algorithm, the first Only one bat in the first Position vector of the next iteration Composed of parameters such as fractional order and loudness attenuation coefficient, velocity The updated formula is: ; in For fractional inertial weights, As of the date The global optimal position in the next iteration. For the first The emission frequency of a single bat. Emission frequency Calculate using the following formula: ; Here Minimum transmission frequency, The maximum transmission frequency, To be taken from the interval Uniform random numbers. Fractional order inertia weights. Introducing fractional-order memory effect can preserve historical movement information of bat populations and prevent the algorithm from prematurely converging to local optima. This effect is defined as: ; in , , For gamma function, Let the fractional order be the order to be optimized. This represents the maximum number of iterations for the fractional-order chaotic bat algorithm. A Logistic chaotic mapping is introduced to perturb the bat's position, leveraging the ergodicity of chaotic motion to enhance the algorithm's global search capability. Chaotic variables... The update is as follows: ; Chaos control parameters At this point, the Logistic mapping is in a completely chaotic state. When the uniform random number is greater than the current pulse emission rate, a chaotic perturbation is applied to the globally optimal position to generate a new position: ; If the random number is less than the loudness and the fitness value of the new position is better than the fitness value corresponding to the current best position, then accept. Simultaneously update loudness and pulse emission rate : ; in This is the loudness attenuation coefficient. This is the pulse emissivity enhancement factor. The initial pulse emission rate. Fitness function. Defined as the maximum autocorrelation coefficient of a data series With stationarity indicators The weighted sum is used to measure the predictability of data after fractional derivative processing: ; In the formula and These represent the standard deviation and mean of the data sequence, respectively. Through optimization of the bat population, the output is... Maximum optimal fractional order .

[0030] Perform preprocessing on each data sequence The fractional-order differential transform, defined by Grünwald-Letnikov, can be directly implemented numerically and is suitable for processing discrete-time series. ; in For time step, This is the floor function. Data processed by fractional derivatives can simultaneously retain the trend information and random details of the sequence, enhance weakly correlated components, and significantly improve the accuracy of subsequent prediction models.

[0031] Construct a time-series hypergraph between device nodes based on the processed data. Node set Includes all servers, cabinets, power distribution cabinets, distributed power supplies, energy storage battery packs, cooling equipment, and environmental sensors. Hyperedge collection. It is composed of three types of coupled relationships: the physical power supply tree topology, the liquid cooling pipe network flow direction, and thermally coupled hyperedges constructed by the heat and airflow connections between server racks. Each hyperedge can connect more than two nodes to express multi-body thermal-electrical-computation coupling relationships, which is impossible to achieve with ordinary graph structures. Edge weight matrix The initial values ​​are given by the absolute values ​​of the Pearson correlation coefficients between nodes. A connection is established when the absolute value is greater than 0.7. Let be the total number of nodes. The edge weights are dynamically optimized using the fractional-order chaotic bat algorithm described above, at which point the fitness function... Defined as the weighted sum of node degree distribution entropy and average clustering coefficient, it is used to measure the structural rationality of the hypergraph: ; in For node degree The probability, For the maximum node degree, represents the average clustering coefficients of the hypergraph. Output the optimized edge weight matrix. The multidimensional data volume after fractional-order differential transformation is denoted as... ,in For time step, For feature dimensions, along with They are transmitted together to S2, S3 and S4.

[0032] S1 receives the prediction error matrix from S2. And the scheduling deviation matrix of S6 This feedback is used to dynamically adjust the fractional order and inertia weight, achieving closed-loop optimization of data preprocessing intensity and subsequent prediction and scheduling performance. First, the node-level average prediction error vector is calculated. : ; Then calculate the weighted average prediction error: ; Finally, update the optimal fractional order and the maximum inertia weight: ; ; This adaptive adjustment mechanism can automatically optimize the parameters of data preprocessing based on the actual effects of prediction and scheduling, thereby improving the robustness of the entire system.

[0033] S2. Computational load probability prediction and uncertainty quantification based on variational mode decomposition and temporal hypergraph neural differential equations: The core algorithm in this step is a variational mode decomposition-guided temporal hypergraph neural differential equation adversarial Bruker prediction, which receives S1's... and And achieve iterative bidirectional interaction with S5 (see reference) Figure 4 Together, they can improve the accuracy of predictions and the robustness of scheduling.

[0034] First, for each server from The equivalent computing load index extracted from The time series data is subjected to variational mode decomposition. The equivalent computing load index is obtained by weighted summation of the server's CPU utilization, GPU utilization, memory bandwidth utilization, and network throughput; the weights are determined using principal component analysis. Decomposed into One intrinsic mode function , Each mode revolves around the center frequency Fluctuations. The constrained variational model for decomposition is: ; in For the Dirac function, For convolution operations, This is a partial derivative operator. A quadratic penalty factor is introduced. and Lagrange multipliers Construct the augmented Lagrange function: ; The iterative solution using the alternating direction multiplier method yields the following update formulas for each variable: 1. Intrinsic mode function update: ;in This represents the Fourier transform.

[0035] 2. Center frequency update: ; 3. Lagrange multiplier update: ;in To update the step size, the value is set to 1.

[0036] The value of and the optimal fractional order of S1 output Directly related, enabling deep integration of data preprocessing and signal decomposition: ; according to The modes are divided into three categories: high frequency, medium frequency, and low frequency, which correspond to the random fluctuations, periodic changes, and trend evolution of the load, respectively.

[0037] For each modality, a temporal hypergraph neural differential equation prediction unit is constructed. The graph structure adopts S1. ,node eigenvectors It is composed of the current value of the corresponding mode, the task queue length, and the temperature of the rack. System hidden state The evolution is characterized by the following neural differential equation, which can accurately model the dynamic behavior of the system in continuous time: ; in For a hypergraph attention network with gating mechanism, a parameterized vector field. To and The set of neighboring nodes belonging to any superedge. For the super-edge attention weights, For message passing matrix, The externally enforced vector is encoded from S1's weather and electricity price forecasts. (Super-edge attention weights) The calculation formula is: ; in For attention vectors, Here is the attention weight matrix. The activation function is a linear rectified function with leakage. A fifth-order Dormand-Prince Runge-Kutta integrator is used to transform the above equation from its initial state. Forward integration to the future Step 1, to obtain the hidden state sequence. Then, perform a multi-scale causal dilatational convolution in parallel on the hidden state sequence, with the dilatation rate set to... The output layer provides five quantiles. Predicted value Forming a probability distribution surface .

[0038] Total loss during model training It consists of three parts.

[0039] The first term is the quantile loss, summed over all servers and time points, used to measure the accuracy of the prediction: ; in For indicator functions, when The value is 1 if it is true, and 0 otherwise.

[0040] The second term is the cost penalty for the quantile bar scheduling generated by the interaction with S5. S5 constructs the Wasserstein uncertainty set based on the prediction quantiles of S2, and its radius... Determined by the width of the prediction interval: ; in A sensitivity coefficient is preset. When S5 solves the day-ahead scheduling under the worst-case distribution, it will obtain the corresponding marginal cost increment. This cost reflects the scheduling risk caused by forecast uncertainty. This cost is fed back into this step as a regularization term: ; For joint weight balancing. In the early stages of training, S5 returns once. The numerical scale is used for calibration After that, the calculation is repeated every ten training epochs. and This enables joint adversarial learning between S2 and S5: S2 attempts to narrow the prediction interval to reduce scheduling robustness costs, while S5 updates the cost function corresponding to the worst distribution accordingly. The two compete with each other to improve the overall performance of the system.

[0041] The third term is the smoothing regularization term. To suppress second-order difference overshoot in the predicted sequence and ensure the smoothness of the prediction results: ; Minimize on historical window data using the Adam optimizer. The learning rate is initially set to 0.001, and the batch size is 128. The early stopping mechanism is triggered when the validation set loss stops decreasing for five consecutive epochs.

[0042] After training, the model receives real-time updates from S1 online. Iteratively outputs the quantile prediction for the next 96 steps. This information is then passed to S5 and S7 respectively. Simultaneously, S5 returns the optimal uncertain radius after each rolling optimization. This is used to fine-tune the steps in this process. S7 provides feedback on the deviation between the actual load and the predicted value at the end of each day, which is used to incrementally fine-tune the parameters of the neural differential equation to ensure the long-term adaptive capability of the model.

[0043] S3. Real-time identification and task elasticity analysis of microgrid inertia and energy storage aging based on improved gray wolf optimization-extended Kalman filter: This step obtains fractional-order differential data and hypergraph structure from S1, and exports computation graph descriptions, service level agreement parameters and resource constraint information of all submitted tasks from the park task scheduler, completing the power-side characteristic identification and computing power-side elastic analysis.

[0044] By using grid frequency change rate and active power data collected by synchronous phasor measurement units, the equivalent inertia constant of the microgrid can be identified in real time. The standard physical model for the inertial response of a microgrid is: ; in For the first The rate of change of the power grid frequency at any given time, This refers to the fluctuation in active power. This is the rated angular frequency of the power grid.

[0045] To accurately estimate from finite noise-containing measurements Simultaneously, the aging parameters of the energy storage battery are identified, and an improved gray wolf optimization algorithm combined with extended Kalman filtering is employed. This method utilizes the global search capability of the gray wolf optimization algorithm to optimize the noise covariance matrix of the extended Kalman filter, solving the problem of decreased estimation accuracy caused by the fixed noise covariance of the traditional extended Kalman filter.

[0046] The initial gray wolf population size is 50, and the maximum number of iterations is 100. An improved nonlinear convergence factor is used. Used to balance global search and local exploitation, it can maintain stronger local search capability in the later stages of iteration compared to the linear convergence factor: ; in This represents the maximum number of iterations for the Grey Wolf optimization algorithm. Grey Wolf position. The update follows: ; in , , For interval A random vector within. Fitness function. Defined as the sum of squared prediction errors of the extended Kalman filter: ; in This is the length of the sliding window, with a value of 100. These are actual measured values. To obtain the predicted values ​​from the extended Kalman filter, the gray wolf search aims to find the optimal process noise covariance. and measurement noise covariance This maximizes the estimation accuracy of the extended Kalman filter.

[0047] Extended Kalman Filter State Vector ,in The equivalent inertia of the microgrid. To store and activate energy, The rated number of energy storage cycles, This is the state of charge (energy storage). State transition function. Defined as: ; in For energy storage charging and discharging efficiency, For the first Energy storage charging and discharging power at all times, This represents the rated capacity of the energy storage. Measurement function. The matched inertia response mechanism is defined as follows: ; This equation, with the rate of frequency change as an observable input, is consistent with the physical equations of microgrid inertia. (Open-circuit voltage of the energy storage battery) The Shepherd combinatorial model is adopted. This is internal resistance.

[0048] State prediction and covariance prediction are as follows: ; ; in State transition function The Jacobian matrix. The Kalman gain, state update, and covariance update are as follows: ; ; ; in For measurement function Jacobian matrix, For measurement vectors, It is an identity matrix. This filter outputs the equivalent inertia of the microgrid in real time. and energy storage health status and provide the aging coefficient. .

[0049] The Grey Wolf optimization algorithm employs an offline pre-training combined with online rolling updates. In the offline phase, historical data is used to train and obtain the initial optimal noise covariance matrix. In the online phase, Grey Wolf optimization is performed hourly, updating the noise covariance matrix using a sliding window of data from the past hour. At other times, the current optimal noise covariance matrix is ​​used to run the extended Kalman filter. This approach ensures both recognition accuracy and real-time performance.

[0050] For each task, static program analysis and historical execution traces are used to categorize its elasticity into one of five completely mutually exclusive categories: rigid uninterruptible tasks, tasks that can be paused briefly, tasks that can be paused for a long time, tasks that can run at reduced frequency, and tasks that can migrate across nodes. Combining this with the median load forecast output from S2, the electrical power that each server can release or absorb is calculated if pause, frequency reduction, or migration is performed on tasks with permissible flexibility at different time periods. This results in a computing power task scheduling flexibility contribution matrix. ,in Indicates the first The server is at the The maximum power value that can be adjusted in one time step.

[0051] Receive the scheduling deviation matrix fed back by S6 Based on this, the convergence factor of the Grey Wolf algorithm is adjusted so that the identification accuracy can be adaptively optimized according to the scheduling effect: ; Identified Energy storage aging coefficient and the flexibility contribution matrix It is passed to S4 and S5 together.

[0052] S4. Construction of a multi-objective Pareto front based on bacterial foraging-quantum particle swarm optimization and non-dominated sorting: This step receives the edge weight matrix from S1, the interval prediction width output by S2, and the edge weight matrix from S3. and A flexible power resource model that takes into account dynamic energy efficiency and carbon emission factors is constructed and embedded into a multi-objective optimization problem to generate a Pareto optimal solution set.

[0053] First, based on historical data of the chiller compressor, water pump, and cooling tower, as well as weather forecasts, a semi-empirical thermodynamic model is used for online data assimilation to identify the dynamic power utilization efficiency prediction sequence. The formula for calculating electrical energy utilization efficiency is: ; in This represents the total power consumption of the park. This refers to the electrical power consumption of information technology equipment. The photovoltaic output adopts a Beta distribution model, with the following probability density function: ; in and Let be the shape parameter of the Beta distribution. The photovoltaic system is set to maximum output. The energy storage equivalent circuit parameters are updated based on the identification results from S3. The average carbon emission factor per kilowatt-hour is calculated using the grid time-of-use carbon emission factor from S1. The upper and lower limits of the adjustable power at the gate are obtained by summarizing. Energy storage state of charge safety range and carbon emission constraint linear formula.

[0054] Based on this, a multi-objective optimization model for computer-electricity collaborative scheduling is established.

[0055] Objective 1 is to minimize the park's electricity operating costs. : ; in For the first Real-time electricity price of the power grid at each time step. For the first The power purchased from the grid at each time step For time step, The total number of time steps in the scheduling cycle. , , The cost coefficient for the diesel generator is obtained through linear regression of the fuel consumption characteristic curve provided by the equipment manufacturer. A typical no-load value example is shown below. =0.0002 yuan / kW 2 , =0.5 yuan / kW =10 yuan / h, the actual value can be dynamically adjusted according to the model of the diesel generator in the park and the current fuel price; For the first The output power of the diesel generator at each time step The unit cost of energy storage charging and discharging, The number of energy storage battery packs, For the first The energy storage battery pack in the first The charging power at each time step For the first The energy storage battery pack in the first Discharge power at each time step For the first The demand response subsidy price at each time step. For the first The demand response power reduction at each time step.

[0056] Objective 2 is to minimize the cumulative frequency deviation of the microgrid in the park. : ;in For the first Microgrid frequency deviation at each time step This is the actual frequency of the microgrid. This is the rated frequency of the power grid.

[0057] Objective 3 is to minimize the energy storage aging rate. : ;in For the first The energy storage battery pack in the first The charge / discharge cycle power at each time step.

[0058] The constraints include power balance constraints, microgrid inertia constraints, distributed generation output constraints, energy storage operation constraints, and computing power task constraints. Among these, the power balance constraint means that the total output of all power sources equals the total power consumption of all loads; the computing power load corresponds to the power load... Using the interval prediction values ​​provided by S2 And based on the predicted interval width Dynamically relax all constraint boundaries: ; in This represents the historical maximum prediction interval width. The dynamic adjustment methods for each constraint are as follows: 1. Microgrid inertia constraints: ; 2. Output constraints of distributed power sources: , , ; 3. Energy storage operation constraints: , , ; 4. Computing power task constraints: .

[0059] To efficiently search the Pareto optimal front, a fusion optimization algorithm combining bacterial foraging and quantum particle swarm optimization is employed, nested with a non-dominated sorting and weighted crowding calculation mechanism. The non-dominated level and weighted crowding are used as the sole criteria for population selection. The population size is set to 80, and the maximum number of iterations is 300.

[0060] Position update fusion chemotaxis operation and quantum particle swarm global attraction: ; in For chemotactic step size, For the direction of chemotaxis, For the first The optimal position of an individual bacterium The optimal position globally. This refers to quantum inertial weights. Linear decay: ; in , , This represents the maximum number of iterations for the fusion algorithm.

[0061] After each iteration, a unified non-dominated sort is performed, dividing all individuals into several non-dominated levels; for individuals within the same level, multi-objective weights are introduced. The weighted congestion degree is calculated using the following formula: ; In the formula , For adjacent individuals at the same level The function value of each objective. , For the Pareto solution set, the first The extreme values ​​of each objective. The preference weights for the corresponding targets.

[0062] The sorting rule is as follows: The population is globally sorted from highest to lowest non-dominated level, and then by weighted crowding density from highest to lowest within the same level. The top half of the sorted individuals are selected as the replication parents, and after replication, they replace the bottom half of the sorted individuals. Subsequent dispersal operations are still performed according to the dispersal probability. Randomly reset the inferior individuals.

[0063] After the algorithm iteration terminates, first-level non-dominated individuals are selected according to the non-dominated level to form a complete Pareto optimal solution set. It is passed to S5.

[0064] Feasible region envelope generation and propagation: Based on the union of all the above constraints, generate the day-ahead scheduling feasible region envelope. Defined as the set of all scheduling schemes that satisfy the constraints of power balance, inertia, power output, energy storage operation, and computing task, mathematically as: ; It is a convex polyhedron defined by linear inequality constraints, containing all scheduling adjustments considered feasible during the day-ahead scheduling phase. and It is passed to S5 along with the other components.

[0065] Receive the scheduling deviation matrix fed back by S6 Dynamically update multi-objective weights The weight update formula is: ; in The daily planned value for the k-th target (k=1,2,3 correspond to power operation cost, microgrid frequency deviation accumulation, and energy storage aging rate, respectively). This refers to the corresponding target value calculated after the actual execution throughout the day.

[0066] Updated By directly substituting it into the weighted congestion calculation of the next iteration, a closed-loop connection is achieved between the actual scheduling effect and multi-objective preferences and population selection rules, and the weight adjustment mechanism is no longer disconnected.

[0067] S5. Multi-objective scheduling decision based on Wasserstein spherical interval type-II fuzzy Nash product: This step will use the predicted probability surface of S2. S3 compliance matrix Pareto front of S4 and feasible domain envelope Deep integration yields the final day-ahead scheduling scheme, forming a two-way adversarial closed loop with S2 (see reference). Figure 4 S4-S5 Optimization Decision Process Reference Figure 3 .

[0068] Using the quantile predictions from S2 output, for each time period and server Constructing an experience reference distribution And define the radius of the Wasserstein-1 sphere. for: ; Uncertain set of joint load ,in For the Cartesian direct product operation, iterate through all server numbers v in the park and couple the independent uncertain sets of each server into a global joint uncertain set; Indicated by Centered on, with radius Wasserstein-1 ball.

[0069] View the park as two competing entities: the computing power scheduling entity focuses on the value of task completion. Power operators are concerned about negative costs. A comprehensive objective is constructed using the generalized Nash bargaining product: ; in This represents the bargaining power coefficient.

[0070] Value of Computing Power Task Completion The specific expression is: ; in This represents the total number of tasks. For the first The priority weight of each task. For the first The completion time of each task. For the first The deadline for each task. For the first Power allocated to each task For the first The maximum power required for each task.

[0071] negative costs of electricity operation The specific expression is: ;in The campus power operation cost is defined in S4.

[0072] Preservation utility is the baseline payoff for both game players operating under a fixed local strategy without cooperative scheduling. It is derived from statistics on the lowest electricity operating cost (i.e., highest negative electricity cost) in the park's historical data over the past 7 days: 1. Retention of computing power as the main body Take the minimum value of task completion within all scheduling cycles over the past 7 days, using the following formula: ; 2. Retention of the main power supply Take the maximum value of the negative electricity cost over all dispatch cycles in the past 7 days, using the following formula: ; Before the start of each day's scheduling, the historical database is automatically updated once to retain the utility value as a fixed input parameter for the Nash bargaining product model for that day. The parameter source, calculation logic, and update cycle are all complete and clear.

[0073] Pareto optimal solution set passed by S4 As a unique set of candidate solutions, for each candidate solution Perform the following robustness assessment: 1. For candidate solutions In Wasserstein's uncertain set Solve the objective function value under the worst-case load distribution. Because... It decreases as the load increases. The value decreases as the load increases, and both correspond to the minimum value in the worst-case scenario. Therefore, the formula is revised as follows: ; ; 2. Substitute the objective value under the worst-case distribution into the Nash bargaining product formula to calculate the comprehensive bargaining score for each candidate solution: ; 3. Simultaneously calculate the robust cost increment corresponding to each candidate solution. Since the cost under the worst distribution will necessarily be higher than the nominal cost, the formula is modified as follows: ; Take all candidate solutions The maximum value is used as feedback to S2. This enables adversarial interaction with S2. Since it is always non-negative, it conforms to the physical meaning of marginal cost increment and can be added to the loss function of S2 as a positive penalty term.

[0074] The Pareto solution set above after robustness evaluation This is the sole source of "a set of candidate solutions" in S5. To select a single scheduling scheme from the candidate solutions, an interval-type II fuzzy logic system is constructed. The input variable is the value corresponding to each candidate solution. , , The corresponding normalized index is output as overall satisfaction. The input variable normalization method is min-max normalization: ; in and These are the minimum and maximum values ​​of the index among all candidate solutions, respectively.

[0075] Both input and output are covered by three interval type II Gaussian membership functions: "low", "medium", and "high". Their upper and lower membership degrees are: ; in The center of the Gaussian function, and These are the Gaussian function widths for the upper and lower membership degrees, respectively. For the normalized input variable, the centers of the three fuzzy sets are 0.2, 0.5, and 0.8, respectively. The value is 0.1. The value is 0.15. The range of the output variable is... The centers of the three fuzzy sets are 0.2, 0.5, and 0.8, respectively. The value is 0.1. The value is 0.15.

[0076] Twenty-seven fuzzy rules are established as follows: 1. If the power operating cost is low, the microgrid frequency deviation is low, and the distributed energy storage aging rate is low, then the overall satisfaction is high; 2. If the power operating cost is low, the microgrid frequency deviation is low, and the distributed energy storage aging rate is moderate, then the overall satisfaction is high; 3. If the power operating cost is low, the microgrid frequency deviation is low, and the distributed energy storage aging rate is high, then the overall satisfaction is moderate; 4. If the power operating cost is low, the microgrid frequency deviation is moderate, and the distributed energy storage aging rate is low, then the overall satisfaction is high; 5. If the power operating cost is low, the microgrid frequency deviation is moderate, and the distributed energy storage aging rate is moderate, then the overall satisfaction is moderate; 6. If the power operating cost is low, the microgrid frequency deviation is moderate, and the distributed energy storage aging rate is high, then the overall satisfaction is moderate. 7. If the power operating cost is low, the microgrid frequency deviation is high, and the distributed energy storage aging rate is low, then the overall satisfaction is moderate; 8. If the power operating cost is low, the microgrid frequency deviation is high, and the distributed energy storage aging rate is moderate, then the overall satisfaction is moderate; 9. If the power operating cost is low, the microgrid frequency deviation is high, and the distributed energy storage aging rate is high, then the overall satisfaction is low; 10. If the power operating cost is moderate, the microgrid frequency deviation is low, and the distributed energy storage aging rate is low, then the overall satisfaction is high; 11. If the power operating cost is moderate, the microgrid frequency deviation is low, and the distributed energy storage aging rate is moderate, then the overall satisfaction is moderate; 12. If the power operating cost is moderate and the microgrid frequency deviation is high, then the overall satisfaction is moderate. 13. If the power operating cost is moderate, the microgrid frequency deviation is moderate, and the distributed energy storage aging rate is low, then the overall satisfaction is moderate; 14. If the power operating cost is moderate, the microgrid frequency deviation is moderate, and the distributed energy storage aging rate is moderate, then the overall satisfaction is moderate; 15. If the power operating cost is moderate, the microgrid frequency deviation is moderate, and the distributed energy storage aging rate is high, then the overall satisfaction is low; 16. If the power operating cost is moderate, the microgrid frequency deviation is high, and the distributed energy storage aging rate is low, then the overall satisfaction is moderate; 17. If the power operating cost is moderate, the microgrid frequency deviation is high, and the distributed energy storage aging rate is moderate, then the overall satisfaction is low; 18. If the power operating cost is moderate, the microgrid frequency deviation is high, and the distributed energy storage aging rate is moderate, then the overall satisfaction is low; 19. If the power operating cost is high, the microgrid frequency deviation is low, and the distributed energy storage aging rate is low, then the overall satisfaction is moderate; 20. If the power operating cost is high, the microgrid frequency deviation is low, and the distributed energy storage aging rate is moderate, then the overall satisfaction is moderate; 21. If the power operating cost is high, the microgrid frequency deviation is low, and the distributed energy storage aging rate is high, then the overall satisfaction is low; 22. If the power operating cost is high, the microgrid frequency deviation is moderate, and the distributed energy storage aging rate is low, then the overall satisfaction is moderate; 23. If the power operating cost is high, the microgrid frequency deviation is moderate, and the distributed energy storage aging rate is moderate, then the overall satisfaction is low; 24.If power operating costs are high, microgrid frequency deviation is moderate, and distributed energy storage aging rate is high, then overall satisfaction is low; 25. If power operating costs are high, microgrid frequency deviation is high, and distributed energy storage aging rate is low, then overall satisfaction is low; 26. If power operating costs are high, microgrid frequency deviation is high, and distributed energy storage aging rate is moderate, then overall satisfaction is low; 27. If power operating costs are high, microgrid frequency deviation is high, and distributed energy storage aging rate is high, then overall satisfaction is low.

[0077] Calculate the applicability range of each rule during inference. : ; Normalize the applicability range: ; The upper center of gravity is obtained using the center of gravity method. and lower center of gravity : ; ; in , The final output is a clear value. Calculate the solution sequentially for all candidate solutions. The solution with the highest satisfaction level is selected as the optimal computer-computer collaborative scheduling scheme. Specifically, this includes the power purchase plan for the next 24 hours in 15-minute increments, the target number of active servers, the task pause migration matrix, energy storage charging and discharging instructions, and chiller unit settings.

[0078] The feasible region envelope passed by S4 Together with the optimal scheduling scheme It is also passed to S6 as a constraint boundary for the real-time scheduling phase.

[0079] S6. Real-time elastic scheduling and online model evolution based on Lyapunov drift and information gap decision-making: This step receives the day-ahead baseline plan from S5. S2 ultra-short-term forecast, S3 compliance matrix and the feasible domain envelope of S5 It executes in a rolling cycle of five minutes and forms a real-time closed loop with S2 and S5.

[0080] Define three virtual queues to track violations of the system's security constraints: Track the cumulative amount of time the server ingress temperature exceeds the safety limit. Track the extent to which the average response time of tracking tasks violates the service level agreement. Track the cumulative deviation of the energy storage state of charge from the reference trajectory. The update rule is: ; ; ; in The server's inlet temperature. To the upper limit of safe temperature, The average response time for the task. The maximum response time specified in the service level agreement. This represents the actual state of charge of the energy storage. Construct the Lyapunov function as a reference state of charge. Its single-step drift It is used to measure the stability of the system state.

[0081] When peaks may occur in the S2 ultra-short-term forecast and real-time electricity price display, information gap decision theory is introduced to quantify the impact of electricity price uncertainty on dispatch costs. A nested iterative two-level optimization structure is used to solve the problem. 1. Outer layer optimization: Fixed-day baseline plan Solve for the maximum permissible uncertain radius. : ;in This refers to the planned power purchase capacity at the border crossing. The maximum acceptable electricity purchase cost is set at 120% of the day-ahead planned electricity purchase cost, or 1.5 times the average cost based on the park's historical operating data. This represents the upper limit of the uncertainty radius for electricity prices. The example value is 0.5, which means that the electricity price can fluctuate upwards by a maximum of 50%. The actual value can be adjusted according to the historical electricity price volatility of the park.

[0082] 2. Inner layer optimization: In The following approach minimizes drift and penalties to reduce scheduling costs while ensuring system stability: ; The constraints include the feasible region envelope of S5. and flexibility matrix Decision variables This covers task pause and resume, graphics processor frequency adjustment, energy storage power, and fine-tuning of chiller setpoints. This small-scale 0-1 integer programming problem can be solved in under 5 seconds using the branch and bound method. The energy storage aging cost for the current period is calculated from the energy storage aging rate term in S4. * The sum is obtained by accumulation.

[0083] A dual closed-loop feedback mechanism is constructed. The inner loop makes rapid adjustments on a one-minute basis, inputting the actual deviation signal into the proportional-integral-derivative (PID) controller. ; in This is the proportionality coefficient. The integral coefficient is... The differential coefficients are... This is a deviation signal. Fine-tuning of scheduling instructions within 15 minutes allows for rapid response to real-time fluctuations. The outer loop, on a daily basis, compares the actual load, temperature, and energy storage trajectory throughout the day with the S2 forecast and S5 plan to generate a prediction error matrix. and scheduling deviation matrix The results are fed back to S1, S3, and S4 respectively, and used as incremental training samples. The neural differential equation parameters in S2 are fine-tuned using the elastic weight consolidation algorithm, and the Lyapunov penalty weights in S6 are adjusted using Bayesian optimization. Nash weights of S5 The loss function of the elastic weight consolidation algorithm is: ; in For the loss of new samples, The regularization coefficient is . These are the diagonal elements of the Fisher information matrix. These are the current model parameters. These are the parameters of the old model. This algorithm can retain old knowledge while learning new knowledge, avoiding catastrophic forgetting.

[0084] S7. Multi-dimensional evaluation of the effectiveness of computer-powered collaborative scheduling and emergency switching under abnormal operating conditions: It receives full-day actual execution data recorded by S6 and performs quantitative evaluation from five dimensions: economy, stability, environmental friendliness, computing power service quality, and energy storage lifespan. Economic indicators include actual electricity operating cost, unit computing power electricity cost, and net demand response revenue. The formula for calculating unit computing power electricity cost is: ; Total computing power output The calculation formula is: ;in For the first The server is at the Graphics processor utilization at each time step For central processing unit utilization, This refers to memory utilization.

[0085] Stability indicators include the maximum frequency deviation, average frequency deviation, standard deviation, and voltage pass rate. The formula for calculating the voltage pass rate is: ; in The time during which the voltage remains within the acceptable range. This refers to the total operating time. Environmental performance indicators include carbon dioxide emissions and diesel consumption. The formula for calculating carbon dioxide emissions is: ; in This refers to the carbon dioxide emissions from diesel combustion. The quality of computing power services includes on-time task completion rate, average response time, and computing power utilization. The formula for calculating the on-time task completion rate is: ; in To the number of tasks completed on time, This represents the total number of missions. Energy storage life indicators include average health status and average aging rate.

[0086] The indicators are compared with preset thresholds. The conditions for triggering abnormal operating conditions include: power operating costs exceeding 20% ​​of the threshold, frequency deviation exceeding ±0.2 Hz, task on-time completion rate below 95%, any energy storage health status below 80%, distributed power output dropping by more than 50%, or the park actively disconnecting from the grid. Four emergency dispatch modes are activated accordingly: economic priority mode, stability priority mode, computing power priority mode, and energy storage protection mode. In economic priority mode, computing power allocation for non-critical tasks is reduced, energy storage discharge power is increased, and electricity purchase from the grid is reduced. In stability priority mode, diesel generators are immediately started, energy storage charging and discharging power is adjusted to provide inertia support, and the execution of non-critical tasks is restricted. In computing power priority mode, computing power requirements for critical tasks are prioritized, electricity purchase from the grid is increased, and energy storage charging is suspended. In energy storage protection mode, battery packs with excessively low health status are taken out of operation, and the charging and discharging plans of other battery packs are adjusted. Simultaneously, an assessment and anomaly handling report is generated and stored in the database for use by S8.

[0087] S8. Data-driven global closed-loop self-evolution: Evaluation data for S7 and historical execution samples from S2, S3, S4, S5, and S6 were collected, and three self-evolutionary operations were performed on the models. Firstly, the non-intrusive load decomposition model and server power-computing power mapping coefficients of S1 were updated by incorporating the latest operational data through online transfer learning. The objective function of the transfer learning was: ; in For the target domain loss, For source domain loss, First, to balance the weights for transfer learning. Second, using the deviation between the actual load and the predicted S2 value as incremental samples, the elastic weight consolidation algorithm is used to fine-tune the parameters of the temporal hypergraph neural differential equation, so that it both preserves historical patterns and adapts to recent changes. Third, Bayesian optimization is used to automatically adjust the Lyapunov penalty coefficient of S6. S4 Multi-Objective Weights And the Nash bargaining factor of S5 The objective function of Bayesian optimization is: ; in , , These are the weighting coefficients. , , For reference only. After all updates are completed, the system seamlessly switches to the next day's S1→S2→…→S7 process, forming a never-ending closed loop of intelligent evolution through computing and power collaboration.

[0088] In some embodiments, the chaos control parameters in the fractional-order chaotic bat algorithm The selection criteria are: when At that time, the Logistic mapping is in a chaotic state, and taking... This approach achieves maximum ergodicity, a value widely adopted in standard chaotic optimization literature. For different operating conditions, the prediction error matrix is ​​obtained by receiving feedback from S6. Adaptive fine-tuning: ; The number and relationships of hyperedges in the temporal hypergraph are fixed by the physical topology and require no optimization. Edge weight matrix. The optimization employs another independent instance of the fractional-order chaotic bat algorithm, with a maximum number of iterations of Population size Fractional order order search range The weighting coefficients in the weighted sum of node degree distribution entropy and average clustering coefficient. Based on historical data of the park, a grid search was used to determine the maximum degree of the hypergraph module.

[0089] In some embodiments, the number of modes in the variational mode decomposition It is not a fixed value, but is adaptively determined based on the frequency domain energy distribution of the load signal: firstly, the computing power load timing is... Perform a Fast Fourier Transform to count the number of local maxima in the spectrum. ,make That is, it must be decomposed into at least 3 modes and at most 8 modes. Simultaneously, it checks whether the center frequencies of each mode overlap; if the difference between the center frequencies of adjacent modes is less than the average center frequency difference... If so, then the two modes are merged and decomposed again.

[0090] The vector field network in the temporal hypergraph neural differential equation The specific structure is as follows: Input layer: Node hidden state Node features Aggregating neighbor information External coercion Dimensions after splicing .

[0091] Three fully connected layers, each with the following number of neurons: , , Each layer is followed by LayerNorm and ReLU activation functions.

[0092] The output layer is a linear layer, and the output... .

[0093] Attention parameters , All of them learn from the data through the Adam optimizer, and are initialized using Xavier.

[0094] In some embodiments, the fitness function of the improved gray wolf optimization algorithm The prediction target is: microgrid frequency change rate. The predicted value. Specifically, the extended Kalman filter estimates the predicted value of the output frequency change rate based on the current state. The gray wolf optimization minimizes the sum of squares between the predicted and actual measurements. (Sliding window length) correspond seconds (sampling period) (seconds), per Updated every minute.

[0095] The computing power task elasticity category to flexibility contribution matrix The mapping rules are as follows: Rigid Uninterruptible Tasks: Compliance Factor Adjustable power ; The task can be paused briefly: Longest pause duration Minutes, adjustable power ; Tasks can be paused for extended periods: Longest pause duration Minutes, adjustable power ; Tasks that can run at reduced frequency: Lowest frequency Rated frequency, adjustable power ; Tasks that can be migrated across nodes: The migration time shall not exceed Minutes, adjustable power .

[0096] in For server During the period Real-time power consumption (obtained through non-intrusive monitoring). .

[0097] In some embodiments, the bacterial foraging-quantum particle swarm fusion algorithm includes: chemotactic step length Adaptive strategy adopted: ,in For individuals The fitness (the inverse of the Pareto non-dominated level). , The best and worst fitness of the current population, Prevent division by zero.

[0098] chemotaxis exist Uniform random sampling is performed within each generation, and the data is regenerated in each generation.

[0099] Dispel probability For a fixed value, the new position is uniformly and randomly generated within the decision variable space during dispersal.

[0100] Quantum inertial weight linear decay from to It is updated after each iteration.

[0101] In some embodiments, the specific method for constructing the Wasserstein uncertain set is as follows: For each server and time period S2 output , , , , Quantiles are used to construct empirical distributions. The distribution is discrete, and the support points are these five quantiles. The probability mass of each point is 1. Wasserstein-1 sphere radius Pick and The larger one, multiplied by the sensitivity coefficient. The final radius is then obtained.

[0102] The worst distribution expectation in the sub-bulb optimization This is achieved by solving the following dual linear programming problem (each candidate solution is solved independently): ; in For cost or value function, For transportation planning. This linear programming problem uses the commercial solver Gurobi. Solve within seconds. When the number of candidate solutions exceeds... At that time, first through - Constraint method selects the most Robustness evaluation is then performed after identifying a non-dominated solution.

[0103] In some embodiments, the prediction error matrix in S6 The element is defined as: ; in For server At any moment The actual load (equivalent computing power load index). This is the median prediction output by S2. To predict the interval width, To prevent division by zero, the matrix is ​​calculated hourly, stored in the database, and fed back to S1 and S2 at the end of each day.

[0104] The scheduling deviation matrix The three lines are defined as follows: First line: Active power deviation at the threshold ; Second line: Energy storage state of charge deviation , ; Third line: Server ingress temperature deviation .

[0105] At the end of each day, and Feedback is sent to S1, S3, and S4 in matrix form for adaptive parameter adjustment.

[0106] In some embodiments, the Fisher information matrix in the elastic weight consolidation algorithm diagonal elements Calculated using the following formula: ; in The number of historical playback samples (take) ), The conditional probability density of the prediction model output given the input (for regression tasks, assume a Gaussian distribution with fixed variance). After each epoch of incremental training, the most recent... The Fisher information matrix is ​​recalculated based on the data from the previous day. Regularization coefficients are then used. Optimize the loss based on the validation set.

[0107] The source domain for the online transfer learning is simulation data and a general server power-computing power model used during the park design phase or initial operation phase, while the target domain is actual operational data. Transfer learning balances weights. Adaptive rules are used: ; in This represents the mean squared error. Each The mapping coefficients are updated once a day.

[0108] In some embodiments, the parameters of the proportional-integral-derivative (PID) controller The Ziegler-Nichols method combined with online adaptive adjustment was employed: a step response experiment was conducted before the first use to obtain the critical gain. and critical period ,set up After that, every time it runs... Fine-tuning was performed using gradient descent over several hours. ; Learning rate The optimization objective is the cumulative absolute deviation integral over the past 24 hours. Synchronous adjustment, maintain , The proportional relationship.

Claims

1. A method for collaborative scheduling of computing power and electricity in intelligent computing parks based on computing power load prediction, characterized in that, include: S1. Collect multi-source heterogeneous data from the computing power side, power side, environment side, and cooling system of the intelligent computing park. After imputing missing values ​​and removing outliers from the multi-source heterogeneous data, use the fractional-order chaotic bat algorithm to determine the optimal fractional-order number, and perform fractional-order differential transformation on each data sequence accordingly. At the same time, construct a time-series hypergraph representing the power supply, pipeline, and thermal coupling relationship between devices, and optimize the edge weight matrix of the time-series hypergraph. S2. Based on the data after fractional-order differential transformation and the edge weight matrix, perform multi-scale variational mode decomposition on the computing load time series, construct a time series hypergraph neural differential equation prediction model, and output a probability prediction surface containing multiple quantiles and the prediction interval width; introduce the robust cost increment fed back by S5 as a regularization term in the training of the prediction model to form an adversarial learning closed loop. S3. Based on the transformed data and the edge weight matrix, the improved gray wolf optimization and extended Kalman filter fusion algorithm is used to identify the equivalent inertia and energy storage aging coefficient of the microgrid in real time, and the computing power tasks are flexibly classified to generate a task flexibility contribution matrix. S4. Based on the edge weight matrix, the prediction interval width, the microgrid equivalent inertia, and the energy storage aging coefficient, with the power operation cost, microgrid frequency deviation accumulation, and energy storage aging rate as objectives, and dynamically relaxing the constraint boundary according to the prediction interval width, the bacterial foraging-quantum particle swarm fusion algorithm is adopted to filter according to the non-dominated level and the weighted congestion degree calculated by the multi-objective preference weight, and to generate the Pareto optimal solution set and scheduling feasible domain envelope. S5. Construct a Wasserstein uncertainty set based on the probability prediction surface, perform robustness evaluation on each candidate solution in the Pareto optimal solution set, calculate the generalized Nash bargaining product score weighted by the Nash bargaining coefficient, and use an interval type II fuzzy logic system to select the optimal day-ahead scheduling scheme; at the same time, calculate the robust cost increment corresponding to each candidate solution, and feed back the maximum value of the robust cost increment to S2; S6. Based on the optimal day-ahead scheduling scheme, and combined with the ultra-short-term prediction results, the task flexibility contribution matrix, and the scheduling feasible region envelope, minute-level dual-closed-loop real-time elastic scheduling is performed using Lyapunov drift plus information gap decision theory to generate a prediction error matrix and a scheduling deviation matrix, which are then fed back to S1, S3, and S4 respectively.

2. The intelligent computing park computing power collaborative scheduling method based on computing power load prediction according to claim 1, characterized in that, In S1, the fractional-order chaotic bat algorithm uses the weighted sum of the maximum autocorrelation coefficient of the data sequence and the stationarity index as the fitness function; the initial value of the edge weight matrix of the time-series hypergraph is determined by the absolute value of the Pearson correlation coefficient between nodes, and is dynamically optimized through another instance of the fractional-order chaotic bat algorithm using the weighted sum of the node degree distribution entropy and the average clustering coefficient as the fitness function; S1 receives the prediction error matrix and the scheduling deviation matrix fed back from S6, and adaptively adjusts the optimal fractional-order order and the maximum inertia weight in the fractional-order chaotic bat algorithm.

3. The intelligent computing park computing power collaborative scheduling method based on computing power load prediction according to claim 1, characterized in that, In S2, the quadratic penalty factor of the variational mode decomposition is positively correlated with the optimal fractional order obtained in S1; the temporal hypergraph neural differential equation prediction model uses a Dormand-Prince fifth-order Runge-Kutta integrator for forward integration, outputting predicted values ​​with quantiles of 0.05, 0.25, 0.50, 0.75, and 0.95; the robust cost increment is added as a regularization term to the training loss function of the prediction model; wherein the robust cost increment is the difference between the power operation cost under the worst distribution and the power operation cost under the nominal distribution.

4. The intelligent computing park computing power collaborative scheduling method based on computing power load prediction according to claim 1, characterized in that, In S3, the improved gray wolf optimization algorithm employs a nonlinear convergence factor and optimizes the process noise covariance and measurement noise covariance of the extended Kalman filter using the sum of squared prediction errors as the fitness function. The improved gray wolf optimization algorithm is executed by combining offline pre-training with online sliding window rolling updates. The elasticity categories of the computing power tasks are divided into five types: rigid uninterruptible tasks, tasks that can be paused for a short time, tasks that can be paused for a long time, tasks that can run at reduced frequency, and tasks that can migrate across nodes. S3 receives the scheduling deviation matrix fed back from S6 and adaptively adjusts the convergence factor of the improved gray wolf optimization algorithm.

5. The intelligent computing park computing power collaborative scheduling method based on computing power load prediction according to claim 1, characterized in that, In S4, the dynamic relaxation of the constraint boundary is achieved by: calculating a dynamic adjustment coefficient based on the predicted interval width, and multiplying the upper and lower limits of the microgrid inertia constraint, distributed power output constraint, energy storage operation constraint, and computing power task constraint by the dynamic adjustment coefficient; the weighted congestion calculation introduces the multi-objective preference weight, and the population screening is performed in the order of non-dominated level from high to low, and weighted congestion of the same level from large to small; S4 receives the scheduling deviation matrix fed back by S6, dynamically updates the multi-objective preference weights, and substitutes them into the weighted congestion calculation for the next iteration.

6. The intelligent computing park computing power collaborative scheduling method based on computing power load prediction according to claim 1, characterized in that, In S5, the specific evaluation of the biblical robustness is as follows: within the Wasserstein uncertainty set, solve for the minimum value of the computational task completion and the minimum negative cost of power operation for each candidate solution under the worst load distribution. The generalized Nash bargaining product takes the minimum value as input and is based on the retention utility obtained from statistics of the lowest power operating cost in the park's history over the past seven days.

7. The intelligent computing park computing power collaborative scheduling method based on computing power load prediction according to claim 1, characterized in that, In S6, the dual closed-loop real-time elastic scheduling is specifically as follows: the inner loop fine-tunes the scheduling instructions through the proportional-integral-derivative controller at the minute level to quickly respond to real-time fluctuations; the outer loop generates the prediction error matrix and the scheduling deviation matrix on a daily basis and feeds them back in the opposite direction. The information gap decision theory adopts a two-layer nested optimization structure. The outer layer solves for the maximum allowable uncertainty radius of electricity price, and the inner layer minimizes the Lyapunov drift plus penalty function under this radius.

8. The intelligent computing park computing power collaborative scheduling method based on computing power load prediction according to claim 1, characterized in that, Also includes: S7. Collect actual execution data throughout the day and evaluate the scheduling effect from five dimensions: economy, stability, environmental protection, computing power service quality and energy storage life. When abnormal operating conditions are triggered, automatically switch to the corresponding emergency scheduling mode. S8. Summarize the entire process operation data and evaluation data, and update the model parameters of each stage through online transfer learning, elastic weight consolidation algorithm and Bayesian optimization method to achieve global closed-loop self-evolution.

9. The intelligent computing park computing power collaborative scheduling method based on computing power load prediction according to claim 8, characterized in that, In S7, the abnormal operating condition judgment conditions include: power operation cost exceeding the threshold by 20%, frequency deviation exceeding ±0.2 Hz, task on-time completion rate below 95%, any energy storage health status below 80%, distributed power output dropping by more than 50%, or the park actively disconnecting from the grid; the emergency dispatch mode includes economic priority mode, stability priority mode, computing power priority mode, and energy storage protection mode.

10. The intelligent computing park computing power collaborative scheduling method based on computing power load prediction according to claim 8, characterized in that, In S8, the global closed-loop self-evolution specifically includes: updating the non-intrusive load decomposition model and server power-computing power mapping coefficients through online transfer learning; fine-tuning the parameters of the time-series hypergraph neural differential equation prediction model through the elastic weight consolidation algorithm; automatically adjusting the Lyapunov penalty coefficient, the multi-objective preference weight, and the Nash bargaining coefficient in the penalty term of the Lyapunov drift through Bayesian optimization; and seamlessly switching to the scheduling process of the next day after all model updates are completed.