A power supply energy efficiency intelligent scheduling power distribution system for a large computing power data center

CN122697348APending Publication Date: 2026-09-04NINGBO SIHONG ELECTRICAL APPLIANCE IND
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611082682.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-07-21
Publication Date
2026-09-04

AI Technical Summary

Technical Problem

现有集中式配电调度存在响应滞后、整体容错性差的问题,传统分布式调度仅依靠固定阈值本地管控,缺少全局用电工况统筹优化能力,两类方案均难以满足当下高密度算力机房综合运营需求

Benefits of technology

本发明通过数据采集层统一获取电力侧运行参数与算力侧运行参数,同步对接算力任务调度平台传递算力任务的预测功耗参数与机柜剩余供电容量,为后续配电调控提供统一的数据基础,部署在各机柜列的列头柜内的边缘协同预调度层依靠对应机柜列的实时采集数据快速输出本地应急配电调节指令,无需将现场数据上传至中心平台集中运算,有效规避传统集中式调度全部数据集中运算带来的处理延迟问题,部署于数据中心管控平台的全局能效调度层根据能够通过历史调度数据持续更新的预设数字孪生配电模型,结合全数据中心的采集数据与市电峰谷时段参数生成全域配电配置数据,弥补传统分布式调度仅依靠固定阈值、无法统筹全中心用电工况的短板,本地应急配电调节指令与全域配电配置数据采用互不干扰的独立通路下发至配电执行层,配电执行层优先处理本地应急配电调节操作,再按照全域配电配置基于全域配电配置数据执行配电回路的通断、电源输出功率调整、冗余电源启停操作,各功能层级相互配合,本地应急处置无需等待全局运算流程,全局配电优化又基于完整实时的电力侧运行参数、算力侧运行参数调整配电策略,两类调控流程互不挤占传输资源与处理资源,局部机柜供电故障仅触发对应本地调控,不会影响整个数据中心调度体系运行,大幅提升整体容错能力,既能在机房出现供电异常时快速介入调节,保障高密度算力业务持续稳定运行,也能结合市电峰谷规律均衡分配各个机柜的用电负荷,充分释放整体配电节能空间,同步提升供电运行稳定性与整体用电经济性。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122697348A_ABST
    Figure CN122697348A_ABST
Patent Text Reader

Abstract

The application discloses a power supply energy efficiency intelligent scheduling power distribution system for a large computing power data center, relates to the field of power supply energy efficiency scheduling of the large computing power data center, and comprises a data acquisition layer, an edge cooperative pre-scheduling layer, a global energy efficiency scheduling layer and a power distribution execution layer. The data acquisition layer synchronously acquires full-amount operating parameters of the power side and the computing power side of the data center, the edge cooperative pre-scheduling layer is arranged in each cabinet column head cabinet, pre-scheduling response is conducted on local sudden computing power load fluctuation, and local emergency power distribution regulation instructions are generated. The global energy efficiency scheduling layer is arranged in a data center management and control platform, generates global power distribution configuration data based on preset preset digital twin power distribution models and in combination with acquisition data of the whole data center and power supply peak-valley period parameters. The application adopts a two-stage cooperative scheduling architecture, rapidly responds to load mutation through the edge side, avoids service interruption caused by overload triggering protection, and improves power supply reliability.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of power energy efficiency scheduling for high-performance data centers, and particularly to a power energy efficiency intelligent scheduling and distribution system for high-performance data centers. Background Technology

[0002] High-power data centers support high-power services such as AI training and high-performance computing, leading to continuously increasing power consumption per rack. The level of power supply scheduling directly affects power supply reliability, energy efficiency, and electricity costs. Existing centralized power distribution scheduling suffers from slow response and poor overall fault tolerance, while traditional distributed scheduling relies solely on local control with fixed thresholds, lacking the ability to comprehensively optimize global power consumption conditions. Neither of these solutions can meet the comprehensive operational needs of today's high-density computing data centers. Summary of the Invention

[0003] This invention proposes a power efficiency intelligent scheduling and distribution system for high-computing-power data centers to solve the problems mentioned in the prior art.

[0004] To achieve the above objectives, the present invention provides a power efficiency intelligent scheduling and distribution system for high-computing-power data centers, comprising a data acquisition layer, an edge collaborative pre-scheduling layer, a global energy efficiency scheduling layer, and a power distribution execution layer; The data acquisition layer is used to synchronously collect the power-side operating parameters and computing-side operating parameters of the high-performance computing data center. The power-side operating parameters include the mains power line electrical parameters, UPS output operating parameters, rack power distribution circuit parameters, and energy storage unit power parameters. The computing-side operating parameters include the real-time computing load, chip operating status, and power consumption parameters of each computing node. The data acquisition layer receives the predicted power consumption parameters of the computing tasks generated by the computing task scheduling platform and sends the remaining power supply capacity of the rack to the computing task scheduling platform. The edge collaborative pre-scheduling layer is deployed in the head cabinet of each rack column to receive the data collected from the corresponding rack column, collect local computing load data in real time, and generate local emergency power distribution adjustment commands. The global energy efficiency scheduling layer is deployed on the data center management platform. It is used to generate multiple sets of power distribution configuration parameters based on the preset digital twin power distribution model that maps the power distribution hardware and computing load operation status of the computer room, combined with the collected data of the entire data center and the peak and valley time parameters of the mains power. It outputs a set of global power distribution configuration data every hour. The global energy efficiency scheduling layer sends the global power distribution configuration data to the computing task scheduling platform and updates the preset digital twin power distribution model with historical scheduling data. The path used to transmit the global power distribution configuration data is independent of the path used to send local emergency power distribution adjustment commands. The power distribution execution layer is used to receive local emergency power distribution adjustment commands and global power distribution configuration data, and prioritizes processing local emergency power distribution adjustment commands. Based on the global power distribution configuration data, it performs power distribution circuit switching, power output power adjustment, and redundant power supply start-up and shutdown operations.

[0005] Preferably, the data acquisition layer includes power sensors, power consumption acquisition probes, and computing load acquisition units. Multiple power sensors are arranged one-to-one at the mains power input end, UPS output end, and each PDU input end. Each CPU computing node and each GPU computing node is individually configured with a power consumption acquisition probe and a computing load acquisition unit.

[0006] Preferably, each edge processing unit in the edge collaborative pre-scheduling layer pre-stores the power consumption baseline curves of all computing nodes in the corresponding rack column. The edge processing unit compares the local computing load fluctuation with the preset fluctuation threshold. When the local computing load fluctuation exceeds the preset fluctuation threshold, it adjusts the output power of redundant power supplies and the power supply priority of non-core loads within the range of the local rack column. The generated local emergency power distribution adjustment command is directly sent to the power distribution execution layer for execution.

[0007] Preferably, the global energy efficiency scheduling layer is configured with a preset digital twin power distribution model. The preset digital twin power distribution model uses historical energy efficiency data of the data center, real-time collected power side and computing side parameters, and power distribution equipment health parameters as the data source. The preset digital twin power distribution model performs simulation calculations on the generated multiple sets of power distribution configuration parameters, outputs the corresponding PUE value and power supply reliability parameters of the entire data center, and outputs a scheduling strategy that matches the power distribution configuration data of the entire domain.

[0008] Preferably, the global energy efficiency scheduling layer performs time-sharing power distribution control based on the scheduling strategy output by the preset digital twin power distribution model; during the off-peak period of the mains power, the global energy efficiency scheduling layer controls the redundant power supply to operate at full load while charging the supporting energy storage unit; during the peak period of the mains power, the energy storage unit is called to supply power to the outside, while the power supply circuit of the redundant power supply module in the no-load state is shut down.

[0009] Preferably, the power distribution execution layer is equipped with a multi-redundant power supply switching module. The multi-redundant power supply switching module is used for online adjustment of redundant power supplies. When the scheduling strategy determines that the redundancy of the redundant power supply exceeds the set threshold, the multi-redundant power supply switching module controls the excess redundant power supply to enter a low-power sleep state. When the collected load data increases, the multi-redundant power supply switching module performs a redundant power supply wake-up operation.

[0010] Preferably, the edge collaborative pre-scheduling layer is also equipped with a power supply fault prediction unit. The power supply fault prediction unit is used to predict whether there is a risk of power supply fault based on the ripple value of the collected power parameters and the trend of equipment temperature change. When a risk of power supply fault is predicted, the corresponding load is migrated to the redundant power supply circuit of the adjacent rack row.

[0011] Preferably, the global energy efficiency scheduling layer is also configured with an efficiency audit module. The efficiency audit module is used to statistically analyze the PUE value, power supply reliability parameters, and electricity cost of the entire data center after each scheduling, and use the statistical results as training samples to update the parameters of the preset digital twin power distribution model.

[0012] Preferably, the global energy efficiency scheduling layer communicates with the computing power task scheduling platform of the data center. The global energy efficiency scheduling layer synchronously outputs scheduling strategies to the computing power task scheduling platform, driving computing power tasks to be allocated to racks that meet the power supply redundancy and energy efficiency standards as determined by the preset digital twin power distribution model.

[0013] This invention also provides a power efficiency intelligent scheduling method for high-performance data centers, applied to the aforementioned power efficiency intelligent scheduling power distribution system for high-performance data centers, comprising the following steps: Step 1: Synchronously collect the power-side operating parameters, computing-side operating parameters, and predicted power consumption parameters of the computing tasks to be sent from the high-performance computing data center through the data acquisition layer. Step 2: The edge collaborative pre-scheduling layer judges the sudden computing load fluctuation of the local rack column. When the load fluctuation exceeds the preset fluctuation threshold, a local emergency power distribution adjustment command is generated and sent to the power distribution execution layer. At the same time, the load fluctuation data is synchronized to the global energy efficiency scheduling layer. Step 3: Based on the preset digital twin power distribution model, the global energy efficiency scheduling layer generates an hourly scheduling strategy by combining the collected data from the entire data center and the peak and valley time parameters of the mains power, and sends it to the power distribution execution layer for execution. Step four: Execute the corresponding scheduling instructions through the power distribution execution layer, and perform power distribution circuit switching, power output adjustment, and redundant power supply start-up and shutdown operations based on the global power distribution configuration data.

[0014] Compared with existing technologies, the beneficial effects of this invention are: This invention obtains power-side and computing-side operating parameters uniformly through a data acquisition layer, and synchronously connects with the computing task scheduling platform to transmit predicted power consumption parameters and remaining power supply capacity of the racks, providing a unified data foundation for subsequent power distribution control. An edge-coordinated pre-scheduling layer deployed in the head racks of each rack column quickly outputs local emergency power distribution adjustment commands based on real-time data collected from the corresponding rack column, eliminating the need to upload field data to a central platform for centralized processing. This effectively avoids the processing delays caused by centralized processing of all data in traditional centralized scheduling. A global energy efficiency scheduling layer deployed on the data center management platform generates global power distribution configuration data based on a preset digital twin power distribution model that is continuously updated through historical scheduling data, combined with data collected from the entire data center and peak / valley time parameters of the mains power. This overcomes the shortcomings of traditional distributed scheduling, which relies solely on fixed thresholds and cannot coordinate the overall power consumption conditions of the entire center. Local emergency power distribution adjustment commands and global power distribution configuration data... Data is sent to the power distribution execution layer via independent, non-interfering pathways. The power distribution execution layer prioritizes local emergency power distribution adjustment operations, and then executes power circuit switching, power output adjustment, and redundant power supply start-up and shutdown operations based on the global power distribution configuration data. Each functional level cooperates with each other, and local emergency response does not need to wait for the global calculation process. Global power distribution optimization adjusts the power distribution strategy based on complete and real-time power-side operating parameters and computing power-side operating parameters. The two types of control processes do not crowd out transmission and processing resources. Local cabinet power supply failures only trigger corresponding local control and will not affect the operation of the entire data center scheduling system, greatly improving the overall fault tolerance. It can quickly intervene and adjust when power supply anomalies occur in the computer room to ensure the continuous and stable operation of high-density computing power services. It can also balance the power load of each cabinet according to the peak and valley patterns of the mains power, fully release the overall power distribution energy-saving space, and simultaneously improve the stability of power supply operation and the overall power economy. Attached Figure Description

[0015] Figure 1 This is an overall flowchart of the intelligent power efficiency scheduling method proposed in this invention; Figure 2 This is a flowchart illustrating the multi-source data acquisition and processing flow on the power and computing sides of this invention. Figure 3 This is a flowchart of the millisecond-level pre-scheduling and fault prediction process for edge collaboration in this invention; Figure 4 This is a flowchart of the global optimal energy efficiency scheduling process based on a preset digital twin model, as described in this invention. Detailed Implementation

[0016] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0017] Reference Figures 1 to 4 The first aspect of this invention discloses a power efficiency intelligent scheduling and distribution system for high-computing-power data centers, comprising a data acquisition layer, an edge collaborative pre-scheduling layer, a global energy efficiency scheduling layer, and a power distribution execution layer. The data acquisition layer and the edge collaborative pre-scheduling layer can interact via industrial Ethernet; the edge collaborative pre-scheduling layer, the global energy efficiency scheduling layer, and the power distribution execution layer can be interconnected based on a 5G deterministic communication link.

[0018] The data acquisition layer is used to synchronously collect the power-side operating parameters and computing-side operating parameters of the high-performance computing data center. The power-side operating parameters include the mains power line electrical parameters, UPS output operating parameters, rack power distribution circuit parameters, and energy storage unit power parameters. The computing-side operating parameters include the real-time computing load, chip operating status, and power consumption parameters of each computing node. The data acquisition layer receives the predicted power consumption parameters of the computing tasks generated by the computing task scheduling platform and sends the remaining power supply capacity of the rack to the computing task scheduling platform.

[0019] For example, the sampling frequency of the data acquisition layer is no less than 1kHz, such as 1.2kHz. The power-side operating parameters can be collected at points covering the three-phase voltage, current, power factor, and 3rd-51st harmonic content parameters at the 10kV mains input terminal; the DC voltage, output current, and remaining battery capacity parameters at the online UPS output terminal; the circuit on / off status and real-time load rate parameters at the PDU input terminals of each rack row; and the SOC value and charging / discharging power parameters of the lithium iron phosphate energy storage unit. Computing-side operating parameters can be synchronously collected through the server BMC interface, CPU MSR register, and GPU SMI interface, covering the real-time computing load, chip operating frequency, real-time power consumption, memory occupancy rate, and power consumption parameters of the supporting liquid cooling module for each CPU and GPU computing node. Predicted power consumption parameters can include the peak power consumption per unit time, continuous running time, and the time node when the power consumption peak occurs for batch AI training tasks and online inference tasks.

[0020] The edge collaborative pre-scheduling layer is deployed in the head cabinet of each rack column. It is used to receive the collected data of the corresponding rack column, collect local computing load data in real time, and generate local emergency power distribution adjustment commands. For example, the edge collaborative pre-scheduling layer may include multiple edge computing units. Each edge computing unit may use an Intel Atom x7000 series processor, run the VxWorks real-time operating system, and manage 8-12 (e.g., 10) computing racks. It is used to receive the collected data of the corresponding rack column, and the response latency to sudden fluctuations in local computing load is no more than 20ms, for example, 18ms, and generates local emergency power distribution adjustment commands.

[0021] The global energy efficiency scheduling layer is deployed on the data center management platform. It is used to generate multiple sets of power distribution configuration parameters based on the preset digital twin power distribution model that maps the power distribution hardware and computing load operation status of the computer room, combined with the collected data of the entire data center and the peak and valley time parameters of the mains power. It outputs a set of global power distribution configuration data every hour. The global energy efficiency scheduling layer sends the global power distribution configuration data to the computing task scheduling platform and updates the preset digital twin power distribution model with historical scheduling data. The path used to transmit the global power distribution configuration data is independent of the path used to send local emergency power distribution adjustment commands.

[0022] The preset digital twin power distribution model is a pre-built digital simulation model that maps one-to-one with the actual power distribution hardware and computing clusters of the data center. It incorporates all electrical parameters, topology, and power consumption characteristics of the power distribution equipment, and can perform operational simulations in conjunction with power-side and computing-side operating parameters. The global energy efficiency scheduling layer combines data collected from the data center with grid peak-valley time parameters, including peak-valley time division standards, electricity prices for each time period, and grid constraints, to deduce multiple sets of power distribution configuration parameters. Grid constraints refer to the maximum load limit, allowable power fluctuation range, and mandatory peak-shaving requirements imposed by the external power grid on the data center. The global energy efficiency scheduling layer compares all power distribution configuration parameters based on energy efficiency indicators and power supply safety requirements, selects the power distribution control scheme with the best overall performance, and then integrates and packages all control parameters within this scheme into standardized, transmittable global power distribution configuration data. This data is sent to the computing task scheduling platform, and historical scheduling data is used to update the preset digital twin power distribution model. The periodic scheduling path for transmitting power distribution configuration data across the entire domain and the path for issuing emergency power distribution adjustment commands in fault scenarios are logically separated. The physical links can respectively adopt 5G deterministic communication and industrial Ethernet, with their bandwidth not competing with each other. This prevents conventional large-capacity scheduling data packets from blocking emergency commands, ensures low-latency transmission of emergency adjustment commands under fault conditions, and improves the power supply safety of the computer room.

[0023] The power distribution execution layer receives local emergency power distribution adjustment commands and global power distribution configuration data, prioritizing the processing of local emergency power distribution adjustment commands. Based on the global power distribution configuration data, it executes operations such as switching power distribution circuits on and off, adjusting power output, and starting / stopping redundant power supplies. For example, the power distribution execution layer can be deployed at the control terminals of each PDU, UPS, energy storage converter, and redundant power supply module, using solid-state relays as the execution devices. Its action delay can be less than 1ms. It receives and executes local emergency power distribution adjustment commands and scheduling strategies, completing operations such as switching power distribution circuits on and off, continuously adjusting power output with 0.1kW accuracy, starting / stopping redundant power supplies and waking them from hibernation, and switching energy storage charging and discharging.

[0024] According to the embodiments of this application, the power efficiency intelligent scheduling and distribution system for high-computing-power data centers uniformly acquires power-side and computing-side operating parameters through a data acquisition layer. It simultaneously connects to the computing task scheduling platform to transmit predicted power consumption parameters and remaining power supply capacity of the racks, providing a unified data foundation for subsequent power distribution control. The edge collaborative pre-scheduling layer deployed in the head racks of each rack column quickly outputs local emergency power distribution adjustment commands based on real-time data collected from the corresponding rack column, eliminating the need to upload field data to the central platform for centralized processing. This effectively avoids the processing delay problem caused by centralized processing of all data in traditional centralized scheduling. The global energy efficiency scheduling layer deployed on the data center management platform generates full-domain power distribution configuration data based on a preset digital twin power distribution model that can be continuously updated through historical scheduling data, combined with collected data from the entire data center and peak / valley time parameters of the mains power. This compensates for the shortcomings of traditional distributed scheduling, which relies solely on fixed thresholds and cannot coordinate the power consumption conditions of the entire center. Local emergency power distribution adjustment commands and global power distribution configuration data are sent to the power distribution execution layer through independent channels that do not interfere with each other. The power distribution execution layer prioritizes local emergency power distribution adjustment operations, and then performs power circuit switching, power output adjustment, and redundant power supply start-up and shutdown operations based on the global power distribution configuration data. The functional levels cooperate with each other, and local emergency response does not need to wait for the global calculation process. Global power distribution optimization adjusts the power distribution strategy based on complete and real-time power-side operating parameters and computing power-side operating parameters. The two types of control processes do not crowd out transmission and processing resources. Local cabinet power supply failures only trigger the corresponding local control and will not affect the operation of the entire data center scheduling system, which greatly improves the overall fault tolerance capability. It can quickly intervene and adjust when the power supply is abnormal in the computer room to ensure the continuous and stable operation of high-density computing power services. It can also balance the power load of each cabinet according to the peak and valley patterns of the mains power, fully release the overall power distribution energy-saving space, and simultaneously improve the stability of power supply operation and the overall power economy.

[0025] In one implementation, the data acquisition layer includes power sensors, power consumption probes, and computing load acquisition units. Multiple power sensors are arranged one-to-one at the mains power input, UPS output, and each PDU input. Each CPU and GPU computing node is individually configured with a power consumption probe and a computing load acquisition unit. For example, the power sensors can be Hall effect voltage / current sensors with a sampling error of 0.38%, meeting the accuracy requirement of within 0.5%. The power consumption probes can be deployed at the BMC IPMI interface of each computing node; the computing load acquisition units can be deployed in the Linux operating system kernel's sched module. The resource utilization of the power consumption probes and computing load acquisition units is less than 1%, avoiding impact on normal computing tasks.

[0026] This implementation method deploys power sensors at the mains power inlet, UPS output, and each PDU input in a layered manner, and independently configures power consumption acquisition probes and computing load acquisition units for each CPU and GPU power node. On the one hand, it can completely cover the entire power supply chain of the data center, measure power loss at each level in a layered manner, and achieve accurate energy consumption traceability. On the other hand, it can collect power consumption and computing load data synchronously for a single node, support accurate fitting of node power consumption benchmark curves, and accurately locate abnormal power consumption of a single machine. At the same time, it forms data cross-verification based on the total power consumption of the upper layer power supply and the aggregated power consumption of the lower layer nodes, which effectively improves the accuracy of the collected data and reduces data errors.

[0027] In one implementation, each edge processing unit in the edge collaborative pre-scheduling layer pre-stores the power consumption baseline curves of all computing nodes in the corresponding rack column. The edge processing unit compares the local computing load fluctuation with a preset fluctuation threshold. When the local computing load fluctuation exceeds the preset fluctuation threshold, it adjusts the output power of redundant power supplies and the power supply priority of non-core loads within the local rack column. The generated local emergency power distribution adjustment command is directly sent to the power distribution execution layer for execution. For example, the power consumption baseline curve can be obtained by the edge processing unit during the 72-hour burn-in phase before the rack column goes online. During the burn-in phase, four operating conditions—no load, half load, full load, and turbo boost—are applied to each computing node in the rack column. Each operating condition runs continuously for 18 hours, and real-time power consumption values ​​are collected every second. Based on the full collection of data, the power consumption baseline curve of each node is obtained. The curve can characterize the correspondence between the node's computing load rate and operating power consumption. Each computing node has independent load-power consumption correspondence characteristics due to hardware differences.

[0028] The preset fluctuation threshold can be set to 15% of the total rated power consumption of all computing nodes in the local rack. This threshold is set based on the statistical rule that the instantaneous power consumption when a large model training task starts is usually 10%-12% of the rated power consumption, with a 3% safety margin reserved to avoid false alarms triggered by instantaneous power consumption fluctuations.

[0029] When the local computing load fluctuation is detected to exceed the preset fluctuation threshold, the redundant power output power and the power supply priority of non-core loads are adjusted within the local cabinet. Non-core loads are divided into three levels according to priority: Level 1 is AI training and online inference core loads, Level 2 is cold storage nodes and background operation and maintenance task nodes, and Level 3 is non-real-time offline computing task nodes. The power supply priority is sorted from high to low. During scheduling, the power supply power of Level 3 loads is reduced or even temporarily suspended.

[0030] The generated local emergency power distribution adjustment command is sent directly to the power distribution execution layer for execution without waiting for the command response from the global energy efficiency scheduling layer. When there is insufficient power redundancy in the local cabinet column, a cross-column collaborative scheduling request can be initiated to the edge processing unit of the adjacent cabinet column through the 10 Gigabit fiber optic link between the column head cabinets. The cross-column scheduling response delay does not exceed 48ms. After the pre-scheduling command is executed, the execution result, the adjusted power value, and the local remaining power supply capacity parameters are automatically sent back to the global energy efficiency scheduling layer.

[0031] In this embodiment, each edge processing unit in the edge collaborative pre-scheduling layer pre-stores the power consumption baseline curves of all computing nodes in the corresponding rack column. The edge processing unit compares the local computing load fluctuation with the preset fluctuation threshold locally. When the load fluctuation exceeds the threshold, it can autonomously adjust the output power of redundant power supplies and adjust the power supply priority of non-core loads within the local rack column. It can directly generate local emergency power distribution adjustment commands and send them to the power distribution execution layer for execution without the need for relaying to the upper-level platform. This significantly shortens the power supply emergency response delay, quickly smooths out instantaneous power consumption shocks, avoids rack power supply overload, and reduces cross-level communication interaction overhead, thereby improving the stability of rack column power supply operation.

[0032] In one implementation, the global energy efficiency scheduling layer is configured with a preset digital twin power distribution model. The preset digital twin power distribution model uses historical energy efficiency data of the data center, real-time collected power-side and computing-side parameters, and power distribution equipment health parameters as the data source. The preset digital twin power distribution model performs simulation calculations on the generated multiple sets of power distribution configuration parameters, outputs the corresponding PUE value and power supply reliability parameters of the entire data center, and outputs a scheduling strategy that matches the power distribution configuration data of the entire domain.

[0033] For example, the preset digital twin power distribution model of the global energy efficiency scheduling layer can be constructed by integrating historical energy efficiency data from the data center over the past three years, real-time collected power-side and computing-side parameters, and power distribution equipment health parameters. Historical energy efficiency data can be stored in the InfluxDB time-series database, aggregated hourly, including PUE data for different seasons, load rates, and periods with different grid electricity prices, as well as power line loss data. Power distribution equipment health parameters include the UPS's cumulative runtime, battery internal resistance, power module fan speed, and capacitor aging degree. Battery internal resistance is collected every 24 hours using a built-in internal resistance meter, and capacitor aging degree is indirectly calculated using the output ripple change rate.

[0034] The pre-defined digital twin power distribution model can simulate the PUE value, power supply reliability parameters, and total electricity cost parameters of the entire data center under at least 12 candidate scheduling strategies. Power supply reliability parameters include the continuous power supply level of core computing power, the probability of overload failure in the power distribution circuit, the success rate of migrating faulty computing power, and the duration of emergency power supply from energy storage, used to determine the power supply stability capability of each candidate scheduling strategy. Specifically, the continuous power supply level of core computing power is the proportion of the total runtime during which core AI, inference, and other critical computing power services operate without power outages throughout the year; the probability of overload failure in the power distribution circuit is the likelihood of the power distribution line load exceeding its rated capacity and triggering a power outage, calculated through simulation; the success rate of migrating faulty computing power is the percentage of successful transfers of affected computing power tasks to normal redundant power supply cabinets and their stable operation when power supply risks occur; and the duration of emergency power supply from energy storage is the duration for which energy storage devices can support the normal operation of core computing power after a mains power failure. The constraints of the simulation process are that the continuous power supply level of the core computing power is not less than 99.999%, the power supply harmonic content is not more than 5%, and the power supply capacity redundancy is not less than 10%. Candidate strategies that do not meet the constraints are directly eliminated, and the optimal scheduling strategy is selected from the remaining candidate strategies.

[0035] In one embodiment, the global energy efficiency scheduling layer performs time-sharing power distribution control based on the scheduling strategy output by the preset digital twin power distribution model. During the off-peak period of the mains power, the global energy efficiency scheduling layer controls the redundant power supply to operate at full load while charging the supporting energy storage unit. During the peak period of the mains power, the energy storage unit is called to supply power to the outside, while the power supply circuit of the redundant power supply module in the no-load state is shut down.

[0036] For example, during off-peak hours, the global energy efficiency dispatch layer controls the redundant power supply to maintain its load rate within the highest energy efficiency range of 80% to 90%. This range is set according to the general energy efficiency curve of the switching power supply, and the conversion efficiency of the switching power supply in this load range can reach over 96%, which is the highest energy efficiency range. At the same time, the energy storage converter is controlled to charge the matching energy storage unit to a state of charge (SOC) of no less than 90% at its rated power to avoid overcharging the energy storage and prolonging the battery cycle life. Off-peak hours are the periods when the local grid's published electricity price is more than 50% lower than the benchmark electricity price. Every 24 hours, the electricity price data for the next 48 hours is synchronized through the grid's publicly available electricity price API.

[0037] During peak mains power periods, energy storage units are prioritized for power supply, with higher priority than mains power. Simultaneously, redundant power modules with a load rate below 20% are shut down. In this load range, the switching power supply's conversion efficiency is typically below 85%, and shutting them down avoids ineffective energy loss. The load rate of the remaining operating power supplies is adjusted to the highest efficiency range of 80%-90%. When the energy storage unit's state of charge (SOC) falls below 20%, it automatically switches back to mains power to prevent over-discharge of the energy storage. Peak mains power periods are defined as times when the local grid's published electricity price is more than 200% higher than the benchmark price. Upon receiving a grid peak-shaving demand response signal, energy storage power is prioritized, while power supply to non-core loads is temporarily suspended.

[0038] In one implementation, the power distribution execution layer is equipped with a multi-redundant power supply switching module. This module is used for online adjustment of redundant power supplies. When the scheduling strategy determines that the redundancy of the redundant power supply exceeds a set threshold, the multi-redundant power supply switching module controls the excess redundant power supply to enter a low-power sleep state. When the collected load data increases, the multi-redundant power supply switching module performs a redundant power supply wake-up operation. For example, the multi-redundant power supply switching module can use an IGBT contactless solid-state switch with an on-resistance of less than 1mΩ, a switching delay of less than 8ms, and no power interruption during the switching process, thus preventing power loss and restart of the computing node.

[0039] Redundancy = (Total capacity of redundant power supply − Total power of real-time load) ÷ Total power of real-time load. When the redundancy is greater than 30%, it is determined that the redundancy exceeds the set threshold. The excess redundant power supply is controlled to enter a low-power sleep state. In the sleep state, the PFC circuit and power output stage of the power supply are turned off, and only the standby control circuit is powered. The power module power consumption in the sleep state is only 4.2% of that in the normal operation mode.

[0040] When a load increase is detected, the redundant power supply is woken up within 12ms. During the wake-up process, the power supply output voltage is established smoothly and will not affect the power supply stability of the existing load. After each switch, the output parameters of the power module are automatically checked. The check includes whether the output voltage error is within ±1% and whether the output ripple is less than 3%. Power modules with abnormalities are marked as faults and are no longer included in the scheduling scope. An operation and maintenance work order is automatically generated and pushed to the operation and maintenance platform.

[0041] In one implementation, the edge collaborative pre-scheduling layer is also equipped with a power supply fault prediction unit. The power supply fault prediction unit is used to predict whether there is a risk of power supply fault based on the ripple value of the collected power parameters and the trend of equipment temperature change. When a risk of power supply fault is predicted, the corresponding load is migrated to the redundant power supply circuit of the adjacent rack row.

[0042] For example, the input features of the algorithm can include four dimensions: voltage ripple value, equipment temperature change rate, current harmonic content, and cumulative power supply runtime. The training samples consist of 1200 fault samples and 10000 normal samples collected over the past two years, divided into training and testing sets in an 8:2 ratio. After training, the model is deployed locally on the edge processing unit without uploading to the cloud. The prediction trigger condition is: the ripple value of the power parameter exceeds 3% of the AC component of the power supply output DC voltage, or the equipment temperature changes by more than 5 degrees Celsius per minute. If either condition is met, it is determined that there is a risk of power supply failure. In advance, a load migration instruction is sent to the computing power task scheduling platform to migrate the core computing power tasks of the corresponding rack column to the idle computing power nodes of the adjacent rack column within 1 minute. At the same time, the load of the original power supply circuit is switched to the redundant power supply circuit of the adjacent column. The redundant power supply circuit is kept in an idle standby state under normal conditions and can be connected to the load at any time to avoid business interruption in the event of a failure.

[0043] In one implementation, the global energy efficiency scheduling layer is further configured with an energy efficiency audit module. This module is used to statistically analyze the PUE value, power supply reliability parameters, and electricity costs of the entire data center after each scheduling operation. The statistical results are then used as training samples to update the parameters of the preset digital twin power distribution model. For example, the statistical dimensions may include the load rate change, power loss, and cumulative runtime of each power module; the PUE value and unit computing power electricity cost of each rack column; and the overall energy efficiency indicators and SLA compliance rate of the entire data center. The statistical results are used as training samples to update the parameters of the preset digital twin power distribution model, with the sample update frequency being once a day.

[0044] Therefore, the performance audit module can statistically analyze the overall PUE, power supply reliability, and electricity cost of the data center after each scheduling, and send the statistical samples back to update the preset digital twin power distribution model parameters to continuously improve the accuracy of multi-strategy simulation of the model. At the same time, it can accumulate all scheduling operation data to support the output of monthly energy efficiency audit reports and assist operation and maintenance personnel in carrying out energy efficiency optimization and electricity cost control.

[0045] In one implementation, the global energy efficiency scheduling layer is communicatively connected to the data center's computing task scheduling platform. The global energy efficiency scheduling layer synchronously outputs scheduling policies to the computing task scheduling platform, driving computing tasks to be allocated to rack rows whose power supply redundancy and energy efficiency indicators meet the standards as determined by a preset digital twin power distribution model. For example, the global energy efficiency scheduling layer can communicate with the data center's computing task scheduling platform via an encrypted TCP communication interface, synchronizing the scheduling policies to the platform at 5-minute intervals. The synchronized scheduling policies include the current power supply redundancy capacity, current power load rate, energy efficiency ratio, and maximum computing load power consumption parameters that each rack row can handle. When sending new computing tasks, the computing task scheduling platform prioritizes allocating tasks to rack rows with sufficient power supply redundancy and power load rates in the highest energy efficiency range. When the power supply redundancy of a rack row falls below 10%, an alarm is issued to the computing task scheduling platform, prohibiting new computing tasks from being sent to that rack row. This achieves bidirectional coordination between power supply scheduling and computing task scheduling, avoiding power overload.

[0046] The second aspect of this invention discloses a power efficiency intelligent scheduling method for high-computing-power data centers, applied to the aforementioned power efficiency intelligent scheduling power distribution system for high-computing-power data centers, comprising the following steps: Step 1: Synchronously collect the power-side operating parameters, computing-side operating parameters, and predicted power consumption parameters of the computing tasks to be sent from the high-performance computing data center through the data acquisition layer. Step 2: The edge collaborative pre-scheduling layer judges the sudden computing load fluctuation of the local rack column. When the load fluctuation exceeds the preset fluctuation threshold, a local emergency power distribution adjustment command is generated and sent to the power distribution execution layer. At the same time, the load fluctuation data is synchronized to the global energy efficiency scheduling layer. Step 3: Based on the preset digital twin power distribution model, the global energy efficiency scheduling layer generates an hourly scheduling strategy by combining the collected data from the entire data center and the peak and valley time parameters of the mains power, and sends it to the power distribution execution layer for execution. Step four: Execute the corresponding scheduling instructions through the power distribution execution layer, and perform power distribution circuit switching, power output adjustment, and redundant power supply start-up and shutdown operations based on the global power distribution configuration data.

[0047] The intelligent power efficiency scheduling method for high-computing-power data centers according to the embodiments of this application realizes hierarchical power distribution management based on a layered collaborative architecture. The data acquisition layer synchronously collects multi-dimensional parameters such as power, computing power, and task prediction to ensure the integrity of the underlying data. The edge collaborative pre-scheduling layer handles instantaneous load impacts on server racks locally, quickly issues emergency power distribution commands, reduces cross-level communication latency, and avoids power supply overload. The global energy efficiency scheduling layer uses a digital twin model combined with the peak and valley patterns of the mains power to generate long-term optimized scheduling strategies, and coordinates the energy efficiency and electricity costs of the entire center. The power distribution execution layer uniformly executes various power distribution adjustment actions, realizing the combination of local emergency response and global long-term energy efficiency optimization, taking into account both the stability of power supply operation and the overall power utilization efficiency of the data center.

[0048] Reference Figure 1 This diagram illustrates the overall closed-loop operation of a power efficiency intelligent scheduling system for high-performance data centers. The data acquisition layer synchronously collects all operating parameters from both the power and computing sides; the edge collaborative pre-scheduling layer monitors local computing load fluctuations in real time, and when fluctuations exceed preset thresholds, immediately generates local emergency power distribution adjustment commands and synchronizes the fluctuation data to the global energy efficiency scheduling layer; the global energy efficiency scheduling layer generates global power distribution configuration data based on a preset digital twin power distribution model; the power distribution execution layer simultaneously receives both types of scheduling commands, prioritizing local emergency adjustment commands, and then executes operations such as power circuit on / off, power output adjustment, and redundant power supply start / stop based on the global power distribution configuration data.

[0049] Reference Figure 2 This diagram details the specific process of emergency control and fault migration performed by the edge collaborative pre-scheduling layer within the local rack column. The edge collaborative pre-scheduling layer first acquires the collected data and pre-stored power consumption baseline curves for the corresponding rack column. On one hand, the edge processing unit compares the local computing power load fluctuation with a preset fluctuation threshold. If the fluctuation exceeds the threshold, it adjusts the output power of redundant power supplies and the power supply priority of non-core loads within the local rack column, directly generating a local emergency power distribution adjustment command and sending it to the power distribution execution layer for execution. On the other hand, the power supply fault prediction unit analyzes the power supply fault risk based on power parameter ripple values ​​and equipment temperature change trends. If a risk is predicted, the corresponding load is migrated to the redundant power supply circuit of an adjacent rack column.

[0050] Reference Figure 3This diagram illustrates the execution steps of the global energy efficiency scheduling layer in optimizing resource scheduling across the entire domain based on a preset digital twin power distribution model. The global energy efficiency scheduling layer collects historical energy efficiency data, real-time operating parameters, and power distribution equipment health parameters. It then uses the preset digital twin power distribution model to perform simulation calculations on multiple sets of candidate power distribution configuration parameters, outputting the PUE value for the entire data center, power supply reliability parameters, and matching scheduling strategies. The system performs time-sharing control based on peak and off-peak power parameters. During off-peak periods, redundant power supplies are controlled to operate efficiently and charge energy storage units. During peak periods, energy storage is used to supply power, and idle, inefficient redundant power supply modules are shut down. The efficiency audit module statistically analyzes energy efficiency, reliability, and electricity cost data after each scheduling operation, using the results as training samples to update the digital twin model parameters. Simultaneously, the global scheduling layer synchronizes the scheduling strategy to the computing task scheduling platform, driving the allocation of computing tasks to racks that meet both power supply redundancy and energy efficiency standards.

[0051] Reference Figure 4 This flowchart fully presents the standardized operating method of the intelligent dispatching system from data acquisition to strategy execution. First, power-side operating parameters and computing power and predicted power consumption parameters are collected synchronously. Second, the edge collaborative pre-scheduling layer judges the sudden computing power load fluctuations of the local rack column. When the fluctuation exceeds the preset fluctuation threshold, a local emergency power distribution adjustment command is immediately generated and sent to the power distribution execution layer. At the same time, the load fluctuation data is synchronized to the global energy efficiency dispatching layer. The global energy efficiency dispatching layer receives the real-time collected data and generates hourly dispatching strategies based on the preset digital twin power distribution model and global data. Finally, the power distribution execution layer executes the dispatching commands and the on / off of power distribution circuits and the start / stop of power supplies.

[0052] The above are merely preferred embodiments of the present invention, but the scope of protection of the present invention is not limited thereto. Any equivalent substitutions or modifications made by those skilled in the art within the scope of the technology disclosed in the present invention, based on the technical solution and inventive concept of the present invention, should be covered within the scope of protection of the present invention.

Claims

1. A power efficiency intelligent scheduling and distribution system for high-computing-power data centers, characterized in that, It includes a data acquisition layer, an edge collaborative pre-scheduling layer, a global energy efficiency scheduling layer, and a power distribution execution layer; The data acquisition layer is used to synchronously collect the power-side operating parameters and computing-side operating parameters of the high-performance computing data center. The power-side operating parameters include the mains power line electrical parameters, UPS output operating parameters, rack power distribution circuit parameters, and energy storage unit power parameters. The computing-side operating parameters include the real-time computing load, chip operating status, and power consumption parameters of each computing node. The data acquisition layer receives the predicted power consumption parameters of the computing tasks generated by the computing task scheduling platform and sends the remaining power supply capacity of the rack to the computing task scheduling platform. The edge collaborative pre-scheduling layer is deployed in the head cabinet of each rack column to receive the collected data of the corresponding rack column, collect local computing load data in real time, and generate local emergency power distribution adjustment commands. The global energy efficiency scheduling layer is deployed on the data center management platform. It is used to generate multiple sets of power distribution configuration parameters based on a preset digital twin power distribution model that maps the operating status of the computer room power distribution hardware and computing load, combined with the collected data of the entire data center and the peak and valley time parameters of the mains power. It outputs a set of global power distribution configuration data every hour. The global energy efficiency scheduling layer sends the global power distribution configuration data to the computing task scheduling platform and updates the preset digital twin power distribution model with historical scheduling data. The path used to transmit the global power distribution configuration data is independent of the path used to send local emergency power distribution adjustment commands. The power distribution execution layer is used to receive local emergency power distribution adjustment commands and global power distribution configuration data, and prioritizes processing local emergency power distribution adjustment commands. Based on the global power distribution configuration data, it performs power distribution circuit switching, power output power adjustment, and redundant power supply start-up and shutdown operations.

2. The intelligent power dispatching and distribution system for high-performance data centers according to claim 1, characterized in that, The data acquisition layer includes power sensors, power consumption acquisition probes, and computing load acquisition units. Multiple power sensors are arranged one-to-one at the mains power input end, UPS output end, and each PDU input end. Each CPU computing node and each GPU computing node is individually configured with a power consumption acquisition probe and a computing load acquisition unit.

3. The intelligent power dispatching and distribution system for high-performance data centers according to claim 1, characterized in that, Each edge processing unit in the edge collaborative pre-scheduling layer pre-stores the power consumption baseline curves of all computing nodes in the corresponding rack column. The edge processing unit compares the local computing load fluctuation with the preset fluctuation threshold. When the local computing load fluctuation exceeds the preset fluctuation threshold, it adjusts the output power of redundant power supplies and the power supply priority of non-core loads within the range of the local rack column. The generated local emergency power distribution adjustment command is directly sent to the power distribution execution layer for execution.

4. The intelligent power dispatching and distribution system for high-performance data centers according to claim 1, characterized in that, The global energy efficiency scheduling layer is configured with a preset digital twin power distribution model. The preset digital twin power distribution model uses historical energy efficiency data of the data center, real-time collected power and computing side parameters, and power distribution equipment health parameters as the data source. The preset digital twin power distribution model performs simulation calculations on the generated multiple sets of power distribution configuration parameters, outputs the corresponding PUE value and power supply reliability parameters of the entire data center, and outputs a scheduling strategy that matches the power distribution configuration data of the entire domain.

5. The intelligent power dispatching and distribution system for high-performance data centers according to claim 4, characterized in that, The global energy efficiency scheduling layer performs time-sharing power distribution control based on the scheduling strategy output by the preset digital twin power distribution model. During the off-peak period of the mains power, the global energy efficiency scheduling layer controls the redundant power supply to operate at full load while charging the supporting energy storage unit. During the peak period of the mains power, the energy storage unit is called to supply power to the outside, while the power supply circuit of the redundant power supply module in the no-load state is shut down.

6. The intelligent power dispatching and distribution system for high-performance data centers according to claim 4, characterized in that, The power distribution execution layer is equipped with a multi-redundant power supply switching module. The multi-redundant power supply switching module is used for online adjustment of redundant power supplies. When the scheduling strategy determines that the redundancy of the redundant power supply exceeds the set threshold, the multi-redundant power supply switching module controls the excess redundant power supply to enter a low-power sleep state. When the collected load data increases, the multi-redundant power supply switching module performs a redundant power supply wake-up operation.

7. The intelligent power dispatching and distribution system for high-performance data centers according to claim 3, characterized in that, The edge collaborative pre-scheduling layer is also equipped with a power supply fault prediction unit. The power supply fault prediction unit is used to predict whether there is a risk of power supply fault based on the ripple value of the collected power parameters and the trend of equipment temperature change. When a risk of power supply fault is predicted, the corresponding load is migrated to the redundant power supply circuit of the adjacent rack column.

8. The intelligent power dispatching and distribution system for high-performance data centers according to claim 4, characterized in that, The global energy efficiency scheduling layer is also configured with an efficiency audit module. The efficiency audit module is used to statistically analyze the PUE value, power supply reliability parameters, and electricity cost of the entire data center after each scheduling, and use the statistical results as training samples to update the parameters of the preset digital twin power distribution model.

9. The intelligent power dispatching and distribution system for high-performance data centers according to claim 4, characterized in that, The global energy efficiency scheduling layer is communicatively connected to the computing power task scheduling platform of the data center. The global energy efficiency scheduling layer synchronously outputs scheduling strategies to the computing power task scheduling platform, driving computing power tasks to be allocated to racks that meet the power supply redundancy and energy efficiency standards as determined by the preset digital twin power distribution model.

10. A power efficiency intelligent scheduling method for high-performance data centers, applied to the power efficiency intelligent scheduling and distribution system for high-performance data centers as described in any one of claims 1 to 9, characterized in that, Includes the following steps: Step 1: Synchronously collect the power-side operating parameters, computing-side operating parameters, and predicted power consumption parameters of the computing tasks to be sent from the high-performance computing data center through the data acquisition layer. Step 2: The edge collaborative pre-scheduling layer judges the sudden computing load fluctuation of the local rack column. When the load fluctuation exceeds the preset fluctuation threshold, a local emergency power distribution adjustment command is generated and sent to the power distribution execution layer. At the same time, the load fluctuation data is synchronized to the global energy efficiency scheduling layer. Step 3: Based on the preset digital twin power distribution model, the global energy efficiency scheduling layer generates an hourly scheduling strategy by combining the collected data from the entire data center and the peak and valley time parameters of the mains power, and sends it to the power distribution execution layer for execution. Step four: Execute the corresponding scheduling instructions through the power distribution execution layer, and perform power distribution circuit switching, power output adjustment, and redundant power supply start-up and shutdown operations based on the global power distribution configuration data.