Multi-loop power adaptive distribution system for edge data center

By constructing an electricity-heat-IT digital model and collaborative scheduling, the risks of circuit breaker tripping and hotspots in the edge data center power adaptive distribution system during rapid power changes were resolved, achieving efficient collaborative management of power resources and fault location.

CN121906459APending Publication Date: 2026-04-21NANTONG TAILIAN DATA TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
NANTONG TAILIAN DATA TECHNOLOGY CO LTD
Filing Date
2025-12-04
Publication Date
2026-04-21

AI Technical Summary

Technical Problem

When faced with rapid power changes in servers, the adaptive power distribution system in edge data centers struggles to respond to power switching within milliseconds, leading to circuit breaker tripping and hotspot risks, and making fault location difficult.

Method used

By employing a data acquisition module, an IT load intent characterization module, a server power prediction module, an initial server power allocation module, and a power command issuance module, an electricity-heat-IT digital model is constructed to perform coordinated scheduling of power resources and cooling capacity, thereby achieving prediction and dynamic adjustment of power resources.

Benefits of technology

It can effectively predict server power surges, reduce circuit breaker tripping, record the causal chain of power distribution decisions, avoid hotspot risks, and improve resource utilization and system stability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121906459A_ABST
    Figure CN121906459A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-loop power adaptive distribution system for an edge data center, and relates to the technical field of power distribution, and the system comprises the steps: building a load portrait library of a server after collecting data, and predicting a power curve and peak power in a future time t when a scheduling intention is captured; based on the historical data, outputting predicted power of future time t, and allocating an initial resource weight and a power budget of each server; constructing an electricity-heat-IT digital model, carrying a resource weight redistribution strategy, and redistributing resource weights and power resources of all servers according to the electricity-heat-IT digital model; and after the electric power resources of the server with the changed electric power resources in the resource weight redistribution strategy are returned to other servers according to the resource tree model, the server resource updating module is operated again for the cabinet with the changed electric power resources. According to the method, the problems that the instantaneous fault of the machine room is difficult to check and local hot spots and thermal runaway risks exist are solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of power distribution technology, specifically a multi-loop adaptive power distribution system for edge data centers. Background Technology

[0002] With the large-scale deployment of edge computing and services such as AI inference training, video processing, and low-latency databases closer to the user, edge data centers are gradually evolving from simple server rooms into highly dense clusters of computing nodes. Compared to traditional centralized server rooms, edge server rooms are typically constrained by stricter power supply space constraints, more multi-circuit power distribution topologies, and higher wake-up concurrency: the probability of a large number of tasks starting up at the near end in a short period of time increases significantly. Currently, adaptive power distribution systems for server rooms face the following challenges: Server CPUs and GPUs can switch from a low-power idle state to a full-load active state in microseconds, causing a sudden surge in power consumption; in contrast, even the fastest static switching switches in power systems have switching times on the order of milliseconds. This order-of-magnitude speed difference means that when a large number of servers in a rack are simultaneously awakened, by the time the adaptive system detects the load change and makes a switching decision, overload may have already occurred, causing partially idle circuit breakers to trip.

[0003] However, if power control performs current limiting or switching on a millisecond scale, the server PSU or power management module may be restarted or reset due to short-term voltage / current disturbances. Such events often don't leave clear evidence in routine power or IT monitoring, creating a gray area: maintenance personnel cannot determine whether the problem stems from the equipment itself, the power supply system's actions, or insufficient cooling triggering thermal protection. This makes attribution of responsibility, recovery decisions, and subsequent improvements extremely difficult.

[0004] To improve overall resource utilization or avoid loop overload, adaptive allocation algorithms may perform clustered allocation of power to several cabinets or rack units at the power level. However, power concentration in the thermal environment usually manifests immediately as localized temperature rise. If the data center's cooling system cannot provide synchronized and precise on-demand airflow, thermal inertia and air conditioning response delays will lead to localized hotspots, equipment frequency reduction, or even system downtime. In other words, simple power optimization without a linked thermal model may sacrifice temperature safety for short-term power balance, which carries high risk and serious consequences.

[0005] Therefore, the present invention provides a multi-loop adaptive power distribution system for edge data centers. Summary of the Invention

[0006] The purpose of this invention is to provide a multi-loop adaptive power distribution system for edge data centers to solve the existing problems mentioned in the background art.

[0007] To achieve the above objectives, the present invention provides the following technical solution: a multi-loop adaptive power distribution system for edge data centers, comprising: The data acquisition module is used to deploy a data acquisition network in the computer room to collect data in batches, including the readings of the air inlet temperature sensor built into each server, power data, and IT load data. The IT load intent profiling module builds a server load profile library. When the scheduling intent is captured, it predicts the power curve and peak power within a future time t. The server power prediction module is used to take the load scheduling intention of each rack as input based on historical data and output the predicted power at a future time t. The initial server power allocation module is used to allocate initial resource weights and power budgets to each server based on predicted power. The server resource update module is used to build an electricity-heat-IT digital model and carry a resource weight redistribution strategy. Based on the electricity-heat-IT digital model, it redistributes the resource weights and power resources of all servers. The power command issuance module is used to transfer the power resources of servers whose power resources have changed in the resource weight redistribution strategy back to other servers according to the resource tree model, and then run the server resource update module again for the racks whose power resources have changed, and send air conditioning air volume commands according to temperature changes.

[0008] A further improvement of the present invention is that the IT load intent characterization module includes: An offline-built server load profile library is constructed by extracting features from the power time series of historical tasks from startup to steady state and obtaining typical profile clusters through a clustering algorithm. Each profile cluster is saved with a parameterized piecewise function. An online two-stage intent matching unit is used to quickly match the input scheduling intent with a profile database using cosine similarity to obtain an initial template. Then, XGBoost is used as a fine-tuner, taking the parameters of the initial template and the current data center context as input, to output a task-level power prediction curve. and uncertainty .

[0009] A further improvement of this invention is that the server power prediction module adopts a hybrid structure, and the specific steps include: Extracting the predicted power of multiple tasks assigned to the same server , k represents the k-th task, calculated according to the resource requisition coefficient. Perform a weighted summation to obtain the baseline power of the i-th server: ; Using the server's historical power sequence, ambient temperature, and IT load data from the most recent M seconds as input, an LSTM network is used to predict residual terms. Ultimately, server-level power prediction was obtained. ; Subsequently, a standard deviation estimate is output based on model integration. And calculate conservative estimates , This represents the safety factor.

[0010] A further improvement of this invention is that the initial server power allocation module formalizes the weight allocation problem into a constrained convex optimization problem and uses a fast solver for rolling solution, wherein the optimization problem includes: Introducing the variable default rate s, with the objective of minimizing the maximum default rate, the conservative estimate is... As a safety power limit, for any server ,satisfy For any loop ,satisfy And satisfy the total power constraint Simultaneously set constraint rates The optimized output is then subjected to linear ramp smoothing; where... Indicates assignment to the first The resource weight of the i-th server is used to determine the current system status as the i-th server. Power budget allocated to each server .

[0011] A further improvement of this invention is that the electro-thermal-IT digital model includes establishing an equivalent heat capacity for each server i. Equivalent thermal resistance Wind ratio coefficient Discretized temperature dynamic equations are used for time step calculation. The forward simulation yields the future time window. Temperature of internal server i ; like and The difference is less than or equal to the set server temperature threshold. At that time, the system will directly execute the current number of times. Power budget allocated to each server ; like and The difference is greater than the set server temperature threshold. At that time, the resource weight redistribution strategy is implemented.

[0012] A further improvement of this invention is that the time step is performed using a discretized temperature dynamic equation. Forward simulation includes: ; in, This represents the power-to-heat conversion coefficient of the i-th server. The system's heat dissipation capability for server i is represented as... This indicates the heat dissipation contribution of the air conditioner to the server. This indicates the air supply temperature of server i. This indicates the return air temperature of server i; This indicates the airflow allocated to server i. Indicates the specific heat of air. This indicates the ambient temperature of the computer room.

[0013] A further improvement of this invention is that the resource weight redistribution strategy is implemented by calculating the temperature excess of server i, wherein the temperature excess of server i is represented as... It is equipped with hysteresis control constraints. When the hysteresis control constraints are met, the decay factor is calculated using an exponential decay function. , where the parameters and This represents the decay parameter, which is used to obtain the updated resource weights for server i. and update the number Power budget allocated to each server .

[0014] A further improvement of this invention is that the hysteresis control constraint is used to introduce hysteresis and minimum hold time. The constraints include Three consecutive values ​​greater than 0 with a duration greater than .

[0015] A further improvement of this invention is that the power command issuance module includes a power release extraction unit and a resource tree model construction unit; the power release extraction unit is used to extract resources released from the overheating node and calculate the power release. ; The resource tree model construction unit is used to allocate the released power to the candidate server set A according to the initial resource weight ratio, wherein the servers in the candidate server set A satisfy the future time window. The difference between the internal server temperature and the current time t is less than or equal to the set server temperature threshold. First, the initial resource weight of the j-th server in server set A is proportional to the sum of the initial resource weights of all servers in server set A. The resource allocation ratio of the j-th server is obtained, and the initial allocation is set based on the resource allocation ratio. .

[0016] A further improvement of the present invention is that the data acquisition module specifically includes N distributed edge acquisition nodes, each of which includes: a power measurement unit, a temperature sensor interface, a management agent interface, and a communication unit; The power measurement unit is used to collect the bus current of the cabinet. ,Voltage The system samples locally at a high frequency, calculates the RMS, average value, and peak value within a fixed window, and then reports them. The temperature sensor interface is used to directly connect to or act as a proxy to read the temperature sensor readings of the server's built-in air inlet temperature sensor. The management agent interface is used to collect server-side IT metrics; The communication unit is used to perform edge preprocessing on the collected data and report it in batches to the edge controller via the TLS channel.

[0017] Compared with the prior art, the beneficial effects of the present invention are: 1. This invention first introduces IT load intention feedforward prediction, conservative power allocation model and local solid-state buffer device into the power distribution strategy, so as to realize that the power system has a safety margin in advance before the server's μs-level power surge; effectively make up for the time difference between server power fluctuation and the power switching device's ms-level response; and solve the problem of circuit breaker tripping and power interruption caused by response lag when large-scale tasks are started. 2. By constructing an integrated monitoring network for electricity, heat, and IT, and introducing transactional command issuance and log auditing mechanisms, a complete causal chain record of power allocation decisions, execution actions, and subsequent equipment status is achieved; power allocation and environmental status can only be reproduced in the event of an anomaly, and the source of the fault can be quickly distinguished as IT, power supply, or cooling; thus solving the problem of difficult location of instantaneous faults and unclear responsibility in traditional systems. 3. By establishing a digital twin model of electricity, heat, and IT, and introducing dynamic weight adjustment triggered by temperature thresholds in power allocation, the coordinated scheduling of power resources and rack cooling capacity is achieved. When the temperature of a local rack exceeds the threshold, the system automatically reduces its power weight and transfers power to racks with larger temperature margins through a resource tree backflow mechanism, while simultaneously adjusting the air conditioning airflow; effectively avoiding the risks of local hotspots and thermal runaway caused by simple power optimization. Attached Figure Description

[0018] Figure 1 This is a framework diagram of a multi-loop adaptive power distribution system for edge data centers according to the present invention. Detailed Implementation

[0019] The technical solution of the present invention will be described in detail below with reference to the accompanying drawings and specific embodiments. It should be understood that the embodiments of the present invention and the specific features in the embodiments are detailed descriptions of the technical solution of the present invention, rather than limitations thereof. In the absence of conflict, the embodiments of the present invention and the technical features in the embodiments can be combined with each other.

[0020] The term "and / or" is merely a description of the relationship between related objects, indicating that there can be three relationships. For example, A and / or B can represent three cases: A exists alone, A and B exist simultaneously, and B exists alone.

[0021] Example 1 Figure 1 This embodiment illustrates a multi-loop adaptive power distribution system for edge data centers, comprising: The data acquisition module is used to deploy a data acquisition network in the computer room to collect data in batches, including the readings of the air inlet temperature sensor built into each server, power data, and IT load data. The IT load intent profiling module builds a server load profile library. When the scheduling intent is captured, it predicts the power curve and peak power within a future time t. The server power prediction module is used to take the load scheduling intention of each rack as input based on historical data and output the predicted power at a future time t. The initial server power allocation module is used to allocate initial resource weights and power budgets to each server based on predicted power. The server resource update module is used to build an electricity-heat-IT digital model and carry a resource weight redistribution strategy. Based on the electricity-heat-IT digital model, it redistributes the resource weights and power resources of all servers. The power command issuance module is used to transfer the power resources of servers whose power resources have changed in the resource weight redistribution strategy back to other servers according to the resource tree model, and then run the server resource update module again for the racks whose power resources have changed, and send air conditioning air volume commands according to temperature changes.

[0022] The data acquisition module specifically includes N distributed edge acquisition nodes, each of which includes: a power measurement unit, a temperature sensor interface, a management agent interface, and a communication unit; The power measurement unit is used to collect the bus current of the cabinet. ,Voltage The RMS, average value, and peak value are calculated and reported within a 100 ms window after local sampling at a high frequency (e.g., 1 kHz). The temperature sensor interface is used to directly connect to or act as a proxy to read the temperature sensor readings of the server's built-in air inlet temperature sensor. The management agent interface is used to collect server-side IT metrics (including but not limited to CPU / GPU utilization, CPU / GPU frequency, number of active cores, process or container identifier, task ID, and scheduling intent flag). The communication unit is used to perform edge preprocessing on the collected data, including median filtering, short missing linear interpolation, and in-window anomaly detection and labeling, and reports the data in batches to the edge controller via the TLS channel with protobuf encoding and HMAC signature. High-resolution current / power / temperature / IT timing data is guaranteed through high-frequency (local 1kHz sampling, 100ms reporting) and PTP time synchronization. This capability is a prerequisite for identifying μs→ms level impulses (through peak / slope characteristics) and associating them with scheduling events.

[0023] Edge caching and local anomaly protection enable local protection to be activated even when the central control center is unreachable, triggering protection actions in milliseconds or faster, thus reducing the probability of circuit breaker tripping.

[0024] High-frequency data and HMAC / TLS audit links provide a chain of evidence for post-event diagnosis, reducing the gray area of ​​"transient failures".

[0025] The IT load intent profiling module includes: An offline-built server load profile library, which extracts features from the power time series of historical tasks from startup to steady state, including peak values. Rise time constant steady-state power Duration and the amplitude of shaking Typical profile clusters are obtained through the K-means clustering algorithm, and each profile cluster is saved using a parameterized piecewise function. An online two-stage intent matching unit is used to quickly match the input scheduling intent (including application type, requested resource vectors such as CPU / GPU quantity, batch size, model size, etc.) with a profile database using cosine similarity to obtain an initial template. Then, XGBoost is used as a fine-tuner, taking the parameters of the initial template and the current data center context (inlet air temperature, current load, historical similar task statistics) as input, and outputting a task-level power prediction curve. and uncertainty .

[0026] By using a profile library and two-stage matching, reliable task-level power curves and uncertainties can be quickly generated upon receiving scheduling intentions, and predictions can be sent to edge controllers in advance for feedforward preparation in the case of pre-notification.

[0027] Preparing buffers and rate limiting before the task is actually woken up is key to mitigating the difference between μs-level surges and ms-level switching; the predicted output has uncertainty, which can be used for conservative quotas to reduce instantaneous overload caused by overly optimistic predictions.

[0028] The server power prediction module adopts a hybrid structure, and the specific steps include: Extracting the predicted power of multiple tasks assigned to the same server , k represents the k-th task, calculated according to the resource requisition coefficient. Perform a weighted summation to obtain the baseline power of the i-th server: ; Using the server's historical power sequence, ambient temperature, and IT load data from the most recent M seconds as input, an LSTM network is used to predict residual terms. Ultimately, server-level power prediction was obtained. ; Subsequently, a standard deviation estimate is output based on model integration. And calculate conservative estimates , This represents the safety factor.

[0029] The resource requisition coefficient is obtained through historical fitting. When the error between the measured peak and the predicted peak exceeds a threshold (e.g., 20%), online fine-tuning or offline retraining is automatically triggered, and the model version and rollback information are recorded for auditing purposes.

[0030] The hybrid model (baseline overlay + LSTM residual) retains interpretability and improves accuracy in the short term. The output confidence is used for subsequent conservative constraints, and the prediction uncertainty is explicitly introduced into the allocation decision, reducing circuit breaker malfunction and thermal risk from the source. Online fine-tuning and rollback records provide model tracking evidence for ex-post analysis and repair of IT faults caused by electrical actions but not detected.

[0031] The initial server power allocation module formalizes the weight allocation problem into a constrained convex optimization problem and uses a fast solver for rolling solution. The optimization problem includes: Introducing the variable default rate *s*, the objective is to minimize the maximum default rate, i.e., to minimize the excess ratio of all servers relative to their rated power; formally, this is... ; The conservative estimate As a safety power limit, for any server ,satisfy For any loop ,satisfy And satisfy the total power constraint For servers with a hard SLA priority, an inviolable lower or upper bound is added to the optimization model to ensure hard constraints; a constraint rate is also set. The optimized output is then subjected to linear ramp smoothing; where... Indicates assignment to the first The resource weight of the i-th server is used to determine the current system status as the i-th server. Power budget allocated to each server .

[0032] The solution process includes constructing an LP matrix to standardize the linear constraints and convert them to standard form, calling the OSQP solver, and if the LP matrix has a solution, retrieving the solution. and If the LP matrix has no solution, first prioritize the low-priority tasks. Set the value to 0 or the minimum, and solve again. Perform rate smoothing on the solution, and then calculate the weights. .

[0033] The goal is to achieve a fair and conservative allocation by minimizing the maximum default rate s: when capacity is limited, priority is given to ensuring that no equipment is pushed to the brink of danger, and to reducing transient failures or overheating of any loop or server due to over-allocation.

[0034] Introducing rate constraints and slope smoothing allows for controllable slope changes in power allocation, preventing the allocation of large amounts of power to a single point at once. This mitigates hotspots caused by thermal coupling hysteresis and reduces the sudden changes experienced by circuit breakers. If a solution proves infeasible, the system automatically degrades to a lower-priority task, ensuring hard SLAs and system safety. This helps to classify system-induced IT failures as "controlled degradation" rather than random failures.

[0035] The electro-thermal-IT digital model includes establishing the equivalent heat capacity for each server i. Equivalent thermal resistance Wind ratio coefficient Discretized temperature dynamic equations are used for time step calculation. The forward simulation yields the future time window. Temperature of internal server i ; like and The difference is less than or equal to the set server temperature threshold. At that time, the system will directly execute the current number of times. Power budget allocated to each server ; like and The difference is greater than the set server temperature threshold. At that time, the resource weight redistribution strategy is implemented.

[0036] Mapping power fluctuations to temperature responses (including airflow, heat capacity, etc.) makes power allocation decisions thermally sensitive, allowing for the prediction of hotspots before allocation; if hotspots are generated, allocation decisions are prevented or mitigated, fundamentally avoiding the situation of "electricity optimization → thermal disaster".

[0037] Digital twins can perform forward simulations and generate reproducible event logs, providing evidence for post-event determination of the causal relationship between IT failures and power operations.

[0038] Using temperature as a constraint can help identify unsafe distributions in advance and prevent circuit breaker tripping.

[0039] The time step is calculated using a discretized temperature dynamic equation. Forward simulation includes: Physically, power is converted into heat, and then thermal resistance and cooling capacity are considered to obtain an engineering approximation of temperature dynamics; discretization facilitates real-time forward simulation; among which, This represents the power-to-heat conversion coefficient of the i-th server. The system's heat dissipation capability for server i is represented as... This indicates the heat dissipation contribution of the air conditioner to the server. This indicates the air supply temperature of server i. This indicates the return air temperature of server i; This indicates the airflow allocated to server i. Indicates the specific heat of air. This represents the ambient temperature of the server room. Thermal coupling between server racks is represented by a state-space expression. express; Physical discrete temperature equations and coupling matrices make simulations more accurate and can identify return / supply air interaction problems caused by power concentration (i.e. hotspot propagation paths). This allows for the addition of global thermal boundary conditions during resource allocation, avoiding global thermal degradation caused by local short-term optimization.

[0040] A refined thermal model can predict the rate of temperature rise, and combined with rate constraints, it can avoid instantaneous thermal overshoot caused by the mismatch between power slope and thermal inertia.

[0041] The resource weight redistribution strategy is implemented by calculating the temperature excess of server i, whereby the temperature excess of server i is represented as... It is equipped with hysteresis control constraints. When the hysteresis control constraints are met, the decay factor is calculated using an exponential decay function. , where the parameters and This indicates that the attenuation parameter is set and adjusted in real time through online identification or operational experience; the updated resource weight of server i is obtained. and update the number Power budget allocated to each server .

[0042] When prediction / simulation shows that the temperature of a node exceeds the threshold, the exponential decay weight rapidly reduces the weight of that node (and the effect intensifies with the degree of excess), thereby actively returning power and quickly suppressing the growth of hot spots.

[0043] Hysteresis and minimum hold time design avoid frequent adjustments (jitter) caused by sensing or short pulses, ensuring system stability and preventing secondary failures caused by control jitter.

[0044] The exponential form ensures a nonlinear sensitive response to overheating (the more severe the overheating, the more drastic the decay), which is an effective engineering means to prevent catastrophic thermal runaway.

[0045] The hysteresis control constraint is used to introduce hysteresis and minimum hold time. The constraints include Three consecutive values ​​greater than 0 with a duration greater than .

[0046] Using multiple consecutive threshold exceedances that consistently exceed the minimum hold time as triggering conditions improves robustness and avoids erroneous backflow decisions caused by measurement noise and short-term bursts (such as a single I / O surge); this protects service availability and reduces misjudgments that lead to load drops.

[0047] The power command issuance module includes a power release extraction unit and a resource tree model construction unit; the power release extraction unit is used to extract resources released from the overheating node and calculate the power release. ; The resource tree model construction unit is used to allocate the released power to the candidate server set A according to the initial resource weight ratio, wherein the servers in the candidate server set A satisfy the future time window. The difference between the internal server temperature and the current time t is less than or equal to the set server temperature threshold. First, the initial resource weight of the j-th server in server set A is proportional to the sum of the initial resource weights of all servers in server set A. The resource allocation ratio of the j-th server is obtained, and the initial allocation is set based on the resource allocation ratio. Then, capacity constraints were checked for each loop. If the conditions are not met, scale proportionally. Until the constraints are satisfied and cross-feed allocation or air conditioning increments are allowed or triggered as necessary, This indicates the maximum capacity of the feed (incoming line).

[0048] Calculate the released power and allocate it within the same feed according to the resource tree priority, minimizing cross-loop actions, reducing grid topology constraints and control delays, and reducing short-term current surges and topology-related malfunctions; the temperature threshold requirement for the candidate set ensures that the return current only occurs to nodes with thermal buffer margins, avoiding the migration of thermal problems from one node to another node prone to hotspots; capacity checks and proportional scaling ensure that no allocation violates physical loop limits, avoiding PDU / loop overload caused by the algorithm, which could indirectly trigger server failures.

[0049] The threshold and weight settings can be set by default according to the present invention, or they can be set by those skilled in the art.

[0050] Those skilled in the art will understand that embodiments of the present invention can be provided as methods, systems, or computer program products. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0051] This invention is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations and / or block diagrams. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0052] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.

[0053] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0054] The embodiments of the present invention have been described above with reference to the accompanying drawings. However, the present invention is not limited to the specific embodiments described above. The specific embodiments described above are merely illustrative and not restrictive. Those skilled in the art can make many other forms under the guidance of the present invention without departing from the spirit and scope of the claims. All of these forms are within the protection scope of the present invention.

Claims

1. A multi-loop adaptive power distribution system for edge data centers, characterized in that: include: The data acquisition module is used to deploy a data acquisition network in the computer room to collect data in batches, including the readings of the air inlet temperature sensor built into each server, power data, and IT load data. The IT load intent profiling module builds a server load profile library. When the scheduling intent is captured, it predicts the power curve and peak power within a future time t. The server power prediction module is used to take the load scheduling intention of each rack as input based on historical data and output the predicted power at a future time t. The initial server power allocation module is used to allocate initial resource weights and power budgets to each server based on predicted power. The server resource update module is used to build an electricity-heat-IT digital model and carry a resource weight redistribution strategy. Based on the electricity-heat-IT digital model, it redistributes the resource weights and power resources of all servers. The power command issuance module is used to transfer the power resources of servers whose power resources have changed in the resource weight redistribution strategy back to other servers according to the resource tree model, and then run the server resource update module again for the racks whose power resources have changed, and send air conditioning air volume commands according to temperature changes.

2. The multi-loop adaptive power distribution system for edge data centers according to claim 1, characterized in that: The IT load intent profiling module includes: An offline-built server load profile library is constructed by extracting features from the power time series of historical tasks from startup to steady state and obtaining typical profile clusters through a clustering algorithm. Each profile cluster is saved with a parameterized piecewise function. An online two-stage intent matching unit is used to quickly match the input scheduling intent with a profile database using cosine similarity to obtain an initial template. Then, XGBoost is used as a fine-tuner, taking the parameters of the initial template and the current data center context as input, to output a task-level power prediction curve. and uncertainty .

3. A multi-loop adaptive power distribution system for edge data centers according to claim 2, characterized in that: The server power prediction module adopts a hybrid structure, and the specific steps include: Extracting the predicted power of multiple tasks assigned to the same server , k represents the k-th task, calculated according to the resource requisition coefficient. Perform a weighted summation to obtain the baseline power of the i-th server: ; Using the server's historical power sequence, ambient temperature, and IT load data from the most recent M seconds as input, an LSTM network is used to predict residual terms. Ultimately, server-level power prediction was obtained. ; Subsequently, a standard deviation estimate is output based on model integration. And calculate conservative estimates , This represents the safety factor.

4. The multi-loop adaptive power distribution system for edge data centers according to claim 3, characterized in that: The initial server power allocation module formalizes the weight allocation problem into a constrained convex optimization problem and uses a fast solver for rolling solution. The optimization problem includes: Introducing the variable default rate s, with the objective of minimizing the maximum default rate, the conservative estimate is... As a safety power limit, for any server ,satisfy For any loop ,satisfy And satisfy the total power constraint Simultaneously set constraint rates The optimized output is then subjected to linear ramp smoothing; where... Indicates assignment to the first The resource weight of the i-th server is used to determine the current system status as the i-th server. Power budget allocated to each server .

5. A multi-loop adaptive power distribution system for edge data centers according to claim 4, characterized in that: The electro-thermal-IT digital model includes establishing the equivalent heat capacity for each server i. Equivalent thermal resistance Wind ratio coefficient Discretized temperature dynamic equations are used for time step calculation. The forward simulation yields the future time window. Temperature of internal server i ; like and The difference is less than or equal to the set server temperature threshold. At that time, the system will directly execute the current number of times. Power budget allocated to each server ; like and The difference is greater than the set server temperature threshold. At that time, the resource weight redistribution strategy is implemented.

6. A multi-loop adaptive power distribution system for edge data centers according to claim 5, characterized in that: The time step is calculated using a discretized temperature dynamic equation. Forward simulation includes: ; in, This represents the power-to-heat conversion coefficient of the i-th server. The system's heat dissipation capability for server i is represented as... This indicates the heat dissipation contribution of the air conditioner to the server. This indicates the air supply temperature of server i. This indicates the return air temperature of server i; This indicates the airflow allocated to server i. Indicates the specific heat of air. This indicates the ambient temperature of the computer room.

7. A multi-loop adaptive power distribution system for edge data centers according to claim 6, characterized in that: The resource weight redistribution strategy is implemented by calculating the temperature excess of server i, whereby the temperature excess of server i is represented as... It is equipped with hysteresis control constraints. When the hysteresis control constraints are met, the decay factor is calculated using an exponential decay function. , where the parameters and This represents the decay parameter, which is used to obtain the updated resource weights for server i. and update the number Power budget allocated to each server .

8. A multi-loop adaptive power distribution system for edge data centers according to claim 7, characterized in that: The hysteresis control constraint is used to introduce hysteresis and minimum hold time. The constraints include Three consecutive values ​​greater than 0 with a duration greater than .

9. A multi-loop adaptive power distribution system for edge data centers according to claim 8, characterized in that: The power command issuance module includes a power release extraction unit and a resource tree model construction unit; the power release extraction unit is used to extract resources released from the overheating node and calculate the power release. ; The resource tree model construction unit is used to allocate the released power to the candidate server set A according to the initial resource weight ratio, wherein the servers in the candidate server set A satisfy the future time window. The difference between the internal server temperature and the current time t is less than or equal to the set server temperature threshold. First, the initial resource weight of the j-th server in server set A is proportional to the sum of the initial resource weights of all servers in server set A. The resource allocation ratio of the j-th server is obtained, and the initial allocation is set based on the resource allocation ratio. .

10. A multi-loop adaptive power distribution system for edge data centers according to claim 9, characterized in that: The data acquisition module specifically includes N distributed edge acquisition nodes, and each distributed edge acquisition node includes: a power measurement unit, a temperature sensor interface, a management agent interface, and a communication unit; The power measurement unit is used to collect the bus current of the cabinet. ,Voltage The system samples locally at a high frequency, calculates the RMS, average value, and peak value within a fixed window, and then reports them. The temperature sensor interface is used to directly connect to or act as a proxy to read the temperature sensor readings of the server's built-in air inlet temperature sensor. The management agent interface is used to collect server-side IT metrics; The communication unit is used to perform edge preprocessing on the collected data and report it in batches to the edge controller via the TLS channel.