A server resource allocation method and system

By acquiring the electrical and thermodynamic signals of the power supply network and combining the laws of electromagnetic induction and thermal resistance degradation, a tiered resource allocation strategy is generated, which solves the problem of power supply network collapse under sudden high concurrent traffic and achieves efficient, safe resource allocation and rapid response.

CN122132185AInactive Publication Date: 2026-06-02ZHONGNAN INFORMATION TECH (SHENZHEN) CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
ZHONGNAN INFORMATION TECH (SHENZHEN) CO LTD
Filing Date
2026-05-06
Publication Date
2026-06-02
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Existing technologies rely on pure software scheduling when faced with sudden high concurrency traffic, which can lead to power network collapse and fail to meet the requirements for ultra-fast response. Furthermore, there is a risk of transient voltage drops and downtime in rack-level power supply networks.

Method used

By acquiring the electrical and thermodynamic signals of the power supply network, and combining the laws of electromagnetic induction and thermal resistance degradation effects, the ideal rise slope and thermal attenuation compensation factor of the power supply network are calculated, and a tiered resource allocation strategy is generated to ensure that the power supply network allocates resources within a safe range.

Benefits of technology

It achieves efficient and secure resource allocation during sudden high-concurrency traffic surges, avoids transient power grid outages, and improves computing power response speed and system long-term operational accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122132185A_ABST
    Figure CN122132185A_ABST
Patent Text Reader

Abstract

This invention relates to the field of data processing, specifically to a server resource allocation method and system. The method includes: acquiring the service request queues and corresponding priority weights of computing nodes; acquiring the electrical parameters and thermodynamic characteristics of the power supply network; calculating the ideal ramp-up slope based on the electrical parameters and equivalent parasitic inductance; evaluating the thermal resistance degradation effect of filter components based on thermodynamic characteristics and calculating a thermal attenuation compensation factor; calculating the global computing power ramp-up constraint steps within the current scheduling cycle based on the ideal ramp-up slope and thermal attenuation compensation factor; generating a tiered resource allocation strategy based on the priority weights of the service request queues and the global computing power ramp-up constraint steps, and then converting it into hardware status instructions to schedule node resources. This invention resolves the technical contradiction between rapid global computing power ramp-up and transient power supply failure of rack power supplies, achieving rapid allocation of computing resources while ensuring the safety of the underlying power supply.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of data processing, and more specifically to a server resource allocation method and system. Background Technology

[0002] With the widespread adoption of cloud-native computing and microservice architecture across various industries, modern hyperscale data centers face severe challenges from tidal effects and sudden surges in high-concurrency traffic. In scenarios such as e-commerce flash sales, breaking news alerts, and real-time financial settlements, massive amounts of business requests flood into the computing cluster within seconds, requiring the underlying system to horizontally scale computing resources and vertically increase physical frequency in a very short time. To reduce the enormous physical energy consumption at night or during off-peak hours, existing technologies typically place the central processing units of a large number of physical servers into a deep sleep state during periods of low traffic, causing them to run at the lowest base frequency.

[0003] However, when faced with morning rush hour or sudden traffic surges, this resource allocation based on a pure software perspective reveals fatal flaws: on the one hand, existing pure software scheduling instructions suffer from severe latency when queuing in the operating system kernel network stack and process queue, which cannot meet the demand for rapid response to sudden traffic surges. On the other hand, upon receiving a traffic surge signal, the cloud scheduler, in order to process the request queue as quickly as possible, will adopt an aggressive concurrent wake-up strategy. That is, within a very short time window, it will simultaneously issue wake-up commands and maximum turbo frequency commands to dozens or even hundreds of physical servers on the same rack. When a large number of microprocessors in the rack simultaneously jump from extremely low power consumption to full load power consumption within a nanosecond-level time window, the current change rate of the rack-level power distribution unit is extremely high. According to Faraday's law of electromagnetic induction, under the transient current surge of up to thousands of amperes per second, the inherent parasitic inductance in the power supply network will generate a huge back electromotive force, which will cause a transient voltage drop on the bus. When a large number of computing nodes are simultaneously woken up at full load, the bus voltage will instantly reach the safety threshold of the server motherboard voltage regulation module, triggering the motherboard's undervoltage protection hardware logic, ultimately causing a large-scale power outage and crash of the physical servers in the entire rack. Summary of the Invention

[0004] To address the problem that existing technologies heavily rely on upper-layer software scheduling and neglect the limits of the underlying physical circuitry, leading to rack-level power supply network failure when faced with sudden high-concurrency traffic, this invention provides a server resource allocation method and system.

[0005] In a first aspect, the present invention provides a server resource allocation method, which adopts the following technical solution: The system acquires the service request queue of the computing nodes, obtains and preprocesses the original electrical and thermodynamic signals of the power supply network to obtain electrical parameters and thermodynamic characteristics. Based on the electrical parameters, combined with a preset undervoltage protection threshold and the equivalent parasitic inductance of the power supply network, the ideal ramp-up slope of the power supply network is calculated. The thermal resistance degradation effect of the filtering components of the power supply network is evaluated according to the thermodynamic characteristics, and then a thermal attenuation compensation factor is calculated. The duration of the current scheduling cycle is acquired, and the global computing power ramp-up constraint steps are calculated based on the ideal ramp-up slope, the thermal attenuation compensation factor, and the duration of the current scheduling cycle. The priority weight of the service request queue is determined, and a tiered resource allocation strategy is generated according to the priority weight and the global computing power ramp-up constraint steps, which is then converted into underlying hardware status instructions to schedule the computing nodes.

[0006] This invention solves the problem of blindness in traditional pure software resource scheduling. It takes the physical capacity limit of the underlying power supply network as a hard constraint on computing power allocation. Under the premise of ensuring that the power supply network will not experience transient downtime due to concurrent wake-ups, it achieves efficient and safe allocation of computing node resources.

[0007] Further, the preprocessing includes: performing discrete quantization conversion on the original electrical signal and the original thermodynamic signal respectively; performing smoothing filtering on the converted original electrical signal and the original thermodynamic signal to obtain smoothed electrical data and smoothed thermodynamic data; and performing timestamp matching on the smoothed electrical data and the smoothed thermodynamic data to obtain the spatiotemporally aligned electrical parameters and thermodynamic features.

[0008] This invention effectively filters out electromagnetic interference white noise caused by high-frequency switching actions of the power supply network by performing discrete quantization conversion and smoothing filtering on the original electrical signal and the original thermodynamic signal respectively, and timestamp matching on the heterogeneous data, thus ensuring the purity of the input data. At the same time, it solves the problem of data misalignment caused by differences in sampling frequency and physical location of different sensors, providing accurate underlying data for subsequent calculations.

[0009] Furthermore, the ideal upward slope satisfies the following relationship:

[0010] in, The ideal upward slope; The reference bus voltage; The preset undervoltage protection threshold; The equivalent parasitic inductance is given.

[0011] Furthermore, the thermal attenuation compensation factor satisfies the following relationship:

[0012] in, This refers to the thermal attenuation compensation factor; The physical temperature in the thermodynamic characteristic; The ambient base temperature in the aforementioned thermodynamic characteristics; This is the highest physical tolerance temperature of the filter component; This is the high-frequency thermal resistance sensitivity coefficient.

[0013] Furthermore, the number of steps required to increase the global computing power satisfies the following relationship:

[0014] in, The number of steps constraining the global computing power increase; The ideal upward slope; This refers to the thermal attenuation compensation factor; The duration of the current scheduling period; Increase the base current increment for a single energy efficiency level of the computing node; This is the high-temperature leakage current compensation coefficient; This is the floor symbol.

[0015] Furthermore, the high-frequency thermal resistance sensitivity coefficient was obtained through an offline physical hardware calibration experiment.

[0016] Furthermore, the generation of the tiered resource allocation strategy specifically includes: extracting request feature identifiers from the business request queue, matching the request feature identifiers with preset priority mapping criteria to determine their corresponding priority weights, arranging the business requests to be processed, and allocating the global computing power boost constraint steps to the computing nodes carrying the corresponding business requests in stages according to the arrangement results, and intercepting the computing power boost requests of the unallocated computing nodes in the current scheduling cycle.

[0017] This invention avoids resource contention by allocating limited security computing power to high-priority services such as core transaction links, thereby preventing invalid or low-priority services from monopolizing resources and achieving tiered, refined scheduling driven by both business value and physical security.

[0018] Furthermore, the server resource allocation method further includes: converting the global computing power boosting constraint steps allocated to the computing node into a low-level performance configuration instruction as the low-level hardware status instruction, and then sending a system interrupt signal with the highest response priority to the computing node; then writing the low-level performance configuration instruction into the performance status control register inside the computing node to change the operating frequency of the computing node.

[0019] This invention avoids the latency loss caused by control instructions queuing in the operating system kernel network stack and process queue, and changes the physical operating frequency by penetrating the hardware communication channel, thereby improving the computing power response speed when facing sudden traffic.

[0020] Furthermore, the server resource allocation method further includes: after outputting the underlying hardware status instruction, obtaining the extreme value of bus voltage drop and the preset allowable drop value for the next scheduling cycle, calculating the drop residual based on the extreme value of bus voltage drop and the preset allowable drop value, inputting the drop residual into a preset gradient descent optimizer, and updating the equivalent parasitic inductance for the next scheduling cycle in real time.

[0021] This invention achieves closed-loop adaptive updating of environmental parasitic parameters, eliminates hardware drift errors caused by aging of internal components and fine-tuning of physical structure, and ensures the accuracy of computing power allocation boundaries under long-term operation.

[0022] Secondly, the present invention provides a server resource allocation system, which adopts the following technical solution: A server resource allocation system includes a processor and a memory, wherein the memory stores computer program instructions, and when the computer program instructions are executed by the processor, the server resource allocation method described above is implemented.

[0023] By adopting the above technical solution, a server resource allocation method is generated into a computer program and stored in a memory, which is then loaded and executed by a processor. This allows for the creation of a terminal device based on the memory and processor, making it convenient to use.

[0024] The present invention has the following technical effects: This invention addresses the pain point of rack power supply avalanche in modern hyperscale data centers when dealing with sudden high concurrency traffic. It introduces the physical characteristics of electromagnetic induction and thermal resistance into the upper-layer logic of cloud-native computing power scheduling. By acquiring microsecond-level electrical and thermodynamic fluctuation data, it constructs a software resource allocation method based on the underlying hardware status. This makes the resource allocation strategy no longer dependent on the utilization rate of the microprocessor or the memory level, but controlled by the transient physical tolerance limit of the power supply network, thereby constructing a reliable hardware security protection mechanism. This invention derives the ideal boundary by introducing the law of electromagnetic induction, and implements nonlinear penalty compensation by combining the high-frequency capacitor thermal resistance degradation effect. This can accurately assess the hidden voltage drop risks caused by the surge in the equivalent impedance of the filter component under high temperature conditions. This invention converts physical quantities in the pure electrical domain into equivalent logic steps of energy efficiency levels that can be executed by a computer microprocessor, and is equipped with a parameter adaptive update mechanism based on voltage error residuals, enabling the entire system to maintain high-precision security detection and computing power release in a dynamically changing physical environment. Attached Figure Description

[0025] Figure 1 This is a flowchart of a server resource allocation method provided in an embodiment of the present invention.

[0026] Figure 2 This is a comparison diagram of the effects of the prior art and the method of the present invention provided in the embodiments of the present invention. Detailed Implementation

[0027] This invention provides a server resource allocation method, referring to... Figure 1 This includes steps S1-S4: S1: Data Acquisition and Preprocessing.

[0028] This step mainly involves synchronously collecting business data from computing nodes and basic physical data from the power supply network within the same fixed time period. Specifically, it involves collecting raw features through multi-dimensional sensors, performing discrete quantization, filtering and noise reduction, and heterogeneous data association operations to eliminate asynchronous errors and high-frequency interference, and finally outputting spatiotemporally aligned electrical parameters and thermodynamic characteristics.

[0029] First, the system allocates resources periodically at fixed time intervals, which is the scheduling period, preferably 10ms in this embodiment. To eliminate the asynchronous data update error caused by continuous time flow, this embodiment adopts a strategy of target scheduling time slices that are strictly aligned with the current scheduling period. The duration of this slice is the duration of the current scheduling period. Within this target scheduling time slice, the system treats all collected physical and business states as a static discrete cross-section. By establishing this global time domain constraint, all subsequent calculation steps are in the same time dimension, thus eliminating the need to introduce redundant continuous time variables.

[0030] To accurately acquire the service request queues of computing nodes and the raw electrical and thermodynamic signals of the power supply network, a hardware-aware network needs to be constructed, as detailed below: In terms of acquiring the original electrical signals, this embodiment uses the I2C out-of-band bus built into the baseboard management controller to read the reference bus voltage waveform of the motherboard voltage regulation module using a microsecond-level polling mechanism, as the original electrical signal of the power supply network. In terms of acquiring the original thermodynamic signal, this embodiment uses the silicon diode thermistor built into the motherboard as the sensing source to acquire the real-time temperature fluctuation signal of the power supply phase sequence area of ​​the motherboard voltage regulation module. At the same time, a high-precision thermocouple is deployed at the air inlet of the rack to collect the temperature signal of the ambient temperature fluctuation. The above two signals are used as the original thermodynamic signal of the power supply network. On the service-aware side, the computing nodes directly pull the service request queues fed back in real time from the top-level switch of the rack.

[0031] Because the data center environment contains electromagnetic interference white noise caused by the high-frequency operation of the switching power supply, this embodiment uses a high-precision analog-to-digital converter to perform discrete quantization conversion on the original electrical signal and the original thermodynamic signal respectively. For the converted original electrical signal, a Kalman filter algorithm is used for smoothing to remove sharp noise and obtain smooth electrical data. For the converted original thermodynamic signal, a moving average filter algorithm is used for smoothing to remove high-frequency thermal noise and obtain smooth thermodynamic data.

[0032] Because the physical rate of change of temperature signals is lower than that of electrical signals, there is a sampling frequency difference and phase misalignment between smoothed electrical data and smoothed thermodynamic data. Therefore, it is necessary to perform timestamp matching between the two, as follows: Using the starting clock edge of the target scheduling time slice as a reference, data frames at the same time are aligned using a hardware timer; After completing the spatiotemporal alignment, feature values ​​are extracted from the aligned data frames. First, the arithmetic mean of the smoothed electrical data within the target scheduling time slice is calculated to obtain the reference bus voltage. In this embodiment, the electrical parameter refers to the reference bus voltage. Next, smooth thermodynamic data at the end of the target scheduling time slice is extracted and physically mapped. Smooth thermodynamic data from the power supply phase sequence region of the motherboard voltage regulation module is defined as physical temperature, and smooth thermodynamic data from the rack air inlet is defined as ambient base temperature. In this embodiment, the thermodynamic feature refers to the set of physical temperature and ambient base temperature.

[0033] Finally, the electrical parameters and thermodynamic characteristics are output after denoising and spatiotemporal alignment.

[0034] It should be noted that the specific hardware communication channels, sensor devices, and filtering algorithms used in this embodiment are merely preferred embodiments for implementing the data acquisition and preprocessing of this invention. In actual engineering deployments, those skilled in the art can, based on differences in the underlying computing architecture or the physical characteristics of the sensors, employ methods such as peripheral component interconnection expansion channels or serial peripheral interfaces to acquire the original electrical and thermodynamic signals. Similarly, median filtering and adaptive digital filtering can be used for smoothing and filtering. For the spatiotemporal alignment and numerical feature extraction of heterogeneous data, software-level dynamic interpolation fitting algorithms and microprocessor-based high-precision network time protocols can also be used. All the aforementioned local hardware replacements or algorithm adjustments based on the same logical rules are included within the scope of protection of this invention.

[0035] S2: Transient power supply safety constraints and ideal pull-up slope calculation.

[0036] This step aims to determine the transient physical tolerance limit of the power supply network within the target scheduling time slice. Specifically, based on the electrical parameters obtained from S1, combined with the equivalent parasitic inductance of the power supply network and the undervoltage protection threshold physically fixed on the motherboard, the maximum current change rate allowed by the rack without triggering the underlying protection logic is derived using the physical law of electromagnetic induction, i.e., the ideal pull-up slope.

[0037] First, based on the motherboard hardware manual corresponding to the computing node, the undervoltage protection threshold that is fixed at the factory is extracted. This undervoltage protection threshold represents the absolute physical dead zone boundary of the forced power-off protection mechanism triggered by the motherboard voltage regulation module when the transient voltage drop on the bus is too large. Next, during the maintenance window when business is idle, an extreme wake-up command is issued to the computing node to instantly jump from the lowest base frequency to the highest turbo frequency. Using the high-frequency analog-to-digital converter built into the baseboard management controller, the current increment of the computing node and the corresponding time during this process are collected synchronously, and the two are divided to obtain the reference current change rate. At the same time, the extreme value of the bus voltage drop during this process is captured. According to Faraday's law of electromagnetic induction, the extreme value of the bus voltage drop is divided by the reference current change rate to obtain the equivalent parasitic inductance of the power supply network.

[0038] When a computing node in a rack is woken up by a sudden surge of concurrent service requests within a short period of time, the energy efficiency level of the underlying microprocessor will jump, causing a rapid increase in transient current in the power supply network. According to Faraday's law of electromagnetic induction, under extremely high transient current change rates, the equivalent parasitic inductance of the power supply network will generate a reverse electromotive force, which will lead to a transient voltage drop on the bus. In order to ensure that the voltage drop caused by concurrent wake-up can still be maintained above the undervoltage protection threshold when it bottoms out, it is necessary to use underlying physical constraints to reverse control the overall current rise rate when the computing node is woken up.

[0039] Based on the aforementioned physical security constraint mechanism, under the global time-domain constraint constructed by the target scheduling time slice, the reference bus voltage is used as the static physical starting point of the current scheduling cycle. Combining the undervoltage protection threshold and the equivalent parasitic inductance of the power supply network, the ideal rise slope of the power supply network is calculated. The specific relationship is as follows:

[0040] in, For the ideal upward slope; The reference bus voltage; This is the undervoltage protection threshold. This is the equivalent parasitic inductance.

[0041] The underlying derivation logic of this relation is a reverse engineering application of Faraday's law of electromagnetic induction. According to the chapter on the volt-ampere characteristics of inductors in the book *Circuits*, the terminal voltage of an inductor is proportional to the rate of change of the current flowing through it, as stated in the standard physical equation: It can be seen that The relation is an algebraic rearrangement of the standard physical equation. The term represents the maximum safe voltage reduction margin of the power supply network within the current scheduling cycle. By dividing this term by the equivalent parasitic inductance, the limiting rate of change of current that the power supply network can withstand, i.e. the ideal pull-up slope, can be solved in reverse.

[0042] S3: Thermoelectric coupling compensation and computing power step conversion.

[0043] Specifically, by introducing an environmental degradation penalty mechanism, the hardware impedance drift error caused by high temperature is repaired, and the physical current change rate is equivalently converted into the number of logic energy efficiency steps that a computer microprocessor can execute.

[0044] In large-scale, high-frequency concurrent wake-up scenarios, the continuous high current output of the power supply network causes a sharp rise in the physical temperature of filtering components such as tantalum capacitors and solid-state capacitor arrays. As the temperature rises, the equivalent series resistance of the filtering components undergoes thermal degradation, which generates a hidden additional voltage drop in the power supply circuit. Therefore, to evaluate the thermal degradation effect of the filtering components in the power supply network, a thermal attenuation compensation factor is constructed, with the specific relationship as follows:

[0045] in, This is the thermal decay compensation factor; Physical temperature is a thermodynamic characteristic. This refers to the environmental baseline temperature in thermodynamic characteristics. The maximum physical tolerance temperature of the filter component is obtained from the manufacturer's hardware specification sheet. This is the high-frequency thermal resistance sensitivity coefficient.

[0046] It can be seen that when the rack is operating healthily at low temperature, the thermal attenuation compensation factor approaches 1, and the system allows the scheduler to increase computing power. When high-frequency concurrency causes the filter components to continuously heat up and approach their tolerance limit, the thermal attenuation compensation factor will shrink smoothly, reducing computing power and avoiding power network downtime caused by sudden wake-up of computing nodes, thus achieving a balance between hardware security and computing power release efficiency.

[0047] It should be noted that the high-frequency thermal resistance sensitivity coefficient was obtained through an offline physical hardware calibration experiment, the specific steps of which are as follows: First, a hardware-in-the-loop offline test topology was constructed. A motherboard voltage regulation module of the same model as the computing node was placed in a programmable constant temperature and humidity test chamber. A programmable DC electronic load was connected to the power supply output terminal of the motherboard voltage regulation module to simulate the burst computing power consumption of the microprocessor during high-frequency concurrent wake-up. At the same time, a high-bandwidth digital storage oscilloscope and an active differential voltage probe were connected to the core bus detection point, and a PT100 high-precision platinum resistance temperature sensor was mounted on the surface of the filter component. Next, a multi-temperature-domain microsecond-level step pull-up test was performed; the initial temperature of the programmable constant temperature and humidity test chamber was set to the above-mentioned ambient base temperature, and the programmable DC electronic load was controlled to perform a standard current step pull-up action, that is, the consumed current was pulled up to the full load peak current at a rate of 1000A / ms; the transient voltage drop was captured and recorded using a high-bandwidth digital storage oscilloscope, and recorded as the transient reference voltage drop; Subsequently, the temperature of the programmable constant temperature and humidity test chamber was gradually increased in 5℃ increments until the maximum physical tolerance temperature specified by the filter component was reached. At each temperature that reached thermodynamic equilibrium, the programmable DC electronic load was controlled to perform a standard current step increase action, and the transient voltage drop at each temperature was recorded simultaneously and recorded as the transient actual voltage drop. Subsequently, a data stripping operation of heterogeneous parameters is performed; the transient actual voltage drop at each temperature is extracted and the transient reference voltage drop is subtracted one by one; since the rate of change of the applied current is the same each time, the inductance voltage drop caused by Faraday's law of electromagnetic induction is eliminated, and a pure thermal resistance voltage drop array is obtained, which is only caused by the surge in the equivalent series resistance of the filter component due to high temperature. Finally, text feature fitting and parameter solidification are performed; the temperature difference between each temperature and the ambient baseline temperature is calculated, and divided by the temperature difference between the highest physical tolerance temperature and the ambient baseline temperature to obtain a percentage constant array characterizing the severity of thermal stress; this percentage constant array is used as the independent variable of the horizontal axis, and the corresponding pure thermal resistance voltage drop array is used as the dependent variable of the vertical axis to construct a two-dimensional scatter plot, which is fitted using the least squares method to obtain a geometric trend line, minimizing the sum of squared vertical errors of all scatter points to this geometric trend line; the slope value of this geometric trend line is extracted, which represents the linear attenuation characteristic of hardware impedance with thermal stress, and this slope value is the high-frequency thermal resistance sensitivity coefficient.

[0048] Next, the continuous current state variables need to be converted into discrete digital logic instruction variables. Based on the first-order Taylor expansion-based linear leakage current thermal dependence model proposed in ICCAD ("Accurate Temperature-Dependent Integrated Circuit Leakage Power Estimation is Easy"), the charge overhead is evaluated by calculating the current and temperature coefficients. Based on the first-order RC equivalent model of chip thermodynamics proposed in ECC ("Nonlinear and Asymmetric Thermal-Aware DVFS Control"), the input thermal power is determined by the supply current. To prevent thermal or power failure within a very short scheduling cycle, the maximum allowable injection current increment is proportional to the current remaining temperature margin. Based on this, the duration of the current scheduling cycle predefined in S1 is obtained. Combining the ideal pull-up slope and thermal decay compensation factor, the number of global computing power pull-up constraint steps within the current scheduling cycle is calculated. The specific relationship is as follows:

[0049] in, To increase the global computing power, a limit number of steps is set. For the ideal upward slope; This is the thermal decay compensation factor; The duration of the current scheduling cycle; The base current increment required to improve the energy efficiency level of a computing node is found in the microprocessor architecture white paper; This is the high-temperature leakage current compensation coefficient, which can be found in the semiconductor transistor temperature drift handbook. This is the floor symbol.

[0050] This represents the maximum current increment that the power supply network can withstand within the current scheduling cycle without triggering an underlying undervoltage shutdown; considering the static leakage current effect inside transistors when silicon-based microprocessors operate at high temperatures, This paper utilizes a high-temperature leakage current compensation coefficient to broaden and correct the originally fixed single-step base current increment, obtaining the maximum current cost required for the microprocessor to jump to a higher physical operating frequency under the current thermal environment; finally, it uses... The calculated result is rounded down to ensure power supply safety.

[0051] It can be seen that this physical boundary-based calculation converts continuous physical quantities such as voltage, impedance, and ampere in the pure electrical domain into an exact, discrete digital token, namely the global computing power boosting constraint step, so that the upper-layer cloud-native scheduler has a safe allocation benchmark, ensuring that all computing power wake-up instructions are restricted within the safe range of the underlying physical hardware in the current scheduling cycle.

[0052] S4: Resource scheduling and parameter updates.

[0053] This step aims to deliver the global computing power boosting constraint steps obtained in S3 to the computing nodes in the most efficient way, and to build a self-evolving system of environmental parameters. Specifically, a tiered resource allocation strategy is generated based on business priorities to avoid network stack latency of the operating system kernel. The allocation steps are converted into underlying hardware status instructions in the form of out-of-band interrupt penetration. After the instructions are issued, a gradient descent optimizer is introduced to adaptively update the hardware parasitic parameters to ensure accuracy in each scheduling cycle.

[0054] During sudden traffic surges, the service value carried by different computing nodes within the rack varies; for example, the value of the core transaction link is higher than that of the edge log collection link. To allocate limited computing power to high-value services, the priority weights corresponding to the service request queues obtained in S1 are first acquired. Based on these priority weights and the global computing power increase constraint steps, a tiered resource allocation strategy is generated, as follows: Considering the complexity of microservices in cloud-native clusters in modern data centers, which cannot be exhaustively mapped using traditional static network port lookup mechanisms, this embodiment uses the native service quality classification standard of the underlying container orchestration system to preset priority mapping criteria; extracts the service quality category bound to the target container group carrying the business request as the request feature identifier in the business request queue; Based on the preset priority criteria, a three-level mapping logic is constructed: business requests belonging to the absolute guarantee category, i.e., those with reserved CPU memory and upper and lower limit constraints, are mapped to the highest priority weight; business requests belonging to the elastic fluctuation category, i.e., those with only basic resource requirements and no strict upper limit, are mapped to the medium priority weight; and business requests belonging to the best-effort category, i.e., those low-priority background tasks without any resource constraint declarations, are mapped to the lowest priority weight; for business requests with the same priority weight, the length of their waiting time is used as the standard, with the longer the waiting time, the higher the priority weight. Based on priority weights, the business requests to be processed are sorted in descending order. According to the sorting result, the queue of business requests in descending order is traversed, and the computing power steps corresponding to the business request with the highest priority weight are extracted. It is then determined whether the global computing power increase constraint steps are greater than or equal to the computing power steps. If so, the computing power steps are fully allocated to the corresponding computing node, and the value is deducted equally from the remaining global computing power increase constraint steps. If not, all computing power is allocated to the computing node, the global computing power increase constraint steps are cleared to zero, the traversal is interrupted, and computing power increase requests from unallocated computing nodes in the current scheduling cycle are intercepted, forming a tiered resource allocation strategy.

[0055] Next, using low-level hardware direct control technology, the above-mentioned tiered resource allocation strategy is converted into low-level hardware status instructions to schedule computing nodes to process resources, as follows: First, the global computing power boosting constraint steps allocated to the computing nodes are converted into P-State transition instructions that conform to the advanced configuration and power interface standards, and these instructions are used as underlying hardware state instructions. Furthermore, using the baseboard management controller as a hardware trigger source, a system interrupt signal with the highest response priority is sent to the computing node through a professional sideband communication channel. This system interrupt signal is a non-maskable interrupt, which can forcibly suspend all regular processes currently being executed by the microprocessor. Next, the underlying hardware status instruction is written into the performance status control register inside the computing node, which changes the operating frequency of the computing node by physically overwriting it, ultimately achieving a zero-latency response for computing power boost.

[0056] Considering that data center racks operate under long-term, high-load conditions, copper busbar aging, connector oxidation, and minor adjustments to the physical structure can all cause physical drift in the equivalent parasitic inductance of the power supply network. Therefore, it is necessary to update the equivalent parasitic inductance in real time, as detailed below: After outputting the underlying hardware status command, the high-frequency analog-to-digital converter built into the baseboard management controller is used to continuously monitor the transient waveform of the bus voltage in the next scheduling cycle, extract the lowest voltage valley value of the waveform, and record its absolute value as the extreme value of the bus voltage drop; the undervoltage protection threshold specified in the motherboard hardware manual is extracted as the preset allowable drop value. The algebraic difference between the extreme value of the bus voltage drop and the allowable drop value is calculated to obtain the drop residual, which reflects the deviation between the equivalent parasitic inductance and the true value in the current scheduling cycle. In order to update the equivalent parasitic inductance in real time, this embodiment pre-sets a gradient descent optimizer, specifically an adaptive compensation operator instruction based on the deviation ratio. Its operation logic is as follows: a fixed learning multiplier is pre-programmed into the register, and the drop residual is multiplied by the learning multiplier to obtain the inductance compensation increment. The equivalent parasitic inductance for the current scheduling cycle is subtracted from the inductance compensation increment to obtain the equivalent parasitic inductance for the next scheduling cycle, ensuring that the safe physical boundary of computing power allocation is always under high-precision control.

[0057] It should be noted that this learning multiplier is derived from the inverse derivation of Faraday's law of electromagnetic induction, as follows: Extract the single-step base current increment corresponding to the microprocessor's transition to a single energy efficiency level, and the shortest physical clock cycle consumed on the hardware bus for this transition action. Divide the two to obtain the reference current change rate. Use the reciprocal of the reference current change rate as the learning multiplier, which is 0.01 in this embodiment.

[0058] Figure 2 This is a comparison diagram of the effects of the prior art and the method of the present invention provided in the embodiments of the present invention. It can be seen that during sudden flow peaks, the prior art, due to the lack of physical constraints on the equivalent parasitic inductance of the power supply network, issues full aggressive computing power commands, which causes extremely high transient current change rates, resulting in second-order decay oscillations of the bus voltage. In contrast, the present invention reduces the uncontrollable power outage risk to a smooth and controlled release of computing power, and the bus voltage is always maintained above the undervoltage protection threshold.

Claims

1. A server resource allocation method, characterized in that, include: The system retrieves the service request queue of the computing node, obtains the original electrical and thermodynamic signals of the power supply network and performs preprocessing to obtain electrical parameters and thermodynamic characteristics. Based on the electrical parameters, combined with the preset undervoltage protection threshold and the equivalent parasitic inductance of the power supply network, the ideal pull-up slope of the power supply network is calculated. The thermal resistance degradation effect of the power supply network's filter components is evaluated based on the thermodynamic characteristics, and then the thermal attenuation compensation factor is calculated; the duration of the current scheduling cycle is obtained, and the global computing power increase constraint steps are calculated based on the ideal increase slope, the thermal attenuation compensation factor, and the duration of the current scheduling cycle; The priority weight of the service request queue is determined, and a tiered resource allocation strategy is generated based on the priority weight and the global computing power increase constraint steps. This strategy is then converted into underlying hardware status instructions to schedule the computing nodes.

2. The server resource allocation method according to claim 1, characterized in that, The preprocessing includes: performing discrete quantization conversion on the original electrical signal and the original thermodynamic signal respectively; performing smoothing filtering on the converted original electrical signal and the original thermodynamic signal to obtain smoothed electrical data and smoothed thermodynamic data; and performing timestamp matching on the smoothed electrical data and the smoothed thermodynamic data to obtain the spatiotemporally aligned electrical parameters and thermodynamic features.

3. The server resource allocation method according to claim 1, characterized in that, The ideal upward slope satisfies the following relationship: in, The ideal upward slope; The reference bus voltage; The preset undervoltage protection threshold; The equivalent parasitic inductance is given.

4. The server resource allocation method according to claim 1, characterized in that, The thermal attenuation compensation factor satisfies the following relationship: in, This refers to the thermal attenuation compensation factor; The physical temperature in the thermodynamic characteristic; The ambient base temperature in the aforementioned thermodynamic characteristics; This is the highest physical tolerance temperature of the filter component; This is the high-frequency thermal resistance sensitivity coefficient.

5. A server resource allocation method according to claim 1, characterized in that, The global computing power boost constraint steps satisfy the following relationship: in, The number of steps constraining the global computing power increase; The ideal upward slope; This refers to the thermal attenuation compensation factor; The duration of the current scheduling period; Increase the base current increment for a single energy efficiency level of the computing node; This is the high-temperature leakage current compensation coefficient; This is the floor symbol.

6. A server resource allocation method according to claim 4, characterized in that, The high-frequency thermal resistance sensitivity coefficient was obtained through an offline physical hardware calibration experiment.

7. A server resource allocation method according to claim 1, characterized in that, The generation of the tiered resource allocation strategy specifically includes: extracting request feature identifiers from the business request queue, matching the request feature identifiers with preset priority mapping criteria to determine their corresponding priority weights, arranging the business requests to be processed, and allocating the global computing power boost constraint steps to the computing nodes carrying the corresponding business requests in stages according to the arrangement results, and intercepting the computing power boost requests of the unallocated computing nodes in the current scheduling cycle.

8. A server resource allocation method according to claim 1, characterized in that, The method further includes: converting the global computing power boosting constraint steps allocated to the computing node into a low-level performance configuration instruction as the low-level hardware status instruction, and then sending a system interrupt signal with the highest response priority to the computing node; then writing the low-level performance configuration instruction into the performance status control register inside the computing node to change the operating frequency of the computing node.

9. A server resource allocation method according to claim 1, characterized in that, The method further includes: after outputting the underlying hardware status instruction, obtaining the extreme value of bus voltage drop and the preset allowable drop value for the next scheduling cycle, calculating the drop residual based on the extreme value of bus voltage drop and the preset allowable drop value, inputting the drop residual into a preset gradient descent optimizer, and updating the equivalent parasitic inductance for the next scheduling cycle in real time.

10. A server resource allocation system, characterized in that, include: A processor and a memory, the memory storing computer program instructions that, when executed by the processor, implement a server resource allocation method according to any one of claims 1-9.