Fan resource dynamic allocation method in heterogeneous computing environment

By establishing a timing analysis framework in the intelligent computing server, predicting temperature and load values, and combining the task scheduling queue to adjust fan resource allocation in advance, the problem of untimely fan resource allocation in the existing technology is solved, ensuring timely heat dissipation of the computing unit and improving the heat dissipation effect.

CN120653453AActive Publication Date: 2025-09-16SHENZHEN STONE TECH HLDG CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202511164218.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-20
Publication Date
2025-09-16
Estimated Expiration
2045-08-20

AI Technical Summary

Technical Problem

In intelligent computing servers, the existing fan resource allocation strategy cannot respond in a timely manner to sudden and significant increases in computing unit load or sharp temperature rises, resulting in poor heat dissipation and untimely resource allocation due to frequent adjustments.

Method used

By obtaining the temperature and load data of the target computing unit, a timing analysis framework is established to predict the temperature and load values ​​at the next moment. Combined with the task scheduling queue, the fan resource allocation is adjusted in advance to ensure reasonable allocation.

Benefits of technology

It achieves timely and sufficient heat dissipation under complex working conditions, avoids the problem of untimely allocation of fan resources, and improves the heat dissipation effect.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120653453A_ABST
    Figure CN120653453A_ABST
Patent Text Reader

Abstract

The invention discloses a fan resource dynamic allocation method in a heterogeneous computing environment, and relates to the technical field of resource allocation. The method comprises the following steps: acquiring temperature data and load data of a target calculation unit at each moment in a target time period; based on the current temperature data and the current load data of the target calculation unit, determining whether the target calculation unit meets a resource redistribution condition; determining a predicted temperature value of the target calculation unit at the next moment based on the temperature data and the load data at each moment under the condition that the target calculation unit meets a resource redistribution condition, and determining a predicted load value of the target calculation unit at the next moment based on a task scheduling queue of the target calculation unit; and based on the predicted temperature value and the predicted load value, adjusting the current fan resource allocation quantity of the target calculation unit to obtain the updated fan resource allocation quantity of the target calculation unit. According to the invention, timely heat dissipation for the computing unit can be ensured, and the heat dissipation effect is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of resource allocation, and in particular to a method for dynamically allocating fan resources in a heterogeneous computing environment. Background Art

[0002] With the continuous evolution of artificial intelligence, deep learning, and big data analytics technologies, intelligent computing servers, as high-performance computing devices, are gaining widespread application in numerous fields due to their ease of operation, flexible expansion, and high reliability. However, when operating under high load, the computing units of intelligent computing servers generate a significant amount of heat. If not promptly addressed, this can damage the computing components and cause significant losses. Therefore, an air-cooled cooling system must be installed within the intelligent computing server to effectively dissipate heat through heat sinks.

[0003] At present, by monitoring the load and temperature data of the computing unit in real time and comparing it with the preset curve of fan resource allocation under different loads with temperature changes, the fan resources required at the current moment can be determined, thereby realizing dynamic allocation of fan resources.

[0004] However, when the computing unit's operating load suddenly increases significantly or the temperature rises sharply, the fan resources allocated based on the current load and temperature data may not be sufficient to cope with the subsequent rapid temperature rise. Frequent adjustments to fan resource allocation are required, which can easily lead to untimely fan resource allocation and affect the heat dissipation effect. Summary of the Invention

[0005] The embodiment of the present invention provides a method for dynamically allocating fan resources in a heterogeneous computing environment, which can ensure timely and sufficient heat dissipation for computing units under various complex working conditions, thereby improving the heat dissipation effect.

[0006] According to a first aspect of an embodiment of the present invention, a method for dynamically allocating fan resources in a heterogeneous computing environment is provided. The method includes: Obtain temperature data and load data of the target computing unit at each moment within the target time period; Determining whether the target computing unit meets the resource reallocation condition based on the current temperature data and the current load data of the target computing unit; When the target computing unit meets the resource reallocation conditions, the predicted temperature value of the target computing unit at the next moment is determined based on the temperature data and load data at each moment, and the predicted load value of the target computing unit at the next moment is determined based on the task scheduling queue of the target computing unit; Based on the predicted temperature value and the predicted load value, the current fan resource allocation amount of the target computing unit is adjusted to obtain an updated fan resource allocation amount of the target computing unit.

[0007] Furthermore, the present application also proposes determining the predicted temperature value of the target computing unit at the next moment based on the temperature data and load data at each moment, including: Based on the temperature and load data at each moment, evaluate the degree to which the current fan resource allocation of the target computing unit is ineffective in suppressing temperature growth; Determine the degree of poor heat dissipation of the target computing unit based on the degree of suppression ineffectiveness and the change trend of the temperature data; When the degree of poor heat dissipation is greater than or equal to a preset threshold, a predicted temperature value of the target computing unit at the next moment is determined based on a change trend of the temperature data and a change trend of the load data.

[0008] Furthermore, the present application also proposes to evaluate the degree to which the current fan resource allocation of the target computing unit is ineffective in suppressing temperature growth based on the temperature data and load data at each moment, including: Constructing a first temperature variation sequence based on the temperature data at each moment, wherein the first temperature variation sequence includes instantaneous variations of the temperature data at each moment within the target time period; Determining a temperature growth ratio of the target computing unit within a target time period based on the first temperature change sequence, and determining a load growth ratio of the target computing unit within the target time period based on load data at each moment within the target time period; Based on the temperature growth ratio and the load growth ratio, the degree to which the current fan resource allocation amount of the target computing unit is ineffective in suppressing the temperature growth is determined.

[0009] Furthermore, the present application also proposes determining the degree to which the current fan resource allocation of the target computing unit is ineffective in suppressing temperature growth based on the temperature growth ratio and the load growth ratio, including: Divide the temperature growth ratio by the load growth ratio to obtain a similarity ratio between the temperature growth ratio and the load growth ratio; Matching the current temperature data and the current load data of the target computing unit with the preset curves of fan resource allocation amount versus temperature under different loads to obtain a first theoretical fan resource allocation amount for the target computing unit; The degree of ineffectiveness of the current fan resource allocation of the target computing unit in suppressing temperature growth is determined by using the similarity ratio and the target resource difference of the target computing unit. The target resource difference is the absolute value of the difference between the current fan resource allocation of the target computing unit and the first theoretical fan resource allocation.

[0010] Furthermore, the present application also proposes determining the degree of poor heat dissipation of the target computing unit based on the degree of suppression ineffectiveness and the change trend of temperature data, including: Extracting instantaneous changes from the first temperature change sequence of the target computing unit in sequence from back to front until the instantaneous changes become negative, thereby obtaining a second temperature change sequence of the target computing unit; Performing mean processing on each instantaneous change in the second temperature change sequence to obtain an average temperature increase of the target computing unit; The degree of poor heat dissipation of the target computing unit is determined based on the average temperature increase, the number of instantaneous temperature changes in the second temperature change sequence, and the degree of suppression ineffectiveness.

[0011] Furthermore, the present application also proposes determining the degree of poor heat dissipation of the target computing unit based on the average temperature increase, the number of instantaneous changes in the second temperature change sequence, and the degree of suppression ineffectiveness, including: Multiply the average temperature increase by the number of instantaneous changes to obtain a first calculation result; Subtract the suppression ineffectiveness degree from the preset coefficient value to obtain a second calculation result; The first calculation result is divided by the second calculation result to obtain the degree of poor heat dissipation of the target computing unit.

[0012] Furthermore, the present application also proposes determining the predicted temperature value of the target computing unit at the next moment based on the change trend of the temperature data and the change trend of the load data, including: Determine the current predicted temperature change of the target computing unit based on the change trend of the temperature data and the change trend of the load data; The current temperature data of the target computing unit is added to the current predicted temperature change to obtain the predicted temperature value of the target computing unit at the next moment.

[0013] Furthermore, the present application also proposes determining whether the target computing unit meets the resource reallocation conditions based on the current temperature data and current load data of the target computing unit, including: Matching the current temperature data and the current load data of the target computing unit with the preset curves of fan resource allocation amount versus temperature under different loads to obtain a first theoretical fan resource allocation amount for the target computing unit; When the current fan resource allocation amount of the target computing unit is less than the first theoretical fan resource allocation amount, it is determined that the target computing unit meets the resource reallocation condition.

[0014] Furthermore, the present application also proposes determining the predicted load value of the target computing unit at the next moment based on the task scheduling queue of the target computing unit, including: Get the task computation amount of the target task in the task scheduling queue. The target task is the task that the target computing unit needs to execute at the next moment. Based on the task calculation amount of the target task, determine the total amount of predicted task calculation of the target computing unit at the next moment; The predicted load value of the target computing unit at the next moment is determined by using the total amount of predicted task calculations and the maximum computing capacity of the target computing unit.

[0015] Furthermore, the present application also proposes adjusting the current fan resource allocation of the target computing unit based on the predicted temperature value and the predicted load value to obtain an updated fan resource allocation of the target computing unit, including: Matching the current temperature data and current load data of the target computing unit, as well as the predicted temperature value and predicted load value, with the curves of fan resource allocation under different preset loads as a function of temperature, to obtain a first theoretical fan resource allocation amount and a second theoretical fan resource allocation amount for the target computing unit; An updated fan resource allocation amount of the target computing unit is determined by using the first theoretical fan resource allocation amount, the second theoretical fan resource allocation amount, and the degree of ineffectiveness of suppressing temperature growth of the current fan resource allocation amount of the target computing unit.

[0016] The present invention has the following beneficial effects: In the method for dynamically allocating fan resources in a heterogeneous computing environment provided by an embodiment of the present invention, the temperature and load data of the target computing unit at each moment in the target time period are first obtained to provide a basis for a comprehensive analysis of change trends. Then, based on the current temperature and load data, whether resource reallocation is necessary is determined to avoid blind adjustments. When the resource reallocation conditions are met, the temperature and load data at each moment in the target time period are used to predict the temperature value at the next moment. Simultaneously, the load value at the next moment is predicted in conjunction with the task scheduling queue. This comprehensively considers past data patterns and future task schedules, making the prediction more forward-looking. Finally, based on these two predicted values, the current fan resource allocation is adjusted to obtain an updated and reasonable fan resource allocation plan. In this way, the present invention breaks away from the limitations of relying solely on current data. Through historical tracing and forward-looking prediction, it predicts changes in the cooling requirements of the computing unit in advance, proactively and accurately allocates fan resources, and effectively avoids the problem of untimely fan resource allocation. This ensures that the computing unit can be cooled in a timely and sufficient manner under various complex working conditions, thereby improving the cooling effect. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] In order to more clearly illustrate the technical solutions and advantages of the embodiments of the present invention or the prior art, the following briefly introduces the drawings required for use in the embodiments or the prior art descriptions. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0018] Figure 1 A schematic diagram of a flow chart of a method for dynamically allocating fan resources in a heterogeneous computing environment provided by one embodiment of the present invention; Figure 2 A schematic diagram of the process of S103 provided in one embodiment of the present invention; Figure 3 A schematic diagram of the process of S201 provided in one embodiment of the present invention; Figure 4 An example graph of a curve showing how fan resource allocation varies with temperature under different preset loads provided by an embodiment of the present invention; Figure 5 This is a flow chart of S202 provided in one embodiment of the present invention. DETAILED DESCRIPTION

[0019] To further illustrate the technical means and effectiveness of the present invention in achieving its intended purpose, the following, in conjunction with the accompanying drawings and preferred embodiments, details the specific implementation, structure, features, and effectiveness of a method for dynamically allocating fan resources in a heterogeneous computing environment proposed by the present invention. In the following description, references to "one embodiment" or "another embodiment" do not necessarily refer to the same embodiment. Furthermore, specific features, structures, or characteristics of one or more embodiments may be combined in any suitable manner.

[0020] Unless defined otherwise, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention belongs.

[0021] It should be noted that the acquisition, storage, use, and processing of data in the technical solution of the present invention comply with the relevant provisions of laws and regulations.

[0022] It should be noted that in the embodiments of the present invention, certain software, components, models and other existing solutions in the industry may be mentioned. They should be regarded as exemplary and their purpose is only to illustrate the feasibility of implementing the technical solution of the present invention, but it does not mean that the applicant has or will necessarily use the solution.

[0023] With the booming development of artificial intelligence, deep learning, and big data analytics technologies, intelligent computing servers, as core equipment in the high-performance computing field, have gained widespread adoption in numerous industries, including financial risk control modeling, weather simulation and forecasting, and intelligent medical image analysis, thanks to their numerous advantages, including ease of use, strong scalability, and stable and reliable operation. When performing high-load tasks, the computing units within intelligent computing servers, like a high-load engine, generate massive amounts of heat. If this heat is not dissipated promptly, it acts like a "heat blanket" wrapped around precision instruments, easily leading to performance degradation, data errors, or even permanent damage due to overheating. This can lead to serious consequences such as system crashes and business interruptions, resulting in immeasurable economic losses. Therefore, it is essential to equip intelligent computing servers with an air-cooling system. This air-cooling system utilizes heat sinks, acting as heat transporters, to efficiently transfer heat generated by the computing units to the outside, ensuring stable operation of the server at an appropriate temperature and achieving effective heat management.

[0024] Currently, the mainstream fan resource allocation strategy is to monitor the load and temperature data of the computing unit in real time, and compare it with the pre-set curve of fan resource allocation under different loads and temperature changes, so as to determine the amount of fan resources required at the current moment and achieve dynamic allocation of fan resources.

[0025] However, this strategy has a significant shortcoming: when the load on the computing unit increases dramatically in a short period of time, or the temperature rises suddenly, the fan resources allocated based on the current instantaneous state data are often unable to meet the cooling needs of the subsequent rapid temperature rise. This is like facing a sudden flood, where flood control supplies prepared based on the normal water volume are far from enough. In this case, the system has to frequently adjust the fan resource allocation. This lagging and passive adjustment method can easily lead to untimely fan resource allocation, which in turn seriously affects the cooling effect.

[0026] In the face of the above problems, this application first addresses the problem of delayed allocation of cooling resources in scenarios where the computing unit experiences sudden load and nonlinear temperature increases, and explores how to introduce thermal state predictions at future moments into a dynamic adjustment mechanism. Traditional methods are based on matching preset curves with current data, and are unable to cope with sudden changes in the temperature rise rate caused by a surge in load. To this end, this application considers establishing a time series analysis framework for load and temperature data, extracting the changing trends of temperature and load within the target time period, predicting the temperature and load values ​​at the next moment, and adjusting the fan resource allocation in advance accordingly.

[0027] In this regard, the present application proposes a method for dynamically allocating fan resources in a heterogeneous computing environment, which can be applied to the server. Figure 1As shown, the fan resource dynamic allocation method in the heterogeneous computing environment may include the following steps S101 to S104: S101, obtaining temperature data and load data of a target computing unit at each moment in a target time period.

[0028] In this embodiment, the target computing unit is a hardware module or logic unit in an intelligent computing server that performs specific computing tasks. For example, it can be a single processor core, a computing group consisting of multiple cores, or a computing accelerator card with specific functions. It is the entity that performs specific computing tasks and generates heat that requires fan cooling.

[0029] The target time period is a pre-defined time span used to collect temperature and load data for the target computing unit. The target time period can be set to a specific length based on actual needs, such as 1 minute or 5 minutes. This provides sufficient historical data for subsequent analysis, allowing for more accurate prediction of future temperature and load changes.

[0030] Temperature data is numerical information reflecting the target computing unit's operating temperature at various times, typically acquired through an installed temperature sensor. The temperature sensor converts the computing unit's temperature signal into an electrical signal, which is then processed through analog-to-digital conversion and other processing to obtain a specific temperature value, typically in degrees Celsius.

[0031] Load data is an indicator of the computing pressure on the target compute unit at each moment. It can be measured in various ways, such as central processing unit (CPU) utilization (for CPU compute units), graphics processing unit (GPU) utilization (for GPU compute units), and memory usage. Load data reflects how busy the compute unit is; higher loads typically generate more heat.

[0032] For example, in an intelligent computing server, a temperature sensor is installed on the target computing unit. This sensor monitors the unit's temperature in real time and converts the temperature signal into an electrical signal. The server then reads these electrical signals at regular intervals (e.g., once per second). After analog-to-digital conversion and data processing, it obtains the temperature data for each moment in the target time period.

[0033] The server also uses different methods to obtain load data based on the target compute unit type. For example, for CPU compute units, CPU usage can be obtained through the operating system's system calls; for GPU compute units, GPU utilization can be obtained using the software development kit (SDK) provided by the GPU vendor; and for memory load, memory occupancy can be obtained by querying the operating system's memory management interface.

[0034] S102 : Determine whether the target computing unit meets a resource reallocation condition based on current temperature data and current load data of the target computing unit.

[0035] In this embodiment, resource reallocation conditions are criteria or rules used to determine whether the target computing unit needs to adjust the current fan resource allocation. These conditions can be set based on actual cooling needs and system performance requirements. For example, if the current temperature exceeds a preset temperature threshold (such as 80°C) or the current load exceeds a preset load threshold (such as 70%), the resource reallocation condition is considered to be met.

[0036] As an example, the server reads the current temperature data and the current load data of the target computing unit from a database or memory storing the temperature data and the load data.

[0037] Compare the current temperature data with the preset temperature threshold, and compare the current load data with the preset load threshold. For example, set the preset temperature threshold to 80°C and the preset load threshold to 70%. If the current temperature exceeds 80°C, the current load exceeds 70%, or if both conditions are met, the target computing unit is determined to meet the resource reallocation conditions. Otherwise, the conditions are not met and the unit continues to operate according to the current fan resource allocation.

[0038] S103, when the target computing unit meets the resource reallocation conditions, the predicted temperature value of the target computing unit at the next moment is determined based on the temperature data and load data at each moment, and the predicted load value of the target computing unit at the next moment is determined based on the task scheduling queue of the target computing unit.

[0039] In this embodiment, the task scheduling queue is a list of tasks currently pending and being processed by the target computing unit. Each task in the queue typically includes attributes such as task type, task size, and estimated execution time. The task scheduling queue reflects the load that the computing unit is likely to experience over a period of time and is an important basis for predicting the load value at the next moment.

[0040] The predicted temperature value is an estimate of the target computing unit's likely temperature at the next moment, based on historical temperature and load data from the target computing unit during the target time period. This predicted temperature value helps prepare for cooling in advance and prevent damage to computing components caused by excessive temperatures.

[0041] The predicted load value is an estimate of the load that the target compute unit will experience at the next moment, based on the task information currently in the task scheduling queue and by analyzing the characteristics and execution order of the tasks. This value reflects the future workload of the compute unit and provides a basis for the optimal allocation of fan resources.

[0042] As an example, the server cleans the temperature data and load data collected during the target time period, removes outliers (such as values ​​that are too high or too low due to sensor failure), and normalizes the data so that it is distributed within a certain range to facilitate subsequent algorithm processing.

[0043] Then, select an appropriate forecasting model, such as a time series analysis model (such as an autoregressive integrated moving average model) or a machine learning model (such as a support vector machine or neural network). Use historical temperature and load data as a training set to train the forecasting model and adjust the model parameters to ensure that the model better fits the changing trends of the historical temperature and load data.

[0044] Finally, the current temperature data and load data are input into the trained prediction model to calculate the predicted temperature value at the next moment.

[0045] At the same time, the server further parses the task scheduling queue of the target computing unit to obtain information such as the type, size, and expected execution time of each task in the queue. Based on the characteristics of the task, an appropriate load prediction algorithm is used. For example, for tasks with fixed execution modes and resource requirements, their load contribution at the next moment can be predicted based on the task's historical execution time and resource usage. For complex tasks, a prediction method based on task similarity can be used to match the current task with historical tasks, and the load contribution of the current task can be predicted by referring to the load conditions of similar tasks. The predicted load contributions of each task are accumulated to obtain the predicted load value of the target computing unit at the next moment.

[0046] S104 : Based on the predicted temperature value and the predicted load value, adjust the current fan resource allocation amount of the target computing unit to obtain an updated fan resource allocation amount of the target computing unit.

[0047] In this embodiment, the current fan resource allocation is the fan speed, number of fans, or fan power and other heat dissipation-related resource parameters allocated to the target computing unit at the current moment. These parameters determine the heat dissipation capacity of the fan.

[0048] Updated fan resource allocations are calculated by adjusting the current fan resource allocations based on the predicted temperature and load. The goal of updating fan resource allocations is to more appropriately match the compute unit's future cooling needs and ensure stable operation at an appropriate temperature.

[0049] As an example, the server pre-defines a fan resource allocation strategy. This strategy can be a mapping table that determines the corresponding fan resource allocation based on different combinations of predicted temperature and predicted load values. For example, when the predicted temperature is high and the predicted load is large, a higher fan speed is allocated or the number of fans is increased; when the predicted temperature is low and the predicted load is small, a lower fan speed is allocated or the number of fans is reduced.

[0050] Then, using the predicted temperature and load as input, the system calculates the updated fan resource allocation based on the fan resource allocation policy. For example, if the predicted temperature exceeds 85°C and the predicted load exceeds 80%, the fan speed is increased to 90% of the maximum speed; if the predicted temperature is between 70°C and 80°C and the predicted load is between 50% and 70%, the fan speed is set to 60% of the maximum speed.

[0051] The server adjusts the speed, quantity, or power of the fans corresponding to the target computing unit based on the calculated updated fan resource allocation, realizing dynamic allocation of fan resources to meet the future cooling needs of the computing unit.

[0052] In the method for dynamic fan resource allocation in a heterogeneous computing environment provided by this embodiment, the temperature and load data of the target computing unit at each moment in the target time period are first obtained to provide a basis for comprehensive analysis of change trends. The current temperature and load data are then used to determine whether resource reallocation is necessary, avoiding blind adjustments. When the resource reallocation conditions are met, the temperature and load data at each moment in the target time period are used to predict the temperature value at the next moment. Simultaneously, the load value at the next moment is predicted using the task scheduling queue, comprehensively considering past data patterns and future task schedules to make the prediction more forward-looking. Finally, based on these two predicted values, the current fan resource allocation is adjusted to obtain an updated and reasonable fan resource allocation plan. In this way, the present invention transcends the limitations of relying solely on current data. Through historical tracing and forward-looking prediction, it anticipates changes in the cooling requirements of the computing unit, proactively and accurately allocates fan resources, and effectively avoids the problem of untimely fan resource allocation. This ensures timely and sufficient cooling of the computing unit under various complex working conditions, thereby improving cooling performance.

[0053] In some of the above-mentioned schemes of the present application, when determining the predicted temperature value of the target computing unit at the next moment based on the temperature data and load data at each moment, due to the lack of quantitative analysis of the temperature data and load data, the predicted temperature value of the target computing unit at the next moment is less accurate.

[0054] In this regard, Figure 2 As shown, the present application further proposes that S103 may specifically include the following S201 to S203: S201, evaluating the degree to which the current fan resource allocation of the target computing unit is ineffective in suppressing temperature growth based on temperature data and load data at various moments; S202, determining a degree of poor heat dissipation of the target computing unit based on the degree of suppression ineffectiveness and a change trend of the temperature data; S203 , when the degree of poor heat dissipation is greater than or equal to a preset threshold, determining a predicted temperature value of the target computing unit at the next moment based on a change trend of the temperature data and a change trend of the load data.

[0055] In this embodiment, the suppression ineffectiveness is used to measure the effectiveness of the current fan resource allocation in controlling temperature growth. Specifically, it measures whether the cooling capacity provided by the current fan can effectively slow or prevent the temperature rise of the target computing unit. If the fan cooling capacity is insufficient and the temperature rises rapidly, the suppression ineffectiveness is high. Conversely, if the fan can effectively control the temperature, the suppression ineffectiveness is low.

[0056] The degree of heat dissipation failure assesses the overall performance of the target computing unit's cooling system by combining the degree to which the current fan resource allocation is ineffective in suppressing temperature increases and the changing trends of temperature data. A high degree of heat dissipation failure indicates that the target computing unit faces significant heat dissipation issues, potentially leading to excessive temperatures that affect performance or even damage the hardware. A low degree of heat dissipation failure indicates that the target computing unit's cooling system is performing well.

[0057] The preset threshold is a pre-set critical value for poor cooling. It determines whether the poor cooling reaches a point where further action is required (i.e., predicting the next temperature to adjust fan resources). When the poor cooling level is greater than or equal to the threshold, it indicates a serious cooling problem and requires action.

[0058] The changing trend of temperature data is to analyze the temperature data at each moment over a period of time to determine whether the temperature is rising, falling or stable, as well as the rate of rise or fall, which helps to predict future temperature changes.

[0059] The load data change trend is to analyze the load data at each moment within a period of time to determine whether the load is increasing, decreasing, or remaining stable, as well as the rate of increase or decrease. The load change trend and temperature change trend are interrelated and jointly affect the temperature prediction at the next moment.

[0060] For example, the server collects temperature and load data of the target computing unit at various times from sensors and other devices. Using machine learning regression models or physics-based models, the model is trained to learn the relationship between temperature and load data and temperature growth under different fan resource allocations.

[0061] Then, the current temperature data, load data, and current fan resource allocation are fed into the trained model to determine the difference between the expected temperature increase and the actual temperature increase under the current fan resource allocation. A larger difference indicates a greater degree of ineffectiveness in suppressing temperature increase with the current fan resource allocation; a smaller difference indicates a lower degree of ineffectiveness. For example, if the model predicts a 2°C temperature increase per hour under the current fan resource allocation, but the actual temperature increase is 5°C per hour, the degree of ineffectiveness is high.

[0062] Then, based on experience or experimental data, a weight is assigned to the degree of suppression ineffectiveness. This weight reflects its importance in assessing the degree of poor heat dissipation. For example, if suppression ineffectiveness is considered to have a significant impact on the degree of poor heat dissipation, a higher weight (e.g., 0.7) can be assigned. Time series analysis is then used to analyze the temperature data to determine whether the temperature is rising, falling, or stable, as well as the rate of rise or fall. A rapid temperature increase is assigned a higher score (e.g., 0.8); a slow increase is assigned a lower score (e.g., 0.3); and a stable or declining temperature is assigned a score of 0. The degree of suppression ineffectiveness is then multiplied by the corresponding weight and added to the score for the temperature data trend to obtain a quantitative value for the degree of poor heat dissipation. For example, if the suppression ineffectiveness score is 0.6, the weight is 0.7, and the rapid temperature increase score is 0.8, then the degree of poor heat dissipation = 0.6 × 0.7 + 0.8 = 1.22. (Also, the calculated results can be normalized to fall within an appropriate range based on actual conditions.)

[0063] The calculated degree of poor heat dissipation is then compared with a preset threshold. If it is greater than or equal to the threshold, the next prediction step is performed. If it is less than the threshold, the heat dissipation condition is acceptable and no further forecast adjustment is required. If the value is greater than or equal to the threshold, trend analysis is performed on both the temperature and load data. For temperature data, a time series forecasting model can be used to predict future temperature trends based on historical temperature data. For load data, a similar time series analysis method can be used to predict future load trends.

[0064] Using the changing trends of temperature and load data as input features, a model is built to predict the target computing unit's temperature value at the next moment. A regression model can be used to perform feature engineering on the changing trends of temperature and load data (e.g., extracting features such as the slope and volatility of the trend), which is then input into the model for training. Finally, the current temperature and load data trends are input into the trained prediction model to obtain the predicted temperature value of the target computing unit at the next moment.

[0065] Through this embodiment, the present application can promptly detect abnormal heat dissipation in the computing unit and predict future temperature trends. This helps to adjust fan resource allocation in advance, avoiding the problem of delayed heat dissipation caused by a sharp temperature rise. At the same time, by considering the changing trends of temperature and load, the prediction results are more accurate, and fan resources can be allocated more precisely, improving the cooling effect.

[0066] In some of the above-mentioned schemes of the present application, when evaluating the ineffectiveness of the current fan resources of the target computing unit in suppressing temperature growth, due to the lack of quantitative analysis of the dynamic correlation between temperature and load, it is impossible to accurately identify the matching deviation between the fan resource allocation and the actual heat dissipation requirements.

[0067] In this regard, Figure 3 As shown, the present application further proposes that S201 may specifically include the following S301 to S303: S301, constructing a first temperature variation sequence based on temperature data at various moments, wherein the first temperature variation sequence includes instantaneous variations of temperature data at various moments within a target time period; S302, determining a temperature growth ratio of a target computing unit within a target time period based on the first temperature change sequence, and determining a load growth ratio of the target computing unit within the target time period based on load data at each moment within the target time period; S303 : Determine, based on the temperature growth ratio and the load growth ratio, the degree to which the current fan resource allocation of the target computing unit is ineffective in suppressing the temperature growth.

[0068] In this embodiment, when constructing the first temperature variation sequence, the instantaneous temperature variation at each moment in the target time period is extracted to form a time series for quantifying the temperature fluctuation trend. For example, the temperature data set in the target time period is , n means that a total of n temperature data are recorded, then the first temperature change sequence in the target time period is ,in It is used to represent the i-th instantaneous change in the first temperature change sequence. That is, during the construction of the first temperature change sequence, each instantaneous change is generated by calculating the difference between the temperature data at adjacent moments.

[0069] The temperature growth rate can be determined by the following formula 1: Formula 1 In formula 1, Used to characterize the temperature growth rate of the target computing unit c within the target time period, Used to characterize the n-1th instantaneous change in the first temperature change sequence, It is used to represent the first instantaneous change in the first temperature change sequence, and n is used to represent that a total of n temperature data are recorded.

[0070] The load data set within the target time period is , s means that a total of s load data are recorded. Specifically, the load growth rate can be determined by the following formula 2: Formula 2 In formula 2, It is used to characterize the load growth ratio of the target computing unit c within the target time period. Used to characterize the s-1th load data, It is used to represent the first load data, and s is used to represent that a total of s load data are recorded.

[0071] As an example, the server constructs a first temperature change sequence based on the temperature data at each moment. The first temperature change sequence includes the instantaneous temperature changes at each moment within the target time period. For example, assuming the target time period is 10 minutes and temperature data is collected once per minute, 10 temperature data points are obtained. By calculating the difference between two adjacent temperature data points, 9 instantaneous temperature changes are obtained. These 9 instantaneous temperature changes constitute the first temperature change sequence.

[0072] Then, based on the first temperature variation sequence, the temperature growth rate of the target computing unit within the target time period is determined using the above formula 1. And based on the load data at each moment within the target time period, the load growth rate of the target computing unit within the target time period is determined using the above formula 2.

[0073] Finally, based on the temperature growth ratio and the load growth ratio, the degree to which the current fan resource allocation for the target computing unit is ineffective in suppressing temperature growth is determined. Furthermore, the temperature growth ratio can be divided by the load growth ratio to obtain a ratio. If this ratio is large, it indicates that the temperature is increasing much faster than the load, indicating that the current fan resource allocation is less effective in suppressing temperature growth.

[0074] This embodiment accurately assesses the degree to which the current fan resource allocation for a target computing unit is ineffective in suppressing temperature growth. This allows for timely detection of insufficient fan resource allocation, providing a basis for subsequent dynamic adjustment of fan resources and effectively avoiding poor cooling performance caused by untimely fan resource allocation. Furthermore, by combining and analyzing temperature and load data, a more comprehensive assessment of cooling performance can be achieved, improving the accuracy and rationality of fan resource allocation.

[0075] In some of the above-mentioned schemes in the present application, when evaluating the ineffectiveness of the current fan resource allocation in suppressing temperature growth by using the temperature growth ratio and the load growth ratio, it is difficult to accurately quantify the degree of deviation between the actual fan resource allocation and the theoretical demand by relying solely on the ratio of the two, resulting in the evaluation results of the ineffectiveness of suppression deviating from the actual heat dissipation demand.

[0076] In this regard, the present application further proposes that S303 may specifically include: Divide the temperature growth ratio by the load growth ratio to obtain a similarity ratio between the temperature growth ratio and the load growth ratio; Matching the current temperature data and the current load data of the target computing unit with the preset curves of fan resource allocation amount versus temperature under different loads to obtain a first theoretical fan resource allocation amount for the target computing unit; The degree of ineffectiveness of the current fan resource allocation of the target computing unit in suppressing temperature growth is determined by using the similarity ratio and the target resource difference of the target computing unit. The target resource difference is the absolute value of the difference between the current fan resource allocation of the target computing unit and the first theoretical fan resource allocation.

[0077] In this embodiment, the similarity ratio reflects the dynamic correlation between temperature growth and load growth; the preset curve contains the mapping relationship between the fan resource allocation amount and temperature change under different loads, and the theoretical allocation benchmark is determined by matching the current temperature data with the current load data; the absolute value of the difference measures the deviation between the actual allocation and the theoretical allocation, and the calculation weight of the suppression ineffectiveness is adjusted in combination with the similarity ratio.

[0078] For example, Figure 4 As shown, a curve example of fan resource allocation under different preset loads is provided. In which, fan speed is used as fan resource allocation, the horizontal axis of the coordinate system is used to represent temperature, and the vertical axis of the coordinate system is used to represent fan speed.

[0079] Specifically, after obtaining the temperature growth ratio and load growth ratio of the target computing unit within the target time period, the two are divided to obtain a similarity ratio, which is used to indicate whether the temperature growth is synchronized with the load growth. Subsequently, based on the current temperature data and the current load data, the corresponding first theoretical fan resource allocation is searched on a preset curve. The absolute value of the difference between the current fan resource allocation and the first theoretical fan resource allocation is used as the target resource difference. This target resource difference is then comprehensively considered with the similarity ratio to obtain the degree of suppression ineffectiveness.

[0080] As an example, the server divides the temperature increase ratio by the load increase ratio to obtain the similarity ratio between the temperature increase ratio and the load increase ratio. For example, if the temperature increase ratio is 0.8 and the load increase ratio is 0.5, the similarity ratio is 1.6.

[0081] Then, the current temperature data and current load data of the target computing unit are matched with a preset curve showing how fan resource allocation varies with temperature under different loads to obtain a first theoretical fan resource allocation for the target computing unit. Specifically, an interpolation method can be used to find the point on the preset curve that is closest to the current temperature data and current load data. The corresponding fan resource allocation is the first theoretical fan resource allocation.

[0082] Finally, the degree of ineffectiveness of the current fan resource allocation of the target computing unit in suppressing temperature growth is determined using the similarity ratio and the target resource difference of the target computing unit, where the target resource difference is the absolute value of the difference between the current fan resource allocation of the target computing unit and the first theoretical fan resource allocation.

[0083] Specifically, the degree to which the current fan resource allocation amount of the target computing unit is ineffective in suppressing temperature growth can be determined by the following formula 3: Formula 3 In formula 3, X is used to represent the degree to which the current fan resource allocation is ineffective in suppressing temperature growth. Used to characterize the target resource difference, Used to characterize the temperature growth rate of the target computing unit c within the target time period, Used to represent the load growth ratio of the target computing unit c within the target time period.

[0084] Among them, the larger the target resource difference, the worse the current fan's suppression of temperature changes, and the higher the degree of ineffectiveness of suppression. The similarity ratio indicates the degree to which the temperature growth ratio is close to the load growth ratio. The larger this value is, the more the temperature growth at this time depends on load changes rather than fan resources, indicating that the current fan's suppression of temperature changes is worse, and the degree of ineffectiveness of suppression is higher.

[0085] This embodiment accurately assesses the effectiveness of the current fan resource allocation in suppressing temperature increases. This allows for timely adjustments to fan resource allocation strategies to prevent computing unit overheating caused by continued temperature increases. Furthermore, by considering the correlation between temperature growth and load growth, the accuracy and rationality of fan resource allocation are improved, reducing unnecessary resource waste.

[0086] In some of the above-mentioned schemes of the present application, when the target computing unit is in a stage of continuous temperature rise, the existing method may cause distortion in the calculation of the average temperature increase because the temperature change sequence contains data from the cooling stage, affecting the assessment accuracy of the degree of poor heat dissipation, and thus causing a lag in fan resource adjustment.

[0087] In this regard, Figure 5 As shown, the present application further proposes that S202 may specifically include the following S501 to S503: S501, extracting instantaneous changes from the first temperature change sequence of the target computing unit from the back to the front until the instantaneous change is negative, thereby obtaining a second temperature change sequence of the target computing unit; S502, performing mean processing on each instantaneous change in the second temperature change sequence to obtain an average temperature increase of the target computing unit; S503 : Determine the degree of poor heat dissipation of the target computing unit based on the average temperature increase, the number of instantaneous temperature changes in the second temperature change sequence, and the degree of suppression ineffectiveness.

[0088] In this embodiment, the extraction logic for the second temperature change sequence is limited to traversing the temperature change data backward from the current moment, ensuring that only the instantaneous changes during the period of continuous temperature increase are retained. The average temperature increase is calculated using the arithmetic mean method, summing all elements in the second temperature change sequence and dividing by the number of elements.

[0089] Specifically, during temperature monitoring, when a temperature rise is detected at multiple consecutive moments, the system reversely extracts the most recent consecutive positive temperature changes to form a second temperature change sequence. For example, if five consecutive temperature increases of -0.2°C, 0.5°C, 0.7°C, 0.6°C, and 0.8°C are detected, the system automatically captures the last four positive temperature changes. To calculate the average temperature increase, the four changes are added and divided by 4, resulting in an average temperature increase of 0.65°C.

[0090] As an example, the server extracts the instantaneous changes from the first temperature change sequence of the target computing unit in sequence, from back to front, until the instantaneous change becomes negative, to obtain the second temperature change sequence of the target computing unit. For example, if the first temperature change sequence is [0.5, 0.8, 1.2, -0.3, 0.6, 1.0], the extracted second temperature change sequence is [0.6, 1.0].

[0091] Then, the instantaneous changes in the second temperature change sequence are averaged to obtain the average temperature increase of the target computing unit. Furthermore, the average temperature increase can be obtained by adding all the instantaneous changes in the second temperature change sequence and dividing by the number of instantaneous changes.

[0092] Finally, the degree of heat dissipation in the target computing unit is determined based on the average temperature increase, the number of instantaneous temperature changes in the second temperature change sequence, and the degree of suppression ineffectiveness. Specifically, the average temperature increase can be multiplied by the number of instantaneous temperature changes to obtain a first calculation result; the degree of suppression ineffectiveness can be subtracted from a preset coefficient value to obtain a second calculation result; and the degree of heat dissipation in the target computing unit can be divided by the second calculation result to obtain the degree of heat dissipation in the target computing unit. Thus, by analyzing the temperature change trend and the degree of suppression ineffectiveness, the heat dissipation status of the computing unit can be accurately assessed, providing a basis for subsequent fan resource allocation.

[0093] This embodiment accurately assesses the heat dissipation status of the target computing unit, avoiding untimely fan resource allocation. By analyzing temperature trends and the degree of suppression ineffectiveness, temperature changes can be predicted in advance, enabling timely adjustment of fan resources, improving heat dissipation, and ensuring the stable operation of the intelligent computing server.

[0094] In some of the above-mentioned schemes of the present application, a method for determining the degree of poor heat dissipation based on the degree of suppression ineffectiveness and the trend of temperature data changes is proposed. However, in a scenario where the temperature continues to rise, relying solely on the simple superposition of the degree of suppression ineffectiveness and the temperature change trend may lead to deviations in the assessment of the degree of poor heat dissipation, and fail to accurately reflect the cumulative effect of temperature increase and the combined impact of resource suppression failure, thereby affecting the adjustment accuracy of subsequent fan resource allocation.

[0095] In this regard, the present application further proposes that S503 may specifically include: Multiply the average temperature increase by the number of instantaneous changes to obtain a first calculation result; Subtract the suppression ineffectiveness degree from the preset coefficient value to obtain a second calculation result; The first calculation result is divided by the second calculation result to obtain the degree of poor heat dissipation of the target computing unit.

[0096] In this embodiment, the degree of poor heat dissipation of the target computing unit can be determined by the following formula 4: Formula 4 In formula 4, is used to characterize the degree of poor heat dissipation of the target computing unit, X is used to characterize the ineffectiveness of the current fan resource allocation of the target computing unit in suppressing temperature growth, and m is used to characterize the number of instantaneous changes in the second temperature change sequence. Used to characterize the average temperature increase of the second temperature variation sequence.

[0097] in, Used to characterize the effectiveness of the current fan resource allocation of the target computing unit in suppressing temperature growth. The smaller the value, the worse the match between the current heat dissipation performance and the state of the target computing unit, and the greater the degree of poor heat dissipation of the target computing unit. The larger the average temperature increase and the number of instantaneous temperature changes in the second temperature change sequence, the more continuously and rapidly the temperature is increasing, and the greater the degree of poor heat dissipation of the target computing unit.

[0098] This embodiment allows for a more accurate assessment of the extent of poor heat dissipation in computing units, providing a more reliable basis for subsequent fan resource allocation. This approach, which considers temperature trends and the effectiveness of current fan resource allocation, can identify heat dissipation issues more promptly, helping to improve the response speed and efficiency of the cooling system.

[0099] In some of the above-mentioned schemes of the present application, when determining the predicted temperature value through the changing trend of temperature data and the changing trend of load data, if the current temperature change is not accurately quantified, it may cause a deviation between the predicted temperature value and the actual temperature value, causing the fan resource adjustment to lag behind the temperature rise rate.

[0100] In this regard, the present application further proposes that S203 may specifically include: Determine the current predicted temperature change of the target computing unit based on the change trend of the temperature data and the change trend of the load data; The current temperature data of the target computing unit is added to the current predicted temperature change to obtain the predicted temperature value of the target computing unit at the next moment.

[0101] In this embodiment, the variation trend of the temperature data is analyzed by the temperature difference sequence at consecutive moments; the variation trend of the load data is analyzed by the load data at consecutive moments.

[0102] As an example, the current predicted temperature change of the target computing unit can be determined by the following formula 5: Formula 5 In formula 5, Used to characterize the predicted temperature change of the target computing unit at the nth moment, Used to characterize the instantaneous change of the target computing unit at the n-1th moment, Used to characterize the load data at the nth moment, Used to represent the load data at the n-1th moment.

[0103] in, To calculate the load change based on the load data, To estimate the temperature change under load, we can get the predicted temperature change at the nth moment: .

[0104] Then, after determining the target computing unit's current predicted temperature change using Formula 5, add the target computing unit's current temperature data to the current predicted temperature change to obtain the target computing unit's predicted temperature value at the next moment. For example, if the current temperature is 50°C and the predicted temperature change is 2°C, the predicted temperature value at the next moment is 52°C.

[0105] This embodiment accurately predicts the target computing unit's temperature at the next moment, providing a basis for subsequent fan resource allocation. This allows for pre-adjustment of fan resource allocation, avoiding untimely cooling caused by a sharp temperature rise, and improving cooling effectiveness and the operational stability of the intelligent computing server.

[0106] In some of the above-mentioned schemes of the present application, it is proposed to determine whether the resource reallocation conditions are met based on the current temperature data and the current load data. However, in this process, relying solely on the instantaneous data at the current moment may not accurately reflect the long-term demand for resource allocation. In particular, when the load or temperature suddenly changes, the difference between the theoretical resource allocation demand and the actual allocation amount is not quantified, resulting in a lag in the triggering of the resource reallocation conditions.

[0107] In this regard, the present application further proposes that S102 may specifically include: Matching the current temperature data and the current load data of the target computing unit with the preset curves of fan resource allocation amount versus temperature under different loads to obtain a first theoretical fan resource allocation amount for the target computing unit; When the current fan resource allocation amount of the target computing unit is less than the first theoretical fan resource allocation amount, it is determined that the target computing unit meets the resource reallocation condition.

[0108] In this embodiment, a curve that shows how fan resource allocation varies with temperature under different loads is generated by fitting historical operating data. For example, when the load is 60% and the temperature is 65°C, the corresponding theoretical fan resource allocation is 2000 rpm. The matching process uses interpolation or least squares to map the coordinates of the current temperature data and the current load data onto the curve, outputting the corresponding theoretical allocation. The absolute value of the difference between the current fan resource allocation and the theoretical value is used to quantify the resource gap. When the gap exceeds a preset threshold, the resource reallocation condition is triggered.

[0109] Specifically, when the target computing unit is in an operating state with a load of 70% and a temperature of 68°C, the corresponding theoretical allocation amount is 2200 rpm by querying the preset curve. If the current actual allocation amount is 2000 rpm, the absolute value of the difference is 200 rpm. If the preset threshold is 150 rpm, the difference exceeds the threshold at this time, and it is determined that resource reallocation needs to be triggered. By introducing the theoretical allocation amount as a benchmark, the dynamic adjustment condition is transformed from a single instantaneous data comparison to a quantitative comparison of the actual allocation amount and the theoretical demand, identifying potential resource shortages in advance and avoiding response delays caused by data mutations. For example, in the early stage when the load suddenly increased to 75%, the actual allocation amount had not yet been adjusted, but the theoretical demand had risen to 2300 rpm. At this time, the difference quickly exceeded the threshold, prompting the system to immediately start the reallocation process.

[0110] As an example, in the heterogeneous computing environment of an intelligent computing server, when performing dynamic allocation of fan resources, the current temperature data of computing unit A is first collected in real time, which is 72°C, and the current load data of computing unit A is obtained as 85%. The above real-time data is input into the preset load-temperature-resource mapping database for matching. The database stores the corresponding curves of temperature and theoretical fan speed under different load conditions calibrated through experiments. The query shows that the theoretical fan speed corresponding to the current load of 85% should be 3200 rpm, while the current actual operating speed is 2800 rpm. By comparison, it is determined that the current actual speed is much lower than the theoretical speed, triggering the resource reallocation condition.

[0111] This embodiment effectively addresses the problem of delayed cooling response caused by sudden load changes in the prior art. By matching the theoretical resource allocation curve with measured operating parameters in real time, an adjustment mechanism is triggered immediately upon detecting a resource allocation gap, avoiding the delay effect of traditional methods that rely on historical data predictions. This ensures real-time synchronization between fan resource allocation and temperature changes by predicting resource gaps in advance and proactively intervening in adjustments, thereby suppressing abnormal temperature increases and ensuring stable operation of the computing unit under high-load conditions.

[0112] In some of the above-mentioned solutions of this application, the current fan resource allocation only depends on the current load data, and fails to consider the amount of tasks to be executed in the task scheduling queue at the next moment, resulting in insufficient accuracy of the predicted load value and inability to respond in advance to the surge in heat dissipation demand caused by instantaneous load changes.

[0113] In this regard, the present application further proposes that S103 specifically includes: Get the task computation amount of the target task in the task scheduling queue. The target task is the task that the target computing unit needs to execute at the next moment. Based on the task calculation amount of the target task, determine the total amount of predicted task calculation of the target computing unit at the next moment; The predicted load value of the target computing unit at the next moment is determined by using the total amount of predicted task calculations and the maximum computing capacity of the target computing unit.

[0114] In this embodiment, the task scheduling queue contains the target tasks to be executed at the next moment. By extracting and accumulating the task calculation amounts of these target tasks, the predicted total task calculation amount is obtained. The ratio of the predicted total task calculation amount to the maximum computing power is used as the predicted load value, which can be specifically calculated by the formula: predicted load value = (predicted total task calculation amount / maximum computing power) × 100%. For example, the task scheduling queue contains two tasks with task calculation amounts of 300 units and 500 units respectively, the maximum computing power is 1000 units, the predicted total task calculation amount is 800 units, and the predicted load value is 80%. The load at the next moment is directly calculated based on the amount of tasks to be processed in the task scheduling queue, avoiding the lag of relying solely on current load data.

[0115] Specifically, after obtaining the target task in the task scheduling queue, the total computing amount that needs to be processed at the next moment is obtained by accumulating the task computing amount of each target task. The total computing amount is compared with the maximum computing power of the target computing unit and converted into a predicted load value in percentage form. For example, when the maximum computing power is 1000 units, the total computing amount is 800 units, corresponding to an 80% load value, while the total computing amount is 1200 units, corresponding to a 120% load value, and an overload warning needs to be triggered at this time. The predicted load value and the temperature prediction value are used together as the basis for adjusting fan resources, so that the load change trend can be predicted before the task is actually executed, and the fan resource allocation can be increased in advance. For example, when the predicted load value reaches 80%, the theoretical fan resource allocation is obtained according to the preset curve matching, and dynamic adjustment is performed in combination with the current degree of suppression ineffectiveness, effectively reducing the risk of temperature out of control due to a sudden increase in load.

[0116] As an example, after the computational load of the tasks to be executed in the task scheduling queue is obtained, the number of floating-point operations and data throughput contained in the image recognition task and matrix operation task to be processed at the next moment are extracted respectively. Assuming that the image recognition task needs to perform 1.2×10^6 floating-point operations and the matrix operation task needs to process 8×10^5 data units. After the computational scale of the two is accumulated, the total computational load of the predicted task is 2.0×10^6 units of computational load. Furthermore, the maximum computing power of the target computing unit is set to process 3.0×10^6 units of computational load per second. By multiplying the ratio of the total computational load of the predicted task to the maximum computing power by 100%, the predicted load value for the next moment is finally obtained as 66.7%.

[0117] This embodiment accurately predicts the future load state of the computing unit, enabling fan resource adjustments to match the upcoming computing intensity in advance. This effectively avoids the delayed cooling response caused by sudden load changes in traditional methods. By establishing a direct mapping between task computational effort and load values, the cooling system ensures resource adaptation before temperatures rise, reducing the risk of hardware overheating and maintaining the continuous and stable operation of the computing unit.

[0118] In some of the above-mentioned schemes of the present application, a method of dynamically adjusting the fan resource allocation based on the current temperature data and the current load data is proposed. However, when the operating load of the computing unit increases significantly or the temperature rises sharply, the theoretical fan resource allocation obtained based on the current data matching cannot accurately reflect the temperature and load change trends at subsequent moments, resulting in the problem of insufficient delay in the adjusted fan resource allocation, which may cause the risk of untimely heat dissipation.

[0119] In this regard, the present application further proposes that S104 specifically includes: Matching the current temperature data and current load data of the target computing unit, as well as the predicted temperature value and predicted load value, with the curves of fan resource allocation under different preset loads as a function of temperature, to obtain a first theoretical fan resource allocation amount and a second theoretical fan resource allocation amount for the target computing unit; An updated fan resource allocation amount of the target computing unit is determined by using the first theoretical fan resource allocation amount, the second theoretical fan resource allocation amount, and the degree of ineffectiveness of suppressing temperature growth of the current fan resource allocation amount of the target computing unit.

[0120] In this embodiment, the first theoretical fan resource allocation amount represents the theoretical amount of fan resources required under the current temperature and load conditions, and the second theoretical fan resource allocation amount represents the theoretical amount of fan resources required under the predicted temperature and predicted load conditions; the degree of suppression ineffectiveness quantifies the degree of deviation between the actual effect of the current fan resource allocation amount on temperature control and the theoretical effect.

[0121] As an example, the updated fan resource allocation amount of the target computing unit can be determined by the following formula 6: Formula 6 In formula 6, The updated fan resource allocation used to characterize the target computing unit, A first theoretical fan resource allocation amount for characterizing the target computing unit, A second theoretical fan resource allocation amount for characterizing the target computing unit, Indicates the degree to which the current fan resource allocation is ineffective in suppressing temperature growth.

[0122] Specifically, the first theoretical fan resource allocation amount and the second theoretical fan resource allocation For reference, the corresponding weight is set based on the ineffectiveness X of the current fan resource allocation on the temperature increase, and the updated fan resource allocation of the target computing unit can be obtained.

[0123] Through this embodiment, the problem of lagging fan resource adjustment due to sudden load changes in the prior art is effectively solved. By dynamically matching the theoretical resource amounts under current temperature and load conditions and predicted temperature and load conditions, and combining the quantitative evaluation of historical resource suppression effects, feedforward optimization of fan resource allocation is achieved. This adjustment mechanism based on multi-dimensional data fusion can significantly shorten the response delay of the cooling system and maintain the thermal stability of the target computing unit during the rapid temperature rise stage, thereby ensuring the continuous and reliable operation of the intelligent computing server under sudden high-load conditions.

[0124] It should be noted that the order in which the embodiments of the present invention are described above is for illustrative purposes only and does not necessarily represent the superiority or inferiority of the embodiments. The processes depicted in the accompanying drawings do not necessarily require the specific order or sequential order shown to achieve the desired results. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0125] The various embodiments in this specification are described in a progressive manner, and the same or similar parts between the various embodiments can be referred to each other. Each embodiment focuses on the differences from other embodiments.

Claims

1. A method for dynamically allocating fan resources in a heterogeneous computing environment, characterized in that: The method comprises: Obtain temperature data and load data of the target computing unit at each moment within the target time period; determining whether the target computing unit meets a resource reallocation condition based on current temperature data and current load data of the target computing unit; If the target computing unit meets the resource reallocation condition, determining a predicted temperature value of the target computing unit at a next moment based on the temperature data and the load data at each moment, and determining a predicted load value of the target computing unit at a next moment based on the task scheduling queue of the target computing unit; Based on the predicted temperature value and the predicted load value, the current fan resource allocation amount of the target computing unit is adjusted to obtain an updated fan resource allocation amount of the target computing unit.

2. The method for dynamic allocation of fan resources in a heterogeneous computing environment according to claim 1, characterized in that: The determining, based on the temperature data and the load data at each moment, a predicted temperature value of the target computing unit at the next moment, includes: evaluating, based on the temperature data and load data at each moment, the degree to which the current fan resource allocation of the target computing unit is ineffective in suppressing temperature growth; determining a degree of poor heat dissipation of the target computing unit based on the degree of ineffective suppression and a change trend of the temperature data; When the degree of poor heat dissipation is greater than or equal to a preset threshold, a predicted temperature value of the target computing unit at the next moment is determined based on a change trend of the temperature data and a change trend of the load data.

3. The method for dynamic allocation of fan resources in a heterogeneous computing environment according to claim 2, characterized in that: The evaluating, based on the temperature data and the load data at each moment, the degree to which the current fan resource allocation of the target computing unit is ineffective in suppressing temperature growth includes: constructing a first temperature variation sequence based on the temperature data at each moment, wherein the first temperature variation sequence includes instantaneous variations of the temperature data at each moment within the target time period; determining a temperature growth ratio of the target computing unit within the target time period based on the first temperature change sequence, and determining a load growth ratio of the target computing unit within the target time period based on load data at each moment within the target time period; Based on the temperature growth ratio and the load growth ratio, the degree of ineffectiveness of suppressing the temperature growth of the target computing unit's current fan resource allocation is determined.

4. The method for dynamic allocation of fan resources in a heterogeneous computing environment according to claim 3, characterized in that: The determining, based on the temperature growth ratio and the load growth ratio, the degree to which the current fan resource allocation amount of the target computing unit is ineffective in suppressing the temperature growth includes: Dividing the temperature increase ratio by the load increase ratio to obtain a similarity ratio between the temperature increase ratio and the load increase ratio; Matching the current temperature data and the current load data of the target computing unit with a preset curve of fan resource allocation amount versus temperature under different loads to obtain a first theoretical fan resource allocation amount for the target computing unit; The similarity ratio and the target resource difference of the target computing unit are used to determine the degree of ineffectiveness of the current fan resource allocation of the target computing unit in suppressing temperature growth. The target resource difference is the absolute value of the difference between the current fan resource allocation of the target computing unit and the first theoretical fan resource allocation.

5. The method for dynamic allocation of fan resources in a heterogeneous computing environment according to claim 2, wherein: The determining, based on the suppression ineffectiveness degree and the change trend of the temperature data, the degree of poor heat dissipation of the target computing unit includes: Extracting instantaneous changes from the first temperature change sequence of the target computing unit in sequence from back to front until the instantaneous changes are negative, thereby obtaining a second temperature change sequence of the target computing unit; performing mean processing on each of the instantaneous changes in the second temperature change sequence to obtain an average temperature increase of the target computing unit; The degree of poor heat dissipation of the target computing unit is determined based on the average temperature increase, the number of instantaneous temperature changes in the second temperature change sequence, and the degree of suppression ineffectiveness.

6. The method for dynamic allocation of fan resources in a heterogeneous computing environment according to claim 5, characterized in that: The determining the degree of poor heat dissipation of the target computing unit based on the average temperature increase, the number of instantaneous temperature changes in the second temperature change sequence, and the degree of suppression ineffectiveness includes: Multiplying the average temperature increase by the number of instantaneous changes to obtain a first calculation result; Subtracting the suppression ineffectiveness degree from the preset coefficient value to obtain a second calculation result; The first calculation result is divided by the second calculation result to obtain the degree of poor heat dissipation of the target computing unit.

7. The method for dynamic allocation of fan resources in a heterogeneous computing environment according to claim 2, wherein: The determining, based on the change trend of the temperature data and the change trend of the load data, a predicted temperature value of the target computing unit at a next moment, includes: determining a current predicted temperature change of the target computing unit based on a change trend of the temperature data and a change trend of the load data; The current temperature data of the target computing unit is added to the current predicted temperature change to obtain the predicted temperature value of the target computing unit at the next moment.

8. The method for dynamically allocating fan resources in a heterogeneous computing environment according to any one of claims 1 to 7, wherein: The determining, based on the current temperature data and the current load data of the target computing unit, whether the target computing unit meets the resource reallocation condition includes: Matching the current temperature data and the current load data of the target computing unit with a preset curve of fan resource allocation amount versus temperature under different loads to obtain a first theoretical fan resource allocation amount for the target computing unit; In a case where the current fan resource allocation amount of the target computing unit is less than the first theoretical fan resource allocation amount, it is determined that the target computing unit meets the resource reallocation condition.

9. The method for dynamically allocating fan resources in a heterogeneous computing environment according to any one of claims 1 to 7, wherein: The determining, based on the task scheduling queue of the target computing unit, a predicted load value of the target computing unit at a next moment, includes: Obtaining a task computation amount of a target task in the task scheduling queue, where the target task is a task that the target computing unit needs to execute at the next moment; Determining the total amount of predicted task calculations of the target computing unit at the next moment based on the task calculation amount of the target task; The predicted load value of the target computing unit at the next moment is determined by using the predicted task computing amount and the maximum computing capacity of the target computing unit.

10. The method for dynamically allocating fan resources in a heterogeneous computing environment according to any one of claims 1 to 7, wherein: The adjusting the current fan resource allocation amount of the target computing unit based on the predicted temperature value and the predicted load value to obtain an updated fan resource allocation amount of the target computing unit includes: Matching the current temperature data and the current load data of the target computing unit, as well as the predicted temperature value and the predicted load value, respectively, with a curve showing a change in fan resource allocation amount with temperature under different preset loads to obtain a first theoretical fan resource allocation amount and a second theoretical fan resource allocation amount for the target computing unit; An updated fan resource allocation amount of the target computing unit is determined by using the first theoretical fan resource allocation amount, the second theoretical fan resource allocation amount, and the degree of ineffectiveness of suppressing temperature growth of the current fan resource allocation amount of the target computing unit.

Citation Information

Patent Citations

  • Hadoop computing task initial allocation method based on load prediction

    CN110262897A

  • Computing power resource processing method

    CN118069380A

  • Methods, apparatus, and systems to dynamically schedule workloads among compute resources based on temperature

    US20200326994A1