A fan resource dynamic allocation method in a heterogeneous computing environment
By predicting the future temperature and load values of the computing unit in the intelligent computing server and dynamically adjusting the fan resource allocation, the problem of untimely fan resource allocation in the existing technology is solved, and effective heat dissipation is achieved under complex working conditions.
Patent Information
- Application Number
- CN202511164218.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-20
- Publication Date
- 2025-10-17
- Estimated Expiration
- 2045-08-20
AI Technical Summary
In intelligent computing servers, the existing fan resource allocation strategy cannot promptly respond to situations where the computing unit load suddenly increases significantly or the temperature rises sharply, resulting in poor heat dissipation and affecting the stability of the computing components.
By acquiring the temperature and load data of the target computing unit, a timing analysis framework is established to predict the temperature and load values at future moments. The fan resource allocation is adjusted based on the predicted values to achieve dynamic allocation.
Ensure timely and sufficient heat dissipation for the computing unit under complex working conditions, improve the heat dissipation effect, and avoid the problem of untimely allocation of fan resources.
Smart Images

Figure CN120653453B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of resource allocation, and particularly relates to a fan resource dynamic allocation method in a heterogeneous computing environment. BACKGROUND
[0002] With the continuous evolution of artificial intelligence, deep learning and big data analysis technology, the intelligent computing server as a high-performance computing device is widely used in many fields due to its simple operation, flexible expansion and high reliability. However, when the computing units of the intelligent computing server are running under high load, a large amount of heat will be generated, which may damage the computing components and cause significant losses if not handled in time. Therefore, a fan cooling system needs to be installed inside the intelligent computing server to conduct heat through the cooling fins and achieve effective heat dissipation.
[0003] At present, the load and temperature data of the computing unit are monitored in real time, and compared with the preset curve of the fan resource allocation amount changing with temperature under different loads to determine the required fan resource at the current time, and then the dynamic allocation of the fan resource is realized.
[0004] However, when the computing unit load instantaneously rises sharply or the temperature rises sharply, the fan resource allocated according to the current load and temperature data may not be sufficient to cope with the subsequent rapid temperature rise, and frequent multiple adjustments of the fan resource allocation are required, which may cause the problem of untimely fan resource allocation and affect the heat dissipation effect. SUMMARY
[0005] The embodiment of the present application provides a fan resource dynamic allocation method in a heterogeneous computing environment, which can ensure that the computing unit is cooled in time and adequately under various complex working conditions, and can improve the heat dissipation effect.
[0006] In a first aspect, the embodiment of the present application provides a fan resource dynamic allocation method in a heterogeneous computing environment, which comprises:
[0007] Obtaining temperature data and load data of the target computing unit at each time in a target time period;
[0008] Based on the current temperature data and the current load data of the target computing unit, determining whether the target computing unit meets the resource reallocation condition;
[0009] In the case that the target computing unit meets the resource reallocation condition, based on the temperature data and the load data at each time, determining a predicted temperature value of the target computing unit at the next time, and based on the task scheduling queue of the target computing unit, determining a predicted load value of the target computing unit at the next time;
[0010] Adjust the current fan resource allocation amount of the target computing unit based on the predicted temperature value and the predicted load value, to obtain an updated fan resource allocation amount of the target computing unit.
[0011] Further, the application also proposes determining the predicted temperature value of the target computing unit at the next time based on the temperature data and the load data at each time, comprising:
[0012] According to the temperature data and the load data at each time, evaluating the invalidity degree of the current fan resource allocation amount of the target computing unit in inhibiting temperature growth;
[0013] Based on the invalidity degree and the change trend of the temperature data, determining the heat dissipation poor degree of the target computing unit;
[0014] In the case where the heat dissipation poor degree is greater than or equal to a preset degree threshold, determining the predicted temperature value of the target computing unit at the next time based on the change trend of the temperature data and the change trend of the load data.
[0015] Further, the application also proposes evaluating the invalidity degree of the current fan resource allocation amount of the target computing unit in inhibiting temperature growth according to the temperature data and the load data at each time, comprising:
[0016] According to the temperature data at each time, constructing a first temperature change amount sequence, and the first temperature change amount sequence includes the instantaneous change amount of the temperature data at each time within a target time period;
[0017] Based on the first temperature change amount sequence, determining the temperature growth rate of the target computing unit within the target time period, and based on the load data at each time within the target time period, determining the load growth rate of the target computing unit within the target time period;
[0018] Based on the temperature growth rate and the load growth rate, determining the invalidity degree of the current fan resource allocation amount of the target computing unit in inhibiting temperature growth.
[0019] Further, the application also proposes determining the invalidity degree of the current fan resource allocation amount of the target computing unit in inhibiting temperature growth based on the temperature growth rate and the load growth rate, comprising:
[0020] Dividing the temperature growth rate by the load growth rate to obtain a similarity degree ratio of the temperature growth rate and the load growth rate;
[0021] Matching the current temperature data and the current load data of the target computing unit with a preset fan resource allocation amount curve varying with temperature under different loads to obtain a first theoretical fan resource allocation amount of the target computing unit;
[0022] The similarity degree ratio and a target resource difference value of the target computing unit are used to determine an invalid degree of inhibition of the current fan resource allocation of the target computing unit on temperature increase, and the target resource difference value is an absolute value of a difference between the current fan resource allocation of the target computing unit and the first theoretical fan resource allocation.
[0023] Further, the application also proposes determining the poor heat dissipation degree of the target computing unit based on the invalid degree of inhibition and the change trend of the temperature data, including:
[0024] The instantaneous change amounts are extracted from the first temperature change amount sequence of the target computing unit in a backward order until the instantaneous change amount is negative, to obtain a second temperature change amount sequence of the target computing unit.
[0025] The mean value of each instantaneous change amount in the second temperature change amount sequence is processed to obtain a temperature average increase amplitude of the target computing unit.
[0026] The poor heat dissipation degree of the target computing unit is determined based on the temperature average increase amplitude, the number of instantaneous change amounts in the second temperature change amount sequence, and the invalid degree of inhibition.
[0027] Further, the application also proposes determining the poor heat dissipation degree of the target computing unit based on the temperature average increase amplitude, the number of instantaneous change amounts in the second temperature change amount sequence, and the invalid degree of inhibition, including:
[0028] The temperature average increase amplitude is multiplied by the number of instantaneous change amounts to obtain a first calculation result.
[0029] A preset coefficient value is subtracted by the invalid degree of inhibition to obtain a second calculation result.
[0030] The first calculation result is divided by the second calculation result to obtain the poor heat dissipation degree of the target computing unit.
[0031] Further, the application also proposes determining the predicted temperature value of the target computing unit at the next moment based on the change trend of the temperature data and the change trend of the load data, including:
[0032] The current predicted temperature change amount of the target computing unit is determined based on the change trend of the temperature data and the change trend of the load data.
[0033] The predicted temperature value of the target computing unit at the next moment is obtained by adding the current predicted temperature change amount to the current temperature data of the target computing unit.
[0034] Further, the application also proposes determining whether the target computing unit meets the resource re-allocation condition based on the current temperature data and the current load data of the target computing unit, including:
[0035] Match the current temperature data and the current load data of the target computing unit with the preset curve of the fan resource allocation amount changing with temperature under different loads to obtain a first theoretical fan resource allocation amount of the target computing unit;
[0036] In a case where the current fan resource allocation amount of the target computing unit is less than the first theoretical fan resource allocation amount, it is determined that the target computing unit meets the resource re-allocation condition.
[0037] Further, the application also proposes that, based on the task scheduling queue of the target computing unit, a predicted load value of the target computing unit at the next moment is determined, comprising:
[0038] Obtaining a task computing amount of a target task in the task scheduling queue, the target task being a task to be executed by the target computing unit at the next moment;
[0039] Based on the task computing amount of the target task, a predicted total task computing amount of the target computing unit at the next moment is determined;
[0040] Using the predicted total task computing amount and the maximum computing capacity of the target computing unit, a predicted load value of the target computing unit at the next moment is determined.
[0041] Further, the application also proposes that, based on the predicted temperature value and the predicted load value, the current fan resource allocation amount of the target computing unit is adjusted to obtain an updated fan resource allocation amount of the target computing unit, comprising:
[0042] Matching the current temperature data and the current load data of the target computing unit and the predicted temperature value and the predicted load value with the preset curve of the fan resource allocation amount changing with temperature under different loads to obtain a first theoretical fan resource allocation amount and a second theoretical fan resource allocation amount of the target computing unit;
[0043] Using the first theoretical fan resource allocation amount, the second theoretical fan resource allocation amount and the current fan resource allocation amount of the target computing unit to determine the updated fan resource allocation amount of the target computing unit.
[0044] The application has the following beneficial effects:
[0045] The fan resource dynamic allocation method in the heterogeneous computing environment provided by the embodiment of the application comprises the following steps: first, temperature data and load data of a target computing unit at each moment in a target time period are acquired, so as to provide a basis for comprehensive analysis of the change trend; then, whether resource reallocation is needed is judged according to the current temperature data and the current load data, so as to avoid blind adjustment. When the resource reallocation condition is met, the temperature value at the next moment is predicted by using the temperature data and the load data at each moment in the target time period, and the load value at the next moment is predicted by combining the task scheduling queue, so that the past data rule and the future task arrangement are comprehensively considered, and the prediction is more prospective. Finally, the current fan resource allocation amount is adjusted based on the two predicted values, and an updated fan resource reasonable allocation scheme is obtained. In this way, the application jumps out of the limitation of relying on the current data, and the change of the cooling demand of the computing unit is predicted in advance by means of historical tracing and prospective prediction, so that the fan resource is actively and accurately allocated, and the problem of untimely fan resource allocation is effectively avoided. Therefore, the computing unit can be cooled in time and sufficiently under various complex working conditions, and the cooling effect can be improved. BRIEF DESCRIPTION OF DRAWINGS
[0046] In order to more clearly illustrate the technical solutions in the embodiments of the application or the prior art, and the advantages thereof, a brief introduction will be given to the drawings needed in the embodiments or the prior art description. Obviously, the drawings in the following description only show some embodiments of the application, and for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0047] Figure 1 A flowchart of a fan resource dynamic allocation method in a heterogeneous computing environment provided by an embodiment of the application is shown in the figure.
[0048] Figure 2 A flowchart of S103 provided by an embodiment of the application is shown in the figure.
[0049] Figure 3 A flowchart of S201 provided by an embodiment of the application is shown in the figure.
[0050] Figure 4 An example of a curve showing that the fan resource allocation amount changes with temperature under different loads provided by an embodiment of the application is shown in the figure.
[0051] Figure 5 A flowchart of S202 provided by an embodiment of the application is shown in the figure. DETAILED DESCRIPTION
[0052] In order to further illustrate the technical means and effects taken by the present application to achieve the predetermined inventive purpose, the following describes in detail the specific implementation, structure, features and effects of a fan resource dynamic allocation method in a heterogeneous computing environment according to the present application, with reference to the accompanying drawings and preferred embodiments. In the following description, different "one embodiment" or "another embodiment" do not necessarily refer to the same embodiment. In addition, the specific features, structures or characteristics in one or more embodiments can be combined in any suitable form.
[0053] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs.
[0054] It should be noted that the acquisition, storage, use, processing, etc. of data in the technical solutions of the present application comply with the relevant provisions of laws and regulations.
[0055] It should be noted that in the embodiments of the present application, some industry existing solutions of software, components, models, etc. may be mentioned, which should be considered as exemplary, and the purpose is only to illustrate the feasibility of the implementation of the technical solutions of the present application, but it does not mean that the applicant has or will necessarily use the solution.
[0056] At present, the mainstream fan resource allocation strategy is to monitor the load and temperature data of the computing unit in real time, and compare them with the pre-set curve of the fan resource allocation amount changing with temperature under different loads, to determine the required fan resource amount at the current time, so as to achieve dynamic allocation of fan resources.
[0057] At present, the mainstream fan resource allocation strategy is to monitor the load and temperature data of the computing unit in real time, and compare them with the pre-set curve of the fan resource allocation amount changing with temperature under different loads, to determine the required fan resource amount at the current time, so as to achieve dynamic allocation of fan resources.
[0058] But this strategy has obvious short board: when the load of the computing unit increases sharply in a short time, or the temperature rises sharply, the fan resources allocated according to the current instantaneous state data often cannot meet the subsequent rapid temperature rise cooling demand. It is like facing a sudden flood, and the flood control materials prepared according to the normal water volume are far from enough. At this time, the system has to adjust the fan resource allocation quantity frequently, and this lagging and passive adjustment method is easy to cause the fan resource allocation not timely, and then seriously affects the cooling effect.
[0059] In the face of the above problems, the present application first aims at the cooling resource lagging allocation problem in the scenario of sudden load and nonlinear temperature rise of the computing unit, explores how to introduce the future time heat state prediction into the dynamic adjustment mechanism. The traditional method cannot cope with the sudden change of temperature rise rate caused by load surge based on the current time data matching the preset curve. For this, the present application considers to establish a time series analysis framework of load and temperature data, predicts the temperature and load values at the next time by extracting the temperature and load change trend in the target time period, and adjusts the fan resource allocation quantity in advance accordingly.
[0060] For this, the present application proposes a fan resource dynamic allocation method in a heterogeneous computing environment, which can be applied to a server. As shown in Figure 1 The fan resource dynamic allocation method in the heterogeneous computing environment can include the following S101-S104:
[0061] S101, obtaining temperature data and load data of a target computing unit at each time in a target time period.
[0062] In this embodiment, the target computing unit is a hardware module or a logical unit in an intelligent computing server that undertakes specific computing tasks. For example, it can be a single processor core, a computing group composed of multiple cores, or a computing acceleration card with specific functions, which is the main body that executes specific computing tasks and generates heat and needs fan cooling.
[0063] The target time period is a pre-set time span for collecting temperature data and load data of the target computing unit. The length of the target time period can be set according to actual needs, for example, it can be 1 minute, 5 minutes, etc. in the past, and its role is to provide enough historical data for subsequent analysis, so as to more accurately predict future temperature changes and load changes.
[0064] The temperature data is numerical information reflecting the working temperature of the target computing unit at each time, which is usually obtained through the installed temperature sensor. The temperature sensor converts the temperature signal of the computing unit into an electrical signal, and then processes it through analog-to-digital conversion, etc. to obtain the specific temperature value, which is generally in Celsius.
[0065] Load data is an indicator representing the computing task pressure that the target computing unit bears at each time point, which can be measured in various ways, such as central processing unit (CPU) usage (for CPU computing units), graphics processing unit (GPU) utilization (for GPU computing units), memory occupancy, etc. Load data reflects the busy degree of the computing unit, and the higher the load, the more heat the computing unit usually generates.
[0066] As an example, in the intelligent computing server, a temperature sensor is installed for the target computing unit, which monitors the temperature of the computing unit in real time and converts the temperature signal into an electrical signal. Then the server reads these electrical signals at certain time intervals (such as once per second), and after analog-to-digital conversion and data processing, the temperature data at each time point in the target time period is obtained.
[0067] At the same time, the server obtains load data in different ways according to the type of the target computing unit. For example, for CPU computing units, CPU usage can be obtained through system calls of the operating system; for GPU computing units, GPU utilization can be obtained using the software development kit (SDK) provided by the GPU manufacturer; for memory load, memory occupancy can be obtained by querying the memory management interface of the operating system.
[0068] S102, based on the current temperature data and the current load data of the target computing unit, determine whether the target computing unit meets the resource reallocation condition.
[0069] In this embodiment, the resource reallocation condition is a standard or rule for judging whether the target computing unit needs to adjust the current fan resource allocation. These conditions can be set according to actual heat dissipation needs and system performance requirements, for example, when the current temperature exceeds a certain preset temperature threshold (such as 80℃), or the current load exceeds a certain preset load threshold (such as 70%), it is considered that the resource reallocation condition is met.
[0070] As an example, the server reads the current temperature data and the current load data of the target computing unit from the database or memory that stores the temperature data and the load data.
[0071] The current temperature data is compared with a preset temperature threshold, and the current load data is compared with a preset load threshold. For example, the preset temperature threshold is set to 80°C, and the preset load threshold is set to 70%. If the current temperature exceeds 80°C, or the current load exceeds 70%, or multiple conditions in the two conditions are met at the same time, it is determined that the target computing unit meets the resource reallocation condition; otherwise, the condition is not met, and the current fan resource allocation amount is continued to be run.
[0072] S103, in the case where the target computing unit meets the resource reallocation condition, based on the temperature data and the load data at each time, the predicted temperature value of the target computing unit at the next time is determined, and based on the task scheduling queue of the target computing unit, the predicted load value of the target computing unit at the next time is determined.
[0073] In this embodiment, the task scheduling queue is a list recording the task information of the target computing unit currently to be processed and being processed. Each task in the queue usually contains task type, task size, estimated execution time and other attributes. The task scheduling queue reflects the load that the computing unit may bear in the future, and is an important basis for predicting the load value at the next time.
[0074] The predicted temperature value is an estimated value of the temperature that the computing unit may reach at the next time, which is obtained by using a specific prediction algorithm based on the historical temperature data and the historical load data of the target computing unit in the target time period. The predicted temperature value helps to prepare for heat dissipation in advance and avoid damage to the computing element caused by high temperature.
[0075] The predicted load value is an estimated value of the load that the computing unit will bear at the next time, which is obtained by analyzing the characteristics and execution order of the tasks in the current task scheduling queue of the target computing unit. The predicted load value can reflect the future work intensity of the computing unit and provide a basis for reasonable allocation of fan resources.
[0076] As an example, the server cleans the temperature data and load data collected in the target time period, removes outliers (such as excessively high or low values caused by sensor failure), and performs normalization processing to make the data distributed within a certain range, facilitating subsequent algorithm processing.
[0077] Then, a suitable prediction model is selected, such as a time series analysis model (such as an autoregressive integrated moving average model), a machine learning model (such as a support vector machine, a neural network, etc.). The historical temperature data and the historical load data are used as a training set to train the prediction model, and the parameters of the model are adjusted so that the model can better fit the trend of the historical temperature data and the historical load data.
[0078] Finally, the temperature data and load data of the current time are input into the trained prediction model to calculate the predicted temperature value of the next time.
[0079] Meanwhile, the server further analyzes the task scheduling queue of the target computing unit to obtain information such as the type, size, and predicted execution time of each task in the queue. According to the characteristics of the task, a suitable load prediction algorithm is adopted. For example, for tasks with fixed execution mode and resource demand, the historical execution time and resource occupation of the task can be used to predict its load contribution at the next time; for complex tasks, a prediction method based on task similarity can be used to match the current task with historical tasks and refer to the load of similar tasks to predict the load contribution of the current task. The predicted load contributions of each task are added to obtain the predicted load value of the target computing unit at the next time.
[0080] In S104, the current fan resource allocation of the target computing unit is adjusted based on the predicted temperature value and the predicted load value to obtain an updated fan resource allocation of the target computing unit.
[0081] In this embodiment, the current fan resource allocation is the fan speed, fan quantity, or fan power and other resource parameters related to heat dissipation allocated to the target computing unit at the current time, which determines the heat dissipation capacity of the fan.
[0082] The updated fan resource allocation is a new fan resource parameter obtained by adjusting the current fan resource allocation according to the predicted temperature value and the predicted load value. The updated fan resource allocation aims to more reasonably match the future heat dissipation needs of the computing unit and ensure stable operation of the computing unit at an appropriate temperature.
[0083] As an example, the server formulates a fan resource allocation strategy in advance, which can be a mapping relationship table to determine the corresponding fan resource allocation according to different combinations of the predicted temperature value and the predicted load value. For example, when the predicted temperature value is high and the predicted load value is large, a high fan speed or an increased fan quantity is allocated; when the predicted temperature value is low and the predicted load value is small, a low fan speed or a reduced fan quantity is allocated.
[0084] Then, the predicted temperature value and the predicted load value are input to calculate the updated fan resource allocation according to the fan resource allocation strategy. For example, if the predicted temperature value exceeds 85℃ and the predicted load value exceeds 80%, the fan speed is increased to 90% of the maximum speed; if the predicted temperature value is between 70℃ and 80℃ and the predicted load value is between 50% and 70%, the fan speed is set to 60% of the maximum speed.
[0085] The server adjusts the speed, quantity or power of the fan corresponding to the target computing unit according to the calculated updated fan resource allocation amount, so as to realize dynamic allocation of the fan resource and meet the future heat dissipation demand of the computing unit.
[0086] In the fan resource dynamic allocation method in the heterogeneous computing environment provided by the embodiment, the temperature data and the load data of the target computing unit at each time in the target time period are obtained first, so as to provide a basis for comprehensive analysis of the change trend; then, whether resource reallocation is needed is judged according to the current temperature data and the current load data, so as to avoid blind adjustment. When the resource reallocation condition is met, the temperature value at the next time is predicted by using the temperature data and the load data at each time in the target time period, and the load value at the next time is predicted by combining the task scheduling queue, so that the past data rule and the future task arrangement are comprehensively considered, and the prediction is more prospective. Finally, the current fan resource allocation amount is adjusted based on the two predicted values, and an updated reasonable fan resource allocation scheme is obtained. In this way, the application jumps out of the limitation of relying on the current data only, and the heat dissipation demand change of the computing unit is predicted in advance through historical tracing and prospective prediction, so that the fan resource is actively and accurately allocated, and the problem of untimely fan resource allocation is effectively avoided. Therefore, the computing unit can be cooled in time and sufficiently under various complex working conditions, and the heat dissipation effect can be improved.
[0087] In some schemes of the application, when the predicted temperature value of the target computing unit at the next time is determined based on the temperature data and the load data at each time, the accuracy of the predicted temperature value of the target computing unit at the next time is poor due to the lack of quantitative analysis of the temperature data and the load data.
[0088] To this end, as shown in Figure 2 The application further provides that S103 specifically can include the following S201 to S203:
[0089] S201, evaluating the invalidity degree of the current fan resource allocation amount of the target computing unit to temperature increase according to the temperature data and the load data at each time;
[0090] S202, determining the heat dissipation poor degree of the target computing unit based on the invalidity degree and the change trend of the temperature data;
[0091] S203, in the case that the heat dissipation poor degree is greater than or equal to a preset degree threshold, determining the predicted temperature value of the target computing unit at the next time based on the change trend of the temperature data and the change trend of the load data.
[0092] In the present embodiment, the invalidation degree of suppression is used to measure the control effect of the current fan resource allocation on temperature growth, i.e., whether the heat dissipation capability provided by the current fan can effectively slow down or stop the temperature rise of the target computing unit. If the fan's heat dissipation capability is insufficient, the temperature grows faster, and the invalidation degree of suppression is high; on the contrary, if the fan can effectively control the temperature, the invalidation degree of suppression is low.
[0093] The heat dissipation deficiency degree is a comprehensive evaluation of the invalidation degree of suppression of the current fan resource allocation on temperature growth and the trend of temperature data to assess the overall performance of the target computing unit's heat dissipation system. A high heat dissipation deficiency degree means that the target computing unit's heat dissipation system is facing serious problems, which may lead to excessive temperature affecting performance or even damaging hardware; a low heat dissipation deficiency degree indicates that the target computing unit's heat dissipation system is running well.
[0094] The preset degree threshold is a pre-set critical value of the heat dissipation deficiency degree, used to determine whether the heat dissipation deficiency degree meets the standard for taking further measures (i.e., predicting the temperature at the next time to adjust the fan resources). When the heat dissipation deficiency degree is greater than or equal to the preset degree threshold, it indicates that the heat dissipation problem is serious and needs to be addressed.
[0095] The trend of temperature data is obtained by analyzing the temperature data at each time within a period of time to determine whether the temperature is rising, falling, or stable, as well as the rate of rise or fall, which helps to predict the future temperature change.
[0096] The trend of load data is obtained by analyzing the load data at each time within a period of time to determine whether the load is increasing, decreasing, or remaining stable, as well as the rate of increase or decrease. The trend of load data is related to the trend of temperature data and jointly affects the temperature prediction at the next time.
[0097] As an example, the server collects the temperature data and load data of the target computing unit at each time from sensors and other devices, and uses a regression model in machine learning or a method based on a physical model. Taking the temperature data and load data as input features and the temperature growth as output target, the model is trained to learn the relationship between the temperature data, load data, and temperature growth under different fan resource allocations.
[0098] Then, the temperature data, load data, and current fan resource allocation at the current time are input into the trained model to obtain the difference between the expected temperature growth and the actual temperature growth under the current fan resource allocation. The larger the difference, the higher the invalidation degree of suppression of the current fan resource allocation on temperature growth; the smaller the difference, the lower the invalidation degree of suppression. For example, if the model predicts that the temperature should rise by 2°C per hour under the current fan resource allocation, but actually rises by 5°C per hour, the invalidation degree of suppression is high.
[0099] Then, according to experience or experimental data, a weight value is given to the inhibition invalidity, which reflects the importance of the inhibition invalidity in the evaluation of the heat dissipation poor degree. For example, if it is considered that the inhibition invalidity has a greater impact on the heat dissipation poor degree, a higher weight (such as 0.7) can be given. Then, the time series analysis method is used to analyze the temperature data to determine whether the temperature is rising, falling or stable, and the rate of rising or falling. If the temperature shows a rapid rising trend, a higher score (such as 0.8) can be given; if the temperature rises slowly, a lower score (such as 0.3) can be given; if the temperature is stable or falling, a score of 0 can be given. Then, the inhibition invalidity is multiplied by the corresponding weight, and the score of the temperature data trend is added to obtain the quantified value of the heat dissipation poor degree. For example, the inhibition invalidity score is 0.6, the weight is 0.7, and the temperature rises rapidly with a score of 0.8, then the heat dissipation poor degree = 0.6 x 0.7 + 0.8 = 1.22 (in addition, the calculation result can be normalized according to the actual situation to fall within a suitable range).
[0100] Then, the calculated heat dissipation poor degree is compared with the preset degree threshold value. If it is greater than or equal to the preset degree threshold value, the next step of prediction is entered; if it is less than the preset degree threshold value, it means that the heat dissipation condition is acceptable, and there is no need to make prediction adjustment. In the case of greater than or equal to the preset degree threshold value, the trend analysis is performed on the temperature data and the load data respectively. For the temperature data, a time series prediction model can be used to predict the temperature trend in the future period according to the change rule of the historical temperature data; for the load data, a similar time series analysis method can also be used to predict the future trend of the load.
[0101] The model of the prediction target calculation unit for the temperature value at the next time is established by taking the change trend of the temperature data and the load data as the input features. A regression model can be used to perform feature engineering processing (such as extracting the slope and volatility rate of the trend) on the change trend of the temperature data and the load data, and then input into the model for training. Finally, the change trend of the temperature data and the change trend of the load data at the current time are input into the trained prediction model to obtain the predicted temperature value of the target calculation unit at the next time.
[0102] Through this embodiment, the application can timely find the heat dissipation abnormality of the calculation unit and predict the future temperature trend. This helps to adjust the fan resource allocation in advance to avoid the problem of heat dissipation not in time caused by rapid temperature rise. At the same time, by considering the change trend of the temperature and the load, the prediction result is more accurate, and the fan resource can be allocated more accurately to improve the heat dissipation effect.
[0103] In some schemes of the application, when evaluating the invalidity degree of the current fan resource of the target computing unit in suppressing temperature growth, due to the lack of quantitative analysis of the dynamic correlation between temperature and load, the matching deviation between the fan resource allocation and the actual heat dissipation demand cannot be accurately identified.
[0104] To this end, as shown in Figure 3 , the application further proposes that S201 specifically can include the following S301 to S303:
[0105] S301, according to the temperature data at each time, a first temperature change sequence is constructed, and the first temperature change sequence includes the instantaneous change of the temperature data at each time within the target time period;
[0106] S302, based on the first temperature change sequence, the temperature growth rate of the target computing unit within the target time period is determined, and based on the load data at each time within the target time period, the load growth rate of the target computing unit within the target time period is determined;
[0107] S303, based on the temperature growth rate and the load growth rate, the invalidity degree of the current fan resource allocation of the target computing unit in suppressing temperature growth is determined.
[0108] In this embodiment, when constructing the first temperature change sequence, the instantaneous temperature change at each time within the target time period is extracted to form a time sequence for quantifying the temperature fluctuation trend. For example, the temperature data set within the target time period is , n represents that n times of temperature data are recorded, and the first temperature change sequence within the target time period is , wherein is used to represent the i-th instantaneous change in the first temperature change sequence. That is, in the construction process of the first temperature change sequence, each instantaneous change is calculated by the difference between the temperature data of adjacent time points.
[0109] The temperature growth rate can be determined by the following formula 1:
[0110] Formula 1
[0111] In formula 1, is used to represent the temperature growth rate of the target computing unit c within the target time period, is used to represent the n-1-th instantaneous change in the first temperature change sequence, is used to represent the first instantaneous change in the first temperature change sequence, and n is used to represent that n times of temperature data are recorded.
[0112] The load data set within the target time period is s represents that s times of load data are recorded together. Specifically, the load growth rate can be determined by the following formula 2:
[0113] Formula 2
[0114] In formula 2, for representing the load growth rate of the target computing unit c in the target time period, for representing the s-1th load data, for representing the 1st load data, and s for representing that s times of load data are recorded together.
[0115] As an example, the server constructs a first temperature change amount sequence according to the temperature data at each time point. The first temperature change amount sequence includes the instantaneous change amount of the temperature data at each time point in the target time period. For example, assuming that the target time period is 10 minutes and the temperature data is collected every minute, 10 temperature data points can be obtained. By calculating the difference between adjacent two temperature data points, 9 instantaneous change amounts can be obtained, which constitute the first temperature change amount sequence.
[0116] Then, based on the first temperature change amount sequence, the temperature growth rate of the target computing unit in the target time period is determined by the above formula 1. And based on the load data at each time point in the target time period, the load growth rate of the target computing unit in the target time period is determined by the above formula 2.
[0117] Finally, based on the temperature growth rate and the load growth rate, the invalidity degree of the current fan resource allocation amount of the target computing unit to the suppression of temperature growth is determined. Further, the temperature growth rate can be divided by the load growth rate to obtain a ratio. If the ratio is large, it means that the temperature growth speed is much faster than the load growth speed, indicating that the current fan resource allocation amount has poor suppression effect on temperature growth.
[0118] Through the embodiment, the invalidity degree of the current fan resource allocation amount of the target computing unit to the suppression of temperature growth can be accurately evaluated. Therefore, the situation of insufficient fan resource allocation can be found in time, providing a basis for subsequent dynamic adjustment of fan resources, and effectively avoiding the problem of poor heat dissipation effect caused by untimely fan resource allocation. Further, by combining the analysis of temperature data and load data, the heat dissipation effect can be more comprehensively evaluated, and the accuracy and rationality of fan resource allocation can be improved.
[0119] In some solutions of the present application, when evaluating the invalidity degree of the current fan resource allocation amount in inhibiting temperature growth by the temperature growth rate and the load growth rate, it is difficult to accurately quantify the deviation degree of the actual fan resource allocation and the theoretical demand by only relying on the ratio of the two, resulting in that the evaluation result of the invalidity degree deviates from the real heat dissipation demand.
[0120] To this end, the present application further proposes that S303 can specifically include:
[0121] The temperature growth rate is divided by the load growth rate to obtain a similarity degree ratio of the temperature growth rate and the load growth rate;
[0122] The current temperature data and the current load data of the target computing unit are matched with a preset curve of the fan resource allocation amount varying with temperature under different loads to obtain a first theoretical fan resource allocation amount of the target computing unit;
[0123] The invalidity degree of the current fan resource allocation amount of the target computing unit in inhibiting temperature growth is determined by using the similarity degree ratio and a target resource difference value of the target computing unit, and the target resource difference value is an absolute value of the difference between the current fan resource allocation amount and the first theoretical fan resource allocation amount of the target computing unit.
[0124] In the present embodiment, the similarity degree ratio reflects the dynamic correlation between temperature growth and load growth; the preset curve contains a mapping relationship of the fan resource allocation amount varying with temperature under different loads, and the theoretical allocation benchmark is determined by matching the current temperature data and the current load data; and the absolute value of the difference measures the deviation amplitude of the actual allocation and the theoretical allocation, and the calculation weight of the invalidity degree is adjusted in combination with the similarity degree ratio.
[0125] For example, as shown in Figure 4 A curve example diagram of the fan resource allocation amount varying with temperature under different loads is provided. In the diagram, the fan speed is taken as the fan resource allocation amount, the horizontal axis of the coordinate system is used to represent the temperature, and the vertical axis of the coordinate system is used to represent the fan speed.
[0126] Specifically, after obtaining the temperature growth rate and the load growth rate of the target computing unit in the target time period, the two are divided to obtain the similarity degree ratio, which is used to represent whether the temperature growth is synchronized with the load growth. Subsequently, the first theoretical fan resource allocation amount is found on the preset curve based on the current temperature data and the current load data. The absolute value of the difference between the current fan resource allocation amount and the first theoretical fan resource allocation amount is taken as the target resource difference value, and the target resource difference value and the similarity degree ratio are comprehensively considered to obtain the invalidity degree.
[0127] As an example, the server divides the temperature growth rate by the load growth rate to obtain a similarity ratio of the temperature growth rate to the load growth rate. For example, the temperature growth rate is 0.8, and the load growth rate is 0.5, so the similarity ratio is 1.6.
[0128] Then, the current temperature data and the current load data of the target computing unit are matched with the preset curve of the fan resource allocation amount changing with temperature under different loads to obtain a first theoretical fan resource allocation amount of the target computing unit. Specifically, an interpolation method can be used to find a point closest to the current temperature data and the current load data on the preset curve, and the corresponding fan resource allocation amount is the first theoretical fan resource allocation amount.
[0129] Finally, the similarity ratio and the target resource difference value of the target computing unit are used to determine the invalidity of the current fan resource allocation amount of the target computing unit in inhibiting temperature growth. The target resource difference value is the absolute value of the difference between the current fan resource allocation amount and the first theoretical fan resource allocation amount of the target computing unit.
[0130] Specifically, the invalidity of the current fan resource allocation amount of the target computing unit in inhibiting temperature growth can be determined by the following formula 3:
[0131] Formula 3
[0132] In formula 3, X is used to represent the invalidity of the current fan resource allocation amount in inhibiting temperature growth, is used to represent the target resource difference value, is used to represent the temperature growth rate of the target computing unit c in the target time period, is used to represent the load growth rate of the target computing unit c in the target time period.
[0133] Wherein, the greater the target resource difference value, the worse the current fan inhibits temperature change, and the higher the invalidity; the greater the similarity ratio, the greater the degree to which the temperature growth rate approaches the load growth rate, indicating that the temperature growth at this time depends more on load change than on fan resources, indicating that the current fan inhibits temperature change worse and the invalidity is higher.
[0134] Through the embodiment, the inhibitory effect of the current fan resource allocation amount on temperature growth can be accurately evaluated. Therefore, the fan resource allocation strategy can be adjusted in time to avoid the problem of overheating of the computing unit caused by continuous temperature rise. At the same time, by considering the correlation between temperature growth and load growth, the accuracy and rationality of fan resource allocation are improved, and unnecessary resource waste is reduced.
[0135] In some solutions of the present application, when the target computing unit is in the temperature continuous rising stage, the existing method may cause the temperature average increase calculation to be distorted due to the data of the temperature change amount sequence containing the temperature falling stage, affecting the evaluation accuracy of the heat dissipation poor degree, and further causing the fan resource adjustment to lag.
[0136] To this end, as shown in Figure 5 the present application further proposes that S202 can specifically include the following S501 to S503:
[0137] S501, from the first temperature change amount sequence of the target computing unit, extract the instantaneous change amount from back to front until the instantaneous change amount is negative, to obtain the second temperature change amount sequence of the target computing unit;
[0138] S502, performing mean value processing on each instantaneous change amount in the second temperature change amount sequence to obtain the temperature average increase of the target computing unit;
[0139] S503, determining the heat dissipation poor degree of the target computing unit based on the temperature average increase, the number of instantaneous change amounts in the second temperature change amount sequence, and the invalidity suppression degree.
[0140] In the present embodiment, the extraction logic of the second temperature change amount sequence is limited to reverse traversal of the temperature change data from the current time, ensuring that only the instantaneous change amount in the temperature continuous rising stage is retained. The calculation of the temperature average increase uses the arithmetic mean value method, which sums all elements in the second temperature change amount sequence and divides the element number.
[0141] Specifically, in the temperature monitoring process, when the temperature rise at consecutive time points is detected, the latest consecutive positive temperature change amount is extracted in reverse to form the second temperature change amount sequence. For example, when the temperature rise amplitudes at consecutive 5 time points are detected as -0.2℃, 0.5℃, 0.7℃, 0.6℃ and 0.8℃, the system automatically intercepts the last 4 positive change amounts. When calculating the temperature average increase, the 4 change amounts are added and divided by 4 to obtain a temperature average increase of 0.65℃.
[0142] As an example, the server extracts the instantaneous change amount from the first temperature change amount sequence of the target computing unit from back to front until the instantaneous change amount is negative, to obtain the second temperature change amount sequence of the target computing unit. For example, the first temperature change amount sequence is [0.5, 0.8, 1.2, -0.3, 0.6, 1.0], and the extracted second temperature change amount sequence is [0.6, 1.0].
[0143] The mean value of each instantaneous change in the second temperature change sequence is processed to obtain the temperature average increase of the target computing unit. Further, all instantaneous changes in the second temperature change sequence can be added and divided by the number of instantaneous changes to obtain the temperature average increase.
[0144] Finally, based on the temperature average increase, the number of instantaneous changes in the second temperature change sequence, and the invalidity degree of suppression, the heat dissipation poor degree of the target computing unit is determined. Specifically, the temperature average increase can be multiplied by the number of instantaneous changes to obtain a first calculation result; a preset coefficient value is subtracted by the invalidity degree of suppression to obtain a second calculation result; and the first calculation result is divided by the second calculation result to obtain the heat dissipation poor degree of the target computing unit. Thus, by analyzing the temperature change trend and the invalidity degree of suppression, the heat dissipation condition of the computing unit can be accurately evaluated to provide a basis for subsequent fan resource allocation.
[0145] Through the embodiment, the heat dissipation condition of the target computing unit can be accurately evaluated to avoid the problem of untimely fan resource allocation. By analyzing the temperature change trend and the invalidity degree of suppression, the temperature change can be predicted in advance, the fan resource can be adjusted in time, the heat dissipation effect can be improved, and the stable operation of the intelligent computing server can be ensured.
[0146] In some of the above schemes of the present application, a method of determining the heat dissipation poor degree based on the invalidity degree of suppression and the temperature data change trend is proposed. However, in the scenario of continuous temperature rise, the simple superposition of the invalidity degree of suppression and the temperature change trend may lead to deviation in the evaluation of the heat dissipation poor degree, which cannot accurately reflect the cumulative effect of the temperature increase and the comprehensive influence of the resource suppression failure, and further affect the adjustment accuracy of the subsequent fan resource allocation.
[0147] To this end, the present application further proposes that S503 specifically can include:
[0148] The temperature average increase is multiplied by the number of instantaneous changes to obtain a first calculation result;
[0149] A preset coefficient value is subtracted by the invalidity degree of suppression to obtain a second calculation result;
[0150] The first calculation result is divided by the second calculation result to obtain the heat dissipation poor degree of the target computing unit.
[0151] In the embodiment, the heat dissipation poor degree of the target computing unit can be determined by the following formula 4:
[0152] Formula 4
[0153] In formula 4, X is used to represent the degree of invalidity of the current fan resource allocation of the target computing unit in inhibiting temperature growth, and m is used to represent the number of instantaneous change amounts in the second temperature change amount sequence, is used to represent the temperature average increase amplitude of the second temperature change amount sequence.
[0154] wherein, is used to represent the degree of invalidity of the current fan resource allocation of the target computing unit in inhibiting temperature growth, and m is used to represent the number of instantaneous change amounts in the second temperature change amount sequence,
[0155] Through the embodiment, the heat dissipation poor degree of the computing unit can be more accurately evaluated, thereby providing a more reliable basis for subsequent fan resource allocation. This method considers the temperature change trend and the effect of the current fan resource allocation, can more timely discover heat dissipation problems, and is helpful to improve the response speed and efficiency of the heat dissipation system.
[0156] In some schemes of the application, when the predicted temperature value is determined by the change trend of the temperature data and the change trend of the load data, if the current temperature change amount cannot be accurately quantified, there may be a deviation between the predicted temperature value and the actual temperature value, so that the fan resource adjustment lags behind the temperature rising rate.
[0157] To this end, the application further proposes that S203 specifically can include:
[0158] determining a current predicted temperature change amount of the target computing unit based on the change trend of the temperature data and the change trend of the load data;
[0159] adding the current temperature data of the target computing unit to the current predicted temperature change amount to obtain a predicted temperature value of the target computing unit at a next time.
[0160] In the embodiment, the change trend of the temperature data is analyzed by a temperature difference value sequence of consecutive time points; and the change trend of the load data is analyzed by load data of consecutive time points.
[0161] As an example, the current predicted temperature change amount of the target computing unit can be specifically determined by the following formula 5:
[0162] Formula 5
[0163] In formula 5, is used to represent the predicted temperature change amount of the target computing unit at the nth time point, a current temperature data of the target computing unit at the n-1th moment, a load data at the n-1th moment, a load data at the n-1th moment.
[0164] wherein, a load change calculated according to the load data, a predicted value of the temperature change combined with the load effect, and .
[0165] Then, after determining the current predicted temperature change of the target computing unit by the above formula 5, the predicted temperature value of the target computing unit at the next moment can be obtained by adding the current predicted temperature change to the current temperature data of the target computing unit. For example, if the current temperature is 50℃ and the predicted temperature change is 2℃, the predicted temperature value at the next moment is 52℃.
[0166] Through the embodiment, the temperature value of the target computing unit at the next moment can be accurately predicted, which provides a basis for subsequent fan resource allocation. Therefore, the fan resource allocation can be adjusted in advance to avoid the problem of delayed heat dissipation caused by rapid temperature rise, thereby improving the heat dissipation effect and the operation stability of the intelligent computing server.
[0167] In some of the above schemes of the present application, it is proposed to determine whether the resource re-allocation condition is met based on the current temperature data and the current load data. However, in this process, only the instantaneous data at the current moment may not accurately reflect the long-term demand of resource allocation, especially when the load or temperature changes suddenly, the difference between the theoretical resource allocation demand and the actual allocation amount is not quantified, resulting in a lag in triggering the resource re-allocation condition.
[0168] To this end, the present application further proposes that S102 specifically can include:
[0169] matching the current temperature data and the current load data of the target computing unit with a curve of fan resource allocation amount changing with temperature under a preset different load to obtain a first theoretical fan resource allocation amount of the target computing unit;
[0170] in a case where the current fan resource allocation amount of the target computing unit is less than the first theoretical fan resource allocation amount, determining that the target computing unit meets the resource re-allocation condition.
[0171] In this embodiment, the preset curve of fan resource allocation amount changing with temperature under different loads is generated by fitting historical operation data. For example, when the load is 60% and the temperature is 65°C, the corresponding theoretical fan resource allocation amount is 2000 rpm. The matching process uses interpolation or least squares method to map the coordinate points of the current temperature data and the current load data to the curve, and outputs the corresponding theoretical allocation amount. The absolute value of the difference between the current fan resource allocation amount and the theoretical value is used to quantify the resource gap, and the resource redistribution condition is triggered when the gap exceeds the preset threshold.
[0172] Specifically, when the target computing unit is in an operating state with a load of 70% and a temperature of 68°C, the corresponding theoretical allocation amount is obtained by querying the preset curve, which is 2200 rpm. If the current actual allocation amount is 2000 rpm, the absolute value of the difference is 200 rpm. If the preset threshold is 150 rpm, the difference exceeds the threshold at this time, and it is determined that resource redistribution needs to be triggered. By introducing the theoretical allocation amount as a reference, the dynamic adjustment condition is converted from a single instantaneous data comparison to a quantitative comparison between the actual allocation amount and the theoretical demand, which identifies potential resource shortages in advance and avoids response delays caused by sudden data changes. For example, in the initial stage of the load surge to 75%, the actual allocation amount has not been adjusted, but the theoretical demand has risen to 2300 rpm, and the difference exceeds the threshold rapidly at this time, prompting the system to immediately start the redistribution process.
[0173] As an example, in the heterogeneous computing environment of the intelligent algorithm server, when performing dynamic allocation of fan resources, first, the current temperature data of computing unit A is collected in real time as 72°C, and the current load data of computing unit A is obtained as 85%. The above real-time data is input into the preset load-temperature-resource mapping database for matching, which stores the corresponding curves of temperature and theoretical fan speed under different load states calibrated by experiments. It is queried that the theoretical fan speed corresponding to the current load of 85% should be 3200 rpm, while the current actual running speed is 2800 rpm. By comparison, it is determined that the current actual speed is much lower than the theoretical speed, triggering the resource redistribution condition.
[0174] Through this embodiment, the problem of heat dissipation response lag caused by load mutation in the prior art is effectively solved. By matching the theoretical resource allocation curve with the measured operating parameters in real time, the adjustment mechanism is triggered as soon as the resource allocation gap is detected, avoiding the delay effect of the traditional method relying on historical data prediction. In this way, by predicting the resource gap in advance and actively intervening in the adjustment, the real-time synchronization of fan resource allocation and temperature change is ensured, thereby suppressing the abnormal rise of temperature and ensuring the stable operation of the computing unit under high load conditions.
[0175] In some solutions of the present application, the current fan resource allocation only relies on the current load data, and fails to consider the amount of tasks to be executed in the task scheduling queue at the next moment, resulting in insufficient accuracy of the predicted load value and inability to cope with the surge in heat dissipation demand caused by instantaneous load changes in advance.
[0176] To this end, the present application further provides that S103 specifically includes:
[0177] Obtaining the task computing amount of the target task in the task scheduling queue, the target task being a task to be executed by the target computing unit at the next moment;
[0178] Determining the predicted total task computing amount of the target computing unit at the next moment based on the task computing amount of the target task;
[0179] Determining the predicted load value of the target computing unit at the next moment using the predicted total task computing amount and the maximum computing capacity of the target computing unit.
[0180] In the present embodiment, the task scheduling queue contains target tasks to be executed at the next moment, and the predicted total task computing amount is obtained by extracting the task computing amount of these target tasks and accumulating them. The ratio of the predicted total task computing amount to the maximum computing capacity is taken as the predicted load value, which can be specifically calculated by the formula: predicted load value = (predicted total task computing amount / maximum computing capacity) x 100%. For example, the task scheduling queue contains two tasks with task computing amounts of 300 units and 500 units respectively, and the maximum computing capacity is 1000 units. The predicted total task computing amount is 800 units, and the predicted load value is 80%. The load at the next moment is directly calculated by the amount of tasks to be processed in the task scheduling queue, avoiding the lag of relying only on the current load data.
[0181] Specifically, after obtaining the target tasks in the task scheduling queue, the total computing amount to be processed at the next moment is obtained by accumulating the task computing amount of each target task. The total computing amount is compared with the maximum computing capacity of the target computing unit to convert it into a predicted load value in percentage form. For example, when the maximum computing capacity is 1000 units, a total computing amount of 800 units corresponds to an 80% load value, and a total computing amount of 1200 units corresponds to a 120% load value, at which point an overload warning needs to be triggered. The predicted load value and the temperature prediction value are used together as the basis for fan resource adjustment, so that the load change trend can be predicted in advance before the actual execution of the task, and the amount of fan resource allocation can be increased in advance. For example, when the predicted load value reaches 80%, the theoretical fan resource allocation amount is obtained according to the preset curve matching, and is dynamically adjusted in combination with the current invalidity suppression degree, effectively reducing the risk of temperature loss of control caused by sudden load surge.
[0182] As an example, after the computational load of the tasks to be executed in the task scheduling queue is obtained, the number of floating-point operations and data throughput contained in the image recognition task and matrix operation task to be processed at the next moment are extracted respectively. Assuming that the image recognition task needs to perform 1.2×10^6 floating-point operations and the matrix operation task needs to process 8×10^5 data units. After the computational scale of the two is accumulated, the total computational load of the predicted task is 2.0×10^6 units of computational load. Furthermore, the maximum computing power of the target computing unit is set to process 3.0×10^6 units of computational load per second. By multiplying the ratio of the total computational load of the predicted task to the maximum computing power by 100%, the predicted load value for the next moment is finally obtained as 66.7%.
[0183] This embodiment accurately predicts the future load state of the computing unit, enabling fan resource adjustments to match the upcoming computing intensity in advance. This effectively avoids the delayed cooling response caused by sudden load changes in traditional methods. By establishing a direct mapping between task computational effort and load values, the cooling system ensures resource adaptation before temperatures rise, reducing the risk of hardware overheating and maintaining the continuous and stable operation of the computing unit.
[0184] In some of the above-mentioned schemes of the present application, a method of dynamically adjusting the fan resource allocation based on the current temperature data and the current load data is proposed. However, when the operating load of the computing unit increases significantly or the temperature rises sharply, the theoretical fan resource allocation obtained based on the current data matching cannot accurately reflect the temperature and load change trends at subsequent moments, resulting in the problem of insufficient delay in the adjusted fan resource allocation, which may cause the risk of untimely heat dissipation.
[0185] In this regard, the present application further proposes that S104 specifically includes:
[0186] Matching the current temperature data and current load data of the target computing unit, as well as the predicted temperature value and predicted load value, with the curves of fan resource allocation under different preset loads as a function of temperature, to obtain a first theoretical fan resource allocation amount and a second theoretical fan resource allocation amount for the target computing unit;
[0187] An updated fan resource allocation amount of the target computing unit is determined by using the first theoretical fan resource allocation amount, the second theoretical fan resource allocation amount, and the degree of ineffectiveness of suppressing temperature growth of the current fan resource allocation amount of the target computing unit.
[0188] In the embodiment, the first theoretical fan resource allocation amount represents a theoretically required fan resource amount under the current temperature and load condition, and the second theoretical fan resource allocation amount represents a theoretically required fan resource amount under the predicted temperature and predicted load condition; and the invalidation degree of inhibition quantifies the deviation degree of the actual effect of the current fan resource allocation amount on temperature control from the theoretical effect.
[0189] As an example, the updated fan resource allocation amount of the target computing unit can be determined by the following formula 6:
[0190] Formula 6
[0191] In formula 6, is used to represent the updated fan resource allocation amount of the target computing unit, is used to represent the first theoretical fan resource allocation amount of the target computing unit, is used to represent the second theoretical fan resource allocation amount of the target computing unit, is used to represent the invalidation degree of inhibition of the current fan resource allocation amount on temperature increase.
[0192] Specifically, the first theoretical fan resource allocation amount and the second theoretical fan resource allocation amount are taken as references, and a corresponding weight is set based on the invalidation degree of inhibition X of the current fan resource allocation amount on temperature increase, so as to obtain the updated fan resource allocation amount of the target computing unit.
[0193] Through the embodiment, the problem of fan resource adjustment lag caused by load mutation in the prior art is effectively solved. By dynamically matching the theoretical resource amount under the current temperature and load condition and the theoretical resource amount under the predicted temperature and load condition, and combining the quantitative evaluation of the historical resource inhibition effect, the feedforward optimization of the fan resource allocation amount is realized. This adjustment mechanism based on multi-dimensional data fusion can significantly shorten the response delay of the heat dissipation system, maintain the thermal stability of the target computing unit in the temperature rapid rising stage, and thus guarantee the continuous and reliable operation of the intelligent computing server under the condition of sudden high load.
[0194] It should be noted that the above-mentioned sequence of the embodiments of the present application is only for description, and does not represent the advantages and disadvantages of the embodiments. The processes depicted in the drawings do not necessarily require the specific order or continuous order shown to achieve the desired results. In some embodiments, multi-task processing and parallel processing are also possible or can be advantageous.
[0195] Each embodiment in the specification is described in a progressive manner, and the same or similar parts between each embodiment can be referred to each other. Each embodiment focuses on the differences from other embodiments.
Claims
1. A method for dynamically allocating fan resources in a heterogeneous computing environment, characterized in that: The method comprises: Obtain temperature data and load data of the target computing unit at each moment within the target time period; determining whether the target computing unit meets a resource reallocation condition based on current temperature data and current load data of the target computing unit; If the target computing unit meets the resource reallocation condition, determining a predicted temperature value of the target computing unit at a next moment based on the temperature data and the load data at each moment, and determining a predicted load value of the target computing unit at a next moment based on the task scheduling queue of the target computing unit; Adjusting a current fan resource allocation amount of the target computing unit based on the predicted temperature value and the predicted load value to obtain an updated fan resource allocation amount of the target computing unit; The method for obtaining the predicted temperature value of the target computing unit at the next moment includes: evaluating, based on the temperature data and load data at each moment, the degree to which the current fan resource allocation of the target computing unit is ineffective in suppressing temperature growth; determining a degree of poor heat dissipation of the target computing unit based on the degree of ineffective suppression and a change trend of the temperature data; When the degree of poor heat dissipation is greater than or equal to a preset threshold, determining a predicted temperature value of the target computing unit at a next moment based on a change trend of the temperature data and a change trend of the load data; Methods for assessing the degree of suppression ineffectiveness include: constructing a first temperature variation sequence based on the temperature data at each moment, wherein the first temperature variation sequence includes instantaneous variations of the temperature data at each moment within the target time period; determining a temperature growth ratio of the target computing unit within the target time period based on the first temperature change sequence, and determining a load growth ratio of the target computing unit within the target time period based on load data at each moment within the target time period; Determining, based on the temperature growth ratio and the load growth ratio, the degree to which the current fan resource allocation amount of the target computing unit is ineffective in suppressing the temperature growth includes: Dividing the temperature increase ratio by the load increase ratio to obtain a similarity ratio between the temperature increase ratio and the load increase ratio; Matching the current temperature data and the current load data of the target computing unit with a preset curve of fan resource allocation amount versus temperature under different loads to obtain a first theoretical fan resource allocation amount for the target computing unit; The similarity ratio and the target resource difference of the target computing unit are used to determine the degree of ineffectiveness of the current fan resource allocation of the target computing unit in suppressing temperature growth. The target resource difference is the absolute value of the difference between the current fan resource allocation of the target computing unit and the first theoretical fan resource allocation.
2. The method for dynamic allocation of fan resources in a heterogeneous computing environment according to claim 1, characterized in that: The determining, based on the suppression ineffectiveness degree and the change trend of the temperature data, the degree of poor heat dissipation of the target computing unit includes: Extracting instantaneous changes from the first temperature change sequence of the target computing unit in sequence from back to front until the instantaneous changes are negative, thereby obtaining a second temperature change sequence of the target computing unit; performing mean processing on each of the instantaneous changes in the second temperature change sequence to obtain an average temperature increase of the target computing unit; The degree of poor heat dissipation of the target computing unit is determined based on the average temperature increase, the number of instantaneous temperature changes in the second temperature change sequence, and the degree of suppression ineffectiveness.
3. The method for dynamic allocation of fan resources in a heterogeneous computing environment according to claim 2, characterized in that: The determining the degree of poor heat dissipation of the target computing unit based on the average temperature increase, the number of instantaneous temperature changes in the second temperature change sequence, and the degree of suppression ineffectiveness includes: Multiplying the average temperature increase by the number of instantaneous changes to obtain a first calculation result; Subtracting the suppression ineffectiveness degree from the preset coefficient value to obtain a second calculation result; The first calculation result is divided by the second calculation result to obtain the degree of poor heat dissipation of the target computing unit.
4. The method for dynamic allocation of fan resources in a heterogeneous computing environment according to claim 1, characterized in that: The determining, based on the change trend of the temperature data and the change trend of the load data, a predicted temperature value of the target computing unit at a next moment, includes: determining a current predicted temperature change of the target computing unit based on a change trend of the temperature data and a change trend of the load data; The current temperature data of the target computing unit is added to the current predicted temperature change to obtain the predicted temperature value of the target computing unit at the next moment.
5. The method for dynamically allocating fan resources in a heterogeneous computing environment according to any one of claims 1 to 4, characterized in that: The determining, based on the current temperature data and the current load data of the target computing unit, whether the target computing unit meets the resource reallocation condition includes: Matching the current temperature data and the current load data of the target computing unit with a preset curve of fan resource allocation amount versus temperature under different loads to obtain a first theoretical fan resource allocation amount for the target computing unit; In a case where the current fan resource allocation amount of the target computing unit is less than the first theoretical fan resource allocation amount, it is determined that the target computing unit meets the resource reallocation condition.
6. The method for dynamically allocating fan resources in a heterogeneous computing environment according to any one of claims 1 to 4, characterized in that: The determining, based on the task scheduling queue of the target computing unit, a predicted load value of the target computing unit at a next moment, includes: Obtaining a task computation amount of a target task in the task scheduling queue, where the target task is a task that the target computing unit needs to execute at the next moment; Determining the total amount of predicted task calculations of the target computing unit at the next moment based on the task calculation amount of the target task; The predicted load value of the target computing unit at the next moment is determined by using the predicted task computing amount and the maximum computing capacity of the target computing unit.
7. The method for dynamically allocating fan resources in a heterogeneous computing environment according to any one of claims 1 to 4, characterized in that: The adjusting the current fan resource allocation amount of the target computing unit based on the predicted temperature value and the predicted load value to obtain an updated fan resource allocation amount of the target computing unit includes: Matching the current temperature data and the current load data of the target computing unit, as well as the predicted temperature value and the predicted load value, respectively, with a curve showing a change in fan resource allocation amount with temperature under different preset loads to obtain a first theoretical fan resource allocation amount and a second theoretical fan resource allocation amount for the target computing unit; An updated fan resource allocation amount of the target computing unit is determined by using the first theoretical fan resource allocation amount, the second theoretical fan resource allocation amount, and the degree of ineffectiveness of suppressing temperature growth of the current fan resource allocation amount of the target computing unit.
Citation Information
Patent Citations
Hadoop computing task initial allocation method based on load prediction
CN110262897A
Computing power resource processing method
CN118069380A