Server control method, control system and computer readable storage medium
By dynamically adjusting the heat dissipation power of the server and adjusting the power of the heat dissipation module according to the temperature and load volume, the problem of heat dissipation mismatch in different states of the server is solved, and stable control of the server temperature and efficient utilization of energy are achieved.
Patent Information
- Application Number
- CN202510486870.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-17
- Publication Date
- 2025-07-22
AI Technical Summary
In the prior art, the server uses a fixed heat dissipation power under different working conditions, resulting in energy waste or heat dissipation lag, and the heat dissipation power cannot be adjusted in time, affecting the server temperature stability.
The heat dissipation power of the heat dissipation module is dynamically adjusted according to the temperature of the server and the load to be processed, and the heat dissipation module is controlled to dissipate heat to the server by the largest of the first heat dissipation power and the second heat dissipation power to ensure that the temperature is within a reasonable range.
It realizes precise adjustment of server temperature, avoids lag in cooling power adjustment, reduces energy consumption, and ensures that the server operates within a reasonable temperature range.
Smart Images

Figure CN120353314A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of servers, and particularly to a control method, a control system, and a computer-readable storage medium for a server. Background Art
[0002] In a large data center, servers usually run continuously for 24 hours. When a server is running, the internal integrated circuits will continuously generate heat.
[0003] In the related art, after a server is started, a fixed heat dissipation power is used to dissipate heat from the server. However, different amounts of heat are generated by the server in different working states, resulting in changes in the temperature of the server. If a fixed heat dissipation power is used, it will cause the problem of excessive heat dissipation power when the server is running at low load or in a low-temperature state, resulting in energy waste, and the problem of insufficient heat dissipation power when the server is running at high load or in a high-temperature state, unable to dissipate heat in time. Summary of the Invention
[0004] This application provides a control method, a control system, and a computer-readable storage medium for a server, improving the heat dissipation effect of the server.
[0005] In a first aspect, this application provides a control method for a server, including: obtaining the temperature of the server in a server cluster and the amount of pending load of the server;
[0006] Determining a first heat dissipation power of a heat dissipation module according to the temperature of the server, and determining a second heat dissipation power of the heat dissipation module according to the amount of pending load of the server;
[0007] Controlling the heat dissipation module to dissipate heat from the server according to the maximum of the first heat dissipation power and the second heat dissipation power.
[0008] In a second aspect, this application further provides a control system for a server, including: a resource allocation module, a heat dissipation module, and a server cluster, where the server cluster includes at least one server;
[0009] Both the heat dissipation module and the servers in the server cluster are communicatively connected to the resource allocation module, and the heat dissipation module is used to dissipate heat from the servers in the server cluster;
[0010] The resource allocation module is configured to obtain the temperature of the server in the server cluster and the amount of pending load of the server, determine a first heat dissipation power of the heat dissipation module according to the temperature of the server, and determine a second heat dissipation power of the heat dissipation module according to the amount of pending load of the server, and control the heat dissipation module to dissipate heat from the server according to the maximum of the first heat dissipation power and the second heat dissipation power.
[0011] In a third aspect, the present application also provides a computer-readable storage medium storing a computer program, wherein when the computer program is executed by a processor, the steps of the control method of the server provided in the first aspect are implemented.
[0012] The control method of the server provided by the embodiments of the present application controls the heat dissipation module to dissipate heat from the server according to the maximum of the first heat dissipation power and the second heat dissipation power. When the first heat dissipation power is greater than the second heat dissipation power, it indicates that the amount of load to be processed is small, and the server will not generate too much heat in the future. However, the current temperature of the server is high, and the heat already generated has not been completely dissipated. Therefore, the heat dissipation module is controlled to dissipate heat from the server according to the first heat dissipation power. When the second heat dissipation power is greater than the first heat dissipation power, the heat dissipation module is controlled to dissipate heat from the server according to the second heat dissipation power, realizing pre-start of heat dissipation for the server and cooling the server in advance. Subsequently, the load to be processed is sent to the server, and the peak temperature of the server will be greatly reduced. Thus, the control method of the server can respond to the changes in the temperature of the server and the temperature to be processed, timely adjust the heat dissipation power of the heat dissipation module, avoid the problem of burning out the server due to lag in heat dissipation power adjustment, realize precise adjustment of the heat dissipation power, and ensure that the working temperature of the server is maintained within a reasonable temperature range. Description of the Drawings
[0013] To more clearly illustrate the embodiments of the present application, the drawings required for use in the embodiments will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0014] Figure 1 It is a schematic flowchart of a control method for a server provided by an embodiment of the present application;
[0015] Figure 2 It is a schematic layout diagram of a temperature detection device in a server provided by an embodiment of the present application;
[0016] Figure 3 It is a schematic diagram of a liquid cooling heat dissipation channel of a server provided by an embodiment of the present application;
[0017] Figure 4 It is a schematic structural diagram of a control system of a server provided by an embodiment of the present application. Detailed Embodiments
[0018] In order to more clearly understand the above objects, features, and advantages of the present application, the solutions of the present application will be further described below. It should be noted that, without conflict, the embodiments of the present application and the features in the embodiments can be combined with each other.
[0019] In the following description, many specific details are set forth in order to provide a thorough understanding of the present application. However, the present application may be implemented in other ways different from those described herein. Obviously, the embodiments in the specification are only a part of the embodiments of the present application, rather than all of the embodiments.
[0020] In order to enable those skilled in the art of the present technology to better understand the solution of the present application, the present application will be further described in detail below in conjunction with the accompanying drawings and specific embodiments.
[0021] In the related art, after the server is started, it will use a fixed heat dissipation power to dissipate heat from the server. However, the server generates different amounts of heat in different working states, resulting in changes in the temperature of the server. If a fixed heat dissipation power is used, it will cause the problem of excessive heat dissipation power when the server is running at low load or in a low-temperature state, resulting in energy waste, and the problem of insufficient heat dissipation power when the server is running at high load or in a high-temperature state, unable to dissipate heat in time.
[0022] To solve the above problems, an embodiment of the present application provides a control method for a server. Figure 1 It is a schematic flowchart of a control method for a server provided by an embodiment of the present application. As Figure 1 shown, the control method of the server includes:
[0023] S101. Obtain the temperature of the server in the server cluster and the amount of pending load of the server.
[0024] Exemplarily, the temperature of the server in the server cluster can be collected by a temperature detection device such as a temperature sensor, and the temperature of the server in the server cluster is sent to a resource allocation module, and the resource allocation module is, for example, a resource allocation server.
[0025] The amount of pending load of the server can include, for example, tasks that have been issued to the server but have not been processed by the server and tasks that are about to be issued to the server.
[0026] S102. Determine a first heat dissipation power of the heat dissipation module according to the temperature of the server, and determine a second heat dissipation power of the heat dissipation module according to the amount of pending load of the server.
[0027] During the operation of the server, the internal integrated circuit continuously dissipates heat. The higher the temperature of the server, the higher the heat dissipation power of the heat dissipation module is required to achieve a good heat dissipation effect. Using a heat dissipation module with a fixed power cannot achieve a good heat dissipation effect. Therefore, it is necessary to determine the first heat dissipation power of the heat dissipation module according to the temperature of the server, and the heat dissipation module uses the first heat dissipation power to dissipate heat from the server to improve the heat dissipation effect on the server.
[0028] As the amount of workload to be processed by the server increases, the heat generated by the server also increases. The cooling module uses a fixed power to cool the server. When the workload is large, the fixed-power cooling module cannot achieve good cooling effect, resulting in the problem of cooling lag. When the workload is small, the fixed-power cooling module leads to high energy consumption. Therefore, it is necessary to determine the second cooling power of the cooling module based on the workload to be processed by the server. By determining the second cooling power of the cooling module based on the workload to be processed by the server, pre-start of the cooling module can be realized, and the server can be cooled in advance, thereby reducing the peak temperature of the server and avoiding the problem of burning out the server due to the lag of cooling power adjustment.
[0029] Exemplarily, in the resource allocation module, for example, the correspondence between the temperature of the server and the first cooling power of the cooling module can be stored. When the temperature of the server is determined, the corresponding first cooling power can be matched by looking up the table.
[0030] In the resource allocation module, for example, the correspondence between the workload to be processed and the second cooling power of the cooling module can also be stored. When the workload to be processed by the server is determined, the corresponding second cooling power can be matched by looking up the table.
[0031] S103. Control the cooling module to cool the server according to the larger value of the first cooling power and the second cooling power.
[0032] Exemplarily, in response to the first cooling power being greater than the second cooling power, it indicates that the workload to be processed at this time is small, and the server will not generate too much heat subsequently, while the temperature of the server is high, and the heat already generated has not been completely dissipated. Therefore, control the cooling module to cool the server according to the first cooling power.
[0033] Exemplarily, in response to the second cooling power being greater than the first cooling power, it indicates that the workload to be processed by the server is large. If the cooling module is controlled to cool the server according to the first cooling power, the heat generated during the subsequent processing tasks of the server cannot be dissipated in time. Therefore, control the cooling module to cool the server according to the second cooling power, realize pre-start of cooling for the server, and then after the workload to be processed is sent to the server, the peak temperature of the server will be greatly reduced.
[0034] The control method of the server provided by the embodiment of the present application controls the heat dissipation module to dissipate heat from the server according to the maximum of the first heat dissipation power and the second heat dissipation power. When the first heat dissipation power is greater than the second heat dissipation power, it indicates that the amount of load to be processed is small, and the server will not generate too much heat subsequently. However, the current temperature of the server is relatively high, and the heat already generated has not been completely dissipated. Therefore, the heat dissipation module is controlled to dissipate heat from the server according to the first heat dissipation power. When the second heat dissipation power is greater than the first heat dissipation power, the heat dissipation module is controlled to dissipate heat from the server according to the second heat dissipation power, so as to realize pre-start of heat dissipation for the server, cool down the server in advance, and then send the load to be processed to the server. The peak temperature of the server will be greatly reduced. Thus, the control method of the server in the embodiment of the present application can respond to the changes in the temperature of the server and the temperature to be processed, timely adjust the heat dissipation power of the heat dissipation module, realize precise adjustment of the heat dissipation power, and ensure that the working temperature of the server is maintained within a reasonable temperature range.
[0035] Optionally, determining the first heat dissipation power of the heat dissipation module according to the temperature of the server includes: in response to the temperature of the server being less than the first temperature threshold, determining the power of the liquid cooling heat dissipation module in the heat dissipation module as the first liquid cooling heat dissipation power, and determining the power of the air cooling heat dissipation module in the heat dissipation module as 0; in response to the temperature of the server being greater than or equal to the first temperature threshold, determining the power of the liquid cooling heat dissipation module in the heat dissipation module as the second liquid cooling heat dissipation power based on the temperature range to which the temperature of the server belongs, and determining the power of the air cooling heat dissipation module in the heat dissipation module as the first air cooling heat dissipation power; wherein, both the second liquid cooling heat dissipation power and the first air cooling heat dissipation power are positively correlated with the temperature of the server.
[0036] In some embodiments, when it is detected that the temperature of the CPU (Central Processing Unit) is less than 50 °C, it indicates that the temperature is relatively low at this time, and there is no need for the heat dissipation module to use too high a heat dissipation power to dissipate heat from the server. Only the liquid cooling heat dissipation module with low energy consumption is enabled, and the power of the liquid cooling heat dissipation module is maintained at 50% of the rated power of the liquid cooling heat dissipation module, and the air cooling heat dissipation module is turned off, such as turning off the fan for heat dissipation, thereby reducing the heat dissipation energy consumption of the server.
[0037] When the detected CPU temperature is greater than or equal to 50°C and less than or equal to 70°C, the power of the liquid cooling module is maintained at 100% of the rated power of the liquid cooling module, and the air cooling power is 40% of the rated power of the air cooling module to assist in heat dissipation; when the detected CPU temperature is greater than or equal to 71°C and less than or equal to 85°C, the power of the liquid cooling module is maintained at 120% of the rated power of the liquid cooling module, and the air cooling power is 80% of the rated power of the air cooling module for rapid heat dissipation; when the detected CPU temperature is greater than or equal to 85°C, the power of the liquid cooling module is maintained at 120% of the rated power of the liquid cooling module, and the air cooling power is 100% of the rated power of the air cooling module for rapid heat dissipation.
[0038] In some other embodiments, when the detected temperature of the GPU (Graphics Processing Unit) is less than 60°C, it indicates that the temperature is relatively low at this time, and there is no need for the cooling module to use too high cooling power to cool the server. Only the liquid cooling module with lower energy consumption is enabled. The power of the liquid cooling module is maintained at 50% of the rated power of the liquid cooling module, and the air cooling module is turned off. For example, the fan cooling is turned off, thereby reducing the cooling energy consumption of the server.
[0039] When the detected GPU temperature is greater than or equal to 60°C and less than or equal to 75°C, the power of the liquid cooling module is maintained at 100% of the rated power of the liquid cooling module, and the air cooling power is 40% of the rated power of the air cooling module to assist in heat dissipation; when the detected GPU temperature is greater than or equal to 76°C and less than or equal to 95°C, the power of the liquid cooling module is maintained at 120% of the rated power of the liquid cooling module, and the air cooling power is 80% of the rated power of the air cooling module for rapid heat dissipation; when the detected GPU temperature is greater than or equal to 95°C, the power of the liquid cooling module is maintained at 120% of the rated power of the liquid cooling module, and the air cooling power is 100% of the rated power of the air cooling module for rapid heat dissipation.
[0040] In still some other embodiments, when the detected motherboard temperature of the server is less than 25°C, the air cooling module is turned off; when the detected motherboard temperature of the server is greater than or equal to 25°C and less than or equal to 35°C, the air cooling power is 40% of the rated power of the air cooling module for heat dissipation; when the detected motherboard temperature of the server is greater than or equal to 36°C and less than or equal to 45°C, the air cooling power is 80% of the rated power of the air cooling module for heat dissipation; when the detected motherboard temperature of the server is greater than 45°C, the air cooling power is 100% of the rated power of the air cooling module for heat dissipation.
[0041] It should be noted that the specific values of the first temperature threshold, temperature range, and heat dissipation power provided in the embodiments of the present application can be set according to the actual usage of the server, and the embodiments of the present application do not limit this.
[0042] Optionally, obtaining the temperature of the server in the server cluster includes: obtaining the motherboard temperature of the server and at least one component temperature; determining the first heat dissipation power of the heat dissipation module according to the temperature of the server includes: determining the first sub-heat dissipation power of the heat dissipation module according to the motherboard temperature of the server, and determining the second sub-heat dissipation power corresponding to the heat dissipation module one by one according to the temperatures of the components; controlling the heat dissipation module to dissipate heat from the server according to the maximum value of the first heat dissipation power and the second heat dissipation power includes: controlling the heat dissipation module to dissipate heat from the server according to the maximum value of the first sub-heat dissipation power, each second sub-heat dissipation power, and the second heat dissipation power.
[0043] Figure 2 It is a layout schematic diagram of a temperature detection device provided in an embodiment of the present application in a server, as Figure 2 shown. A plurality of first temperature detection devices 101 are embedded on the motherboard of the server to detect the temperature on the motherboard of the server, for example, to detect the heat generated by capacitors, resistors, components, etc. on the motherboard. An air-cooled heat dissipation module, such as a fan, is provided on the motherboard. A second temperature detection device 102 is also separately provided on the key components of the server to detect the temperature of the key components of the server. A liquid-cooled heat dissipation module, such as a cold plate liquid-cooled heat dissipation module, can be provided on the key components. The cold plate liquid-cooled heat dissipation module supports hot-swap maintenance, which is convenient for fault replacement. Figure 2 Taking the GPU inference server as an example, three first temperature detection devices 101 provided on the motherboard and second temperature detection devices 102 provided on the key components on the CPU1, CPU2, and GPU are shown. Exemplarily, the air-cooled heat dissipation module can be provided on the motherboard, and the liquid-cooled heat dissipation module can be provided in the CPU1, CPU2, and GPU.
[0044] Figure 3 It is a schematic diagram of a liquid-cooled heat dissipation channel of a server provided in an embodiment of the present application, as Figure 3As shown, the CPU1, CPU2, and GPU can all be set with the same water inlet and the same water outlet. CPU1 corresponds to the first liquid cooling channel, CPU2 corresponds to the second liquid cooling channel, and GPU corresponds to the third liquid cooling channel. The cooling medium enters through the main pipe of the water inlet, and then flows into the corresponding liquid cooling channels respectively, passing through the heat sinks of at least one of the devices of CPU1, CPU2, and GPU, taking away the heat of the devices, and finally converging into the main pipe of the water outlet to achieve heat interaction. Compared with setting different water inlets and water outlets for the cooling channels of the CPU and GPU respectively, in the embodiment of the present application, the CPU1, CPU2, and GPU are all set with the same water inlet and the same water outlet, making the cooling channels more concentrated and facilitating unified control.
[0045] Determine the first sub-cooling power of the cooling module according to the motherboard temperature of the server, determine the second sub-cooling power corresponding to the cooling module one by one according to the temperatures of each component, and control the cooling module to cool the server according to the maximum value among the first sub-cooling power, the second sub-cooling power, and the second cooling power, so as to improve the cooling efficiency of the server, avoid the problem that the temperature of the server is too high and the cooling power is too low, resulting in the server being unable to dissipate heat in time and affecting the normal operation of the server, and improve the cooling effect of the server.
[0046] Exemplarily, when the detected motherboard temperature of the server is 40 °C, the corresponding air-cooling power is 80% of the rated power of the air-cooling module; when the detected temperature of the CPU of the server is 60 °C, the corresponding liquid-cooling power is 100% of the rated power of the liquid-cooling module, and the air-cooling power is 40% of the rated power of the air-cooling module; when the detected load to be processed is 80 °C, the corresponding liquid-cooling power is 120% of the rated power of the liquid-cooling module, and the air-cooling power is 100% of the rated power of the air-cooling module. At this time, control the cooling module to cool the server according to the liquid-cooling power of 120% of the rated power of the liquid-cooling module and the air-cooling power of 100% of the rated power of the air-cooling module.
[0047] It should be noted that the specific values of the motherboard temperature and the cooling power of the server can be set according to the actual use situation of the server, and the embodiment of the present application does not limit this.
[0048] Optionally, determining the second heat dissipation power of the heat dissipation module according to the load to be processed by the server includes: in response to the load to be processed by the server being less than the load threshold, determining the power of the liquid cooling heat dissipation module in the heat dissipation module as the third liquid cooling heat dissipation power, and determining the power of the air cooling heat dissipation module in the heat dissipation module as 0; in response to the load to be processed by the server being greater than or equal to the load threshold, determining the power of the liquid cooling heat dissipation module in the heat dissipation module as the fourth liquid cooling heat dissipation power based on the load range to which the load to be processed by the server belongs, and determining the power of the air cooling heat dissipation module in the heat dissipation module as the second air cooling heat dissipation power; wherein, both the fourth liquid cooling heat dissipation power and the second air cooling heat dissipation power are positively correlated with the load to be processed by the server.
[0049] Exemplarily, when it is detected that the load to be processed by the server is less than 50%, it indicates that the load is relatively low at this time, and there is no need for the heat dissipation module to use too high heat dissipation power to dissipate heat from the server. Only the liquid cooling heat dissipation module with lower energy consumption is enabled, and the power of the liquid cooling heat dissipation module is maintained at 50% of the rated power of the liquid cooling heat dissipation module, and the air cooling heat dissipation module is turned off, for example, the fan heat dissipation is turned off, thereby reducing the heat dissipation energy consumption of the server.
[0050] When it is detected that the load to be processed by the server is greater than or equal to 50% and less than or equal to 80%, the power of the liquid cooling heat dissipation module is maintained at 100% of the rated power of the liquid cooling heat dissipation module, and the air cooling heat dissipation power is 40% of the rated power of the air cooling heat dissipation module to assist in heat dissipation; when it is detected that the load to be processed by the server is greater than 80%, the power of the liquid cooling heat dissipation module is maintained at 120% of the rated power of the liquid cooling heat dissipation module, and the air cooling heat dissipation power is 100% of the rated power of the air cooling heat dissipation module for rapid heat dissipation.
[0051] It can be understood that the air cooling heat dissipation module in the embodiments of the present disclosure can adjust its power by adjusting its rotation speed, and the rotation speed of the fan inside the server is linearly adjusted based on the load or temperature, gradually increasing or decreasing the rotation speed.
[0052] Optionally, obtaining the load to be processed by the server includes: determining a prediction task based on historical data; splitting the prediction task and the current task into multiple subtasks according to the task type; obtaining the processing progress of each server in the server cluster; according to the processing progress of each server in the server cluster, allocating the multiple subtasks to the corresponding servers; using the subtasks allocated to the server and the current load of the server as the load to be processed by the server.
[0053] Exemplarily, the prediction unit, such as the AI load prediction unit, uses an LSTM (Long Short-Term Memory) neural network to process time series data. The AI load prediction unit can capture the periodicity (such as daily traffic fluctuations) and suddenness (such as instant high-concurrency requests) of the load. The AI prediction unit refreshes historical data monthly, and statistically calculates data such as the utilization rate, I / O throughput, and network traffic of each server in the server cluster, so as to determine the prediction task. The task allocation unit is, for example, an FPGA (Field-Programmable Gate Array) chip card as the "traffic control center" of each server. The FPGA chip card is built with a multi-channel high-speed bus and externally connected to the server cluster in the computer room, including but not limited to servers such as CPU computing servers, GPU inference servers, memory storage pool servers, hard disk BOX storage pool servers, and Switch BOX switching pool servers, and is responsible for scheduling the work content of the servers in the server cluster.
[0054] The FPGA chip card simultaneously receives real-time current tasks. The FPGA chip card splits the prediction task and the current task into multiple subtasks, such as task types like query (such as database query), calculation and statistics (such as tax settlement), parsing (such as video analysis), and inference (such as AI inference). The calculation and statistics tasks are assigned to the CPU computing server cluster; the query tasks are assigned to the hard disk BOX storage pool server cluster; the parsing and inference tasks are assigned to the GPU inference server cluster, etc. The FPGA chip card performs resource allocation according to the load type predicted by the prediction unit. When periodic or sudden loads are predicted, the FPGA chip card statistics the load arrival duration and reports it to the processing unit in the resource allocation module.
[0055] The processing unit obtains the processing progress of each server in the server cluster. According to the processing progress of each server in the server cluster, it assigns multiple subtasks to the corresponding server. For example, for the calculation and statistics tasks, it obtains the processing progress of each CPU computing server in the server cluster, and according to the processing progress of each CPU computing server, it assigns the subtasks to the corresponding CPU computing server. The subtasks assigned to the CPU computing server and the current load of the CPU computing server are used as the workload to be processed by the CPU computing server.
[0056] In some embodiments, the prediction unit predicts that a sudden load will arrive in 2 minutes. The processing unit in the resource allocation service module counts the processing progress of each server in the current computing server cluster and finds that server A is in a dormant state or will complete the current task in 1 minute and 30 seconds. Then, the processing unit will not issue tasks to server A, and server A will handle the upcoming sudden load. If the sudden load does not arrive within 3 minutes, the processing unit will reissue tasks to server A or put it into a dormant state according to the allocation logic again.
[0057] Meanwhile, the FPGA chip card will control the components of the idle servers and put the idle components into the "deep sleep" state to reduce unnecessary energy consumption.
[0058] Thus, the processing path of the remaining tasks is automatically switched according to the processing progress of each computing server, realizing the dynamic allocation of energy consumption and resources, achieving the collaborative work of the servers, improving the working efficiency of the servers, and avoiding the problem of uneven server resource allocation leading to task backlog on the servers.
[0059] Optionally, the processing progress of the server distributes multiple subtasks to the corresponding servers, including: in response to the existence of at least one idle server in the server cluster, determining the remaining working duration of the servers in the working state in the server cluster; in response to the existence of a server in the working state and the remaining working duration corresponding to the processing progress being less than the preset duration, distributing the subtasks to the server in the working state and the remaining working duration corresponding to the processing progress being less than the preset duration; in response to the remaining working duration corresponding to the processing progress of the servers in the working state being greater than or equal to the preset duration, distributing the subtasks to the idle servers.
[0060] Exemplarily, when it is detected that there is at least one idle CPU computing server in the server cluster, the remaining working duration of the CPU computing servers in the working state in the server cluster is counted. The preset duration is, for example, 15s. If it is detected that there is a CPU computing server with a remaining duration less than 15s, it means that there is a CPU computing server with non-urgent work. Then, the subtasks are distributed to the server in the working state and the remaining working duration corresponding to the processing progress being less than 15s, and the idle servers are reserved to handle subsequent sudden loads.
[0061] If it is detected that the remaining working duration of the CPU computing servers in the working state is greater than or equal to 15s, it means that the task volume of the CPU computing servers in the working state is large. Then, the subtasks are distributed to the idle servers to disperse the task volume of each CPU computing server.
[0062] Exemplarily, the amount of tasks issued to the CPU computing server in the working state is 10% of the maximum load of the CPU computing server, and the amount of tasks of the issued subtasks is 30% of the maximum load of the CPU computing server. Then, the load to be processed by the CPU computing server is 40%, and the second cooling power is determined according to the 40% load to be processed.
[0063] Optionally, distributing multiple subtasks to corresponding servers according to the processing progress of each server includes: in response to the absence of idle servers in the server cluster, determining the remaining working duration corresponding to the processing progress of each server in the server cluster; and distributing the subtasks to the server with the shortest remaining working duration.
[0064] Exemplarily, when it is detected that there is no idle CPU computing server in the server cluster, determining the remaining working duration corresponding to the processing progress of each CPU computing server in the server cluster, and distributing the subtasks of the calculation and statistics type to the CPU computing server with the shortest remaining working duration, so as to make the distribution of the task amount more balanced.
[0065] It can be understood that when it is detected that there is an abnormal process in the server cluster, such as processing timeout or waiting for the next instruction, etc., or there is a CPU computing server with a relatively long current task processing time, for example, the single-task processing time exceeds 5 minutes, then the abnormal CPU computing server is removed and then the subtasks are distributed. The specific distribution logic can refer to the above embodiments, and the embodiments of the present application will not be elaborated here.
[0066] The embodiment of the present application also provides a control system for a server. Figure 4 As shown in the structural schematic diagram of the control system for a server provided by the embodiment of the present application, Figure 4 As shown, it includes: a resource allocation module 10, a heat dissipation module 20, and a server cluster 30. The server cluster 30 includes at least one server; both the heat dissipation module 20 and the servers in the server cluster are communicatively connected to the resource allocation module 10. The heat dissipation module 20 is used to dissipate heat from the servers in the server cluster; the resource allocation module 10 is configured to obtain the temperature of the servers in the server cluster and the load to be processed by the servers, determine the first cooling power of the heat dissipation module 20 according to the temperature of the servers, and determine the second cooling power of the heat dissipation module 20 according to the load to be processed by the servers, and control the heat dissipation module 20 to dissipate heat from the servers according to the maximum value of the first cooling power and the second cooling power.
[0067] Figure 4Exemplarily, it is shown in the figure that the server cluster 30 includes a CPU computing server, a GPU computing server, a memory storage pool server, and a hard disk storage pool server. In other embodiments, the server cluster 30 may also include other types of servers, which are not limited in the embodiments of the present application.
[0068] Exemplarily, the heat dissipation module 20 may include, for example, a liquid cooling heat dissipation control module 201, an air cooling heat dissipation control module 202, an air cooling heat dissipation module, and a liquid cooling heat dissipation module. The specific installation positions of the air cooling heat dissipation module and the liquid cooling heat dissipation module can be understood with reference to the positions provided in the above embodiments, Figure 4 which are not specifically shown in the figure. The resource allocation module 10 is configured to obtain the first heat dissipation power and the second heat dissipation power, and determine the heat dissipation power of the air cooling heat dissipation module and / or the liquid cooling heat dissipation module according to the maximum of the first heat dissipation power and the second heat dissipation power. The resource allocation module 10 is communicatively connected to the liquid cooling heat dissipation control module 201 and the air cooling heat dissipation control module 202 in the heat dissipation module 20. The resource allocation module 10 sends the heat dissipation power to the liquid cooling heat dissipation module through the liquid cooling heat dissipation control module 201, and sends the heat dissipation power to the air cooling heat dissipation module through the air cooling heat dissipation control module 202 to control the heat dissipation of the server by the air cooling heat dissipation module and / or the liquid cooling heat dissipation module.
[0069] Optionally, as Figure 4 shown, the resource allocation module 10 includes: a prediction unit 11, a task allocation unit 12, and a processing unit 13; both the processing unit 13 and the prediction unit 11 are communicatively connected to the task allocation unit 12; the prediction unit 11 is configured to determine a prediction task based on the historical data sent by the task allocation unit 12 and send the prediction task to the task allocation unit 12; the task allocation unit 12 is configured to split the prediction task and the current task into multiple subtasks according to the task type and send the multiple subtasks to the processing unit 13; the processing unit 13 is configured to obtain the processing progress of each server in the server cluster, allocate the multiple subtasks to the corresponding server according to the processing progress of each server, and use the subtasks allocated to the server and the current task amount of the server as the pending load amount of the server, and determine the second heat dissipation power of the heat dissipation module 20 according to the pending load amount of the server; the processing unit 13 is further configured to determine the first heat dissipation power of the heat dissipation module 20 according to the temperature of the server, and control the heat dissipation of the server by the heat dissipation module 20 according to the maximum of the first heat dissipation power and the second heat dissipation power.
[0070] Specifically, the prediction unit 11 includes, for example, an AI load prediction unit. The task allocation unit 12 is, for example, an FPGA chip card. The AI load prediction unit 11 calculates in real time the types of tasks split in the FPGA chip card and the tasks that have been processed and completed in the server cluster as historical data. The AI load prediction unit 11 determines prediction tasks based on the historical data. The FPGA chip card receives the current tasks and prediction tasks transmitted from the outside. The FPGA chip card can also statistically calculate the load status of the server cluster in real time, such as the GPU video memory occupancy rate, the queue depth of the SSD (Solid State Drive), etc., optimize the task processing path, and implement dynamic weight allocation. The processing unit 13 is also communicatively connected to each server in the server cluster, and is used to obtain the processing progress of each server in the server cluster. According to the processing progress of each server in the server cluster, a plurality of subtasks are allocated to the corresponding server. The processing unit 13 also takes the subtasks allocated to the server and the current load of the server as the to-be-processed load of the server, and determines the second heat dissipation power of the heat dissipation module 20 according to the to-be-processed load of the server. The processing unit 13 also obtains the temperature of the server, determines the first heat dissipation power of the heat dissipation module 20 according to the temperature of the server, and controls the heat dissipation module 20 to dissipate heat from the server according to the maximum value of the first heat dissipation power and the second heat dissipation power.
[0071] In some embodiments, the control system of the server further includes a cluster fault prediction module 40 and a cluster intrusion warning module 50. The cluster fault prediction module 40 is responsible for monitoring in real time the usage of components of the servers in the computing service cluster. The cluster fault prediction module 40 optimizes the load distribution weight in real time based on the fault point information of the components, the service life of the components, etc., and extends the service life of the components. For example, when the SSD life decays, its load distribution weight is automatically reduced, and when the GPU is continuously abnormally hot, its load distribution weight is automatically reduced.
[0072] The cluster intrusion warning module 50 is responsible for monitoring in real time the sensors of the servers in the computing service cluster. When it detects that the sensors of a server are abnormal, such as unauthorized removal of the hard disk, abnormal opening of the server cover, access of an untrusted storage device, etc., it immediately triggers the intrusion warning function, encrypts and migrates the data to the remaining secure server nodes. At the same time, the functions of the server are locked, and the abnormal events are combined with the computer room monitoring and blockchain records to achieve the traceability of the abnormal warning.
[0073] The embodiment of the present application also provides a computer-readable storage medium, in which a computer program is stored. Wherein, the computer program is configured to execute the steps in any one of the above embodiments of the control method of the server when running.
[0074] In an exemplary embodiment, the computer-readable storage medium may include, but is not limited to, various media capable of storing computer programs, such as USB flash drives, read-only memory (ROM), random access memory (RAM), external hard drives, magnetic disks, or optical discs.
[0075] An embodiment of the present application also provides a computer program product. The computer program product includes a computer program, and when the computer program is executed by a processor, it implements the steps in any of the above-described control method embodiments of the server.
[0076] Another embodiment of the present application also provides a computer program product, including a non-volatile computer-readable storage medium. The non-volatile computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, it implements the steps in any of the above-described control method embodiments of the server.
[0077] The control method, control system, and computer-readable storage medium of the server provided by the present application control the heat dissipation module to dissipate heat from the server according to the maximum of the first heat dissipation power and the second heat dissipation power. When the first heat dissipation power is greater than the second heat dissipation power, it indicates that the amount of load to be processed is small, and the server will not generate excessive heat in the future. However, the current temperature of the server is high, and the heat already generated has not been completely dissipated. Therefore, the heat dissipation module is controlled to dissipate heat from the server according to the first heat dissipation power. When the second heat dissipation power is greater than the first heat dissipation power, the heat dissipation module is controlled to dissipate heat from the server according to the second heat dissipation power, realizing pre-start of heat dissipation for the server and cooling the server in advance. Subsequently, the load to be processed is sent to the server, and the peak temperature of the server will be greatly reduced. Thus, the control method of the server in the embodiment of the present application can respond to the changes in the temperature of the server and the temperature to be processed, timely adjust the heat dissipation power of the heat dissipation module, achieve precise adjustment of the heat dissipation power, and ensure that the operating temperature of the server is maintained within a reasonable temperature range.
[0078] It should be noted that, in this text, relational terms such as "first" and "second" are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, such that a process, method, article or device comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the presence of additional identical elements in the process, method, article or device comprising the element.
[0079] The above are only specific embodiments of the present application, enabling those skilled in the art to understand or implement the present application. Various modifications to these embodiments will be obvious to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present application. Therefore, the present application will not be limited to these embodiments herein, but rather to the broadest scope consistent with the principles and novel features disclosed herein.
Claims
1. A control method for a server, characterized in that, Including: Obtaining the temperature of the servers in the server cluster and the amount of pending load of the servers; Determining a first heat dissipation power of the heat dissipation module according to the temperature of the servers, and determining a second heat dissipation power of the heat dissipation module according to the amount of pending load of the servers; Controlling the heat dissipation module to dissipate heat from the servers according to the maximum value of the first heat dissipation power and the second heat dissipation power.
2. The control method of the server according to claim 1, wherein The determining the first heat dissipation power of the heat dissipation module according to the temperature of the servers includes: In response to the temperature of the servers being less than a first temperature threshold, determining the power of the liquid cooling heat dissipation module in the heat dissipation module as a first liquid cooling heat dissipation power, and determining the power of the air cooling heat dissipation module in the heat dissipation module as 0; In response to the temperature of the servers being greater than or equal to the first temperature threshold, determining the power of the liquid cooling heat dissipation module in the heat dissipation module as a second liquid cooling heat dissipation power based on the temperature range to which the temperature of the servers belongs, and determining the power of the air cooling heat dissipation module in the heat dissipation module as a first air cooling heat dissipation power; Wherein, both the second liquid cooling heat dissipation power and the first air cooling heat dissipation power are positively correlated with the temperature of the servers.
3. The control method of the server according to claim 1, characterized in that, The obtaining the temperature of the servers in the server cluster includes: obtaining the main board temperature of the servers and at least one component temperature; The determining the first heat dissipation power of the heat dissipation module according to the temperature of the servers includes: determining a first sub-heat dissipation power of the heat dissipation module according to the main board temperature of the servers, and determining a second sub-heat dissipation power corresponding to the heat dissipation module one by one according to each component temperature; The controlling the heat dissipation module to dissipate heat from the servers according to the maximum value of the first heat dissipation power and the second heat dissipation power includes: controlling the heat dissipation module to dissipate heat from the servers according to the maximum value of the first sub-heat dissipation power, each second sub-heat dissipation power, and the second heat dissipation power.
4. The control method of the server according to claim 1, characterized in that The determining the second heat dissipation power of the heat dissipation module according to the amount of pending load of the servers includes: In response to the amount of pending load of the servers being less than a load threshold, determining the power of the liquid cooling heat dissipation module in the heat dissipation module as a third liquid cooling heat dissipation power, and determining the power of the air cooling heat dissipation module in the heat dissipation module as 0; In response to the amount of pending load of the servers being greater than or equal to the load threshold, determining the power of the liquid cooling heat dissipation module in the heat dissipation module as a fourth liquid cooling heat dissipation power based on the load range to which the amount of pending load of the servers belongs, and determining the power of the air cooling heat dissipation module in the heat dissipation module as a second air cooling heat dissipation power; Wherein, both the fourth liquid cooling heat dissipation power and the second air cooling heat dissipation power are positively correlated with the amount of pending load of the servers.
5. The control method of the server according to claim 1, characterized in that, The obtaining the amount of pending load of the servers includes: Determining a prediction task based on historical data; Splitting the prediction task and the current task into multiple subtasks according to the task type; Obtaining the processing progress of each server in the server cluster; According to the processing progress of each server in the server cluster, allocating the multiple subtasks to the corresponding servers; Taking the subtasks allocated to the server and the current load of the server as the amount of pending load of the server.
6. The control method of the server according to claim 5, wherein Allocating the multiple subtasks to corresponding servers according to the processing progress of each server includes: In response to at least one idle server existing in the server cluster, determining the remaining working duration of the servers in the server cluster that are in a working state; In response to a server in a working state whose remaining working duration corresponding to the processing progress is less than a preset duration, allocating the subtask to the server in a working state whose remaining working duration corresponding to the processing progress is less than the preset duration; In response to the remaining working durations corresponding to the processing progress of the servers in a working state all being greater than or equal to the preset duration, allocating the subtask to the idle server.
7. The control method of the server according to claim 5, characterized in that Allocating the multiple subtasks to corresponding servers according to the processing progress of each server includes: In response to no idle server existing in the server cluster, determining the remaining working duration corresponding to the processing progress of each server in the server cluster; Allocating the subtask to the server with the shortest remaining working duration.
8. A control system for a server, characterized in that, Including: A resource allocation module, a heat dissipation module, and a server cluster, where the server cluster includes at least one server; Both the heat dissipation module and the servers in the server cluster are communicatively connected to the resource allocation module, and the heat dissipation module is used to dissipate heat from the servers in the server cluster; The resource allocation module is configured to obtain the temperature of the servers in the server cluster and the pending workload of the servers, determine the first heat dissipation power of the heat dissipation module according to the temperature of the servers, and determine the second heat dissipation power of the heat dissipation module according to the pending workload of the servers, and control the heat dissipation module to dissipate heat from the servers according to the maximum of the first heat dissipation power and the second heat dissipation power.
9. The control system of the server according to claim 8, characterized in that, The resource allocation module includes: a prediction unit, a task allocation unit, and a processing unit; both the processing unit and the prediction unit are communicatively connected to the task allocation unit; The prediction unit is configured to determine a prediction task based on the historical data sent by the task allocation unit and send the prediction task to the task allocation unit; The task allocation unit is configured to split the prediction task and the current task into multiple subtasks according to the task type and send the multiple subtasks to the processing unit; The processing unit is configured to obtain the processing progress of each server in the server cluster, allocate the multiple subtasks to corresponding servers according to the processing progress of each server, and use the subtasks allocated to the server and the current task amount of the server as the pending workload of the server, and determine the second heat dissipation power of the heat dissipation module according to the pending workload of the server; The processing unit is further configured to determine the first heat dissipation power of the heat dissipation module according to the temperature of the server, and control the heat dissipation module to dissipate heat from the server according to the maximum of the first heat dissipation power and the second heat dissipation power.
10. A computer-readable storage medium, characterized in that, A computer program is stored in the computer-readable storage medium, where when the computer program is executed by a processor, the steps of the control method of the server according to any one of claims 1 to 7 are implemented.