Task scheduler device and program
The task scheduler device optimizes task assignments to reduce cooling inefficiencies and power consumption by distributing or aggregating tasks across computing devices, addressing suboptimal fan utilization in server cooling systems.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-09-18
- Publication Date
- 2026-03-26
AI Technical Summary
Existing server cooling systems experience inefficiencies due to unnecessary operation of cooling fans, leading to increased power consumption without considering the physical configuration and load distribution among computing devices, resulting in suboptimal fan utilization.
A task scheduler device that collects cooling status, determines inefficiencies, and adjusts task assignments to minimize fan output by distributing or aggregating tasks across computing devices, optimizing fan utilization based on thermal zones.
Reduces cooling inefficiencies and overall power consumption without affecting application performance by equalizing load and minimizing fan operation.
Smart Images

Figure JP2024033306_26032026_PF_FP_ABST
Abstract
Description
Task Scheduler Device and Program
[0001] The present invention relates to a task scheduler device and a program.
[0002] In vRAN (virtual Radio Access Network) - AI (Artificial Intelligence) inference technology, CPU (Central Processing Unit) operations are frequently used. For example, in applications that perform signal and media processing on a CPU (such as vRAN L1 signal processing and Deep-Learning), multiple computing devices are used for high-throughput computing processing. The multiple computing devices are computing accelerator devices (ACC: Accelerator) such as a CPU, an ASIC (Application Specific Integrated Circuit), an FPGA (Field Programmable Gate Array), and a GPU (Graphics Processing Unit).
[0003] In a server, a single cooling fan may cool multiple computing devices (cooling targets such as a CPU, GPU, ASIC, and FPGA).
[0004] FIG. 12 is a schematic configuration diagram of a server of a passive cooling system in which a single cooling fan cools multiple computing devices. The server 1 of the existing system shown in FIG. 12 includes hardware (HW: hardware) 10, an OS (Operating System) 20, a userland 30 having applications (APL#1, APL#2), and a fan controller 70.
[0005] The hardware 10 includes multiple computing devices (cooling targets such as a CPU, GPU, ASIC, and FPGA) 51 to 54 (computing device #1 to #4), a cooling fan (hereinafter referred to as a fan) 61 (fan #1) that cools the computing devices 51 and 52 by blowing air, and a fan 62 (fan #2) that cools the computing devices 53 and 54 by blowing air.
[0006] The computing devices 51 and 52 are cooled by a single fan 61 (fan #1), and the computing devices 51 and 52 and fan 61 (fan #1) constitute a thermal zone (TZ) 71. The computing devices 53 and 54 are cooled by a fan 62 (fan #2), and the computing devices 53 and 54 and fan 62 (fan #2) constitute a thermal zone (TZ) 72. In other words, thermal zone 71 is a collection of computing devices 51 and 52 cooled by a single fan 61 (fan #1), and thermal zone 72 is a collection of computing devices 53 and 54 cooled by a single fan 62 (fan #2). Note that the configuration of thermal zones 71 and 72 varies depending on the structure of the server 1's enclosure, etc.
[0007] OS20 includes device drivers, middleware, etc. OS20 does not assign computing devices considering the physical configuration.
[0008] The applications (APL#1, APL#2) are located in userland 30 and, via the OS 20, cause each computing device 51-54 (computing devices #1-#4) to execute tasks. In Figure 12, APL#1 issues a task to the OS 20, and the OS 20 assigns this task to computing device 51 (computing device #1) (arrow aa in Figure 12). Similarly, APL#2 issues a task to the OS 20, and the OS 20 assigns this task to computing device 54 (computing device #4) (arrow bb in Figure 12).
[0009] The fan controller 70 controls the rotation speed of fans 61 and 62 (fans #1 and #2) based on temperature information (code cc in Figure 12) from thermal zones 71 and 72. The fan controller 70 outputs a control signal (fan output) to control the rotation speed of fans 61 and 62 (fans #1 and #2) (code dd in Figure 12).
[0010] As shown by the symbol ee in Figure 12, the computing device 51 (computing device #1) housed in the thermal zone 71 is under heavy load due to the assignment from the OS 20 that received the task APL #1. To cope with this heavy load, computing device 51 (computing device #1) requires 80% fan output (referred to as 80% required fan output, and the same notation will be used hereafter). On the other hand, computing device 52 (computing device #2) housed in the same thermal zone 71 has no task assignment and requires 0% fan output. Therefore, in the thermal zone 71, fan 61 (fan #1) provides airflow (cooling) to computing devices 51 and 52 (computing devices #1 and #2) at 80% fan output (symbol ff in Figure 12).
[0011] Furthermore, as shown by the symbol gg in Figure 12, the computing device 54 (computing device #4) housed in the thermal zone 72 is under heavy load due to the assignment from the OS 20 that received the task APL #2. To cope with this heavy load, computing device 54 (computing device #4) is running at 80% of its required fan output. On the other hand, computing device 53 (computing device #3), also housed in the same thermal zone 72, has no task assignment and requires 0% fan output. Therefore, in the thermal zone 72, fan 62 (fan #2) provides airflow (cooling) to computing devices 53 and 54 (computing devices #3 and #4) at 80% fan output (symbol hh in Figure 12).
[0012] In the allocation pattern for server 1 of the existing system shown in Figure 12, two fans 61 and 62 (fans #1 and #2) are required for cooling.
[0013] Non-patent document 1 proposes a power saving method using existing task assignment. The technology in non-patent document 1 aims to reduce cooling power consumption by minimizing the temperature rise of the core and package.
[0014] F. Beneventi, A. Bartolini, C. Cavazzoni and L. Benini, "Cooling-Aware Node-level Task Allocation for Next-Generation Green HPC Systems", [online], [Retrieved August 1, 2020], Internet〈URL: https: / / ieeexplore.ieee.org / document / 7568402〉
[0015] In the existing system shown in Figure 12, Server 1 has a problem where the task assignment pattern causes unnecessary operation of the cooling fan, leading to increased power consumption. Hereafter, the occurrence of unnecessary operation of the cooling fan and the resulting increase in power consumption will be defined as "cooling inefficiency." Specific examples of cooling inefficiency will be described by dividing them into inefficiency pattern <1> and inefficiency pattern <2>.
[0016] - Inefficient Pattern <1> Figure 13 is a diagram illustrating the inefficient pattern <1> in server 1 of the existing system shown in Figure 12. The same reference numerals are used for components identical to those in Figure 12.
[0017] As shown by the symbol ii in Figure 13, the computing device 51 (computing device #1) housed in the thermal zone 71 is under heavy load due to the assignment of task 33 (for example, the assignment of task 33 from the OS 20 after receiving the task issuance of APL #1 in Figure 12) (symbol ee in Figure 13). To cope with this heavy load, the computing device 51 (computing device #1) is required to have 80% fan output. Therefore, in the thermal zone 71, the fan 61 (fan #1) performs airflow (cooling) at 80% fan output (white arrow ff in Figure 13).
[0018] Inefficient pattern <1>, in thermal zone 71, one computing device 51 (computing device #1) is at high output, while the other computing device 52 (computing device #2) has no task assigned and requires 0% fan output. Therefore, in thermal zone 71, it is possible to assign a new task to computing device 52 (computing device #2) without increasing the fan output of fan 61 (fan #1) (even if a load is placed on computing device #2, the fan output will not change). However, such control is not being performed.
[0019] As indicated by the symbol jj in Figure 13, the computing device 53 (computing device #3) housed in the thermal zone 72 is under low load due to the assignment of a low-load task 34 (for example, the assignment of a low-load task 34 from the OS 20 after receiving the task APL #2 in Figure 12). Because it is under low load, the computing device 53 (computing device #3) requires a fan output of 1%. On the other hand, the computing device 54 (computing device #4) housed in the same thermal zone 72 has no task assignment and requires a fan output of 0%. Therefore, in the thermal zone 72, the fan 62 (fan #2) performs airflow (cooling) at a fan output of 1% (white arrow kk in Figure 13).
[0020] Inefficient pattern <1> part 2 is a state in thermal zone 72 where fan 62 (fan #2) is operating at 1% fan output to meet the 1% fan output requirement of computing device 53 (computing device #3) processing the low-load task 34 (white arrow kk in Figure 13). Although fan 62 (fan #2) could be stopped by moving the low-load task 34 to another thermal zone 71, such control is not performed.
[0021] In summary, inefficiency pattern <1> part 1 does not consider a situation where a high-power computing device is always present in a single thermal zone, and new tasks can be assigned without increasing fan output. Inefficiency pattern <1> part 2 does not consider a situation where fans in multiple thermal zones are running for devices handling low-load tasks.
[0022] - Inefficient Pattern <2> Figure 14 is a diagram illustrating the inefficient pattern <2> in server 1 of the existing system shown in Figure 12. The same reference numerals are used for the same components as in Figure 12. Figure 15 is a diagram showing the relationship between fan speed and fan power consumption.
[0023] As shown by the symbol ll in Figure 14, the computing device 51 (computing device #1) housed in thermal zone 71 is under heavy load due to the assignment of task 35 (for example, the assignment of task 35 from OS 20 after receiving task issuance for APL #1 in Figure 12) (symbol mm in Figure 14). To cope with this heavy load, computing device 51 (computing device #1) is operating at 100% fan output. Therefore, in thermal zone 71, fan 61 (fan #1) is operating at 100% fan output for airflow (cooling) (white arrow nn in Figure 14). In addition, computing devices 53 and 54 (computing devices #3 and #4) housed in thermal zone 72 have no task assignments and are operating at 0% fan output (white dashed arrow oo in Figure 14).
[0024] Inefficiency pattern <2> occurs when the fan output is abnormally high, as shown by the symbol pp in Figure 14, and the power consumption is higher than when multiple fans are running to distribute the task. In Figure 14, the computing device 51 (computing device #1) is under heavy load and requires 100% fan output, so fan 61 (fan #1) operates at 100% fan output (symbol qq in Figure 14). With such 100% fan output, the following inefficiency pattern <2> occurs.
[0025] As shown in the shaded area rr of Figure 15, fan power consumption increases sharply as the fan speed increases (fan power consumption increases in proportion to the cube of the rotation speed). For this reason, in high-output conditions, the overall power consumption may be lower if multiple fans 61, 62 (fans #1, #2) are operated by distributing the fans rather than operating a single fan 61 (fan #1) in the high-speed range.
[0026] The technology described in Non-Patent Document 1 aims to reduce cooling power consumption by minimizing the temperature rise of the core and package. Although the technology in Non-Patent Document 1 does not result in performance degradation, there is room for further optimization as it does not take into account the efficiency of fan utilization. Furthermore, it cannot support multiple computing devices, and therefore does not satisfy the goal of minimizing power consumption for multiple devices.
[0027] In light of this background, the present invention was made, and its objective is to reduce cooling inefficiencies without affecting application performance and to reduce the overall power consumption of the server.
[0028] To solve the aforementioned problems, a task scheduler device for a server having a thermal zone consisting of one or more computing devices cooled by a single fan is provided, characterized by comprising: a cooling status collection unit that collects the usage status of the computing devices; an inefficiency determination unit that obtains the usage status of the computing devices from the cooling status collection unit and determines the thermal zone where cooling inefficiency is occurring based on the acquired information; and an assignment change instruction unit that instructs a change in task assignment to the computing devices based on the content of the inefficiency notified by the inefficiency determination unit.
[0029] According to the present invention, cooling inefficiencies can be reduced without affecting application performance, thereby reducing the overall power consumption of the server.
[0030] This is a schematic diagram of a server in a passive cooling system equipped with a task scheduler device according to an embodiment of the present invention. This is a block diagram showing the schematic configuration of the task scheduler device according to an embodiment of the present invention. This is a diagram showing an example configuration in which the task scheduler device according to an embodiment of the present invention is deployed on the OS. This is a flowchart showing the determination of distribution and aggregation of the task scheduler device according to this embodiment. This is a flowchart showing the process of assigning new tasks by aggregation of the task scheduler device according to this embodiment. This is a diagram illustrating the operation of the task scheduler device when the inefficiency of the inefficiency pattern <1> occurring in the passive cooling system of Figure 1 is resolved by "aggregation". This is a flowchart illustrating the distribution process of distributing tasks of the task scheduler device according to this embodiment to multiple computing devices. This is a diagram illustrating the inefficiency of the inefficiency pattern <2> of the task scheduler device according to this embodiment. This is a diagram illustrating the operation of the task scheduler device when the inefficiency of the inefficiency pattern <2> occurring in the passive cooling system of Figure 8 is resolved by ensuring that the heat generated by each CPU is even. This is a flowchart showing a modified example of the distribution process of distributing tasks of the task scheduler device according to this embodiment to multiple computing devices. This is a hardware configuration diagram showing an example of a computer that implements the functions of a task scheduler device in a computing system according to an embodiment of the present invention. This is a schematic configuration diagram of a server in a passive cooling system where a single cooling fan cools multiple computing devices. This is a diagram illustrating inefficiency pattern <1> in an existing system server shown in Figure 12. This is a diagram illustrating inefficiency pattern <2> in an existing system server shown in Figure 12. This is a diagram showing the relationship between fan rotation speed and fan power consumption.
[0031] The following describes a task scheduler device and the like in an embodiment of the present invention (hereinafter referred to as "this embodiment") with reference to the drawings. (Embodiment) [Overall Configuration] Figure 1 is a schematic configuration diagram of a server in a passive cooling system equipped with a task scheduler device 100 according to an embodiment of the present invention. The same reference numerals are used for the same components as in Figure 12. The server 1000 of the system shown in Figure 1 comprises hardware (HW) 10, OS 20, userland 30 having applications (APL#1, APL#2), and a fan controller 70.
[0032] The hardware 10 includes multiple computing devices (CPU, GPU, ASIC, FPGA, etc., which are to be cooled) 51 to 54 (computing device #1 to #4), a fan 61 (fan #1) that cools computing devices 51 and 52 by blowing air, and a fan 62 (fan #2) that cools computing devices 53 and 54 by blowing air.
[0033] The computing devices 51 and 52 are cooled by a single fan 61 (fan #1), and the computing devices 51 and 52 and fan 61 (fan #1) constitute a thermal zone (TZ) 71. The computing devices 53 and 54 are cooled by a fan 62 (fan #2), and the computing devices 53 and 54 and fan 62 (fan #2) constitute a thermal zone (TZ) 72. In other words, thermal zone 71 is a group of one or more computing devices 51 and 52 cooled by a single fan 61 (fan #1), and thermal zone 72 is a group of one or more computing devices 53 and 54 cooled by a single fan 62 (fan #2). Note that the form of thermal zones 71 and 72 varies depending on the structure of the server 1000 enclosure, etc.
[0034] OS20 includes device drivers, middleware, etc. OS20 does not perform computing device allocation that takes the physical configuration into account.
[0035] The applications (APL#1, APL#2) are located in userland 30 and, via the OS 20, cause each computing device 51-54 (computing devices #1-#4) to execute tasks. As shown in Figure 1, APL#1 issues a task to the OS 20, and the OS 20 assigns this task to computing device 51 (computing device #1) (arrow a in Figure 1). Similarly, APL#2 issues a task to the OS 20, and the OS 20 assigns this task to computing device 54 (computing device #4) (arrow b in Figure 1).
[0036] A task scheduler device 100 is located in userland 30. The task scheduler device 100 is a task assignment device that assigns tasks considering cooling efficiency. The task scheduler device 100 includes a cooling status collection unit 110, an inefficiency determination unit 120, and an assignment change instruction unit 130 (details are shown later in Figure 2). The cooling status collection unit 110 of the task scheduler device 100 receives device usage status and fan output from the OS 20.
[0037] The fan controller 70 controls the rotation speed of fans 61 and 62 (fans #1 and #2) based on temperature information (indicated as c in Figure 1) from thermal zones 71 and 72. The fan controller 70 outputs a control signal (fan output) to control the rotation speed of fans 61 and 62 (fans #1 and #2) (indicated as d in Figure 1).
[0038] As shown by the symbol e in Figure 1, the computing device 51 (computing device #1) housed in the thermal zone 71 is under heavy load due to an assignment from the OS 20 that received the task APL #1. To cope with this heavy load, computing device 51 (computing device #1) is cooled by airflow at 80% of the required fan output. On the other hand, computing device 52 (computing device #2), also housed in the same thermal zone 71, has no task assignment and requires 0% fan output. Therefore, in the thermal zone 71, fan 61 (fan #1) provides airflow (cooling) to computing devices 51 and 52 (computing devices #1 and #2) at 80% of the fan output (symbol f in Figure 1).
[0039] Furthermore, as shown by the symbol g in Figure 1, the computing device 54 (computing device #4) housed in the thermal zone 72 is under heavy load due to the assignment from the OS 20 that received the task APL #2. To cope with this heavy load, computing device 54 (computing device #4) is required to have 80% fan output. On the other hand, computing device 53 (computing device #3), also housed in the same thermal zone 72, has no task assignment and requires 0% fan output. Therefore, in the thermal zone 72, fan 62 (fan #2) provides airflow (cooling) to computing devices 53 and 54 (computing devices #3 and #4) at 80% fan output (symbol h in Figure 1).
[0040] In the allocation pattern of the existing system server 1000 shown in Figure 1, two fans 61 and 62 (fans #1 and #2) are required for cooling.
[0041] [Configuration and Prerequisites of Server 1000 in a Passive Cooling System] - Server 1000 in a passive cooling system, in which fans 61, 62 (cooling fans) and cooling targets are separated, comprises one or more pairs of fans 61, 62 whose output can be changed and specified, and one or more computing devices (cooling targets such as CPU, GPU, ASIC, FPGA, etc.) that are cooled by them. - Tasks run on computing devices 51 to 54, and the assignment of tasks to computing devices 51 to 54 can be changed via middleware or OS 20.
[0042] The requirements for assigning tasks in a passive cooling system server 1000, taking cooling efficiency into consideration, are as follows: • Requirement 1: (Minimizing power consumption of multiple computing devices 51-54) Minimize the overall power consumption of the server in a passive cooling system enclosure equipped with multiple fans that simultaneously cool multiple computing devices 51-54. • Requirement 2: (Performance impact) Avoid significant degradation in the performance (e.g., processing delay, throughput) of applications running on the system. • Requirement 3: (Eliminating cooling inefficiencies) Eliminate cooling inefficiencies caused by the positional relationship between fans 61, 62 and the devices being cooled.
[0043] [Task Scheduler Device 100] FIG. 2 is a block diagram showing a schematic configuration of a task scheduler device 100 according to an embodiment of the present invention. The task scheduler device 100 shown in FIG. 2 includes a cooling state collection unit 110, an inefficiency determination unit 120, and an allocation change instruction unit 130.
[0044] <Cooling State Collection Unit 110> The cooling state collection unit 110 collects the usage status of arithmetic devices (cooling targets such as CPUs, GPUs, ASICs, FPGAs, etc.) 51 to 54 (arithmetic devices #1 to #4) (FIG. 1). Specifically, the cooling state collection unit 110 uses various interfaces provided by the OS 20 (including middleware and device drivers) (FIG. 1) (for example, obtaining the CPU usage rate from the OS 20, obtaining the power consumption from the device driver of the GPU, etc.) to collect the usage status of the arithmetic devices 51 to 54 (arithmetic devices #1 to #4).
[0045] The cooling state collection unit 110 includes a thermal zone definition information storage unit 111, a sensor data acquisition processing unit 112, and a usage status collection processing unit 113.
[0046] The thermal zone definition information storage unit 111 stores thermal zone definition information for realizing task allocation. The thermal zone definition information holds cooling layout information (which arithmetic device is cooled by which fan).
[0047] The sensor data acquisition processing unit 112 signal-processes sensor information from each sensor (not shown) installed in the arithmetic devices 51 to 54 (FIG. 1) and the fans 61 and 62 (FIG. 1), and acquires the device temperature of the arithmetic devices 51 to 54, the power consumption of the arithmetic devices 51 to 54, and the fan rotation speeds of the fans 61 and 62 (FIG. 1). The device temperature of the arithmetic devices 51 to 54 is preferably the output of the sensors installed near each of the arithmetic devices 51 to 54, but may also be the output of the sensors installed in the thermal zones 71 and 72 (FIG. 1). Further, the power consumption of the arithmetic devices 51 to 54 may be obtained by calculation from the power sensor output of each of the arithmetic devices 51 to 54 or the detected current of each of the arithmetic devices 51 to 54. Also, the fan rotation speeds of the fans 61 and 62 (FIG. 1) are obtained by acquiring the control signal of the fan controller 70 (FIG. 1).
[0048] When the arithmetic devices 51 to 54 are CPUs, the usage status collection processing unit 113 collects the usage status from the usage rates of each CPU core and the like. When the arithmetic devices 51 to 54 are compute accelerator devices, the usage status is collected from the response time of the offload processing and the like.
[0049] <Inefficiency determination unit 120> The inefficiency determination unit 120 acquires the usage status of the arithmetic devices 51 to 54 from the cooling state collection unit 110, and based on the acquired information, detects (determines) the thermal zones 71 and 72 (FIG. 1) where inefficiencies occur, and notifies the allocation change instruction unit 130. Specifically, the inefficiency determination unit 120 refers to the cooling layout information (thermal zone definition information) held by the thermal zone definition information storage unit 111, detects an improvable cooling inefficiency, and notifies the allocation change instruction unit 130 of an allocation change instruction. Here, the method for eliminating the inefficiency location includes dispersion (FIG. 7 etc.) and aggregation (FIG. 5 etc.). When aggregating, the inefficiency determination unit 120 determines a plurality of thermal zones in which only a single arithmetic device is in a predetermined high-load state in a plurality of thermal zones.
[0050] When dispersing, the inefficiency determination unit 120 determines the arithmetic devices in a predetermined high-load state within the thermal zones 71 and 72.
[0051] <Assignment Change Instruction Unit 130> Based on the details of the inefficiency notified by the inefficiency determination unit 120, the assignment change instruction unit 130 instructs the OS 20 and middleware (not shown) to change the task assignment to the computing devices 51 to 54.
[0052] When consolidating, the assignment change instruction unit 130 assigns tasks from computing devices that are in a high-load state to computing devices belonging to fewer thermal zones. When distributing, the assignment change instruction unit 130 distributes and assigns tasks from computing devices that are in a high-load state to multiple computing devices. Here, the assignment change instruction unit 130 also uses the thermal zone definition information from the thermal zone definition information storage unit 111 of the cooling state collection unit 110 as needed. The assignment change instruction unit 130 can use the thermal zone definition information without waiting for instructions from the inefficiency determination unit 120, thereby shortening the response time.
[0053] [Example of Task Scheduler Device Application] Figure 3 shows an example configuration in which the task scheduler device 100 is placed on the OS 20. The same reference numerals are used for the same components as in Figure 1. In Figure 1, the task scheduler device 100 is placed on the userland 30, but the task scheduler device 100 may also be placed on the middleware of the OS 20.
[0054] The operation of the task scheduler device 100 configured as described above will be explained below. The task scheduler device 100 performs "distribution" (Figure 7, etc.) to eliminate inefficient parts of inefficient pattern <1> (Figure 13), or performs "aggregation" (Figure 5, etc.) to eliminate inefficient parts of inefficient pattern <2> (Figure 14).
[0055] <Basic Concept> The task scheduler device 100 distributes tasks as much as possible to distribute them, thereby equalizing the load and reducing fan output. Furthermore, for tasks that require the continuous occupation of the same resource, which cannot be distributed, the task scheduler device 100 concentrates them into the fewest possible thermal zones to minimize the number of operating fans.
[0056] <Decision on Distribution / Aggregation> Figure 4 is a flowchart showing the decision on distribution / aggregation. In step S1, the task scheduler device 100 (Figures 1 and 3) determines whether it is possible to move or split the task. If it is possible to move or split the task (S1: Yes), it proceeds to the [Distribution] process in step S2. If it is not possible to move or split the task (S1: No), it proceeds to the [Aggregation] process in step S3. The task scheduler device 100 basically selects distribution first. In addition, for tasks that cannot be moved or that would be time-consuming to move, the task scheduler device 100 reduces the number of operating fans by aggregating them.
[0057] In the [distribution] process of step S2, the inefficiency determination unit 120 completes the processing of this flow by distributing the task across multiple devices. The [distribution] process suppresses the temperature rise of one device and reduces the required fan output.
[0058] In step S3, the inefficiency determination unit 120 completes the processing of this flow by consolidating the tasks into the fewest possible number of thermal zones 71, 72 (Figures 1 and 3) so that the number of operating fans is minimized. Tasks that require the continuous occupation of the same resource cannot be distributed, so they are consolidated into the fewest possible number of thermal zones 71, 72 so that the number of operating fans is minimized. Here, tasks that require the continuous occupation of a specific computing device cannot be distributed, so the number of operating fans is reduced by consolidation.
[0059] <Aggregation> The process of assigning new tasks by aggregation will be explained with reference to Figures 1, 5, and 6. Figure 5 is a flowchart of the process of assigning new tasks by aggregation. In step S11, the task scheduler device 100 (Figures 1 and 3) determines whether or not there are thermal zones 71 and 72 (Figures 1 and 3) with fans in operation. If there are thermal zones 71 and 72 with fans in operation (S11: Yes), in step S12 the task scheduler device 100 (Figures 1 and 3) determines whether or not there are available devices. If there are available devices, in step S13 the inefficiency determination unit 120 assigns a task to an available device in the thermal zone in use and completes the processing of this flow.
[0060] If there are no thermal zones 71, 72 with fans running in step S11 (S11: No), or if there are no available devices in step S11 (S12: No), the inefficiency determination unit 120 assigns a computing device to an unused thermal zone to increase the number of running fans and completes the processing of this flow.
[0061] <Explanation of Aggregation Operation> Figure 6 is a diagram illustrating the operation of the task scheduler device 100 when the inefficient parts of the inefficient pattern <1> that occur in Figure 1 are resolved by "aggregation". In the state before the control operation of the task scheduler device 100 shown in Figure 1, two fans 61 and 62 (fans #1 and #2) are required: fan 61 (fan #1) for 80% of the required fan output of the computing device 51 (computing device #1) housed in the thermal zone 71, and fan 62 (fan #2) for 80% of the required fan output of the computing device 54 (computing device #4) housed in the thermal zone 72.
[0062] In Figure 1, the computing device 51 (computing device #1) housed in thermal zone 71 and the computing device 54 (computing device #4) housed in thermal zone 72 correspond to a single computing device in a predetermined high-load state.
[0063] When the flow in Figure 4 is executed and [Aggregate] is selected, and the new task assignment process by aggregation shown in Figure 5 is executed, the task scheduler device 100 performs the operation shown in Figure 6. As shown in Figure 6, the usage status collection processing unit 113 (Figure 2) of the cooling status collection unit 110 collects the usage status of the computing devices 51 to 54 (computing devices #1 to #4) using the CPU usage rate and other data from the OS 20. In addition, the sensor data acquisition processing unit 112 (Figure 2) of the cooling status collection unit 110 acquires the device temperature of computing devices 51 to 54, the power consumption of computing devices 51 to 54, and the fan rotation speed of fans 61 and 62 (Figure 6).
[0064] The inefficiency determination unit 120 of the task scheduler device 100 acquires the usage status of the computing devices 51 to 54 from the cooling status collection unit 110, and based on the acquired information, it refers to the cooling layout information (thermal zone definition information) held by the thermal zone definition information storage unit 111 (Figure 2) to detect (determine) the thermal zone 72 (Figure 6) where inefficiency is occurring, and notifies the assignment change instruction unit 130.
[0065] In Figure 6, the inefficiency determination unit 120 determines that the computing device 51 (computing device #1) housed in thermal zone 71 and the computing device 54 (computing device #4) housed in thermal zone 72 are a single computing device in a predetermined high-load state. In this case, the single computing device 51, 54 in a high-load state belongs to thermal zones 71 and 72. The inefficiency determination unit 120 detects (determines) the thermal zone 72 (Figure 6) where inefficiency is occurring and concentrates the computing devices in thermal zone 71 (Figure 6) so as not to use the fan 62 (Figure 6) of thermal zone 72. Alternatively, the computing devices in thermal zone 71 may be concentrated in thermal zone 72 (Figure 6) so as not to use the fan 61 (Figure 6) of thermal zone 71.
[0066] The assignment change instruction unit 130 assigns the tasks of the computing device 54 (computing device #4), which has been determined to be in a high-load state, to computing devices 52 (computing device #2) that belong to fewer thermal zones 71.
[0067] Specifically, the assignment change instruction unit 130 instructs the OS 20 and middleware to change the task assignment to the computing devices 52 and 54 (computing devices #2 and #4) based on the content of the inefficiency notified by the inefficiency determination unit 120 (bold arrow j in Figure 6). Upon receiving the resource assignment change instruction from the assignment change instruction unit 130 to the computing devices 52 and 54 (computing devices #2 and #4), the OS 20 and middleware change the task issuance of APL#2 (arrow b in Figure 6) from computing device 54 (computing device #4) to computing device 52 (computing device #2) (indicated by k in Figure 6). The OS 20 and middleware then assign the changed task to computing device 52 (computing device #2) (arrow l in Figure 6). Since the task issuance for APL#2 (arrow b in Figure 6) has been changed to arithmetic device 52 (arithmetic device #2), there will be no task assignment to arithmetic device 54 (arithmetic device #4) (dashed arrow in Figure 6).
[0068] In Figure 6, the fan controller 70 outputs a fan output change instruction to fan 62 (fan #2), reducing the output of unnecessary fans (in this case, the fan output of fan #2 is set to 0%).
[0069] Fan 61 (fan #1) in thermal zone 71 (Figure 6) provides airflow (cooling) at 80% fan output to both computing devices 51 (computing device #1) and computing device 52 (computing device #2), which has been newly assigned a task and started operating (white arrow m in Figure 6). As shown in Figure 1, in thermal zone 71, fan 61 (fan #1) was originally providing airflow (cooling) at 80% fan output to computing device 51 (computing device #1), so the fan output of 80% remains unchanged and continues.
[0070] On the other hand, the fan 62 (fan #2) in thermal zone 72 (Figure 6) has its fan output reduced to 0% (fan #2 stops) (dashed white arrow n in Figure 6). As a result, cooling is performed only by the fan 61 (fan #1) in thermal zone 71 (Figure 6), minimizing cooling power consumption.
[0071] <Distribution> Referring to Figures 7, 8, and 9, task assignment, which distributes tasks across multiple computing devices, will be explained. When computing devices 51 to 54 are CPUs, the task scheduler device 100 periodically acquires tasks that run on the CPUs and targets tasks with high loads for control. Next, it changes the combination of CPU cores used by the task in question so that the heat generated by each CPU is evenly distributed. However, if the degree of parallelism of the task in question is small and it is not possible to distribute the task across all CPUs, the CPUs are allocated in such a way that the number of thermal zones used is minimized, and the number of operating fans is minimized.
[0072] Figure 7 is a flowchart illustrating distributed processing in which tasks are distributed across multiple computing devices. As shown in the connector in step S20, the task scheduler device 100 (Figures 8 and 9) periodically acquires tasks to be run on the CPU (the processing result of S24 is input to the connector in S20). In step S21, the task scheduler device 100 (Figures 1 and 3) determines whether or not it has detected a task to be controlled. If it does not detect a task to be controlled (S21: No), it terminates the processing of this flow.
[0073] If a task to be controlled is detected (S21: Yes), in step S22, the cooling state collection unit 110 of the task scheduler device 100 collects task and hardware information such as the CPU usage rate and CPU temperature of the task to be controlled.
[0074] In step S23, the inefficiency determination unit 120 of the task scheduler device 100 performs a task transfer destination calculation to allocate tasks so that the heat generated by each CPU is evenly distributed.
[0075] Here, if the degree of parallelism when allocating tasks so that the heat generated by each CPU is even is less than or equal to the number of CPUs in the thermal zone, the tasks are distributed to each CPU in a single thermal zone (example of operation in Figure 6).
[0076] Otherwise, increase the number of operating fans and distribute them as evenly as possible across multiple thermal zones. Variations of the allocation method include: - Using the same number of cores for each CPU. - Reducing the number of cores used by CPUs with high temperatures and moving them to CPU cores with lower temperatures. - Using dynamic programming based on CPU power consumption and the utilization rate of each core, predict the heat generated by each worker thread of a task and arrange each thread so that the heat generated by each CPU is uniform.
[0077] In step S24, the assignment change instruction unit 130 of the task scheduler device 100 issues an assignment change instruction to the OS 20, specifying the CPU core to be used by the task, and then returns to step S20.
[0078] In the task destination calculation in step S23 described above, there are the following variations. Specifically, as a variation of the task destination calculation that determines which CPU to use, there is a method in which the task is actually moved while measuring the temperature of each CPU to equalize the heat generation.
[0079] <Explanation of Distributed Operation> Figures 8 and 9 are explanatory diagrams for distributing tasks across multiple computing devices. Figure 8 is a diagram illustrating the inefficiency of inefficient pattern <2>, and Figure 9 is a diagram illustrating the operation of the task scheduler device 100 when the inefficiency of inefficient pattern <2> in Figure 8 is resolved by ensuring that the heat generated by each CPU is equal. In the state before the control operation of the task scheduler device 100 shown in Figure 8, the fan 61 (fan #1) is operating at 100% fan output because the computing device 51 (computing device #1) housed in the thermal zone 71 requires 100% fan output. At such 100% fan output, the fan speed is high, and inefficient pattern <2> occurs where fan power consumption rises sharply (Figure 15). In Figure 8, the computing device 51 (computing device #1) housed in the thermal zone 71 corresponds to a computing device in a predetermined high-load state.
[0080] When the [Distributed] option is selected during the execution of the flow in Figure 4, and a distributed process is executed to distribute the tasks shown in Figure 7 to multiple computing devices, the task scheduler device 100 performs the operation shown in Figure 9. As shown in Figure 9, the usage status collection processing unit 113 (Figure 2) of the cooling status collection unit 110 collects the usage status of computing devices 51 to 54 (computing devices #1 to #4) using the CPU usage rate and other data from the OS 20. In addition, the sensor data acquisition processing unit 112 (Figure 2) of the cooling status collection unit 110 acquires the device temperature of computing devices 51 to 54, the power consumption of computing devices 51 to 54, and the fan speed of fans 61 and 62 (Figure 8).
[0081] The inefficiency determination unit 120 of the task scheduler device 100 collects the usage status of computing devices 51 to 54 from the cooling state collection unit 110, and based on the collected information, detects (determines) a thermal zone 71 (Figure 8) where a computing device 51 that requires a high fan speed exists and inefficiency occurs due to the high fan speed, and notifies the assignment change instruction unit 130.
[0082] In Figure 9, the inefficiency determination unit 120 reduces the fan speed of the fan 61 (Figure 9) in the thermal zone 71 (Figure 8), where the fan speed is high and inefficiency is occurring, by distributing the tasks of the computing device in the thermal zone 71 to multiple computing devices.
[0083] The assignment change instruction unit 130 instructs the OS 20 and middleware to change the task assignment to the computing devices 51 and 52 (computing devices #1 and #2) based on the details of the inefficiency notified by the inefficiency determination unit 120 (bold arrow o in Figure 9). Upon receiving the resource assignment change instruction from the assignment change instruction unit 130 to the computing devices 51 and 52 (computing devices #1 and #2), the OS 20 and middleware change the task issuance of APL#1 (arrow a in Figure 9) from a single computing device 51 (computing device #1) to computing devices 51 and 52 (computing devices #1 and #2) (indicated by the symbol p in Figure 9). The OS 20 and middleware then assign the changed task to computing device 51 (computing device #1) and computing device 52 (computing device #2) (arrows q and r in Figure 9). Since the task issuance for APL#1 (arrow a in Figure 9) has been changed to arithmetic device 51 (arithmetic device #1) and arithmetic device 52 (arithmetic device #2), the required fan output of arithmetic device 51 (arithmetic device #1) is reduced from 100% to 50%.
[0084] In Figure 9, the fan controller 70 outputs a fan output change instruction to fan 61 (fan #1), reducing the output of fan 61 (fan #1) by half (making the fan output of fan #1 50%).
[0085] This distributes the task across multiple computing devices 51 and 52 (computing device #1 and #2), thereby suppressing the temperature rise of each individual computing device 51 and 52 (computing device #1 and #2), and consequently reducing the required fan output.
[0086] <Modified Version> In the modified version, if load balancing is possible to determine the computing device to be allocated for each request, the temperature and fan output of the computing devices in use are obtained, and when a new task is generated, the available computing devices are adjusted based on the power consumption forecast, and then the task is allocated to the computing device that generates the least heat.
[0087] Figure 10 is a flowchart showing a modified example of distributed processing in which tasks are distributed across multiple computing devices. When the modified new request generation program starts, in step S31, the usage status collection processing unit 113 (Figure 2) of the cooling status collection unit 110 (Figure 9) collects the usage status of computing devices 51 to 54 (computing devices #1 to #4) using the CPU usage rate and other data from the OS 20. In addition, the sensor data acquisition processing unit 112 (Figure 2) of the cooling status collection unit 110 acquires the device temperature of computing devices 51 to 54, the power consumption of computing devices 51 to 54, and the fan rotation speed of fans 61 and 62 (Figure 9).
[0088] In step S32, the inefficiency determination unit 120 (Figure 9) determines the allocation destination to minimize the number of operating fans by sequentially assigning one thermal zone to each CPU in order, so as to minimize the number of thermal zones, if the degree of parallelism of the tasks is less than or equal to the total number of CPUs.
[0089] Here, if the degree of parallelism of the tasks is less than or equal to the total number of CPUs, the tasks are distributed to each CPU in a single thermal zone (example of operation in Figure 6).
[0090] If the degree of parallelism of a task is greater than the total number of CPUs, it will be distributed as evenly as possible across all thermal zones.
[0091] The following are variations of the allocation method: • Each CPU uses the same number of cores. • Allocate to the available core of the CPU with the lowest temperature.
[0092] In step S33, the assignment change instruction unit 130 of the task scheduler device 100 issues an assignment change instruction to the OS 20, specifying the CPU core to be used by the task in question, and then terminates the processing of this flow.
[0093] [Hardware Configuration] The server 1000 (Figures 1, 3, 6, 8, and 9) equipped with the task scheduler device 100 according to the above embodiment is realized by a computer 900 having a configuration such as that shown in Figure 11. Figure 11 is a hardware configuration diagram showing an example of a computer 900 that realizes the functions of the task scheduler device 100 (Figures 1, 2, 3, 6, 8, and 9). The computer 900 has a CPU 901, ROM 902, RAM 903, HDD 904, communication interface (I / F) 906, input / output interface (I / F) 905, and media interface (I / F) 907.
[0094] The CPU 901 operates based on programs stored in the ROM 902 or HDD 904, and controls various parts of the task scheduler device 100 (Figure 1). The ROM 902 stores boot programs executed by the CPU 901 when the computer 900 starts up, as well as programs that depend on the computer 900's hardware.
[0095] The CPU 901 controls input devices 910, such as a mouse or keyboard, and output devices 911, such as a display, via the input / output interface 905. The CPU 901 acquires data from the input devices 910 and outputs the generated data to the output devices 911 via the input / output interface 905. In addition to the CPU 901, a GPU (Graphics Processing Unit) or the like may also be used as a processor.
[0096] The HDD 904 stores programs executed by the CPU 901 and data used by those programs. The communication I / F 906 receives data from other devices via a communication network (e.g., NW (Network) 922) and outputs it to the CPU 901, and also transmits data generated by the CPU 901 to other devices via the communication network.
[0097] The media interface 907 reads a program or data stored in the recording medium 912 and outputs it to the CPU 901 via the RAM 903. The CPU 901 loads the program related to the desired processing from the recording medium 912 onto the RAM 903 via the media interface 907 and executes the loaded program. The recording medium 912 is an optical recording medium such as a DVD (Digital Versatile Disc) or PD (Phase change rewritable Disk), a magneto-optical recording medium such as an MO (Magneto Optical disk), a magnetic recording medium, a conductive memory tape medium, or a semiconductor memory.
[0098] For example, when computer 900 functions as a task scheduler device 100 (Figures 1, 2, 6, 8, and 9) configured as one of the devices according to this embodiment, the CPU 901 of computer 900 realizes the functions of the task scheduler device 100 by executing a program loaded on RAM 903. The HDD 904 stores the data in RAM 903. The CPU 901 reads and executes a program related to the target process from the recording medium 912. Alternatively, the CPU 901 may read a program related to the target process from another device via a communication network (NW 922).
[0099] [Effect] As described above, the task scheduler device 100 (Figures 1, 2, 3, 6, 8, 9) of the server 1000 (Figures 1, 3, 6, 8, 9) has a group of one or more computing devices 51 to 54 (computing devices #1 to #4) that are cooled by a single fan 61, 62 (fan #1, fan #2) (Figures 1, 3, 6, 8, 9) as a thermal zone 71, 72 (Figures 1, 3, 6, 8, 9), wherein the computing devices 51 to 54 (computing devices #1 to #4) The system includes a cooling state collection unit 110 that collects usage information for (Figures 1, 3, 6, 8, and 9), an inefficiency determination unit 120 that obtains usage information for computing devices 51 to 54 from the cooling state collection unit 110 and determines which thermal zones 71 and 72 (Figures 1, 3, 6, 8, and 9) are experiencing cooling inefficiency based on the acquired information, and an assignment change instruction unit 130 that instructs a change in task assignment to computing devices 51 to 54 based on the content of the inefficiency notified by the inefficiency determination unit 120.
[0100] In this way, the task scheduler device 100 arranges tasks to minimize cooling fan power consumption, taking into account the physical arrangement within the chassis. Furthermore, the task scheduler device 100 moves tasks to computing devices located in positions where they can be cooled by fewer high-power fans. As a result, cooling inefficiencies are reduced without affecting application performance, and the overall power consumption of the server can be minimized.
[0101] The task scheduler device 100 also satisfies the following requirements. The cooling status collection unit 110 comprehensively collects necessary information such as the usage status (utilization rate), power consumption, temperature, and fan output of each fan for all computing devices 51 to 54 (computing devices #1 to #4). The inefficiency determination unit 120 detects thermal zones where cooling inefficiency is occurring based on the information collected by the cooling status collection unit 110 and notifies the assignment change instruction unit 130. As a result, the task scheduler device 100 satisfies requirement 1: [Minimizing power consumption of multiple devices]. The inefficiency determination unit 120 satisfies requirement 2: [Performance impact] by determining the destination of the task to which it should move has resources equivalent to the source. Based on the notified details of the inefficiency, the assignment change instruction unit 130 instructs the OS 20 and middleware to change the task assignment to the device, thereby eliminating the inefficiency, and the task scheduler device 100 satisfies requirement 3: [Elimination of cooling inefficiency].
[0102] In the task scheduler device 100 (Figures 1, 2, 3, 6, 8, and 9), the inefficiency determination unit 120 determines which of the multiple thermal zones 71 and 72 (Figures 1 and 3) have only a single computing device 51 to 54 (computing device #1 to #4) in a predetermined high-load state, and the assignment change instruction unit 130 assigns the tasks of the determined high-load computing devices 51 and 52 (computing device #51 and #52) to computing devices 51 and 52 (computing device #51 and #52) belonging to fewer thermal zones 71.
[0103] In this way, the task scheduler device 100 detects multiple thermal zones 71, 72 where only a single computing device 51 (computing device #51) is under high load, and assigns it to the fewest possible computing devices 51, 52 (computing devices #51, #52) in the thermal zone 71. This makes effective use of fans operating at high output and reduces the number of operating fans, thereby minimizing cooling power consumption.
[0104] In the task scheduler device 100 (Figures 1, 2, 3, 6, 8, and 9), the inefficiency determination unit 120 determines which computing device 51 (computing device #51) (Figure 8) is in a predetermined high-load state within the thermal zones 71 and 72 (Figures 1, 3, 6, 8, and 9), and the assignment change instruction unit 130 distributes and assigns the tasks of the determined high-load computing device 51 (computing device #51) to multiple computing devices 51 and 52 (computing devices #51 and #52) (Figure 9).
[0105] In this way, the task scheduler device 100 can suppress the temperature rise of one computing device by distributing tasks to multiple computing devices 51, 52 (computing device #51, #52) (Figure 9), and as a result can reduce the required fan output.
[0106] In the task scheduler device 100 (Figures 1, 2, 3, 6, 8, and 9), the cooling state collection unit 110 (Figure 2) includes a thermal zone definition information storage unit 111 (Figure 2) that holds cooling layout information, a sensor data acquisition processing unit 112 (Figure 2) that acquires the device temperature of the computing devices 51 to 54 (computing devices #1 to #4), the power consumption of the computing devices 51 to 54, and the fan rotation speed of the fans 61 and 62 (fan #1 and fan #2), and a usage status collection processing unit 113 (Figure 2) that collects the usage status of the computing devices 51 to 54.
[0107] In this way, the task scheduler device 100 can collect the usage status of computing devices using various interfaces provided by the OS 20 (including middleware and device drivers). As a result, the inefficiency determination unit 120 can detect cooling inefficiencies that can be improved by combining this information with the cooling layout information (which computing device is cooled by which fan) that is internally held by the thermal zone definition information storage unit 111.
[0108] Furthermore, among the processes described in the above embodiments and each modified example, all or part of the processes described as being performed automatically can be performed manually, or all or part of the processes described as being performed manually can be performed automatically by known methods. In addition, the processing procedures, control procedures, specific names, and information including various data and parameters shown in the above documents and drawings can be arbitrarily changed unless otherwise specified. Moreover, each component of each illustrated device is a functional concept and does not necessarily have to be physically configured as shown. In other words, the specific forms of distribution and integration of each device are not limited to those shown, and all or part of them can be functionally or physically distributed and integrated in any unit according to various loads and usage conditions.
[0109] Furthermore, each of the above configurations, functions, processing units, and processing means may be implemented in hardware, either partially or entirely, by designing them as integrated circuits, for example. Alternatively, each of the above configurations and functions may be implemented in software that allows the processor to interpret and execute programs that implement each function. Information such as programs, tables, and files that implement each function can be stored in memory, a recording device such as a hard disk or SSD (Solid State Drive), or a recording medium such as an IC (Integrated Circuit) card, an SD (Secure Digital) card, or an optical disc.
[0110] 10 Hardware (HW) 20 OS 30 Userland 70 Fan Controller 51-54 Computing Devices (Computing Devices #1-#4) 61, 62 Fans (Fan #1, Fan #2) (Cooling Fans) 71, 72 Thermal Zone (TZ) 100 Task Scheduler Device 110 Cooling Status Collection Unit 120 Inefficiency Determination Unit 130 Assignment Change Instruction Unit 111 Thermal Zone Definition Information Storage Unit 112 Sensor Data Acquisition Processing Unit 113 Usage Status Collection Processing Unit 1000 Server APL#1, APL#2 Application
Claims
1. A task scheduler device for a server having a thermal zone comprising a group of one or more computing devices cooled by a single fan, the task scheduler device comprising: a cooling status collection unit that collects the usage status of the computing devices; an inefficiency determination unit that obtains the usage status of the computing devices from the cooling status collection unit and determines the thermal zone where cooling inefficiency is occurring based on the acquired information; and an assignment change instruction unit that instructs a change in task assignment to the computing devices based on the content of the inefficiency notified by the inefficiency determination unit.
2. The task scheduler device according to claim 1, characterized in that the inefficiency determination unit determines that in a plurality of thermal zones only one computing device is in a predetermined high-load state, and the assignment change instruction unit assigns the task of the computing device in the determined high-load state to a smaller number of computing devices belonging to the thermal zones.
3. The task scheduler device according to claim 1, characterized in that the inefficiency determination unit determines a computing device that is in a predetermined high-load state within the thermal zone, and the assignment change instruction unit distributes and assigns the tasks of the determined computing device that is in a high-load state to a plurality of computing devices.
4. A program for a computer acting as a task scheduler device for a server having a thermal zone consisting of one or more computing devices cooled by a single fan, to execute the following steps: a procedure for collecting the usage status of the computing devices; a procedure for obtaining the usage status of the computing devices and determining the thermal zone where cooling inefficiency is occurring based on the obtained information; and a procedure for instructing a change in task assignment to the computing devices based on the determined inefficiency.
Citation Information
Patent Citations
Job scheduling system for parallel computer
JP2004126968A
Server operation system and server operation method
JP2011076158A
Calculation processing system, and job distribution arrangement method and program thereof
JP2012073784A
Electronic apparatus cooling system
JP2014183061A
Task scheduler in distributed processing system
WO2003083693A1