Task scheduling method and device for processor, storage medium and program product

By predicting the future temperature and operating parameters of the processor cluster, and combining the target parameters of the liquid cooling system and the target parameters of the phase change material, the output of the coolant is adjusted, and the problem of uneven heat dissipation efficiency of the processor cluster is solved, achieving a more uniform and efficient heat dissipation effect.

CN120179055AActive Publication Date: 2025-06-20INSPUR SUZHOU INTELLIGENT TECH CO LTD

Patent Information

Application Number
CN202510654914.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-21
Publication Date
2025-06-20
Estimated Expiration
2045-05-21

AI Technical Summary

Technical Problem

The heat dissipation efficiency of each processor in the processor cluster is uneven, resulting in a high temperature of some processors and a low temperature of the other part, which is low overall heat dissipation efficiency.

Method used

By predicting future temperature and operating parameters based on the current temperature and operating parameters of the processor cluster, combining the target parameters of the liquid cooling system and the target parameters of the phase change material, the expected output of the coolant is determined, and the actual output is adjusted based on the control mechanism, and finally task scheduling is performed based on the cooling error.

Benefits of technology

Improve the heat dissipation efficiency of the processor cluster, ensure that each processor can evenly dissipate heat, and reduce equipment losses.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120179055A_ABST
    Figure CN120179055A_ABST
Patent Text Reader

Abstract

The invention provides a task scheduling method and equipment of a processor, a storage medium and a program product, which can be applied to the technical field of heat dissipation. The task scheduling method for the processors comprises the steps of determining predicted temperatures and predicted operation parameters of the processors according to current temperatures and current operation parameters of the processors in a processor cluster; according to the target system parameters of the liquid cooling system, the target material parameters of the phase change material, the predicted temperature and the predicted operation parameters, the expected output quantity of the cooling liquid in the liquid cooling system is determined; on the basis of a control mechanism of the liquid cooling system, according to the expected output quantity, the predicted temperature and the current temperature, the actual output quantity of the cooling liquid at the future moment is obtained; and performing task scheduling on the plurality of processors according to the cooling error between the expected output quantity and the actual output quantity, the predicted temperature and the predicted operation parameters. The control of the liquid cooling system and the task scheduling of the processors are combined to perform heat dissipation on the processor cluster, so that the heat dissipation efficiency of the processor cluster is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of heat dissipation, and more particularly to a task scheduling method, device, storage medium, and program product for a processor. Background Art

[0002] Multiple processors in a processor cluster can work together to achieve high-concurrency task execution or complex algorithm acceleration. Since the tasks of each processor in the processor cluster are different, the operating conditions are different, and the heat generated by each processor is also different.

[0003] When using a liquid cooling system to dissipate heat from the processor cluster, there are some processors with still relatively high temperatures after heat dissipation, and the temperatures of some other processors are relatively low after heat dissipation, resulting in low heat dissipation efficiency of the processor cluster. Summary of the Invention

[0004] In view of the above problems, the present invention provides a task scheduling method, device, storage medium, and program product for a processor.

[0005] According to a first aspect of the present invention, there is provided a task scheduling method for a processor, including: determining a predicted temperature and predicted operating parameters of the processor at a future moment according to the current temperatures and current operating parameters of multiple processors in a processor cluster; determining an expected output amount of a coolant in a liquid cooling system according to target system parameters of the liquid cooling system, target material parameters of a phase change material, the predicted temperature, and the predicted operating parameters, where the liquid cooling system is used to dissipate heat from the processor cluster; obtaining an actual output amount of the coolant at a future moment based on a control mechanism of the liquid cooling system according to the expected output amount, the predicted temperature, and the current temperature; and performing task scheduling on multiple processors according to a cooling error between the expected output amount and the actual output amount, the predicted temperature, and the predicted operating parameters.

[0006] A second aspect of the present invention provides a task scheduling device for a processor, including: a first determination module for determining a predicted temperature and predicted operating parameters of the processor at a future moment according to the current temperatures and current operating parameters of multiple processors in a processor cluster; a second determination module for determining an expected output amount of a coolant in a liquid cooling system according to target system parameters of the liquid cooling system, target material parameters of a phase change material, the predicted temperature, and the predicted operating parameters, where the liquid cooling system is used to dissipate heat from the processor cluster; an obtaining module for obtaining an actual output amount of the coolant at a future moment based on a control mechanism of the liquid cooling system according to the expected output amount, the predicted temperature, and the current temperature; and a task scheduling module for performing task scheduling on multiple processors according to a cooling error between the expected output amount and the actual output amount, the predicted temperature, and the predicted operating parameters.

[0007] A third aspect of the present invention provides an electronic device, including: one or more processors; a memory for storing one or more computer programs, wherein the one or more processors execute the one or more computer programs to implement the steps of the above method.

[0008] A fourth aspect of the present invention further provides a computer-readable storage medium, on which a computer program or instruction is stored, and when the computer program or instruction is executed by a processor, the steps of the above method are implemented.

[0009] A fifth aspect of the present invention further provides a computer program product, including a computer program or instruction, and when the computer program or instruction is executed by a processor, the steps of the above method are implemented.

[0010] According to an embodiment of the present invention, by determining the predicted temperature and predicted operating parameters of a processor at a future moment based on the current temperatures and current operating parameters of multiple processors in a processor cluster, anticipatory heat dissipation of the processor can be performed based on the predicted temperature and predicted operating parameters. According to the target system parameters of the liquid cooling system, the target material parameters of the phase change material, the predicted temperature, and the predicted operating parameters, the desired output volume of the coolant in the liquid cooling system is determined, such that the desired output volume synthesizes the output volumes of the coolant required for the operation of multiple processors. Based on the control mechanism of the liquid cooling system, the actual output volume of the coolant at a future moment is obtained according to the desired output volume, the predicted temperature, and the current temperature. According to the cooling error between the desired output volume and the actual output volume, the predicted temperature, and the predicted operating parameters, task scheduling is performed on multiple processors in the processor cluster. Combining the control of the liquid cooling system with the task scheduling of the processor for heat dissipation of the processor cluster improves the heat dissipation efficiency of the processor cluster. BRIEF DESCRIPTION OF THE DRAWINGS

[0011] Through the following description of the embodiments of the present invention with reference to the drawings, the above content and other objects, features, and advantages of the present invention will become clearer. In the drawings:

[0012] Figure 1 The application scenario diagram of the task scheduling method of the processor according to the embodiment of the present invention is shown;

[0013] Figure 2 The flowchart of the task scheduling method of the processor according to the embodiment of the present invention is shown;

[0014] Figure 3 The hierarchical interaction flowchart of the NoF protocol according to the embodiment of the present invention is shown;

[0015] Figure 4A The schematic diagram of the gold finger groove of the lower housing according to the embodiment of the present invention is shown;

[0016] Figure 4BShows a schematic diagram of the gold finger bumps of the upper housing according to an embodiment of the present invention;

[0017] Figure 5 Shows a schematic diagram of the detachable member according to an embodiment of the present invention;

[0018] Figure 6 Shows a schematic diagram of the liquid cooling quick connector according to an embodiment of the present invention;

[0019] Figure 7 Shows a block diagram of the task scheduling device of the processor according to an embodiment of the present invention; and

[0020] Figure 8 Shows a block diagram of the electronic device suitable for implementing the task scheduling method of the processor according to an embodiment of the present invention. Detailed Embodiments

[0021] Hereinafter, embodiments of the present invention will be described with reference to the accompanying drawings. However, it should be understood that these descriptions are merely exemplary and are not intended to limit the scope of the present invention. In the following detailed description, for the sake of explanation, many specific details are set forth to provide a thorough understanding of the embodiments of the present invention. However, obviously, one or more embodiments can also be implemented without these specific details. In addition, in the following description, descriptions of well-known structures and technologies are omitted to avoid unnecessarily obscuring the concepts of the present invention.

[0022] The terms used herein are merely for describing specific embodiments and are not intended to limit the present invention. The terms "including", "comprising", etc. used herein indicate the presence of the described features, steps, operations, and / or components, but do not exclude the presence or addition of one or more other features, steps, operations, or components.

[0023] All terms used herein (including technical and scientific terms) have the meanings commonly understood by those skilled in the art, unless otherwise defined. It should be noted that the terms used herein should be interpreted as having a meaning consistent with the context of this specification and should not be interpreted in an idealized or overly rigid manner.

[0024] In the case of using expressions such as "at least one of A, B, and C, etc.", generally, it should be interpreted according to the meaning commonly understood by those skilled in the art (for example, "a system having at least one of A, B, and C" should include, but is not limited to, a system having only A, only B, only C, having A and B, having A and C, having B and C, and / or having A, B, and C).

[0025] Multiple processors in a processor cluster can work together to achieve high-concurrency task execution or complex algorithm acceleration. Since the tasks of the processors in the processor cluster are different, their operating conditions are different, and the heat generated by each processor is also different.

[0026] When using a liquid cooling system to dissipate heat from the processor cluster, there are some processors with still relatively high temperatures after cooling, while the temperatures of other processors are relatively low, resulting in low heat dissipation efficiency of the processor cluster.

[0027] In view of this, embodiments of the present invention provide a task scheduling method for processors, including: determining the predicted temperature and predicted operating parameters of the processors at a future moment according to the current temperatures and current operating parameters of multiple processors in the processor cluster; determining the desired output volume of the coolant in the liquid cooling system according to the target system parameters of the liquid cooling system, the target material parameters of the phase change material, the predicted temperature, and the predicted operating parameters, where the liquid cooling system is used to dissipate heat from the processor cluster; based on the control mechanism of the liquid cooling system, obtaining the actual output volume of the coolant at a future moment according to the desired output volume, the predicted temperature, and the current temperature; and performing task scheduling for multiple processors according to the cooling error between the desired output volume and the actual output volume, the predicted temperature, and the predicted operating parameters.

[0028] Figure 1 The application scenario diagram of the task scheduling method for processors according to an embodiment of the present invention is shown.

[0029] As Figure 1 shown, the application scenario according to this embodiment may include a processor cluster 110, a liquid cooling system 120, and a circuit board 130.

[0030] The processor cluster 110 includes multiple processors 111, and the multiple processors 111 can be plugged and unplugged via mounting slots on the coolant pipes of the liquid cooling system 120.

[0031] The liquid cooling system 120 may include a phase change heat storage tank 121, a main circulation pump 122, a manifold distributor 123, and a microchannel cold plate with coolant pipes distributed thereon. The microchannel cold plate is not shown in Figure 1 , and multiple processors 111 in the processor cluster 110 are mounted on the microchannel cold plate. The phase change heat storage tank 121 can deliver the phase change material into the phase change material coating. The pump speed of the main circulation pump 122 is controlled based on the control mechanism of the liquid cooling system to deliver coolant to each coolant pipe in the microchannel cold plate, and the pump speed can be 200 to 5000 revolutions per minute.

[0032] The circuit board 130 can be used to obtain the current temperature and current operating parameters from the processor cluster 110 and send them to the server that executes the task scheduling method of the processor. It should be noted that components capable of supporting the task scheduling method of the processor are provided in the circuit board 130, and the circuit board 130 can also be used to perform task scheduling on the processor cluster.

[0033] The circuit board 130 communicates with the processor cluster 110 via an optical fiber interface array, and uses an optoelectronic conversion unit to realize the mutual conversion of optical signals and electrical signals, so as to send the relevant data of the processor cluster 110 to the circuit board 130. The circuit board 130 is also provided with a PCIe (Peripheral Component Interconnect Express) switch chip to achieve high-speed data interaction.

[0034] An LED (Light-Emitting Diode) array for monitoring the temperature of the processor cluster 110 is provided on the circuit board 130.

[0035] It should be understood that Figure 1 the numbers of the processor cluster, liquid cooling system, circuit board, etc. in are only illustrative. According to the implementation requirements, there can be any number of processor clusters, liquid cooling systems, and circuit boards.

[0036] Figure 2 A flowchart of the task scheduling method of the processor according to an embodiment of the present invention is shown.

[0037] As Figure 2 shown, the task scheduling method of the processor in this embodiment includes operations S210 to S240.

[0038] In operation S210, based on the current temperatures and current operating parameters of multiple processors in the processor cluster, determine the predicted temperature and predicted operating parameters of the processors at a future moment.

[0039] According to an embodiment of the present invention, the processor (Central Processing Unit) can be an accelerated processor, a heterogeneous computing processor, etc. The processor cluster can be an accelerated processor cluster, a heterogeneous computing processor cluster, etc. The heterogeneous computing processor cluster can be a hybrid cluster of a central processor and an image processor.

[0040] According to an embodiment of the present invention, the current temperature can be obtained through a temperature sensor integrated inside the processor or measured through an external temperature probe. For example, the acquisition temperature frequency of the internally integrated temperature sensor can be 10 kHz.

[0041] The current temperature and current operating parameters can be obtained from a server management tool. The server management tool can remotely query the temperature, operating parameters, etc. of the processor.

[0042] The current operating parameters can include current power consumption, current load, current frequency, current usage rate, etc. For example, the acquisition frequency of the current power consumption can be 1 megahertz, and the acquisition frequency of the current load can be 100 hertz.

[0043] Exemplarily, taking the current operating parameters as constraints, the current temperature of the processor is input into a temperature fitting model to obtain the predicted temperature of the processor at a future moment. The temperature fitting model can be obtained by fitting the historical temperature of the processor based on the historical operating constraints of the processor.

[0044] Taking the current temperature of the processor as a constraint, the current operating parameters of the processor are input into an operating parameter fitting model to obtain the predicted operating parameters of the processor at a future moment. The operating parameter fitting model can be obtained by fitting the historical operating parameters of the processor based on the temperature constraints of the processor.

[0045] Exemplarily, the current temperature and current operating parameters of the processor can be input into a machine learning model to obtain the predicted temperature and predicted operating parameters of the processor at a future moment.

[0046] The machine learning model can be a neural network model (such as a long short-term memory neural network), a random forest model, etc. The historical temperature sequence can be the historical temperature within 5 minutes, and the predicted temperature is the predicted temperature for the next 3 seconds.

[0047] In operation S220, according to the target system parameters of the liquid cooling system, the target material parameters of the phase change material, the predicted temperature, and the predicted operating parameters, the desired output volume of the coolant in the liquid cooling system is determined.

[0048] According to an embodiment of the present invention, the liquid cooling system is used to dissipate heat from a processor cluster. The liquid cooling system can include a microchannel cold plate and other emergency cooling devices. The width of the coolant pipeline of the microchannel cold plate is 100 micrometers.

[0049] According to an embodiment of the present invention, the phase change material can absorb the heat of the processor.

[0050] According to an embodiment of the present invention, the target system parameters characterize the system performance change of the liquid cooling system at a future moment. The target material parameters characterize the material performance change of the phase change material at a future moment. For example, the target system parameters can be a system adaptation coefficient. As the pump speed of the liquid cooling system increases, the system adaptation coefficient decreases. The target material parameters can be a material adaptation coefficient. When the phase change material enters a saturated state, the material adaptation coefficient can increase.

[0051] According to an embodiment of the present invention, based on a processor heat dissipation control strategy, the desired output volume of the coolant in the liquid cooling system can be determined according to the target system parameters of the liquid cooling system, the target material parameters of the phase change material, the predicted temperature, and the predicted operating parameters.

[0052] For example, the processor heat dissipation control strategy can be to achieve coordinated heat dissipation according to the temperature and operating parameters of the processor through an adaptive control algorithm.

[0053] In operation S230, based on the control mechanism of the liquid cooling system, the actual output volume of the coolant at a future moment is obtained according to the desired output volume, the predicted temperature, and the current temperature.

[0054] According to an embodiment of the present invention, the control mechanism of the liquid cooling system can be PID (Proportional-Integral-Derivative) control, feedforward control, fuzzy control, etc.

[0055] Exemplarily, the control mechanism of the liquid cooling system can be PID control, and the formula of the control mechanism of PID control is as follows:

[0056] (1)

[0057] represents the actual output volume of the coolant at a future moment, T current represents the current temperature T target represents the predicted temperature. Terr represents the temperature error value (i.e., the error value between the current temperature and the predicted temperature). K p represents the proportional coefficient (the value range can be 0.5 - 1.2). K p determines the response speed of the liquid cooling system, K p the larger the value, the more aggressive the adjustment. K i represents the integral coefficient (the value range can be 0.01 - 0.1). K i controls the error correction intensity, K i the larger the value, the higher the steady-state accuracy. K p (T current - T target ) is the proportional term, and the larger the real-time temperature difference, the larger the flow rate adjustment amplitude. K i ∫(Terr)dt is the integral term, continuously accumulating the temperature error to eliminate the steady-state error (long-term precision). It can suppress temperature fluctuations (such as sudden changes in the load of the processor), and compensate for changes in the ambient temperature through the integral term. The PID control algorithm can adjust the rotation speed of the liquid cooling pump, thereby controlling the output volume of the coolant. The rotation speed of the liquid cooling pump can be from 200 to 5000 revolutions per minute.

[0058] The coolant flow rate of the liquid cooling system is limited. There is a situation where the heat generated by the processor during operation exceeds the heat dissipation capacity of the maximum coolant flow rate (for example, the heat generated during operation is 100 joules and the heat dissipation is 80 joules), which may cause the processor to be in a high-temperature condition and result in equipment loss.

[0059] At the same time, the control mechanism of the liquid cooling system remains unchanged, and the actual output volume of the coolant in the liquid cooling system is the same. Due to the different operating conditions of different processors, the required coolant output volumes (desired output volumes) for multiple processors to operate are different.

[0060] In operation S240, according to the cooling error, predicted temperature, and predicted operating parameters between the desired output volume and the actual output volume, task scheduling is performed for multiple processors.

[0061] According to an embodiment of the present invention, the cooling error between the desired output volume and the actual output volume can be the cooling temperature difference, the error between the desired output volume and the actual output volume.

[0062] For example, the cooling temperature difference can be determined by cooling the processor with the coolants of the desired output volume and the actual output volume respectively, and based on the difference between the two cooled temperatures.

[0063] According to an embodiment of the present invention, by determining the predicted temperature and predicted operating parameters of the processor at a future moment based on the current temperature and current operating parameters of multiple processors in the processor cluster, anticipatory heat dissipation can be performed on the processor based on the predicted temperature and predicted operating parameters. According to the target system parameters of the liquid cooling system, the target material parameters of the phase change material, the predicted temperature, and the predicted operating parameters, the desired output volume of the coolant in the liquid cooling system is determined, so that the desired output volume synthesizes the coolant output volumes required for multiple processors to operate. Based on the control mechanism of the liquid cooling system, according to the desired output volume, predicted temperature, and current temperature, the actual output volume of the coolant at a future moment is obtained. According to the cooling error, predicted temperature, and predicted operating parameters between the desired output volume and the actual output volume, task scheduling is performed for multiple processors in the processor cluster. Combining the control of the liquid cooling system with the task scheduling of the processor for heat dissipation of the processor cluster improves the heat dissipation efficiency of the processor cluster.

[0064] According to an embodiment of the present invention, determining the desired output volume of the coolant in the liquid cooling system according to the target system parameters of the liquid cooling system, the target material parameters of the phase change material, the predicted temperature, and the predicted operating parameters includes: determining the heat dissipation driven by power consumption of the processor according to the target system parameters and the predicted operating parameters; determining the temperature compensation amount of the processor according to the target material parameters, the predicted temperature, and the temperature safety threshold condition; and determining the desired output volume according to the heat dissipation driven by power consumption and the temperature compensation amount.

[0065] According to an embodiment of the present invention, in the processor heat dissipation control strategy, the expected output Q cool has the following formula:

[0066] (2)

[0067] k1 represents the target system parameter, k2 represents the target material parameter, P represents the predicted operating parameter, T target represents the predicted temperature, T safe represents the temperature safety threshold, and the heat dissipation driven by power consumption is , and the temperature compensation amount is . The exponent can reflect the non-linear characteristics of heat accumulation in the liquid cooling system. For every 1°C exceeding the safety temperature, the heat dissipation demand increases exponentially.

[0068] For example, the temperature safety threshold condition can be that the predicted temperature is greater than the temperature safety threshold. The temperature safety threshold can be 85% of the processor junction temperature, such as 85°C.

[0069] For example, the temperature safety threshold condition can be that the temperature difference between multiple adjacent processors is greater than the temperature difference threshold.

[0070] The target system parameter and the target material parameter can be obtained by calculating through heat model construction. For example, for a processor cluster that needs continuous heat dissipation, k1 = 0.8 and k2 = 0.3 can be set; for strengthening the temperature control of the processor transient, k1 = 1.2 and k2 = 0.5 can be set; for the heat dissipation energy-saving mode, k1 = 0.3 and k2 = 0.1 can be set.

[0071] According to an embodiment of the present invention, since the operating conditions of different processors are different, the heat dissipation driven by the power consumption of the processor can be determined according to the target system parameter and the predicted operating parameter. At the same time, as the phase change material changes with the temperature of the processor, the heat dissipation property of the phase change material may also change. Therefore, the temperature compensation amount of the processor is determined according to the target material parameter, the predicted temperature, and the temperature safety threshold condition. Thus, the expected output amount is determined according to the heat dissipation driven by power consumption and the temperature compensation amount, so that the expected output amount synthesizes the coolant output amounts required for the operation of multiple processors.

[0072] According to an embodiment of the present invention, according to the cooling error between the expected output amount and the actual output amount, the predicted temperature, and the predicted operating parameter, task scheduling is performed on multiple processors, including: determining a first processor whose predicted temperature satisfies the temperature safety threshold condition from multiple processors according to the cooling error; determining a task to be migrated from at least one task to be processed by the first processor according to the predicted temperature of the first processor; determining a second processor that meets the migration condition from multiple processors according to the predicted temperatures and predicted operating parameters of multiple processors; and migrating the task to be migrated to the second processor for processing.

[0073] Exemplarily, according to the cooling error, a processor that cannot be cooled in time can be determined from multiple processors. It is determined whether the predicted temperature of the processor that cannot be cooled in time meets the temperature safety threshold condition.

[0074] Exemplarily, the heat dissipation difference can be determined according to the predicted temperature and the actual output of the first processor; according to the heat dissipation difference, a task to be migrated can be determined from at least one task to be processed by the first processor. At least one task has a corresponding running heat generation.

[0075] Exemplarily, according to the predicted temperatures of multiple processors, a processor with a large temperature difference from the first processor can be determined, and according to the predicted operating parameters of multiple processors, a relatively idle processor can be determined. A second processor that meets the migration condition is determined from the processor with a large temperature difference from the first processor and the relatively idle processor.

[0076] According to an embodiment of the present invention, determining a first processor whose predicted temperature meets the temperature safety threshold condition from multiple processors according to the cooling error includes: when there is a processor in the processor cluster with a cooling error greater than the error threshold, determining a first processor whose predicted temperature meets the temperature safety threshold condition from the processor cluster.

[0077] According to an embodiment of the present invention, when the cooling error of the processor is greater than the error threshold, it indicates that the operating temperature of the processor is difficult to be cooled in time. For example, the cooling error is 10% and the error threshold is 2%, and the actual heat dissipation of the processor is less than the expected heat dissipation.

[0078] According to an embodiment of the present invention, when the cooling error of the processor is less than the error threshold, it indicates that the operating temperature of the processor can be cooled in time.

[0079] For example, determining a first processor whose predicted temperature is greater than the temperature safety threshold from the processor cluster; and / or determining a first processor with a temperature difference greater than the temperature difference threshold between multiple adjacent processors.

[0080] A predicted temperature matrix can be constructed according to the predicted temperatures of each processor. A safety threshold matrix can be constructed according to the safety temperature thresholds of multiple processors. The predicted temperature matrix can be subtracted from the safety threshold matrix to determine a first processor that exceeds the safety temperature threshold from multiple processors.

[0081] Each row or column in the predicted temperature matrix can represent processors at different positions. By taking the difference between adjacent rows or columns in the predicted temperature matrix, a first processor with a temperature difference greater than the temperature difference threshold between multiple adjacent processors can be determined.

[0082] For example, a predicted temperature field is established based on the three-dimensional positions and temperatures of multiple processors. When it is monitored that the predicted temperature field is greater than the safe temperature threshold, the tasks to be processed by the processors can be migrated.

[0083] According to an embodiment of the present invention, by determining whether there is a processor in the processor cluster with a cooling error greater than the error threshold, it can be determined whether multiple processors can be timely cooled in the future. In the case where there is a processor in the processor cluster with a cooling error greater than the error threshold, a first processor whose predicted temperature meets the temperature safety threshold condition is determined from the processor cluster, and a first processor that needs to adjust the task amount to reduce the operating heat can be determined, thereby protecting the processor and reducing equipment loss.

[0084] According to an embodiment of the present invention, there are multiple second processors, the predicted operating parameters include at least one of the following: predicted power consumption and predicted computing load, and the migration conditions include a temperature difference condition and an operating condition; determining a second processor that meets the migration conditions from multiple processors according to the predicted temperatures and predicted operating parameters of the multiple processors includes: determining candidate processors that meet the temperature difference condition from the multiple processors according to the predicted temperatures of the multiple processors; determining a second processor that meets the operating condition from the multiple candidate processors according to the position distance between the candidate processor and the first processor, the predicted operating parameters of the candidate, and the time slice length, and the time slice length is determined according to the number of processing cores of the processor and the time required to process at least one task.

[0085] According to an embodiment of the present invention, the predicted power consumption may be the electric power consumed when the processor operates at a future time. The predicted computing load may be the amount of tasks to be processed.

[0086] According to an embodiment of the present invention, the temperature difference condition may be that the predicted temperature difference from the first processor is greater than a preset temperature difference value.

[0087] The predicted temperature distribution differences of different processors are monitored in real time. When the predicted temperature exceeds the safe temperature threshold, the tasks to be migrated are dynamically migrated to candidate processors with lower temperatures, which can form an intelligent scheduling similar to "thermal convection".

[0088] According to an embodiment of the present invention, the operating condition may be that the evaluation value determined based on the predicted operating parameters, the position distance, and the time slice length is greater than a preset evaluation threshold. The weights of the predicted computing load, the position distance, and the time slice length can be determined according to the task attributes of the tasks to be migrated.

[0089] For example, if the attribute of the task to be migrated is real-time service processing, the weight of the time slice length can be increased to determine a second processor from multiple candidate processors that can quickly process and complete the task to be migrated.

[0090] For example, the attribute of the task to be migrated can be a batch operation. The weight of the location distance can be increased, and the second processor with the minimum transmission distance can be determined from multiple candidate processors, which can reduce the data transmission time.

[0091] For example, if the task volume of the task attribute of the task to be migrated is relatively large, the weights of the candidate predicted operation parameters (such as predicted operation load and predicted power consumption) can be reduced to avoid migrating the task to be migrated to a processor with a high load or power consumption.

[0092] For example, in task scheduling, the number of processing cores for the same task to be migrated is less than the preset number of cores, so as to reduce the number of cross-core processing and avoid complicating the processing process of the task to be migrated.

[0093] According to an embodiment of the present invention, the time slice length t slice The formula is as follows:

[0094] (3)

[0095] The time slice length t slice takes the larger value of 0.1 ms (millisecond) and . Ttask represents the task processing duration; Ncore represents the number of processing cores.

[0096] According to an embodiment of the present invention, based on the predicted temperatures of multiple processors, candidate processors that meet the temperature difference condition are determined from multiple processors, so as to realize dynamically migrating the task to be migrated to a candidate processor with a lower temperature, which can form an intelligent scheduling similar to "thermal convection". According to the location distance between the candidate processor and the first processor, the candidate predicted operation parameters, and the time slice length, the second processor that meets the operation condition is determined from multiple candidate processors, so as to realize that the task to be migrated can be normally processed on the second processor and improve the service processing efficiency.

[0097] According to an embodiment of the present invention, the task to be migrated can be decomposed into discrete subtasks (such as 100 millisecond granularity), and combined with the time slice cutting technology to realize microsecond-level task migration scheduling, ensuring the atomic allocation and rapid recombination ability of computing power resources.

[0098] Based on the White Rabbit protocol, through physical layer timestamp marking and fiber optic delay compensation, the task to be migrated is migrated to the second processor, achieving a global clock synchronization accuracy of ±0.5 nanoseconds.

[0099] In task migration, modulation technology can be used to optimize data processing and communication. The modulation technology suppresses the clock jitter to the <0.1 picosecond level, ensuring the time consistency of task migration across processors.

[0100] According to an embodiment of the present invention, determining a task to be migrated from at least one task processed by a first processor according to a predicted temperature of the first processor includes: determining the amount of tasks completed by the first processor under temperature safety threshold conditions according to the predicted temperature of the first processor; and determining the task to be migrated from at least one task processed by the first processor according to the actual output amount and the amount of tasks.

[0101] According to an embodiment of the present invention, the temperature safety threshold condition may be that the predicted temperature is greater than the temperature safety threshold, and determining the amount of tasks completed by the first processor when it is less than or equal to the temperature safety threshold; and / or the temperature safety threshold condition may be that the temperature difference between the first processor and the adjacent processor is greater than the temperature difference threshold, and determining the amount of tasks completed by the first processor when it is less than or equal to the temperature difference threshold.

[0102] According to an embodiment of the present invention, determining a reference task amount of the first processor corresponding to the actual output amount, and when the amount of tasks is less than the reference task amount, determining the task to be migrated from at least one task processed by the first processor.

[0103] By comparing the amount of tasks with the reference task amount, it is possible to accurately determine that the first processor can obtain effective heat dissipation, and at the same time accurately lock the task to be migrated.

[0104] According to an embodiment of the present invention, since the operating conditions of different processors are different, by determining the amount of tasks completed by the first processor under temperature safety threshold conditions according to the predicted temperature of the first processor; and determining the task to be migrated from at least one task processed by the first processor according to the actual output amount and the amount of tasks, the amount of tasks to be migrated can be accurately determined, so that each processor in the processor cluster can be accurately cooled, and the heat dissipation efficiency is high.

[0105] According to an embodiment of the present invention, the target system parameters and target material parameters are determined based on the following steps: determining a combined data set of multiple initial system parameters and initial material parameters according to the simulation temperature and simulation operating parameters, and the simulation operating parameters and simulation temperature are obtained by performing a simulation using a simulation model of the processor; inputting the multiple combined data sets into a thermal model respectively to obtain multiple thermal model function values; comparing the multiple thermal model function values to obtain a target combination set that meets the preset function conditions, and the target combination set includes the target system parameters and target material parameters.

[0106] According to an embodiment of the present invention, the simulation model may include a simulation liquid cooling system and a simulation processor. The simulation liquid cooling system and the simulation processor are respectively obtained by performing three-dimensional modeling on the liquid cooling system and the processor. The simulation liquid cooling system is configured with a control mechanism of the liquid cooling system, and the simulation processor is configured with the same tasks and operating states as the processor. At the same time, the parameters of the phase change material are also set in the simulation to simulate the heat dissipation of the liquid cooling system to the processor under different temperatures and operating parameters.

[0107] For example, to make the simulation more accurate, the simulation model can be meshed, and the mesh size can be 0.1 mm. That is, the spatial discretization accuracy of the simulation model can be a 0.1 mm mesh.

[0108] According to an embodiment of the present invention, both the simulated temperature and the simulated operating parameters can be represented by a variety of initial system parameters and initial material parameters. For example, the simulated temperature can be T’(k 1, k2), and the simulated operating parameters can be P’(k 1, k2). Therefore, based on a large number of simulated temperatures and simulated operating parameters, a combined data set of a variety of initial system parameters and initial material parameters can be determined.

[0109] Exemplarily, the formula of the thermal model F can be as follows:

[0110] F = α·T’(k 1, k2)+ β·P’(k 1, k2)+γ QoS’(k 1, k2) (4)

[0111] According to an embodiment of the present invention, QoS’(k 1, k2) represents the simulated service quality of the simulation processor (such as simulation delay, simulation throughput, etc.). α, β, and γ are the weights of the simulated temperature, the simulated operating parameters, and the simulated service quality, respectively. For example, α = 0.6, β = 0.3, and γ = 0.1.

[0112] According to an embodiment of the present invention, the thermal model can be adjusted according to the actual situation. For example, the thermal model can be improved by considering the heat generation due to eddy current loss.

[0113] According to an embodiment of the present invention, the preset function condition can be: the minimum value of multiple thermal model function values.

[0114] According to an embodiment of the present invention, the simulated operating parameters and the simulated temperature are obtained by performing a simulation using the simulation model of the processor. Based on the simulated temperature and the simulated operating parameters, a combined data set of a variety of initial system parameters and initial material parameters can be determined, which can reduce the time for experiments using a liquid cooling system and a processor cluster and improve the efficiency of obtaining the combined data set. Inputting the multiple combined data sets into the thermal model respectively to obtain multiple thermal model function values; comparing the multiple thermal model function values to obtain the target system parameters and target material parameters that meet the preset function condition. Therefore, the target system parameters and target material parameters that meet the preset function condition can be obtained using the simulated operating parameters and the simulated temperature simulated by the simulation model, which is convenient for subsequent task scheduling of the processor and improves the task scheduling efficiency.

[0115] According to an embodiment of the present invention, determining the predicted temperature and predicted operating parameters of a processor at a future moment based on the current temperatures and current operating parameters of multiple processors in a processor cluster includes: inputting the current temperatures and current operating parameters into a prediction model to obtain the predicted temperature and predicted operating parameters, where the prediction model is trained based on simulated operating parameters and simulated temperatures.

[0116] According to an embodiment of the present invention, the training data of the prediction model can be 100,000 sets of simulated temperatures and simulated operating parameters. The types of the simulated operating parameters are the same as those of the current operating parameters. For example, the current operating parameters can be the current power consumption and the current operating load. The simulated operating parameters can be the simulated power consumption and the simulated operating load.

[0117] According to an embodiment of the present invention, the prediction model can be a machine learning model. For example, the machine learning model has a three-layer long short-term memory neural network with 128 hidden units.

[0118] Prediction constraint conditions can be set. For example, the predicted temperature is less than 85 degrees Celsius, the predicted energy consumption is less than 1.2, and the QoS (Quality of Service) compliance rate is maintained at greater than or equal to 99.9%.

[0119] Exemplarily, the processor cluster is deployed in a cabinet, and there are 640 processors in the processor cluster. The number of processing cores of the processor cluster can be 61,440.

[0120] The liquid cooling system controls the speed of the main circulation pump (such as a magnetic levitation centrifugal pump) through PID to output the coolant. The flow rate of the coolant can be 3000 liters per minute, and the head is 60 meters.

[0121] The coolant can be composed of engineered deionized water with a thermal conductivity of 0.6 watts per meter per Kelvin and nano-aluminum oxide particles with a concentration of 2%.

[0122] For task scheduling, the processor cluster is connected to a circuit board or a server via a dedicated computing power transmission interface (Compute Fabric Interface, CFI) through a data physical link. The circuit board or the server is used to execute the task scheduling method of the processor. For example, the server can be Kubernetes, and the scheduling period can be shortened to 50 microseconds.

[0123] The dedicated computing power transmission interface can be an optoelectronic fusion interface. The electrical signal channel is PCIe 6.0 ×16 (64 gigabits per second transmission); the optical signal channel is a silicon optical engine (wavelength of 1310 nanometers, bandwidth of 800 gigabits per second).

[0124] The computing power transmission interface is communicatively connected to the ports of the circuit board using an optical switching matrix. The optical switching matrix can be a 32×32 port silicon optical switch (single wavelength of 1.6 terabits per second). The electrical transmission channel can be a PCIe 6.0×64 aggregated link.

[0125] Based on the processor heat dissipation control strategy, determine the desired output volume of the coolant in the liquid cooling system according to the target system parameters of the liquid cooling system, the target material parameters of the phase change material, the predicted temperature, and the predicted operating parameters. The processor heat dissipation control strategy can be Model Predictive Control (MPC), with a prediction horizon of 5 seconds and a control horizon that can be 2 seconds.

[0126] The NoF (NVMe over Fabric, Non-Volatile Memory Host Controller Interface Specification based on Fibre Channel) protocol can be used to schedule tasks for multiple processors. First, initialize the communication protocol, start the NoF protocol stack, load the kernel module, configure the cache coherence domain, and establish a virtual NUMA (Non-Uniform Memory Access) topology.

[0127] Figure 3 Shows the hierarchical interaction flowchart of the NoF protocol according to an embodiment of the present invention.

[0128] As Figure 3 shown, the NoF protocol is divided into a physical layer 310, a topology layer 320, a cache layer 330, and an application layer 340.

[0129] Use the task classifier in the application layer 340 of the NoF protocol to classify the tasks to be processed received by the processor cluster.

[0130] For example, identify the computing type of the task. The computing type can be FP32, INT8, control instructions, etc. Package and send the task source data to the cache layer 330.

[0131] Use the improved MESI (Modified-Exclusive-Shared-Invalid) protocol state machine in the cache layer 330 to determine whether the task being processed by the processor is completed. In the case where the task processing is completed, transfer the task to be processed to the topology layer 320.

[0132] Utilize a virtual NUMA domain builder within the topology layer 320. The virtual NUMA domain builder can be a tool or component for simulating a physical non-uniform memory access (NUMA) architecture in a virtualized environment. Divide multiple processors in a processor cluster into multiple virtual NUMA nodes, enabling virtual machines to perceive and utilize the NUMA topology, thereby optimizing resource allocation and performance.

[0133] Utilize the virtual NUMA domain builder to obtain the positional distance between a candidate processor and the first processor, thereby determining the processor corresponding to the task to be processed according to the task scheduling result obtained by the task scheduling method for the server to execute the processor.

[0134] Utilize the physical layer 310 to transfer the task to be processed to the corresponding processor.

[0135] The physical layer 310 includes a hybrid signal encoder and a forward error correction module. The forward error correction coding scheme can be Reed-Solomon(255, 239). For example, the hybrid signal encoder can be a hybrid transmission scheme that applies PAM4 (Pulse Amplitude Modulation with 4 Levels, four-level pulse amplitude modulation) electrical signals combined with NRZ (Non-Return-to-Zero Optical Signal, non-return-to-zero code optical signal) optical signals.

[0136] The topology layer 320 includes a fault domain isolation module. The heartbeat detection period of the fault domain isolation module is 10 microseconds. For example, in the case where the action time of the electronic switch is less than 10 microseconds, the power supply and data link of the faulty processor cluster can be quickly cut off to prevent the fault from spreading to other modules, while protecting the backup link from being contaminated. The topology reconstruction delay is less than 100 microseconds.

[0137] The cache layer 330 has a cache line locking mechanism across processor clusters. The timeout detection window of the cache line locking mechanism is 100 nanoseconds.

[0138] The application layer 340 has a QoS marking engine that can define priorities for tasks to be processed. For example, the tasks to be processed can be divided into 3 levels of priorities, from high to low are the tasks to be processed with real-time <5 milliseconds, the tasks to be processed with high throughput <50 milliseconds, and the tasks to be processed with the attribute of background processing.

[0139] The layers interact through standard interfaces. The hot label is a metadata identifier carrying task characteristics, used to transfer key information related to task scheduling and resource allocation. In the interaction between the application layer 340 and the buffer layer 330, cross-layer information transfer is achieved by adding an 8-byte fixed-length field to the task metadata header. The 8-byte fixed-length field can include task priority, resource requirements, and timestamp, etc.

[0140] The cross-layer interaction mechanism between the topology layer 320 and the physical layer 310, the core of which is to feedback the physical link quality status through the bit error rate index monitored in real time by the physical layer, so as to ensure the high reliability of the data transmission of the task to be processed.

[0141] When it is detected that the heartbeat loss of the processor cluster is greater than 3 times, the fault tolerance process is triggered. The fault domain isolation module (the action time of the electronic switch is less than 10 microseconds) can quickly cut off the power supply and data link of the faulty processor cluster. After isolating the faulty module, the task to be processed is dynamically migrated to the backup processor cluster, the data path is re-established, and the system function is restored, and the topology reconstruction delay is less than 100 milliseconds.

[0142] Through end-to-end cyclic redundancy check, a specified polynomial is used. The task to be processed is written into the main storage medium and the mirror storage medium at the same time to ensure the consistency of the dual-copy data, and the total delay of the dual-write operation does not exceed 15 nanoseconds.

[0143] According to an embodiment of the present invention, the above method further includes, when a failure of the liquid cooling system is detected: starting the sprinkler device; and / or controlling multiple processors to reduce the operating frequency.

[0144] According to an embodiment of the present invention, the cooling setting for the processor cluster can be: when the liquid cooling system is operating normally, the coolant in the liquid cooling system is used for heat dissipation; when the liquid cooling system fails, the sprinkler device is started.

[0145] For example, the failure of the liquid cooling system can be blockage of the coolant pipeline in the liquid cooling system, failure of the control unit, insufficient coolant, etc.

[0146] For example, when it is detected that the current temperature is greater than 85 degrees Celsius, the sprinkler device can be automatically started. The over-temperature protection response time can be set to be less than 10 milliseconds.

[0147] According to an embodiment of the present invention, due to the failure of the liquid cooling system, the heat dissipation of multiple processors in the processor cluster cannot be timely performed, and the sprinkler device can be started so that the processors can still be cooled.

[0148] For example, the nozzle aperture of the sprinkler device can be 50 microns, and the array density can be 200 per square centimeter. The coolant of the sprinkler device can be perfluoromethyl hexanone, the boiling point of perfluoromethyl hexanone is 49 degrees Celsius, and the latent heat of vaporization is 140 kJ per kilogram.

[0149] According to an embodiment of the present invention, due to the failure of the liquid cooling system, multiple processors can be controlled to reduce the operating frequency, thereby avoiding losses of the processors.

[0150] According to an embodiment of the present invention, the above method further includes: determining that the liquid cooling system fails when the temperature change rate of the current temperature meets a preset temperature fault condition.

[0151] According to an embodiment of the present invention, the predicted temperature fault condition may be that the temperature change rate is greater than 15 degrees.

[0152] For example, when it is detected that the temperature change rate of the current temperature is greater than 15 degrees, it is determined that the liquid cooling system fails.

[0153] Exemplarily, the liquid cooling system may include a microchannel cold plate with coolant pipes on it. The cross-sectional size of the pipes may be 100 microns wide × 300 microns deep. The surface roughness of the pipes is less than or equal to 0.8 microns. The material of the pipes may be copper-tungsten alloy, and when the thickness of the copper-tungsten alloy is 1 meter, the temperature difference between both sides is 1 Kelvin, and the heat transferred per second through an area of 1 square meter is 320 joules.

[0154] The coolant circulation path may be an asymmetric flow channel, and the inlet pipe diameter of the coolant pipe may be 8 milliseconds, and the outlet pipe diameter is 12 milliseconds. There is an anti-corrosion coating inside the coolant pipe. For example, the anti-corrosion coating may be a nickel-phosphorus alloy layer with a thickness of 50 microns.

[0155] According to an embodiment of the present invention, multiple processors are respectively encapsulated in a detachable member with a phase change material coating, and mounting grooves are formed on the surface of the detachable member for fixing the detachable member relative to the liquid cooling system.

[0156] According to an embodiment of the present invention, mounting grooves are formed on the surface of the detachable member, and the insertion and extraction force may be less than or equal to 15 Newtons, thus realizing an anti-misinsertion design. For example, the mounting grooves are used to mount the detachable member on the coolant pipe. The shape of the mounting grooves matches the shape of the coolant pipe.

[0157] For example, the size of the detachable member may be 75 mm × 150 mm, the surface may be a liquid metal filling layer with a thickness of 0.2 mm, and a temperature sensor may be embedded in the back. The accuracy of the temperature sensor may be ±0.5 degrees Celsius.

[0158] For example, mounting grooves suitable for multiple processors may be provided. The processor is configured with a 12V DC input bidirectional Buck-Boost (step-down and step-up) voltage regulator. The dynamic voltage regulation range is 0.6V - 1.8V, and the step accuracy is 10 mV. The processor is also configured with a signal regenerator, supporting 0 - 15 dB insertion loss compensation to ensure high-speed signal integrity.

[0159] According to an embodiment of the present invention, the phase change material coating may be provided on the contact surface of the processor located in the detachable member.

[0160] For example, the phase change material can be gallium-based liquid metal (phase change point: 29.8 °C), and the activation threshold of the phase change material can be that the temperature difference between the current temperature and the phase change point is greater than or equal to 15 °C, at which point the gallium-based liquid metal begins the phase change process to achieve heat absorption or release.

[0161] The phase change material coating can be injected with 300 L of gallium-based alloy with a purity of 99.999%. When the temperature difference between the current temperature and the phase change point is equal to 20 °C, the heat storage density is 800 kJ per cubic meter.

[0162] There can be gallium-based liquid metal in the microchannel flow path, and the pressure difference in the microchannel flow path is maintained in the range of 0.2 - 0.5 MPa.

[0163] According to an embodiment of the present invention, the processor cluster can be mechanically connected to the liquid cooling system through a liquid cooling quick connector. Under sealed conditions, the detected internal pressure is greater than 3 bar.

[0164] According to an embodiment of the present invention, the processor cluster uses an MPO-24 multi-core connector to achieve high-speed data transmission. The loss of a single connection point is less than 0.2 dB. dB is used to represent the degree of power attenuation during signal transmission. 0.2 dB indicates that the signal loss at each connection point is small and the transmission efficiency is high.

[0165] According to an embodiment of the present invention, the liquid cooling plate in the liquid cooling system can have a quick-release heat conduction connector, and the contact thermal resistance of the quick-release heat conduction connector is less than 0.05 °C·cm² / W.

[0166] According to an embodiment of the present invention, the detachable component includes an upper housing and a lower housing. The gold finger bumps on the upper housing are adapted to the gold finger grooves on the lower housing, and the processor is electrically connected to the circuit board via the gold finger contacts in the gold finger grooves.

[0167] According to an embodiment of the present invention, the circuit board can be used to execute the task scheduling method of the processor. An optoelectronic hybrid signal repeater is encapsulated on the circuit board. The electrical signal gain of the optoelectronic hybrid signal repeater is +6 dB, and the optical signal pre-emphasis is 3 dB. The circuit board is also encapsulated with a priority arbiter, which can be a 128-level queue management based on FPGA, that is, it arbitrates the requests of 128 input channels for priority to achieve queue management, conflict resolution, and orderly scheduling.

[0168] Figure 4A Shows a schematic diagram of the gold finger groove of the lower housing according to an embodiment of the present invention.

[0169] As Figure 4AAs shown, there may be gold finger contacts in the gold finger groove 411 on the lower housing 410 to transmit the operating parameters and temperature of the processor to the circuit board or server. There is an anti-fooling mark 412 on the lower housing 410. The anti-fooling mark 412 can be located at the 1 / 5 position of each side when rotating in the clockwise direction, which can prevent the installation error of the detachable component.

[0170] Figure 4B The schematic diagram of the gold finger bump of the upper housing according to the embodiment of the present invention is shown.

[0171] As Figure 4B shown, the gold finger bump 421 protrudes relative to other parts of the upper housing 420 to be adapted to the gold finger groove 411.

[0172] It should be noted that the upper housing 420 can also be of other shapes. For example, it is a rectangular housing with the same shape and size as Figure 4A the lower housing 410.

[0173] Figure 5 The schematic diagram of the detachable component according to the embodiment of the present invention is shown.

[0174] As Figure 5 shown, the detachable component 400 has an installation groove 430 for installing the detachable component on the coolant pipeline. The groove shape of the installation groove matches the shape of the coolant pipeline.

[0175] Figure 6 The schematic diagram of the liquid cooling quick connector according to the embodiment of the present invention is shown.

[0176] As Figure 6 shown, multiple coolant pipelines can be connected through the liquid cooling quick connector 610. There are also 3 groups of Hall sensors 620 on the coolant pipeline to ensure that the sealing pressure is greater than 2.5 MPa.

[0177] Figure 7 The structural block diagram of the task scheduling device of the processor according to the embodiment of the present invention is shown.

[0178] As Figure 7 shown, the task scheduling device 700 of the processor in this embodiment includes a first determination module 710, a second determination module 720, an acquisition module 730, and a task scheduling module 740.

[0179] The first determination module 710 is used to determine the predicted temperature and predicted operating parameters of the processor at a future moment according to the current temperature and current operating parameters of multiple processors in the processor cluster. In one embodiment, the first determination module 710 can be used to execute the operation S210 described above, which will not be elaborated here.

[0180] The second determination module 720 is configured to determine an expected output of the coolant in the liquid cooling system according to the target system parameters of the liquid cooling system, the target material parameters of the phase change material, the predicted temperature, and the predicted operating parameters. The liquid cooling system is used to dissipate heat from the processor cluster. In one embodiment, the second determination module 720 may be configured to perform the operation S220 described above, which will not be elaborated here.

[0181] The obtaining module 730 is configured to obtain an actual output of the coolant at a future moment based on the control mechanism of the liquid cooling system, according to the expected output, the predicted temperature, and the current temperature. In one embodiment, the obtaining module 730 may be configured to perform the operation S230 described above, which will not be elaborated here.

[0182] The task scheduling module 740 is configured to perform task scheduling for a plurality of processors according to a cooling error between the expected output and the actual output, the predicted temperature, and the predicted operating parameters. In one embodiment, the task scheduling module 740 may be configured to perform the operation S240 described above, which will not be elaborated here.

[0183] According to an embodiment of the present invention, the task scheduling module 740 includes a first determination sub-module, a second determination sub-module, a third determination sub-module, and a migration sub-module. The first determination sub-module is configured to determine a first processor whose predicted temperature satisfies the temperature safety threshold condition from the plurality of processors according to the cooling error. The second determination sub-module is configured to determine a task to be migrated from at least one task to be processed by the first processor according to the predicted temperature of the first processor. The third determination sub-module is configured to determine a second processor that satisfies the migration condition from the plurality of processors according to the predicted temperatures and predicted operating parameters of the plurality of processors. The migration sub-module is configured to migrate the task to be migrated to the second processor for processing.

[0184] According to an embodiment of the present invention, the first determination sub-module includes a first determination unit. The first determination unit is configured to determine a first processor whose predicted temperature satisfies the temperature safety threshold condition from the processor cluster when there is a processor in the processor cluster with a cooling error greater than the error threshold.

[0185] According to an embodiment of the present invention, there are multiple second processors, the predicted operating parameters include at least one of the following: predicted power consumption and predicted computing load, and the migration condition includes a temperature difference condition and an operating condition. The third determination sub-module includes a second determination unit and a third determination unit. The second determination unit is configured to determine candidate processors that satisfy the temperature difference condition from the plurality of processors according to the predicted temperatures of the plurality of processors. The third determination unit is configured to determine a second processor that satisfies the operating condition from the multiple candidate processors according to the positional distance between the candidate processor and the first processor, the predicted operating parameters of the candidate, and the time slice length. The time slice length is determined according to the number of processing cores of the processor and the duration required to process at least one task.

[0186] According to an embodiment of the present invention, the second determination sub-module includes a fourth determination unit and a fifth determination unit. The fourth determination unit is configured to determine the amount of tasks completed by the first processor under the temperature safety threshold condition according to the predicted temperature of the first processor, and the fifth determination unit is configured to determine the task to be migrated from at least one task processed by the first processor according to the actual output amount and the amount of tasks.

[0187] According to an embodiment of the present invention, the second determination module 720 includes a fourth determination sub-module, a fifth determination sub-module, and a sixth determination sub-module. The fourth determination sub-module is configured to determine the heat dissipation driven by power consumption of the processor according to the target system parameters and the predicted operating parameters. The fifth determination sub-module is configured to determine the temperature compensation amount of the processor according to the target material parameters, the predicted temperature, and the temperature safety threshold condition. The sixth determination sub-module is configured to determine the expected output amount according to the heat dissipation driven by power consumption and the temperature compensation amount.

[0188] According to an embodiment of the present invention, the target system parameters and the target material parameters are determined based on the following steps: determining a combined data set of multiple initial system parameters and initial material parameters according to the simulated temperature and the simulated operating parameters, where the simulated operating parameters and the simulated temperature are obtained by performing a simulation using a simulation model of the processor; inputting the multiple combined data sets into the thermal model respectively to obtain multiple thermal model function values; comparing the multiple thermal model function values to obtain a target combination set that satisfies the preset function condition, and the target combination set includes the target system parameters and the target material parameters.

[0189] According to an embodiment of the present invention, the first determination module 710 includes an acquisition sub-module. The acquisition sub-module is configured to input the current temperature and the current operating parameters into the prediction model to obtain the predicted temperature and the predicted operating parameters, and the prediction model is trained based on the simulated operating parameters and the simulated temperature.

[0190] According to an embodiment of the present invention, the above-mentioned device further includes a start module and a control module. The start module is configured to, when detecting a failure of the liquid cooling system: start the spraying device; and / or the control module is configured to control multiple processors to reduce the operating frequency.

[0191] According to an embodiment of the present invention, the above-mentioned device further includes a third determination module, and the third determination module is configured to determine that the liquid cooling system has a failure when the temperature change rate of the current temperature satisfies the preset temperature failure condition.

[0192] According to an embodiment of the present invention, multiple processors are respectively encapsulated in detachable components with a phase change material coating, and mounting grooves are formed on the surface of the detachable components, and the mounting grooves are used to fix the detachable components relative to the liquid cooling system.

[0193] According to an embodiment of the present invention, the detachable member includes an upper housing and a lower housing. The gold finger bumps of the upper housing are adapted to the gold finger grooves of the lower housing. The processor is electrically connected to the circuit board via the gold finger contacts in the gold finger grooves.

[0194] According to an embodiment of the present invention, any multiple of the first determination module 710, the second determination module 720, the acquisition module 730, and the task scheduling module 740 can be combined and implemented in one module, or any one of them can be split into multiple modules. Or, at least part of the functions of one or more of these modules can be combined with at least part of the functions of other modules and implemented in one module. According to an embodiment of the present invention, at least one of the determination module 710, the prediction module 720, the acquisition module 730, and the task scheduling module 740 can be at least partially implemented as a hardware circuit, such as a field programmable gate array (FPGA), a programmable logic array (PLA), a system on chip, a system on substrate, a system on package, an application specific integrated circuit (ASIC), or can be implemented by any other reasonable way of integrating or packaging circuits, etc., by hardware or firmware, or implemented in any one of the three implementation manners of software, hardware, and firmware, or in an appropriate combination of any several of them. Or, at least one of the first determination module 710, the second determination module 720, the acquisition module 730, and the task scheduling module 740 can be at least partially implemented as a computer program module, and when the computer program module is run, it can execute the corresponding functions.

[0195] Figure 8 A block diagram of an electronic device suitable for implementing a task scheduling method of a processor according to an embodiment of the present invention is shown.

[0196] As Figure 8 shown, the electronic device 800 according to an embodiment of the present invention includes a processor 801, which can perform various appropriate actions and processes according to a program stored in a read only memory (ROM) 802 or a program loaded from a storage section 808 into a random access memory (RAM) 803. The processor 801 can include, for example, a general microprocessor (such as a CPU), an instruction set processor, and / or a related chipset, and / or a dedicated microprocessor (such as an application specific integrated circuit (ASIC)), etc. The processor 801 can also include on-board memory for caching purposes. The processor 801 can include a single processing unit or multiple processing units for performing different actions of the method flow according to an embodiment of the present invention.

[0197] In the RAM 803, various programs and data required for the operation of the electronic device 800 are stored. The processor 801, the ROM 802, and the RAM 803 are connected to each other via a bus 804. The processor 801 performs various operations of the method flow according to the embodiments of the present invention by executing the programs in the ROM 802 and / or the RAM 803. It should be noted that the programs may also be stored in one or more memories other than the ROM 802 and the RAM 803. The processor 801 may also perform various operations of the method flow according to the embodiments of the present invention by executing the programs stored in the one or more memories.

[0198] According to an embodiment of the present invention, the electronic device 800 may further include an input / output (I / O) interface 805, and the input / output (I / O) interface 805 is also connected to the bus 804. The electronic device 800 may further include one or more of the following components connected to the input / output (I / O) interface 805: an input portion 806 including a keyboard, a mouse, etc.; an output portion 807 including a cathode ray tube (CRT), a liquid crystal display (LCD), etc. and a speaker, etc.; a storage portion 808 including a hard disk, etc.; and a communication portion 809 including a network interface card such as a LAN card, a modem, etc. The communication portion 809 performs communication processing via a network such as the Internet. A drive 810 is also connected to the input / output (I / O) interface 805 as needed. A removable medium 811, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 810 as needed so that a computer program read from it can be installed into the storage portion 808 as needed.

[0199] The present invention also provides a computer-readable storage medium, which may be included in the device / apparatus / system described in the above embodiments; or may exist alone without being assembled into the device / apparatus / system. The above computer-readable storage medium carries one or more programs, and when the one or more programs are executed, the method according to the embodiments of the present invention is implemented.

[0200] According to an embodiment of the present invention, the computer-readable storage medium may be a non-volatile computer-readable storage medium, for example, it may include but is not limited to: portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the above. In the present invention, the computer-readable storage medium may be any tangible medium that contains or stores a program, and this program can be used by or in conjunction with an instruction execution system, apparatus, or device. For example, according to an embodiment of the present invention, the computer-readable storage medium may include the above-described ROM 802 and / or RAM 803 and / or one or more memories other than ROM 802 and RAM 803.

[0201] An embodiment of the present invention further includes a computer program product, which includes a computer program, and the computer program contains program code for executing the method shown in the flowchart. When the computer program product runs in a computer system, the program code is used to enable the computer system to implement the task scheduling method of the processor provided by the embodiment of the present invention.

[0202] When the computer program is executed by the processor 801, it executes the above functions defined in the system / apparatus of the embodiment of the present invention. According to an embodiment of the present invention, the above-described systems, apparatuses, modules, units, etc. can be implemented by computer program modules.

[0203] In one embodiment, the computer program can rely on tangible storage media such as optical storage devices and magnetic storage devices. In another embodiment, the computer program can also be transmitted and distributed in the form of a signal on a network medium, and is downloaded and installed through the communication part 809, and / or installed from the removable medium 811. The program code contained in the computer program can be transmitted by any suitable network medium, including but not limited to: wireless, wired, etc., or any suitable combination of the above.

[0204] In such an embodiment, the computer program can be downloaded and installed from the network through the communication part 809, and / or installed from the removable medium 811. When the computer program is executed by the processor 801, it executes the above functions defined in the system of the embodiment of the present invention. According to an embodiment of the present invention, the above-described systems, devices, apparatuses, modules, units, etc. can be implemented by computer program modules.

[0205] According to embodiments of the present invention, program code for executing the computer programs provided by the embodiments of the present invention can be written in any combination of one or more programming languages. Specifically, these computing programs can be implemented using high-level procedural and / or object-oriented programming languages, and / or assembly / machine languages. The programming languages include, but are not limited to, such as Java, C++, Python, the "C" language, or similar programming languages. The program code can be executed entirely on the user's computing device, partially on the user's device, partially on a remote computing device, or entirely on a remote computing device or server. In the case of a remote computing device, the remote computing device can be connected to the user's computing device through any type of network, including a local area network (LAN) or a wide area network (WAN), or it can be connected to an external computing device (for example, by using an Internet service provider to connect through the Internet).

[0206] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in the flowchart or block diagram may represent a module, a program segment, or a part of code, and the above-mentioned module, program segment, or part of code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than that marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram or flowchart, as well as the combination of blocks in the block diagram or flowchart, can be implemented by a dedicated hardware-based system for performing the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.

[0207] Those skilled in the art can understand that the features described in the various embodiments of the present invention can be combined and / or combined in various ways, even if such combinations or combinations are not explicitly described in the present invention. In particular, without departing from the spirit and teachings of the present invention, the features described in the various embodiments of the present invention can be combined and / or combined in various ways. All such combinations and / or combinations fall within the scope of the present invention.

[0208] The above describes the embodiments of the present invention. However, these embodiments are for illustrative purposes only and are not intended to limit the scope of the present invention. Although the embodiments are described separately above, this does not mean that the measures in each embodiment cannot be used advantageously in combination. Without departing from the scope of the present invention, those skilled in the art can make various substitutions and modifications, and all such substitutions and modifications should fall within the scope of the present invention.

Claims

1. A task scheduling method for a processor, characterized in that: The method comprises: Determining predicted temperatures and predicted operating parameters of the processors at a future time according to current temperatures and current operating parameters of the plurality of processors in the processor cluster; Determining an expected output of a coolant in the liquid cooling system according to a target system parameter of the liquid cooling system, a target material parameter of the phase change material, the predicted temperature, and the predicted operating parameter, wherein the liquid cooling system is used to dissipate heat for the processor cluster; Based on the control mechanism of the liquid cooling system, according to the expected output, the predicted temperature and the current temperature, the actual output of the coolant at the future time is obtained; Task scheduling is performed on the plurality of processors according to a cooling error between the expected output and the actual output, the predicted temperature, and the predicted operating parameters.

2. The method according to claim 1, characterized in that: The step of performing task scheduling on the plurality of processors according to the cooling error between the expected output and the actual output, the predicted temperature, and the predicted operating parameter comprises: According to the cooling error, determining, from the plurality of processors, a first processor whose predicted temperature satisfies a temperature safety threshold condition; Determining, according to the predicted temperature of the first processor, a task to be migrated from at least one task to be processed by the first processor; Determining a second processor that meets a migration condition from among the plurality of processors according to the predicted temperatures and the predicted operating parameters of the plurality of processors; Migrate the task to be migrated to the second processor for processing.

3. The method according to claim 2, characterized in that: The step of determining, according to the cooling error, from the plurality of processors, a first processor whose predicted temperature satisfies a temperature safety threshold condition, comprises: In the case that there is a processor in the processor cluster whose cooling error is greater than the error threshold, a first processor whose predicted temperature satisfies the temperature safety threshold condition is determined from the processor cluster.

4. The method according to claim 2, characterized in that: The second processor includes a plurality of processors, the predicted operating parameters include at least one of the following: predicted power consumption and predicted computing load, and the migration conditions include temperature difference conditions and operating conditions; The step of determining a second processor that meets a migration condition from the plurality of processors according to the predicted temperatures and the predicted operating parameters of the plurality of processors includes: Determine, according to the predicted temperatures of the plurality of processors, a candidate processor that satisfies a temperature difference condition from among the plurality of processors; Based on the location distance between the candidate processor and the first processor, the candidate's predicted operating parameters and the time slice length, the second processor that meets the operating conditions is determined from multiple candidate processors, and the time slice length is determined based on the number of processing cores of the processor and the time required to process the at least one task.

5. The method according to claim 2, characterized in that: The step of determining the task to be migrated from at least one task processed by the first processor according to the predicted temperature of the first processor includes: determining, according to the predicted temperature of the first processor, an amount of tasks to be completed by the first processor under the temperature safety threshold condition; The task to be migrated is determined from at least one task processed by the first processor according to the actual output amount and the task amount.

6. The method according to claim 1, characterized in that: The step of determining the expected output of the coolant in the liquid cooling system according to the target system parameter of the liquid cooling system, the target material parameter of the phase change material, the predicted temperature and the predicted operating parameter comprises: Determining the power consumption driving heat dissipation of the processor according to the target system parameters and the predicted operating parameters; Determining a temperature compensation amount for a processor according to the target material parameter, the predicted temperature and a temperature safety threshold condition; The expected output amount is determined according to the power consumption driving heat dissipation amount and the temperature compensation amount.

7. The method according to claim 1, characterized in that: The target system parameters and the target material parameters are determined based on the following steps: Determining a combined data set of multiple initial system parameters and initial material parameters according to a simulation temperature and a simulation operating parameter, wherein the simulation operating parameter and the simulation temperature are obtained by performing simulation using a simulation model of a processor; Inputting the plurality of combined data sets into a thermal model respectively to obtain a plurality of thermal model function values; A plurality of thermal model function values ​​are compared to obtain a target combination set that meets a preset function condition, wherein the target combination set includes the target system parameters and the target material parameters.

8. The method according to claim 7, characterized in that: The step of determining predicted temperatures and predicted operating parameters of the processors at a future time according to current temperatures and current operating parameters of the multiple processors in the processor cluster includes: The current temperature and the current operating parameters are input into a prediction model to obtain the predicted temperature and the predicted operating parameters, wherein the prediction model is trained based on simulation operating parameters and simulation temperature.

9. The method according to claim 1, characterized in that: The method further comprises, in the event that a failure of the liquid cooling system is detected: Activate the sprinklers; and / or The plurality of processors are controlled to reduce operating frequencies.

10. The method according to claim 9, characterized in that The method further comprises: When the temperature change rate of the current temperature meets a preset temperature fault condition, it is determined that a fault occurs in the liquid cooling system.

11. The method according to claim 1, characterized in that: The plurality of processors are respectively encapsulated in a detachable component having a phase change material coating, and a mounting groove is formed on a surface of the detachable component, and the mounting groove is used to fix the detachable component relative to the liquid cooling system.

12. The method according to claim 11, characterized in that: The detachable component comprises an upper shell and a lower shell, the gold finger bumps of the upper shell are matched with the gold finger grooves of the lower shell, and the processor is electrically connected to the circuit board via the gold finger contacts in the gold finger grooves.

13. An electronic device comprising: one or more processors; a memory for storing one or more computer programs, It is characterized in that the one or more processors execute the one or more computer programs to implement the steps of the method according to any one of claims 1 to 12.

14. A computer-readable storage medium having a computer program or instruction stored thereon, characterized in that: When the computer program or instruction is executed by a processor, the steps of the method according to any one of claims 1 to 12 are implemented.

15. A computer program product comprising a computer program or instructions, characterized in that When the computer program or instruction is executed by a processor, the steps of the method according to any one of claims 1 to 12 are implemented.

Citation Information

Patent Citations

  • Task scheduling method and device

    CN113254172A

  • Task scheduling method, migration method and system

    CN115114016A

  • Task scheduling method and device, equipment and storage medium

    CN117519919A

  • Liquid cooling server intelligent temperature control method based on local software monitoring and liquid cooling server

    CN119536482A

  • Host power consumption and temperature balance control method and system, medium and program product

    CN119536958A

Cited By

  • Heat dissipation management method, electronic equipment, storage medium and program product

    CN120762509A

  • Heat dissipation management method, electronic device, storage medium, and program product

    CN120762509B

  • AI computer server and use method thereof

    CN121326113A