Processor task scheduling method, device, storage medium and program product

By predicting the future temperature and operating parameters of the processor, and adjusting the coolant output in combination with the parameters of the liquid-cooled system and phase-change material, uniform heat dissipation of the processor cluster is achieved, solving the problem of high temperatures in some processors and improving the heat dissipation efficiency.

CN120179055BActive Publication Date: 2025-08-12INSPUR SUZHOU INTELLIGENT TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510654914.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-05-21
Publication Date
2025-08-12
Estimated Expiration
2045-05-21

AI Technical Summary

Technical Problem

The temperature after heat dissipation of some processors in the processor cluster is high, while the temperature of the other processors is low, resulting in low heat dissipation efficiency.

Method used

By predicting future temperatures and operating parameters based on the current temperature and operating parameters of the processor cluster, determining the expected output of the coolant in combination with the target system and phase change material parameters of the liquid cooling system, and adjusting the actual output based on the control mechanism of the liquid cooling system, and performing task scheduling to improve heat dissipation efficiency.

Benefits of technology

Improve the heat dissipation efficiency of the processor cluster, ensure that each processor dissipates heat evenly, and avoids equipment losses.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120179055B_ABST
    Figure CN120179055B_ABST
Patent Text Reader

Abstract

The present invention provides a task scheduling method, device, storage medium, and program product for a processor, which can be applied to the field of heat dissipation technology. The task scheduling method for the processor includes: determining the predicted temperature and predicted operating parameters of the processors based on the current temperature and current operating parameters of multiple processors in a processor cluster; determining the expected output of the coolant in the liquid cooling system based on the target system parameters of the liquid cooling system, the target material parameters of the phase change material, the predicted temperature, and the predicted operating parameters; obtaining the actual output of the coolant at a future time based on the expected output, the predicted temperature, and the current temperature based on the control mechanism of the liquid cooling system; and scheduling tasks for multiple processors based on the cooling error between the expected output and the actual output, the predicted temperature, and the predicted operating parameters. Combining the control of the liquid cooling system with the task scheduling of the processors to dissipate heat for the processor cluster improves the heat dissipation efficiency of the processor cluster.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of heat dissipation technology, and more particularly to a task scheduling method, device, storage medium and program product for a processor. Background Art

[0002] Multiple processors in a processor cluster can work together to achieve high-concurrency task execution or complex algorithm acceleration. Because each processor in the processor cluster has different tasks and thus different operating conditions, the heat generated by each processor is also different.

[0003] The processor cluster is cooled by a liquid cooling system. However, the temperature of some processors remains high after cooling, while the temperature of other processors remains low after cooling, resulting in low cooling efficiency of the processor cluster. Summary of the Invention

[0004] In view of the above problems, the present invention provides a task scheduling method, device, storage medium and program product for a processor.

[0005] According to a first aspect of the present invention, a method for scheduling processor tasks is provided, comprising: determining predicted temperatures and predicted operating parameters of multiple processors in a processor cluster at a future time based on current temperatures and current operating parameters of the processors; determining an expected output of coolant in a liquid cooling system based on target system parameters of the liquid cooling system, target material parameters of a phase change material, the predicted temperature, and the predicted operating parameters, the liquid cooling system being used to dissipate heat from the processor cluster; obtaining an actual output of the coolant at a future time based on the expected output, the predicted temperature, and the current temperature, based on a control mechanism of the liquid cooling system; and scheduling tasks for the multiple processors based on a cooling error between the expected output and the actual output, the predicted temperature, and the predicted operating parameters.

[0006] The second aspect of the present invention provides a task scheduling device for a processor, comprising: a first determination module, used to determine the predicted temperature and predicted operating parameters of multiple processors in a processor cluster at a future time based on the current temperature and current operating parameters of multiple processors in the processor cluster; a second determination module, used to determine the expected output of the coolant in the liquid cooling system based on the target system parameters of the liquid cooling system, the target material parameters of the phase change material, the predicted temperature and the predicted operating parameters, the liquid cooling system being used to dissipate heat for the processor cluster; an acquisition module, used to obtain the actual output of the coolant at a future time based on the expected output, the predicted temperature and the current temperature based on the control mechanism of the liquid cooling system; a task scheduling module, used to schedule tasks for multiple processors based on the cooling error between the expected output and the actual output, the predicted temperature and the predicted operating parameters.

[0007] A third aspect of the present invention provides an electronic device, comprising: one or more processors; and a memory for storing one or more computer programs, wherein the one or more processors execute the one or more computer programs to implement the steps of the above method.

[0008] The fourth aspect of the present invention further provides a computer-readable storage medium having a computer program or instructions stored thereon, which implements the steps of the above method when the computer program or instructions are executed by a processor.

[0009] The fifth aspect of the present invention further provides a computer program product, comprising a computer program or instructions, which implement the steps of the above method when executed by a processor.

[0010] According to an embodiment of the present invention, by determining the predicted temperature and predicted operating parameters of the processors at a future time based on the current temperature and current operating parameters of multiple processors in a processor cluster, the processors can be heat-dissipated predictively based on the predicted temperature and predicted operating parameters. The expected output of the coolant in the liquid cooling system is determined based on the target system parameters of the liquid cooling system, the target material parameters of the phase change material, the predicted temperature, and the predicted operating parameters, so that the expected output integrates the coolant output required for the operation of multiple processors. Based on the control mechanism of the liquid cooling system, the actual output of the coolant at a future time is obtained based on the expected output, the predicted temperature, and the current temperature. Task scheduling is performed on multiple processors in the processor cluster based on the cooling error between the expected output and the actual output, the predicted temperature, and the predicted operating parameters. The control of the liquid cooling system is combined with the task scheduling of the processors to dissipate heat for the processor cluster, thereby improving the heat dissipation efficiency of the processor cluster. BRIEF DESCRIPTION OF THE DRAWINGS

[0011] The above contents and other objects, features and advantages of the present invention will become more apparent through the following description of the embodiments of the present invention with reference to the accompanying drawings, in which:

[0012] Figure 1 A diagram showing an application scenario of a task scheduling method for a processor according to an embodiment of the present invention is shown;

[0013] Figure 2 A flowchart of a task scheduling method for a processor according to an embodiment of the present invention is shown;

[0014] Figure 3 A NoF protocol layered interaction flow chart according to an embodiment of the present invention is shown;

[0015] Figure 4A A schematic diagram showing a gold finger groove of a lower housing according to an embodiment of the present invention is shown;

[0016] Figure 4BA schematic diagram showing a gold finger bump of an upper housing according to an embodiment of the present invention is shown;

[0017] Figure 5 A schematic diagram showing a detachable component according to an embodiment of the present invention;

[0018] Figure 6 A schematic diagram of a liquid cooling quick connector according to an embodiment of the present invention is shown;

[0019] Figure 7 A structural block diagram of a task scheduling device for a processor according to an embodiment of the present invention is shown; and

[0020] Figure 8 A block diagram of an electronic device suitable for implementing a task scheduling method for a processor according to an embodiment of the present invention is shown. DETAILED DESCRIPTION

[0021] Hereinafter, embodiments of the present invention will be described with reference to the accompanying drawings. However, it should be understood that these descriptions are exemplary only and are not intended to limit the scope of the present invention. In the following detailed description, for ease of explanation, many specific details are set forth to provide a comprehensive understanding of embodiments of the present invention. However, it is apparent that one or more embodiments may also be implemented without these specific details. In addition, in the following description, descriptions of known structures and technologies are omitted to avoid unnecessary confusion of the concept of the present invention.

[0022] The terms used herein are only for describing specific embodiments and are not intended to limit the present invention. The terms "comprise", "include", etc. used herein indicate the presence of the features, steps, operations and / or components, but do not exclude the presence or addition of one or more other features, steps, operations or components.

[0023] All terms used herein (including technical and scientific terms) have the meanings commonly understood by those skilled in the art unless otherwise defined. It should be noted that the terms used herein should be interpreted as having a meaning consistent with the context of this specification and should not be interpreted in an idealized or overly rigid manner.

[0024] When expressions such as "at least one of A, B, and C, etc." are used, they should generally be interpreted in accordance with the meaning commonly understood by those skilled in the art (for example, "a system having at least one of A, B, and C" should include but is not limited to a system having A alone, B alone, C alone, A and B, A and C, B and C, and / or A, B, C, etc.).

[0025] Multiple processors in a processor cluster can work together to achieve high-concurrency task execution or complex algorithm acceleration. Because each processor in the processor cluster has different tasks and thus different operating conditions, the heat generated by each processor is also different.

[0026] The processor cluster is cooled by a liquid cooling system. However, the temperature of some processors is still high after cooling, while the temperature of other processors is low, resulting in low cooling efficiency of the processor cluster.

[0027] In view of this, an embodiment of the present invention provides a task scheduling method for a processor, including: determining the predicted temperature and predicted operating parameters of multiple processors in a processor cluster at a future time based on the current temperature and current operating parameters of multiple processors in the processor cluster; determining the expected output of the coolant in the liquid cooling system based on the target system parameters of the liquid cooling system, the target material parameters of the phase change material, the predicted temperature and the predicted operating parameters, and the liquid cooling system is used to dissipate heat for the processor cluster; based on the control mechanism of the liquid cooling system, obtaining the actual output of the coolant at a future time based on the expected output, the predicted temperature and the current temperature; and scheduling tasks for multiple processors based on the cooling error between the expected output and the actual output, the predicted temperature and the predicted operating parameters.

[0028] Figure 1 An application scenario diagram of a task scheduling method for a processor according to an embodiment of the present invention is shown.

[0029] like Figure 1 As shown, the application scenario according to this embodiment may include a processor cluster 110 , a liquid cooling system 120 , and a circuit board 130 .

[0030] The processor cluster 110 includes a plurality of processors 111 , and the plurality of processors 111 can be plugged into and unplugged from the cooling liquid pipes of the liquid cooling system 120 via mounting slots.

[0031] The liquid cooling system 120 may include a phase change heat storage tank 121, a main circulation pump 122, a manifold distributor 123 and a microchannel cold plate with coolant pipes distributed thereon. Figure 1 As shown in the figure, multiple processors 111 in processor cluster 110 are mounted on a microchannel cold plate. A phase change heat storage tank 121 can deliver phase change material to the phase change material coating. The liquid cooling system's control mechanism controls the pump speed of a main circulation pump 122, which delivers coolant to the coolant pipes in the microchannel cold plate. The pump speed can range from 200 to 5000 revolutions per minute.

[0032] Circuit board 130 can be used to obtain the current temperature and current operating parameters from processor cluster 110 and send them to a server that executes the task scheduling method for the processors. It should be noted that circuit board 130 is equipped with components that can support the task scheduling method for the processors, and circuit board 130 can also be used to schedule tasks for the processor cluster.

[0033] Circuit board 130 communicates with processor cluster 110 via an optical fiber interface array, using optoelectronic conversion units to convert optical signals into electrical signals, thereby transmitting data from processor cluster 110 to circuit board 130. Circuit board 130 also includes a PCIe (Peripheral Component Interconnect Express) switch chip to enable high-speed data exchange.

[0034] An LED (Light-Emitting Diode) array for monitoring the temperature of the processor cluster 110 is provided on the circuit board 130 .

[0035] It should be understood that Figure 1 The number of processor clusters, liquid cooling systems, and circuit boards shown in the figure is merely illustrative. Any number of processor clusters, liquid cooling systems, and circuit boards may be provided as needed.

[0036] Figure 2 A flowchart of a task scheduling method for a processor according to an embodiment of the present invention is shown.

[0037] like Figure 2 As shown, the task scheduling method of the processor of this embodiment includes operations S210 to S240.

[0038] In operation S210 , predicted temperatures and predicted operating parameters of the processors at a future time are determined based on current temperatures and current operating parameters of a plurality of processors in a processor cluster.

[0039] According to an embodiment of the present invention, a processor (Central Processing Unit) may be an accelerator processor, a heterogeneous computing processor, etc. A processor cluster may be an accelerator processor cluster, a heterogeneous computing processor cluster, etc. A heterogeneous computing processor cluster may be a hybrid cluster of central processing units and graphics processors.

[0040] According to an embodiment of the present invention, the current temperature may be obtained by a temperature sensor integrated into the processor or measured by an external temperature probe. For example, the temperature acquisition frequency of the internally integrated temperature sensor may be 10 kHz.

[0041] The current temperature and current operating parameters can be obtained from a server management tool, which can remotely query the processor's temperature, operating parameters, etc.

[0042] The current operating parameters may include current power consumption, current load, current frequency, current usage rate, etc. For example, the acquisition frequency of the current power consumption may be 1 MHz, and the acquisition frequency of the current load may be 100 Hz.

[0043] For example, the current operating parameters are used as constraints, and the current temperature of the processor is input into the temperature fitting model to obtain the predicted temperature of the processor at a future time. The temperature fitting model can be obtained by fitting the processor's historical operating constraints using the processor's historical temperature.

[0044] Using the processor's current temperature as a constraint, the processor's current operating parameters are input into the operating parameter fitting model to obtain the processor's predicted operating parameters at a future time. The operating parameter fitting model can be obtained by fitting the processor's historical operating parameters based on the processor's temperature constraint.

[0045] For example, the current temperature and current operating parameters of the processor can be input into the machine learning model to obtain the predicted temperature and predicted operating parameters of the processor at a future time.

[0046] The machine learning model can be a neural network model (such as a long short-term memory neural network), a random forest model, etc. The historical temperature series can be the historical temperature within 5 minutes, and the predicted temperature is the predicted temperature in the next 3 seconds.

[0047] In operation S220 , a desired output of the coolant in the liquid cooling system is determined based on target system parameters of the liquid cooling system, target material parameters of the phase change material, the predicted temperature, and the predicted operating parameters.

[0048] According to an embodiment of the present invention, a liquid cooling system is used to dissipate heat from a processor cluster. The liquid cooling system may include a microchannel cold plate and other emergency cooling devices. The width of the coolant channel of the microchannel cold plate is 100 microns.

[0049] According to an embodiment of the present invention, the phase change material is capable of absorbing heat from the processor.

[0050] According to an embodiment of the present invention, a target system parameter represents a change in the system performance of a liquid cooling system at a future time. A target material parameter represents a change in the material performance of a phase change material at a future time. For example, the target system parameter may be a system adaptability coefficient, which decreases as the pump speed of the liquid cooling system increases. The target material parameter may be a material adaptability coefficient, which increases as the phase change material enters a saturated state.

[0051] According to an embodiment of the present invention, the expected output of the coolant in the liquid cooling system can be determined based on the processor heat dissipation control strategy, the target system parameters of the liquid cooling system, the target material parameters of the phase change material, the predicted temperature and the predicted operating parameters.

[0052] For example, the processor heat dissipation control strategy may be to implement coordinated heat dissipation according to the temperature and operating parameters of the processor through an adaptive control algorithm.

[0053] In operation S230 , based on the control mechanism of the liquid cooling system, the actual output amount of the cooling liquid at a future time is obtained according to the expected output amount, the predicted temperature, and the current temperature.

[0054] According to an embodiment of the present invention, the control mechanism of the liquid cooling system may be PID (Proportional-Integral-Derivative) control, feedforward control, fuzzy control, or the like.

[0055] For example, the control mechanism of the liquid cooling system may be PID control, and the control mechanism formula of PID control is as follows:

[0056] (1)

[0057] Indicates the actual output of coolant at the future moment, T current Indicates the current temperature T target Represents the predicted temperature. Terr represents the temperature error value (i.e. the error value between the current temperature and the predicted temperature). K p Indicates the proportional coefficient (the value range can be 0.5-1.2). p Determines the response speed of the liquid cooling system, K p The larger the value, the more radical the adjustment. i Indicates the integral coefficient (the value range can be 0.01-0.1). K i Control error correction strength, K i The larger the value, the higher the steady-state accuracy. p (T current -T target ) is a proportional term. The greater the real-time temperature difference, the greater the flow adjustment range. K i ∫(Terr)dt is the integral term, which continuously accumulates temperature errors to eliminate steady-state errors (long-term accuracy). This suppresses temperature fluctuations (such as sudden changes in processor load) by compensating for ambient temperature changes through the integral term. The PID control algorithm adjusts the speed of the liquid cooling pump, thereby controlling the coolant output. The liquid cooling pump speed can range from 200 to 5000 revolutions per minute.

[0058] The coolant flow rate of the liquid cooling system is limited. If the heat generated by the processor exceeds the maximum coolant flow rate (for example, if the heat generated is 100 joules and the heat dissipation is 80 joules), the processor may reach a high temperature and cause equipment damage.

[0059] At the same time, the control mechanism of the liquid cooling system remains unchanged, and the actual output of the coolant in the liquid cooling system is the same. However, due to the different operating conditions of different processors, the coolant output required (expected output) for multiple processors is different.

[0060] In operation S240 , tasks are scheduled for the plurality of processors based on the cooling error between the expected output and the actual output, the predicted temperature, and the predicted operating parameters.

[0061] According to an embodiment of the present invention, the cooling error between the expected output and the actual output may be a cooling temperature difference or an error between the expected output and the actual output.

[0062] For example, the cooling temperature difference may be determined based on the difference between the two cooled temperatures obtained by cooling the processor using a desired output amount and an actual output amount of the coolant.

[0063] According to an embodiment of the present invention, by determining the predicted temperature and predicted operating parameters of the processors at a future time based on the current temperature and current operating parameters of multiple processors in a processor cluster, the processors can be heat-dissipated predictively based on the predicted temperature and predicted operating parameters. The expected output of the coolant in the liquid cooling system is determined based on the target system parameters of the liquid cooling system, the target material parameters of the phase change material, the predicted temperature, and the predicted operating parameters, so that the expected output integrates the coolant output required for the operation of multiple processors. Based on the control mechanism of the liquid cooling system, the actual output of the coolant at a future time is obtained based on the expected output, the predicted temperature, and the current temperature. Task scheduling is performed on multiple processors in the processor cluster based on the cooling error between the expected output and the actual output, the predicted temperature, and the predicted operating parameters. The control of the liquid cooling system is combined with the task scheduling of the processors to dissipate heat for the processor cluster, thereby improving the heat dissipation efficiency of the processor cluster.

[0064] According to an embodiment of the present invention, the expected output of the coolant in the liquid cooling system is determined based on the target system parameters of the liquid cooling system, the target material parameters of the phase change material, the predicted temperature and the predicted operating parameters, including: determining the power consumption driven heat dissipation of the processor based on the target system parameters and the predicted operating parameters; determining the temperature compensation amount of the processor based on the target material parameters, the predicted temperature and the temperature safety threshold conditions; and determining the expected output based on the power consumption driven heat dissipation and the temperature compensation amount.

[0065] According to an embodiment of the present invention, the expected output quantity Q in the processor heat dissipation control strategy is cool The formula is as follows:

[0066] (2)

[0067] k1 represents the target system parameter, k2 represents the target material parameter, P represents the predicted operating parameter, T target Represents the predicted temperature, T safe Indicates the temperature safety threshold, and the power consumption drives the heat dissipation to be , the temperature compensation is The exponent reflects the nonlinear nature of heat accumulation in liquid cooling systems. For every 1°C above the safe temperature, the heat dissipation demand increases exponentially.

[0068] For example, the temperature safety threshold condition may be that the predicted temperature is greater than a temperature safety threshold, which may be 85% of the processor junction temperature, such as 85 degrees Celsius.

[0069] For example, the temperature safety threshold condition may be that the temperature difference between a plurality of adjacent processors is greater than a temperature difference threshold.

[0070] Target system and material parameters can be calculated using a thermal model. For example, if a processor cluster requires continuous heat dissipation, k1 can be set to 0.8 and k2 to 0.3; to enhance transient processor temperature control, k1 can be set to 1.2 and k2 to 0.5; and for energy-saving heat dissipation mode, k1 can be set to 0.3 and k2 to 0.1.

[0071] According to embodiments of the present invention, due to the varying operating conditions of different processors, the power-driven heat dissipation of the processors can be determined based on target system parameters and predicted operating parameters. Furthermore, as the heat dissipation properties of the phase-change material change with processor temperature, the temperature compensation of the phase-change material can also change. Therefore, the temperature compensation amount for the processors is determined based on the target material parameters, the predicted temperature, and the temperature safety threshold. Consequently, the expected output is determined based on the power-driven heat dissipation and the temperature compensation amount, thereby integrating the coolant output required for the operation of multiple processors.

[0072] According to an embodiment of the present invention, task scheduling is performed on multiple processors based on a cooling error between an expected output and an actual output, a predicted temperature, and predicted operating parameters, including: determining, from multiple processors, a first processor whose predicted temperature satisfies a temperature safety threshold condition; determining, from at least one task to be processed by the first processor, a task to be migrated, based on the predicted temperature of the first processor; determining, from multiple processors, a second processor that meets the migration condition, based on the predicted temperatures and predicted operating parameters of the multiple processors; and migrating the task to be migrated to the second processor for processing.

[0073] For example, based on the cooling error, a processor that cannot be cooled in time may be determined from the plurality of processors, and it is determined whether the predicted temperature of the processor that cannot be cooled in time meets a temperature safety threshold condition.

[0074] For example, a heat dissipation difference may be determined based on the predicted temperature and actual output of the first processor; and a task to be migrated is determined from at least one task to be processed by the first processor based on the heat dissipation difference. The at least one task has a corresponding amount of heat generated during operation.

[0075] For example, based on the predicted temperatures of the multiple processors, a processor with a larger temperature difference from the first processor can be determined, and based on the predicted operating parameters of the multiple processors, a relatively idle processor can be determined. A second processor that meets the migration condition is determined from the processor with a larger temperature difference from the first processor and the relatively idle processor.

[0076] According to an embodiment of the present invention, based on the cooling error, the first processor whose predicted temperature meets the temperature safety threshold condition is determined from multiple processors, including: when there is a processor with a cooling error greater than the error threshold in the processor cluster, the first processor whose predicted temperature meets the temperature safety threshold condition is determined from the processor cluster.

[0077] According to an embodiment of the present invention, if the processor cooling error is greater than the error threshold, it indicates that the processor's operating temperature is difficult to dissipate in a timely manner. For example, if the cooling error is 10% and the error threshold is 2%, the actual heat dissipation of the processor is less than the expected heat dissipation.

[0078] According to an embodiment of the present invention, when the cooling error of the processor is less than the error threshold, it indicates that the operating temperature of the processor can be dissipated in time.

[0079] For example, a first processor is determined from the processor cluster, whose predicted temperature is greater than a temperature safety threshold; and / or a first processor is determined, whose temperature difference among a plurality of adjacent processors is greater than a temperature difference threshold.

[0080] A predicted temperature matrix can be constructed based on the predicted temperatures of each processor. A safety threshold matrix can be constructed based on the safety temperature thresholds of multiple processors. The predicted temperature matrix can be subtracted from the safety threshold matrix to determine the first processor from the multiple processors that exceeds the safety temperature threshold.

[0081] Each row or column in the predicted temperature matrix may represent a processor at a different location. By calculating the difference between adjacent rows or columns in the predicted temperature matrix, the first processor having a temperature difference greater than a temperature difference threshold among multiple adjacent processors may be determined.

[0082] For example, a predicted temperature field can be established based on the three-dimensional positions and temperatures of multiple processors. When the predicted temperature field is detected to be greater than a safe temperature threshold, the processor's pending tasks can be migrated.

[0083] According to an embodiment of the present invention, by determining whether there are processors in a processor cluster with a cooling error greater than an error threshold, it is possible to determine whether multiple processors will be able to receive timely heat dissipation in the future. If there are processors in the processor cluster with a cooling error greater than the error threshold, the first processor in the processor cluster whose predicted temperature meets the temperature safety threshold is determined. This processor can then be identified as the first processor to adjust its workload to reduce operating heat, thereby protecting the processors and reducing equipment losses.

[0084] According to an embodiment of the present invention, the second processor includes multiple processors, and the predicted operating parameters include at least one of the following: predicted power consumption and predicted computing load, and the migration conditions include temperature difference conditions and operating conditions; based on the predicted temperatures and predicted operating parameters of the multiple processors, the second processor that meets the migration conditions is determined from the multiple processors, including: based on the predicted temperatures of the multiple processors, determining the candidate processor that meets the temperature difference condition from the multiple processors; based on the position distance between the candidate processor and the first processor, the candidate's predicted operating parameters and the time slice length, determining the second processor that meets the operating conditions from the multiple candidate processors, and the time slice length is determined based on the number of processing cores of the processor and the time required to process at least one task.

[0085] According to an embodiment of the present invention, the predicted power consumption may be the electric power consumed by the processor when it is running at a future time. The predicted computing load may be the amount of tasks to be processed.

[0086] According to an embodiment of the present invention, the temperature difference condition may be that the predicted temperature difference between the device and the first device is greater than a preset temperature difference value.

[0087] The predicted temperature distribution differences of different processors are monitored in real time. When the predicted temperature exceeds the safe temperature threshold, the task to be migrated is dynamically migrated to the candidate processor with lower temperature, forming an intelligent scheduling similar to "thermal convection".

[0088] According to an embodiment of the present invention, the operating condition may be that an evaluation value determined based on the predicted operating parameters, location distance, and time slice length is greater than a preset evaluation threshold. The weights of the predicted computing load, location distance, and time slice length may be determined based on the task attributes of the task to be migrated.

[0089] For example, if the attribute of the task to be migrated is real-time business processing, the weight of the time slice length may be increased to determine a second processor that can quickly process and complete the task to be migrated from multiple candidate processors.

[0090] For example, the attribute of the task to be migrated may be batch business, and the weight of the location distance may be increased. The second processor with the smallest transmission distance may be determined from multiple candidate processors, thereby reducing the data transmission time.

[0091] For example, if the task attribute of the task to be migrated has a large task volume, the weight of the candidate predicted operating parameters (such as predicted operating load and predicted power consumption) is reduced to avoid migrating the task to be migrated to a processor with a higher load or power consumption.

[0092] For example, the number of cores that process the same task to be migrated is less than the preset number of cores in task scheduling. This can reduce the number of cross-core processing times and avoid complicating the processing of the task to be migrated.

[0093] According to an embodiment of the present invention, the time slice length t slice The formula is as follows:

[0094] (3)

[0095] Time slice length t slice Take 0.1ms (milliseconds) and Ttask represents the task processing time; Ncore represents the number of processing cores.

[0096] According to an embodiment of the present invention, based on the predicted temperatures of multiple processors, a candidate processor that meets the temperature difference condition is identified from the multiple processors, thereby dynamically migrating the task to be migrated to the candidate processor with a lower temperature, thus achieving intelligent scheduling similar to "thermal convection." Based on the location distance between the candidate processor and the first processor, the candidate's predicted operating parameters, and the time slice length, a second processor that meets the operating conditions is identified from the multiple candidate processors, thereby ensuring that the task to be migrated can be processed normally on the second processor, thereby improving business processing efficiency.

[0097] According to an embodiment of the present invention, the task to be migrated can be decomposed into discrete subtasks (such as 100 millisecond granularity), and time slice cutting technology can be used to achieve microsecond-level task migration scheduling, ensuring the atomic allocation and rapid reorganization capabilities of computing resources.

[0098] Based on the White Rabbit protocol, tasks to be migrated are migrated to the second processor through physical layer timestamp marking and fiber delay compensation, achieving a global clock synchronization accuracy of ±0.5 nanoseconds.

[0099] Modulation technology can be used during task migration to optimize data processing and communication. Modulation technology suppresses clock jitter to <0.1 picoseconds, ensuring time consistency during task migration across processors.

[0100] According to an embodiment of the present invention, based on the predicted temperature of the first processor, a task to be migrated is determined from at least one task processed by the first processor, including: based on the predicted temperature of the first processor, determining the amount of tasks completed by the first processor under temperature safety threshold conditions; based on the actual output amount and task amount, determining the task to be migrated from at least one task processed by the first processor.

[0101] According to an embodiment of the present invention, the temperature safety threshold condition may be that the predicted temperature is greater than the temperature safety threshold, and the amount of tasks completed by the first processor is determined to be less than or equal to the temperature safety threshold; and / or the temperature safety threshold condition may be that the temperature difference between the first processor and the adjacent processor is greater than the temperature difference threshold, and the amount of tasks completed by the first processor is determined to be less than or equal to the temperature difference threshold.

[0102] According to an embodiment of the present invention, a reference task amount of the first processor corresponding to the actual output amount is determined, and when the task amount is less than the reference task amount, a task to be migrated is determined from at least one task processed by the first processor.

[0103] By comparing the task amount with the reference task amount, it can be accurately determined whether the first processor can obtain effective heat dissipation, and the task to be migrated can also be accurately locked.

[0104] According to an embodiment of the present invention, since the operating conditions of different processors are different, the amount of tasks completed by the first processor under the temperature safety threshold condition is determined based on the predicted temperature of the first processor; based on the actual output amount and task amount, the task to be migrated is determined from at least one task processed by the first processor, and the amount of tasks to be migrated can be accurately determined, so that the processors in the processor cluster can all be accurately cooled, resulting in high heat dissipation efficiency.

[0105] According to an embodiment of the present invention, the target system parameters and the target material parameters are determined based on the following steps: according to the simulation temperature and the simulation operation parameters, a combination data set of multiple initial system parameters and initial material parameters is determined, and the simulation operation parameters and the simulation temperature are obtained by performing simulation using the simulation model of the processor; the multiple combination data sets are input into the thermal model respectively to obtain multiple thermal model function values; the multiple thermal model function values are compared to obtain a target combination set that meets the preset function conditions, and the target combination set includes the target system parameters and target material parameters.

[0106] According to an embodiment of the present invention, a simulation model may include a simulated liquid cooling system and a simulated processor. The simulated liquid cooling system and the simulated processor are obtained by three-dimensionally modeling the liquid cooling system and the processor, respectively. The simulated liquid cooling system is configured with a control mechanism for the liquid cooling system, and the simulated processor is configured with the same tasks and operating states as the processor. The simulation also sets parameters for the phase change material to simulate the liquid cooling system dissipating heat from the processor under different temperatures and operating parameters.

[0107] For example, in order to make the simulation more accurate, the simulation model can be divided into grids with a grid size of 0.1 mm. That is, the spatial discretization accuracy of the simulation model can be 0.1 mm grid.

[0108] According to an embodiment of the present invention, both the simulation temperature and the simulation operation parameters can be represented by a variety of initial system parameters and initial material parameters. For example, the simulation temperature can be T'(k 1, k2), the simulation operation parameter can be P'(k 1, k2). Therefore, based on a large number of simulation temperatures and simulation operation parameters, a combination data set of multiple initial system parameters and initial material parameters can be determined.

[0109] For example, the formula of thermal model F may be as follows:

[0110] F=α·T'(k 1, k2) + β·P'(k 1, k2) + γ QoS' (k 1, k2)(4)

[0111] According to an embodiment of the present invention, QoS'(k 1, k2) represents the simulation service quality of the simulation processor (such as simulation latency, simulation throughput, etc.). α, β, and γ are the weights of simulation temperature, simulation operating parameters, and simulation service quality, respectively. For example, α = 0.6, β = 0.3, and γ = 0.1.

[0112] According to an embodiment of the present invention, the thermal model may be adjusted according to actual conditions. For example, the thermal model may be improved by taking into account the heat generated by eddy current loss.

[0113] According to an embodiment of the present invention, the preset function condition may be: a minimum value of a plurality of thermal model function values.

[0114] According to an embodiment of the present invention, simulation operating parameters and simulated temperatures are obtained by performing simulation using a simulation model of a processor. Based on the simulated temperatures and simulated operating parameters, a combined dataset of multiple initial system parameters and initial material parameters can be determined, thereby reducing the time required to conduct experiments using a liquid cooling system and a processor cluster and improving the efficiency of obtaining combined datasets. The multiple combined datasets are input into the thermal model to obtain multiple thermal model function values; these multiple thermal model function values are compared to obtain target system parameters and target material parameters that meet preset function conditions. Therefore, the simulated operating parameters and simulated temperatures obtained by simulation using the simulation model can be used to obtain target system parameters and target material parameters that meet preset function conditions, facilitating subsequent processor task scheduling and improving task scheduling efficiency.

[0115] According to an embodiment of the present invention, based on the current temperatures and current operating parameters of multiple processors in a processor cluster, the predicted temperatures and predicted operating parameters of the processors at a future moment are determined, including: inputting the current temperatures and current operating parameters into a prediction model to obtain the predicted temperatures and predicted operating parameters, where the prediction model is trained based on simulation operating parameters and simulation temperatures.

[0116] According to an embodiment of the present invention, the training data of the prediction model can be 100,000 sets of simulated temperatures and simulated operating parameters. The types of simulated operating parameters are the same as the types of current operating parameters. For example, the current operating parameters can be current power consumption and current operating load. The simulated operating parameters can be simulated power consumption and simulated operating load.

[0117] According to an embodiment of the present invention, the prediction model may be a machine learning model, for example, a machine learning model having a 3-layer long short-term memory neural network and 128 hidden units.

[0118] Prediction constraints can be set, for example, predicting the temperature to be less than 85 degrees Celsius, predicting the energy consumption to be less than 1.2, and maintaining a QoS (Quality of Service) compliance rate greater than or equal to 99.9%.

[0119] For example, the processor cluster is deployed in a cabinet, and there are 640 processors in the processor cluster. The number of processing cores in the processor cluster may be 61,440.

[0120] The liquid cooling system uses PID to control the speed of the main circulation pump (such as a magnetic levitation centrifugal pump) to output coolant. The coolant flow rate can be 3000 liters per minute and the head is 60 meters.

[0121] The coolant may be composed of engineered deionized water having a thermal conductivity of 0.6 Watt-Kelvin per meter and a 2% concentration of nano-alumina particles.

[0122] To schedule tasks, the processor cluster connects to a circuit board or server via a dedicated compute fabric interface (CFI) via a physical data link. The circuit board or server executes the processor's task scheduling method. For example, the server could be Kubernetes, shortening the scheduling cycle to as little as 50 microseconds.

[0123] The dedicated interface for computing power transmission can be an optoelectronic fusion interface. The electrical signal channel is PCIe 6.0 × 16 (64 gigabits per second); the optical signal channel is a silicon photonic engine (wavelength 1310 nanometers, bandwidth 800 gigabits per second).

[0124] An optical switching matrix connects the computing power transmission interface to the ports on the circuit board. The optical switching matrix can be a 32x32-port silicon photonic switch (single wavelength 1.6 terabits per second). The electrical transmission channel can be a PCIe 6.0x64 aggregate link.

[0125] Based on the processor's thermal control strategy, the expected coolant output of the liquid cooling system is determined based on the target system parameters of the liquid cooling system, the target material parameters of the phase change material, the predicted temperature, and the predicted operating parameters. The processor thermal control strategy can be Model Predictive Control (MPC), with a prediction horizon of 5 seconds and a control horizon of 2 seconds.

[0126] The NoF (NVMe over Fabric) protocol can be used to schedule tasks across multiple processors. This involves initializing the communication protocol, starting the NoF protocol stack, loading kernel modules, configuring cache coherence domains, and establishing a virtual NUMA (Non-Uniform Memory Access) topology.

[0127] Figure 3 A NoF protocol layered interaction flow chart according to an embodiment of the present invention is shown.

[0128] like Figure 3 As shown, the NoF protocol is divided into a physical layer 310 , a topology layer 320 , a cache layer 330 , and an application layer 340 .

[0129] The task classifier in the application layer 340 of the NoF protocol is used to classify the tasks to be processed received by the processor cluster.

[0130] For example, the computing type of the task is identified, which can be FP32, INT8, control instructions, etc. The task source data is packaged and sent to the cache layer 330.

[0131] The improved MESI (Modified-Exclusive-Shared-Invalid) protocol state machine in the cache layer 330 is used to determine whether the task being processed by the processor is completed. If the task processing is completed, the task to be processed is transmitted to the topology layer 320.

[0132] The topology layer 320 includes a virtual NUMA domain builder, which can be a tool or component used to emulate a physical non-uniform memory access (NUMA) architecture in a virtualized environment. This builder divides multiple processors in a processor cluster into multiple virtual NUMA nodes, enabling virtual machines to perceive and utilize the NUMA topology, thereby optimizing resource allocation and performance.

[0133] The virtual NUMA domain builder is used to obtain the position distance between the candidate processor and the first processor, so as to determine the processor corresponding to the task to be processed according to the task scheduling result obtained by the task scheduling method of the server execution processor.

[0134] The physical layer 310 is used to transmit the tasks to be processed to the corresponding processor.

[0135] Physical layer 310 includes a hybrid signal encoder and a forward error correction module. The forward error correction coding scheme may be Reed-Solomon (255, 239). For example, the hybrid signal encoder may be a hybrid transmission scheme that uses PAM4 (Pulse Amplitude Modulation with 4 Levels) electrical signals combined with NRZ (Non-Return-to-Zero Optical Signal) optical signals.

[0136] Topology layer 320 includes a fault domain isolation module with a heartbeat detection period of 10 microseconds. For example, if the electronic switch actuation time is less than 10 microseconds, the power supply and data link of the faulty processor cluster can be quickly disconnected, preventing the fault from spreading to other modules while protecting the backup link from contamination. Topology reconstruction latency is less than 100 microseconds.

[0137] The cache layer 330 has a cache line locking mechanism across processor clusters, and the timeout detection window of the cache line locking mechanism is 100 nanoseconds.

[0138] The application layer 340 has a QoS marking engine that prioritizes pending tasks. For example, pending tasks can be divided into three levels of priority: real-time tasks with a latency of <5 milliseconds, high-throughput tasks with a latency of <50 milliseconds, and tasks with background processing attributes.

[0139] Each layer interacts through a standard interface. A hot tag is a metadata identifier that carries task characteristics and is used to convey key information related to task scheduling and resource allocation. In interactions between the application layer 340 and the cache layer 330, cross-layer information transfer is achieved by adding an 8-byte fixed-length field to the task metadata header. This 8-byte fixed-length field can include task priority, resource requirements, and a timestamp.

[0140] The core of the cross-layer interaction mechanism between the topology layer 320 and the physical layer 310 is to feedback the physical link quality status through the bit error rate indicator monitored in real time by the physical layer, thereby ensuring high reliability of data transmission for pending tasks.

[0141] If a processor cluster loses heartbeats more than three times, the fault-tolerance process is triggered. The fault domain isolation module (with an electronic switch operating time of less than 10 microseconds) quickly disconnects the power and data link to the faulty processor cluster. After isolating the faulty module, pending tasks are dynamically migrated to the backup processor cluster, reestablishing data paths and restoring system functionality. Topology reconstruction latency is less than 100 milliseconds.

[0142] Through end-to-end cyclic redundancy check (CRC), using a specified polynomial, pending tasks are written simultaneously to the primary storage medium and the mirror storage medium, ensuring data consistency in both copies. The total latency of the dual write operation does not exceed 15 nanoseconds.

[0143] According to an embodiment of the present invention, the above method further includes, when a failure of the liquid cooling system is detected: starting a spray device; and / or controlling multiple processors to reduce operating frequencies.

[0144] According to an embodiment of the present invention, the cooling setting for the processor cluster may be: when the liquid cooling system operates normally, the coolant in the liquid cooling system is used for heat dissipation; when the liquid cooling system fails, the spray device is started.

[0145] For example, a failure in the liquid cooling system may be caused by a blockage in the coolant pipe in the liquid cooling system, a control unit failure, insufficient coolant, etc.

[0146] For example, if the current temperature is detected to be greater than 85 degrees Celsius, the sprinkler can be automatically activated. The over-temperature protection response time can be set to less than 10 milliseconds.

[0147] According to an embodiment of the present invention, if a liquid cooling system fails and cannot dissipate heat for multiple processors in a processor cluster in a timely manner, a spray device may be activated to allow the processors to still dissipate heat.

[0148] For example, the nozzle aperture of the spray device may be 50 microns, and the array density may be 200 per square centimeter. The coolant of the spray device may be perfluorohexanone, which has a boiling point of 49 degrees Celsius and a latent heat of vaporization of 140 kilojoules per kilogram.

[0149] According to an embodiment of the present invention, due to a failure of the liquid cooling system, multiple processors can be controlled to reduce their operating frequencies, thereby avoiding loss of the processors.

[0150] According to an embodiment of the present invention, the method further includes: determining that a failure occurs in the liquid cooling system when a temperature change rate of the current temperature satisfies a preset temperature failure condition.

[0151] According to an embodiment of the present invention, the predicted temperature fault condition may be that the temperature change rate is greater than 15 degrees.

[0152] For example, when it is detected that the temperature change rate of the current temperature is greater than 15 degrees, it is determined that the liquid cooling system has failed.

[0153] For example, the liquid cooling system may include a microchannel cold plate with coolant pipes. The pipes may have a cross-sectional dimension of 100 microns wide by 300 microns deep. The pipes may have a surface roughness of 0.8 microns or less. The pipes may be made of copper-tungsten alloy, which has a thermal conductivity of 320 joules per second per square meter when the material is 1 meter thick and the temperature difference across the material is 1 Kelvin.

[0154] The coolant circulation path can be an asymmetric flow channel, the inlet diameter of the coolant pipe can be 8 millimeters, and the outlet diameter can be 12 millimeters. The coolant pipe has an anti-corrosion coating, for example, the anti-corrosion coating can be a nickel-phosphorus alloy layer with a thickness of 50 microns.

[0155] According to an embodiment of the present invention, multiple processors are respectively encapsulated in detachable components with a phase change material coating, and mounting grooves are formed on the surface of the detachable components for fixing the detachable components relative to the liquid cooling system.

[0156] According to an embodiment of the present invention, a mounting groove is formed on the surface of the removable component, and the insertion and removal force can be less than or equal to 15 Newtons, thereby preventing incorrect insertion. For example, the mounting groove is used to mount the removable component on a coolant pipe. The shape of the mounting groove matches the shape of the coolant pipe.

[0157] For example, the size of the detachable component can be 75 mm × 150 mm, the surface can be a liquid metal filling layer with a thickness of only 0.2 mm, and a temperature sensor can be embedded in the back, and the accuracy of the temperature sensor can be ±0.5 degrees Celsius.

[0158] For example, various processor mounting slots are available. The processor is equipped with a 12V DC input bidirectional Buck-Boost (step-down and step-up) regulator. Dynamic voltage adjustment range is 0.6V to 1.8V, with a step accuracy of 10mV. The processor also features a signal regenerator that supports 0-15dB insertion loss compensation to ensure high-speed signal integrity.

[0159] According to an embodiment of the present invention, a phase change material coating layer may be provided on the contact surface of the processor located on the detachable component.

[0160] For example, the phase change material can be gallium-based liquid metal (phase change point 29.8 degrees Celsius), and the activation threshold of the phase change material can be that the temperature difference between the current temperature and the phase change point is greater than or equal to 15 degrees Celsius, at which point the gallium-based liquid metal begins the phase change process and realizes heat absorption or release.

[0161] The phase change material coating can be injected with 300L of gallium-based alloy with a purity of 99.999%. When the temperature difference between the current temperature and the phase change point is equal to 20 degrees Celsius, the heat storage density is 800 kilojoules per cubic meter.

[0162] Gallium-based liquid metal can be passed through the microchannel flow channel, and the pressure difference in the microchannel flow channel is maintained in the range of 0.2-0.5 MPa.

[0163] According to an embodiment of the present invention, the processor cluster may be mechanically connected to the liquid cooling system via a liquid cooling quick connector. In a sealed state, the detected internal pressure is greater than 3 bar.

[0164] According to an embodiment of the present invention, the processor cluster uses MPO-24 multi-core connectors to achieve high-speed data transmission, with a loss of less than 0.2dB per connection point. dB represents the degree of power attenuation during signal transmission. 0.2dB indicates low signal loss per connection point and high transmission efficiency.

[0165] According to an embodiment of the present invention, a liquid cooling plate in a liquid cooling system may have a quick-release thermal conductive joint, and the contact thermal resistance of the quick-release thermal conductive joint is less than 0.05 degrees Celsius square centimeter per watt.

[0166] According to an embodiment of the present invention, the detachable component includes an upper shell and a lower shell, the gold finger bumps of the upper shell are adapted to the gold finger grooves of the lower shell, and the processor is electrically connected to the circuit board via the gold finger contacts in the gold finger grooves.

[0167] According to an embodiment of the present invention, a circuit board can be used to execute a task scheduling method for a processor. The circuit board includes an optoelectronic hybrid signal repeater with a +6dB electrical signal gain and 3dB optical signal pre-emphasis. The circuit board also includes a priority arbiter, which can be an FPGA-based 128-level queue manager. This prioritizes requests from 128 input channels, enabling queue management, conflict resolution, and orderly scheduling.

[0168] Figure 4A A schematic diagram of a gold finger groove of a lower shell according to an embodiment of the present invention is shown.

[0169] like Figure 4AAs shown, gold finger contacts can be located within gold finger grooves 411 on lower housing 410 to transmit processor operating parameters and temperature to a circuit board or server. Lower housing 410 also has a foolproof mark 412 located about one-fifth of the way along each side, rotating clockwise, to prevent incorrect installation of detachable components.

[0170] Figure 4B A schematic diagram of a gold finger bump of an upper housing according to an embodiment of the present invention is shown.

[0171] like Figure 4B As shown, the gold finger protrusion 421 is protruding relative to other parts of the upper shell 420 so as to be matched with the gold finger groove 411 .

[0172] It should be noted that the upper shell 420 may also be in other shapes, for example, Figure 4A The lower shell 410 is a rectangular shell with the same shape and size.

[0173] Figure 5 A schematic diagram showing a detachable component according to an embodiment of the present invention is shown.

[0174] like Figure 5 As shown, the detachable component 400 has a mounting groove 430, which is used to mount the detachable component on the coolant pipe. The shape of the mounting groove matches the shape of the coolant pipe.

[0175] Figure 6 A schematic diagram of a liquid cooling quick connector according to an embodiment of the present invention is shown.

[0176] like Figure 6 As shown, multiple coolant pipes can be connected through liquid cooling quick connectors 610. Three sets of Hall sensors 620 are also provided on the coolant pipes to ensure that the sealing pressure is greater than 2.5 MPa.

[0177] Figure 7 A structural block diagram of a task scheduling device for a processor according to an embodiment of the present invention is shown.

[0178] like Figure 7 As shown, the task scheduling device 700 of the processor of this embodiment includes a first determination module 710 , a second determination module 720 , an acquisition module 730 and a task scheduling module 740 .

[0179] The first determination module 710 is used to determine the predicted temperature and predicted operating parameters of the processors at a future time based on the current temperatures and current operating parameters of the multiple processors in the processor cluster. In one embodiment, the first determination module 710 can be used to perform the operation S210 described above, which will not be repeated here.

[0180] The second determination module 720 is configured to determine a desired output of coolant in the liquid cooling system based on the target system parameters of the liquid cooling system, the target material parameters of the phase change material, the predicted temperature, and the predicted operating parameters. The liquid cooling system is configured to dissipate heat from the processor cluster. In one embodiment, the second determination module 720 may be configured to execute operation S220 described above, which will not be further described herein.

[0181] The acquisition module 730 is used to obtain the actual output of the coolant at a future time based on the desired output, the predicted temperature, and the current temperature based on the control mechanism of the liquid cooling system. In one embodiment, the acquisition module 730 can be used to perform the operation S230 described above, which will not be repeated here.

[0182] The task scheduling module 740 is used to schedule tasks for multiple processors based on the cooling error between the expected output and the actual output, the predicted temperature, and the predicted operating parameters. In one embodiment, the task scheduling module 740 can be used to perform the operation S240 described above, which will not be repeated here.

[0183] According to an embodiment of the present invention, the task scheduling module 740 includes a first determination submodule, a second determination submodule, a third determination submodule, and a migration submodule. The first determination submodule is configured to determine, from a plurality of processors, a first processor whose predicted temperature satisfies a temperature safety threshold condition based on a cooling error. The second determination submodule is configured to determine, from at least one task to be processed by the first processor, a task to be migrated based on the predicted temperature of the first processor. The third determination submodule is configured to determine, from a plurality of processors, a second processor that meets the migration condition based on the predicted temperatures and predicted operating parameters of the plurality of processors. The migration submodule is configured to migrate the task to be migrated to the second processor for processing.

[0184] According to an embodiment of the present invention, the first determining submodule includes a first determining unit configured to determine a first processor in the processor cluster whose predicted temperature meets a temperature safety threshold condition when a processor with a cooling error greater than an error threshold exists in the processor cluster.

[0185] According to an embodiment of the present invention, the second processor includes multiple processors, and the predicted operating parameters include at least one of the following: predicted power consumption and predicted computing load, and the migration conditions include temperature difference conditions and operating conditions. The third determination submodule includes a second determination unit and a third determination unit. The second determination unit is used to determine a candidate processor that meets the temperature difference condition from multiple processors based on the predicted temperatures of the multiple processors. The third determination unit is used to determine a second processor that meets the operating conditions from multiple candidate processors based on the location distance between the candidate processor and the first processor, the candidate's predicted operating parameters and the time slice length. The time slice length is determined based on the number of processing cores of the processor and the time required to process at least one task.

[0186] According to an embodiment of the present invention, the second determining submodule includes a fourth determining unit and a fifth determining unit. The fourth determining unit is configured to determine, based on the predicted temperature of the first processor, an amount of tasks to be completed by the first processor under a temperature safety threshold condition, and the fifth determining unit is configured to determine, based on the actual output amount and the amount of tasks, a task to be migrated from at least one task processed by the first processor.

[0187] According to an embodiment of the present invention, the second determination module 720 includes a fourth determination submodule, a fifth determination submodule, and a sixth determination submodule. The fourth determination submodule is configured to determine the power consumption-driven heat dissipation of the processor based on target system parameters and predicted operating parameters. The fifth determination submodule is configured to determine the temperature compensation amount of the processor based on target material parameters, predicted temperature, and temperature safety threshold conditions. The sixth determination submodule is configured to determine the expected output amount based on the power consumption-driven heat dissipation amount and the temperature compensation amount.

[0188] According to an embodiment of the present invention, the target system parameters and the target material parameters are determined based on the following steps: according to the simulation temperature and the simulation operation parameters, a combination data set of multiple initial system parameters and initial material parameters is determined, and the simulation operation parameters and the simulation temperature are obtained by performing simulation using the simulation model of the processor; the multiple combination data sets are input into the thermal model respectively to obtain multiple thermal model function values; the multiple thermal model function values are compared to obtain a target combination set that meets the preset function conditions, and the target combination set includes the target system parameters and target material parameters.

[0189] According to an embodiment of the present invention, the first determination module 710 includes an acquisition submodule configured to input the current temperature and the current operating parameters into a prediction model to obtain the predicted temperature and the predicted operating parameters, wherein the prediction model is trained based on the simulated operating parameters and the simulated temperature.

[0190] According to an embodiment of the present invention, the device further includes a startup module and a control module. The startup module is configured to start the spray device upon detecting a failure in the liquid cooling system; and / or the control module is configured to control the multiple processors to reduce their operating frequencies.

[0191] According to an embodiment of the present invention, the apparatus further includes a third determining module configured to determine that a fault occurs in the liquid cooling system when a temperature change rate of a current temperature satisfies a preset temperature fault condition.

[0192] According to an embodiment of the present invention, multiple processors are respectively encapsulated in detachable components with a phase change material coating, and mounting grooves are formed on the surface of the detachable components for fixing the detachable components relative to the liquid cooling system.

[0193] According to an embodiment of the present invention, the detachable component includes an upper shell and a lower shell, the gold finger bumps of the upper shell are adapted to the gold finger grooves of the lower shell, and the processor is electrically connected to the circuit board via the gold finger contacts in the gold finger grooves.

[0194] According to an embodiment of the present invention, any multiple modules among the first determination module 710, the second determination module 720, the acquisition module 730, and the task scheduling module 740 can be combined into a single module, or any one of these modules can be split into multiple modules. Alternatively, at least part of the functionality of one or more of these modules can be combined with at least part of the functionality of other modules and implemented in a single module. According to an embodiment of the present invention, at least one of the first determination module 710, the second determination module 720, the acquisition module 730, and the task scheduling module 740 can be at least partially implemented as a hardware circuit, such as a field programmable gate array (FPGA), a programmable logic array (PLA), a system on a chip, a system on a substrate, a system on a package, an application-specific integrated circuit (ASIC), or can be implemented in hardware or firmware through any other reasonable means of circuit integration or packaging, or can be implemented in any one of the three implementation methods of software, hardware, and firmware, or any appropriate combination of these. Alternatively, the first determination module 710, the second determination module 720, the acquisition module 730, and the task scheduling module 740. At least one of the above may be at least partially implemented as a computer program module, which may perform the corresponding function when the computer program module is executed.

[0195] Figure 8 A block diagram of an electronic device suitable for implementing a task scheduling method for a processor according to an embodiment of the present invention is shown.

[0196] like Figure 8 As shown, an electronic device 800 according to an embodiment of the present invention includes a processor 801, which can perform various appropriate actions and processes based on programs stored in a read-only memory (ROM) 802 or programs loaded from a storage unit 808 into a random access memory (RAM) 803. The processor 801 may include, for example, a general-purpose microprocessor (e.g., a CPU), an instruction set processor and / or related chipsets and / or a special-purpose microprocessor (e.g., an application-specific integrated circuit (ASIC)), etc. The processor 801 may also include onboard memory for caching purposes. The processor 801 may include a single processing unit or multiple processing units for performing different actions of the method flow according to an embodiment of the present invention.

[0197] Various programs and data required for the operation of the electronic device 800 are stored in the RAM 803. The processor 801, ROM 802, and RAM 803 are connected to each other via a bus 804. The processor 801 executes the programs in the ROM 802 and / or RAM 803 to perform various operations according to the method flow of the embodiment of the present invention. It should be noted that the programs may also be stored in one or more memories other than the ROM 802 and RAM 803. The processor 801 may also execute the programs stored in the one or more memories to perform various operations according to the method flow of the embodiment of the present invention.

[0198] According to an embodiment of the present invention, electronic device 800 may further include an input / output (I / O) interface 805, which is also connected to bus 804. Electronic device 800 may also include one or more of the following components connected to I / O interface 805: an input section 806 including a keyboard, mouse, etc.; an output section 807 including devices such as a cathode ray tube (CRT), liquid crystal display (LCD), and speakers; a storage section 808 including a hard disk; and a communication section 809 including a network interface card such as a LAN card or modem. Communication section 809 performs communication processing via a network such as the Internet. A drive 810 is also connected to I / O interface 805 as needed. Removable media 811, such as a magnetic disk, optical disk, magneto-optical disk, semiconductor memory, etc., is installed in drive 810 as needed, so that computer programs read from the removable media can be installed into storage section 808 as needed.

[0199] The present invention also provides a computer-readable storage medium, which may be included in the device / apparatus / system described in the above embodiments, or may exist independently and not incorporated into the device / apparatus / system. The computer-readable storage medium carries one or more programs, which, when executed, implement the method according to the embodiments of the present invention.

[0200] According to an embodiment of the present invention, a computer-readable storage medium may be a non-volatile computer-readable storage medium, and may include, for example, but not limited to: a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof. In the present invention, a computer-readable storage medium may be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. For example, according to an embodiment of the present invention, a computer-readable storage medium may include the ROM 802 and / or RAM 803 described above, and / or one or more memories other than ROM 802 and RAM 803.

[0201] An embodiment of the present invention further includes a computer program product comprising a computer program containing program code for executing the method shown in the flowchart. When the computer program product is executed in a computer system, the program code is used to cause the computer system to implement the task scheduling method for a processor provided in an embodiment of the present invention.

[0202] The computer program executes the above functions defined in the system / device of the embodiment of the present invention when executed by the processor 801. According to the embodiment of the present invention, the system, device, module, unit, etc. described above can be implemented by a computer program module.

[0203] In one embodiment, the computer program may be stored on a tangible storage medium such as an optical storage device or a magnetic storage device. In another embodiment, the computer program may be transmitted and distributed in the form of a signal on a network medium, downloaded and installed via the communication portion 809, and / or installed from a removable medium 811. The program code contained in the computer program may be transmitted using any appropriate network medium, including but not limited to wireless, wired, or any suitable combination thereof.

[0204] In such an embodiment, the computer program can be downloaded and installed from a network via the communication section 809 and / or installed from a removable medium 811. When the computer program is executed by the processor 801, the above-described functions defined in the system of the embodiment of the present invention are performed. According to the embodiment of the present invention, the systems, devices, means, modules, units, etc. described above can be implemented by computer program modules.

[0205] According to an embodiment of the present invention, the program code for executing the computer program provided by the embodiment of the present invention can be written in any combination of one or more programming languages. Specifically, these computer programs can be implemented using high-level procedural and / or object-oriented programming languages, and / or assembly / machine languages. Programming languages include, but are not limited to, languages such as Java, C++, Python, "C" or similar programming languages. The program code can be executed entirely on the user computing device, partially on the user device, partially on a remote computing device, or entirely on a remote computing device or server. In the case of a remote computing device, the remote computing device can be connected to the user computing device through any type of network, including a local area network (LAN) or a wide area network (WAN), or can be connected to an external computing device (for example, using an Internet service provider to connect via the Internet).

[0206] The flowcharts and block diagrams in the accompanying drawings illustrate the possible implementation architecture, functions and operations of the systems, methods and computer program products according to various embodiments of the present invention. In this regard, each box in the flowchart or block diagram can represent a module, program segment, or a part of code, and the above-mentioned module, program segment, or a part of code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in an order different from that marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram or flowchart, and the combination of boxes in the block diagram or flowchart, can be implemented with a dedicated hardware-based system that performs the specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.

[0207] It will be understood by those skilled in the art that the features described in the various embodiments of the present invention may be combined and / or coupled in various ways, even if such combinations or couplings are not explicitly described in the present invention. In particular, the features described in the various embodiments of the present invention may be combined and / or coupled in various ways without departing from the spirit and teachings of the present invention. All such combinations and / or couplings fall within the scope of the present invention.

[0208] The above describes embodiments of the present invention. However, these embodiments are for illustrative purposes only and are not intended to limit the scope of the present invention. Although each embodiment has been described separately above, this does not mean that the measures in each embodiment cannot be advantageously used in combination. Without departing from the scope of the present invention, those skilled in the art may make various substitutions and modifications, which should all fall within the scope of the present invention.

Claims

1. A task scheduling method for a processor, characterized in that: The method comprises: Determining predicted temperatures and predicted operating parameters of the processors at a future time based on current temperatures and current operating parameters of the plurality of processors in the processor cluster; determining a desired output of coolant in the liquid cooling system according to target system parameters of the liquid cooling system, target material parameters of the phase change material, the predicted temperature, and the predicted operating parameters, wherein the liquid cooling system is used to dissipate heat from the processor cluster; Based on the control mechanism of the liquid cooling system, obtaining the actual output of the coolant at the future time according to the expected output, the predicted temperature, and the current temperature; performing task scheduling on the plurality of processors according to a cooling error between the expected output and the actual output, the predicted temperature, and the predicted operating parameters; The performing task scheduling on the plurality of processors according to the cooling error between the expected output and the actual output, the predicted temperature, and the predicted operating parameters includes: determining, based on the cooling error, a first processor from the plurality of processors, the first processor having the predicted temperature meeting a temperature safety threshold condition; determining, based on the predicted temperature of the first processor, a task to be migrated from at least one task to be processed by the first processor; determining, from the plurality of processors, a second processor that meets a migration condition based on the predicted temperatures and the predicted operating parameters of the plurality of processors; Migrate the task to be migrated to the second processor for processing.

2. The method according to claim 1, characterized in that The step of determining, based on the cooling error, a first processor from the plurality of processors whose predicted temperature satisfies a temperature safety threshold condition comprises: In a case where there is a processor in the processor cluster whose cooling error is greater than the error threshold, a first processor whose predicted temperature satisfies the temperature safety threshold condition is determined from the processor cluster.

3. The method according to claim 1, characterized in that The second processor includes a plurality of processors, the predicted operating parameters include at least one of the following: predicted power consumption and predicted computing load, and the migration conditions include temperature difference conditions and operating conditions; The determining, from the plurality of processors, a second processor that meets a migration condition based on the predicted temperatures and the predicted operating parameters of the plurality of processors includes: determining, based on the predicted temperatures of the plurality of processors, a candidate processor that satisfies a temperature difference condition from among the plurality of processors; Based on the location distance between the candidate processor and the first processor, the candidate's predicted operating parameters and the time slice length, the second processor that meets the operating conditions is determined from multiple candidate processors, and the time slice length is determined based on the number of processing cores of the processor and the time required to process the at least one task.

4. The method according to claim 1, characterized in that The determining, based on the predicted temperature of the first processor, the task to be migrated from at least one task processed by the first processor, includes: determining, based on the predicted temperature of the first processor, an amount of tasks to be completed by the first processor under the temperature safety threshold condition; The task to be migrated is determined from at least one task processed by the first processor according to the actual output amount and the task amount.

5. The method according to claim 1, characterized in that: The determining, based on target system parameters of the liquid cooling system, target material parameters of the phase change material, the predicted temperature, and the predicted operating parameters, of the expected output of the coolant in the liquid cooling system includes: determining a power consumption driven heat dissipation amount of a processor based on the target system parameters and the predicted operating parameters; determining a temperature compensation amount for a processor according to the target material parameters, the predicted temperature, and a temperature safety threshold condition; The expected output amount is determined according to the power consumption driving heat dissipation amount and the temperature compensation amount.

6. The method according to claim 1, characterized in that The target system parameters and the target material parameters are determined based on the following steps: determining a combined data set of multiple initial system parameters and initial material parameters based on a simulation temperature and a simulation operating parameter, wherein the simulation operating parameter and the simulation temperature are obtained by performing simulation using a simulation model of a processor; Inputting the plurality of combined data sets into a thermal model respectively to obtain a plurality of thermal model function values; A plurality of thermal model function values are compared to obtain a target combination set that meets a preset function condition, wherein the target combination set includes the target system parameters and the target material parameters.

7. The method according to claim 6, characterized in that The step of determining predicted temperatures and predicted operating parameters of the processors at a future time based on current temperatures and current operating parameters of the processors in the processor cluster includes: The current temperature and the current operating parameters are input into a prediction model to obtain the predicted temperature and the predicted operating parameters, wherein the prediction model is trained based on the simulation operating parameters and the simulation temperature.

8. The method according to claim 1, characterized in that: The method further includes, when a failure of the liquid cooling system is detected: Activate sprinklers; and / or The plurality of processors are controlled to reduce operating frequencies.

9. The method according to claim 8, characterized in that The method further comprises: When the temperature change rate of the current temperature meets a preset temperature failure condition, it is determined that a failure occurs in the liquid cooling system.

10. The method according to claim 1, characterized in that: The plurality of processors are respectively encapsulated in a detachable component having a phase change material coating. A mounting groove is formed on a surface of the detachable component, and the mounting groove is used to fix the detachable component relative to the liquid cooling system.

11. The method according to claim 10, characterized in that: The detachable component includes an upper shell and a lower shell. The gold finger bumps of the upper shell are adapted to the gold finger grooves of the lower shell. The processor is electrically connected to the circuit board via the gold finger contacts in the gold finger grooves.

12. An electronic device comprising: one or more processors; a memory for storing one or more computer programs, It is characterized in that the one or more processors execute the one or more computer programs to implement the steps of the method according to any one of claims 1 to 11.

13. A computer-readable storage medium having a computer program or instruction stored thereon, characterized in that: When the computer program or instruction is executed by a processor, the steps of the method according to any one of claims 1 to 11 are implemented.

14. A computer program product comprising a computer program or instructions, characterized in that When the computer program or instruction is executed by a processor, the steps of the method according to any one of claims 1 to 11 are implemented.

Citation Information

Patent Citations

  • Task scheduling method and device

    CN113254172A

  • Liquid cooling server intelligent temperature control method based on local software monitoring and liquid cooling server

    CN119536482A