Server temperature control method, system and device, and storage medium
Through the combination of parallel processing equipment and substrate management controller, the server temperature information is quickly and accurately obtained, especially the timely control of high-temperature sensitive components, solving the problem of time-consuming in traditional methods and improving the service life and reliability of the server.
Patent Information
- Application Number
- PCT/CN2024/122233
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-01-17
- Filing Date
- 2024-09-29
- Publication Date
- 2025-07-24
AI Technical Summary
Traditional server temperature control methods take a long time and cannot control high-temperature sensitive components in a timely and accurate manner, resulting in high fan noise and high energy loss, which affects server life and reliability.
The parallel processing equipment is used to connect to the substrate management controller, and the temperature information is sent in parallel by reading and sorting it, and the temperature information of the graphics processor is processed first, thereby shortening the temperature data acquisition time.
Improve the accuracy and efficiency of server temperature control and extend the service life and reliability of the server.
Smart Images

Figure CN2024122233_24072025_PF_FP_ABST
Abstract
Description
A temperature control method, system, device and storage medium for a server
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS
[0002] This application claims priority to a Chinese patent application filed with the Patent Office of China on January 17, 2024, with application number 202410066855.8, entitled “A temperature control method, system, device and storage medium for a server”, the entire contents of which are incorporated herein by reference. Technical Field
[0003] The present application relates to the field of server control technology, and in particular to a temperature control method, system, device, and storage medium for a server. Background Art
[0004] Generally speaking, when a server is cooling, the BMC (Board Management Controller) usually uses I2C (Inter-Integrated Circuit, a two-wire serial bus) to poll various components out of band to obtain parameter values of peripheral components.
[0005] Refer to Figure 1, which is a schematic diagram of a current server design. The BMC in Figure 1 can access a total of 11 PCIe (Peripheral Component Interconnect express, a high-speed serial computer expansion bus standard) cards through two two-wire serial bus switch chips (I2C switches), labeled PCIe card 0 to PCIe card 10 in Figure 1. If reading the temperature information of each PCIe card takes 100 to 300 milliseconds (different components may have different time consumption), then the BMC will take 1.1 to 3.3 seconds to read the temperature information of all PCIe cards on the backplane in Figure 1, which is a long time. In actual applications, the density of servers is increasing, and the number of components contained therein is increasing. The BMC may have several blocks to poll. After polling all of them, the overall time consumption may take 3 to 5 seconds or even longer. If some components are highly temperature-sensitive, timely and accurate temperature control may not be possible, causing the temperature of these components to rise rapidly. In turn, the fan power needs to be adjusted to a larger level to cool them down. This can easily lead to loud fan noise, high energy loss, frequent fan fluctuations, and shortened fan life, ultimately affecting the service life and reliability of the server.
[0006] In summary, how to more accurately and effectively control the temperature of the server and ensure the service life and reliability of the server is a technical problem that technical personnel in this field urgently need to solve.
[0007] Summary of the Invention
[0008] The purpose of this application is to provide a server temperature control method, system, device and storage medium to more accurately and effectively control the temperature of the server and ensure the service life and reliability of the server.
[0009] To solve the above technical problems, this application provides the following technical solutions:
[0010] A temperature control method for a server is provided. A baseboard management controller is connected to a preset parallel processing device, and the parallel processing device is connected to N components. The temperature control method for the server is applied to the parallel processing device, and includes:
[0011] After the server is powered on, determining the component type of each of the N components one by one and sending the component type to the baseboard management controller;
[0012] In the first stage of each parameter reading cycle, the temperature information of each of the N components is read simultaneously through its own N threads in a parallel reading manner;
[0013] In the second phase of each parameter reading cycle, temperature information of each component whose component type is a graphics processor is sent to the baseboard management controller;
[0014] In the third phase of each parameter reading cycle, temperature information of each component other than the graphics processor is sent to the baseboard management controller, so that the baseboard management controller performs temperature control of the server based on the temperature information of each of the N components.
[0015] In some implementations, the parallel processing device includes a single first controller having at least N threads.
[0016] In some embodiments, the first controller has a first interface and a second interface for connecting to the baseboard management controller;
[0017] In the second phase of each parameter reading cycle, the temperature information of each component whose component type is a graphics processor is sent to the baseboard management controller, including:
[0018] In the second phase of each parameter reading cycle, the temperature information of each component whose component type is a graphics processor is sent to the baseboard management controller through the first interface of the first controller itself;
[0019] In the third phase of each parameter reading cycle, the temperature information of each component of a non-graphics processor type is sent to the baseboard management controller, including:
[0020] In the third phase of each parameter reading cycle, the temperature information of each component whose component type is not a graphics processor is sent to the baseboard management controller through the second interface of the first controller itself.
[0021] In some embodiments, among the N components, there are K types of components excluding the graphics processor, where K is a positive integer not less than 2, and the third phase of the parameter reading cycle is divided into K sub-phases, where i is a positive integer and 1≤i≤K;
[0022] In the third phase of each parameter reading cycle, the temperature information of each component of a non-GPU component type is sent to the baseboard management controller via the second interface of the first controller itself, including:
[0023] In the i-th sub-phase of the third phase of each parameter reading cycle, the temperature information of the i-th component of a non-graphics processor type is sent to the baseboard management controller through the second interface of the first controller itself.
[0024] In some embodiments, the parallel processing device includes M second controllers, the total number of threads of the M second controllers is greater than or equal to N, M is a positive integer not less than 2, any one of the second controllers is connected to at least one of the N components, and any one of the N components is connected to at most one second controller.
[0025] In some embodiments, the device models of the M second controllers are the same, each second controller has a threads, a is a positive integer and a×M≥N, and each second controller has a first interface and a second interface for connecting to the baseboard management controller.
[0026] In some embodiments, in the first phase of each parameter reading cycle, the temperature information of each of N components is read simultaneously in a parallel reading manner by its own N threads, including:
[0027] In the first stage of each parameter reading cycle, each of the second controllers reads the temperature information of each component connected to itself simultaneously through its own a threads in a parallel reading manner.
[0028] In some embodiments, the second stage is divided into M sub-stages, j is a positive integer and 1≤j≤M;
[0029] In the second phase of each parameter reading cycle, the temperature information of each component whose component type is a graphics processor is sent to the baseboard management controller, including:
[0030] In the jth sub-phase of the second phase of each parameter reading cycle, the jth second controller among the M second controllers sends the temperature information of each component of the graphics processor type connected to the jth second controller itself to the baseboard management controller through its own first interface.
[0031] In some embodiments, among the N components, there are K types of components excluding the graphics processor, where K is a positive integer not less than 2, and the third phase of the parameter reading cycle is divided into K rounds, each round is divided into M sub-phases, where i is a positive integer and 1≤i≤K, and j is a positive integer and 1≤j≤M;
[0032] In the third phase of each parameter reading cycle, the temperature information of each component of a non-graphics processor type is sent to the baseboard management controller, including:
[0033] During the j-th sub-period of the i-th round of the third phase of each parameter reading cycle, the j-th second controller among the M second controllers sends temperature information of the i-th component of a non-graphics processor type connected to the j-th second controller to the baseboard management controller via its second interface.
[0034] In some embodiments, the parallel processing device is a parallel processing device based on a microcontroller unit, or a parallel processing device based on a field programmable gate array, or a parallel processing device based on a complex programmable logic device.
[0035] In some embodiments, further comprising:
[0036] When a fault signal is received from any one component, the sending of the temperature information of the current stage is suspended and the fault signal is sent to the baseboard management controller, and the sending of the temperature information of the current stage is continued after the sending of the fault signal is completed.
[0037] A temperature control system for a server includes a baseboard management controller (BMC) and a preset parallel processing device connected to the BMC, wherein the parallel processing device is connected to N components and includes:
[0038] A power-on detection module, configured to determine the component type of each of the N components one by one and send the component type to the baseboard management controller after the server is powered on;
[0039] The first-stage execution module is used to read the temperature information of N components simultaneously in a parallel reading manner through its own N threads in the first stage of each parameter reading cycle;
[0040] a second-stage execution module, configured to send temperature information of each component whose component type is a graphics processor to the baseboard management controller in the second stage of each parameter reading cycle;
[0041] The third stage execution module is configured to send temperature information of each component of a non-graphics processor type to the baseboard management controller during the third stage of each parameter reading cycle, so that the baseboard management controller performs temperature control of the server based on the temperature information of each of the N components.
[0042] A temperature control device for a server comprises:
[0043] memory for storing computer programs;
[0044] The processor is configured to execute the computer program to implement the steps of the temperature control method for the server as described above.
[0045] A computer-readable storage medium stores a computer program, which, when executed by a processor, implements the steps of the server temperature control method as described above.
[0046] A temperature control method for a server is provided. A baseboard management controller is connected to a preset parallel processing device, and the parallel processing device is connected to N components. The temperature control method for the server is applied to the baseboard management controller, and includes:
[0047] Powering on the server and, after determining the component types of the N components one by one through the parallel processing device, receiving the component types of the N components sent by the parallel processing device;
[0048] In the second phase of each parameter reading cycle, receiving temperature information of each component of the graphics processor type sent by the parallel processing device;
[0049] In the third phase of each parameter reading cycle, receiving temperature information of each component of a non-graphics processor type sent by the parallel processing device;
[0050] performing temperature control of the server based on temperature information of each of the N components;
[0051] In the first stage of each parameter reading cycle, the parallel processing device reads the temperature information of each of the N components simultaneously in a parallel reading manner through its own N threads. BRIEF DESCRIPTION OF THE DRAWINGS
[0052] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0053] FIG1 is a schematic diagram of a current server design;
[0054] FIG2 is a schematic structural diagram of a temperature control system of a server in a specific embodiment of the present application;
[0055] FIG3 is a flowchart of an implementation of a temperature control method for a server in the present application applied to a parallel processing device;
[0056] FIG4 is a schematic structural diagram of a parallel processing device in a specific embodiment of the present application;
[0057] FIG5 is a schematic structural diagram of a parallel processing device in another specific embodiment of the present application;
[0058] FIG6 is a schematic diagram of the module structure of a parallel processing device in another specific embodiment of the present application;
[0059] FIG7 is a schematic structural diagram of a temperature control device for a server in a specific embodiment of the present application;
[0060] FIG8 is a schematic diagram of the structure of a computer-readable storage medium in this application;
[0061] FIG9 is a flowchart of an implementation of a temperature control method for a server in the present application applied to a baseboard management controller. DETAILED DESCRIPTION
[0062] The core of this application is to provide a temperature control method for a server, which can effectively shorten the time it takes for a baseboard management controller to obtain temperature information of each component, and can promptly obtain temperature information of a graphics processor with high temperature sensitivity, so that the solution of this application can more accurately and effectively realize temperature control of the server, thereby ensuring the service life and reliability of the server.
[0063] In order to enable those skilled in the art to better understand the present application, the present application is further described in detail below in conjunction with the accompanying drawings and specific embodiments. Obviously, the embodiments described are only a part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without making any creative efforts are within the scope of protection of the present application.
[0064] Please refer to Figure 2, which is a schematic diagram of the structure of a server temperature control system in a specific embodiment of the present application. The server temperature control system includes a baseboard management controller and a parallel processing device. In Figure 2, the baseboard management controller is connected to a preset parallel processing device, and the parallel processing device is connected to N components. Please refer to Figure 3, which is a flowchart of an implementation method of a server temperature control method in the present application. This server temperature control method can be applied to the parallel processing device and includes the following steps:
[0065] Step S301: After the server is powered on, the component types of N components are determined one by one and sent to a baseboard management controller.
[0066] Since in the solution of the present application, temperature information of different types of components is subsequently sent from the parallel processing device to the baseboard management controller in sequence according to the different component types, it is necessary to first determine the component types of the N components.
[0067] Step S301 only needs to be executed once after the server is powered on. That is, during each subsequent parameter reading cycle, the baseboard management controller will know the component types of each of the N components. Of course, in some cases, components may be updated. For example, if a new GPU (Graphics Processing Unit) is added to a slot on a server backplane, the new component insertion operation is usually performed after the server is powered off. Of course, if live operation is allowed in some cases, step S301 can be executed once after the component update.
[0068] The N components described in this application are typically PCIe cards installed on a backplane. Parallel processing devices can also be installed on the backplane. Of course, in some cases, the server may have multiple backplanes. In this case, the solution of this application can be applied to each backplane, allowing the baseboard management controller to quickly read the temperature information of each component based on the parallel processing devices installed on the backplane. For example, in some cases, the server has a front window hard disk backplane, a built-in SSD (Solid State Disk) backplane, a front window expansion board, a rear window PCIE card backplane, and so on.
[0069] After the server is powered on, the baseboard management controller may poll the N components connected to the parallel processing device through the parallel processing device, thereby determining the component type of each of the N components one by one.
[0070] Step S302: In the first phase of each parameter reading cycle, the temperature information of each of the N components is read simultaneously in a parallel reading manner through its own N threads.
[0071] In the first stage of each parameter reading cycle, the parallel processing device uses its own N threads to simultaneously read the temperature information of N components in a parallel reading manner. In this way, even if the value of N is large, the time consumed by the solution of the present application in the first stage of the parameter reading cycle will not increase.
[0072] There are many ways to implement a parallel processing device. For example, in one embodiment of the present application, considering that MCU (Micro Controller Unit), CPLD (Complex Programmable Logic Device), and FPGA (Field Programmable Gate Array) all have the ability to implement multithreading, in one embodiment of the present application, the parallel processing device can be a parallel processing device based on a microcontroller unit, or a parallel processing device based on a field programmable gate array, or a parallel processing device based on a complex programmable logic device. For example, the first controller described in the embodiments below can be specifically an MCU, or a CPLD, or an FPGA. Similarly, each second controller described in the embodiments below can be an MCU, or a CPLD, or an FPGA.
[0073] Furthermore, the parallel processing device can be implemented by a single chip or by multiple chips. Both methods have their own advantages. Both implementation methods will be described in detail below.
[0074] Step S303: In the second phase of each parameter reading cycle, the temperature information of each component whose component type is a graphics processor is sent to the baseboard management controller.
[0075] In the first phase, the parallel processing device simultaneously reads the temperature information of N components through its own N threads in a parallel reading manner. Since GPUs (Graphics Processing Units) are highly temperature-sensitive components and are prone to overheating, the solution of the present application specifically sends the temperature information of each component of the graphics processor type to the baseboard management controller in the second phase of each parameter reading cycle. In other words, the temperature information of each GPU among the N components is preferentially sent to the baseboard management controller so that the baseboard management controller can promptly obtain the temperature information of each GPU and perform temperature control, which is conducive to ensuring the stability of GPU temperature.
[0076] Step S304: in the third phase of each parameter reading cycle, the temperature information of each component other than the graphics processor is sent to the baseboard management controller, so that the baseboard management controller performs temperature control of the server based on the temperature information of each of the N components.
[0077] After the temperature information of each GPU is sent, the third stage of each parameter reading cycle can be entered. In this stage, the parallel processing device can send the temperature information of each component except the GPU to the baseboard management controller, so that the baseboard management controller can control the temperature of the server based on the temperature information of each N component.
[0078] When the baseboard management controller controls the temperature of the server based on the temperature information of each of N components, there can be multiple specific implementation methods, which can be set and adjusted according to actual needs without affecting the implementation of this application. For example, at any time, when the baseboard management controller determines that a component is overheated, it can immediately control the speed of the relevant fan to increase to reduce the temperature of the component.
[0079] For the temperature information of any component, the specific content included in the temperature information of the component can also be set and adjusted according to actual needs. It usually includes the temperature of the component. In some cases, there may also be other temperature information of the component. For example, the load size of the component can also be regarded as a reflection of the temperature of the component, and therefore can also be used as the temperature information of the component. Of course, the setting of the specific content included in the temperature information usually needs to be coordinated with the relevant algorithms used by the baseboard management controller for temperature control to realize the temperature control of the server to ensure that each component operates within the appropriate temperature range.
[0080] In a specific embodiment of the present application, the parallel processing device includes a single first controller having at least N threads.
[0081] As described above, the parallel processing device can be implemented by a single chip or multiple chips. In this embodiment, the parallel processing device is composed of a single controller, that is, implemented by a single chip, and this controller is called the first controller.
[0082] When the parallel processing device is composed of the first controller, it is required that the first controller has at least N threads. This implementation can significantly reduce the time spent reading temperature information. The first controller usually uses N sets of independently operated I2C buses to connect N components respectively.
[0083] Further, in a specific embodiment of the present application, the first controller has a first interface and a second interface for connecting to a baseboard management controller;
[0084] Accordingly, step S303 specifically includes:
[0085] In the second phase of each parameter reading cycle, the temperature information of each component whose component type is a graphics processor is sent to the baseboard management controller through the first interface of the first controller itself;
[0086] Accordingly, in the third phase of each parameter reading cycle described in step S304, sending the temperature information of each component of a non-GPU component type to the baseboard management controller may specifically include:
[0087] In the third phase of each parameter reading cycle, the temperature information of each component whose component type is not the graphics processor is sent to the baseboard management controller through the second interface of the first controller itself.
[0088] In this implementation, it is taken into account that for a certain backplane, the baseboard management controller usually reserves one or two I2C channels to read the temperature information of each component on the backplane, and in the solution of the present application, the temperature information of each GPU needs to be sent to the baseboard management controller in a timely manner. Therefore, in this implementation, the first controller has a first interface and a second interface for connecting to the baseboard management controller. This implementation is shown in Figure 4. That is, in Figure 4, the parallel processing device specifically adopts the first controller, and the first controller is connected to the baseboard management controller through two I2C buses, one of which uses the first interface to connect to the BMC, marked as I2C-13 in Figure 4, and the other uses the second interface to connect to the BMC, marked as I2C-14 in Figure 4.
[0089] Since only the first controller is needed in this implementation, the BMC will poll I2C-13 and I2C-14 in each parameter reading cycle. Specifically, the temperature information of each GPU is read through I2C-13 in the second phase of each parameter reading cycle. At this time, the first controller can send the temperature information of each component of the GPU type to the BMC through the first interface of the first controller itself.
[0090] The BMC can read the temperature information through I2C-14 in the third phase of each parameter reading cycle. At this time, the first controller can send temperature information of each component of a non-GPU component type to the BMC through the second interface of the first controller itself.
[0091] It should also be noted that in the embodiment of FIG4 , the BMC is connected to the first controller via two I2C buses, one of which is dedicated to receiving GPU temperature information sent by the first controller, while the other is used to receive temperature information from other components sent by the first controller. If, in some cases, the BMC only reserves one I2C bus for connection to the first controller, then the first controller can simply connect to the BMC via a single interface, and temperature information for all types of components is then sent to the BMC via that interface. Of course, in actual applications, the BMC typically reserves two I2C buses. This is because different types of components use different messages for temperature data transmission. If all different types of components were sent to the BMC via a single I2C bus, configuring that bus would be more complex, and prioritizing GPU temperature information would be necessary. However, in this embodiment of the present application, a dedicated bus is used to transmit temperature information for each GPU. The BMC only needs to prioritize reading temperature information from each GPU through that bus each time, thus ensuring efficient transmission of temperature information for each GPU and simplifying design.
[0092] In a specific embodiment of the present application, among the N components, there are K types of components excluding the graphics processor, where K is a positive integer not less than 2, and the third phase of the parameter reading cycle is divided into K sub-phases, where i is a positive integer and 1≤i≤K;
[0093] Accordingly, in the third phase of each parameter reading cycle, sending the temperature information of each component of a non-GPU type to the baseboard management controller via the second interface of the first controller itself may specifically include:
[0094] In the i-th sub-phase of the third phase of each parameter reading cycle, the temperature information of the i-th component of a non-graphics processor type is sent to the baseboard management controller through the second interface of the first controller itself.
[0095] This implementation takes into account that when the first controller sends component temperature information to the BMC, the temperature information of components of the same type can be sent all at once. For example, in the above example, the first interface of the first controller is dedicated to transmitting GPU temperature information. Therefore, when the BMC reads through I2C-13 in the second phase of each parameter reading cycle, the first controller can send the temperature information of each GPU to the baseboard management controller all at once through its own first interface.
[0096] When the BMC reads through I2C-14 in the third phase of each parameter reading cycle, since there may be one or more component types other than the GPU, it is necessary to send the temperature information in different sub-phases according to the different component types. That is, in the i-th sub-phase of the third phase of each parameter reading cycle described in this embodiment, the temperature information of the i-th component type other than the GPU is sent to the baseboard management controller through the second interface of the first controller itself.
[0097] Table 1 is used as an example for explanation. Table 1 is a comparison table of three different examples of the implementation of FIG. 4 and a traditional solution.
[0098] Table 1:
[0099] The time values in Table 1 are all in milliseconds. It can be seen that in the traditional solution, after the BMC sequentially polls 11 PCIe cards, it takes 11×100=1100 ms to obtain the temperature information of all 11 PCIe cards.
[0100] In Example 1 of Table 1, it is assumed that the first controller of Figure 4 is connected to 10 GPUs and 1 non-GPU type component. The 10 GPUs are, for example, GPU0 to GPU9 in Figure 4, and the 1 non-GPU type component is, for example, AIC2 in Figure 4. AIC is also known as Add-In Card. This application uses AIC to represent non-GPU type components, such as network cards, RAID (Redundant Arrays of Independent Disks) cards, HCA (Host Channel Adapter) cards, HBA (Host Bus Adapter) cards, etc.
[0101] In Example 1 of Table 1, the first stage of the parameter reading cycle takes 100ms. Within these 100ms, the first controller obtains the temperature information of all 11 PCIe cards. The second stage is 5ms. At this time, the BMC reads I2C-13, that is, reads the first interface of the first controller, so that the BMC can obtain the temperature information of each GPU at one time. The third stage is 5ms. At this time, the BMC reads I2C-14, that is, reads the second interface of the first controller, so that the BMC can obtain the temperature information of AIC2. It can be seen that in Example 1, the parameter reading cycle is a total of 110 milliseconds. Compared with the traditional architecture, (1100-110) / 1100=90%, which means that the efficiency can be increased by about 90%, which means that the original time consumption is reduced by 90%.
[0102] In Example 2 of Table 1, it is assumed that the first controller in FIG4 is connected to 9 GPUs and 3 non-GPU components. The 9 GPUs are, for example, GPU1 to GPU8 in FIG4 , and the 3 non-GPU components are AIC0, AIC1, and AIC2 in FIG4 . It is also assumed that AIC0, AIC1, and AIC2 are components of the same type, for example, the same RAID card.
[0103] In Example 2 of Table 1, the first phase of the parameter reading cycle takes 100ms. Within these 100ms, the first controller obtains the temperature information of all 11 PCIe cards. The second phase is 5ms. At this time, the BMC reads I2C-13, that is, reads the first interface of the first controller, so that the BMC can obtain the temperature information of these 9 GPUs at one time. Since AIC0, AIC1, and AIC2 are the same type of components, the third phase of Example 2 is 5ms. At this time, the BMC reads I2C-14, that is, reads the second interface of the first controller, so that the BMC can obtain the temperature information of AIC0, AIC1, and AIC2 at one time. In Example 2, the parameter reading cycle is a total of 110 milliseconds. Compared with the traditional architecture, (1100-110) / 1100=90%, which can increase efficiency by about 90%.
[0104] In Example 3 of Table 1, it is assumed that the first controller in Figure 4 is connected to 9 GPUs and 3 non-GPU type components. The 9 GPUs are, for example, GPU1 to GPU8 in Figure 4, and the 3 non-GPU type components are AIC0, AIC1, and AIC2 in Figure 2. The difference from Example 2 is that in Example 3, it is assumed that AIC0, AIC1, and AIC2 are components of different types.
[0105] In Example 3 of Table 1, the first and second phases of the parameter read cycle are the same as in Example 2 and will not be repeated. Since AIC0, AIC1, and AIC2 are different types of components, the third phase of Example 3 is three 5ms, that is, the third phase requires 15ms. At this time, the BMC obtains the temperature information of AIC0, AIC1, and AIC2 in sequence by reading I2C-14, that is, reading the second interface of the first controller. In Example 3, the parameter read cycle is a total of 120 milliseconds. Compared with the traditional architecture, (1100-120) / 1100≈89%, which means that the efficiency can be increased by about 89%, which means that the original time consumption is reduced by 89%.
[0106] In a specific embodiment of the present application, the parallel processing device includes M second controllers, the total number of threads of the M second controllers is greater than or equal to N, and M is a positive integer not less than 2.
[0107] As described above, the parallel processing device can be implemented by a single chip or by multiple chips, and the parallel processing device implemented by a single chip is described in detail above. In this embodiment, the parallel processing device is implemented by multiple chips, that is, the parallel processing device includes M controllers, and these M controllers are referred to as second controllers. M is a positive integer not less than 2. Since there are M second controllers, the total number of threads of the M second controllers is required to be greater than or equal to N.
[0108] When M second controllers are used, it can be understood that any one second controller is connected to at least one component among the N components, and any one component among the N components is connected to at most one second controller.
[0109] In the aforementioned implementation using the first controller, since a single chip is required to connect N components, the space requirement is high, requiring a continuous circuit board area for the layout of the first controller. Of course, this approach simplifies wiring. However, this implementation is equivalent to implementing a distributed parallel processing device through M second controllers. These M second controllers can be arranged in a distributed manner, eliminating the need for a continuous circuit board area and reducing the board requirements, but it does increase overall wiring complexity.
[0110] In a specific embodiment of the present application, the device models of the M second controllers are the same, each second controller has a threads, a is a positive integer and a×M≥N, and each second controller has a first interface and a second interface for connecting to the baseboard management controller.
[0111] As described above, the baseboard management controller usually reserves one or two I2C channels, and in the solution of the present application, the temperature information of each GPU needs to be sent to the baseboard management controller in a timely manner. Therefore, for the implementation method using M second controllers, each second controller is provided with a first interface and a second interface for connecting to the baseboard management controller. After adopting such a setting, for each second controller, a dedicated I2C channel can be used to realize the transmission of temperature information of each GPU connected to the second controller, so as to ensure the transmission efficiency of the temperature information of each GPU.
[0112] FIG5 shows three second controllers, ie, M=3, and each second controller can be connected to a maximum of four PCIe components.
[0113] In this implementation, the device models of the M second controllers are all the same, and each second controller has a threads. This is because this implementation is convenient for expansion and reduces the firmware adaptation workload of the staff. For example, in the case of Figure 4 above, there are 11 external components, so it is necessary to set up a first controller with at least N threads, which is connected to these 11 components through N groups of independently running I2C buses. Assuming that 5 more components are added during subsequent operation, the first controller cannot meet this requirement, and the staff needs to redesign a new first controller with more groups of independent I2Cs.
[0114] Taking Figure 5 as an example, since the device models of the M second controllers are all the same, if 5 components are added, it is only necessary to add another completely identical second controller on the basis of Figure 5, so that each second controller is connected to 4 components. In this way, there is no need for staff to redesign the firmware, that is, to use one more second controller of the same model, which makes this implementation method very flexible.
[0115] In a specific implementation of the present application, step S302 may specifically include:
[0116] In the first stage of each parameter reading cycle, each second controller reads the temperature information of each component connected to itself simultaneously through its own a threads in a parallel reading manner.
[0117] In this embodiment, since a solution of M second controllers is adopted, in the first stage of each parameter reading cycle, each second controller can read the temperature information of each component connected to itself simultaneously through its own a threads in a parallel reading manner. For example, in actual applications, since the I2C bus is usually used for connection, each second controller can be connected to a maximum of a components through a group of independently working I2C buses, thereby simultaneously obtaining the temperature information of each component connected to itself.
[0118] In a specific embodiment of the present application, the second stage is divided into M sub-stages, j is a positive integer and 1≤j≤M;
[0119] Step S303 may specifically include:
[0120] In the jth sub-phase of the second phase of each parameter reading cycle, the jth second controller among the M second controllers sends the temperature information of each component of the graphics processor type connected to the jth second controller itself to the baseboard management controller through its own first interface.
[0121] In the aforementioned embodiment, since it is a parallel processing device implemented by a single controller, in the second phase, the first controller can directly send the temperature information of each GPU to the BMC. In this embodiment, since it is a parallel processing device implemented by M second controllers, the BMC cannot communicate with the M second controllers at the same time, and each second controller may be connected to a GPU. Therefore, in this embodiment, the BMC needs to poll the M second controllers in the second phase, that is, the second phase needs to be divided into M sub-phases, so that in the j-th sub-phase of the second phase of each parameter reading cycle, the j-th second controller among the M second controllers sends the temperature information of each GPU it has obtained to the BMC. Of course, if a second controller is not connected to any GPU, the BMC will skip this second controller in the second phase.
[0122] In a specific embodiment of the present application, among the N components, there are K types of components excluding the graphics processor, where K is a positive integer not less than 2, and the third phase of the parameter reading cycle is divided into K rounds, each round is divided into M sub-phases, where i is a positive integer and 1≤i≤K, and j is a positive integer and 1≤j≤M;
[0123] Step S304, in the third phase of each parameter reading cycle, sends the temperature information of each component of a non-GPU type to the baseboard management controller, which may specifically include:
[0124] During the jth sub-period of the i-th round of the third phase of each parameter reading cycle, the j-th second controller among the M second controllers sends temperature information of the i-th component of a non-graphics processor type connected to the j-th second controller to the baseboard management controller via its second interface.
[0125] In this implementation, since a solution of M second controllers is adopted and there may be multiple component types other than the GPU, the third stage of each parameter reading cycle needs to be divided into K rounds, and each round is divided into M sub-stages to realize the transmission of temperature information of different component types of different second controllers.
[0126] Table 2 is used as an example for explanation. Table 2 is a comparison table of three different examples of the implementation of FIG. 5 and a traditional solution.
[0127] Table 2:
[0128] The time values in Table 2 are all in milliseconds. It can be seen that in the traditional solution, after the BMC sequentially polls 11 PCIe cards, it takes 11×100=1100 ms to obtain the temperature information of all 11 PCIe cards.
[0129] In Example 1 of Table 2, it is assumed that the three second controllers in FIG5 are connected to a total of 10 GPUs and 1 non-GPU type component. The 10 GPUs are GPU0 to GPU9 in FIG5 , and the 1 non-GPU type component is, for example, AIC2 in FIG5 .
[0130] In Example 1 of Table 2, the first phase of the parameter reading cycle takes 100ms. During this 100ms, each of the three second controllers obtains the temperature information of the components connected to it. The second phase is divided into three sub-phases. First, the BMC reads I2C-13, at this time polling the first interface of the first second controller in Figure 5, so that the BMC can obtain the temperature information of each GPU connected to the first second controller at one time. Then, the BMC reads I2C-13, at this time polling the first interface of the second second controller in Figure 5, so that the BMC can obtain the temperature information of each GPU connected to the second second controller at one time. Finally, the BMC reads I2C-13, at this time polling the first interface of the third second controller in Figure 5, so that the BMC can obtain the temperature information of each GPU connected to the third second controller at one time. At this point, the second phase of the parameter reading cycle ends.
[0131] In Example 1 in Table 2, there is only one non-GPU component. The information obtained in step S301 indicates that this component is connected to the third second controller. Therefore, during the third phase of the parameter read cycle, the BMC reads I2C-14, specifically the second interface of the third second controller, to obtain the temperature information of AIC2. As can be seen, in Example 1, the parameter read cycle is a total of 120 milliseconds. Compared to the traditional architecture, (1100-120) / 1100≈89%, which represents an 89% increase in efficiency and an 89% reduction in time.
[0132] In Example 2 of Table 2, it is assumed that the three second controllers in FIG5 are connected to nine GPUs and three non-GPU components. The nine GPUs are, for example, GPU1 to GPU8 in FIG5 , and the three non-GPU components are AIC0, AIC1, and AIC2 in FIG5 . It is also assumed that AIC0, AIC1, and AIC2 are components of the same type, for example, the same RAID card.
[0133] In Example 2 in Table 2, the first phase of the parameter reading cycle takes 100ms. During this 100ms, each of the three secondary controllers acquires the temperature information of its connected components. The second phase is divided into three sub-phases, which, like Example 1, also take 15ms and will not be repeated here.
[0134] In Example 2 of Table 2, there are three non-GPU components. Based on the information obtained in step S301, we know that one of the three non-GPU components is connected to the first second controller, and the other two are connected to the third second controller. Therefore, during the third phase of the parameter read cycle, the BMC can read I2C-14 and first poll the second interface of the first second controller, then the second interface of the third second controller, taking a total of 10ms. As can be seen, in Example 2 of Table 2, the parameter read cycle is a total of 125 milliseconds. Compared to the traditional architecture, (1100-125) / 1100≈88.6%, which roughly increases efficiency by 88.6%.
[0135] In Example 3 of Table 2, it is assumed that the three second controllers in Figure 5 are connected to 9 GPUs and 3 non-GPU type components. The 9 GPUs are, for example, GPU1 to GPU8 in Figure 5, and the 3 non-GPU type components are AIC0, AIC1, and AIC2 in Figure 5, and it is assumed that AIC0, AIC1, and AIC2 are components of different types.
[0136] In Example 3 of Table 2, the first and second phases of the parameter reading cycle are the same as those of Example 2 of Table 2 and are not described again. Since AIC0, AIC1, and AIC2 are different types of components, and the BMC can know that one of the three non-GPU type components is connected to the first second controller, and the other two are connected to the third second controller, therefore, in the first sub-period of the first round of the third phase of the parameter reading cycle, the BMC can poll the second interface of the first second controller by reading I2C-14, thereby reading the temperature data of AIC0. Since there are no components of the same type, the first round of the third phase ends, and then the BMC can directly enter the third sub-period of the second round of the third phase. At this time, the BMC can poll the second interface of the third second controller by reading I2C-14, thereby reading the temperature data of AIC1. Since there are no components of the same type, the second round of the third phase ends. Finally, we can directly enter the third sub-period of the third round of the third phase. At this point, the BMC can poll the second interface of the third second controller by reading I2C-14, thereby reading the temperature data of AIC2. Since there are no components of the same type, the third round of the third phase ends. As can be seen, the third phase in this example takes a total of 15ms. As can be seen in Example 3 in Table 2, the parameter reading cycle is a total of 130 milliseconds. Compared with the traditional architecture, (1100-130) / 1100≈88%, which means that the efficiency can be increased by approximately 88.6%.
[0137] It can be understood that the temperature is the sub-stages or sub-periods described above in this application. They are only used to distinguish the temperature data of different types of components sent by different controllers, and do not mean that each sub-stage or sub-period has temperature data that needs to be sent to the BMC. In particular, when adopting an implementation method of M second controllers, the types of components connected to different second controllers may not be exactly the same. For sub-stages or sub-periods that do not need to send any temperature data, they can be skipped directly. For example, in Example 3 of Table 2 above, there are 3 types of components other than the GPU, that is, K=3, and the number of second controllers is 3, that is, M=3. Therefore, in the third stage, theoretically, there are a maximum of 9 sub-periods that need to send temperature data to the BMC. However, in Example 3 of Table 2 above, since there are only 3 components other than the GPU, the third stage ends after only 3 sub-periods.
[0138] In addition, the examples in Tables 1 and 2 also show that whether a single controller or multiple controllers are used to implement parallel processing devices, efficiency can be effectively increased, that is, the time required to read BMC temperature data is effectively reduced. Of course, when a single controller is used, the efficiency increase is slightly higher than when multiple controllers are used.
[0139] In a specific embodiment of the present application, it may further include:
[0140] When a fault signal is received from any component, the sending of the temperature information of the current stage is suspended and the fault signal is sent to the baseboard management controller, and the sending of the temperature information of the current stage is continued after the sending of the fault signal is completed.
[0141] This embodiment takes into account that the parallel processing device of the present application can also be used to receive fault signals sent by various components. If a fault signal is received from any one component, in order to ensure the priority of sending the fault signal, in this embodiment, the sending of the temperature information of the current stage can be paused and the fault signal can be sent to the baseboard management controller, and the sending of the temperature information of the current stage can be continued after the fault signal is sent.
[0142] Applying the technical solution provided in the embodiments of the present application, considering that the baseboard management controller takes a long time to collect temperature information due to the large number of components caused by the baseboard management controller polling each component, the solution of the present application specifically provides a parallel processing device connected to the baseboard management controller to assist the baseboard management controller in acquiring temperature data and thus achieving temperature control of the server. Specifically, the parallel processing device is connected to N components. In the first phase of each parameter reading cycle, the parallel processing device uses its own N threads to simultaneously read the temperature information of each of the N components in a parallel reading manner. It can be seen that regardless of the number N, since the parallel processing device uses its own N threads to simultaneously read the temperature information of each of the N components in a parallel reading manner, even if the number of components is large or small, the time consumed in the first phase of the parameter reading cycle will not be increased. In the second and third phases, the parallel processing device can send the temperature information of each of the N components to the baseboard management controller, so that the baseboard management controller can control the temperature of the server. The time consumed in both the second and third phases is very short. Furthermore, the present application takes into account that different types of components have different temperature sensitivities. Compared to other components, the graphics processor is a component with high temperature sensitivity. Therefore, in the solution of the present application, the temperature information of each component whose component type is a graphics processor will be sent to the baseboard management controller in the second stage, so that the baseboard management controller can promptly obtain the temperature information of each graphics processor and then perform temperature control, which is conducive to ensuring the temperature stability of the graphics processor. Of course, since the solution of the present application needs to distinguish between component types, it is necessary to determine the component type of each of N components one by one after the server is powered on and send them to the baseboard management controller. Compared with traditional solutions, this operation is indeed an additional operation and is time-consuming, but it only needs to be performed once after the server is powered on, and does not affect the parameter reading cycle during the subsequent operation of the server.
[0143] In summary, the solution of the present application can effectively shorten the time taken by the baseboard management controller to obtain the temperature information of each component, and can timely obtain the temperature information of the highly temperature-sensitive graphics processor, so that the solution of the present application can more accurately and effectively realize the temperature control of the server and ensure the service life and reliability of the server.
[0144] Corresponding to the above method embodiment, an embodiment of the present application further provides a temperature control system for a server, as shown in FIG2 , which may include a baseboard management controller and a preset parallel processing device connected to the baseboard management controller. The parallel processing device is connected to N components, as shown in FIG6 . The parallel processing device may include:
[0145] A power-on detection module 601 is used to determine the component type of each of N components one by one and send the component type to the baseboard management controller after the server is powered on;
[0146] The first phase execution module 602 is configured to read the temperature information of N components simultaneously in a parallel reading manner through its own N threads in the first phase of each parameter reading cycle;
[0147] The second phase execution module 603 is configured to send temperature information of each component whose component type is a graphics processor to a baseboard management controller in the second phase of each parameter reading cycle;
[0148] The third phase execution module 604 is configured to send temperature information of each component other than the graphics processor to the baseboard management controller during the third phase of each parameter reading cycle, so that the baseboard management controller performs temperature control of the server based on the temperature information of each of the N components.
[0149] In a specific embodiment of the present application, the parallel processing device includes a single first controller having at least N threads.
[0150] In a specific embodiment of the present application, the first controller has a first interface and a second interface for connecting to a baseboard management controller;
[0151] Accordingly, the second-stage execution module 603 is specifically configured to:
[0152] In the second phase of each parameter reading cycle, the temperature information of each component whose component type is a graphics processor is sent to the baseboard management controller through the first interface of the first controller itself;
[0153] Accordingly, the third stage execution module 604 is specifically used to:
[0154] In the third phase of each parameter reading cycle, temperature information of each component of a non-graphics processor type is sent to the baseboard management controller through the second interface of the first controller itself, so that the baseboard management controller performs temperature control of the server based on the temperature information of each of the N components.
[0155] In a specific embodiment of the present application, there are K types of components excluding the graphics processor, where K is a positive integer not less than 2, and the third phase of the parameter reading cycle is divided into K sub-phases, where i is a positive integer and 1≤i≤K;
[0156] Accordingly, the third stage execution module 604 is specifically used to:
[0157] In the i-th sub-stage of the third stage of each parameter reading cycle, the temperature information of the i-th component of the component type other than the graphics processor is sent to the baseboard management controller through the second interface of the first controller itself, so that the baseboard management controller controls the temperature of the server based on the temperature information of each of the N components.
[0158] In a specific embodiment of the present application, the parallel processing device includes M second controllers, the total number of threads of the M second controllers is greater than or equal to N, M is a positive integer not less than 2, any second controller is connected to at least one of the N components, and any one of the N components is connected to at most one second controller.
[0159] In a specific embodiment of the present application, the device models of the M second controllers are the same, each second controller has a threads, a is a positive integer and a×M≥N, and each second controller has a first interface and a second interface for connecting to the baseboard management controller.
[0160] In a specific embodiment of the present application, the first phase execution module 602 is specifically configured to:
[0161] In the first stage of each parameter reading cycle, each second controller reads the temperature information of each component connected to itself simultaneously through its own a threads in a parallel reading manner.
[0162] In a specific embodiment of the present application, the second stage is divided into M sub-stages, j is a positive integer and 1≤j≤M;
[0163] Accordingly, the second-stage execution module 603 is specifically configured to:
[0164] In the jth sub-phase of the second phase of each parameter reading cycle, the jth second controller among the M second controllers sends the temperature information of each component of the graphics processor type connected to the jth second controller itself to the baseboard management controller through its own first interface.
[0165] In a specific embodiment of the present application, there are K types of components excluding the graphics processor, where K is a positive integer not less than 2, and the third phase of the parameter reading cycle is divided into K rounds, each round is divided into M sub-phases, where i is a positive integer with 1≤i≤K, and j is a positive integer with 1≤j≤M.
[0166] Accordingly, the third stage execution module 604 is specifically used to:
[0167] During the j-th sub-period of the i-th round of the third phase of each parameter reading cycle, the j-th second controller among the M second controllers sends temperature information of the i-th component of a non-graphics processor type connected to the j-th second controller to the baseboard management controller via its second interface, so that the baseboard management controller performs temperature control of the server based on the temperature information of each of the N components.
[0168] In a specific embodiment of the present application, the parallel processing device is a parallel processing device based on a microcontroller unit, or a parallel processing device based on a field programmable gate array, or a parallel processing device based on a complex programmable logic device.
[0169] In a specific embodiment of the present application, it also includes:
[0170] The fault signal processing module is used to suspend the transmission of temperature information of the current stage and send the fault signal to the baseboard management controller when receiving a fault signal sent by any component, and continue the transmission of temperature information of the current stage after the fault signal is sent.
[0171] Corresponding to the above method and system embodiments, embodiments of the present application further provide a temperature control device for a server and a computer-readable storage medium, which can be referenced in correspondence with the above.
[0172] Referring to FIG7 , the temperature control device of the server may include:
[0173] Memory 701, used for storing computer programs;
[0174] The processor 702 is configured to execute a computer program to implement the steps of the temperature control method for a server in any of the above embodiments.
[0175] Referring to FIG8 , a computer-readable storage medium 80 stores a computer program 81. When executed by a processor, the computer program 81 implements the steps of the server temperature control method described in any of the above embodiments. The computer-readable storage medium 80 herein includes random access memory (RAM), internal memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, hard disks, removable disks, CD-ROMs, or any other form of storage medium known in the art.
[0176] Referring to FIG. 9 , which is a flowchart illustrating an implementation of a temperature control method for a server in the present application, a baseboard management controller is connected to a preset parallel processing device, which is connected to N components. The temperature control method for the server is applied to the baseboard management controller and includes:
[0177] Step S901: Power on the server, and after determining the component types of N components one by one through the parallel processing device, receive the component types of the N components sent by the parallel processing device;
[0178] Step S902: In the second phase of each parameter reading cycle, receiving temperature information of each component of the graphics processor type sent by the parallel processing device;
[0179] Step S903: In the third phase of each parameter reading cycle, receiving temperature information of each component of a non-GPU type sent by the parallel processing device;
[0180] Step S904: performing temperature control on the server based on the temperature information of each of the N components;
[0181] In the first stage of each parameter reading cycle, the parallel processing device reads the temperature information of each of the N components simultaneously in a parallel reading manner through its own N threads.
[0182] It should also be noted that, in this document, relational terms such as first and second, etc., are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "comprises," "comprising," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or device. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of additional identical elements in the process, method, article, or device comprising the element.
[0183] Professionals may further appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the above description has generally described the components and steps of each example according to their functions. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professionals and technicians may use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0184] Specific examples are used herein to illustrate the principles and implementation methods of this application. The description of the above embodiments is only intended to help understand the technical solution and core ideas of this application. It should be noted that, for those skilled in the art, without departing from the principles of this application, various improvements and modifications may be made to this application, and such improvements and modifications also fall within the scope of protection of this application.
Claims
1. A temperature control method for a server, characterized in that, The baseboard management controller is connected to a preset parallel processing device, and the parallel processing device is connected to N components. The temperature control method of the server is applied to the parallel processing device and includes: After the server is powered on, determine the component type of each of the N components one by one and send it to the baseboard management controller; In the first stage of each parameter reading cycle, simultaneously read the temperature information of each of the N components in a parallel reading manner through N threads of itself; In the second stage of each parameter reading cycle, send the temperature information of each component whose component type is a graphics processor to the baseboard management controller; In the third stage of each parameter reading cycle, send the temperature information of each component whose component type is not a graphics processor to the baseboard management controller, so that the baseboard management controller performs temperature control of the server based on the temperature information of each of the N components.
2. The temperature control method of the server according to claim 1, wherein The parallel processing device includes a single first controller having at least N threads.
3. The temperature control method of the server according to claim 2, characterized in that, The first controller has a first interface and a second interface for connecting to the baseboard management controller; In the second stage of each parameter reading cycle, sending the temperature information of each component whose component type is a graphics processor to the baseboard management controller includes: In the second stage of each parameter reading cycle, send the temperature information of each component whose component type is a graphics processor to the baseboard management controller through the first interface of the first controller itself; In the third stage of each parameter reading cycle, sending the temperature information of each component whose component type is not a graphics processor to the baseboard management controller includes: In the third stage of each parameter reading cycle, send the temperature information of each component whose component type is not a graphics processor to the baseboard management controller through the second interface of the first controller itself.
4. The temperature control method of the server according to claim 3, wherein Among the N components, there are K types of component types other than the graphics processor. K is a positive integer not less than 2. The third stage of the parameter reading cycle is divided into K sub-stages, and i is a positive integer and 1 ≤ i ≤ K; In the third stage of each parameter reading cycle, sending the temperature information of each component whose component type is not a graphics processor to the baseboard management controller through the second interface of the first controller itself includes: In the i-th sub-stage of the third stage of each parameter reading cycle, send the temperature information of the i-th type of component whose component type is not a graphics processor to the baseboard management controller through the second interface of the first controller itself.
5. The temperature control method of the server according to claim 1, wherein The parallel processing device includes M second controllers. The total number of threads of the M second controllers is greater than or equal to N. M is a positive integer not less than 2. Any one of the second controllers is connected to at least one of the N components, and any one of the N components is connected to at most one of the second controllers.
6. The temperature control method of the server according to claim 5, characterized in that, The device models of the M second controllers are all the same. Each second controller has a threads. a is a positive integer and a×M ≥ N. And each second controller has a first interface and a second interface for connecting to the baseboard management controller.
7. The temperature control method of the server according to claim 6, wherein, In the first stage of each parameter reading cycle, the temperature information of each of the N components is simultaneously read in a parallel reading manner through N threads of its own, including: In the first stage of each parameter reading cycle, each of the second controllers reads the temperature information of each component connected to itself in a parallel reading manner through a threads of its own.
8. The temperature control method of the server according to claim 6, characterized in that, The second stage is divided into M sub-stages, where j is a positive integer and 1 ≤ j ≤ M; In the second stage of each parameter reading cycle, the temperature information of each component with a component type of graphics processor is sent to the baseboard management controller, including: In the j-th sub-stage of the second stage of each parameter reading cycle, the j-th second controller among the M second controllers sends the temperature information of each component with a component type of graphics processor connected to the j-th second controller itself to the baseboard management controller through its first interface.
9. The temperature control method of the server according to claim 6, wherein Among the N components, there are K types of component types other than the graphics processor, where K is a positive integer not less than 2. The third stage of the parameter reading cycle is divided into K rounds, and each round is divided into M sub-stages. i is a positive integer and 1 ≤ i ≤ K, and j is a positive integer and 1 ≤ j ≤ M; In the third stage of each parameter reading cycle, the temperature information of each component with a component type other than the graphics processor is sent to the baseboard management controller, including: In the j-th sub-period of the i-th round of the third stage of each parameter reading cycle, the j-th second controller among the M second controllers sends the temperature information of the i-th type of component with a component type other than the graphics processor connected to the j-th second controller itself to the baseboard management controller through its second interface.
10. The temperature control method of the server according to claim 1, characterized in that, The parallel processing device is a parallel processing device based on a microcontroller unit, or a parallel processing device based on a field programmable gate array, or a parallel processing device based on a complex programmable logic device.
11. The temperature control method of the server according to any one of claims 1 to 9, characterized in that, It also includes: When a failure signal sent by any one component is received, the transmission of the temperature information in the current stage is paused, and the failure signal is sent to the baseboard management controller, and after the failure signal is sent, the transmission of the temperature information in the current stage is continued.
12. The temperature control method of the server according to claim 1, characterized in that, The step of individually determining the component type of each of the N components includes: polling the N components connected to the parallel processing device through the parallel processing device.
13. The temperature control method of the server according to claim 3, characterized in that, It also includes: In response to the temperature information of each component of the graphics processor and the temperature information of each component of the non-graphics processor being sent to the baseboard management controller through a single I2C bus, priorities are set for the temperature information of each component of the graphics processor.
14. The temperature control method of the server according to claim 3, characterized in that, The first interface is configured to be dedicated to transmitting the temperature information of each component with a component type of graphics processor at one time.
15. The temperature control method of the server according to claim 6, wherein Each of the second controllers is configured to use a dedicated I2C to implement the transmission of the temperature information of each component with a component type of graphics processor connected to itself.
16. The temperature control method of the server according to claim 1, characterized in that The parallel processing device is implemented by a single chip or multiple chips.
17. A temperature control system for a server, characterized in that, It includes a baseboard management controller and a preset parallel processing device connected to the baseboard management controller. The parallel processing device is connected to N components. The parallel processing device includes: The power-on detection module is used to determine the component types of N components one by one after the server is powered on and send them to the baseboard management controller; The first-stage execution module is used to simultaneously read the temperature information of N components in parallel through its N threads in the first stage of each parameter reading cycle; The second-stage execution module is used to send the temperature information of each component with the component type of graphics processor to the baseboard management controller in the second stage of each parameter reading cycle; The third-stage execution module is used to send the temperature information of each component with a component type other than the graphics processor to the baseboard management controller in the third stage of each parameter reading cycle, so that the baseboard management controller performs temperature control of the server based on the temperature information of N components; 18. A temperature control device for a server, characterized in that, It includes: A memory for storing computer programs; A processor for executing the computer program to implement the steps of the server temperature control method according to any one of claims 1 to 16; 19. A computer-readable storage medium, characterized in that, A computer program is stored on the computer-readable storage medium, and when the computer program is executed by the processor, the steps of the server temperature control method according to any one of claims 1 to 16 are implemented; 20. A temperature control method for a server, characterized in that, The baseboard management controller is connected to a preset parallel processing device, and the parallel processing device is connected to N components. The server temperature control method is applied to the baseboard management controller and includes: After the server is powered on and the component types of N components are determined one by one through the parallel processing device; Receiving the component types of N components sent by the parallel processing device; Receiving the temperature information of each component with the component type of graphics processor sent by the parallel processing device in the second stage of each parameter reading cycle; Receiving the temperature information of each component with a component type other than the graphics processor sent by the parallel processing device in the third stage of each parameter reading cycle; Performing temperature control of the server based on the temperature information of N components; Wherein, in the first stage of each parameter reading cycle, the parallel processing device simultaneously reads the temperature information of N components in parallel through its N threads.
Citation Information
Patent Citations
Memory temperature reading method and system
CN111949465A
Bus scheduling method, device and equipment, medium and substrate management control chip
CN116185915A
Server and heat dissipation control method thereof
CN116466810A
Heat dissipation regulation and control optimization method, device and equipment of GPU acceleration card, medium and product
CN117149567A
Temperature control method, system and equipment of server and storage medium
CN117591378A