Control Method, Device and Computer Readable Storage Medium of Heat Dissipation Device

By deploying a dual operating system in the server, the first operating system quickly takes over the control of the cooling device after the controller is restarted, solving the problem of low timeliness of the cooling device control after the controller is restarted, ensuring the cooling effect and efficiency of the server.

CN119292438BActive Publication Date: 2025-07-08INSPUR SUZHOU INTELLIGENT TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202411832335.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-12
Publication Date
2025-07-08
Estimated Expiration
2044-12-12

AI Technical Summary

Technical Problem

In the prior art, the timeliness of the cooling device control of the server after the controller is restarted is low, resulting in a decrease in the server operation efficiency.

Method used

A dual operating system (first operating system and second operating system) is deployed in the server. The first operating system quickly takes over the control of the cooling device after the controller is restarted, and finely controls it according to the ambient temperature and component temperature; after the second operating system is started, the control power is handed over to the second operating system to continue to control.

Benefits of technology

It realizes that the cooling device is quickly resumed after the controller is restarted, ensuring the heat dissipation effect and efficiency of the server, and improving the timeliness and accuracy of the control of the cooling device.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119292438B_ABST
    Figure CN119292438B_ABST
Patent Text Reader

Abstract

An embodiment of the present application provides a control method, device, and computer-readable storage medium for a heat dissipation device. Among them, the server includes: server components, a first controller, and a heat dissipation device. The first operating system and the second operating system are deployed on the first controller. The method includes: when the first controller restarts during the operation of the server, the first operating system controls the operation of the heat dissipation device according to the environmental temperature of the server and the component temperature of the first category of server components, and the first operating system detects the system running state of the second operating system; when the detected system running state is the started state, the first operating system stops controlling the operation of the heat dissipation device; the second operating system controls the operation of the heat dissipation device according to the environmental temperature of the server and the component temperature of the second category of server components. It solves the problem of low timeliness in controlling the heat dissipation device.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of the present application relate to the field of computers, and more particularly, to a control method, apparatus, and computer-readable storage medium for a heat dissipation device. Background Art

[0002] The heat dissipation condition on a server has a crucial impact on the component lifespan and service operation of the server. During the operation of the server, the controller used to control the operation of the heat dissipation device on the server may perform a restart operation. However, in the current related technologies, in this case, the heat dissipation device on the server can only be controlled after the controller restart is completed, and the heat dissipation device will be in a state where it cannot dissipate heat for the server for a long time, thereby affecting the operation efficiency of the server.

[0003] In view of the problems in the related technologies, such as the low timeliness of controlling the heat dissipation device, no effective solution has been proposed yet. Summary of the Invention

[0004] Embodiments of the present application provide a control method, apparatus, and computer-readable storage medium for a heat dissipation device, so as to at least solve the problem of low timeliness of controlling the heat dissipation device in the related technologies.

[0005] According to an embodiment of the present application, a control method for a heat dissipation device is provided. The server includes: server components, a first controller, and a heat dissipation device. A first operating system and a second operating system are deployed on the first controller. The method includes:

[0006] When the first controller restarts during the operation of the server, the first operating system controls the operation of the heat dissipation device according to the environmental temperature of the server and the component temperature of the first type of server components, and the first operating system detects the system operation state of the second operating system, where the first type of server components are the server components that allow the first operating system to collect component temperatures;

[0007] When it is detected that the system operation state is the started state, the first operating system stops controlling the operation of the heat dissipation device; the second operating system controls the operation of the heat dissipation device according to the environmental temperature of the server and the component temperature of the second type of server components, where the second type of server components are the server components that allow the second operating system to collect component temperatures.

[0008] In an exemplary embodiment, a first control curve and a first control configuration are further configured on the first controller. The first control curve is used to indicate a first corresponding relationship between the environmental temperature of the server and the operating parameters of the heat dissipation device. The first control configuration is used to indicate the parameter configuration in the operating parameter algorithm corresponding to the server components of the first category. The operating parameter algorithm is used to calculate the operating parameters of the heat dissipation device according to the component temperature of the server components;

[0009] The first operating system controls the operation of the heat dissipation device according to the environmental temperature of the server and the component temperature of the server components of the first category, including:

[0010] The first operating system calls the first control curve to convert the environmental temperature of the server into a first operating parameter, and the first operating system calls the first control configuration to calculate a second operating parameter according to the component temperature of the server components of the first category;

[0011] The first operating system determines a target operating parameter according to the first operating parameter and the second operating parameter;

[0012] The first operating system controls the operation of the heat dissipation device according to the target operating parameter.

[0013] In an exemplary embodiment, the first operating system calls the first control curve to convert the environmental temperature of the server into a first operating parameter, and the first operating system calls the first control configuration to calculate a second operating parameter according to the component temperature of the server components of the first category, including:

[0014] During the startup process of the second operating system, the second operating system collects the environmental temperature of the server and the component temperature of the server components of the first category;

[0015] The second operating system sends the collected environmental temperature of the server and the component temperature of the server components of the first category to the first operating system;

[0016] The first operating system calls the first control curve to convert the environmental temperature of the server into a first operating parameter, and the first operating system calls the first control configuration to calculate a second operating parameter according to the component temperature of the server components of the first category.

[0017] In an exemplary embodiment, a target storage space is also deployed on the server, and the target storage space allows the first operating system and the second operating system to access it. The second operating system sending the collected ambient temperature of the server and the component temperatures of the server components of the first category to the first operating system includes:

[0018] The second operating system writes the collected ambient temperature of the server and the component temperatures of the server components of the first category into the target storage space;

[0019] The second operating system sends an interrupt request to the first operating system;

[0020] The first operating system responds to the interrupt request and reads the ambient temperature of the server and the component temperatures of the server components of the first category from the target storage space.

[0021] In an exemplary embodiment, a second control curve and a second control configuration are also configured on the first controller. The second control curve is used to indicate a second corresponding relationship between the ambient temperature of the server and the operating parameters of the heat dissipation device, and the second control configuration is used to indicate the parameter configuration in the operating parameter algorithm corresponding to the server components of the second category. The operating parameter algorithm is used to calculate the operating parameters of the heat dissipation device according to the component temperatures of the server components;

[0022] The second operating system controlling the operation of the heat dissipation device according to the ambient temperature of the server and the component temperatures of the server components of the second category includes:

[0023] The second operating system calls the second control curve to convert the ambient temperature of the server into a third operating parameter, and the second operating system calls the second control configuration to calculate a fourth operating parameter according to the component temperatures of the server components of the second category;

[0024] The second operating system determines a reference operating parameter according to the third operating parameter and the fourth operating parameter;

[0025] The second operating system controls the operation of the heat dissipation device according to the reference operating parameter.

[0026] In an exemplary embodiment, the heat dissipation performance of the heat dissipation device corresponding to the operating parameter under the first corresponding relationship for the same ambient temperature of the server is higher than that of the heat dissipation device corresponding to the operating parameter under the second corresponding relationship.

[0027] In an exemplary embodiment, the heat dissipation types of the server include multiple heat dissipation types, and the multiple heat dissipation types include the target heat dissipation type to which the heat dissipation device belongs. The second control curve includes a first curve and a second curve. For the same ambient temperature of the server, the heat dissipation performance of the heat dissipation device corresponding to the operating parameters under the first curve is higher than that corresponding to the operating parameters under the second curve.

[0028] The conversion of the ambient temperature of the server into the third operating parameter by the second operating system invoking the second control curve includes:

[0029] The second operating system detects the heat dissipation type adopted by the server among the multiple heat dissipation types;

[0030] When the heat dissipation type adopted by the server is the target heat dissipation type, the second operating system invokes the first curve to convert the ambient temperature of the server into the third operating parameter;

[0031] When the heat dissipation type adopted by the server is not the target heat dissipation type, the second operating system invokes the second curve to convert the ambient temperature of the server into the third operating parameter.

[0032] In an exemplary embodiment, the server components of the second category are divided into multiple sets of server components with different risk levels. The second control curve includes multiple parameter curves corresponding to the multiple risk levels one by one. The higher the risk level, the greater the impact of the component temperature on the server. For server components with a higher risk level, at the same ambient temperature of the server, the heat dissipation performance of the heat dissipation device corresponding to the operating parameters on the corresponding control curve is higher;

[0033] The conversion of the ambient temperature of the server into the third operating parameter by the second operating system invoking the second control curve includes:

[0034] The second operating system detects whether there are faulty components among the server components of the second category;

[0035] When it is detected that there are faulty components among the server components of the second category, the second operating system detects the target set of server components to which the faulty components belong in the multiple sets of server components with different risk levels;

[0036] The second operating system invokes the reference curve corresponding to the target risk level of the target set of server components from the multiple parameter curves to convert the ambient temperature of the server into the third operating parameter.

[0037] In an exemplary embodiment, the set of server components of multiple risk levels includes: a set of high heat dissipation risk components and a set of low heat dissipation risk components. The impact degree of the component temperature of the server components in the set of high heat dissipation risk components on the server is higher than that of the component temperature of the server components in the set of low heat dissipation risk components on the server. The multiple parameter curves include a third curve and a fourth curve. The heat dissipation performance of the heat dissipation device corresponding to the operating parameters of the same server's ambient temperature under the fourth curve is higher than that corresponding to the operating parameters under the third curve.

[0038] The second operating system detecting the target server component set to which the faulty component belongs in the set of server components of multiple risk levels includes: the second operating system detecting that the faulty component is a high heat dissipation risk component or a low heat dissipation risk component.

[0039] The second operating system calling the reference curve corresponding to the target risk level of the target server component set from the multiple parameter curves to convert the ambient temperature of the server into the third operating parameter includes: in the case where the faulty component is detected to be a high heat dissipation risk component, the second operating system calling the fourth curve to convert the ambient temperature of the server into the third operating parameter.

[0040] In the case where the faulty component is detected to be a low heat dissipation risk component, the second operating system calls the third curve to convert the ambient temperature of the server into the third operating parameter.

[0041] In an exemplary embodiment, the heat dissipation performance of the heat dissipation device corresponding to the operating parameters of the same server's ambient temperature under the third curve is higher than that corresponding to the operating parameters under the first curve.

[0042] In an exemplary embodiment, the heat dissipation types of the server include multiple heat dissipation types, and the multiple heat dissipation types include the target heat dissipation type to which the heat dissipation device belongs. The server components of the second category include heat dissipation high-risk components and heat dissipation low-risk components. The impact degree of the component temperature of the heat dissipation high-risk components on the server is higher than that of the component temperature of the heat dissipation low-risk components on the server. The second control curve includes a fifth curve, a sixth curve, a seventh curve, and an eighth curve. The heat dissipation performance of the heat dissipation device corresponding to the operating parameters of the same server at the environmental temperature under the fifth curve is higher than that corresponding to the operating parameters under the sixth curve. The heat dissipation performance of the heat dissipation device corresponding to the operating parameters of the same server at the environmental temperature under the eighth curve is higher than that corresponding to the operating parameters under the seventh curve. The heat dissipation performance of the heat dissipation device corresponding to the operating parameters of the same server at the environmental temperature under the seventh curve is higher than that corresponding to the operating parameters under the fifth curve;

[0043] The conversion of the environmental temperature of the server into the third operating parameter by the second operating system invoking the second control curve includes:

[0044] The second operating system detects whether there are faulty components among the server components of the second category;

[0045] In the case where faulty components are detected among the server components of the second category, the second operating system detects whether the faulty components are the heat dissipation high-risk components or the heat dissipation low-risk components. In the case where the faulty components are detected as the heat dissipation high-risk components, the second operating system invokes the eighth curve to convert the environmental temperature of the server into the third operating parameter. In the case where the faulty components are detected as the heat dissipation low-risk components, the second operating system invokes the seventh curve to convert the environmental temperature of the server into the third operating parameter;

[0046] In the case where no faulty components are detected among the server components of the second category, the second operating system detects the heat dissipation type adopted by the server among the multiple heat dissipation types. In the case where the heat dissipation type adopted by the server is the target heat dissipation type, the second operating system invokes the fifth curve to convert the environmental temperature of the server into the third operating parameter. In the case where the heat dissipation type adopted by the server is not the target heat dissipation type, the second operating system invokes the sixth curve to convert the environmental temperature of the server into the third operating parameter.

[0047] In an exemplary embodiment, the server components of the second category include: a plurality of server components, and the second control curve includes a plurality of component curves corresponding to the plurality of server components one by one. The higher the degree of influence of the component temperature on the server, the higher the heat dissipation performance of the heat dissipation device corresponding to the operating parameters on the corresponding control curve at the same ambient temperature of the server;

[0048] The conversion of the ambient temperature of the server into the third operating parameter by the second operating system calling the second control curve includes:

[0049] The second operating system detects whether there is a faulty component among the server components of the second category;

[0050] In the case where a faulty component is detected among the server components of the second category, the second operating system extracts the target server component belonging to the faulty component from the plurality of server components;

[0051] The second operating system calls the target curve corresponding to the target server component from the plurality of component curves to convert the ambient temperature of the server into the third operating parameter.

[0052] In an exemplary embodiment, the operating parameter algorithm is used to calculate the operating parameter of the heat dissipation device according to the component temperature at the sampling moment, the component temperature at the historical moment, and the operating parameter corresponding to the historical sampling moment of the server component;

[0053] The conversion of the second operating parameter by the first operating system calling the first control configuration according to the component temperature of the server components of the first category includes:

[0054] The first operating system calls the first control configuration and substitutes it into the operating parameter algorithm to obtain a target operating parameter algorithm;

[0055] The first operating system obtains the first component temperature at the sampling moment, the second component temperature at the historical moment, and the historical operating parameter corresponding to the historical sampling moment of the server components of the first category, where the historical operating parameter is the operating parameter corresponding to the server components of the first category calculated by the operating parameter algorithm at the historical sampling moment;

[0056] The first operating system substitutes the first component temperature, the second component temperature, and the historical operating parameter into the target operating parameter algorithm to obtain the second operating parameter.

[0057] In an exemplary embodiment, the operating parameter algorithm is PWM(k)=PWM(k - 1)+▽PWM;

[0058] Among them, ▽PWM = Kp * [T(k) - T(k - 1)] + Ki(T(k) - Tsp) + Kd * [[T(k) - T(k - 1)] - [T(k - 1) - T(k - 2)]], PWM(k) is the operating parameter corresponding to the current sampling moment, PWM(k - 1) is the operating parameter corresponding to the historical sampling moment, T(k) is the component temperature at the current sampling moment, the component temperatures at historical moments include T(k - 1) and T(k - 2), T(k - 1) is the component temperature at the previous sampling moment of the current sampling moment, T(k - 2) is the component temperature at the previous sampling moment of the previous sampling moment, the parameter configuration in the operating parameter algorithm includes: Kp, Ki, Kd, and Tsp, and Tsp is the temperature value that the component temperature of the server component is allowed to reach at most during operation.

[0059] In an exemplary embodiment, the first operating system detecting the system operating state of the second operating system includes:

[0060] The first operating system detecting whether the heartbeat of the second operating system is normal;

[0061] When the heartbeat of the second operating system is detected, the first operating system detecting the duration of the heartbeat of the second operating system;

[0062] When the duration of the heartbeat of the second operating system is equal to or exceeds the target duration, the first operating system determines that the detected system operating state is the started state.

[0063] In an exemplary embodiment, a second controller is further deployed on the server, and the method further includes:

[0064] When the first controller restarts, the second controller controls the operation of the heat dissipation device and monitors the startup result of the first controller;

[0065] When the startup result indicates that the startup of the first controller fails, the second controller identifies the operating state of the server;

[0066] When the identified operating state is the running state, the second controller continues to control the operation of the heat dissipation device.

[0067] In an exemplary embodiment, after the second controller monitors the startup result of the first controller, the method further includes:

[0068] When the startup result is used to indicate that the startup of the first controller is successful, the second controller stops controlling the operation of the heat dissipation device, and the first controller controls the operation of the heat dissipation device.

[0069] In an exemplary embodiment, the first controller is connected to the heat dissipation device through the second controller;

[0070] The second controller controls the operation of the heat dissipation device and monitors the startup result of the first controller, including: the second controller checks whether the heartbeat of the first operating system is normal; when it is detected that the heartbeat of the first operating system is abnormal, the second controller controls the operation of the heat dissipation device and checks whether the heartbeat of the second operating system is normal; when it is detected that the heartbeat of the second operating system is abnormal, it is determined that the startup of the first controller fails; when it is detected that the heartbeat of the first operating system is normal, it is determined that the startup of the first controller is successful.

[0071] In an exemplary embodiment, when the startup result is used to indicate that the startup of the first controller fails, the second controller identifies the operating state of the server, including: when the startup result is used to indicate that the startup of the first controller fails, the second controller discards the received operating parameters sent by the first controller and identifies the operating state of the server;

[0072] When the startup result is used to indicate that the startup of the first controller is successful, the second controller stops controlling the operation of the heat dissipation device, and the first controller controls the operation of the heat dissipation device, including: when the startup result is used to indicate that the startup of the first controller is successful, the second controller sends the received operating parameters sent by the first controller to the heat dissipation device.

[0073] According to another embodiment of the present application, a control device for a heat dissipation device is provided. The server includes: server components, a first controller, and a heat dissipation device. The first controller is deployed with a first operating system and a second operating system. The device includes:

[0074] A first control module, configured to, when the first controller restarts during the operation of the server, the first operating system controls the operation of the heat dissipation device according to the environmental temperature of the server and the component temperature of the first type of server components, and the first operating system detects the system operating state of the second operating system, where the first type of server components are server components that allow the first operating system to collect component temperatures;

[0075] A second control module, configured to, when detecting that the system operating state is the started state, stop the first operating system from controlling the operation of the heat dissipation device; and control the operation of the heat dissipation device by the second operating system according to the ambient temperature of the server and the component temperatures of the server components of the second category, where the server components of the second category are the server components that allow the second operating system to collect component temperatures.

[0076] According to another embodiment of the present application, there is also provided a computer-readable storage medium storing a computer program, where the computer program is configured to execute the steps in any one of the above method embodiments when running.

[0077] According to another embodiment of the present application, there is also provided an electronic device including a memory and a processor, where a computer program is stored in the memory, and the processor is configured to run the computer program to execute the steps in any one of the above method embodiments.

[0078] According to another embodiment of the present application, there is also provided a computer program product including a computer program, where the computer program implements the steps in any one of the above method embodiments when executed by a processor.

[0079] Through the present application, a dual operating system (i.e., a first operating system and a second operating system) is deployed on the first controller in the server. If the first controller restarts during the operation of the server, the first operating system can control the operation of the heat dissipation device in a short time, so that the operation of the heat dissipation device will not be interrupted for a long time, ensuring the heat dissipation effect of the server. And the first operating system can control the heat dissipation device according to the ambient temperature of the server and the component temperatures it can collect, making the operation control of the heat dissipation device fine and accurate during the control stage of the first operating system. The first operating system can also detect the system operating state of the second operating system. If the second operating system has been started, the control right of the heat dissipation device is handed over to the second operating system, and the second operating system controls the operation of the heat dissipation device according to the ambient temperature of the server and the component temperatures it can collect. Therefore, the technical problem of low timeliness in controlling the heat dissipation device can be solved, and the effect of improving the timeliness of controlling the heat dissipation device is achieved. Description of the Drawings

[0080] Figure 1 is a hardware structure block diagram of a server device of a control method for a heat dissipation device according to an embodiment of the present application;

[0081] Figure 2 is a flowchart of a control method for a heat dissipation device according to an embodiment of the present application;

[0082] Figure 3 is a flowchart of an operation method of a heat dissipation device according to an embodiment of the present application;

[0083] Figure 4 is a schematic diagram of an RTOS system controlling a server heat dissipation system according to an embodiment of the present application;

[0084] Figure 5 is a schematic diagram of the architecture of a BMC according to an embodiment of the present application;

[0085] Figure 6 is a schematic diagram of the control and operation of a heat dissipation device according to an embodiment of the present application;

[0086] Figure 7 is a block diagram of the structure of a control device of a heat dissipation device according to an embodiment of the present application. Specific Embodiments

[0087] The embodiments of the present application will be described in detail below with reference to the drawings and in conjunction with the embodiments.

[0088] It should be noted that the terms "first", "second", etc. in the specification, claims and drawings of the present application are used to distinguish similar objects, and do not necessarily describe a specific order or sequence.

[0089] The method embodiments provided in the embodiments of the present application can be executed on a server device or a similar computing device. Taking running on a server device as an example, Figure 1 is a hardware block diagram of a server device for a control method of a heat dissipation device according to an embodiment of the present application. As Figure 1 shown, the server device may include one or more ( Figure 1 only one is shown in the figure) processors 102 (the processors 102 may include, but are not limited to, processing devices such as a microprocessor MCU or a programmable logic device FPGA) and a memory 104 for storing data. Among them, the above server device may further include a transmission device 106 for communication functions and an input / output device 108. Those of ordinary skill in the art can understand that, Figure 1 the structure shown in the figure is only schematic and does not limit the structure of the above server device. For example, the server device may further include more or fewer components than those Figure 1 shown in the figure, or have a different configuration from that Figure 1 shown in the figure.

[0090] The memory 104 can be used to store computer programs, for example, software programs and modules of application software, such as the computer program corresponding to the control method of the heat dissipation device in the embodiments of the present application. The processor 102 executes various functional applications and data processing by running the computer program stored in the memory 104, that is, the above-mentioned method is implemented. The memory 104 may include a high-speed random access memory, and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memories, or other non-volatile solid-state memories. In some instances, the memory 104 may further include a memory remotely disposed relative to the processor 102, and these remote memories can be connected to the server device through a network. Examples of the above-mentioned network include but are not limited to the Internet, an enterprise intranet, a local area network, a mobile communication network, and combinations thereof.

[0091] The transmission device 106 is used to receive or send data via a network. Specific examples of the above-mentioned network may include a wireless network provided by a communication provider of the server device. In one instance, the transmission device 106 includes a network adapter (Network Interface Controller, abbreviated as NIC), which can be connected to other network devices through a base station and thus communicate with the Internet. In one instance, the transmission device 106 may be a radio frequency (RF) module, which is used to communicate with the Internet wirelessly.

[0092] In this embodiment, a control method for a heat dissipation device is provided. The server includes: server components, a first controller, and a heat dissipation device. A first operating system and a second operating system are deployed on the first controller. Figure 2 It is a flowchart of the control method for the heat dissipation device according to the embodiments of the present application, as Figure 2 shown, and the process includes the following steps:

[0093] Step S202, when the first controller restarts during the operation of the server, the first operating system controls the operation of the heat dissipation device according to the environmental temperature of the server and the component temperatures of the first category of server components, and the first operating system detects the system operation status of the second operating system, where the first category of server components is the server components that allow the first operating system to collect component temperatures;

[0094] Step S204, when it is detected that the system operation status is the started state, the first operating system stops controlling the operation of the heat dissipation device; the second operating system controls the operation of the heat dissipation device according to the environmental temperature of the server and the component temperatures of the second category of server components, where the second category of server components is the server components that allow the second operating system to collect component temperatures.

[0095] Through the above steps, a dual operating system (i.e., the first operating system and the second operating system) is deployed on the first controller in the server. If the first controller restarts during the operation of the server, the first operating system can control the operation of the heat dissipation device in a short time, so that the operation of the heat dissipation device will not be interrupted for a long time, ensuring the heat dissipation effect of the server. Moreover, the first operating system can control the heat dissipation device according to the environmental temperature of the server and the component temperatures it can collect, making the operation control of the heat dissipation device fine and precise during the control stage of the first operating system. The first operating system can also detect the system operation status of the second operating system. If the second operating system has been started, the control right of the heat dissipation device will be handed over to the second operating system, and the second operating system will control the operation of the heat dissipation device according to the environmental temperature of the server and the component temperatures it can collect. Therefore, the technical problem of low timeliness in controlling the heat dissipation device can be solved, and the timeliness effect of controlling the heat dissipation device can be improved.

[0096] Optionally, in this embodiment, the above first operating system and second operating system can be, but are not limited to, two heterogeneous or homogeneous operating systems, that is, the types of the first operating system and the second operating system can be the same or different.

[0097] Taking the first operating system and the second operating system as heterogeneous operating systems as an example, the first operating system and the second operating system can be operating systems with different sensitivities to response time. For example, the first operating system is more sensitive to response time than the second operating system. Or, the first operating system and the second operating system can be operating systems with different resource occupancies. For example, the first operating system occupies less resources for the service than the second operating system. The first operating system is an operating system that starts faster than the second operating system.

[0098] The above-mentioned first operating system and second operating system can be, but are not limited to, two heterogeneous operating systems deployed on the controller of an embedded system, that is, an embedded operating system. Embedded operating systems can be classified into real-time operating systems (RTOS) and non-real-time operating systems according to the sensitivity to response time. Real-time operating systems can be, but are not limited to, Free RTOS (Free Real-Time Operating System) and RT Linux (RealTime Linux). Non-real-time operating systems can be, but are not limited to, contiki (Contiki Operating System), HeliOS (Helix Operating System), and Linux (Linux Operating System), etc.

[0099] The above-mentioned first controller can be, but are not limited to, the baseboard management controller (BMC) of a server. The second operating system is the main operating system of the BMC, which can be, but is not limited to, Linux. The first operating system is the auxiliary operating system of the BMC, which can be, but is not limited to, RTOS. The above-mentioned heat dissipation device can be, but is not limited to, a fan, a heat sink, a heat pipe, etc. The operation control of the first controller over the heat dissipation device can be, but is not limited to, the control of the operation parameters of the heat dissipation device by the first controller. For example, if the heat dissipation device is a fan, the first controller can control the duty cycle of the fan speed. If the heat dissipation device is a heat sink, a heat pipe, etc., the first controller can control operation parameters such as the flow rate or velocity of the condensing substance.

[0100] The server components of the first category are server components that allow the first operating system to collect component temperatures, such as: CPU, memory, PCH, network card, etc. The server components of the second category are server components that allow the second operating system to collect component temperatures, such as: CPU, memory, hard disk, network card, raid card, HBA card, HCA card, GPU, etc.

[0101] In this embodiment, a method for operating a heat dissipation device is provided. The server includes: server components, a first controller, and a heat dissipation device. The first controller deploys the first operating system and the second operating system. Among the server components that can be configured on the server, there are: server components of the first category and server components of the second category. The server components of the first category are server components that allow the first operating system to collect component temperatures, and the server components of the second category are server components that allow the second operating system to collect component temperatures. Figure 3It is a flowchart of an operation method for a heat dissipation device according to an embodiment of the present application. As Figure 3 shown, the process includes the following steps:

[0102] Step S302, during the operation of the server, the first controller is restarted;

[0103] Step S304, during the restart of the first controller, the heat dissipation device operates according to the operation parameters output by the first operating system, where the operation parameters output by the first operating system are determined according to the environmental temperature of the server and the component temperatures of the first category of server components;

[0104] Step S306, when the first controller completes the restart, the heat dissipation device operates according to the operation parameters output by the second operating system, where the operation parameters output by the second operating system are determined according to the environmental temperature of the server and the component temperatures of the second category of server components.

[0105] Through the above steps, a dual operating system (i.e., the first operating system and the second operating system) is deployed on the first controller in the server. If the first controller is restarted during the operation of the server, the heat dissipation device can operate according to the operation parameters output by the first operating system in a short time, so that the operation of the heat dissipation device will not be interrupted for a long time, ensuring the heat dissipation effect of the server. And the operation parameters output by the first operating system are determined according to the environmental temperature of the server and the component temperatures it can collect, making the operation control of the heat dissipation device in this stage also fine and precise. If the second operating system has been started, that is, the first controller has completed the restart, the heat dissipation device can operate according to the operation parameters output by the second operating system, and the operation parameters output by the second operating system are determined according to the environmental temperature of the server and the component temperatures it can collect, and the operation control of the heat dissipation device will be more fine and precise. Therefore, the technical problem of low timeliness in controlling the heat dissipation device can be solved, and the effect of improving the timeliness of controlling the heat dissipation device is achieved.

[0106] The restart of the first controller can be, but is not limited to, performed by a restart operation executed on the first controller, or can be, but is not limited to, a restart automatically triggered by the operation of the services on the first controller. For example, some faults on the first controller may be solved by the automatic restart of the first controller.

[0107] Optionally, in this embodiment, the startup duration of the first operating system is less than that of the second operating system. For example: The first operating system is an operating system that can be quickly started within 3 to 5 seconds.

[0108] The restart of the first controller can, but is not limited to, refer to the restart of the entire first controller, that is, both the first operating system and the second operating system are restarted. In this case, the first operating system can control the heat dissipation device immediately after a short startup duration. The heat dissipation effect of the heat dissipation device will only be interrupted for a very short time.

[0109] The restart of the first controller can also, but is not limited to, refer to the restart of the main operating system of the first controller, that is, the second operating system is restarted. In this case, the first operating system can control the heat dissipation device immediately. The heat dissipation effect of the heat dissipation device will not be interrupted.

[0110] In an optional example, a first control curve and a first control configuration are also configured on the first controller. The first control curve is used to indicate the first correspondence between the ambient temperature of the server and the operating parameters of the heat dissipation device. The first control configuration is used to indicate the parameter configuration in the operating parameter algorithm corresponding to the server components of the first category. The operating parameter algorithm is used to calculate the operating parameters of the heat dissipation device according to the component temperature of the server components. In the above step S202, the first operating system can control the operation of the heat dissipation device according to the ambient temperature of the server and the component temperature of the server components of the first category in the following ways, but is not limited to these ways: the first operating system calls the first control curve to convert the ambient temperature of the server into the first operating parameter, and the first operating system calls the first control configuration to calculate the second operating parameter according to the component temperature of the server components of the first category; the first operating system determines the target operating parameter according to the first operating parameter and the second operating parameter; the first operating system controls the operation of the heat dissipation device according to the target operating parameter.

[0111] Optionally, in this embodiment, each control curve can, but is not limited to, adopt any curve form that conforms to the server heat dissipation law. For example: the first correspondence can, but is not limited to, include a linear relationship, then the first control curve can, but is not limited to, be a straight line, or a piecewise straight line. The first correspondence can, but is not limited to, include an exponential relationship, then the first control curve can, but is not limited to, be an exponential curve. The operating parameter algorithm can, but is not limited to, adopt any function algorithm that conforms to the server heat dissipation law. For example: the Proportion Integration Differentiation (PID) control speed regulation algorithm.

[0112] Optionally, in this embodiment, determining the target operating parameter according to the first operating parameter and the second operating parameter can, but is not limited to, include determining the larger value of the first operating parameter and the second operating parameter as the target operating parameter.

[0113] Taking the first controller as the BMC, the first operating system as the RTOS real-time operating system, and the second operating system as Linux as an example, in order to make full use of the computing power resources of the BMC, it deploys the RTOS real-time operating system. According to the BMC Linux / RTOS heterogeneous dual-system parallel technology, data exchange between multiple cores and multiple systems is achieved. The ambient temperature and component temperature are collected through the PECI bus, and signal acquisition control at the millisecond level can be realized. Subsequently, dynamic control of temperature and heat dissipation is achieved through closed-loop feedback, so as to accurately control the heat dissipation resources to meet the minimum resource of heat dissipation requirements and reduce ineffective power consumption. For example: Figure 4 It is a schematic diagram of an RTOS system controlling the server heat dissipation system according to an embodiment of the present application, as Figure 4 shown, the RTOS system collects the component temperature from the temperature sensors deployed on various hardware components such as the CPU deployed on the server, and collects the ambient temperature from the temperature sensors deployed in the server's heat dissipation system. The fan speed of the fan deployed in the heat dissipation system is adjusted according to the component temperature and the ambient temperature. At the same time, the RTOS system also has the function of regulating the performance of various hardware components.

[0114] In an optional example, from the perspective of the operation of the heat dissipation device, the operation parameters output by the first operating system are the first operation parameters or the second operation parameters. Among them, the ambient temperature of the server and the first operation parameters conform to the first control curve, and the first control curve is used to indicate the first corresponding relationship between the ambient temperature of the server and the operation parameters of the heat dissipation device. The second operation parameter is calculated according to the component temperature of the first type of server components.

[0115] In this embodiment, the corresponding relationships indicated by each control curve can be designed according to, but not limited to, the server configuration, the model of the heat dissipation device used, and the operation performance of the heat dissipation device. For example: If the heat dissipation device is a fan, then the corresponding relationships indicated by each control curve can be designed according to, but not limited to, the server configuration, the model of the fan, and the PQ characteristic curve. The PQ characteristic curve is one of the key curves describing the performance of the heat dissipation fan. Here, P represents the static pressure (Pressure) generated by the fan, and Q represents the air volume (Airflow) of the fan. The PQ characteristic curve shows the static pressure value that the fan can provide at different air volumes. Specifically, when the air volume increases, the static pressure of the fan usually decreases; conversely, when the air volume decreases, the static pressure of the fan increases. This relationship is shown as a downward-sloping curve on the PQ curve.

[0116] For example: As shown in Table 1, the corresponding relationship between the ambient temperature and the duty cycle parameter of the fan speed is presented. This corresponding relationship can be, but is not limited to, the curve PWM = k * Inlet + b, where PWM is the duty cycle of the fan speed, Inlet is the ambient temperature, and in a certain corresponding relationship, multiple straight lines can be spliced simultaneously. Under different states and different configurations, the air volume requirements of the heat dissipation components are different, and different logics can be achieved by adjusting the magnitudes of k and b; for example, in the S5 state (i.e., the state where the server is powered on but not started), when there is an intelligent network card in the configuration, the air volume requirement is greater than that when there is an OCP network card, and the value of b can be increased correspondingly; when there is an intelligent network card in the configuration, during the power-on process of each component, the air volume requirement is greater than that in the S5 state, and the values of both k and b can be adjusted simultaneously. When the temperature is lower than the minimum temperature value defined by the Inlet curve, the corresponding PWM takes the PWM corresponding to the defined minimum temperature; when the temperature is higher than the minimum value defined by the Inlet curve, the corresponding PMW takes the PWM corresponding to the defined maximum temperature. When multiple curves or algorithms are called simultaneously in a certain state and a certain configuration, the maximum value of the fan speed output by parallel merging of multiple curves or algorithms can be given to the fan.

[0117] Table 1

[0118]

[0119] In an optional example, when the first operating system controls the operation of the heat dissipation device, if the second operating system is starting up and has the ability to collect temperature parameters during its startup process, then the parameters such as the ambient temperature and component temperature required by the first operating system can be, but are not limited to, collected by the second operating system and transmitted to the first operating system through inter-core communication. For example, the first operating system can, but is not limited to, call the first control curve to convert the ambient temperature of the server into the first operating parameter in the following manner, and the first operating system can call the first control configuration to calculate the second operating parameter based on the component temperature of the first category of server components: During the startup process of the second operating system, the second operating system collects the ambient temperature of the server and the component temperature of the first category of server components; the second operating system sends the collected ambient temperature of the server and the component temperature of the first category of server components to the first operating system; the first operating system calls the first control curve to convert the ambient temperature of the server into the first operating parameter, and the first operating system calls the first control configuration to calculate the second operating parameter based on the component temperature of the first category of server components.

[0120] In an optional example, the method of inter-core communication can be, but is not limited to, a method of combining shared memory with interrupt requests. For example: a target storage space is also deployed on the server, and the target storage space allows the first operating system and the second operating system to access it. The second operating system can send the collected server ambient temperature and the component temperature of the first category of server components to the first operating system in the following manners, but is not limited to: the second operating system writes the collected server ambient temperature and the component temperature of the first category of server components into the target storage space; the second operating system sends an interrupt request to the first operating system; the first operating system responds to the interrupt request and reads the server ambient temperature and the component temperature of the first category of server components from the target storage space.

[0121] For example, the first controller is BMC, the first operating system is RTOS real-time operating system, and the second operating system is Linux. Figure 5 is a schematic diagram of the architecture of a BMC according to an embodiment of the present application, such as Figure 5 As shown in the figure, RTOS and Linux are deployed in BMC, and the two communicate through shared memory combined with inter-core communication. BMC is connected to the flash memory (NOR Flash), interface controller (SD / EMMC), hard disk (DDR), out-of-band interface (MAC), central processing unit (CPU), fan (FAN), sensor (Sensor) and integrated south bridge (PCH) on the server.

[0122] Based on the real-time acquisition of sensor temperature data through the RTOS system, before the BMC Linux system is started, the RTOS collects and controls the system temperature; during the BMC Linux system startup phase, Linux collects sensor temperature information, and then passes it to the RTOS through shared memory, and the RTOS performs fan control; when the BMC Linux system is fully started, the Linux system collects sensor temperature and implements temperature control itself. At the same time, when the BMC Linux system fails or restarts, the RTOS can quickly take over the cooling system and accurately control the fan speed.

[0123] Among the three control firmwares of RTOS, BMC Linux and CPLD, BMC has the advantages of high control accuracy and high logic complexity, but BMC takes a long time to start up. When BMC has not started up, fan control cannot be achieved. Both RTOS and CPLD have the characteristics of fast startup, and can start up and control the fan in 3-5S. In addition, RTOS can also obtain the ambient temperature and the temperature of some important components. After obtaining the temperature, RTOS can perform refined control calculations and fan speed output according to these temperatures.

[0124] In an optional embodiment, taking the first controller as the BMC, the first operating system as the RTOS real-time operating system, and the second operating system as Linux as an example, after the server is powered on and boots up, during the running process, the BMC restarts. The RTOS starts first. If the RTOS starts successfully and successfully takes over the fan control right, then at this time, the RTOS conducts the dominant control of the fan. The RTOS differentiates and adjusts the speed according to the current state of the server, the configuration, and the component temperatures that can be obtained. When the RTOS recognizes according to the state that the server is in normal operation after booting up, it can call different fan speeds according to the obtained ambient temperature. Moreover, the RTOS can obtain the core temperatures of some components, such as CPU temperature, memory temperature, PCH temperature, and network card temperature. After obtaining the temperatures, it can conduct refined fan control according to PID. When the server is in the boot state, all components support normal use. However, compared with the BMC, the RTOS cannot obtain the core temperatures of all important components. Therefore, when the RTOS conducts fan control, in the linear relationship between the corresponding ambient temperature and the speed, the fan speed values corresponding to the same ambient temperature are higher than those when the BMC conducts control, so as to prevent the components for which the RTOS cannot obtain the temperatures from overheating during the server operation.

[0125] The RTOS conducts fan control and simultaneously checks the BMC status. After the BMC starts up, the heartbeat signal corresponding to the BMC starts and lasts for more than a certain period. When the RTOS recognizes that the BMC heartbeat is normal, the RTOS hands over the fan control right to the BMC. The BMC differentiates and adjusts the speed according to the current state of the server, the configuration, and the component temperatures that can be obtained. The BMC can obtain the ambient temperature and the component temperatures of most components in the server, can recognize the air-cooling and liquid-cooling information of the server, the status of each component, etc., and conducts detailed speed adjustment calculations according to the temperature and configuration information.

[0126] In an optional example, a second control curve and a second control configuration are also configured on the first controller. The second control curve is used to indicate a second corresponding relationship between the ambient temperature of the server and the operating parameters of the heat dissipation device. The second control configuration is used to indicate the parameter configuration in the operating parameter algorithm corresponding to the server components of the second category. The operating parameter algorithm is used to calculate the operating parameters of the heat dissipation device based on the component temperature of the server components. In the above step S204, the second operating system can control the operation of the heat dissipation device according to the ambient temperature of the server and the component temperature of the server components of the second category in, but not limited to, the following manner: The second operating system calls the second control curve to convert the ambient temperature of the server into a third operating parameter, and the second operating system calls the second control configuration to calculate a fourth operating parameter according to the component temperature of the server components of the second category. The second operating system determines a reference operating parameter according to the third operating parameter and the fourth operating parameter. The second operating system controls the operation of the heat dissipation device according to the reference operating parameter.

[0127] Optionally, in this embodiment, determining the reference operating parameter according to the third operating parameter and the fourth operating parameter may include, but is not limited to, determining the larger value of the third operating parameter and the fourth operating parameter as the reference operating parameter.

[0128] In an optional example, from the perspective of the operation of the heat dissipation device, the operating parameter output by the second operating system is the third operating parameter or the fourth operating parameter. Among them, the ambient temperature of the server and the third operating parameter conform to the second control curve, and the second control curve is used to indicate the second corresponding relationship between the ambient temperature of the server and the operating parameters of the heat dissipation device. The fourth operating parameter is calculated according to the component temperature of the server components of the second category.

[0129] Optionally, in this embodiment, the heat dissipation performance of the heat dissipation device corresponding to the operating parameter under the first corresponding relationship is higher than that of the heat dissipation device corresponding to the operating parameter under the second corresponding relationship for the same ambient temperature of the server.

[0130] The operating parameter output by the first operating system at the same ambient temperature is higher than that when the second operating system is in control, so as to prevent the server components for which the first operating system cannot obtain the temperature from overheating during the operation of the server.

[0131] In an alternative example, the second control curve can be adjusted according to, but not limited to, the heat dissipation type adopted by the server. For example, the heat dissipation types of the server include multiple heat dissipation types, and the multiple heat dissipation types include the target heat dissipation type to which the heat dissipation device belongs. The second control curve includes a first curve and a second curve. The heat dissipation performance of the heat dissipation device corresponding to the operating parameters of the same server's ambient temperature under the first curve is higher than that of the heat dissipation device corresponding to the operating parameters under the second curve. The second operating system can convert the server's ambient temperature into a third operating parameter by, but not limited to, the following method: The second operating system detects the heat dissipation type adopted by the server among the multiple heat dissipation types. When the heat dissipation type adopted by the server is the target heat dissipation type, the second operating system calls the first curve to convert the server's ambient temperature into a third operating parameter. When the heat dissipation type adopted by the server is not the target heat dissipation type, the second operating system calls the second curve to convert the server's ambient temperature into a third operating parameter.

[0132] The multiple heat dissipation types can include, but are not limited to, air cooling and liquid cooling, etc. The heat dissipation type of the heat dissipation device belongs to the target heat dissipation type among the multiple heat dissipation types. If the heat dissipation type adopted by the server is the target heat dissipation type, the operating parameters output to the heat dissipation device are determined by using the control curve with higher heat dissipation performance. If the heat dissipation type adopted by the server is not the target heat dissipation type, the operating parameters output to the heat dissipation device are determined by using the control curve with lower heat dissipation performance.

[0133] For example, if the heat dissipation device is a fan, then the target heat dissipation type is air cooling. If the server is using air cooling for heat dissipation, the rotation speed output to the fan is determined by using the control curve with higher heat dissipation performance. If the server is using liquid cooling for heat dissipation, the rotation speed output to the fan is determined by using the control curve with lower heat dissipation performance.

[0134] In an alternative example, from the perspective of the operation of the heat dissipation device, the heat dissipation types of the server include multiple heat dissipation types, and the multiple heat dissipation types include the target heat dissipation type to which the heat dissipation device belongs. The second control curve includes a first curve and a second curve. The heat dissipation performance of the heat dissipation device corresponding to the operating parameters of the same server's ambient temperature under the first curve is higher than that of the heat dissipation device corresponding to the operating parameters under the second curve. When the heat dissipation type adopted by the server is the target heat dissipation type, the server's ambient temperature and the third operating parameter conform to the first curve. When the heat dissipation type adopted by the server is not the target heat dissipation type, the server's ambient temperature and the third operating parameter conform to the second curve.

[0135] In an alternative example, the second operating system can also perform differentiated and refined control of the cooling device by identifying the fault information of components according to the degree of influence of component temperature on the server, so that the operation of the cooling device better conforms to the operation of the server. For example: The server components of the second category are divided into a set of server components with multiple risk levels. The second control curve includes: multiple parameter curves corresponding one-to-one to multiple risk levels. The higher the risk level, the greater the degree of influence of component temperature on the server. For server components with a higher risk level, the cooling performance of the cooling device corresponding to the operating parameters on the corresponding control curve is higher under the same ambient temperature of the server. The second operating system can, but is not limited to, call the second control curve to convert the ambient temperature of the server into the third operating parameter in the following manner: The second operating system detects whether there are faulty components among the server components of the second category; when it detects that there are faulty components among the server components of the second category, the second operating system detects the target set of server components to which the faulty components belong in the set of server components with multiple risk levels; the second operating system calls the reference curve corresponding to the target risk level of the target set of server components from the multiple parameter curves to convert the ambient temperature of the server into the third operating parameter.

[0136] The server components of the second category are divided into multiple risk levels to obtain a set of server components for each risk level. The multiple risk levels correspond one-to-one to multiple parameter curves. The risk level indicates the degree of influence of component temperature on the server. The higher the risk level, the greater the degree of influence of component temperature on the server. For example: The risk level of the CPU is higher than that of the hard disk.

[0137] If the faulty components are multiple types of server components, the resulting third operating parameter may be multiple parameters. In this case, the maximum value can be extracted from the multiple parameters as the final third operating parameter.

[0138] In an alternative example, from the perspective of the operation of the cooling device, the server components of the second category are divided into a set of server components with multiple risk levels. The second control curve includes: multiple parameter curves corresponding one-to-one to multiple risk levels. The higher the risk level, the greater the degree of influence of component temperature on the server. For server components with a higher risk level, the cooling performance of the cooling device corresponding to the operating parameters on the corresponding control curve is higher under the same ambient temperature of the server. When there are faulty components among the server components of the second category, the ambient temperature of the server and the third operating parameter conform to the reference curve among the multiple parameter curves. Among them, the faulty components belong to the target set of server components in the set of server components with multiple risk levels, and the target risk level of the target set of server components corresponds to the reference curve.

[0139] In an alternative example, the set of server components with multiple risk levels may include: a set of high heat dissipation risk components and a set of low heat dissipation risk components. The impact of the component temperature of the server components in the high heat dissipation risk component set on the server is higher than that of the server components in the low heat dissipation risk component set. The multiple parameter curves include a third curve and a fourth curve. The heat dissipation performance of the heat dissipation device corresponding to the operating parameters of the same server's ambient temperature under the fourth curve is higher than that under the operating parameters of the third curve. The second operating system can detect the faulty component in the target server component set to which the set of server components with multiple risk levels belongs in the following ways, but not limited to: The second operating system detects that the faulty component is a high heat dissipation risk component or a low heat dissipation risk component. The second operating system can convert the server's ambient temperature into the third operating parameter by calling the reference curve corresponding to the target risk level of the target server component set from the multiple parameter curves in the following ways, but not limited to: When it is detected that the faulty component is a high heat dissipation risk component, the second operating system calls the fourth curve to convert the server's ambient temperature into the third operating parameter; when it is detected that the faulty component is a low heat dissipation risk component, the second operating system calls the third curve to convert the server's ambient temperature into the third operating parameter.

[0140] Optionally, the risk levels of server components can be divided into two levels, high heat dissipation risk and low heat dissipation risk, but not limited to this. If the risk level of the faulty server component is high heat dissipation risk, the fourth curve with higher heat dissipation performance can be called to determine the third operating parameter, but not limited to this. If the risk level of the faulty server component is low heat dissipation risk, the third curve with lower heat dissipation performance can be called to determine the third operating parameter, but not limited to this.

[0141] In an alternative example, from the perspective of the operation of the heat dissipation device, the set of server components with multiple risk levels includes: a set of high heat dissipation risk components and a set of low heat dissipation risk components. The impact of the component temperature of the server components in the high heat dissipation risk component set on the server is higher than that of the server components in the low heat dissipation risk component set. The multiple parameter curves include a third curve and a fourth curve. The heat dissipation performance of the heat dissipation device corresponding to the operating parameters of the same server's ambient temperature under the fourth curve is higher than that under the operating parameters of the third curve. When the faulty component is a low heat dissipation risk component, the server's ambient temperature and the third operating parameter conform to the third curve. When the faulty component is a high heat dissipation risk component, the server's ambient temperature and the third operating parameter conform to the fourth curve.

[0142] Optionally, in this embodiment, the heat dissipation performance of the heat dissipation device corresponding to the operating parameters of the same server at the third curve is higher than that of the heat dissipation device corresponding to the operating parameters at the first curve.

[0143] In an optional example, the heat dissipation types of the server include multiple heat dissipation types. The multiple heat dissipation types include the target heat dissipation type to which the heat dissipation device belongs. The second category of server components includes heat dissipation high-risk components and heat dissipation low-risk components. The influence degree of the component temperature of the heat dissipation high-risk components on the server is higher than that of the component temperature of the heat dissipation low-risk components on the server. The second control curve includes the fifth curve, the sixth curve, the seventh curve, and the eighth curve. The heat dissipation performance of the heat dissipation device corresponding to the operating parameters of the same server at the fifth curve is higher than that of the heat dissipation device corresponding to the operating parameters at the sixth curve. The heat dissipation performance of the heat dissipation device corresponding to the operating parameters of the same server at the eighth curve is higher than that of the heat dissipation device corresponding to the operating parameters at the seventh curve. The heat dissipation performance of the heat dissipation device corresponding to the operating parameters of the same server at the seventh curve is higher than that of the heat dissipation device corresponding to the operating parameters at the fifth curve. Optionally, but not limited to, the second operating system can call the second control curve to convert the ambient temperature of the server into the third operating parameter in the following manner: The second operating system detects whether there are faulty components in the second category of server components; when it is detected that there are faulty components in the second category of server components, the second operating system detects whether the faulty component is a heat dissipation high-risk component or a heat dissipation low-risk component; when it is detected that the faulty component is a heat dissipation high-risk component, the second operating system calls the eighth curve to convert the ambient temperature of the server into the third operating parameter; when it is detected that the faulty component is a heat dissipation low-risk component, the second operating system calls the seventh curve to convert the ambient temperature of the server into the third operating parameter; when it is detected that there are no faulty components in the second category of server components, the second operating system detects the heat dissipation type adopted by the server among the multiple heat dissipation types; when the heat dissipation type adopted by the server is the target heat dissipation type, the second operating system calls the fifth curve to convert the ambient temperature of the server into the third operating parameter; when the heat dissipation type adopted by the server is not the target heat dissipation type, the second operating system calls the sixth curve to convert the ambient temperature of the server into the third operating parameter.

[0144] Optionally, in this embodiment, the heat dissipation device can be controlled by combining the risk level of the server components and the heat dissipation type of the server, but not limited to this. The second operating system first identifies the faulty components in the server components. If there are faulty components, the heat dissipation device is controlled according to the risk level of the faulty components. If there are no faulty components, the heat dissipation device is controlled according to the heat dissipation type of the server.

[0145] In an optional example, from the perspective of the operation of the heat dissipation device, the heat dissipation types of the server include: multiple heat dissipation types, and the multiple heat dissipation types include the target heat dissipation type to which the heat dissipation device belongs. The server components of the second category include: high-risk heat dissipation components and low-risk heat dissipation components. The influence degree of the component temperature of the high-risk heat dissipation components on the server is higher than that of the component temperature of the low-risk heat dissipation components on the server. The second control curve includes: the fifth curve, the sixth curve, the seventh curve, and the eighth curve. The heat dissipation performance of the heat dissipation device corresponding to the operating parameters of the same server's ambient temperature under the fifth curve is higher than that of the heat dissipation device corresponding to the operating parameters under the sixth curve. The heat dissipation performance of the heat dissipation device corresponding to the operating parameters of the same server's ambient temperature under the eighth curve is higher than that of the heat dissipation device corresponding to the operating parameters under the seventh curve. The heat dissipation performance of the heat dissipation device corresponding to the operating parameters of the same server's ambient temperature under the seventh curve is higher than that of the heat dissipation device corresponding to the operating parameters under the fifth curve. When there is a faulty component in the server components of the second category and the faulty component is a high-risk heat dissipation component, the ambient temperature of the server and the third operating parameter conform to the eighth curve. When there is a faulty component in the server components of the second category and the faulty component is a low-risk heat dissipation component, the ambient temperature of the server and the third operating parameter conform to the seventh curve. When there is no faulty component in the server components of the second category and the heat dissipation type adopted by the server is the target heat dissipation type, the ambient temperature of the server and the third operating parameter conform to the fifth curve. When there is no faulty component in the server components of the second category and the heat dissipation type adopted by the server is not the target heat dissipation type, the ambient temperature of the server and the third operating parameter conform to the sixth curve.

[0146] In an optional example, the server components of the second category include: multiple server components, and the second control curve includes multiple component curves corresponding one by one to the multiple server components. The higher the influence degree of the component temperature on the server, the higher the heat dissipation performance of the heat dissipation device corresponding to the operating parameters of the corresponding control curve at the same ambient temperature of the server. The second operating system can, but is not limited to, call the second control curve in the following way to convert the ambient temperature of the server into the third operating parameter: The second operating system detects whether there is a faulty component in the server components of the second category. When it is detected that there is a faulty component in the server components of the second category, the second operating system extracts the target server component belonging to the faulty component from the multiple server components. The second operating system calls the target curve corresponding to the target server component from the multiple component curves to convert the ambient temperature of the server into the third operating parameter.

[0147] Optionally, in this embodiment, different server components may correspond to different component curves, and the second operating system may, but is not limited to, determine the third operating parameter according to the component curve corresponding to the failed component. If there are multiple third operating parameters, the maximum value is extracted therefrom as the final third operating parameter.

[0148] In an optional example, from the perspective of the operation of the heat dissipation device, the second type of server components includes: multiple server components, the second control curve includes multiple component curves corresponding to the multiple server components one by one, and the higher the degree of influence of the component temperature on the server, the higher the heat dissipation performance of the heat dissipation device corresponding to the operating parameter on the corresponding control curve at the same ambient temperature of the server; in the case where a failed component is detected in the second type of server components, the ambient temperature of the server and the third operating parameter conform to the target curve, where the target server component belonging to the failed component among the multiple server components corresponds to the target curve among the multiple component curves.

[0149] In an optional implementation manner, taking the first controller as the BMC, the first operating system as the RTOS, the second operating system as Linux (representing the BMC), and the heat dissipation device as the fan as an example, when the BMC recognizes that the server is in the power-on completed state, it further discriminates between air cooling and liquid cooling configurations. By judging whether there is a corresponding liquid cooling in-place signal, it determines whether the server is in an air cooling configuration or a liquid cooling configuration, and calls different ambient temperature curves respectively. And compared with the RTOS, the BMC supports obtaining the component temperatures of more server components (such as CPU temperature, memory temperature, hard disk temperature, network card temperature, raid card temperature, HBA card temperature, HCA card temperature, GPU temperature, etc.). After obtaining the component temperatures, fine-grained fan control can be performed according to the PID. The BMC can also identify the status of each component, classify the component temperatures in the server into high heat dissipation risk components and low heat dissipation risk components. When a server component fails, resulting in the component being in place but the BMC cannot obtain the component temperature, the BMC can call different Inlet curves according to the heat dissipation risk level to which the component belongs to ensure that the component does not overheat.

[0150] For example: When the BMC recognizes that the server is in the state of booting completed based on the status, it calls different control curves as shown in Table 2 according to the obtained Inlet temperature (ambient temperature) and the configured status information, and performs PID and linear calculation speed regulation based on the obtained component temperatures of each server component: When the BMC detects that there is no in-position signal corresponding to the liquid cooling configuration, it determines that it is an air cooling configuration at this time and calls the corresponding Inlet-1; when the BMC detects that there is an in-position signal corresponding to the liquid cooling configuration, it determines that it is a liquid cooling configuration at this time and calls the corresponding Inlet-2; when the BMC detects that there is a low-risk component in the server but the temperature cannot be obtained, it calls the corresponding abnormal speed regulation Inlet-3; when the BMC detects that there is a high-risk component in the server but the temperature cannot be obtained, it calls the corresponding abnormal speed regulation Inlet-4. According to different air volume requirements, the corresponding fan speeds from low to high at the same ambient temperature for the Inlet curves are: Inlet-2, Inlet-1, Inlet-3, Inlet-4. After the BMC obtains the CPU temperature, memory temperature, hard disk temperature, network card temperature, raid card temperature, HBA card temperature, HCA card temperature, and GPU temperature, it performs PID and linear calculation speed regulation.

[0151] Table 2

[0152]

[0153] In an optional example, the operating parameter algorithm is used to calculate the operating parameters of the heat dissipation device according to the component temperatures of the server components at the sampling moment, the component temperatures at the historical moment, and the operating parameters corresponding to the historical sampling moment; it can but is not limited to the following method for the first operating system to call the first control configuration to calculate the second operating parameters according to the component temperatures of the first category of server components: The first operating system calls the first control configuration and substitutes it into the operating parameter algorithm to obtain the target operating parameter algorithm; the first operating system obtains the first component temperature of the first category of server components at the sampling moment, the second component temperature at the historical moment, and the historical operating parameters corresponding to the historical sampling moment, where the historical operating parameters are the operating parameters corresponding to the first category of server components calculated by the operating parameter algorithm at the historical sampling moment; the first operating system substitutes the first component temperature, the second component temperature, and the historical operating parameters into the target operating parameter algorithm to obtain the second operating parameters.

[0154] Optionally, in this embodiment, the operating parameter algorithm may, but is not limited to, be PWM(k) = PWM(k - 1) + ▽PWM; where ▽PWM = Kp * [T(k) - T(k - 1)] + Ki(T(k) - Tsp) + Kd * [[T(k) - T(k - 1)] - [T(k - 1) - T(k - 2)]], PWM(k) is the operating parameter corresponding to the current sampling moment, PWM(k - 1) is the operating parameter corresponding to the historical sampling moment, T(k) is the component temperature at the current sampling moment, the component temperatures at historical moments include T(k - 1) and T(k - 2), T(k - 1) is the component temperature at the previous sampling moment before the current sampling moment, T(k - 2) is the component temperature at the previous sampling moment before the previous sampling moment, the parameter configuration in the operating parameter algorithm includes: Kp, Ki, Kd, and Tsp, and Tsp is the temperature value that the component temperature of the server component is allowed to reach at most during operation.

[0155] For the component temperature of the server component in the server, PID speed regulation may, but is not limited to, be adopted. In the PID calculation, the formula is PWM(k) = PWM(k - 1) + ▽PWM; ▽PWM = Kp * [T(k) - T(k - 1)] + Ki(T(k) - Tsp) + Kd * [[T(k) - T(k - 1)] - [T(k - 1) - T(k - 2)]], PWM(k) is the PWM at the kth moment, PWM(k - 1) is the virtual PWM calculated by this temperature sensor at the (k - 1)th moment (i.e., the historical operating parameter corresponding to the above historical sampling moment), Tsp is the set regulation point, that is, during operation, the temperature value that the component temperature is expected to reach at most, and T(k), T(k - 1), and T(k - 2) are the temperature values of the temperature sensor at the kth, (k - 1)th, and (k - 2)th moments respectively.

[0156] Among them, Kp * [T(k) - T(k - 1)] is the differential term, which represents the difference between the temperature at the kth moment and the temperature at the (k - 1)th moment. When the component temperature changes, this term can be used to respond quickly: when the component temperature rises, this term is positive, which can increase the corresponding fan speed; when the component temperature drops, this term is negative, which can decrease the corresponding fan speed; for components such as CPU and GPU, when the power consumption changes greatly, the driven temperature will also change greatly. Generally, a relatively large value (usually 6 - 12) is set for Kp in this term to increase the fan speed change brought by this term and avoid temperature overshoot.

[0157] Ki(T(k) - Tsp) is the integral term, which represents the difference between the temperature at the kth moment and the regulation point temperature: when the component temperature is higher than the regulation point, this term is positive, which can increase the corresponding fan speed; when the component temperature is lower than the regulation point, this term is negative, which can decrease the corresponding fan speed; the regulation point in this term determines the stable value of the temperature during the regulation process.

[0158] Kd * [[T(k) - T(k - 1)] - [T(k - 1) - T(k - 2)]] is the differential term, representing the temperature difference at time k and the temperature difference at time k - 1, which is the difference between these two.

[0159] In an optional embodiment, taking the first controller as the BMC, the first operating system as RTOS, the second operating system as Linux (representing the BMC), and the cooling device as a fan as an example, when the server is running normally, the BMC can control the fan in the following ways but is not limited to: Each control curve is shown in Table 3. When the server is in an air-cooled configuration and all components are in an untested state, the fan speed is relatively low. At this time, the BMC calls the Inlet-1 curve to increase the ambient temperature of the server, and the fan speed of the server also increases accordingly. When the ambient temperature rises to a certain value (such as 40°C), the fan speed no longer increases; when the ambient temperature of the server is decreased, the fan speed of the server decreases accordingly. When the ambient temperature drops to a certain temperature value (such as 20°C), the fan speed no longer decreases but maintains a stable speed. Moreover, in the same configuration, when the server is in a liquid-cooled configuration, the BMC calls the Inlet-2 curve to compare the fan speeds corresponding to different ambient temperatures when all components are air-cooled, and the fan speeds at different ambient temperatures when some components are liquid-cooled (for example, the CPU is liquid-cooled, and there is a corresponding liquid leakage detection line wound around its liquid-cooled plate, and the terminals of the liquid leakage detection line are installed on the motherboard). The fan speed corresponding to the same ambient temperature should be greater than or equal to the speed when there are liquid-cooled components in the configuration. When stress testing the components in the server, it is found that as the temperature of the components increases, the fan speed increases accordingly, and finally the temperature and speed reach a stable value. The BMC can perform refined PID speed regulation based on the obtained core temperature of the components. It can be found that the stable value of the component temperature is less than the spec value of the component temperature. When the ambient temperature of the server is increased, the component temperature and the corresponding fan speed increase and tend to be stable, but within a certain range of ambient temperatures, the stable temperature value of the components is the same; when the ambient temperature is very low and the corresponding fan speed no longer decreases, the stable value of the component temperature after stress testing will be lower than the stable value mentioned above; when the ambient temperature is very high and the corresponding stable fan speed has reached the maximum speed and no longer continues to increase, the stable value of the component temperature after stress testing will be higher than the temperature stable value mentioned above. Such components include CPU, memory, VR, hard disk, PCIe network card, OCP network card, raid card, GPU, etc. The parameter settings of some components are shown in Table 4.Log in to the WEB BMC of the server, or view the sdr information under the OS. When checking and finding that when installing a certain component in the configuration, under normal circumstances, the temperature of this component can be displayed. When an abnormality occurs and the temperature value of this component is displayed as an abnormal value, at this time, call Inlet-3 or Inlet-4. It is found that even if no stress test is performed on this component, the fan speed is higher than that when no stress test is carried out, and when the ambient temperature of the server is changed, the fan speed will also increase correspondingly; and when temperature abnormalities occur in different types of components, the fan speeds at the same ambient temperature are inconsistent. When temperature abnormalities occur in components with low power consumption and low heat dissipation risk, the corresponding speeds are relatively low (corresponding to the Inlet-3 curve), and when temperature abnormalities occur in components with high power consumption and high heat dissipation risk, the corresponding speeds are relatively high (corresponding to the Inlet-4 curve).

[0160] Table 3

[0161]

[0162] Table 4

[0163]

[0164] Use the command to restart the BMC. Within a certain period of time after the restart, the fan speed remains at the speed before the BMC restart. At this time, the CPLD maintains the previous speed for output. After a certain period of time (such as 5S), the RTOS starts up prior to the BMC and controls the fan. At this time, the RTOS cannot obtain the core temperatures of all components. Therefore, it calls the Inlet-4 curve with a higher speed to increase the ambient temperature of the server, and the server fan speed also increases accordingly. When the ambient temperature rises to a certain value, the fan speed no longer increases; when the ambient temperature of the server is decreased, the server fan speed decreases accordingly. When the ambient temperature drops to a certain temperature value, the fan speed no longer decreases but remains at a stable speed. At the same ambient temperature, the speed after restarting the BMC for a certain period of time is the same as the speed corresponding to the temperature anomaly of the components with high power consumption and high heat dissipation risk mentioned above; during the BMC restart process, the RTOS can obtain the CPU temperature and perform PID speed regulation. If the CPU is stressed, it is found that as the CPU temperature rises, the fan speed will further increase, and finally the CPU temperature and speed reach a stable state. The stable value of the CPU temperature is less than the spec value of the CPU temperature. When the ambient temperature of the server is increased, the CPU temperature and the corresponding fan speed increase and tend to be stable. However, within a certain range of ambient temperatures, the stable temperature value of the CPU is the same; when the ambient temperature is very low, the corresponding fan speed no longer decreases. At this time, the stable value of the CPU temperature after stress testing will be lower than the stable value mentioned above; when the ambient temperature is very high, after the corresponding stable fan speed has reached the maximum speed and no longer increases, the stable value of the component temperature after stress testing at this time will be higher than the temperature stable value mentioned above. During the BMC restart process, the RTOS cannot obtain the GPU temperature. If the GPU is stressed, it is found that as the GPU temperature rises, the fan speed does not change. However, by obtaining the GPU temperature through tools under the OS, it can be found that the temperature value of the GPU is lower than the temperature spec. When the BMC is restarting, the RTOS controls the fan, can obtain the ambient temperature for fan control, and can also obtain the component temperatures of some components and perform refined control, but it cannot obtain the component temperatures of all components. However, through the fan speed corresponding to the ambient temperature, it can ensure that the temperatures of all components do not exceed the limit. After restarting the BMC, start stress testing the GPU. After a relatively long period of time, it is found that the fan speed will further increase, and the GPU temperature and fan speed tend to a stable value again. At this time, the BMC starts up and takes over the fan control right to start fan control.

[0165] Remove the BMC chip. It is found that after the boot is completed, after a certain period of time (such as 10S), the CPLD determines that the BMC fails to start. The CPLD takes over the control right of the fan and controls the fan according to the Inlet-4 curve. Raise the ambient temperature of the server, and the fan speed of the server will also increase. When the ambient temperature rises to a certain value, the fan speed will no longer increase; lower the ambient temperature of the server, and the fan speed of the server will decrease accordingly. When the ambient temperature drops to a certain temperature value, the fan speed will no longer decrease but maintain a stable speed. At the same ambient temperature, the fan speed after removing the BMC chip and powering on is the same as the speed corresponding to the abnormal temperature of the components with high power consumption and high heat dissipation risk described above; at this time, when using tools to stress test each component under the OS, it will be found that the temperature of each component rises, and the corresponding fan speed is not affected, but the temperature of each component can be kept under the temperature spec.

[0166] In an optional example, the first operating system can detect the system running state of the second operating system in the following ways, but not limited to: the first operating system detects whether the heartbeat of the second operating system is normal; in the case of detecting the heartbeat of the second operating system, the first operating system detects the duration of the heartbeat of the second operating system; in the case of detecting that the duration of the heartbeat of the second operating system is equal to or exceeds the target duration, the first operating system determines that the detected system running state is the started state.

[0167] Optionally, in this embodiment, the first operating system can determine the system running state of the second operating system by, but not limited to, the heartbeat signal of the second operating system and the duration of the heartbeat signal.

[0168] In an optional example, a second controller is also deployed on the server. In the case of a restart of the first controller, the second controller controls the operation of the cooling device and monitors the startup result of the first controller; in the case of the startup result indicating that the first controller fails to start, the second controller identifies the running state of the server; in the case of identifying the running state as the running state, the second controller continues to control the operation of the cooling device.

[0169] In an optional example, after the second controller monitors the startup result of the first controller, in the case of the startup result indicating that the first controller starts successfully, the second controller stops controlling the operation of the cooling device, and the first controller controls the operation of the cooling device.

[0170] Optionally, in this embodiment, the second controller may, but is not limited to, include a Complex Programmable Logic Device (CPLD for short). When the first controller restarts, the second controller can immediately take over the control of the heat dissipation device and transfer the control right of the heat dissipation device according to the startup situation of the first controller. For example, if the first operating system has been successfully started, the second controller can determine the startup result of the first controller to indicate that the first controller has started successfully, and the second controller can hand over the control right of the heat dissipation device to the first operating system of the first controller.

[0171] In an optional example, from the perspective of the operation of the heat dissipation device, a second controller is also deployed on the server. When the first controller restarts, the heat dissipation device operates according to the operation parameters output by the second controller; when the first controller fails to start, the heat dissipation device continues to operate according to the operation parameters output by the second controller; when the first controller starts successfully, the heat dissipation device stops operating according to the operation parameters output by the second controller and operates according to the operation parameters output by the first controller.

[0172] In an optional example, the first controller is connected to the heat dissipation device through the second controller; the operation of the heat dissipation device can be controlled by the second controller and the startup result of the first controller can be monitored by the second controller in the following ways, but is not limited to: the second controller checks whether the heartbeat of the first operating system is normal; when it is detected that the heartbeat of the first operating system is abnormal, the second controller controls the operation of the heat dissipation device and checks whether the heartbeat of the second operating system is normal; when it is detected that the heartbeat of the second operating system is abnormal, it is determined that the first controller fails to start; when it is detected that the heartbeat of the first operating system is normal, it is determined that the first controller starts successfully.

[0173] In an optional example, when the startup result indicates that the first controller fails to start, the second controller can identify the running state of the server in the following ways, but is not limited to: when the startup result indicates that the first controller fails to start, the second controller discards the operation parameters sent by the first controller received and identifies the running state of the server; when the startup result indicates that the first controller starts successfully, the second controller can stop controlling the operation of the heat dissipation device and the first controller controls the operation of the heat dissipation device in the following ways, but is not limited to: when the startup result indicates that the first controller starts successfully, the second controller sends the operation parameters sent by the first controller received to the heat dissipation device.

[0174] Optionally, in this embodiment, the second controller may be deployed between the first controller and the heat dissipation device. The second controller determines whether to transmit the operating parameters provided by the first controller to the heat dissipation device according to the startup result and operating conditions of the first controller, so as to ensure that the operating parameters generated during the failure process of the first controller are not transmitted to the heat dissipation device.

[0175] In an alternative embodiment, taking the first controller as the BMC, the first operating system as RTOS, the second operating system as Linux (representing the BMC), the second controller as the CPLD, and the heat dissipation device as the fan as an example, after the server is powered on and starts up normally, if the BMC restarts, the CPLD immediately takes over the control of the heat dissipation device. The RTOS starts first, and the CPLD checks whether the heartbeat of the RTOS is normal. If the heartbeat of the RTOS is abnormal for a continuous period of time, it proves that the RTOS does not have the ability to control the fan. If the CPLD determines that the RTOS heartbeat is abnormal for a continuous period of time, the CPLD continues to take over the fan control right. If the CPLD successfully takes over the fan control right, the CPLD then conducts the main control of the fan. The CPLD can obtain the ambient temperature, so the CPLD will call the corresponding Inlet curve according to the ambient temperature to ensure that each component in the server does not overheat.

[0176] The CPLD controls the fan and checks the BMC status synchronously. When the BMC has finished starting, the heartbeat signal corresponding to the BMC starts and lasts for more than a certain period. The CPLD determines that the BMC has completed starting. When the CPLD recognizes that the BMC heartbeat is normal, the CPLD transfers the fan control right to the BMC. The BMC performs differentiated speed regulation according to the current state of the server, the configuration, and the component temperatures that can be obtained. The BMC can obtain the ambient temperature and the core temperatures of most important components in the server, can identify information such as air cooling and liquid cooling of the server, and the status of each component, and performs detailed speed regulation calculations according to the temperature and configuration information. When the BMC recognizes that the server is in the state of having completed booting, it further discriminates the air cooling and liquid cooling configurations. By judging whether there is a corresponding liquid cooling in-place signal, it determines whether the server is in an air cooling configuration or a liquid cooling configuration, and calls different ambient temperature curves respectively. Compared with the RTOS, the BMC supports obtaining the core temperatures of more components (such as CPU temperature, memory temperature, hard disk temperature, network card temperature, raid card temperature, HBA card temperature, HCA card temperature, GPU temperature, etc.). After obtaining the temperature, it can perform refined fan control according to PID. The BMC can also identify the status of each component, classify the component temperatures in the server into high-risk components for heat dissipation and low-risk components for heat dissipation. When a component fails, resulting in the component being in place but the BMC cannot obtain the component temperature, the BMC can call different Inlet curves according to the heat dissipation risk level to which the component belongs to ensure that the component does not overheat.

[0177] The CPLD controls the fan and checks the BMC status synchronously. If the BMC still has no heartbeat, it further obtains the heartbeat status of the RTOS. If the heartbeat of the RTOS is normal and lasts for a certain period, it can be determined that the RTOS has started successfully, and the CPLD transfers the fan control right to the RTOS. If the RTOS starts successfully and takes over the fan control right, the RTOS then performs the main control of the fan. The RTOS performs differentiated speed regulation according to the current state of the server, the configuration, and the component temperatures that can be obtained. When the RTOS recognizes that the server is in the state of having completed booting and is running normally according to the status, it can call different fan speeds according to the obtained ambient temperature. Also, the RTOS can obtain the core temperatures of some components, such as CPU temperature, memory temperature, PCH temperature, network card temperature. After obtaining the temperature, it can perform refined fan control according to PID. When the server is in the booting state, each component supports normal use. However, compared with the BMC, the RTOS cannot obtain the core temperatures of all important components. Therefore, when the RTOS controls the fan, in the linear relationship between the corresponding ambient temperature and the speed, the fan speed value corresponding to the same ambient temperature is higher than that when the BMC controls, to prevent the components for which the RTOS cannot obtain the temperature from overheating during the operation of the server.

[0178] After the RTOS controls the fan, it synchronously checks the status of the BMC. After the BMC starts up successfully, the heartbeat signal corresponding to the BMC starts and lasts for more than a certain period. When the RTOS recognizes that the BMC heartbeat is normal, the RTOS hands over the fan control right to the BMC, and the BMC performs differentiated speed regulation according to the current state of the server, the configuration, and the component temperatures that can be obtained. The BMC can obtain the ambient temperature and the core temperatures of most important components in the server, can identify the air-cooling and liquid-cooling information of the server, the status of each component, etc., and performs detailed speed regulation calculations based on the temperature and configuration information.

[0179] When the BMC recognizes that the server is in the state of completing startup, it discriminates the air-cooling and liquid-cooling configurations. By judging whether there is a corresponding liquid-cooling in-place signal, it determines whether the server is in an air-cooling configuration or a liquid-cooling configuration, and calls different ambient temperature curves respectively. And compared with the RTOS, the BMC supports obtaining the core temperatures of more components (such as CPU temperature, memory temperature, hard disk temperature, network card temperature, raid card temperature, HBA card temperature, HCA card temperature, GPU temperature, etc.). After obtaining the temperature, it can perform refined fan control according to PID. Further, the BMC can identify the status of each component, classify the component temperatures in the server into components with high heat dissipation risk and components with low heat dissipation risk. When a component fails, resulting in the component being in place but the BMC cannot obtain the component temperature, the BMC can call different Inlet curves according to the heat dissipation risk level to which the component belongs to ensure that the component does not overheat.

[0180] The CPLD controls the fan. When it checks that the BMC still has no heartbeat and the RTOS has no heartbeat, the CPLD continues to maintain the fan control.

[0181] In summary, when the BMC controls the fan, it can obtain more component temperatures, the control logic is more complex, and the fan regulation is more refined; the RTOS can start up and control the fan in a short time, and can obtain the ambient temperature and the temperatures of some core components, and can achieve partial refined speed regulation; the CPLD can immediately control the fan when the BMC restarts. Through the coordinated action of the RTOS, CPLD, and BMC, it can not only ensure the stable operation of the server, but also ensure the most refined fan control of the server in different situations, ensure that each component does not overheat and operates with high performance, and can also ensure the reduction of system noise and fan power consumption.

[0182] In an alternative embodiment, taking the first controller as the BMC, the first operating system as RTOS, the second operating system as Linux (representing the BMC), the second controller as CPLD, and the heat dissipation device as a fan as an example, when the BMC restarts during the operation of the server, the control and operation of the heat dissipation device are as follows: Figure 6 It is a schematic diagram of the control and operation of a heat dissipation device according to an embodiment of the present application, as Figure 6 shown. After the server is powered on and starts up stably, if the BMC restarts, if RTOS starts successfully, then RTOS controls the fan and monitors whether the BMC has completed startup. If RTOS fails to start successfully, then CPLD controls the fan and monitors whether the BMC has completed startup. If the BMC has completed startup, then BMCLinux controls the fan. During the process of CPLD controlling the fan and monitoring whether the BMC has completed startup, if RTOS starts successfully, then CPLD hands over the control right of the fan to RTOS, and RTOS controls the fan.

[0183] Through the description of the above embodiments, those skilled in the art can clearly understand that the method according to the above embodiments can be implemented by means of software plus a necessary general hardware platform. Of course, it can also be implemented by hardware, but in many cases the former is a better implementation method. Based on such an understanding, the technical solution of the present application, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disc), and includes several instructions for causing a terminal device (which can be a mobile phone, a computer, a server, or a network device, etc.) to execute the methods described in various embodiments of the present application.

[0184] In this embodiment, a control device for a heat dissipation device is further provided. The server includes: server components, a first controller, and a heat dissipation device. The first operating system and the second operating system are deployed on the first controller. This device is used to implement the above embodiments and preferred implementation methods, and those that have been described will not be repeated. As used below, the term "module" can be a combination of software and / or hardware that can achieve a predetermined function. Although the devices described in the following embodiments are preferably implemented in software, implementation in hardware, or a combination of software and hardware is also possible and contemplated.

[0185] Figure 7 It is a structural block diagram of a control device for a heat dissipation device according to an embodiment of the present application, as Figure 7 shown. The device includes:

[0186] The first control module 702 is configured to, when the first controller restarts during the operation of the server, control the operation of the heat dissipation device according to the environmental temperature of the server and the component temperatures of the server components of the first category by the first operating system, and detect the system operation status of the second operating system by the first operating system, where the server components of the first category are the server components that allow the first operating system to collect component temperatures;

[0187] The second control module 704 is configured to, when it is detected that the system operation status is the started state, stop the first operating system from controlling the operation of the heat dissipation device; control the operation of the heat dissipation device according to the environmental temperature of the server and the component temperatures of the server components of the second category by the second operating system, where the server components of the second category are the server components that allow the second operating system to collect component temperatures.

[0188] The control device of the heat dissipation device is further configured to execute the steps in any of the above method embodiments.

[0189] It should be noted that the above-mentioned modules can be implemented by software or hardware. For the latter, it can be implemented in the following ways, but not limited to: the above-mentioned modules are all located in the same processor; or, the above-mentioned modules are respectively located in different processors in any combination form.

[0190] An embodiment of the present application further provides a computer-readable storage medium, in which a computer program is stored, where the computer program is configured to execute the steps in any of the above method embodiments when running.

[0191] In an exemplary embodiment, the above computer-readable storage medium may include, but is not limited to: USB flash drives, read-only memories (ROM for short), random access memories (RAM for short), mobile hard disks, magnetic disks, or optical disks and other media that can store computer programs.

[0192] An embodiment of the present application further provides an electronic device, including a memory and a processor, where a computer program is stored in the memory, and the processor is configured to run the computer program to execute the steps in any of the above method embodiments.

[0193] In an exemplary embodiment, the above electronic device may further include a transmission device and an input / output device, where the transmission device is connected to the above processor, and the input / output device is connected to the above processor.

[0194] Embodiments of the present application also provide a computer program product. The computer program product includes a computer program, and when the computer program is executed by a processor, the steps in any of the above method embodiments are implemented.

[0195] Embodiments of the present application also provide another computer program product, including a non-volatile computer-readable storage medium. The non-volatile computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps in any of the above method embodiments are implemented.

[0196] Embodiments of the present application also provide a computer program. The computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the steps in any of the above method embodiments.

[0197] Specific examples in this embodiment may refer to the examples described in the above embodiments and exemplary embodiments, and will not be elaborated herein.

[0198] Obviously, those skilled in the art should understand that the above modules or steps of the present application can be implemented by a general-purpose computing device. They can be concentrated on a single computing device or distributed on a network composed of multiple computing devices. They can be implemented by program codes executable by the computing device, so that they can be stored in a storage device and executed by the computing device. And in some cases, the steps shown or described can be executed in a different order than here, or they can be separately made into individual integrated circuit modules, or multiple modules or steps among them can be made into a single integrated circuit module to implement. Thus, the present application is not limited to any specific combination of hardware and software.

[0199] The above are only the preferred embodiments of the present application and are not used to limit the present application. For those skilled in the art, the present application can have various changes and modifications. Any modification, equivalent replacement, improvement, etc. made within the principle of the present application shall be included in the protection scope of the present application.

Claims

1. A control method for a heat dissipation device, characterized in that: The server includes: server components, a first controller, and a heat dissipation device. A first operating system and a second operating system are deployed on the first controller. The method includes: When the first controller restarts during the operation of the server, the first operating system controls the operation of the heat dissipation device according to the environmental temperature of the server and the component temperature of the first category of server components, and the first operating system detects the system operation status of the second operating system. Among them, the first category of server components are server components that allow the first operating system to collect component temperatures; When it is detected that the system operation status is the started state, the first operating system stops controlling the operation of the heat dissipation device; the second operating system controls the operation of the heat dissipation device according to the environmental temperature of the server and the component temperature of the second category of server components. Among them, the second category of server components are server components that allow the second operating system to collect component temperatures; Among them, the first operating system detecting the system operation status of the second operating system includes: The first operating system detecting whether the heartbeat of the second operating system is normal; When the heartbeat of the second operating system is detected, the first operating system detects the duration of the heartbeat of the second operating system; When it is detected that the duration of the heartbeat of the second operating system is equal to or exceeds the target duration, the first operating system determines that the detected system operation status is the started state.

2. The method according to claim 1, characterized in that: A first control curve and a first control configuration are further configured on the first controller. The first control curve is used to indicate the first corresponding relationship between the environmental temperature of the server and the operation parameters of the heat dissipation device. The first control configuration is used to indicate the parameter configuration in the operation parameter algorithm corresponding to the first category of server components. The operation parameter algorithm is used to calculate the operation parameters of the heat dissipation device according to the component temperature of the server components; The first operating system controlling the operation of the heat dissipation device according to the environmental temperature of the server and the component temperature of the first category of server components includes: The first operating system calls the first control curve to convert the environmental temperature of the server into a first operation parameter, and the first operating system calls the first control configuration to calculate a second operation parameter according to the component temperature of the first category of server components; The first operating system determines the target operation parameter according to the first operation parameter and the second operation parameter; The first operating system controls the operation of the heat dissipation device according to the target operation parameter.

3. The method according to claim 2, characterized in that: The first operating system calls the first control curve to convert the ambient temperature of the server into a first operating parameter, and the first operating system calls the first control configuration to calculate a second operating parameter according to the component temperature of the server components of the first category, including: During the startup process of the second operating system, the second operating system collects the ambient temperature of the server and the component temperature of the server components of the first category; The second operating system sends the collected ambient temperature of the server and the component temperature of the server components of the first category to the first operating system; The first operating system calls the first control curve to convert the ambient temperature of the server into a first operating parameter, and the first operating system calls the first control configuration to calculate a second operating parameter according to the component temperature of the server components of the first category.

4. The method according to claim 3, wherein A target storage space is also deployed on the server, and the target storage space allows the first operating system and the second operating system to access. The second operating system sending the collected ambient temperature of the server and the component temperature of the server components of the first category to the first operating system includes: The second operating system writes the collected ambient temperature of the server and the component temperature of the server components of the first category into the target storage space; The second operating system sends an interrupt request to the first operating system; The first operating system responds to the interrupt request and reads the ambient temperature of the server and the component temperature of the server components of the first category from the target storage space.

5. The method according to claim 1 or 2, wherein A second control curve and a second control configuration are further configured on the first controller. The second control curve is used to indicate a second corresponding relationship between the ambient temperature of the server and the operating parameter of the heat dissipation device. The second control configuration is used to indicate the parameter configuration in the operating parameter algorithm corresponding to the server components of the second category. The operating parameter algorithm is used to calculate the operating parameter of the heat dissipation device according to the component temperature of the server components; The second operating system controlling the operation of the heat dissipation device according to the ambient temperature of the server and the component temperature of the server components of the second category includes: The second operating system calls the second control curve to convert the ambient temperature of the server into a third operating parameter, and the second operating system calls the second control configuration to calculate a fourth operating parameter according to the component temperature of the server components of the second category; The second operating system determines a reference operating parameter according to the third operating parameter and the fourth operating parameter; The second operating system controls the operation of the heat dissipation device according to the reference operating parameter.

6. The method according to claim 5, wherein The heat dissipation performance of the heat dissipation device corresponding to the operating parameters of the same server's ambient temperature under the first correspondence relationship is higher than that of the heat dissipation device corresponding to the operating parameters under the second correspondence relationship.

7. The method according to claim 5, wherein The heat dissipation types of the server include: multiple heat dissipation types, the multiple heat dissipation types include the target heat dissipation type to which the heat dissipation device belongs, the second control curve includes a first curve and a second curve, and the heat dissipation performance of the heat dissipation device corresponding to the operating parameters of the same server's ambient temperature under the first curve is higher than that of the heat dissipation device corresponding to the operating parameters under the second curve; The conversion of the server's ambient temperature into the third operating parameter by the second operating system calling the second control curve includes: The second operating system detects the heat dissipation type adopted by the server among the multiple heat dissipation types; When the heat dissipation type adopted by the server is the target heat dissipation type, the second operating system calls the first curve to convert the server's ambient temperature into the third operating parameter; When the heat dissipation type adopted by the server is not the target heat dissipation type, the second operating system calls the second curve to convert the server's ambient temperature into the third operating parameter.

8. The method according to claim 5, wherein The second type of server components is divided into a set of server components with multiple risk levels. The second control curve includes: multiple parameter curves corresponding to the multiple risk levels one by one. The higher the risk level, the greater the impact of the component temperature on the server. The heat dissipation performance of the heat dissipation device corresponding to the operating parameters of the server components with a higher risk level on the corresponding control curve is higher under the same ambient temperature of the server; The conversion of the server's ambient temperature into the third operating parameter by the second operating system calling the second control curve includes: The second operating system detects whether there are faulty components among the second type of server components; When it is detected that there are faulty components among the second type of server components, the second operating system detects the target server component set to which the faulty components belong in the set of server components with multiple risk levels; The second operating system calls the reference curve corresponding to the target risk level of the target server component set from the multiple parameter curves to convert the server's ambient temperature into the third operating parameter.

9. The method according to claim 8, wherein The set of server components with multiple risk levels includes: a set of components with high heat dissipation risk and a set of components with low heat dissipation risk. The impact of the component temperature of the server components in the set of components with high heat dissipation risk on the server is higher than that of the server components in the set of components with low heat dissipation risk. The multiple parameter curves include a third curve and a fourth curve. The heat dissipation performance of the heat dissipation device corresponding to the operating parameters of the same server's ambient temperature under the fourth curve is higher than that under the third curve; The second operating system detecting the target server component set to which the faulty component belongs in the set of server components with multiple risk levels includes: the second operating system detecting whether the faulty component is a component with high heat dissipation risk or a component with low heat dissipation risk; The second operating system calling the reference curve corresponding to the target risk level of the target server component set from the multiple parameter curves to convert the ambient temperature of the server into the third operating parameter includes: in the case where the faulty component is detected to be a component with high heat dissipation risk, the second operating system calling the fourth curve to convert the ambient temperature of the server into the third operating parameter; In the case where the faulty component is detected to be a component with low heat dissipation risk, the second operating system calls the third curve to convert the ambient temperature of the server into the third operating parameter.

10. The method according to claim 9, wherein The heat dissipation performance of the heat dissipation device corresponding to the operating parameters of the same server's ambient temperature under the third curve is higher than that under the first curve.

11. The method according to claim 5, wherein The heat dissipation types of the server include: multiple heat dissipation types, the multiple heat dissipation types include the target heat dissipation type to which the heat dissipation device belongs. The second category of server components includes: components with high heat dissipation risk and components with low heat dissipation risk. The impact of the component temperature of the components with high heat dissipation risk on the server is higher than that of the components with low heat dissipation risk. The second control curves include: a fifth curve, a sixth curve, a seventh curve, and an eighth curve. The heat dissipation performance of the heat dissipation device corresponding to the operating parameters of the same server's ambient temperature under the fifth curve is higher than that under the sixth curve. The heat dissipation performance of the heat dissipation device corresponding to the operating parameters of the same server's ambient temperature under the eighth curve is higher than that under the seventh curve. The heat dissipation performance of the heat dissipation device corresponding to the operating parameters of the same server's ambient temperature under the seventh curve is higher than that under the fifth curve; The second operating system calls the second control curve to convert the ambient temperature of the server into a third operating parameter, including: The second operating system detects whether there is a faulty component among the server components of the second category; When it is detected that there is a faulty component among the server components of the second category, the second operating system detects whether the faulty component is a high-risk heat dissipation component or a low-risk heat dissipation component; when it is detected that the faulty component is a high-risk heat dissipation component, the second operating system calls the eighth curve to convert the ambient temperature of the server into the third operating parameter; when it is detected that the faulty component is a low-risk heat dissipation component, the second operating system calls the seventh curve to convert the ambient temperature of the server into the third operating parameter; When it is detected that there is no faulty component among the server components of the second category, the second operating system detects the heat dissipation type adopted by the server among the multiple heat dissipation types; when the heat dissipation type adopted by the server is the target heat dissipation type, the second operating system calls the fifth curve to convert the ambient temperature of the server into the third operating parameter; when the heat dissipation type adopted by the server is not the target heat dissipation type, the second operating system calls the sixth curve to convert the ambient temperature of the server into the third operating parameter.

12. The method according to claim 5, wherein: The server components of the second category include: a plurality of server components, and the second control curve includes a plurality of component curves corresponding to the plurality of server components one by one. The higher the degree of influence of the component temperature on the server, the higher the heat dissipation performance of the heat dissipation device corresponding to the operating parameter on the corresponding control curve at the same ambient temperature of the server; The second operating system calls the second control curve to convert the ambient temperature of the server into a third operating parameter, including: The second operating system detects whether there is a faulty component among the server components of the second category; When it is detected that there is a faulty component among the server components of the second category, the second operating system extracts the target server component belonging to the faulty component from the plurality of server components; The second operating system calls the target curve corresponding to the target server component from the plurality of component curves to convert the ambient temperature of the server into the third operating parameter.

13. The method according to claim 5, wherein: The operating parameter algorithm is used to calculate the operating parameter of the heat dissipation device according to the component temperature of the server component at the sampling moment, the component temperature at the historical moment, and the operating parameter corresponding to the historical sampling moment; The first operating system calls the first heat dissipation configuration to calculate the second operating parameter according to the component temperature of the server components of the first category, including: The first operating system calls the first heat dissipation configuration and substitutes it into the operating parameter algorithm to obtain a target operating parameter algorithm; Obtain the first component temperature of the server components of the first category at the sampling moment, the second component temperature at the historical moment, and the historical operating parameters corresponding to the historical sampling moment from the first operating system, where the historical operating parameters are the operating parameters corresponding to the server components of the first category calculated by the operating parameter algorithm at the historical sampling moment; Substitute the first component temperature, the second component temperature, and the historical operating parameters into the target operating parameter algorithm by the first operating system to obtain the second operating parameter.

14. The method according to claim 13, wherein: The operating parameter algorithm is PWM(k)=PWM(k - 1)+▽PWM; where ▽PWM = Kp*[T(k)-T(k - 1)]+Ki(T(k)-Tsp)+Kd*[[T(k)-T(k - 1)]-[T(k - 1)-T(k - 2)]], PWM(k) is the operating parameter corresponding to the current sampling moment, PWM(k - 1) is the operating parameter corresponding to the historical sampling moment, T(k) is the component temperature at the current sampling moment, the component temperatures at the historical moments include T(k - 1) and T(k - 2), T(k - 1) is the component temperature at the previous sampling moment before the current sampling moment, T(k - 2) is the component temperature at the sampling moment before the previous sampling moment, the parameter configuration in the operating parameter algorithm includes: Kp, Ki, Kd, and Tsp, and Tsp is the temperature value that the component temperature of the server component is allowed to reach at most during operation.

15. The method according to claim 1, wherein: A second controller is also deployed on the server, and the method further includes: In the case where the first controller restarts, control the heat dissipation device to operate by the second controller, and monitor the startup result of the first controller by the second controller; In the case where the startup result indicates that the startup of the first controller fails, identify the operating state of the server by the second controller; In the case where the identified operating state is the running state, continue to control the heat dissipation device to operate by the second controller.

16. The method according to claim 15, wherein: After the second controller monitors the startup result of the first controller, the method further includes: In the case where the startup result indicates that the startup of the first controller is successful, stop controlling the heat dissipation device to operate by the second controller, and control the heat dissipation device to operate by the first controller.

17. The method according to claim 16, wherein: The first controller is connected to the heat dissipation device through the second controller; The second controller controls the operation of the heat dissipation device and monitors the startup result of the first controller, including: the second controller checks whether the heartbeat of the first operating system is normal; when it is detected that the heartbeat of the first operating system is abnormal, the second controller controls the operation of the heat dissipation device and checks whether the heartbeat of the second operating system is normal; when it is detected that the heartbeat of the second operating system is abnormal, it is determined that the startup of the first controller fails; when it is detected that the heartbeat of the first operating system is normal, it is determined that the startup of the first controller is successful.

18. The method according to claim 17, wherein when the startup result is used to indicate that the startup of the first controller fails, the second controller identifies the operating state of the server, including: when the startup result is used to indicate that the startup of the first controller fails, the second controller discards the received operating parameters sent by the first controller and identifies the operating state of the server; when the startup result is used to indicate that the startup of the first controller is successful, the second controller stops controlling the operation of the heat dissipation device, and the first controller controls the operation of the heat dissipation device, including: when the startup result is used to indicate that the startup of the first controller is successful, the second controller sends the received operating parameters sent by the first controller to the heat dissipation device.

19. A control device for a heat dissipation device, wherein the server includes: server components, a first controller, and a heat dissipation device, and the first controller deploys a first operating system and a second operating system. The device includes: a first control module, configured to, when the first controller restarts during the operation of the server, the first operating system controls the operation of the heat dissipation device according to the environmental temperature of the server and the component temperature of the first type of server components, and the first operating system detects the system operating state of the second operating system, wherein the first type of server components are server components that allow the first operating system to collect component temperatures; a second control module, configured to, when it is detected that the system operating state is the started state, the first operating system stops controlling the operation of the heat dissipation device; the second operating system controls the operation of the heat dissipation device according to the environmental temperature of the server and the component temperature of the second type of server components, wherein the second type of server components are server components that allow the second operating system to collect component temperatures; wherein, the first control module is configured to: the first operating system detects whether the heartbeat of the second operating system is normal; when the heartbeat of the second operating system is detected, the first operating system detects the duration of the heartbeat of the second operating system; In the case where the duration of the heartbeat of the second operating system is detected to be equal to or exceed the target duration, the first operating system determines that the system operating state is detected as the started state.

20. A computer-readable storage medium, characterized in that a computer program is stored in the computer-readable storage medium, wherein when the computer program is executed by a processor, the steps of the method described in any one of claims 1 to 18 are implemented.

21. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that when the processor executes the computer program, the steps of the method described in any one of claims 1 to 18 are implemented.

22. A computer program product, comprising a computer program, characterized in that when the computer program is executed by a processor, the steps of the method described in any one of claims 1 to 18 are implemented.

Citation Information

Patent Citations

  • Control method and system of heat dissipation equipment, program product and storage medium

    CN118885061A