Fan control method and apparatus
By working together with the master node, slave nodes, and fan control module, the fan speed is dynamically adjusted to achieve zone control, which solves the problem of insufficient heat dissipation efficiency and flexibility when nodes fail in the existing technology, and improves the heat dissipation efficiency and stability of the server.
Patent Information
- Application Number
- CN202511232718.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-29
- Publication Date
- 2025-11-07
- Estimated Expiration
- 2045-08-29
AI Technical Summary
Existing server fan control methods rely on only single temperature sensor parameters and fan parameters, which cannot achieve redundant control in the event of node failure, resulting in poor heat dissipation efficiency and flexibility.
By working together with the master node, slave nodes, and fan control module, the backplane temperature and hard drive operating status information of the hard drive backplane area are obtained, and the fan speed is dynamically adjusted to achieve partition control. In the event of a node failure, the slave node or fan control module takes over control to ensure that the fan speed matches the hard drive operating status.
It improves the heat dissipation efficiency and flexibility of high-density servers, reduces system power consumption, and ensures stable heat dissipation in the event of node failure.
Smart Images

Figure CN120739727B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of heat dissipation technology, and in particular to a fan control method and device. BACKGROUND
[0002] With the increasing read-write load of hard disks of high-density servers, the requirements for the heat dissipation system continue to increase to ensure the stability and reliability of the system during high-performance operation.
[0003] Currently, the adjustment mechanism of the current fan is to monitor the data of the temperature sensor in real time through the substrate control element (such as BMC, Baseboard Management Controller), and dynamically adjust the fan speed according to the pre-set temperature and speed matching rule; in addition, the control logic of the fan adopts a distributed decision-making mode, so that each node controller independently generates fan control instructions and sends them to the fan board, and the fan board arbitrates the priority of multiple source instructions, and finally executes the parameters issued by the node controller with the highest priority.
[0004] However, the adjustment mechanism of the current fan only adjusts the fan speed according to the global monitoring data of a single temperature sensor, lacks differentiated perception of the temperature of different regions of the device, and results in limited heat dissipation efficiency; in addition, when the node control module fails, the fan controller will switch to the highest speed mode by default, which guarantees the basic heat dissipation demand, but brings unnecessary energy consumption overhead.
[0005] In summary, the existing server fan control method only relies on a single temperature sensor parameter and fan parameter to control the fan speed, cannot realize redundant control when the node fails, and has poor heat dissipation efficiency and flexibility, which needs to be solved urgently. SUMMARY
[0006] The present application provides a fan control method and device to at least solve the technical problem in the related art that only relying on a single temperature sensor parameter and fan parameter to control the fan speed cannot realize server fan redundant control when the node fails, and has poor heat dissipation efficiency and flexibility.
[0007] The application provides a fan control method applied to a server, and the server comprises a master node and at least one slave node, and the method comprises the following steps: acquiring, by the master node, backboard temperature and hard disk working state information of a hard disk backboard area in the server under the condition that the master node meets normal working requirements; acquiring, by the slave node, the backboard temperature and the hard disk working state information of the hard disk backboard area under the condition that the master node does not meet the normal working requirements; acquiring, by a fan control module, the backboard temperature and the hard disk working state information of the hard disk backboard area under the condition that neither the master node nor the slave node meets the normal working requirements; and controlling the working state of a fan based on the backboard temperature and the hard disk working state information.
[0008] The application further provides a fan control device applied to a server, and the server comprises a master node and at least one slave node, and the device comprises: a first acquisition module, which is used for acquiring, by the master node, backboard temperature and hard disk working state information of a hard disk backboard area in the server under the condition that the master node meets normal working requirements; a second acquisition module, which is used for acquiring, by the slave node, the backboard temperature and the hard disk working state information of the hard disk backboard area under the condition that the master node does not meet the normal working requirements; a third acquisition module, which is used for acquiring, by a fan control module, the backboard temperature and the hard disk working state information of the hard disk backboard area under the condition that neither the master node nor the slave node meets the normal working requirements; and a control module, which is used for controlling the working state of a fan based on the backboard temperature and the hard disk working state information.
[0009] The application further provides an electronic device, comprising: a memory for storing a computer program; and a processor for executing the computer program to implement the steps of any of the above fan control methods.
[0010] The application further provides a nonvolatile computer readable storage medium, wherein the nonvolatile computer readable storage medium stores a computer program, and the computer program is executed by a processor to implement the steps of any of the above fan control methods.
[0011] The application further provides a computer program product comprising a computer program, and the computer program is executed by a processor to implement the steps of any of the above fan control methods.
[0012] Through the application, the backboard temperature and the hard disk working state information of the hard disk backboard area in the server can be acquired by the master node under the condition that the master node meets the normal working requirement; the backboard temperature and the hard disk working state information of the hard disk backboard area can be acquired by the slave node under the condition that the master node does not meet the normal working requirement; the backboard temperature and the hard disk working state information of the hard disk backboard area can be acquired by the fan control module under the condition that the master node and the slave node do not meet the normal working requirement; the working state of the fan is controlled based on the backboard temperature and the hard disk working state information, thus the technical problem that the fan speed is controlled only by relying on single temperature sensor parameter and fan parameter in the related art, the server fan redundancy control cannot be realized when the node fails, and the efficiency and flexibility of heat dissipation are poor can be solved, and the technical effect that the fan speed is dynamically adjusted according to the hard disk working state, the read-write speed and the real-time temperature, and the partition control is realized, so that the heat dissipation efficiency of the double-control high-density server is effectively improved is achieved. BRIEF DESCRIPTION OF DRAWINGS
[0013] In order to more clearly illustrate the embodiments of the present application, the drawings needed in the embodiments will be briefly introduced. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor.
[0014] Figure 1 A flow chart of a fan control method according to an embodiment of the present application is provided.
[0015] Figure 2 An execution block diagram of a fan control method according to an embodiment of the present application is provided.
[0016] Figure 3 A schematic diagram of a corresponding relationship between a fan and a hard disk according to an embodiment of the present application is provided.
[0017] Figure 4 A hard disk fan control logic diagram when the master and slave nodes are in a normal working state according to an embodiment of the present application is provided.
[0018] Figure 5 A hard disk fan control logic diagram when the master and slave nodes both fail according to an embodiment of the present application is provided.
[0019] Figure 6 An execution logic diagram of a double-control high-density server hard disk fan control method according to an embodiment of the present application is provided.
[0020] Figure 7 An example diagram of a fan control device according to an embodiment of the present application is provided.
[0021] Wherein, 10-fan control device, 100-first acquisition module, 200-second acquisition module, 300-third acquisition module, 400-control module. DETAILED DESCRIPTION
[0022] The technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, rather than all the embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without any creative work fall within the protection scope of the present application.
[0023] It should be noted that, in the description of the present application, the terms “comprise”, “contain” or any other variants thereof are intended to cover non-exclusive inclusion, so that the process, method, article or device comprising a series of elements not only includes those elements, but also includes other elements not explicitly listed or inherent to such process, method, article or device. The terms “first”, “second” and the like in the present application are used to distinguish similar objects, and are not used to describe a specific order or sequence.
[0024] In order for those skilled in the art to better understand the present application, the present application will be further described in detail below with reference to the drawings and specific embodiments.
[0025] In combination with the specific application environment architecture or specific hardware architecture on which the fan control method is executed, the specific application environment architecture or specific hardware architecture is described here.
[0026] The embodiments of the present application provide a fan control method.
[0027] As shown in the flowchart of the fan control method of the embodiments of the present application, the fan control method comprises the following steps: Figure 1
[0028] In step S101, in the case that the master node meets the normal working requirements, the backboard temperature and hard disk working state information of the hard disk backboard area in the server are acquired through the master node.
[0029] Those skilled in the art should appreciate that, with the rapid development of cloud computing and big data technology, the server performance demand continues to rise, and the dual-control high-density server is increasingly popular in data centers. During the operation of such servers, especially when performing high-density hard disk read-write operations, a large amount of heat will be generated, posing a severe challenge to the cooling system. The current mainstream fan control technology mainly relies on a single temperature parameter for speed regulation, and has problems such as response lag and insufficient regulation accuracy, which is difficult to meet the stringent requirements of modern high-density servers for cooling efficiency. Therefore, it is urgent to develop a new type of intelligent fan control technology that can real-time perceive the hard disk workload changes and dynamically adjust the cooling strategy accordingly, thereby significantly improving the overall cooling performance of the server system, ensuring that the server can maintain the best operating state under various workloads, breaking through the limitations of the traditional temperature control mode, realizing the precise coordination between the cooling system and the storage load, and providing reliable protection for the stable operation of the high-density server.
[0030] Figure 2 The figure is an execution diagram of the fan control method. As shown in Figure 2 , the server hard disk fan control method of the present application is described by taking one slave node as an example. As can be seen from Figure 2 , the fan board in the embodiment of the present application includes a fan control module and a fan connector; the fan is connected with the fan board through the fan connector, and the fan control module reads and controls the fan speed by controlling the PWM (Pulse Width Modulation) signal and TACH (Tachometer) signal of the fan; the CPLD (Complex Programmable Logic Device) and temperature sensor are loaded on the hard disk backboard area, wherein the working state of the hard disk is stored in the CPLD, and the working state includes standby state, light read-write state, continuous high-load read-write state, random high-concurrency read-write state and fault degradation state, and the temperature sensor can detect the temperature of the hard disk in the area; the 9548 chip is an I2C (Inter-Integrated Circuit) multiplexing switch chip, and the node controller communicates with the hard disk backboard area and the fan backboard through the I2C bus, and the I2C bus numbers for communicating with the hard disk backboard area and the fan backboard are different numbers, and both node 1 and node 2 can read all temperature and hard disk state information in the CPLD (the node can be set as the master node by default, and node 2 is set as the slave node by default), to calculate the fan control parameters and send them to the fan control module on the fan board to control the fan rotation.
[0031] Specifically, in actual implementation, the embodiment of the application can first perform fault analysis on the master node and the slave node of the high-density hard disk storage area of the double-controlled high-density server to determine whether the master node and the slave node can work normally.
[0032] Therefore, the embodiment of the application provides reliable guidance and basis for selecting appropriate components and devices (such as the master node, the slave node, the fan control module, etc.) to obtain the backboard temperature of the hard disk backboard area and the hard disk working state information by detecting whether the master node and the slave node have faults.
[0033] Further, the embodiment of the application can control the master node to communicate with the temperature sensor and the CPLD on the hard disk backboard area through the arbitration controller to read the backboard temperature of each hard disk backboard area and the hard disk working state information in the server, thereby providing reliable data basis for fan rotation control under the condition that the master node has no fault.
[0034] Optionally, in an embodiment of the application, the method further includes: obtaining log information of the master node and the slave node respectively, and determining whether the heartbeat packet between the master node and the slave node is lost, wherein when the heartbeat packet between the master node and the slave node is lost, the number of consecutive times of loss of the heartbeat packet between the master node and the slave node is determined; it is determined whether the communication timeout data is recorded in the log information of the master node and the slave node; if the communication timeout data is not recorded in the log information of the master node and / or the slave node, and the number of consecutive times of loss of the heartbeat packet is less than a preset threshold, it is determined that the master node and / or the slave node meets the normal working requirement; if the communication timeout data is recorded in the log information of the master node and / or the slave node, and the number of consecutive times of loss of the heartbeat packet is greater than or equal to the preset threshold, it is determined that the master node and / or the slave node does not meet the normal working requirement.
[0035] It should be noted that in the double-controlled high-density server, the master node and the slave node usually maintain communication through the heartbeat packet to ensure the availability and data consistency of the cluster. However, network fluctuations, hardware failures or software abnormalities may cause the heartbeat packet to be lost, thereby affecting the system stability. The traditional method usually only relies on the heartbeat packet to detect the node state, which is easy to misjudge (such as temporary network jitter leading to mis-switching). The embodiment of the application can more accurately determine whether the node works normally by combining log analysis with heartbeat loss number statistics strategy, thereby reducing misjudgment and improving the reliability of the system.
[0036] Specifically, the process of determining whether the master node and the slave node in the server meet the normal working requirement in the embodiment of the application is as follows:
[0037] 1. Log information acquisition:
[0038] Collect system logs of the master node and the slave node respectively, and check whether key events such as communication timeout or connection exception are recorded.
[0039] 2. Heartbeat packet loss detection:
[0040] (1) Real-time monitoring of whether the heartbeat packet between the master node and the slave node is lost;
[0041] (2) Statistics of the number of consecutive losses (such as 3 consecutive times of not receiving the heartbeat packet).
[0042] 3. The embodiment of the application can determine and illustrate the node state through the following two cases:
[0043] (1) Case 1 (normal node):
[0044] When there is no communication timeout record in the log, and the number of consecutive losses of the heartbeat packet is less than the preset threshold (such as 3 times), the node still meets the work requirements, and it may be a temporary network fluctuation, and no switching is required.
[0045] (2) Case 2 (abnormal node):
[0046] When there is a communication timeout record in the log, and the number of consecutive losses of the heartbeat packet reaches or exceeds the preset threshold, the node may be faulty, and the master-slave switching or alarm needs to be triggered.
[0047] For example, in a double-control high-density server, the master node (Node-M) and the slave node (Node-S) can be set to send a heartbeat packet every 1 second, and the preset consecutive loss threshold is 5 times. If the heartbeat packet is lost for 2 times (not reaching the threshold), and there is no timeout record in the log, it can be determined that the network is temporarily fluctuating, and the node is still working normally, and no switching is triggered. If the heartbeat packet is lost for 6 times (exceeding the threshold), and the master node log shows "TCP (Transmission Control Protocol) connection timeout", it can be determined that the master node communication is abnormal, and the slave node is triggered to be promoted. If the heartbeat packet is lost for 5 times (reaching the threshold), and the slave node log shows "process crash", it can be determined that the slave node is unavailable, the system alarms and attempts to recover.
[0048] Therefore, the embodiment of the application improves the fault determination accuracy of the distributed storage system through the double detection mechanism of heartbeat packet loss statistics and log analysis strategy, and can effectively distinguish between temporary network problems and real node faults, reduce misoperation, and is suitable for production environments with high stability requirements.
[0049] Optionally, in an embodiment of the present application, before acquiring the log information of the master node and the slave node respectively, further comprising: partitioning the hard disk area of the server to obtain a plurality of hard disk backboard areas; numbering the plurality of hard disk backboard areas and the hard disks in the hard disk backboard areas to obtain corresponding hard disk backboard area numbers and hard disk numbers, and numbering the fans in the server to obtain corresponding fan numbers; determining the corresponding relationship between the fans and the hard disks based on the fan numbers, the hard disk backboard area numbers and the hard disk numbers, so that the fans cool the corresponding hard disks according to the corresponding relationship.
[0050] It should be noted that, Figure 3 is a schematic diagram of the corresponding relationship between the fans and the hard disks. As Figure 3 shown, the embodiment of the present application first partitions the hard disk area of the server to obtain a plurality of hard disk backboard areas; secondly, the embodiment of the present application numbers each hard disk backboard area and all the hard disks in each hard disk backboard area in turn; in addition, the hard disk areas cooled by each fan are separated by a flow guide strip, so that the cooling area of each fan is independent.
[0051] For example, as Figure 3 shown, fan 1 corresponds to hard disks 1 to 21 of hard disk backboard area 1, fan 2 corresponds to hard disks 22 to 35 of hard disk backboard area 1 and hard disks 1 to 7 of hard disk backboard area 2, and so on, so as to determine the corresponding relationship between each fan and each hard disk, so as to control each fan to cool each hard disk according to the corresponding relationship.
[0052] In actual execution process, the number of hard disk backboard areas, the number of fans, and the corresponding relationship between the fans and the hard disks can be determined according to the size of the fan and the position of the hard disk, which is not specifically limited here.
[0053] Therefore, the embodiment of the present application partitions the hard disk area, numbers each fan, each hard disk area (i.e. hard disk backboard area) after partitioning, and all the hard disks in each hard disk area, so as to determine the corresponding relationship between the fan and each hard disk according to the corresponding number, so as to determine the corresponding cooling object of each fan, realize partition control, and improve the efficiency of hard disk cooling.
[0054] Optionally, in an embodiment of the present application, partitioning the hard disk area of the server to obtain a plurality of hard disk backboard areas comprises: collecting airflow temperature distribution data corresponding to the hard disk area of the server to construct a target thermodynamic model according to the airflow temperature distribution data; obtaining a plurality of current working parameters of the hard disks in the server based on the target thermodynamic model, and constructing a multi-objective optimization model through the plurality of current working parameters; solving the multi-objective optimization model to obtain a partitioning result corresponding to the hard disk area of the server, and determining the plurality of hard disk backboard areas according to the partitioning result.
[0055] As an implementable way, the specific process of partitioning the hard disk region of the server by the embodiment of the application is as follows:
[0056] 1. Three-dimensional thermal field modeling stage:
[0057] (1) Collect the temperature distribution of the surface and internal airflow of the hard disk group through a high-density temperature sensor array;
[0058] (2) Establish a three-dimensional thermodynamic model containing temperature gradient, airflow velocity, and structural resistance.
[0059] 2. Dynamic load characteristic analysis stage:
[0060] (1) Real-time monitoring of the composite working parameters (i.e., multiple current working parameters) of each hard disk, including:
[0061] 1) Thermal load characteristics of read-write operations;
[0062] 2) Spatiotemporal distribution characteristics of data access patterns;
[0063] 3) Correlation between cache hit rate and queue depth.
[0064] It should be noted that the temperature gradient distribution of the three-dimensional thermal field model can be used as a benchmark input for thermal load characteristic analysis, the airflow velocity field data can be used to correct the spatiotemporal distribution calculation of the data access pattern, and the structural resistance parameter can jointly affect the weight distribution of the cache hit rate with the queue depth monitoring, forming the basis of thermal-electric composite analysis.
[0065] 3. Self-adaptive optimization partitioning stage:
[0066] (1) Based on the genetic algorithm, a multi-objective optimization problem (i.e., a multi-objective optimization model) is solved to output an optimal partitioning scheme that meets the heat dissipation requirements, and the objective functions of the multi-objective optimization problem include:
[0067] 1) Objective function 1: Minimize the maximum temperature difference within the partition (≤4℃);
[0068] 2) Objective function 2: Equalize the load difference between partitions (≤10%);
[0069] 3) Objective function 3: Maximize the utilization efficiency of the cooling airflow.
[0070] In actual implementation, the heat load characteristics can be converted into a temperature constraint condition of the objective function 1, the load balancing parameters of the objective function 2 can be generated through the space-time distribution characteristics, the air flow efficiency coefficient of the objective function 3 can be derived through the cache-queue association relationship, and the heat load characteristics, the space-time distribution characteristics and the cache-queue association relationship can be normalized through the fitness function of the genetic algorithm, so as to ensure the synergy of multi-objective optimization.
[0071] Therefore, through the three-dimensional heat field modeling, the dynamic load characteristic analysis and the adaptive optimization partitioning, the embodiments of the present application can intelligently partition the server hard disk region, so as to realize the deep synergy partitioning of the thermodynamic characteristics and the work load, and effectively improve the heat dissipation efficiency of the subsequent server.
[0072] Optionally, in an embodiment of the present application, before the log information of the master node and the slave node is acquired respectively, the method further comprises: acquiring priority parameters of each node pre-configured in a baseboard management controller of the server, so as to determine the master node and the slave node according to the priority parameters; performing health degree evaluation on the master node every preset time length, so as to generate corresponding evaluation data, and calculating a priority score corresponding to the master node according to the evaluation data, wherein the evaluation data comprises communication response speed, historical control accuracy and current resource load of the master node; and judging whether the priority score is less than a preset priority switching score threshold, wherein if the priority score is less than the priority switching score threshold, the master node is switched to a new slave node, and the slave node with the highest priority score is switched to a new master node.
[0073] Those skilled in the art should appreciate that the existing distributed storage system usually adopts a master-slave architecture to manage storage nodes, wherein the master node is responsible for coordinating core tasks such as data read-write and load balancing. However, the traditional master node election method is usually static, which cannot dynamically adapt to node performance changes, and may cause the overall system efficiency to decrease. The embodiments of the present application can improve the stability and reliability of the system by dynamically evaluating the health degree of the master node and automatically switching the master-slave roles when necessary.
[0074] Specifically, the process of determining the master-slave nodes according to the priority is as follows:
[0075] 1. Acquire node priority parameters:
[0076] (1) Read the priority parameters of each storage node from the baseboard management controller, which may include hardware performance, historical stability, network delay, etc.
[0077] (2) According to the priority parameters, initially select the master node and the slave node.
[0078] 2. Master node health assessment:
[0079] (1) Every preset time (e.g. 5 minutes), the health of the master node is detected, and the evaluation indicators include:
[0080] 1) Communication response speed (e.g. RPC (Remote Procedure Call) call time);
[0081] 2) Historical control accuracy (e.g. task execution success rate in the past period of time);
[0082] 3) Current resource load (e.g. CPU (Central Processing Unit), memory, disk I / O (Input / Output) usage);
[0083] 4) Calculate the priority score of the master node (e.g. 0-100 points) according to these indicators.
[0084] 3. Master-slave switching judgment:
[0085] (1) If the priority score of the master node is lower than the preset threshold (e.g. 60 points), trigger the master-slave switching:
[0086] 1) The original master node is downgraded to a slave node;
[0087] 2) The node with the highest priority in the slave node is upgraded to the new master node;
[0088] 3) If the master node still meets the requirements, maintain the current architecture.
[0089] For example, if a distributed server contains 2 nodes (Node1, Node2), the initial priority is 90, 85 respectively, so Node1 is selected as the master node, and the specific node dynamic priority assessment is as follows:
[0090] 1. Initial state:
[0091] (1) Master node: Node1 (priority 90);
[0092] (2) Slave node: Node2 (85).
[0093] 2. Health detection:
[0094] After 5 minutes, it is detected that the communication response speed of Node1 decreases (the delay rises from 10ms to 50ms), the historical control accuracy decreases from 98% to 92%, and the CPU load reaches 90%; thus, the priority score can be calculated, which decreases from 90 to 55 (lower than the threshold of 60 points).
[0095] 3. Trigger master-slave switching:
[0096] (1) degrade Node1 to a slave node;
[0097] (2) select the highest priority Node2 (85 points) as the new master node.
[0098] 4. Post-switch state:
[0099] New master node: Node2; slave node: Node1.
[0100] Therefore, the embodiments of the present application can adapt to node performance changes, improve overall reliability and efficiency, and better cope with uncertainties in actual production environments by dynamic priority evaluation combined with automatic master-slave switching mechanism, and can be well applied to high-availability distributed storage systems.
[0101] Optionally, in an embodiment of the present application, it further comprises: when the master node and the slave node both meet the normal working requirements, judging whether there is a communication link fault between the master node and the slave node and the fan control module; if there is no communication link fault between the master node and the slave node and the fan control module, obtaining the backplane temperature and the hard disk working state information of the hard disk backplane area through the master node to calculate the corresponding fan control parameters; if there is a communication link fault between the master node and the slave node and the fan control module, so that the fan control module does not receive the corresponding fan control parameters of the master node and the slave node within a preset response waiting time, obtaining the backplane temperature and the hard disk working state information of the hard disk backplane area through the fan control module to control the corresponding fan of the hard disk backplane area to rotate based on the backplane temperature and the hard disk working state information.
[0102] In the specific implementation process, when the master node and the slave node both meet the normal working requirements (i.e., the master and slave nodes are both in normal working state), the embodiments of the present application can judge whether there is a communication link fault between the master node and the slave node and the fan control module.
[0103] If there is no communication link fault between the master node and the slave node and the fan control module, the embodiments of the present application can control the master node and the slave node to communicate with the temperature sensor and the CPLD on the hard disk backplane area through the arbitration controller to read the backplane temperature and the hard disk working state information (such as hard disk working mode, hard disk read-write rate, etc.) of the hard disk backplane area, and according to the set priority, the highest priority master node calculates the corresponding fan control parameters and sends them to the fan control module on the fan board to control the fan to rotate.
[0104] If the communication between the main controller and the fan control module fails to send fan control information such as backboard temperature and hard disk working state information to the fan control module on the fan board, but the main and standby priorities in the BMC are still the main control node, the embodiment of the application can take over the fan control right through the slave node; if the main and slave nodes are both faulty or cannot communicate (that is, there is a communication link failure), so that the fan control module cannot receive the information of the node control module, that is, the fan control module does not receive the fan control parameters corresponding to the main and slave nodes within a preset response waiting time, the embodiment of the application can send a communication link failure signal to the arbitration controller through the fan control module to obtain the communication right, so as to communicate with the hard disk backboard area through the fan control module through I2C, obtain the corresponding hard disk temperature and hard disk working state information, and then calculate the fan control parameters to realize the rotation control of the fan.
[0105] Therefore, the embodiment of the application realizes dual-channel redundant control and autonomous fault-tolerant mechanism through multi-level cooperation of the main and slave nodes and the fan module, ensures that precise heat dissipation can be maintained when the communication is abnormal, so that intelligent wind speed adjustment can be performed according to the actual working condition of the hard disk.
[0106] In step S102, the backboard temperature and hard disk working state information of the hard disk backboard area are obtained through the slave node when the main control node does not meet the normal working requirements.
[0107] In step S103, the backboard temperature and hard disk working state information of the hard disk backboard area are obtained through the fan control module when neither the main control node nor the slave node meets the normal working requirements.
[0108] Further, the embodiment of the application can obtain the backboard temperature and hard disk working state information of each hard disk backboard area through the slave node when the main control node has a fault and the slave node (which is the slave node with the highest priority among all slave nodes in the multi-slave node server or the single slave node in the embodiment of the application) does not have a fault; in addition, the embodiment of the application can obtain the backboard temperature and hard disk working state information of the hard disk backboard area through the fan control module in the server when both the main and slave nodes have faults.
[0109] Therefore, the embodiment of the application can obtain the backboard temperature and hard disk working state information of the hard disk backboard area through the slave node or the fan control module when the main control node has a fault or both the main and slave nodes have faults, so that the rotation operation of the fan can be efficiently and accurately controlled under different node states.
[0110] Optionally, in one embodiment of the present application, when neither the master node nor the slave node meets the normal working requirement, the fan control module is used to obtain the backboard temperature of the hard disk backboard area and the hard disk working state information, including: if neither the master node nor the slave node meets the normal working requirement, the slave node is used to send a slave node failure signal to the fan control module to control the fan control module to send a communication right request signal to the preset arbitration controller; after the arbitration controller receives the communication right request signal, the fan control module is controlled to read the backboard temperature of the hard disk backboard area and the hard disk working state information through the preset bidirectional binary synchronous serial bus.
[0111] It should be noted that if both the master node and the slave node fail, the slave node is controlled to send a slave node failure signal to the fan control module, and after the fan control module receives the slave node failure signal, the fan control module is used to send a communication right request signal to the arbitration controller to obtain the communication right with the hard disk backboard area, so that the fan control module reads the backboard temperature and the hard disk working state information through the I2C, and then calculates the fan control parameter to control the fan to rotate.
[0112] Therefore, the embodiment of the present application can read the temperature and the hard disk working state information through the fan control module after the node fails to calculate the fan control parameter, thereby improving the stability of the server, reducing the system power consumption, and solving the problems of the traditional technology, such as the large noise and power consumption of the fan controller controlling the fan to rotate at full speed after the node fails.
[0113] Optionally, in one embodiment of the present application, when the master node does not meet the normal working requirement, the slave node is used to obtain the backboard temperature of the hard disk backboard area and the hard disk working state information, including: if the master node does not meet the normal working requirement, the master node is used to send a master node failure signal to the slave node to control the slave node to obtain the backboard temperature of the hard disk backboard area and the hard disk working state information; if the master node does not meet the normal working requirement and the master node does not send the master node failure signal to the slave node, the slave node is used to perform communication detection on the heartbeat signal of the master node to obtain a corresponding detection result; if the slave node determines that the master node does not meet the normal working requirement according to the detection result, the slave node is controlled to obtain the backboard temperature of the hard disk backboard area and the hard disk working state information.
[0114] In actual execution process, if the master node fails and the slave node meets the normal working requirement, the master node is used to send a master node failure signal to the slave node, so that the slave node obtains the fan control right to read the backboard temperature and the hard disk working state information.
[0115] If the master node fails but the master node does not send a master node failure signal to the slave node, the master node and the slave node can detect whether each other is normal through heartbeat signals, so that the slave node detects the heartbeat signal of the master node through inter-node communication to determine whether the master node has a failure.
[0116] For example, in a dual-control high-density server operation, the master node is abnormal due to power fluctuation (without triggering a failure signal), the slave node detects that the heartbeat signal of the master node is lost more than 5 times (a preset threshold is 3 times) through continuous detection, and the temperature data of the master node is not updated for 8 seconds, so that the embodiment of the application controls the slave node to immediately take over the fan control right to collect the backboard temperature and the hard disk working state information in real time.
[0117] Therefore, the embodiment of the application actively identifies the node abnormality without alarm through the heartbeat signal, and when the master node has a failure, timely controls the slave node to autonomously decide to take over the fan control right, thereby greatly improving the node failure discovery rate, avoiding the heat dissipation vacuum period, and improving the heat dissipation efficiency of the server.
[0118] In step S104, the working state of the fan is controlled based on the backboard temperature and the hard disk working state information.
[0119] After the backboard temperature and the hard disk working state information are acquired, the embodiment of the application can dynamically calculate the fan control parameter and send it to the fan controller to control the rotation of the fan, thereby improving the heat dissipation efficiency of the server.
[0120] Optionally, in an embodiment of the application, the working state of the fan is controlled based on the backboard temperature and the hard disk working state information, including: inputting the backboard temperature and the hard disk working state information into a pre-constructed multi-parameter weight model to output the target rotation speed of the fan corresponding to the hard disk backboard area; determining the current read-write operation frequency of the server and calculating the corresponding rotation speed compensation value according to the current read-write operation frequency to correct the target rotation speed through the rotation speed compensation value to obtain the final fan rotation speed; determining the corresponding fan control parameter based on the final fan rotation speed and sending the fan control parameter to the fan control module to control the fan corresponding to the hard disk backboard area to perform the corresponding rotation operation according to the fan control parameter through the fan control module.
[0121] It should be noted that after the backboard temperature and the hard disk working state information (such as read-write speed, working state (such as running / standby / failure)) are read, the embodiment of the application can input the backboard temperature and the hard disk working state information into a preset multi-parameter weight model to calculate the optimal rotation speed (i.e., the target rotation speed) of each fan area, determine the current read-write operation frequency of the server to calculate the corresponding rotation speed compensation value, and correct the target rotation speed through the rotation speed compensation value to obtain the final fan rotation speed.
[0122] For example, in the embodiments of the present application, the fan adjustment is 200 revolutions per 1℃ change in the hard disk temperature, and the current read-write rate value can be converted into a speed compensation value in proportion to one thousandth, while being dynamically modified by referring to the 15-minute load trend. When the temperature of a certain hard disk area exceeds the threshold or the read-write rate suddenly increases, the corresponding fan speed is immediately increased by 30%, and the adjacent hard disk area is assisted to increase by 10%, so as to ensure that the temperature gradient does not exceed 15℃.
[0123] Further, the embodiments of the present application can determine the corresponding fan control parameters based on the final fan speed of each hard disk area, and pack the fan control parameters into control instructions, which are sent to the fan control module through a special communication channel. Finally, the fan control module can analyze the corresponding instructions and drive each fan motor to operate at the set speed.
[0124] As an implementable way, the embodiments of the present application can first obtain the basic speed setting value of each fan area by inputting the hard disk backboard area temperature field distribution characteristics and the hard disk working state parameters into a preset multi-parameter weight model. At the same time, the frequency characteristics of the current read-write operation of the server are analyzed, and a dynamic speed compensation coefficient is generated by using a fuzzy reasoning algorithm. The compensation coefficient and the basic speed are subjected to nonlinear superposition operation, and the final speed control instruction is output after adaptive filtering processing.
[0125] Thus, the embodiments of the present application innovatively realize the collaborative decoupling calculation of the temperature field gradient change and the storage load characteristics, so that the fan speed can accurately match the actual heat dissipation demand, effectively improve the heat dissipation efficiency, and reduce the system power consumption on the premise of ensuring stable operation of the equipment.
[0126] In addition, it can be understood that the embodiments of the present application calculate the fan control parameters by the hard disk temperature and the hard disk working state information (hard disk working mode, hard disk read-write rate, etc.), without relying on a single temperature parameter, thereby greatly improving the heat dissipation efficiency and realizing accurate partition heat dissipation, while ensuring the stability of the hard disk working temperature and significantly reducing the overall system power consumption.
[0127] It should be noted that the specific process of constructing the multi-parameter weight model by the embodiments of the present application is as follows:
[0128] 1. Input parameter fusion:
[0129] Integrate the backboard temperature (real-time monitoring value, historical trend), hard disk working state (read-write load, IOPS (Input / Output Operations Per Second), fault warning signal) and environmental parameters (cabinet air flow speed, environmental temperature and humidity) to form a multi-dimensional feature vector;
[0130] 2. Dynamic weight distribution:
[0131] Adopt fuzzy logic or adaptive neural network to dynamically adjust the weight of each parameter according to real-time working conditions, for example, increase the weight of backboard temperature during high temperature period, and increase the IOPS influence coefficient during high load period;
[0132] 3. Nonlinear relationship modeling:
[0133] Learn the complex nonlinear mapping of temperature-speed through machine learning (such as XGBoost (eXtreme Gradient Boosting) or lightweight deep learning), and introduce hysteresis effect compensation (such as fan response delay) to ensure smooth and stable speed decision;
[0134] 4. Feedback optimization mechanism:
[0135] Combine reinforcement learning to dynamically correct model parameters according to actual cooling effect to achieve continuous iterative optimization.
[0136] Therefore, the multi-parameter weight model of the embodiments of the present application intelligently integrates multi-dimensional data, dynamically adjusts parameter weights, accurately models nonlinear relationships, and introduces real-time feedback mechanisms to achieve adaptive optimization of the hard disk cooling system, thereby effectively improving cooling accuracy, ensuring that key components meet temperature control standards, and dynamically adapting to different load scenarios to optimize energy efficiency. In addition, the embodiments of the present application can continuously improve the control strategy through continuous learning, and can balance response speed and long-term stability, effectively extending the service life of the equipment.
[0137] Optionally, in an embodiment of the present application, the backboard temperature and hard disk working state information are input into a pre-constructed multi-parameter weight model, including: performing shunt collection and time alignment operation on the backboard temperature and hard disk working state information, and binding the backboard temperature and the hard disk working state information corresponding to the backboard temperature in the same collection period; determining the smoothing window size corresponding to the preset dynamic window smoothing strategy according to the read-write frequency in the hard disk working state information, to perform dynamic window smoothing processing on the bound backboard temperature based on the smoothing window size, to obtain a smoothed backboard temperature; performing effectiveness verification and abnormal marking on the hard disk working state information corresponding to the smoothed backboard temperature, to identify invalid state information of the bound hard disk working state information, and performing missing value filling on the hard disk working state information corresponding to the invalid state information, to obtain hard disk working state filling information; input the smoothed backboard temperature and the hard disk working state information or the hard disk working state filling information corresponding to the smoothed backboard temperature into the multi-parameter weight model.
[0138] It should be noted that before the backboard temperature and hard disk working state information are input into the pre-constructed multi-parameter weight model, corresponding preprocessing operations are also required, and the specific process is as follows:
[0139] Step 1, the acquired hard disk backboard temperature data and hard disk working state information are collected and time-aligned, and the temperature data and the working state information of the corresponding hard disk in the same collection period are bound, wherein the hard disk working state information includes the read-write frequency, data transmission volume and error check code of the hard disk;
[0140] Step 2, the bound temperature data are dynamically window-smoothed, and the smoothing window size is determined according to the read-write frequency in the hard disk working state information; the higher the read-write frequency, the smaller the window, so as to retain the temperature rapid change characteristics and filter transient pulse noise;
[0141] Step 3, the bound hard disk working state information is verified for effectiveness and marked for abnormality, invalid state information is identified based on a preset state code rule library, and missing value filling is performed on the state information marked as abnormal based on a Markov chain model of the historical state sequence of the same type of hard disk;
[0142] Then, the embodiment of the application can input the smoothed backboard temperature and the filled hard disk working state information corresponding to the smoothed backboard temperature (i.e. hard disk working state filling information) or hard disk working state information without filling into the multi-parameter weight model.
[0143] In addition, as a realizable way, the embodiment of the application can also construct temperature-state time sequence association characteristics based on the temperature data processed in step 2 and the working state information processed in step 3, calculate the Pearson correlation coefficient of the temperature change rate per unit time and the hard disk read-write frequency, and supplement the coefficient as a new feature to the original parameter set; then, the embodiment of the application can perform dynamic range standardization processing on the supplemented parameter set, divide the load interval according to the data transmission volume in the hard disk working state information, and perform standardization according to the interval mean and standard deviation for different load intervals, so as to ensure the consistency of parameter distribution in different load scenarios, and obtain further optimized preprocessed data which can be directly input into the multi-parameter weight model.
[0144] Therefore, before the backboard temperature and hard disk working state information are input into the pre-constructed multi-parameter weight model, the embodiment of the application preprocesses the backboard temperature and hard disk working state information, thereby greatly improving the quality of the data and effectively guaranteeing the accuracy and reliability of the output data of the multi-parameter weight model.
[0145] Optionally, in an embodiment of the present application, the corresponding rotation speed compensation value is calculated according to the current read-write operation frequency, comprising: collecting the current read-write operation frequency of the server in real time, and distinguishing the read operation frequency and the write operation frequency in the current read-write operation frequency, and recording the continuous execution time length and interval characteristics of the read operation frequency and the write operation frequency; calculating the frequency ratio of the read operation frequency and the write operation frequency, so as to calculate the read-write operation comprehensive load coefficient according to the frequency ratio and the continuous execution time length; obtaining the current temperature baseline of the hard disk of the server, and determining the temperature sensitive factor according to the current temperature baseline, multiplying the read-write operation comprehensive load coefficient and the temperature sensitive factor to obtain the basic compensation value; obtaining the new current read-write operation frequency corresponding to the server, so as to calculate the new frequency ratio according to the new current read-write operation frequency, and dynamically correcting the basic compensation value through the new frequency ratio to obtain the rotation speed compensation value.
[0146] It should be noted that the specific process of calculating the fan rotation speed compensation value according to the read-write operation frequency of the server is as follows:
[0147] Step 1, collecting the current read-write operation frequency of the server in real time, distinguishing the read operation frequency and the write operation frequency, and recording the continuous execution time length and interval characteristics of the two types of operations;
[0148] Step 2, calculating the read-write operation comprehensive load coefficient according to the ratio of the read operation frequency and the write operation frequency, combined with the respective continuous execution time length, wherein the write operation corresponds to a higher load weight, and the longer the continuous execution time length, the greater the cumulative growth amplitude of the load coefficient;
[0149] Step 3: determining the temperature sensitive factor based on the current temperature baseline of the server hard disk, which increases with the increase of the difference between the current temperature of the hard disk and the baseline temperature;
[0150] Step 4: multiplying the comprehensive load coefficient obtained in step 2 and the temperature sensitive factor obtained in step 3 to obtain the basic rotation speed compensation value, so as to reflect the additional contribution of the current read-write load to the heat dissipation demand;
[0151] Step 5: dynamically correcting the basic rotation speed compensation value according to the change rate of the read-write operation frequency, if the frequency shows an upward trend, the compensation value is increased by a preset proportion, if the frequency shows a downward trend, the compensation value is decreased by a preset proportion, and the correction amplitude is adjusted accordingly with the increase of the absolute value of the change rate;
[0152] Step 6: taking the compensation value after dynamic correction as the final rotation speed compensation value, which is used to correct the target rotation speed of the fan.
[0153] Therefore, the embodiment of the present application improves the accuracy of the compensation value calculation, optimizes the fan rotating speed, and improves the heat dissipation efficiency of the server by calculating the fan rotating speed compensation value according to the read-write operation frequency of the server.
[0154] Optionally, in an embodiment of the present application, the method further comprises: acquiring the actual running state of the fan corresponding to the plurality of hard disk backboard regions in real time, wherein the actual running state comprises the actual rotating speed and the current characteristic of the driving circuit; determining the expected running state and the expected heat dissipation requirement of the fan according to the fan control parameter, and acquiring the dynamic change trend of the backboard temperature; comparing the actual running state with the expected running state, and judging whether the fan has a fault in combination with the dynamic change trend of the backboard temperature, wherein if the deviation between the actual running state and the expected running state exceeds the preset normal range, and the change of the backboard temperature does not meet the expected heat dissipation requirement, it is determined that the fan has a fault.
[0155] Specifically, the embodiment of the present application first acquires the backboard temperature and the hard disk working state information of the server hard disk backboard region in real time, calculates the corresponding fan control parameter according to the backboard temperature and the hard disk working state information, wherein the hard disk working state information comprises the hard disk read-write load rate and the data transmission rate; secondly, the embodiment of the present application drives the fan to rotate according to the fan control parameter, and acquires the actual running parameter of the target fan in real time, wherein the actual running parameter comprises the actual rotating speed and the driving current; thirdly, the embodiment of the present application compares the actual running parameter with the theoretical running parameter corresponding to the fan control parameter, and judges whether the fan has a fault in combination with the change rate of the backboard temperature, wherein if the deviation between the actual rotating speed and the theoretical rotating speed exceeds the preset threshold value, the driving current is abnormal, and the change rate of the backboard temperature does not meet the preset heat dissipation trend, it is determined that the fan has a fault.
[0156] Therefore, the embodiment of the present application judges whether the fan has a fault, thereby guaranteeing the integrity of the server fan control method, and improving the robustness and reliability of the fan heat dissipation.
[0157] Optionally, in an embodiment of the present application, if the deviation between the actual running state and the expected running state exceeds the preset normal range, and the change of the backboard temperature does not meet the expected heat dissipation requirement, it is determined that the fan has a fault, comprising: determining the fan heat dissipation requirement corresponding to the hard disk backboard region according to the expected heat dissipation requirement; if the fan heat dissipation requirement is the fan enhanced heat dissipation requirement, but the actual rotating speed data in the actual running state of the fan is not improved as expected, and it is determined according to the dynamic change trend that the backboard temperature presents an upward trend, it is determined that the fan has a mechanical fault; if the fan heat dissipation requirement is the fan maintained heat dissipation requirement, but the actual rotating speed data of the fan appears non-periodic fluctuation, and the backboard temperature presents synchronous abnormal fluctuation with the actual rotating speed data, it is determined that the fan has a driving circuit fault.
[0158] In actual execution, the embodiment of the application can determine whether the target fan has a fault in combination with the dynamic change trend of the backboard temperature. If the fan control parameter indicates that the fan needs to enhance the heat dissipation capacity, the actual rotation speed does not increase as expected, and the backboard temperature presents an upward trend or the cooling rate is lower than expected for a specific duration, it is determined that the fan has a mechanical fault. If the fan control parameter indicates that the fan maintains a stable heat dissipation state, the actual rotation speed presents non-periodic fluctuation, and the backboard temperature presents synchronous abnormal fluctuation with the rotation speed fluctuation, it is determined that the fan has a drive circuit fault.
[0159] Therefore, the embodiment of the application can provide reliable data guidance and basis for subsequent fault repair and alarm by performing specific fault analysis on the fan.
[0160] Optionally, in an embodiment of the application, after determining that the fan has a fault, the method further includes: sending a reset instruction to the fan control module to reacquire new backboard temperature and new hard disk working state information corresponding to the hard disk backboard area, and calculating new fan control parameters according to the new backboard temperature and the new hard disk working state information; controlling the fan to rotate according to the new fan control parameters to obtain a current running state and a new expected running state of the fan, and calculating a state deviation between the current running state and the new expected running state, and determining whether the state deviation exceeds a preset normal range, wherein, in the case that the state deviation exceeds the preset normal range, a fault reminding function is triggered, and the server is controlled to perform a corresponding fault reminding operation.
[0161] It should be noted that, if it is determined that the fan has a fault, the embodiment of the application can first send a reset instruction to the fan control module to clear the abnormal state and reload the fan control parameters. If the actual rotation speed still does not reach the normal range of the theoretical rotation speed, the amplitude or frequency characteristics of the drive signal are gradually adjusted, and the actual running state is observed after each adjustment for a preset duration. If the deviation between the actual running state and the theoretical running state (i.e., the expected running state) is still not reduced to the normal range after a preset number of adjustments, the repair operation is terminated, and the subsequent fault reminding operation is performed.
[0162] Therefore, the embodiment of the application repairs the fault and timely performs fault reminding in the case of repair failure, thereby effectively guaranteeing the intelligent degree of fan heat dissipation and providing solid technical support for realizing efficient fan heat dissipation.
[0163] Optionally, in an embodiment of the present application, in the case that the state deviation exceeds the preset normal range, a fault reminding function is triggered, and the server performs corresponding fault reminding operation, including: determining the fault fan number corresponding to the fault fan and the hard disk backplane area fault number and the hard disk fault number corresponding to the fault fan number, and generating corresponding visual fault identification according to the fault fan number, the hard disk backplane area fault number and the hard disk fault number; obtaining the current running state of the fault fan and the dynamic change trend of the backplane temperature, and generating corresponding fan running state record and backplane temperature change characteristic description according to the current running state and the dynamic change trend of the backplane temperature; based on the fan running state record and the backplane temperature change characteristic description, the corresponding fault alarm information is constructed, and the visual fault identification is displayed on the local management interface of the server, and the fault alarm information is sent to the preset remote management terminal.
[0164] In actual execution process, the embodiment of the present application can display the fault identification on the local management interface of the server, which contains the hard disk backplane area where the fault fan is located and the corresponding physical slot number (i.e. the fault fan number and the hard disk backplane area fault number and the hard disk fault number corresponding to the fault fan number); at the same time, the embodiment of the present application can also send the fault alarm information to the preset remote management terminal, which is accompanied by the fan running state record and the characteristic description of the backplane temperature change curve (i.e. the backplane temperature change characteristic description) during the fault occurrence period, to assist fault positioning.
[0165] Thus, the embodiment of the present application realizes rapid fault diagnosis and alarm by accurately positioning the fault source and generating visual identification, and combines real-time state monitoring and trend analysis, thereby significantly improving the server operation and maintenance efficiency and system reliability.
[0166] Optionally, in an embodiment of the present application, it also includes: based on the actual running state and the expected running state of the fan, calculating the corresponding running state deviation value, and based on the dynamic change trend of the backplane temperature and the expected heat dissipation requirement, calculating the heat dissipation error; adjusting the backplane temperature and the hard disk working state information according to the running state deviation value and the heat dissipation error, and fine-tuning the pre-constructed multi-parameter weight model through the adjusted backplane temperature and the hard disk working state information.
[0167] It should be noted that the embodiment of the present application can compare the actual running state and the expected running state of the fan to calculate the corresponding running state deviation value; in addition, the embodiment of the present application also needs to monitor the backplane temperature change rate or trend synchronously, and compare the expected heat dissipation requirement to calculate the corresponding heat dissipation error.
[0168] Secondly, the embodiment of the present application can intelligently adjust the backboard temperature sampling frequency and the hard disk load distribution strategy according to the corresponding deviation value and the heat dissipation error, balance the heat dissipation demand and the equipment life, obtain a plurality of adjusted backboard temperatures and hard disk working state information, construct a corresponding fine-tuning training data set to fine-tune the multi-parameter weight model, that is, the embodiment of the present application can feed the adjusted temperature and load data to the multi-parameter weight model, update the weight coefficient through the incremental learning mechanism, and improve the adaptability of the rotation speed decision.
[0169] Therefore, the embodiment of the present application can optimize the multi-parameter weight model by adjusting the backboard temperature and the hard disk working state information, so as to obtain more reliable fan rotation speed data and improve the heat dissipation efficiency of the fan.
[0170] The following describes the control logic of the hard disk fan when the master and slave control nodes are in the normal working state in the present application by combining the drawings.
[0171] Figure 4 The control logic of the hard disk fan when the master and slave control nodes are in the normal working state is shown in FIG. 1. As shown in FIG. 1, the control process of the hard disk fan when the master and slave control nodes are in the normal working state is described as follows. Figure 4
[0172] S401: Each hard disk and fan is numbered in advance to determine the correspondence between the hard disk and the fan according to the hard disk number and the fan number;
[0173] S402: In the normal working state, the arbitration controller controls the master node and the slave node to read the hard disk temperature and the hard disk working state information;
[0174] S403: According to the pre-set priority order, the master node controls the fan and adjusts the speed;
[0175] S404: Based on the pre-set multi-parameter weight model, the fan control parameter is obtained by calculating the hard disk temperature and the hard disk working state information corresponding to each fan according to a certain proportion;
[0176] S405: The fan control parameter is sent to the fan control module on the fan board;
[0177] S406: The fan control module controls the fan to rotate according to the fan control parameter.
[0178] Secondly, the following describes the control logic of the hard disk fan when the master and slave control nodes are both faulty in the present application by combining the drawings.
[0179] Figure 5 The control logic of the hard disk fan when the master and slave control nodes are both faulty is shown in FIG. 2. As shown in FIG. 2, the control process of the hard disk fan when the master and slave control nodes are both faulty is described as follows. Figure 5 As shown, the hard disk fan control logic of the application when both master and slave nodes fail is as follows:
[0180] S501: the fan control module sets the maximum response waiting time;
[0181] S502: if the maximum response waiting time is exceeded, the fan control module still does not receive relevant data from the master node or the slave node, and it is determined that both the master and slave nodes have failed;
[0182] S503: the fan control module sends a communication right request signal to the arbitration controller to obtain the communication right;
[0183] S504: the fan control module communicates with the hard disk backplane area to obtain hard disk temperature information and hard disk working state information;
[0184] S505: the fan control module calculates the fan control parameters according to the hard disk temperature information and the hard disk working state information to control the rotation of the fan.
[0185] Finally, the following describes the control method of the double-control high-density server hard disk fan of the application by combining the accompanying drawings.
[0186] Figure 6 The execution logic diagram of the control method of the double-control high-density server hard disk fan is shown in FIG. 1. Figure 6 As shown, the specific execution process of the control method of the double-control high-density server hard disk fan of the application is as follows:
[0187] S601: according to the hard disk number and the fan number, the correspondence between the hard disk and the fan is determined;
[0188] S602: it is determined whether the master node has a failure, if there is a failure, go to S603, otherwise go to S604;
[0189] S603: it is determined whether the master node sends a master node failure signal to the slave node, if yes, go to S605, otherwise go to S606;
[0190] S604: the master node obtains the fan control right, and the arbitration controller defaults that the master node communicates with the hard disk backplane area, and goes to S6012;
[0191] S605: it is determined whether the slave node has a failure, if there is a failure, go to S607, otherwise go to S608;
[0192] S606: it is determined whether the master node heartbeat light is normal, if the master node heartbeat light is normal, go to S604, otherwise go to S605;
[0193] S607: The slave node sends a slave node failure signal to the fan control module, and goes to S6010;
[0194] S608: The slave node obtains the fan control right, and sends a communication right request signal to the arbitration controller to obtain the corresponding communication right;
[0195] S609: The slave node reads the hard disk temperature information and hard disk working state information of each hard disk backboard area, and goes to S6013;
[0196] S6010: The fan control module sends a communication right request signal to the arbitration controller to obtain the communication right;
[0197] S6011: The fan control module reads the hard disk temperature information and hard disk working state information through the bidirectional binary synchronous serial bus to calculate the fan control parameter, and goes to S6015;
[0198] S6012: The master node obtains the hard disk temperature information and hard disk working state information of each hard disk backboard area;
[0199] S6013: The hard disk temperature information and hard disk working state information are input into a pre-constructed multi-parameter weight model, and corresponding calculation is performed according to a certain proportion to obtain the corresponding fan control parameter;
[0200] S6014: The fan control parameter is sent to the fan control module on the fan board;
[0201] S6015: The fan control module controls the fan to rotate according to the fan control parameter.
[0202] Through the description of the above embodiments, those skilled in the art can clearly understand that the method according to the above embodiments can be realized by means of software and the necessary general hardware platform, of course, it can also be realized by hardware, but in many cases the former is a better embodiment.
[0203] The embodiment of the application also provides a fan control device.
[0204] As shown in Figure 7 The fan control device 10 comprises a first acquisition module 100, a second acquisition module 200, a third acquisition module 300 and a control module 400.
[0205] The first acquisition module 100 is configured to acquire the backboard temperature and the hard disk working state information of the hard disk backboard area in the server through the master node when the master node meets the normal working requirement.
[0206] The second obtaining module 200 is configured to obtain the backboard temperature and the hard disk working state information of the hard disk backboard area through the slave node when the master node does not meet the normal working requirement.
[0207] The third obtaining module 300 is configured to obtain the backboard temperature and the hard disk working state information of the hard disk backboard area through the fan control module when neither the master node nor the slave node meets the normal working requirement.
[0208] The control module 400 is configured to control the working state of the fan based on the backboard temperature and the hard disk working state information.
[0209] Optionally, in an embodiment of the present application, the fan control device 10 further comprises a first judging module, a second judging module, a first determining module and a second determining module.
[0210] The first judging module is configured to obtain the log information of the master node and the slave node respectively, and to judge whether the heartbeat packet is lost between the master node and the slave node, wherein when the heartbeat packet is lost between the master node and the slave node, the number of times of continuous loss of the heartbeat packet between the master node and the slave node is determined.
[0211] The second judging module is configured to judge whether the communication timeout data is recorded in the log information of the master node and the slave node.
[0212] The first determining module is configured to determine that the master node and / or the slave node meets the normal working requirement if the communication timeout data is not recorded in the log information of the master node and / or the slave node, and the number of times of continuous loss of the heartbeat packet is less than a preset threshold.
[0213] The second determining module is configured to determine that the master node and / or the slave node does not meet the normal working requirement if the communication timeout data is recorded in the log information of the master node and / or the slave node, and the number of times of continuous loss of the heartbeat packet is greater than or equal to the preset threshold.
[0214] Optionally, in an embodiment of the present application, the third obtaining module 300 comprises a first redundancy control unit and a reading unit.
[0215] The first redundancy control unit is configured to send a slave node fault signal from the slave node to the fan control module through the slave node if neither the master node nor the slave node meets the normal working requirement, so as to control the fan control module to send a communication right request signal to a preset arbitration controller.
[0216] The reading unit is configured to control the fan control module to read the backboard temperature and the hard disk working state information of the hard disk backboard area through a preset bidirectional binary synchronous serial bus after the arbitration controller receives the communication right request signal.
[0217] Optionally, in one embodiment of the present application, the fan control device 10 further comprises a fourth acquisition module, an evaluation module and a priority switching module.
[0218] The fourth acquisition module is configured to acquire the priority parameters of the nodes preconfigured in the baseboard management controller of the server before acquiring the log information of the master node and the slave node respectively, so as to determine the master node and the slave node according to the priority parameters.
[0219] The evaluation module is configured to evaluate the health degree of the master node every preset time length to generate corresponding evaluation data, and calculate the priority score corresponding to the master node according to the evaluation data, wherein the evaluation data includes the communication response speed, the historical control accuracy and the current resource load of the master node.
[0220] The priority switching module is configured to judge whether the priority score is less than a preset priority switching score threshold, wherein if the priority score is less than the priority switching score threshold, the master node is switched to a new slave node, and the slave node with the highest priority score is switched to a new master node.
[0221] Optionally, in one embodiment of the present application, the control module 400 comprises an input unit, a compensation unit and an execution unit.
[0222] The input unit is configured to input the backboard temperature and the hard disk working state information into a pre-constructed multi-parameter weight model to output the target rotating speed of the fan corresponding to the hard disk backboard area.
[0223] The compensation unit is configured to determine the current read-write operation frequency of the server, and calculate a corresponding rotating speed compensation value according to the current read-write operation frequency, so as to correct the target rotating speed by the rotating speed compensation value to obtain the final fan rotating speed.
[0224] The execution unit is configured to determine the corresponding fan control parameter based on the final fan rotating speed, and send the fan control parameter to the fan control module, so as to control the fan corresponding to the hard disk backboard area to perform corresponding rotating operation according to the fan control parameter through the fan control module.
[0225] Optionally, in one embodiment of the present application, the fan control device 10 further comprises a discrimination module, a communication normal module and a communication fault module.
[0226] The discrimination module is configured to judge whether there is a communication link fault between the master node and the slave node and the fan control module when the master node and the slave node both meet the normal working requirements.
[0227] The communication normal module is configured to, if there is no communication link failure between the master node and the slave node and the fan control module, acquire, by the master node, backboard temperature and hard disk working state information of the hard disk backboard area to calculate corresponding fan control parameters.
[0228] The communication failure module is configured to, if there is a communication link failure between the master node and the slave node and the fan control module, acquire, by the fan control module, backboard temperature and hard disk working state information of the hard disk backboard area to control corresponding fans of the hard disk backboard area to rotate based on the backboard temperature and the hard disk working state information, if the fan control module does not receive the corresponding fan control parameters of the master node and the slave node within a preset response waiting time.
[0229] Optionally, in an embodiment of the present application, the second acquisition module 200 comprises a second redundancy control unit, a communication detection unit and a parameter calculation unit.
[0230] The second redundancy control unit is configured to, if the master node does not meet the normal working requirement, send a master node failure signal to the slave node by the master node to control the slave node to acquire backboard temperature and hard disk working state information of the hard disk backboard area.
[0231] The communication detection unit is configured to, if the master node does not meet the normal working requirement and the master node does not send the master node failure signal to the slave node, perform communication detection on a heartbeat signal of the master node by the slave node to obtain a corresponding detection result.
[0232] The parameter calculation unit is configured to, if the slave node determines that the master node does not meet the normal working requirement according to the detection result, control the slave node to acquire backboard temperature and hard disk working state information of the hard disk backboard area.
[0233] Optionally, in an embodiment of the present application, the fan control device 10 further comprises a partition module, a numbering module and a matching module.
[0234] The partition module is configured to partition the hard disk area of the server to obtain a plurality of hard disk backboard areas before acquiring log information of the master node and the slave node respectively.
[0235] The numbering module is configured to number the plurality of hard disk backboard areas and the hard disks in the hard disk backboard areas to obtain corresponding hard disk backboard area numbers and hard disk numbers, and number the fans in the server to obtain corresponding fan numbers.
[0236] The matching module is configured to determine a corresponding relationship between the fans and the hard disks based on the fan numbers, the hard disk backboard area numbers and the hard disk numbers, so that the fans cool the corresponding hard disks according to the corresponding relationship.
[0237] Optionally, in an embodiment of the present application, the partition module comprises: an acquisition unit, a construction unit and a solving unit.
[0238] The acquisition unit is configured to acquire airflow temperature distribution data corresponding to the hard disk region of the server, and to construct a target thermodynamic model based on the airflow temperature distribution data.
[0239] The construction unit is configured to obtain a plurality of current working parameters of the hard disk in the server based on the target thermodynamic model, and to construct a multi-objective optimization model based on the plurality of current working parameters.
[0240] The solving unit is configured to solve the multi-objective optimization model to obtain a partition result corresponding to the hard disk region of the server, and to determine a plurality of hard disk backboard regions based on the partition result.
[0241] Optionally, in an embodiment of the present application, the input unit comprises: a binding subunit, a smoothing subunit, a marking subunit and an information input subunit.
[0242] The binding subunit is configured to perform shunt acquisition and time alignment operations on the backboard temperature and the hard disk working state information, and to bind the backboard temperature and the hard disk working state information corresponding to the backboard temperature in the same acquisition cycle.
[0243] The smoothing subunit is configured to determine a smoothing window size corresponding to a preset dynamic window smoothing strategy based on the read-write frequency in the hard disk working state information, to perform dynamic window smoothing processing on the bound backboard temperature based on the smoothing window size, and to obtain a smoothed backboard temperature.
[0244] The marking subunit is configured to perform validity verification and abnormality marking on the hard disk working state information corresponding to the smoothed backboard temperature, to identify invalid state information of the bound hard disk working state information, and to perform missing value filling on the hard disk working state information corresponding to the invalid state information, to obtain hard disk working state filling information.
[0245] The information input subunit is configured to input the smoothed backboard temperature and the hard disk working state information or the hard disk working state filling information corresponding to the smoothed backboard temperature to the multi-parameter weight model.
[0246] Optionally, in an embodiment of the present application, the fan control device 10 further comprises: a fifth acquisition module, a sixth acquisition module and a comparison module.
[0247] The fifth acquisition module is configured to acquire actual running states of the fans corresponding to the plurality of hard disk backboard regions in real time, wherein the actual running states include actual rotating speeds and current characteristics of the drive circuits.
[0248] The sixth obtaining module is configured to determine the expected running state and the expected heat dissipation requirement of the fan according to the fan control parameter, and obtain the dynamic change trend of the backboard temperature.
[0249] The comparison module is configured to compare the actual running state with the expected running state, and determine whether the fan has a fault in combination with the dynamic change trend of the backboard temperature, wherein if the deviation between the actual running state and the expected running state exceeds a preset normal range, and the change of the backboard temperature does not meet the expected heat dissipation requirement, it is determined that the fan has a fault.
[0250] Optionally, in an embodiment of the present application, the comparison module comprises a requirement determining unit, a first fault analysis unit and a second fault analysis unit.
[0251] The requirement determining unit is configured to determine the fan heat dissipation requirement of the hard disk backboard region according to the expected heat dissipation requirement.
[0252] The first fault analysis unit is configured to, if the fan heat dissipation requirement is the fan enhanced heat dissipation requirement, but the actual rotating speed data in the actual running state of the fan does not increase as expected, and it is determined according to the dynamic change trend that the backboard temperature presents an upward trend, determine that the fan has a mechanical fault.
[0253] The second fault analysis unit is configured to, if the fan heat dissipation requirement is the fan maintained heat dissipation requirement, but the actual rotating speed data of the fan appears non-periodic fluctuation, and the backboard temperature presents synchronous abnormal fluctuation with the actual rotating speed data, determine that the fan has a drive circuit fault.
[0254] Optionally, in an embodiment of the present application, the fan control device 10 further comprises a reset module and a trigger module.
[0255] The reset module is configured to send a reset instruction to the fan control module after it is determined that the fan has a fault, so as to re-obtain the new backboard temperature and the new hard disk working state information of the hard disk backboard region, and calculate new fan control parameters according to the new backboard temperature and the new hard disk working state information.
[0256] The trigger module is configured to control the fan to rotate according to the new fan control parameters, so as to obtain the current running state and the new expected running state of the fan, calculate the state deviation between the current running state and the new expected running state, and determine whether the state deviation exceeds a preset normal range, wherein in the case where the state deviation exceeds the preset normal range, a fault reminding function is triggered, and the server is controlled to perform a corresponding fault reminding operation.
[0257] Optionally, in an embodiment of the present application, the trigger module comprises an identification generating unit, a trend generating unit and an alarm unit.
[0258] The identification generation unit is configured to determine a fault fan number corresponding to the fault fan, a hard disk backboard area fault number corresponding to the fault fan number, and a hard disk fault number, and generate a corresponding visual fault identification according to the fault fan number, the hard disk backboard area fault number, and the hard disk fault number.
[0259] The trend generation unit is configured to acquire a current running state of the fault fan and a dynamic change trend of the backboard temperature, and generate a corresponding fan running state record and backboard temperature change characteristic description according to the current running state and the dynamic change trend of the backboard temperature.
[0260] The alarm unit is configured to construct corresponding fault alarm information based on the fan running state record and the backboard temperature change characteristic description, display the visual fault identification on a local management interface of the server, and send the fault alarm information to a preset remote management terminal.
[0261] Optionally, in an embodiment of the present application, the fan control device 10 further comprises a deviation calculation module and a fine adjustment module.
[0262] The deviation calculation module is configured to calculate a corresponding running state deviation value based on an actual running state and an expected running state of the fan, and calculate a heat dissipation error based on a dynamic change trend of the backboard temperature and an expected heat dissipation requirement.
[0263] The fine adjustment module is configured to adjust corresponding backboard temperature and hard disk working state information according to the running state deviation value and the heat dissipation error, and fine tune a pre-constructed multi-parameter weight model through the adjusted backboard temperature and hard disk working state information.
[0264] Optionally, in an embodiment of the present application, the compensation unit comprises a recording subunit, a load coefficient calculation subunit, a temperature baseline acquisition subunit, and a dynamic correction subunit.
[0265] The recording subunit is configured to acquire a current read-write operation frequency of the server in real time, distinguish a read operation frequency and a write operation frequency in the current read-write operation frequency, and record a continuous execution time length and interval characteristics of the read operation frequency and the write operation frequency.
[0266] The load coefficient calculation subunit is configured to calculate a frequency ratio of the read operation frequency and the write operation frequency, and calculate a read-write operation comprehensive load coefficient according to the frequency ratio and the continuous execution time length.
[0267] The temperature baseline acquisition subunit is configured to acquire a current temperature baseline of a hard disk of the server, determine a temperature sensitive factor according to the current temperature baseline, multiply the read-write operation comprehensive load coefficient and the temperature sensitive factor to obtain a basic compensation value.
[0268] The dynamic correction subunit is configured to obtain a new current read-write operation frequency corresponding to the server, calculate a new frequency ratio value according to the new current read-write operation frequency, and dynamically correct the basic compensation value by using the new frequency ratio value to obtain the rotation speed compensation value.
[0269] The features of the embodiments of the fan control device can be understood by referring to the related descriptions of the embodiments of the fan control method, which will not be repeated here.
[0270] The embodiments of the present application also provide an electronic device, which comprises a memory and a processor, the memory stores a computer program, and the processor is configured to run the computer program to perform the steps in any of the above-mentioned embodiments of the fan control method.
[0271] The embodiments of the present application also provide a non-volatile computer readable storage medium, which stores a computer program, and the computer program is configured to perform the steps in any of the above-mentioned embodiments of the fan control method when running.
[0272] In an example embodiment, the above-mentioned non-volatile computer readable storage medium can include, but is not limited to, a U disk, a read-only memory (ROM), a random access memory (RAM), a mobile hard disk, a magnetic disk or an optical disk, and various media that can store computer programs.
[0273] The embodiments of the present application also provide a computer program product, which comprises a computer program, and the computer program is executed by a processor to implement the steps in any of the above-mentioned embodiments of the fan control method.
[0274] The embodiments of the present application also provide another computer program product, which comprises a non-volatile computer readable storage medium, and the non-volatile computer readable storage medium stores a computer program, and the computer program is executed by a processor to implement the steps in any of the above-mentioned embodiments of the fan control method.
[0275] The skilled person can further realize that the units and algorithm steps of the examples described in conjunction with the embodiments disclosed herein can be realized by electronic hardware, computer software or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the components and steps of the examples have been described in general terms in the above description. Whether the functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. The skilled person can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of the present application.
[0276] The above describes in detail a fan control method, device, equipment and medium provided by the present application. The principles and implementation manners of the present application are described by applying specific examples, and the above description of the examples is only used to help understand the method of the present application and its core idea. It should be pointed out that, for ordinary skilled persons in the technical field, some improvements and modifications can be made to the present application without departing from the principles of the present application, and these improvements and modifications also fall within the protection scope of the claims of the present application.
Claims
1. A fan control method applied to a server, characterized by, The server comprises a master node and at least one slave node, and comprises the following steps: In the case that the master node meets normal working requirements, the backboard temperature and hard disk working state information of a hard disk backboard area in the server are acquired through the master node; In the case that the master node does not meet normal working requirements, the backboard temperature and hard disk working state information of the hard disk backboard area are acquired through a slave node; In the case that neither the master node nor the slave node meets normal working requirements, the backboard temperature and hard disk working state information of the hard disk backboard area are acquired through a fan control module; The working state of a fan is controlled based on the backboard temperature and the hard disk working state information; The working state of the fan is controlled based on the backboard temperature and the hard disk working state information, comprising: The backboard temperature and the hard disk working state information are input into a pre-constructed multi-parameter weight model to output the target rotating speed of the fan corresponding to the hard disk backboard area; The current read-write operation frequency of the server is determined, and a corresponding rotating speed compensation value is calculated according to the current read-write operation frequency, so that the target rotating speed is corrected by the rotating speed compensation value to obtain the final fan rotating speed; Based on the final fan rotating speed, corresponding fan control parameters are determined, and the fan control parameters are sent to the fan control module, so that the fan corresponding to the hard disk backboard area performs corresponding rotating operation according to the fan control parameters through the fan control module.
2. The fan control method according to claim 1, characterized by, Further comprising: The log information of the master node and the slave node is acquired respectively, and it is judged whether the heartbeat packet is lost between the master node and the slave node, wherein when the heartbeat packet is lost between the master node and the slave node, the number of consecutive times of loss of the heartbeat packet between the master node and the slave node is determined; It is judged whether the communication timeout data is recorded in the log information of the master node and the slave node; If the communication timeout data is not recorded in the log information of the master node and / or the slave node, and the number of consecutive times of loss of the heartbeat packet is less than a preset threshold, it is determined that the master node and / or the slave node meets the normal working requirements; If the communication timeout data is recorded in the log information of the master node and / or the slave node, and the number of consecutive times of loss of the heartbeat packet is greater than or equal to the preset threshold, it is determined that the master node and / or the slave node does not meet the normal working requirements.
3. The fan control method of claim 1, wherein In the case that neither the master node nor the slave node meets normal working requirements, the backboard temperature and hard disk working state information of the hard disk backboard area are acquired through a fan control module, comprising: If the master node and the slave node do not meet the normal working requirements, a slave node failure signal is sent to the fan control module through the slave node to control the fan control module to send a communication right request signal to a preset arbitration controller; After the arbitration controller receives the communication right request signal, the fan control module is controlled to read the backboard temperature and hard disk working state information of the hard disk backboard area through a preset bidirectional binary synchronous serial bus.
4. The fan control method according to claim 2, characterized by, Before acquiring the log information of the master node and the slave node respectively, further comprising: Acquiring the priority parameters of each node pre-configured in the baseboard management controller of the server to determine the master node and the slave node according to the priority parameters; Every preset time length, the health degree of the master node is evaluated to generate corresponding evaluation data, and the priority score corresponding to the master node is calculated according to the evaluation data, wherein the evaluation data includes the communication response speed, historical control accuracy and current resource load of the master node; Determine whether the priority score is less than a preset priority switching score threshold, wherein if the priority score is less than the priority switching score threshold, the master node is switched to a new slave node, and the slave node with the highest priority score is switched to a new master node.
5. The fan control method of claim 1, wherein Further comprising: When the master node and the slave node both meet the normal working requirements, it is judged whether there is a communication link fault between the master node, the slave node and the fan control module; If there is no communication link fault between the master node, the slave node and the fan control module, the backboard temperature and hard disk working state information of the hard disk backboard area are acquired through the master node to calculate the corresponding fan control parameters; If there is a communication link fault between the master node, the slave node and the fan control module, so that the fan control module does not receive the fan control parameters corresponding to the master node and the slave node within a preset response waiting time, the backboard temperature and hard disk working state information of the hard disk backboard area are acquired through the fan control module to control the rotation of the fan corresponding to the hard disk backboard area based on the backboard temperature and the hard disk working state information.
6. The fan control method of claim 1, wherein The backboard temperature and hard disk working state information of the hard disk backboard area are acquired through the slave node when the master node does not meet the normal working requirements, comprising: If the master node does not meet the normal working requirements, the master node failure signal is sent to the slave node through the master node to control the slave node to acquire the backboard temperature and hard disk working state information of the hard disk backboard area; If the master node does not meet the normal working requirements, and the master node does not send the master node failure signal to the slave node, the heartbeat signal of the master node is detected through the slave node to obtain the corresponding detection result; If the slave node determines that the master node does not meet the normal working requirements according to the detection result, the slave node is controlled to acquire the backboard temperature and hard disk working state information of the hard disk backboard area.
7. The fan control method according to claim 2, wherein Before acquiring the log information of the master node and the slave node respectively, further comprising: Partitioning a hard disk area of the server to obtain a plurality of hard disk backboard areas; Numbering the plurality of hard disk backboard areas and hard disks in the hard disk backboard areas to obtain corresponding hard disk backboard area numbers and hard disk numbers, and numbering fans in the server to obtain corresponding fan numbers; Based on the fan numbers, the hard disk backboard area numbers and the hard disk numbers, determining a corresponding relationship between the fans and the hard disks, so that the fans cool the corresponding hard disks according to the corresponding relationship.
8. The fan control method of claim 7, wherein, The partitioning of the hard disk area of the server to obtain a plurality of hard disk backboard areas comprises: Collecting airflow temperature distribution data corresponding to the hard disk area of the server to construct a target thermodynamic model according to the airflow temperature distribution data; Based on the target thermodynamic model, obtaining a plurality of current working parameters of the hard disks in the server, and constructing a multi-objective optimization model through the plurality of current working parameters; Solving the multi-objective optimization model to obtain a partitioning result corresponding to the hard disk area of the server, and determining the plurality of hard disk backboard areas according to the partitioning result.
9. The fan control method of claim 1, wherein, The inputting of the backboard temperature and the hard disk working state information into the pre-constructed multi-parameter weight model comprises: Collecting and time-aligning the backboard temperature and the hard disk working state information, and binding the backboard temperature and the hard disk working state information corresponding to the backboard temperature in the same collection cycle; Determining a smoothing window size corresponding to a preset dynamic window smoothing strategy according to the read-write frequency in the hard disk working state information, and performing dynamic window smoothing processing on the bound backboard temperature based on the smoothing window size to obtain a smoothed backboard temperature; Performing validity verification and abnormality marking on the hard disk working state information corresponding to the smoothed backboard temperature to identify invalid state information of the bound hard disk working state information, and performing missing value filling on the hard disk working state information corresponding to the invalid state information to obtain hard disk working state filling information; Inputting the smoothed backboard temperature and the hard disk working state information or the hard disk working state filling information corresponding to the smoothed backboard temperature into the multi-parameter weight model.
10. The fan control method of claim 1, wherein Further comprising: Real-time acquisition of actual running states of fans corresponding to a plurality of hard disk backboard areas, wherein the actual running states include actual rotating speeds and current characteristics of drive circuits; Determination of expected running states and expected cooling requirements of the fans corresponding to the fan control parameters, and acquisition of a dynamic change trend of the backboard temperature; Comparison of the actual running states and the expected running states, and judgment of whether the fans have faults in combination with the dynamic change trend of the backboard temperature, wherein if a deviation between the actual running states and the expected running states exceeds a preset normal range, and the change of the backboard temperature does not conform to the expected cooling requirements, it is determined that the fans have faults.
11. The fan control method of claim 10, wherein, If the deviation of the actual running state and the expected running state exceeds a preset normal range, and the backboard temperature change does not meet the expected heat dissipation requirement, it is determined that the fan has a fault, including: Determining the fan heat dissipation requirement corresponding to the hard disk backboard area according to the expected heat dissipation requirement; If the fan heat dissipation requirement is a fan enhanced heat dissipation requirement, but the actual rotating speed data in the actual running state of the fan does not increase according to the expected running state, and it is determined that the backboard temperature presents an upward trend according to the dynamic change trend, it is determined that the fan has a mechanical fault; If the fan heat dissipation requirement is a fan maintained heat dissipation requirement, but the actual rotating speed data of the fan appears non-periodic fluctuation, and the backboard temperature presents synchronous abnormal fluctuation with the actual rotating speed data, it is determined that the fan has a drive circuit fault.
12. The fan control method of claim 10, wherein, After it is determined that the fan has a fault, further including: Sending a reset instruction to the fan control module to reacquire new backboard temperature and new hard disk working state information corresponding to the hard disk backboard area, and calculating new fan control parameters according to the new backboard temperature and the new hard disk working state information; Controlling the fan to rotate according to the new fan control parameters to obtain a current running state and a new expected running state corresponding to the fan, and calculating a state deviation between the current running state and the new expected running state, and determining whether the state deviation exceeds the preset normal range, wherein, in the case that the state deviation exceeds the preset normal range, a fault reminding function is triggered, and the server is controlled to perform corresponding fault reminding operation.
13. The fan control method of claim 12, wherein, In the case that the state deviation exceeds the preset normal range, the fault reminding function is triggered, and the server is controlled to perform corresponding fault reminding operation, including: Determining a fault fan number corresponding to the fault fan, and a hard disk backboard area fault number and a hard disk fault number corresponding to the fault fan number, and generating a corresponding visual fault identifier according to the fault fan number, the hard disk backboard area fault number and the hard disk fault number; Obtaining a current running state of the fault fan and a dynamic change trend of the backboard temperature, and generating a corresponding fan running state record and backboard temperature change characteristic description according to the current running state and the dynamic change trend of the backboard temperature; Based on the fan running state record and the backboard temperature change characteristic description, a corresponding fault alarm information is constructed, and the visual fault identifier is displayed on a local management interface of the server, and the fault alarm information is sent to a preset remote management terminal.
14. The fan control method of claim 10, wherein, Further including: Based on the actual running state of the fan and the expected running state, a corresponding running state deviation value is calculated, and based on the dynamic change trend of the backboard temperature and the expected heat dissipation requirement, a heat dissipation error is calculated; The backboard temperature and the hard disk working state information are adjusted according to the running state deviation value and the heat dissipation error, and the multi-parameter weight model constructed in advance is fine-tuned through the adjusted backboard temperature and hard disk working state information.
15. The fan control method of claim 1, wherein, The calculating of the corresponding rotation speed compensation value according to the current read-write operation frequency comprises: Real-time collection of the current read-write operation frequency of the server, and distinguishing the read operation frequency and the write operation frequency in the current read-write operation frequency, and recording the continuous execution time length and interval characteristics of the read operation frequency and the write operation frequency; Calculation of the frequency ratio of the read operation frequency and the write operation frequency, so as to calculate the read-write operation comprehensive load coefficient according to the frequency ratio and the continuous execution time length; Acquisition of the current temperature baseline of the hard disk of the server, and determination of the temperature sensitive factor according to the current temperature baseline, multiplication of the read-write operation comprehensive load coefficient and the temperature sensitive factor to obtain a basic compensation value; Acquisition of the corresponding new current read-write operation frequency of the server, so as to calculate the corresponding new frequency ratio according to the new current read-write operation frequency, and dynamically correct the basic compensation value through the new frequency ratio to obtain the rotation speed compensation value.
16. A fan control device applied to a server, characterized by, The server comprises a master node and at least one slave node, comprising: The first acquisition module is used to acquire the backboard temperature and hard disk working state information of the hard disk backboard area in the server through the master node when the master node meets the normal working requirements; The second acquisition module is used to acquire the backboard temperature and hard disk working state information of the hard disk backboard area through the slave node when the master node does not meet the normal working requirements; The third acquisition module is used to acquire the backboard temperature and hard disk working state information of the hard disk backboard area through the fan control module when neither the master node nor the slave node meets the normal working requirements; The control module is used to control the working state of the fan based on the backboard temperature and the hard disk working state information; The control of the working state of the fan based on the backboard temperature and the hard disk working state information comprises: Inputting the backboard temperature and the hard disk working state information into a pre-constructed multi-parameter weight model to output the target rotation speed of the fan corresponding to the hard disk backboard area; Determination of the current read-write operation frequency of the server, and calculation of the corresponding rotation speed compensation value according to the current read-write operation frequency, so as to correct the target rotation speed through the rotation speed compensation value to obtain the final fan rotation speed; Based on the final fan rotation speed, the corresponding fan control parameters are determined, and the fan control parameters are sent to the fan control module to control the fan corresponding to the hard disk backboard area to perform corresponding rotation operation according to the fan control parameters.
17. An electronic device, comprising: Comprise: Memory for storing computer programs; The processor is used to execute the computer program to realize the steps of the fan control method in any one of claims 1 to 15.
18. A computer-readable storage medium, characterized in that, The computer readable storage medium stores a computer program, wherein the computer program is executed by the processor to realize the steps of the fan control method in any one of claims 1 to 15. The computer readable storage medium stores a computer program, wherein the computer program is executed by the processor to realize the steps of the fan control method in any one of claims 1 to 15.
19. A computer program product comprising a computer program, characterized in that, The computer program, when executed by a processor, implements the steps of the fan control method according to any one of claims 1 to 15.
Citation Information
Patent Citations
Hard disk heat dissipation coordination control method, system and equipment, medium and storage server
CN116088652A
Fan rotating speed adjusting method and server
CN118601927A