Server hard disk monitoring device
Through the bidirectional data interaction between the substrate management controller and the hard disk box controller, multiple designated components in the hard disk box are monitored in real time, solving the limitations of traditional hard disk monitoring, and realizing the intelligent management and stability of components in the hard disk box.
Patent Information
- Application Number
- CN202521474966.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Utility models(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-15
- Publication Date
- 2025-08-22
- Estimated Expiration
- 2035-07-15
AI Technical Summary
Traditional server hard disk expansion methods cannot fully monitor various components in the hard disk box, resulting in low hard disk reliability and lack of intelligent management and timely response capabilities.
The two-way data interaction between the substrate management controller and the hard disk box controller is adopted to monitor multiple designated components in the hard disk box in real time, generate status control signals through the hard disk box controller, adjust the operating status and parameters of the components, and combine it with the control module for intelligent management.
It realizes all-round intelligent monitoring of various components in the hard disk box, improves the reliability and stability of the server hard disk, and enhances the monitoring capabilities of the storage system.
Smart Images

Figure CN223260174U_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of computer equipment, and in particular to a server hard disk monitoring device. Background Art
[0002] As data storage needs grow, businesses and individuals need to store more and more data. Traditional server hard drives are no longer able to meet this growing storage demand. To address this issue, related technologies typically rely on external hard drive enclosures, which increase storage capacity through simple physical connections and manual management.
[0003] However, the above-mentioned method of directly connecting the expansion hard disk using a physical interface can only adapt to specific types of server hard disks, and can only monitor simple parameters such as the temperature of the hard disk box and the expansion hard disk. It cannot comprehensively monitor the various components of the hard disk in the hard disk box, and cannot respond to hard disk abnormalities in a timely manner, thereby causing technical problems of low reliability of server hard disks. Utility Model Content
[0004] The present application provides a server hard disk monitoring device to at least solve the problem of low hard disk reliability caused by overly limited monitoring of server hard disks in related technologies.
[0005] The present application also provides a server hard disk monitoring device, comprising: a baseboard management controller, a hard disk box controller and a control module, wherein the baseboard management controller is connected to the hard disk box controller, and the baseboard management controller and the hard disk box controller perform two-way data exchange; the baseboard management controller is used to display information when monitoring the status of the server hard disk in the hard disk box; the hard disk box controller is used to obtain status data of multiple specified components inside the server hard disk, and generate status control signals based on the status data of the multiple specified components, wherein the status control signals are used to control the operating status and component parameters of the multiple specified components; the control module is used to adjust the operating status and component parameters of the multiple specified components in response to the received status control signals.
[0006] Through the two-way data exchange between the baseboard management controller and the hard disk enclosure controller in the embodiment of the present application, multiple designated components of each server hard disk in the hard disk enclosure can be intelligently managed and maintained, ensuring real-time monitoring of the status data of each designated component and improving the comprehensiveness and integrity of the monitoring information. At the same time, the hard disk enclosure controller collects and analyzes the status data of key components in the hard disk enclosure and generates status control signals, accurately regulating the operating status and parameters of each designated component, enhancing the monitoring capabilities of the server storage system and improving the reliability and stability of the server hard disk. BRIEF DESCRIPTION OF THE DRAWINGS
[0007] In order to more clearly illustrate the embodiments of the present application, the following is a brief introduction to the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0008] Figure 1 A structural diagram of an optional server hard disk monitoring device provided in an embodiment of the present application.
[0009] Figure 2 An overall schematic diagram of an optional server hard disk monitoring device provided in an embodiment of the present application.
[0010] Figure 3 A hardware design diagram of an optional hard disk monitoring system provided in an embodiment of the present application.
[0011] Figure 4 A schematic diagram of an optional connection between a hard disk backplane and a hard disk provided in an embodiment of the present application.
[0012] Figure 5 A schematic diagram of the connection between an optional hard disk enclosure and a main BMC provided in an embodiment of the present application. DETAILED DESCRIPTION
[0013] The following will be combined with the accompanying drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of them. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.
[0014] It should be noted that, in the description of this application, the terms "comprises," "includes," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or device. The terms "first," "second," etc., in this application are used to distinguish similar objects, and are not used to describe a particular order or sequence.
[0015] Against the backdrop of increasing server performance and storage requirements, especially for high-performance servers, higher requirements are placed on the capacity of their extended hard drives. In order to ensure the stability and reliability of the server's extended hard drive, it is necessary to monitor the status data of each component of the extended server hard drive in real time to ensure that the performance of the extended server hard drive remains normal. Traditional server hard drive expansion usually relies on external hard drive boxes, in which storage capacity is increased through simple physical connections and manual management. Implementation methods mainly include directly connecting the hard drive to the server motherboard using a physical interface, or providing additional hard drive slots through an external hard drive box. However, these methods have obvious limitations in monitoring the server hard drive in the hard drive box, and lack intelligent management and monitoring methods.
[0016] In order to solve the above problems, the embodiment of the present application proposes a server hard disk monitoring device, which is designed to detect abnormal conditions of the hard disk system in a timely manner and take corresponding measures to deal with them through the coordinated work of the server main BMC (Baseboard Management Controller) and the hard disk enclosure controller when abnormal conditions occur in different types of server hard disks in the hard disk enclosure. At the same time, the stability and reliability of the server hard disk are improved. The structure of the server hard disk monitoring device is as follows: Figure 1 As shown, Figure 1 As shown, the server hard disk monitoring device includes: a server main BMC, a hard disk box controller in the hard disk box, a control module and a server hard disk.
[0017] The baseboard management controller is connected to the hard disk enclosure controller (also known as the slave BMC) and can achieve two-way communication. For example, the master BMC establishes a connection with the hard disk enclosure controller via a high-speed communication bus (such as I2C), and the two continuously exchange data to ensure real-time updates of monitoring information. This connection is based on the IPMI (Intelligent Platform Management Interface) protocol, allowing the master BMC to send control commands to the hard disk enclosure controller and receive status feedback from the hard disk enclosure controller.
[0018] As a key component of the server, the baseboard management controller (BMC) communicates bidirectionally with the hard drive enclosure controller to monitor real-time health, temperature, RAID level, and other key information for all hard drives within the enclosure, displaying it intuitively on the management interface. If an abnormality is detected, such as excessive temperature or hard drive failure, the main BMC can immediately trigger an alarm, notifying the system administrator for timely intervention.
[0019] The hard drive enclosure includes but is not limited to the hard drive enclosure controller, hardware backplane, temperature sensor, RAID card (Redundant Array of Independent Disks, disk array card or independent disk redundant array card), cooling device, and server hard drive, etc.
[0020] Hard drive enclosures typically feature a rectangular structure with a high-strength aluminum alloy casing, offering excellent heat dissipation and interference resistance. The overall dimensions are optimized based on the number and layout of hard drives, ensuring optimal assembly of internal components while facilitating portability and deployment. The front panel of a hard drive enclosure typically houses the hard drive indicator lights, a power switch, and status indicators.
[0021] Among them, each hard disk corresponds to an indicator light, which intuitively displays the working status of the hard disk through different colors (such as green for normal operation and red for fault); the power switch can be used, but is not limited to, to control the power on and off of the hard disk box, and adopts a large-size button design for easy operation; the status indicator light can be used, but is not limited to, to display the overall working status of the hard disk box, such as power status, communication status, etc.
[0022] The rear panel of the hard drive enclosure houses various interfaces, including Ethernet, a control port, SATA / SAS (SATA is a computer bus interface, SAS is a high-performance serial interface), SMbus (System Management Bus, a two-wire interface), I2C (Inter-Integrated Circuit, a serial communication protocol), and a power connector. The interfaces are modular in design for easy insertion and maintenance.
[0023] The left and right panels of the hard drive enclosure are equipped with large heat dissipation holes, which adopt a honeycomb design to increase the heat dissipation area and improve heat dissipation efficiency. At the same time, dust screens are installed inside the heat dissipation holes to prevent dust from entering the hard drive enclosure.
[0024] The top cover of the hard drive enclosure is equipped with a handle for easy movement and transport. The handle is ergonomically designed for easy grip. The base of the hard drive enclosure features rubber feet that provide excellent anti-slip and shock absorption, protecting the internal components from vibration and impact.
[0025] The hard drive enclosure controller uses a dedicated interface to obtain status data for each component of the server's hard drive, as well as status data for components connected to the hard drive, such as the hard drive backplane, RAID card, temperature sensor, and fan. The hard drive enclosure controller not only reads hard drive status data but also intelligently manages hard drive power by controlling the EFUSE switch.
[0026] The hard drive enclosure controller, through close communication with the hard drive backplane, RAID card, and temperature sensor, collects real-time data on the hard drive's operating status, including but not limited to the hard drive's in-place status, RAID card operating mode, hard drive temperature, fault indicator information, hard drive type, interface type, manufacturer information, location information, and capacity information. Based on this data, the hard drive enclosure controller uses a built-in algorithm to analyze the hard drive's health and generates corresponding status control signals based on preset safety thresholds.
[0027] The control module, as the execution unit, directly responds to status control signals from the hard drive enclosure controller and adjusts the hard drive's operating status and parameters. For example, based on information from the temperature sensor, the control module adjusts the fan speed to optimize cooling.
[0028] When receiving status control signals from the hard drive enclosure controller, the control module quickly responds and adjusts the operating status and parameters of the hard drive backplane, RAID card, and hard drive. For example, if the hard drive temperature is detected to be outside the normal range, the control module controls the fan speed by adjusting the PWM (Pulse Width Modulation) signal to achieve dynamic cooling and maintain the optimal operating temperature for the hard drive.
[0029] In order to understand the above server hard disk monitoring device more clearly, Figure 2 The overall schematic diagram shown further describes it.
[0030] like Figure 2 As shown in the figure, the hard drive monitoring system primarily consists of a master BMC, a hard drive enclosure controller (also known as a slave BMC), a hard drive backplane, a RAID card, hard drives, temperature sensors, cooling devices (such as fans), and a communication bus. The master BMC, as the core management unit of the entire system, is responsible for receiving and processing data from the hard drive enclosure controller and issuing control commands. The hard drive enclosure controller, acting as the local monitoring and control center, directly interacts with the hard drive backplane, RAID card, temperature sensors, and cooling devices. The following describes the connections between these components in detail.
[0031] (1) Main BMC and hard disk enclosure controller: connected via the high-speed communication bus I2C, enabling bidirectional data transmission. The main BMC can send control commands to the hard disk enclosure controller, such as hard disk power on and off instructions; the hard disk enclosure controller then feeds back the collected hard disk information, RAID card status, temperature data, etc. to the main BMC.
[0032] (2) Hard disk box controller and hard disk backplane: connected through I2C, the hard disk box controller can control the power switch on the hard disk backplane according to the instructions of the main BMC or local policy to realize the power on and off operation of the hard disk.
[0033] (3) Hard disk box controller and RAID card: connected through the dedicated management interface SMbus, the hard disk box controller can read the status information of the RAID card, such as RAID level, disk array health, read and write performance, etc., and perform necessary configuration and management on the RAID card.
[0034] (4) Hard disk enclosure controller and temperature sensor: The temperature sensor is installed on the hard disk surface or at a key position inside the hard disk enclosure, and is connected to the hard disk enclosure controller through an analog or digital interface to transmit the hard disk temperature data to the hard disk enclosure controller in real time.
[0035] (5) Hard disk box controller and heat dissipation device: connected through PWM (pulse width modulation) signal line, the hard disk box controller adjusts the duty cycle of the PWM signal according to the hard disk temperature, thereby controlling the speed of the cooling fan and realizing heat dissipation control of the hard disk box.
[0036] A baseboard management controller (BMC) is a management controller specifically designed for servers and other hardware devices. It's part of the Intelligent Platform Management Interface (IPMI), allowing system administrators to monitor and control the server's hardware status through a separate management network. Typically integrated into the server's motherboard, the BMC has its own processor, memory, and storage, and runs independent management software.
[0037] The hard drive box can be used for, but not limited to, external expansion and data reading of the server hard drive, high-performance external storage of the server hard drive, and flexible storage location of the server hard drive.
[0038] Obviously, it should be noted that the hard disk itself also has a heat dissipation device, such as a fan. The hard disk box controller determines whether to adjust the operating parameters of its heat dissipation device by collecting the temperature of the hard disk itself.
[0039] In an alternative embodiment, assume a server has an external hard drive enclosure containing multiple hard drives. The main BMC communicates with the enclosure controller via the I2C bus, sending control commands, such as hard drive power on / off requests. Upon receiving the commands, the enclosure controller interacts with the CPLD (Complex Programmable Logic Device) on the hard drive backplane via I2C to control the EFUSE switch, enabling refined hard drive power management. This ensures that the drives are powered on and off at appropriate times, reducing energy consumption and extending drive life.
[0040] Among them, the EFUSE switch (electronic fuse) is an integrated circuit based on semiconductor technology, used for circuit protection and power management. It controls the on and off of the circuit through electronic signals, replacing the functions of traditional mechanical fuses and relays.
[0041] At the same time, the enclosure controller periodically (for example, every three seconds) communicates with the temperature sensor on the hard drive surface via I2C to obtain temperature data. For NVMe drives, the enclosure controller directly reads the drive's internal sensor data via the I2C bus; for non-NVMe drives, the enclosure controller obtains drive temperature information indirectly through the RAID card. If the temperature exceeds a preset threshold, the enclosure controller controls the fan inside the enclosure via PWM signals, adjusting the speed to accommodate the temperature change and maintain a stable operating environment for the drive.
[0042] NVMe is a storage interface protocol based on the PCI Express (PCIe) bus, designed specifically for non-volatile memory (such as flash memory). It fully leverages PCIe's high-speed transmission characteristics to achieve higher performance and lower latency. NVMe hard drives support multiple queues and high concurrency, significantly improving data transfer speeds. Their random read and write performance far exceeds that of traditional hard drives. Non-NVMe hard drives, which do not use the NVMe protocol, are more suitable for various application scenarios, including personal computers, small servers, and enterprise-level storage.
[0043] In terms of RAID card monitoring, the hard drive enclosure controller uses the SMbus interface to read the RAID card's status register to obtain information such as the RAID level, disk array health, and read and write error counts to ensure the normal operation of the RAID system. If an anomaly is found, the hard drive enclosure controller promptly reports it to the main BMC to assist in fault location and recovery.
[0044] To sum up, the server hard disk monitoring device provided in the embodiment of the present application realizes all-round and intelligent monitoring of the hard disks in the server hard disk box through the efficient cooperation of the main BMC and the hard disk box controller, and the precise regulation of the hard disk box controller and the control module, thereby effectively improving the storage reliability and management efficiency of the server.
[0045] Through the embodiments of the present application, bidirectional data exchange between the baseboard management controller and the hard disk enclosure controller enables intelligent management and maintenance of multiple designated components of each server hard disk in the hard disk enclosure, ensuring real-time monitoring of the status data of each designated component and improving the comprehensiveness and integrity of the monitoring information. At the same time, the hard disk enclosure controller collects and analyzes the status data of key components in the hard disk enclosure and generates status control signals, accurately regulating the operating status and parameters of each designated component, enhancing the monitoring capabilities of the server storage system and improving the reliability and stability of the server hard disk.
[0046] In an exemplary embodiment, Figure 1As shown, the server hard disk monitoring device also includes a hard disk backplane, and the hard disk box controller is connected to the hard disk backplane through a multi-way analog switch; a disk array controller card, and the disk array controller card is connected to the hard disk box controller through a bus interface, and is used to perform read and write operations on the server hard disk in the hard disk box.
[0047] like Figure 3 As shown, the hard disk box controller is connected to at least one hard disk backplane through a multi-way analog switch, specifically connected to the I2C switch PCA9548 through the I2C6 pin of the hard disk box controller, and connected to the PCA9555 (input and output expander) through the I2C10 pin of the hard disk box controller. Different GPIOs (General Purpose Input / Output) of PCA9555 correspond to the in-place signals of the hard disk backplane, which helps to detect the connection status of the hard disk backplane.
[0048] Different hard drive backplanes are linked to different channels of the I2C switch PCA9548 module. The I2C switch PCA9548 module plays the role of signal switching and distribution, enabling the hard drive box controller to communicate with multiple hard drive backplanes.
[0049] It should be noted that a hard drive box includes multiple hard drive backplanes, and the multiple hard drive backplanes are connected to the hard drive box controller through the I2Cswitch PCA9548 module, receiving the control signal of the hard drive box controller and feeding back data to the hard drive box controller.
[0050] The above device also includes a process module for the hard disk backplane, which mainly includes (1) a hard disk backplane monitoring process: which mainly runs in the hard disk box controller and is responsible for interacting with the hard disk backplane to read data and control the hard disk backplane. For example, it can obtain information such as the temperature and voltage of the hard disk backplane, and can also perform operations such as resetting the backplane; (2) a sensor monitoring process: which interacts with the hard disk backplane monitoring process through the redis data (buffered data) library. It is mainly responsible for monitoring sensor data, which may be distributed on the hard disk backplane or the hard disk surface, such as temperature sensors, fan speed sensors, etc.; (3) a heat dissipation process: which interacts with the hard disk backplane monitoring process through the redis database. Based on the temperature and other information obtained from the backplane, the heat dissipation strategy is adjusted, for example, the fan speed is controlled to maintain a suitable temperature; (4) an alarm process: which also interacts with the backplane monitoring process through the redis database. When the backplane monitoring process detects abnormal data, such as excessive temperature or loss of the backplane in-position signal, the alarm process will trigger a corresponding alarm mechanism, such as sending an email or displaying an alarm message on the management interface.
[0051] The data exchange process between the hard drive backplane and other components includes the following: The hard drive backplane monitoring process communicates with the hard drive backplane via the I2C6PCA9548 connection to obtain hard drive backplane-related data, such as hardware status and sensor data. The hard drive backplane monitoring process stores the acquired data in the Redis database. The sensor monitoring process, cooling process, and alarm process read relevant data from the Redis database. The sensor monitoring process determines whether the sensor status is normal based on the read data. The cooling process adjusts the cooling strategy based on temperature and other data, and the alarm process triggers alarms based on abnormal data.
[0052] In addition, if Figure 3 As shown, the hard drive enclosure controller connects to the management chip of the RAID card (disk array controller card) via the SMBus interface to enable read and write operations on the server's hard drive data. The software implementation process includes: the hard drive enclosure controller periodically sends query commands to the RAID card via the IPMB protocol, reads the RAID card's status register, and obtains information such as the RAID level, disk array health, and read and write error counts. If a RAID card anomaly is detected, such as a disk failure or array rebuild, the hard drive enclosure controller promptly sends the relevant information to the main BMC and takes appropriate measures based on pre-set policies, such as issuing an alarm or automatically switching to a spare disk.
[0053] Through the connection mechanism and functional implementation of the hard disk enclosure controller, hard disk backplane, and disk array controller card described in this embodiment, and through the use of multi-way analog switches and bus interfaces, the hard disk enclosure controller can flexibly and efficiently monitor and manage server hard disks, ensuring a stable storage environment and efficient operation of the hard disk system. This brings substantial improvements to the server hard disk monitoring process, enhancing data storage security and management efficiency.
[0054] In an exemplary embodiment, the hard disk backplane further includes: a power control switch for controlling the power supply of the server hard disk in response to the switch closed state sent by the hard disk box controller; or for controlling the power supply of the server hard disk in response to the switch open state sent by the hard disk box controller.
[0055] like Figure 3As shown in the figure, taking a hard disk backplane F_BP0_Dev_J11 as an example, its internal hardware link includes the BP_Present pin. The status of the GPIO pin connected to it indicates the hard disk backplane presence signal. The CPLD_UFM pin connected to BP0_UFM indicates the CPLD (Complex Programmable Logic Device) controller program. The CPLD_Data pin connected to CPLD_Reg indicates the connection signal inside the hard disk backplane.
[0056] like Figure 4 As shown, a hard disk backplane is usually connected to multiple server hard disks through slots. Each hard disk power interface on the hard disk backplane is equipped with a power control switch (EFUSE switch). The hard disk box controller controls the power control switch to turn the hard disk power on and off.
[0057] In other words, the power control switch is a key executive element of the hard drive monitoring system in the embodiments of this application. Its function is to accurately control the power on and off of the server hard drive after receiving instructions from the hard drive enclosure controller. At the hardware level, the power control switch usually refers to an EFUSE switch or other form of solid-state relay. They can quickly respond to instructions from the hard drive enclosure controller and achieve instantaneous control of the hard drive power without physical contact.
[0058] Specifically, the EFUSE switch is in a low-impedance state under normal operating conditions, allowing current to flow, thereby powering the hard drive. When the hard drive enclosure controller decides to disconnect the power supply based on the monitored hard drive status, the EFUSE switch is set to a high-impedance state, preventing current from flowing, and the hard drive enters a power-off state.
[0059] The power control switch is designed with hardware safety in mind, ensuring stable power to the hard drives even in the event of a controller failure or communication interruption. For example, if the hard drive controller fails, the switch will enter a default safe state and will not randomly turn the hard drives on or off, preventing accidents.
[0060] The integration of the power control switch matches the communication design of the hard drive enclosure controller, making it easy to implement within existing server hardware architecture without requiring large-scale hardware modifications. Furthermore, by simply increasing the number of power control switches and optimizing the control strategy of the hard drive enclosure controller, the embodiments of this application can easily cope with the challenges posed by future increases in the number of hard drives.
[0061] The combination of the power control switch and the hard drive enclosure controller not only enables intelligent adjustment of the server hard drive power status, but also ensures a secure and stable data storage environment, thereby improving the reliability and efficiency of the entire server system. This refined power management provides new concepts and technical support for server hard drive monitoring and maintenance.
[0062] In an exemplary embodiment, the above-mentioned disk array controller card also includes: a status register, which is used to store the health status and read and write error record information of the disk array controller card; a data read and write management chip, which is connected to the hard disk box controller through the bus interface and is used to perform read and write operations on the status data of multiple specified components inside the server hard disk.
[0063] As described in the preceding embodiments, the hard drive enclosure controller connects to the data read / write management chip of the RAID card (disk array controller card) via the SMBus interface to implement data read / write operations. The software implementation involves the hard drive enclosure controller periodically sending query commands to the RAID card via the IPMB protocol, reading the RAID card's status register to obtain information such as the RAID level, disk array health, and read / write error counts.
[0064] If an abnormality is detected in the RAID card, such as disk failure, array reconstruction, etc., the hard disk box controller will promptly send relevant information to the main BMC and take corresponding measures according to the preset strategy, such as alarm, automatic switching of spare disks, etc.
[0065] The status register is a hardware component that stores information about the RAID card's health and read / write error logs. Located inside the RAID card, it records all important status information. By reading the status register, the hard drive enclosure controller can understand the RAID card's current operating status, including the RAID level, the health of the disk array, and error counts for read and write operations. This is crucial for identifying potential faults and planning maintenance in advance.
[0066] The data read / write management chip is one of the core components on a RAID card, responsible for handling read and write operations on the hard drives. It connects to the hard drive enclosure controller via a bus interface, allowing the enclosure controller to monitor and manage the RAID card's operating status. In addition to basic data exchange, the hard drive enclosure controller can also use the read / write management chip to obtain detailed information about the hard drives, such as temperature, capacity, and serial number. It can also perform advanced functions such as powering on and off the hard drives and dynamically adjusting the RAID configuration, enabling in-depth monitoring and intelligent management of the hard drive system.
[0067] The introduction of status registers and data read / write management chips enriches the capabilities of hard drive monitoring systems. Through the status registers, the hard drive enclosure controller can monitor the health and error history of the RAID card, ensuring data storage security and predicting maintenance needs. The data read / write management chip not only simplifies the data exchange process between the hard drive enclosure controller and the hard drive, but also empowers the hard drive enclosure controller with stronger management authority, enabling it to directly control the hard drive's power supply and adjust its configuration.
[0068] The aforementioned status register and data read / write management chip can monitor the hard drive system status in real time, intelligently respond to various possible fault conditions, optimize the hard drive's operating environment, and improve the stability and efficiency of data storage. This also reduces server hard drive maintenance costs and enhances system availability and data security.
[0069] In an exemplary embodiment, the hard disk enclosure controller includes a data collection module and a data transmission module, wherein the data collection module is used to obtain status data of multiple specified components inside the server hard disk according to a data reading instruction.
[0070] In this embodiment, the hard disk information that needs to be collected includes the hard disk model, serial number, capacity, power-on time, health status (including the status of the fault indicator light), interface type, manufacturer information, location information, and slot information.
[0071] The data collection module is used to read the hard drive information when the hard drive enclosure controller sends SCSI commands to the hard drive through the SATA / SAS interface. At the same time, it combines the monitoring data of the RAID card to conduct a comprehensive assessment of the overall status of the hard drive.
[0072] The data transmission module is used to transmit the hard disk information, RAID card status data, cooling device data, etc. obtained by the hard disk box controller to the main BMC, so that users can view the monitoring data directly on the server interface.
[0073] Specifically, the IPMIB communication protocol is used to package collected hard drive information, RAID card status, temperature data, and other information into fixed-format data packets, which are then sent to the master BMC via the communication bus. The specific transmission mechanism includes an acknowledgement mechanism and a retransmission mechanism to ensure data transmission reliability. After the hard drive enclosure controller BMC sends a data packet, it waits for an acknowledgement from the master BMC. If no acknowledgement is received within a specified time, the packet is resent.
[0074] The use of data collection and data transmission modules ensures that the hard drive monitoring system can acquire and transmit hard drive operating status data in real time, enabling the baseboard management controller to fully understand the health of the hard drive system, identify potential problems promptly, and take appropriate action. This data collection and transmission mechanism design demonstrates the real-time, intelligent, and reliable nature of the server hard drive monitoring system.
[0075] Through the data collection module, the hard drive enclosure controller can periodically scan the status of the hard drive's internal components, covering all aspects of hard drive operation, such as temperature, RAID configuration, and hard drive health. The efficient transmission capabilities of the data transmission module ensure that the BMC can receive this status information immediately, providing system administrators with a real-time monitoring interface for easy monitoring and management.
[0076] In practical applications, the introduction of a hard drive enclosure controller not only enables monitoring of hard drive status without requiring direct hardware access, facilitating data center operations and maintenance, but also enables the system to rapidly respond to hard drive anomalies, such as overtemperature or RAID failures, through real-time data transmission, taking timely action to prevent data loss.
[0077] In an exemplary embodiment, the above-mentioned hard disk box controller includes a first bus interface pin and a second bus interface pin; the first bus interface pin is connected to a multi-way analog switch, wherein the multi-way analog switch is used for signal switching and distribution; the second bus interface pin is connected to an input-output expander, wherein the input-output expander is used to detect the connection status of the hard disk backplane.
[0078] like Figure 3 As shown, the hard disk box controller is connected to at least one hard disk backplane through a multi-way analog switch, specifically through the I2C6 pin (first bus interface pin) of the hard disk box controller to the I2C switch PCA9548, and the I2C10 pin (second bus interface pin) of the hard disk box controller is connected to the PCA9555 (input and output expander). Different GPIOs (General Purpose Input / Output) of PCA9555 serve as the presence signals of the corresponding hard disk backplane, which helps to detect the connection status of the hard disk backplane.
[0079] Different hard drive backplanes are linked to different channels of the I2C switch PCA9548 module. The I2C switch PCA9548 module plays the role of signal switching and distribution, enabling the hard drive box controller to communicate with multiple hard drive backplanes.
[0080] In this embodiment, the multi-way analog switch performs signal switching and distribution. It receives signals from the first bus interface pin of the hard drive enclosure controller and directs them to specific hard drive backplanes or sensors as needed, thereby enabling flexible communication between the hard drive enclosure controller and multiple hard drive backplanes and sensors. The advantage of this design is that a single hard drive enclosure controller can simultaneously monitor and control multiple hard drive backplanes, significantly improving the efficiency and flexibility of system monitoring.
[0081] The I / O expander, connected to the second bus interface pin of the hard drive enclosure controller, detects the presence of hard drive backplanes, ensuring that all backplanes are correctly identified and managed by the hard drive enclosure controller. It works by monitoring designated pin signals. When a hard drive backplane is inserted or removed, the signal state changes, which the I / O expander detects and reports to the hard drive enclosure controller.
[0082] Through the coordinated work of multi-channel analog switches and the first bus interface pins, flexible switching and efficient distribution of signals are achieved, ensuring that the hard drive box controller can accurately communicate with multiple hard drive backplanes and sensors, thereby monitoring key indicators such as hard drive temperature and power status in real time, effectively improving the monitoring coverage and response speed.
[0083] By connecting the IO Expander to the second bus interface pins, the system automatically detects the presence of the hard drive backplane and instantly updates the monitoring list, eliminating blind spots in device management. This not only simplifies hardware layout and reduces maintenance complexity, but also enhances the stability and reliability of the hard drive system, enabling comprehensive, intelligent management of the server's hard drive health.
[0084] In an exemplary embodiment, the above-mentioned control module includes a temperature sensor, a fan module and a speed regulation module, wherein the temperature sensor is arranged on the surface of the server hard disk or at a designated position inside the hard disk box, and is used to collect temperature data of the multiple designated components; the fan module includes fan units of multiple server hard disks, and one server hard disk includes at least one fan unit; the speed regulation module is used to adjust the rotational speed of the fan unit of at least one server hard disk among the multiple server hard disks according to a logical correspondence, wherein the logical correspondence between the temperature of the multiple server hard disks and the expected rotational speed is pre-stored in the speed regulation module.
[0085] In this embodiment, the server hard disks in the hard disk box include two types, one is an NVMe hard disk, and the other is a non-NVMe hard disk. In other words, the embodiment of the present application supports monitoring of multiple types of server hard disks, solving the problem in related technologies that only a single type or model of hard disk can be monitored.
[0086] Specifically, the software implementation of temperature control involves the hard drive enclosure controller reading the hard drive's temperature sensor status data at regular intervals (e.g., 3 seconds) and converting it into actual temperature values. A temperature threshold is also set. When the hard drive temperature exceeds the threshold, the enclosure controller sends a temperature alarm to the main BMC and initiates appropriate cooling measures. For NVMe drives, data is read directly from the drive via I2C. For non-NVMe drives, temperature data is obtained from the RAID card (disk array controller card) by interacting with the drive.
[0087] like Figure 5 As shown, assuming that the fan of the server hard disk exists in the form of a fan module, and a server hard disk includes at least one fan unit, then after the hard disk box controller monitors the data of the temperature sensor of the server hard disk, it determines through analysis whether it is necessary to adjust the fan speed of at least one fan unit of the current server hard disk (partial adjustment or full adjustment).
[0088] Specifically, the speed control module searches for the logical correspondence between the current server hard disk temperature and the expected speed (which may be an interval value). If the current temperature does not match the expected speed found, the fan speed is adjusted.
[0089] It should be noted that in addition to monitoring the server hard drive's temperature, this embodiment also includes a heat dissipation control system for the hard drive enclosure. Multiple PWM speed-controlled fans are installed within a hard drive enclosure, with the fans' PWM control pins connected to the PWM output ports of the hard drive enclosure controller. In addition to the hard drive temperature sensor, the hard drive enclosure also includes multiple temperature monitoring points, including air inlet temperature, power supply temperature, RAID card temperature, and backplane temperature.
[0090] The software control strategy for the hard drive enclosure's thermal management is as follows: the enclosure controller dynamically adjusts the fan speed based on the various internal temperatures. Each temperature has a corresponding thermal parameter. For example, an air inlet temperature of 20°C corresponds to a speed of 30%, and 23°C corresponds to a speed of 60%. For another example, when the hard drive temperature is below 40°C, the fan speed runs at 30%; when the temperature is between 50°C and 60%, the fan speed increases to 60%; and when the temperature exceeds 50°C, the fan speed runs at 100%. After collecting various temperatures, the corresponding fan speed is calculated and the fan speed is increased based on the highest fan speed to avoid thermal issues. Furthermore, a temperature hysteresis is set to prevent frequent fan starts and stops. If the hard drive enclosure controller restarts or fails, the RTOS (Real-Time Operating System) takes over thermal management.
[0091] Temperature sensors collect real-time temperature data from key components within the enclosure. The fan module, tightly integrated with the speed control module, intelligently adjusts the speed of at least one fan for each drive based on a pre-set logical relationship between temperature and speed. This design not only enables precise monitoring of drive temperatures but also efficiently and energy-efficiently maintains the ideal operating temperature within the enclosure, enhancing the drive's stability and lifespan while reducing failure rates due to overheating.
[0092] In an exemplary embodiment, the above-mentioned control module also includes an exception determination module and an exception handling module, wherein the exception determination module is used to determine the target designated component in an abnormal state among the multiple designated components inside the server hard disk based on one data or a combination of multiple data in the status data of the multiple designated components; the exception handling module is used to adjust the state of the target designated component according to the state control signal.
[0093] The abnormality determination module is used to determine the target specified component in an abnormal state based on the hard disk status data collected by the hard disk enclosure controller. For example, based on the status signal collected from the fault indicator (indicating a hard disk failure), the faulty component in the current hard disk is determined.
[0094] Or, based on the collected status signal of the fault indicator and the hard disk location information, it is possible to quickly determine what kind of fault has occurred in the hard disk at a specific location, for example, the temperature of the hard disk 1 connected to the hard disk backplane is too high.
[0095] The exception handling module is used to perform fault handling on the faulty hard disk determined by the exception determination module according to the status control signal. For example, when it is determined that the temperature of the hard disk 1 connected to the hard disk backplane is too high, the fan speed of at least one fan unit of the hard disk 1 is increased.
[0096] In other words, after identifying an abnormal component, the exception handling module generates a status control signal to direct it to adjust its status. For example, if a hard drive overheats, the module will send a command to the speed control module to increase the fan speed to accelerate heat dissipation. If a hard drive failure is detected, the module will notify the RAID card to perform data redundancy transfer to ensure data security. This immediate response mechanism ensures rapid response and action when an abnormal situation occurs.
[0097] The exception handling module's intelligent analysis and immediate response enable accurate identification and timely handling of abnormal statuses within the server's hard drive enclosure components. Based on status data, the module quickly identifies and locates abnormal components. Whether it's an overheating hard drive, power fluctuations, or hard drive health warnings, it generates timely status control signals to guide the speed control module to adjust its cooling strategy or instruct the RAID card to implement data migration and redundancy protection. This proactive exception management and status adjustment mechanism not only improves the operational stability of the hard drive enclosure, but also effectively prevents data loss and hardware failures, enhancing the security of server data storage.
[0098] In an exemplary embodiment, the apparatus further comprises: a data transmission module configured to transmit the status data of the plurality of designated components acquired by the hard disk enclosure controller to the baseboard management controller.
[0099] The data transmission module serves as a bridge between the hard drive enclosure controller and the baseboard management controller, responsible for collecting and transmitting status data. Working closely with the hard drive enclosure controller and utilizing communication protocols, it gathers status information from components such as hard drives, temperature sensors, and RAID cards. The data transmission module then organizes and encrypts this information, securely and quickly transmitting it to the baseboard management controller via Ethernet or other high-speed communication buses. This provides comprehensive monitoring data support for the baseboard management controller (master BMC), ensuring that decisions can be made based on real-time, accurate status data.
[0100] To ensure data transmission security and efficiency, the data transmission module uses a standardized data packet format, such as the IPMI protocol data packet. This format defines the data structure and transmission rules, making the data less susceptible to tampering or loss during transmission.
[0101] The data transmission module also employs an acknowledgement and retransmission mechanism. After sending a data packet, it waits for a confirmation response from the baseboard management controller. If no response is received, the module automatically retransmits the packet until it confirms successful delivery. This mechanism improves data transmission reliability, ensuring accurate transmission of status data even in unstable network conditions.
[0102] It should be noted that the main BMC not only receives the monitoring data sent by the hard disk box controller and Figure 5 In addition to displaying the monitoring data on the display of the server shown, the server can also analyze and send control instructions based on the monitoring data received from the server hard disk to control the operating status or component parameters of each component in the hard disk box.
[0103] The data transmission module ensures real-time and reliable transmission of status data between the hard drive enclosure controller and the baseboard management controller. Data packets are encrypted and sent to the baseboard management controller via the IPMI protocol, ensuring complete data delivery even in poor network conditions. This enhances server hard drive monitoring capabilities, improving server hard drive enclosure system stability and data storage security.
[0104] In an exemplary embodiment, the connection mode of the hard disk enclosure is adjustable so that the hard disk enclosure can adapt to servers of different specifications.
[0105] In this embodiment, modular design of the hard drive enclosure and highly compatible software design can be used, but is not limited to, to enable the hard drive enclosure to be migrated as a whole to other servers. Specifically, the modular design of the hardware includes, but is not limited to, standardized interfaces, pluggable design, adjustable size and shape, and adaptive heat dissipation solutions.
[0106] In addition, the upper cover and base of the hard drive enclosure are designed with adjustable ventilation holes, the size of which can be adjusted according to the heat dissipation requirements of the server, ensuring that the hard drive enclosure can effectively dissipate heat in any environment.
[0107] For example, interface standardization primarily refers to the use of standardized Ethernet, control, and power interfaces in hard drive enclosures. These interfaces are designed in accordance with industry standards, enabling seamless integration with servers of varying specifications. For example, they can connect not only to NVMe hard drives but also to non-NVMe drives, ensuring compatibility with server storage hardware.
[0108] The pluggable design means that all interfaces are modular in design, which makes it easy for users to adjust according to the actual interface type of the server without the need for complex hardware modifications; the hard drive box provides a corresponding interface conversion module, and users only need to simply replace the interface module to achieve connection matching with the server.
[0109] Software compatibility design includes but is not limited to protocol support diversity, driver universality, intelligent identification system, software update mechanism, etc.
[0110] For example, protocol diversity includes software-level support. For example, the hard drive enclosure controller supports multiple communication protocols, including but not limited to IPMI and I2C, to achieve compatibility with different server management systems. This means that no matter which management software the server uses, it can exchange data with the hard drive enclosure through the appropriate protocol to obtain monitoring information or issue control commands.
[0111] Through the above-mentioned flexible design at the hardware and software levels, the hard disk box in the embodiment of the present application can flexibly adapt to servers of different specifications. Whether it is the old architecture or the emerging high-performance computing environment, it can achieve efficient and stable monitoring and management, thereby improving the applicability of the hard disk box, improving the storage performance and reliability of the server, and also reducing the operation and maintenance costs and complexity.
[0112] Professionals may further appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the above description has generally described the components and steps of each example according to their functions. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professionals and technicians may use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0113] The above is a detailed introduction to the server hard disk monitoring device provided by this application. This article uses specific examples to illustrate the principles and implementation methods of this application. The description of the above examples is only intended to help understand the method and core concept of this application. It should be noted that for those skilled in the art, without departing from the principles of this application, various improvements and modifications can be made to this application, and such improvements and modifications also fall within the scope of protection of the claims of this application.
Claims
1. A server hard disk monitoring device, characterized in that: include: A baseboard management controller, a hard disk enclosure controller, and a control module, wherein the baseboard management controller is connected to the hard disk enclosure controller, and the baseboard management controller and the hard disk enclosure controller perform two-way data exchange; The baseboard management controller is used to display information when monitoring the status of the server hard disk in the hard disk box; The hard disk enclosure controller is used to obtain status data of multiple specified components inside the server hard disk, and generate status control signals based on the status data of the multiple specified components, wherein the status control signals are used to control the operating status and component parameters of the multiple specified components; The control module is configured to adjust the operating states and component parameters of the plurality of designated components in response to the received state control signal.
2. The device according to claim 1, characterized in that The device further comprises: The hard disk backplane, the hard disk box controller is connected to the hard disk backplane through a multi-way analog switch; A disk array controller card is connected to the hard disk enclosure controller via a bus interface and is used to perform read and write operations on the server hard disk in the hard disk enclosure.
3. The device according to claim 2, characterized in that The hard disk backplane also includes: A power control switch is used to control the power supply of the server hard disk in response to the switch closed state sent by the hard disk box controller; or to control the power supply of the server hard disk in response to the switch open state sent by the hard disk box controller.
4. The device according to claim 2, characterized in that The disk array controller card also includes: A status register, the status register being used to store the health status of the disk array controller card and read / write error record information; A data read and write management chip is connected to the hard disk box controller through the bus interface and is used to read and write status data of multiple specified components inside the server hard disk.
5. The device according to claim 1, characterized in that The hard disk enclosure controller includes a data collection module, wherein the data collection module is used to obtain status data of multiple specified components inside the server hard disk according to a data reading instruction.
6. The device according to claim 1, characterized in that The hard disk enclosure controller includes a first bus interface pin and a second bus interface pin; The first bus interface pin is connected to a multi-way analog switch, wherein the multi-way analog switch is used for signal switching and distribution; The second bus interface pin is connected to an input / output expander, wherein the input / output expander is used to detect the connection status of the hard disk backplane.
7. The device according to claim 1, characterized in that The control module includes a temperature sensor, a fan module and a speed control module, wherein: The temperature sensor is provided on the surface of the server hard disk or at a designated location inside the hard disk box, and is used to collect temperature data of the plurality of designated components; The fan module includes a plurality of fan units of server hard disks, and one server hard disk includes at least one fan unit; The speed regulation module is used to adjust the rotational speed of a fan unit of at least one of the multiple server hard disks according to a logical correspondence, wherein the speed regulation module pre-stores the logical correspondence between the temperatures of the multiple server hard disks and the expected rotational speeds.
8. The device according to claim 1, characterized in that The control module also includes an abnormality determination module and an abnormality processing module, wherein: The abnormality determination module is configured to determine a target designated component in an abnormal state among the multiple designated components inside the server hard disk based on one data or a combination of multiple data in the status data of the multiple designated components; The exception handling module is used to adjust the state of the target designated component according to the state control signal.
9. The device according to claim 1, characterized in that The device further comprises: The data transmission module is used to transmit the status data of the multiple specified components obtained by the hard disk enclosure controller to the baseboard management controller.
10. The device according to any one of claims 1 to 9, characterized in that The connection mode of the hard disk box is adjustable so that the hard disk box can adapt to servers of different specifications.