Expander card cascade system of disk cluster and supervision method

By introducing a baseboard management controller and programmable logic devices into the disk cluster system, real-time out-of-band monitoring and reset control of the Expander card is achieved, solving the problem that out-of-band monitoring and control cannot be achieved in the prior art, and improving the system's fault tolerance and business continuity.

CN121832849APending Publication Date: 2026-04-10XIAMEN YUANCHOU INTELLIGENT COMPUTING TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-31
Publication Date
2026-04-10

AI Technical Summary

Technical Problem

Existing dual expander redundancy designs cannot achieve out-of-band monitoring and control of the expander state, resulting in the inability to achieve seamless path switching and fault response.

Method used

The expander card cascade system using disk clusters enables real-time out-of-band monitoring and reset control of the first and second level expander cards through the baseboard management controller and programmable logic devices. Fault detection and switching are performed using complex programmable logic devices (CPLDs) to ensure the high availability and stability of the system.

Benefits of technology

It enables real-time monitoring and failover of Expander status, avoids data access interruption, improves system fault tolerance and overall efficiency, reduces maintenance costs, and ensures business continuity.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121832849A_ABST
    Figure CN121832849A_ABST
Patent Text Reader

Abstract

The invention discloses a disk cluster expander card cascade system and a supervision method, and relates to the technical field of disk cluster expander card cascade, the disk cluster expander card cascade system comprises a substrate, the substrate comprises a substrate management controller and a complex programmable logic device, and the substrate management controller is electrically connected with the complex programmable logic device; a plurality of primary expander cards, wherein each primary expander card is electrically connected with the complex programmable logic device; and a plurality of secondary expander cards, each primary expander card is connected with at least one secondary expander card, and each secondary expander card is respectively connected with a plurality of disks. The system realizes out-of-band management and monitoring of the expanders through a complex programmable logic device of a substrate, and transmits information of the expanders to a substrate management controller, so that the problem that an existing dual-Expander redundancy design scheme cannot realize out-of-band direct monitoring and control of Expander states is solved, interruption of data access is avoided, and the reliability of data access is improved. And the service continuity is ensured.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of expander card cascading of disk clusters, and particularly relates to an expander card cascading system and a monitoring method of a disk cluster. BACKGROUND

[0002] In modern enterprise-level storage systems, JBOD architecture has become an important solution for mass data storage due to its flexibility and scalability. As the core component of JBOD, the expander plays a key role in system connectivity, reliability, and manageability. Its main function is to expand the connectivity of the storage system, solving the problem of limited number of ports of the host controller (such as HBA or RAID card).

[0003] In terms of cascading technology, for the dual Expander redundancy design, the current main method is to rely on the host HBA for fault monitoring and Expander path switching. There is no direct communication and synchronization between the two Expanders. In the process of path switching, complete seamless connection cannot be achieved. The communication between the host HBA and the JBOD is mainly in-band, and it is impossible to realize the out-of-band monitoring and control of the Expander state.

[0004] In summary, the existing dual Expander redundancy design scheme has the problem that it cannot realize the out-of-band monitoring and control of the Expander state. SUMMARY

[0005] The present application provides an expander card cascading system of a disk cluster, a monitoring method and an electronic device, to at least solve the problem that the existing dual Expander redundancy design scheme in the related art cannot realize the out-of-band monitoring and control of the Expander state.

[0006] The present application provides an expander card cascading system of a disk cluster, comprising: a substrate, the substrate comprising a substrate management controller and a programmable logic device, the substrate management controller being electrically connected with the programmable logic device; a plurality of primary expander cards, each primary expander card being electrically connected with the programmable logic device; a plurality of secondary expander cards, each primary expander card being connected with at least one secondary expander card, and each secondary expander card being connected with a plurality of disks.

[0007] This application also provides a monitoring method for an expander card cascade system, used to monitor an expander card cascade system of any type of disk cluster. The method includes: upon detecting fault information in the expander card cascade system, determining the faulty expander card based on the fault information; if the faulty expander card is a first-level expander card, sending the fault information to a programmable logic device (PLD) to reset the abnormal first-level expander card through the PLD, wherein the abnormal first-level expander card is the faulty first-level expander card; if the faulty expander card is a second-level expander card, sending the fault information to the PLD to reset the abnormal second-level expander card through the PLD, wherein the abnormal second-level expander card is the faulty second-level expander card.

[0008] This application discloses a disk cluster expander card cascading system that, through programmable logic devices (PLDs) and a baseboard management controller, enables real-time out-of-band monitoring, reset control, and fault switching for first- and second-level expander card failures, effectively ensuring the high availability and stability of the disk cluster system. Specifically, the PLDs mutually monitor the status of the first-level expander cards. When a failure is detected in any first-level expander card, the PLD can quickly initiate and execute a reset command, simultaneously transmitting the fault information to the baseboard management controller. When a second-level expander card fails, the first-level expander card connected to it can be monitored and reset via the PLD. In this way, when any expander card fails, the system can react quickly, using a backup expander card to take over the responsibilities of the failed unit, avoiding data access interruption and ensuring service continuity. This mechanism significantly improves the system's fault tolerance, reduces maintenance costs, and allows for replacement of faulty components without system downtime, greatly improving overall system efficiency and user experience. Through the above technical solution, this application solves the problem that the existing dual expander redundancy design scheme in related technologies cannot achieve out-of-band direct monitoring and control of the expander state. Attached Figure Description

[0009] To more clearly illustrate the embodiments of this application, the accompanying drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0010] Figure 1 A schematic diagram of a disk cluster expander card cascading system provided in this application embodiment;

[0011] Figure 2A schematic diagram of the structure of a specific disk cluster expander card cascading system provided in the embodiments of this application;

[0012] Figure 3 A schematic diagram illustrating the connection between redundant primary and secondary expander cards provided in this embodiment of the application;

[0013] Figure 4 This is a diagram of the cascaded architecture of the expander card for the disk cluster provided in the embodiments of this application;

[0014] Figure 5 A system architecture block diagram of the first-level Expander anomaly provided in the embodiments of this application;

[0015] Figure 6 A system architecture block diagram of the second-level Expander exception provided in the embodiments of this application.

[0016] The above figures include the following reference numerals:

[0017] 10. Baseboard; 11. Baseboard Management Controller; 12. Programmable Logic Device; 20. Level 1 Expander Card; 21. Connection Port; 22. Redundant Port; 30. Level 2 Expander Card; 40. Disk; 50. Host Bus Adapter. Detailed Implementation

[0018] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the protection scope of this application.

[0019] It should be noted that, in the description of this application, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. The terms "first," "second," etc., in this application are used to distinguish similar objects and are not used to describe a specific order or sequence.

[0020] To enable those skilled in the art to better understand the present application, the present application will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0021] For ease of description, the following explains some of the nouns or terms used in the embodiments of this application:

[0022] JBOD: Just a Bunch of Disks. It's a storage configuration that combines multiple independent physical hard drives into a single logical storage unit.

[0023] Expander: In storage architecture, the expander is a key component used to extend the connectivity of the storage system, enabling a single host interface (such as a SAS or SATA controller) to manage more hard drives.

[0024] SAS, Serial Attached SCSI, is a common enterprise-level hard drive interface standard in storage systems.

[0025] SATA, or Serial Advanced Technology Attachment, is primarily used to connect storage devices such as hard disk drives (HDDs) and solid-state drives (SSDs).

[0026] CPLD: Complex Programmable Logic Device.

[0027] BMC: Baseboard Management Controller.

[0028] HBA: Host Bus Adapter. It connects the host (server / workstation) to storage devices (such as hard drives, SSDs, JBODs, etc.) and is responsible for data transfer.

[0029] Embodiments of this application provide a disk cluster expander card cascading system, such as Figure 1 As shown, it includes:

[0030] The substrate 10 includes a substrate management controller 11 and a programmable logic device 12, wherein the substrate management controller 11 is electrically connected to the programmable logic device 12.

[0031] Multiple primary expander cards 20, each of which is electrically connected to the programmable logic device 12;

[0032] Multiple secondary expander cards 30, each of the aforementioned primary expander cards 20 is connected to at least one of the aforementioned secondary expander cards 30, and each of the aforementioned secondary expander cards 30 is connected to multiple disks 40 respectively.

[0033] Among them, the programmable logic device can be a complex programmable logic device (CPLD).

[0034] In this embodiment, the disk cluster expander card cascading system, through a complex programmable logic device (CPLD) and a baseboard management controller (BMC) on the substrate, achieves real-time out-of-band monitoring, reset control, and fault switching for first- and second-level expander card failures, effectively ensuring the high availability and stability of the disk cluster system. Specifically, the CPLDs monitor the status of the first-level expander cards. When a failure is detected in any first-level expander card, the CPLD can quickly initiate and execute a reset command, simultaneously transmitting the fault information to the baseboard management controller. When a second-level expander card fails, the first-level expander card connected to it can be monitored and reset via the CPLD. In this way, when any expander card fails, the system can react quickly, using a backup expander card to take over the responsibilities of the failed unit, avoiding data access interruption and ensuring service continuity. This mechanism significantly improves the system's fault tolerance, reduces maintenance costs, and allows for replacement of faulty components without system downtime, greatly improving the overall system efficiency and user experience. Through the above technical solution, this application solves the problem that the existing dual expander redundancy design scheme in related technologies cannot achieve out-of-band direct monitoring and control of the expander state.

[0035] Furthermore, in this embodiment, the programmable logic device is connected to each of the first-level expander cards via a serial communication protocol to monitor the operating status of each of the first-level expander cards.

[0036] In this embodiment, the programmable logic device (CPLD) establishes communication connections with multiple Level 1 expander cards via the I2C serial communication protocol. This design enables the CPLD to monitor the operational status of the Level 1 expander cards in real time. Specifically, as the monitoring core, the CPLD receives status information from the Level 1 expander cards. Once a fault is detected in a Level 1 expander, the CPLD immediately performs a reset action, attempting to restore the normal operation of the faulty expander. Simultaneously, the CPLD transmits the fault information to the Baseboard Management Controller (BMC) on the management board and notifies the host via in-band communication, ensuring that the host can adjust its data path in a timely manner to prevent the faulty expander from affecting data transmission. This integrated monitoring and control mechanism not only improves the reliability and stability of the JBOD system but also simplifies the fault handling process, enabling fault detection and repair of expanders without downtime, thereby ensuring service continuity.

[0037] Furthermore, in this embodiment, each of the aforementioned first-level expander cards is connected to at least one of the aforementioned second-level expander cards via a point-to-point serial communication protocol.

[0038] In this embodiment, the primary expander card establishes a connection with at least one secondary expander card via a point-to-point serial communication protocol such as SAS (Serial Attached SCSI) or SATA (Serial Advanced Technology Attachment). The core of this design lies in the fact that, through the collaborative interaction and interconnection mechanism between the primary expander cards, real-time monitoring and fault takeover functions are achieved, ensuring that even if one of the primary expander cards fails, the system can seamlessly maintain its connectivity and data transmission efficiency. Specifically, the primary expander card can leverage the advantages of point-to-point serial communication to directly monitor the status of the secondary expander card. Once an anomaly is detected in the secondary expander card, communication is immediately established through a Complex Programmable Logic Device (CPLD) on the substrate, enabling the CPLD to reset the faulty secondary expander card, preventing data transmission interruption and significantly enhancing the stability and reliability of the system.

[0039] Furthermore, this strategy extends beyond fault response to support hardware maintenance or upgrades without impacting normal system operation. When a faulty expander card is detected, the system responds quickly and performs a reset. If the reset fails, the CPLD control mechanism forcibly locks the faulty expander card, allowing maintenance personnel to replace it without system downtime. Once the faulty expander card is repaired or replaced, the system automatically detects the new availability status and reconnects, quickly restoring business continuity. This highly flexible design enables enterprises to effectively manage hardware resources while ensuring business stability, reducing the risk of business interruption due to maintenance and faults, and improving the overall availability and cost-effectiveness of the JBOD system.

[0040] Furthermore, in this embodiment, the aforementioned baseboard management controller is connected to the aforementioned programmable logic device via a serial communication protocol, and is used to receive the operating status information of each of the aforementioned first-level expander cards and the aforementioned second-level expander cards sent by the aforementioned programmable logic device.

[0041] In this embodiment, the Baseboard Management Controller (BMC) establishes a communication connection with the Complex Programmable Logic Device (CPLD) via a serial communication protocol, receiving operational status information of the Level 1 and Level 2 expander cards from the CPLD. This design aims to enhance the fault detection capability and response speed of the JBOD system. By monitoring the status of the Level 1 and Level 2 expanders in real time through the CPLD, once an anomaly is detected, the information is quickly fed back to the BMC, which then takes corresponding fault handling measures, such as resetting or switching to a backup path, to ensure that the system can maintain stable operation without affecting business continuity even when encountering partial failures. This serial communication-based monitoring mechanism not only simplifies signal transmission between systems but also improves the accuracy and efficiency of fault monitoring, providing strong support for the reliable operation of large-scale enterprise-level storage systems.

[0042] The technical solution in this embodiment automates fault monitoring and rapid response by introducing a CPLD as a bridge between the primary and secondary expanders, thereby significantly improving the overall robustness and maintenance convenience of the JBOD system. This design is particularly important for large data centers and high-availability storage environments, as it minimizes downtime caused by hardware failures and ensures data continuity and security. Furthermore, since the communication between the CPLD and BMC is based on a serial protocol, this solution also possesses good scalability and compatibility, enabling it to adapt to different types of storage architectures.

[0043] Furthermore, in this embodiment, as Figure 2 As shown, the system also includes a host bus adapter 50, which is electrically connected to each of the first-level expander cards 20 and is used to transfer data in the disk 40 through the first-level expander cards 20 and the second-level expander cards 30.

[0044] In this embodiment, the system integrates a Host Bus Adapter (HBA), which establishes electrical connections with multiple Level 1 Expander cards. Through the collaborative work of these two levels of expanders, the HBA can effectively interact with the disks under the Level 2 Expander cards. The Level 1 Expander cards not only expand the host's connectivity but also play a crucial role in fault monitoring and dynamic takeover. Specifically, when one of the Level 1 Expanders detects a fault, a Complex Programmable Logic Device (CPLD) on the substrate can be used to reset the faulty expander and, if necessary, take over the disks it manages, ensuring seamless data path switching. Furthermore, the CPLD can monitor the status of the entire cascaded architecture and transmit information to the Baseboard Management Controller (BMC), enabling system administrators to promptly understand the health status of the JBODs. This design not only enhances system reliability but also simplifies the fault repair process, enabling maintenance or replacement of faulty expanders without business interruption. For Level 2 expander failures, Level 1 expanders can also be monitored and taken over via the CPLD, ensuring stable system operation.

[0045] This CPLD-based out-of-band management and monitoring mechanism, combined with a dual Expander redundancy design, provides unprecedented flexibility and robustness for enterprise-level storage systems. Especially when facing large-scale data storage demands, it significantly improves the scalability and data transmission efficiency of the storage system, while reducing the risk of data loss and business interruption due to hardware failures. Through the technical solution of this invention, storage system maintenance becomes more convenient and efficient, allowing for the replacement of critical components without downtime, greatly improving the operational efficiency of data centers and the user experience.

[0046] Furthermore, in this embodiment, as Figure 3 As shown, each of the aforementioned primary expander cards 20 is redundant in pairs. The redundant ports 22 of two redundant primary expander cards 20 are respectively connected to the corresponding secondary expander card 30 of the other primary expander card 20. The connection port 21 of each primary expander card is connected to the corresponding secondary expander card 30.

[0047] In the event that any first-level expander card is detected to be abnormal and the reset fails, the redundant port of the first-level expander card that is redundant with the abnormal first-level expander card is enabled, so that the first-level expander card can be connected to the second-level expander card corresponding to the abnormal first-level expander card, thereby maintaining the stable operation of the entire system without interrupting services.

[0048] In this embodiment, each primary expander card implements a pairwise redundant design, where the redundant ports of the redundant primary expander cards are connected to the corresponding secondary expander cards of the other primary expander card. This design allows the other redundant primary expander card to quickly take over its operation when a primary expander card fails. Real-time monitoring and control via CPLD ensure seamless switching of communication data paths between expanders, thereby maintaining the stable operation of the entire JBOD system without interrupting services. In the event of a primary or secondary expander card failure, the reset mechanism of the failed expander can be triggered immediately, allowing the healthy expander to take over the affected disk links. This enables fault repair or replacement operations without shutting down the system, significantly enhancing system reliability and maintenance efficiency.

[0049] Specifically, when a Level 1 Expander detects a failure in a peer Expander, they communicate with each other via CPLDs to quickly reset and update the status of the failed Expander. For Level 2 Expander failures, the Level 1 Expander monitors via CPLDs and takes appropriate reset and takeover measures to ensure the connectivity of all disk links and the continuity of data transmission. This two-tiered cascading and fault monitoring mechanism not only optimizes the resource utilization of the storage system but also enhances data security, enabling storage services to maintain stable operation even in the face of hardware failures.

[0050] Furthermore, in this embodiment, as Figure 2 As shown, the programmable logic device 12 is electrically connected to each of the first-level expander cards 20 and each of the second-level expander cards 30 via control lines, and is used to reset the first-level expander cards 20 and the second-level expander cards 30.

[0051] In this embodiment, the Complex Programmable Logic Device (CPLD) is electrically connected to each primary and secondary expander card via control lines, enabling reset processing and status monitoring of the primary and secondary expander cards. This design allows the primary expander card to quickly execute a reset operation via the CPLD when a secondary expander card fault is detected, thereby attempting to re-establish a stable connection path. Similarly, when a primary expander card encounters an abnormal situation, another redundant expander card of the same level can take over the responsibilities of the faulty component through the coordination of the CPLD, ensuring the continuity of data communication and the high availability of the system. As the central element of this mechanism, the CPLD not only strengthens the communication and cooperation between expanders at all levels but also provides flexible fault response methods, enabling hardware maintenance and upgrades without interrupting the operation of the JBOD system, significantly improving the stability and operational efficiency of the storage system.

[0052] Specifically, when a Level 1 expander card detects a fault or performance degradation in a Level 2 expander card, it sends a command to the CPLD. Upon receiving the signal, the CPLD immediately performs a reset on the Level 2 expander card. This process requires no host intervention, enabling rapid fault response and handling. Simultaneously, the CPLD monitors the status of the Level 1 expander card. Upon detecting an anomaly or fault, it can proactively trigger a reset process and notify a healthy Level 1 expander card to take over the connection of the faulty device. This out-of-band management and fault recovery mechanism via the CPLD ensures that the system can react quickly when a problem occurs with either a Level 1 or Level 2 expander card, avoiding service interruption and guaranteeing seamless and secure data transfer. Furthermore, this design supports hot-swapping, allowing faulty expander cards to be replaced during system operation, further enhancing system resilience and maintenance convenience. Overall, this solution significantly improves the robustness and business continuity of the JBOD system by enhancing mutual monitoring and fault response capabilities among expanders.

[0053] The specific application environment architecture or specific hardware architecture on which the execution of the supervisory method for the combined extender card cascade system depends is described here.

[0054] Embodiments of this application provide a method for monitoring an expander card cascading system, used to monitor any of the aforementioned expander card cascading systems of disk clusters. The method includes:

[0055] If a fault is detected in the expander card cascade system, the expander card that has failed can be identified based on the fault information.

[0056] If the faulty expander card is a primary expander card, the fault information is sent to a programmable logic device (PLD) so that the PLD can reset the faulty primary expander card.

[0057] If the faulty expander card is a secondary expander card, the fault information is sent to the programmable logic device so that the programmable logic device can perform the reset process on the faulty secondary expander card.

[0058] Specifically, when a fault is detected in the expander card cascade system, this method first identifies the specific expander card that is faulty. If the fault occurs in a first-level expander card, the fault information is transmitted to a Complex Programmable Logic Device (CPLD), which then resets the faulty first-level expander card to restore its normal operation. If the fault is in a second-level expander card, the fault information is similarly sent to the CPLD, which resets the faulty second-level expander card to eliminate the impact of the fault and ensure the continuity and stability of the data transmission path. This embodiment enables fault monitoring and handling of expander cards without system downtime, improving the availability and stability of the JBOD system and reducing the risk of service interruption due to hardware failure. Furthermore, this method supports hot-swapping, meaning that after a first- or second-level expander card is repaired or replaced, it can be put back into use without restarting the entire system, further enhancing the system's maintenance convenience and flexibility, and achieving a more real-time and efficient fault response mechanism.

[0059] In the specific implementation process, after resetting the abnormal first-level expander card through the above-mentioned programmable logic device, the above method further includes: detecting whether the abnormal first-level expander card has recovered from the fault; if it is determined that the abnormal first-level expander card has not recovered from the fault, determining a target first-level expander card that is redundant with the abnormal first-level expander card, and enabling the redundant port of the target first-level expander card, so that the second-level expander card connected to the abnormal first-level expander card is connected to the target first-level expander card.

[0060] In this embodiment, after a Level 1 expander card detects a fault and is reset by a Complex Programmable Logic Device (CPLD), the system further checks whether the Level 1 expander card has returned to normal operation. If the Level 1 expander card fails to recover after the reset, the system automatically identifies and enables a target Level 1 expander card with which it forms a redundant configuration. Specifically, the redundant ports of the target Level 1 expander card are enabled. This operation aims to ensure that Level 2 expander cards previously connected to the faulty Level 1 expander card can seamlessly switch to the target Level 1 expander card, maintaining the connectivity and data transmission stability of the JBOD system. Through the monitoring and control of the CPLD, the switching between Level 1 expander cards can be completed efficiently without affecting business continuity, ensuring high availability and rapid recovery capabilities of the system in the face of faults. In addition, this design also supports hot-swapping, allowing for repair or replacement without interrupting system operation after the faulty Level 1 expander card is removed, significantly enhancing the system's maintenance convenience and expansion flexibility.

[0061] Specifically, after enabling the redundant port of the target primary expander card so that the secondary expander card connected to the abnormal primary expander card can connect to the target primary expander card, the method further includes: generating a first alarm message, wherein the first alarm message is used to remind the abnormal primary expander card to be repaired or replaced; after detecting that the abnormal primary expander card has been repaired or replaced, resetting the expander card cascade system, and controlling the redundant port of the target primary expander card to be closed so that the secondary expander card connected to the abnormal primary expander card can be disconnected from the target primary expander card, thereby restoring the connection between the abnormal primary expander card and the corresponding secondary expander card.

[0062] In this embodiment, when a primary expander card detects a fault in a cascaded primary expander card, the target primary expander card can quickly respond through the CPLD mutual monitoring mechanism. It opens a previously redundant port to take over and connect to the secondary expander card originally connected to the faulty expander card, ensuring the business continuity and data access stability of the JBOD system. This dynamic switching mechanism does not interrupt normal system operation, achieving seamless data path transfer and significantly improving system availability and reliability. Subsequently, the system automatically generates a first alarm message to alert maintenance personnel about the status of the faulty primary expander card, enabling timely repair or replacement measures to minimize the impact of downtime on services. Once the faulty primary expander card is repaired or replaced and its normal operating status is detected, the system will reset the expander card cascade system. At this time, the target primary expander card will close its previously opened redundant port, sever the temporary connection with the secondary expander card of the original faulty expander card, restore the normal architectural layout, and re-establish the connection between the faulty primary expander card and the corresponding secondary expander card. All of these operations are completed automatically without interrupting business operations, demonstrating the system's intelligence and high efficiency. At the same time, through modular design and optimized cascading architecture, the JBOD storage system's scalability and ability to cope with sudden failures are enhanced, ensuring the stable operation of enterprise-level applications.

[0063] Furthermore, after sending the aforementioned fault information to the aforementioned programmable logic device to perform the aforementioned reset process on the abnormal secondary expander card through the aforementioned programmable logic device, the aforementioned method further includes: detecting whether the aforementioned abnormal secondary expander card has recovered from the fault; if it is determined that the aforementioned abnormal secondary expander card has not recovered from the fault, continuously resetting the aforementioned abnormal secondary expander card for a preset duration through the aforementioned programmable logic device to power off the aforementioned abnormal secondary expander card, and generating a second alarm message, wherein the aforementioned second alarm message is used to remind the aforementioned abnormal secondary expander card to be repaired or replaced.

[0064] In this embodiment, when a primary expander card detects a fault in a connected secondary expander card, it not only immediately sends fault information via a complex programmable logic device (CPLD), but also further performs a reset operation on the faulty secondary expander card via the CPLD in an attempt to restore its normal function. Subsequently, the system checks whether the faulty secondary expander card has successfully recovered. If the secondary expander card fails to return to normal, to prevent it from further affecting system stability, this embodiment will also perform a reset process on it via the CPLD for a preset duration, essentially powering off the faulty card. Simultaneously, the system generates a second alarm message, clearly indicating that the faulty secondary expander card needs repair or replacement. This mechanism ensures the high availability and maintenance efficiency of the JBOD system, enabling rapid location and isolation of faulty devices without affecting business continuity, reducing potential data loss risks, and improving overall service quality and user experience.

[0065] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods according to the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method.

[0066] To enable those skilled in the art to better understand the technical solution of this application, the implementation process of the disk cluster expander card cascading system of this application will be described in detail below with reference to specific embodiments.

[0067] This embodiment relates to a specific disk cluster expander card cascading system architecture, such as... Figure 4As shown, a two-stage Expander cascaded architecture is adopted, consisting of a baseboard (including a Baseboard Management Controller (BMC) and a Complex Programmable Logic Device (CPLD), two Level 1 Expander cards, and four Level 2 Expander cards. The BMC and CPLD are connected via the I2C communication protocol. The two Level 1 Expanders (Level 1 EXP1 and Level 1 EXP2) are connected to the CPLD via the I2C communication protocol. The CPLD is electrically connected to each Level 1 Expander (Level 1 EXP1 and Level 1 EXP2) and Level 2 Expanders (Level 2 EXP1, Level 2 EXP2, Level 2 EXP3, and Level 2 EXP4) via a reset control line (RST) to achieve reset processing for each Expander. The Level 1 Expanders adopt a dual-Expander redundancy design, connecting four Level 2 Expanders respectively. The Level 2 Expanders have a total of 40 SAS ports, which can connect up to 32 SAS / SATA hard drives for downlink and two 4-port uplinks. The Level 1 Expanders have a total of 24 SAS ports, with 8 ports allocated to the uplink head unit and 16 ports allocated to the Level 2 Expanders. In this way, through the expansion of two levels, a single host interface (HBA 4 Port) can connect up to 128 hard drives when fully configured. The entire cascaded architecture can be managed and monitored out-of-band via the CPLD (Complex Programmable Logic Device) on the substrate, and the information is transmitted to the BMC (Baseboard Management Controller). Furthermore, the first-level expander can monitor and reset itself via the CPLD. The second-level expander can be monitored and reset by the first-level expander. This ensures that if either the first or second-level expander fails, it will be immediately reset without affecting the normal operation of the replacement expander. This truly achieves seamless switching of communication data paths, ensuring stable and continuous service operation.

[0068] Figure 5 This is a system architecture diagram for a first-level Expander anomaly. Figure 5 As shown, the Host Bus Adapter (HBA) communicates with each Level 1 Expander, the CPLD communicates with each Level 1 Expander via I2C, and the CPLD is connected to each Level 1 Expander via the Reset Control Line (RST). This enables monitoring, takeover, and replacement of Level 1 Expanders in case of malfunction, specifically including the following:

[0069] 1) Each of the two first-level expanders can connect to four second-level expanders simultaneously, managing a full JBOD configuration of 128 hard drives. During normal use, in default mode, each expander can connect to two second-level expanders, managing 64 hard drives respectively. The downstream ports (indicated by dashed lines) for the other two second-level expanders are disabled by default.

[0070] 2) If the first-stage Expander2 malfunctions, the head unit will acquire the fault information in-band and notify the first-stage Expander1 via in-band. After learning of the Expander2 malfunction, Expander1 will proactively send the Expander2 fault information to the CPLD via the I2C protocol, and the CPLD will then reset Expander2.

[0071] 3) After the CPLD resets Expander2, it sends a reset notification to Expander1 via I2C and simultaneously informs the management board BMC via I2C. Expander1 receives this information and then informs the head unit via in-band communication.

[0072] 4) After the head unit receives the information that Expander2 has been reset, it attempts to link Expander2. If the link is successful and the system operates normally, the head unit notifies Expander1 that Expander2 has recovered. Expander1 then informs the management board BMC via I2C that JBOD services are proceeding normally.

[0073] 5) If the head unit fails to link with Expander2 again, it will notify Expander1 via I2C. Expander1 will then notify the CPLD to force a reset and lock Expander2 via I2C, and simultaneously notify the BMC management board via I2C.

[0074] 6) After the CPLD forcibly resets Expander2, it notifies Expander1 via I2C. Expander1 then notifies the head unit in-band. After receiving the information, the head unit sends a command to open the downstream dashed port of Expander1 and take over the remaining 64 hard drives.

[0075] 7) At this point, the Expander2 can be removed for repair or replacement.

[0076] 8) After Expander2 is repaired or replaced with a new board, it can be directly inserted into the JBOD system (both the first-stage and second-stage Expanders support hot-swapping). Once the CPLD on the Expander board detects a change in the presence status of the first-stage Expander2, it will transmit the Expander2 insertion information to Expander1, which in turn informs the head unit. The head unit then commands the downlink dashed port of Expander1 to be closed.

[0077] 9) After the downlink port of Expander1 changes, it notifies the CPLD to release the reset signal of Expander2. At the same time, the CPLD transmits the released reset signal information to Expander1 and the management board BMC, and Expander1 notifies the head unit in-band.

[0078] 10) After the head unit obtains the Expander2 ready information, it relinks Expander2, and the service is restored.

[0079] Figure 6 This is a system architecture diagram for a second-level Expander anomaly. Figure 6 As shown, the Baseboard Management Controller (BMC) communicates with the CPLD via the I2C protocol. The CPLD communicates with each primary Expander via the I2C protocol. The primary Expanders communicate with the secondary Expanders via the SAS protocol. Each secondary Expander is connected to a corresponding HDD on the disk backplane. The CPLD is connected to each primary and secondary Expander via the reset control line (RST) for monitoring, taking over, and replacing abnormal secondary Expanders. Specifically, this includes the following:

[0080] 1) When the two expanders in the first stage detect a fault in a certain expander in the second stage (for example, expander1 in the first stage detects a fault in expander2 in the second stage), they will notify the CPLD on the expander board via I2C.

[0081] 2) The CPLD performs a reset action on the level 2 Expander2 and notifies the level 1 Expander1 via I2C.

[0082] 3) Level 1 Expander1 links to Level 2 Expander3 again. If the link is successful and runs normally, the service will be restored.

[0083] 4) If the link still fails, Level 1 Expander1 will be forcibly reset via CPLD and drag Level 2 Expander3 to its death.

[0084] 5) At this point, the second-stage Expander2 can be removed for repair or replacement.

[0085] 6) After the Level 2 Expander3 is repaired or replaced with a new board, it can be directly inserted into the JBOD system. After the CPLD detects the change in the presence status of Expander3, it transmits the presence information to the Level 1 Expander1.

[0086] 7) Level 1 Expander1 out-of-band notification CPLD releases Level 2 Expander3 reset signal and attempts to reconnect to LinkExpander3. The connection is successful and the system runs normally, and the service is restored.

[0087] This embodiment employs a cascaded architecture consisting of one baseboard, two primary expanders, and four secondary expanders. The primary expanders utilize a dual-expander redundancy design. This enables mutual monitoring, supervision, and replacement between the primary expander cards. It also allows for real-time monitoring and supervision of the secondary expanders by the primary expanders, ensuring real-time monitoring of expander performance and reliability without powering off or interrupting services.

[0088] Embodiments of this application also provide a monitoring device for an expander card cascade system, used to monitor any of the above-mentioned expander card cascade systems of disk clusters, the device comprising:

[0089] The determining unit is used to determine the faulty expander card based on the fault information when a fault information of the expander card cascade system is detected.

[0090] The first reset processing unit is configured to send the fault information to the programmable logic device when the faulty expander card is a first-level expander card, so as to reset the abnormal first-level expander card through the programmable logic device, wherein the abnormal first-level expander card is the faulty first-level expander card.

[0091] The second reset processing unit is used to send the fault information to the programmable logic device when the faulty expander card is a secondary expander card, so as to perform the reset processing on the abnormal secondary expander card through the programmable logic device, wherein the abnormal secondary expander card is the secondary expander card that has failed.

[0092] When this device detects a fault in the expander card cascade system, it first identifies the specific expander card that is faulty. If the fault occurs in a primary expander card, the fault information is transmitted to a Complex Programmable Logic Device (CPLD), which then resets the faulty primary expander card to restore its normal operation. If the fault is in a secondary expander card, the fault information is similarly sent to the CPLD, which resets the faulty secondary expander card to eliminate the impact of the fault and ensure the continuity and stability of the data transmission path. This embodiment enables fault monitoring and handling of expander cards without system downtime, improving the availability and stability of the JBOD system and reducing the risk of service interruption due to hardware failure. Furthermore, this method supports hot-swapping, meaning that after a primary or secondary expander card is repaired or replaced, it can be put back into use without restarting the entire system, further enhancing the system's maintenance convenience and flexibility, and achieving a more real-time and efficient fault response mechanism.

[0093] As an optional solution, the device further includes a first detection unit and a determination unit; the first detection unit is used to detect whether the faulty first-level expander card has recovered after the faulty first-level expander card has been reset by the above-mentioned programmable logic device; the determination unit is used to determine a target first-level expander card that is redundant with the faulty first-level expander card when it is determined that the faulty first-level expander card has not recovered, and to enable the redundant port of the target first-level expander card so that the second-level expander card connected to the faulty first-level expander card is connected to the target first-level expander card.

[0094] In one optional embodiment, the device further includes a first generation unit and a control unit; the first generation unit is configured to generate a first alarm message after enabling the redundant port of the target primary expander card so that the secondary expander card connected to the abnormal primary expander card is connected to the target primary expander card, wherein the first alarm message is used to remind the abnormal primary expander card to be repaired or replaced; the control unit is configured to, after detecting that the abnormal primary expander card has been repaired or replaced, reset the expansion card cascade system and control the redundant port of the target primary expander card to be closed, so that the secondary expander card connected to the abnormal primary expander card is disconnected from the target primary expander card, and the abnormal primary expander card is reconnected to the corresponding secondary expander card.

[0095] In one optional embodiment, the device further includes a second detection unit and a second generation unit. The second detection unit is used to detect whether the faulty secondary expander card has recovered after sending the fault information to the programmable logic device to perform the reset process on the faulty secondary expander card through the programmable logic device. The second generation unit is used to continuously reset the faulty secondary expander card for a preset duration through the programmable logic device when it is determined that the faulty secondary expander card has not recovered, so as to power off the faulty secondary expander card and generate a second alarm message, wherein the second alarm message is used to remind the faulty secondary expander card to be repaired or replaced.

[0096] For a description of the features in the embodiment corresponding to the monitoring device of the expander card cascade system, please refer to the relevant description of the embodiment corresponding to the monitoring method of the expander card cascade system, which will not be repeated here.

[0097] Embodiments of this application also provide an electronic device including a memory and a processor, the memory storing a computer program, the processor being configured to run the computer program to perform the steps in any of the above-described embodiments of the supervisory method for an extender card cascading system.

[0098] Embodiments of this application also provide a computer-readable storage medium storing a computer program, wherein the computer program is configured to execute the steps in any of the above-described embodiments of the supervisory method for an extender card cascading system at runtime.

[0099] In one exemplary embodiment, the aforementioned computer-readable storage medium may include, but is not limited to, various media capable of storing computer programs, such as a USB flash drive, read-only memory (ROM), random access memory (RAM), portable hard disk, magnetic disk, or optical disk.

[0100] Embodiments of this application also provide a computer program product, which includes a computer program that, when executed by a processor, implements the steps in the monitoring method embodiments of any of the above-described expander card cascading systems.

[0101] Embodiments of this application also provide another computer program product, including a non-volatile computer-readable storage medium storing a computer program, which, when executed by a processor, implements the steps in the monitoring method embodiments of any of the above-described extender card cascading systems.

[0102] Those skilled in the art will further recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of the various examples have been generally described in terms of functionality in the foregoing description. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0103] The foregoing has provided a detailed description of a disk cluster expander card cascading system, monitoring method, and electronic device provided in this application. Specific examples have been used to illustrate the principles and implementation methods of this application. The descriptions of the embodiments above are merely for the purpose of helping to understand the method and core ideas of this application. It should be noted that those skilled in the art can make various improvements and modifications to this application without departing from its principles, and these improvements and modifications also fall within the protection scope of the claims of this application.

Claims

1. A disk cluster expander card cascading system, characterized in that, include: A substrate, the substrate including a substrate management controller and a programmable logic device, the substrate management controller being electrically connected to the programmable logic device; Multiple first-level expander cards, each of which is electrically connected to the programmable logic device; Multiple secondary expander cards, each of the primary expander cards is connected to at least one of the secondary expander cards, and each of the secondary expander cards is connected to multiple disks respectively.

2. The disk cluster expander card cascading system according to claim 1, characterized in that, The programmable logic device is connected to each of the first-level expander cards via a serial communication protocol to monitor the operating status of each of the first-level expander cards.

3. The disk cluster expander card cascading system according to claim 1, characterized in that, Each of the first-level expander cards is connected to at least one of the second-level expander cards via a point-to-point serial communication protocol.

4. The disk cluster expander card cascading system according to claim 1, characterized in that, The baseboard management controller communicates with the programmable logic device via a serial communication protocol and is used to receive the operating status information of each of the first-level expander cards and the second-level expander cards sent by the programmable logic device.

5. The disk cluster expander card cascading system according to claim 1, characterized in that, The system also includes a host bus adapter, which is electrically connected to each of the first-level expander cards and is used to transfer data in the disk through the first-level expander cards and the second-level expander cards.

6. The disk cluster expander card cascading system according to claim 1, characterized in that, Each of the first-level expander cards is redundant in pairs, and the redundant ports of the two redundant first-level expander cards are respectively connected to the second-level expander card corresponding to the other first-level expander card; the programmable logic device is electrically connected to each of the first-level expander cards and each of the second-level expander cards through control lines, and is used to reset the first-level expander cards and the second-level expander cards.

7. A monitoring method for an extender card cascading system, characterized in that, The method for monitoring an expander card cascading system of any one of claims 1 to 6 includes: If a fault is detected in the expander card cascade system, the expander card that has failed is determined based on the fault information. If the faulty expander card is a primary expander card, the fault information is sent to a programmable logic device (PLD) so that the PLD can reset the faulty primary expander card. If the faulty expander card is a secondary expander card, the fault information is sent to the programmable logic device so that the programmable logic device can perform the reset process on the faulty secondary expander card.

8. The regulatory method according to claim 7, characterized in that, After resetting the faulty first-level expander card via the programmable logic device, the method further includes: Check whether the faulty first-level expander card has been recovered; If it is determined that the abnormal primary expander card has not recovered from the fault, a target primary expander card that is redundant with the abnormal primary expander card is identified, and the redundant port of the target primary expander card is enabled so that the secondary expander card connected to the abnormal primary expander card can be connected to the target primary expander card.

9. The regulatory method according to claim 8, characterized in that, After enabling the redundant port of the target primary expander card so that the secondary expander card connected to the abnormal primary expander card can be connected to the target primary expander card, the method further includes: Generate a first alarm message, wherein the first alarm message is used to remind the user to repair or replace the abnormal first-level expander card; After the abnormal primary expander card is detected to be repaired or replaced, the redundant port of the target primary expander card is controlled to be closed after the expander card cascade system is reset, so that the secondary expander card connected to the abnormal primary expander card is disconnected from the target primary expander card, and the abnormal primary expander card is restored to the connection with the corresponding secondary expander card.

10. The regulatory method according to claim 7, characterized in that, After sending the fault information to the programmable logic device to perform the reset process on the faulty secondary expander card via the programmable logic device, the method further includes: Check whether the faulty secondary expander card has been recovered; If it is determined that the faulty secondary expander card has not been recovered, the programmable logic device continuously resets the faulty secondary expander card for a preset duration to power off the faulty secondary expander card and generate a second alarm message, wherein the second alarm message is used to remind the faulty secondary expander card to be repaired or replaced.