RAID (Redundant Array of Independent Disks) card high-availability device and method based on multi-path cluster

By using a multi-RAID card collaborative design and an intelligent fault handling mechanism, the single point of failure and single path problem of traditional RAID cards are solved, realizing a highly reliable and high-performance RAID card system that supports multi-path load sharing and redundant backup.

CN120872248APending Publication Date: 2025-10-31CHENGDU HUARUI SHUXIN TECHNOLOGY CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202511041670.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-28
Publication Date
2025-10-31

AI Technical Summary

Technical Problem

Traditional RAID cards suffer from single point of failure risk, limited path, and weak anti-interference capabilities. Existing high availability technologies are costly and suffer significant performance losses, failing to effectively address the redundancy and cluster management issues at the RAID card level.

Method used

A high-availability device employing multiple RAID cards working in tandem is connected to the server motherboard via a PCIe bus. It incorporates front-end and back-end multi-path modules and an intelligent fault handling mechanism to achieve multi-path load balancing and redundant backup. Combined with a cluster management module, it maintains heartbeat information and enables automatic fault switching.

Benefits of technology

It improves system reliability and performance, eliminates the risk of single point of failure, enables IO load sharing and redundant access, reduces operation and maintenance costs, supports Active-Active access mode, and optimizes read and write performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120872248A_ABST
    Figure CN120872248A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of storage equipment, and particularly discloses an RAID card high-availability device and method based on a multi-path cluster, and the device comprises at least two RAID cards; the RAID engine is used for processing RAID read-write data, configuring RAID levels and processing hard disk hot plug; the front-end multi-path module is used for establishing a plurality of data channels with a host and exposing a virtual disk with the same NGUID to the host, so that the host multi-path module can access the virtual disk in an aggregated manner; the rear-end multi-path module is used for establishing a multi-path data channel with the dual-port hard disk, transmitting read-write data and commands of the dual-port hard disk, and realizing disk access link redundancy and fault automatic switching; the system management module is used for processing internal and external management control events, state conversion and internal processes of the RAID card; the cluster management module is used for establishing communication channels among the RAID cards, maintaining heartbeat information among the RAID cards and managing cluster node states; the problem of single-point failure of a traditional RAID card is solved, and the reliability and performance of a system are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of storage device technology, and specifically to a high-availability device and method for RAID cards based on multi-path clustering. Background Technology

[0002] Currently, traditional RAID cards use a single controller design, which has the following drawbacks: Single point of failure risk: If the controller fails, all hard drives it manages will become inaccessible, leading to data loss or business interruption; Single path: The front-end host can only access the RAID card through a single path, making it impossible to achieve multi-path load balancing and redundant backup; Weak anti-interference capability: Under harsh environments (high radiation / electromagnetic interference), the risk of data access interruption is high, and the single access path is easily affected by interference, impacting data reliability.

[0003] Existing high-availability technologies rely on dedicated hardware arbitration modules, which are costly, have long I / O paths, and suffer from significant performance losses. They also fail to effectively address the redundancy and cluster management issues at the RAID card level. Therefore, there is an urgent need for a high-availability solution that supports multi-RAID card collaboration and has automatic failover capabilities. Summary of the Invention

[0004] To overcome the aforementioned technical problems in the prior art, this invention provides a high-availability device and method for RAID cards based on multi-path clustering. By employing multi-RAID card collaboration, front-end and back-end multi-path design, and intelligent fault handling mechanisms, it solves the single-point-of-failure problem of traditional RAID cards and improves the reliability and performance of the system.

[0005] To achieve the above objectives, embodiments of the present invention provide a high-availability device for RAID cards based on multi-path clustering, comprising: at least two RAID cards connected to a server motherboard via a PCIe bus; a RAID engine for processing RAID read / write data, configuring RAID levels, and handling hard drive hot-swapping; a front-end multi-path module for establishing multiple data channels with the host, transmitting read / write data and commands, and exposing virtual disks with the same NGUID to the host, enabling the host multi-path module to aggregate access to the virtual disks; a back-end multi-path module for establishing multiple data channels with dual-port hard drives, transmitting read / write data and commands from the dual-port hard drives, and achieving disk access link redundancy and automatic fault switching; a system management module for handling management and control events, state transitions, and internal processes of the RAID cards; and a cluster management module for establishing communication channels between RAID cards, maintaining heartbeat information between RAID cards, and managing the status of cluster nodes.

[0006] Preferably, the backend multipath module isolates IO conflicts, specifies the working controller node attributes of the virtual disk, and supports automatic switching in case of link failure. When a link failure is detected, it automatically switches to the backup access path.

[0007] Preferably, the cluster management module is used to establish a communication channel between RAID cards, maintain heartbeat information between RAID cards, and manage the status of cluster nodes. Specifically, it includes: establishing a communication channel between RAID cards through the PCIe P2P protocol, wherein the communication channel transmits cluster heartbeat messages and cluster configuration messages between RAID cards; performing a master-slave node election through a master-slave competition process to determine a unique master node for accessing DDF metadata, serializing events, and synchronizing configurations; responding to a split-brain event, designating a dual-port hard drive as an arbitration disk, competing for the arbitration disk lock through the SCSI Reserve command, with the winning RAID card surviving and the losing RAID card leaving the cluster.

[0008] Preferably, it also includes a cache region, which is only effective for HDDs and is used to accelerate read and write operations; for SSDs, the cache region is disabled.

[0009] Accordingly, the present invention also provides a high availability method for RAID cards based on multi-path clustering, comprising the following steps: S1: Combining multiple RAID cards into a cluster, and performing master-slave node election through the cluster management module to determine the master node and slave node, with the master node responsible for metadata management and event processing; S2: Exposing virtual disks with the same NGUID to the host through the front-end multi-path module, enabling the host to access the virtual disks through the multi-path module; S3: Connecting a dual-port hard drive to the back-end multi-path module, specifying the working controller of the virtual disk through the ALUA mechanism, and restricting the host to access the corresponding virtual disk only through the working controller; S4: Maintaining heartbeat information between RAID cards through the communication channel and monitoring the node status in real time; S5: When a RAID card failure is detected, a failover operation is triggered, switching the virtual disk belonging to the failed RAID card to the normal RAID card for management; when the failed RAID card recovers, a failback operation is triggered, and the virtual disk returns to its original RAID card.

[0010] Preferably, S1 specifically includes: the primary / backup node election initiates a primary / backup competition process through the cluster management module, determines the election result through an election algorithm, and the election result is used to allocate the work controllers of virtual disks to achieve IO load sharing; in S3, the ALUA mechanism avoids multiple RAID card access conflicts by statically allocating virtual disk ownership.

[0011] Preferably, in step S4, the heartbeat information is achieved by periodically sending status identifiers and timestamps. If no heartbeat information is received for a preset number of consecutive times, the corresponding node is determined to be faulty. In step S5, during the fault switching and fault return process, the virtual disk status and configuration data are synchronized through the system management module.

[0012] Preferably, the system also includes a cluster configuration synchronization step. When a user initiates a configuration operation through the management tool, the configuration command is sent to the master node. After the master node modifies the physical disk metadata, it notifies the slave node through the cluster management module. The slave node then reads and synchronizes the configuration data from the physical disk metadata area.

[0013] Preferably, the communication channel between RAID cards can be replaced with host shared memory based on CXL.mem technology. The host shared memory includes a cluster message queue, an IO forwarding queue, an arbitration data area, and a shared cache area, supports the Active-Active access mode of virtual disks, and allocates IO requests to the corresponding RAID cards through a hash algorithm.

[0014] The present invention has at least the following technical effects through the technical solution provided by the present invention: By employing a multi-RAID card cluster design, front-end and back-end multi-path redundancy, and intelligent fault handling mechanisms, several beneficial effects are achieved: It completely eliminates the single-point-of-failure risk inherent in traditional single-RAID cards; through heartbeat monitoring, master / slave election, and automatic failover and fallback mechanisms, it ensures uninterrupted data access even when any RAID card or link fails, significantly improving system reliability; it leverages front-end multi-path (host-side aggregated access) and back-end multi-path (dual-port hard drive redundant links), combined with the ALUA mechanism, to achieve IO load sharing; and in alternative solutions, it supports Active-Active access mode through CXL.mem technology, further enhancing IO throughput and fully utilizing hardware resources; it simplifies operation and maintenance management and reduces manual intervention costs through the cluster module's configuration synchronization function (automatic synchronization to slave nodes after master node metadata modification) and flexible arbitration mechanisms (arbitration disk or host memory); and it optimizes read and write performance through host memory caching (HDD only) and CXL.mem shared caching, while eliminating the need for SSD caching, achieving a balance between reliability and cost. Attached Figure Description

[0015] The accompanying drawings are provided to further illustrate embodiments of the present invention and form part of the specification. They are used together with the following detailed description to explain the embodiments of the present invention, but do not constitute a limitation thereof. In the drawings: Figure 1 This is a structural diagram of the RAID card high availability device based on multi-path clustering provided in the embodiments of the present invention; Figure 2 This is a flowchart of a RAID card high availability method based on multi-path clustering provided in an embodiment of the present invention; Figure 3 This is a flowchart of the fault switching operation in an embodiment of the present invention; Figure 4 This is a flowchart of the fault triggering and switching back operation in an embodiment of the present invention; Figure 5 This is a schematic diagram of the cluster configuration synchronization steps in an embodiment of the present invention; Figure 6 This is a schematic diagram of the structure based on CXL.mem technology in an embodiment of the present invention. Detailed Implementation

[0016] The specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings. It should be understood that the specific embodiments described herein are for illustration and explanation only and are not intended to limit the scope of the present invention.

[0017] In this invention, the terms "system" and "network" are used interchangeably. "Multiple" refers to two or more; therefore, in this invention, "multiple" can also be understood as "at least two." "And / or" describes the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A alone, A and B simultaneously, or B alone. Additionally, the character " / ", unless otherwise specified, generally indicates that the preceding and following related objects have an "or" relationship. Furthermore, it should be understood that in the description of this invention, terms such as "first" and "second" are used only for descriptive purposes and should not be construed as indicating or implying relative importance or order.

[0018] This invention provides a high-availability RAID card device based on multi-path clustering, comprising: at least two RAID cards connected to a server motherboard via a PCIe bus; a RAID engine for processing RAID read / write data, configuring RAID levels, and handling hard drive hot-swapping; a front-end multi-path module for establishing multiple data channels with the host, transmitting read / write data and commands, and exposing virtual disks with the same NGUID to the host, enabling the host multi-path module to aggregate access to the virtual disks; a back-end multi-path module for establishing multiple data channels with dual-port hard drives, transmitting read / write data and commands from the dual-port hard drives, and achieving disk access link redundancy and automatic fault switching; a system management module for handling internal and external management and control events, state transitions, and internal processes of the RAID cards; and a cluster management module for establishing communication channels between RAID cards, maintaining heartbeat information between RAID cards, and managing the status of cluster nodes.

[0019] In embodiments of the present invention, such as Figure 1As shown, RAID card A and RAID card B are both connected to the motherboard's PCIe bus. The host processor (hereinafter referred to as CPU) sends IO requests to RAID card A and RAID card B via the motherboard's PCIe bus through various protocols such as NVMe / SATA / SAS. Both RAID card A and RAID card B include a RAID engine, a front-end multipath module, a back-end multipath module, a system management module, and a cluster management module. The RAID engine is used to process RAID read and write data, configure RAID levels (supporting RAID0 / 1 / 10 / 5 / 6 / 50 / 60), and handle hard drive hot-swapping. The system management module is used to handle internal and external management control events, state transitions, and internal processes of the RAID card, and is responsible for the corresponding RAID card's own configuration (such as cache and IO policies), status monitoring (temperature and faults), and management interaction with the host or cluster.

[0020] Specifically, both RAID card A and RAID card B communicate with the host CPU through their own front-end multipath modules, receiving I / O requests from the CPU. The front-end multipath modules expose virtual disks with the same NGUID to the host, allowing the host's multipath modules to see two virtual disks with identical content. These virtual disks can be accessed via aggregated access from different paths, and the access method can be a preferred mode. Furthermore, all dual-port hard drives (HDDs / SSDs) are connected to the server's backplane PCIe or SAS bus (backplane PCIe / SAS bus). RAID card A and RAID card B communicate through the back-end multipath modules and via NVMe / SATA / SAS... The protocol establishes multiple data channels for communication with dual-port hard drives (HDD / SSD), transmitting disk read / write data and commands. Since the same disk can be accessed by two RAID cards, it is necessary to distinguish between RAID card I / O conflicts. The backend multipath module can isolate I / O conflicts through the ALUA mechanism, specifying the working controller node attributes of the virtual disk. The host can only access the corresponding virtual disk on the specified working controller. Different virtual disks have different working controller nodes to ensure that the same virtual disk is only "actively" accessed by one controller. Multiple virtual disks can be statically assigned to different controllers to achieve the purpose of distributing I / O load. The backend multipath module also supports automatic switching in case of link failure, that is, when a link failure is detected, the disk access path can be automatically switched to the backup access path.

[0021] Furthermore, the cluster management modules of RAID card A and RAID card B are interconnected via a PCIe network, enabling P2P communication. This serves as a channel for RAID card cluster communication, transmitting heartbeat messages and configuration management synchronization messages, and managing the status of cluster nodes, thereby establishing a high-availability RAID card cluster. Due to the special nature of the DDF metadata of RAID card A and RAID card B, only one controller node can access this DDF metadata; this controller node is the master node. The determination of the master node is achieved through a master-slave competition process, using an election algorithm to elect a master / slave node. Once a master node is identified, and there is only one master node, the rest are slave nodes. This single master node is used to access DDF metadata, serialize event processing, and synchronize configurations. This means that subsequent system management processes such as event processing and task scheduling within the cluster are all completed in the master node, ensuring serialized event processing and uniqueness of states. It also supports failover; for example, when the master node fails, the slave nodes automatically take over the virtual disks of the master node to ensure uninterrupted service. It also supports configuration synchronization; after the master node modifies the virtual disk configuration (such as RAID level and caching strategy), it synchronizes the new configuration information to the slave nodes to maintain cluster consistency.

[0022] In this embodiment of the invention, the RAID card high-availability device based on multi-path clustering also includes a cache area. RAID card A and RAID card B each have their own corresponding cache area in the host memory as a buffer for read and write acceleration. This cache area only applies to HDDs, while for SSDs, due to their high performance, this cache area is disabled. A disk is fixedly selected as the arbitration disk in the management backend hard drive. In cluster mode, the arbitration disk is used as the arbitration device for the "master-slave election" between RAID card A and RAID card B. In the event of a split-brain event, the cluster management module competes for the arbitration disk lock through the SCSI Reserve command. The winning RAID card survives, and the losing RAID card leaves the cluster.

[0023] Please see Figure 2 Based on the same inventive concept, embodiments of the present invention provide a method for high availability of RAID cards based on multi-path clustering, applied to the RAID card high availability device provided in embodiments of the present invention, comprising the following steps: S1: Multiple RAID cards are combined into a cluster. The cluster management module performs master and slave node election to determine the master node and slave node. The master node is responsible for metadata management and event processing. S2: The front-end multipath module exposes virtual disks with the same NGUID to the host, enabling the host to access the virtual disks through the multipath module. S3: The backend multipath module connects to a dual-port hard drive and specifies the working controller of the virtual disk through the ALUA mechanism, restricting the host to access the corresponding virtual disk only through the working controller. S4: The cluster management module maintains heartbeat information between RAID cards through communication channels and monitors node status in real time; S5: When a RAID card failure is detected, a failover operation is triggered, switching the virtual disk belonging to the failed RAID card to the normal RAID card for management; when the failed RAID card recovers, a failback operation is triggered, and the virtual disk returns to its original RAID card.

[0024] In this embodiment of the invention, for step S1, multiple RAID cards are grouped into a cluster. The multiple RAID cards are interconnected via PCIe and can communicate in P2P, serving as a communication channel between the RAID cards in the cluster. In the cluster, the cluster management module performs master and slave node election. Specifically, the cluster management module initiates a master and slave contention process and determines the election result through an election algorithm. The election result is used to divide the virtual disk's working controller to achieve IO load sharing. Specifically, the master node and slave node in the cluster are determined. The master node is responsible for metadata management and event processing. The important function of the cluster management module is to maintain the state of the cluster, the entry and exit of nodes in the cluster, the switching of master and slave identities of controllers, and the maintenance of the cluster state. The interaction between cards is mainly carried by messages on the cluster link. If the cluster link fails, the cluster is in a split-brain state. At this time, the backend needs an arbitration medium to determine which controller node needs to live and which node needs to quit the cluster. This embodiment of the invention adopts a simple arbitration disk method. A disk is fixedly selected in the management backend hard drive as the arbitration disk to ensure that both RAID cards can access it. By storing the reserve arbitration command, the RAID card that first obtains the disk is the master node, and the remaining nodes quit the cluster.

[0025] In this embodiment of the invention, for S2, virtual disks with the same NGUID are exposed to the host through the front-end multipath module, so that the host can access the virtual disks through the multipath module. That is, by exposing virtual disks with the same NGUID to the host, the host's multipath module can see two virtual disks with the same content and can access them from different paths. The access method can be the preferred mode.

[0026] In this embodiment of the invention, for S3, a dual-port hard drive (HDD / SSD) is connected through a backend multipath module, and the working controller of the virtual disk is specified through the ALUA mechanism to restrict the host to access the corresponding virtual disk only through the working controller. Specifically, the ALUA mechanism avoids access conflicts between multiple RAID cards by statically assigning ownership of virtual disks. For S4, the cluster management module maintains heartbeat information between RAID cards through a communication channel to monitor the node status in real time. Specifically, the heartbeat information of the RAID card is implemented by periodically sending status identifiers and timestamps. A preset number of times is set, and if no heartbeat information is received for a preset number of consecutive times, the corresponding node is determined to be faulty.

[0027] In this embodiment of the invention, for S5, when a RAID card failure is detected, a failover operation is triggered, switching the virtual disk belonging to the failed RAID card to the managed RAID card; when the failed RAID card recovers, a failback operation is triggered, and the virtual disk returns to its original RAID card; for example... Figure 3 As shown, the host initiates a "write to disk 0" request and recognizes that this request requires access to virtual disk 0 of RAID card A. The host initiates a "write to disk 1" request and recognizes that this request requires access to virtual disk 1 of RAID card B. The host initiates a "write to disk 2" request and recognizes that this request requires access to virtual disk 2 of RAID card A. The host initiates a "write to disk 3" request and recognizes that this request requires access to virtual disk 3 of RAID card B. The read and write operations of virtual disk 0 and virtual disk 2 are mapped to physical disk 0 for execution, while the read and write operations of virtual disk 1 and virtual disk 3 are mapped to physical disk 2 for execution. Therefore, virtual disks 0 and 2 belong to RAID card A, and virtual disks 1 and 3 belong to RAID card B. The dashed line represents the state before RAID card B fails, and the solid line represents the state after RAID card B fails. When RAID card B fails, a failover operation is triggered, switching virtual disks 1 and 3 from RAID card B to RAID card A for management. At this time, virtual disks 0, 2, 1, and 3 all belong to RAID card A, meaning that read and write operations on virtual disks 1 and 3 are also mapped to physical disk 0 for execution. Figure 4 As shown, the logical access relationships on the host side and the logical association of the RAID card mapping virtual disk read / write operations to actual physical disks are... Figure 3 The dashed line represents the state of RAID card B before recovery, and the solid line represents the state of RAID card B after recovery. When the faulty RAID card B recovers, a failover operation is triggered, returning virtual disk 1 and virtual disk 3 to their original RAID card B. At this time, virtual disk 0 and virtual disk 2 belong to RAID card A, and virtual disk 1 and virtual disk 3 belong to RAID card B again. The read and write operations of virtual disk 1 and virtual disk 3 are once again mapped to physical disk 2 for execution.

[0028] In this embodiment of the invention, the high availability method for RAID cards based on multi-path clusters provided by the present invention further includes a cluster configuration synchronization step. When a user initiates a configuration operation through a management tool, the configuration command is sent to the master node. After the master node modifies the physical disk metadata, it notifies the slave nodes through the cluster management module. The slave nodes read and synchronize the configuration data from the physical disk metadata area. Specifically, as shown... Figure 5 As shown, the server initiates a "Write to Disk 0" request and recognizes that this request needs to access virtual disk 0 of RAID card A. The server initiates a "Write to Disk 1" request and recognizes that this request needs to access virtual disk 1 of RAID card B. The server initiates a "Write to Disk 2" request and recognizes that this request needs to access virtual disk 2 of RAID card A. The server initiates a "Write to Disk 3" request and recognizes that this request needs to access virtual disk 3 of RAID card B. The read and write operations of virtual disk 0 and virtual disk 2 are mapped to physical disk 0 for execution, and the read and write operations of virtual disk 1 and virtual disk 3 are mapped to physical disk 2 for execution. The user connects to RAID card A and RAID card B through the RAID card management tool on the server. The configuration steps are as follows: (1) If the RAID card management tool selects the cluster management module of RAID card A as the cluster management master node, then the cluster management module of RAID card B is the cluster management slave node, and sends configuration operation commands to the cluster management master node; (2) The cluster management master node receives the configuration command and sends the configuration command to the card A system management module; (3) The system management module of card A parses the configuration command, writes the configuration modification to the metadata area of ​​the physical disk, and returns the configuration result to the cluster management master node; (4) The cluster management master node of card A sends a configuration change notification to the cluster management slave node of card B; (5) The cluster management of card B is notified by the node that the configuration of the card B system management module has changed; (6) The system management module of card B reads the configuration data written by card A from the metadata area of ​​the physical disk.

[0029] In this embodiment of the invention, in order to simplify the design of the RAID card, an SSD cache solution is used instead of a DDR cache inside the RAID card, which also simplifies the circuit design of the supercapacitor.

[0030] In embodiments of the present invention, such as Figure 6As shown, when the communication channel between RAID cards is replaced with host shared memory based on CXL.mem technology, the host memory is shared via PCIe, and consistent memory access can be maintained. This allows for real-time I / O forwarding. Simultaneously, the cluster's heartbeat and arbitration media are also implemented through shared host memory. This allows the virtual disk to be divided into multiple fixed-size fragments, with controller assignment determined by the allocated address. The front-end multipath module calculates I / O assignment using a simple hash algorithm and forwards the I / O to the corresponding controller node, supporting Active-Active access mode for the virtual disk. Therefore, from the host's perspective, the virtual disk has no working controller, making the multipath strategy more flexible. Besides configuring it in preferred mode, it can also be configured in round-robin mode.

[0031] Furthermore, using CXL.mem technology, a dedicated memory area is selected in the host memory. This dedicated memory area is used for communication between RAID cards, and the communication data between RAID cards includes the following: Cluster message queue: Uses two circular queues to achieve bidirectional message communication; IO forwarding queue: The front-end multipath module forwards host IO requests that do not belong to the local machine to the remote end, supports the non-ownership Active-Active access mode of virtual disks, and also uses two circular queues to achieve bidirectional data forwarding. Arbitration data area: The atomic memory access instruction atomic_fetch_add defined in CXL3.0 is used for operation. The cluster master node is responsible for initialization. When the cluster message queue fails, the RAID card node avoids calling atomic_fetch_add=0 and determines whether it can grab the group master role first. Shared buffer: Used to cache RAID card read and write data, accelerating read and write response speed, and can be shared by multiple RAID card nodes at the same time.

[0032] In this embodiment of the invention, using CXL.mem as the cluster management and high-speed data forwarding channel can better achieve scalability. With NVM or SAS switching networks used in the backend, the number of RAID card cluster nodes can be further expanded. The corresponding cluster message queues and IO forwarding queues in the CXL.mem host memory need to be increased accordingly, with the corresponding queue numbers being [number missing]. , where n is the number of RAID card nodes in the cluster.

[0033] This invention provides a high-availability device and method for RAID cards based on multi-path clustering. Through multi-RAID card clustering design, front-end and back-end multi-path redundancy, and intelligent fault handling mechanisms, it brings several beneficial effects: First, it completely eliminates the single-point-of-failure risk of traditional single-RAID cards. Through heartbeat monitoring, master / slave election, and automatic failover and failback mechanisms, it ensures uninterrupted data access when any RAID card or link fails, significantly improving system reliability. Second, it leverages front-end multi-path (host-side aggregated access) and back-end multi-path (dual-port hard drive redundant links) combined with the ALUA mechanism to achieve IO load sharing. It also supports Active-Active communication through CXL.mem technology. The system offers several advantages: First, it improves IO throughput and fully utilizes hardware resources. Second, it simplifies operation and maintenance by using the cluster module's configuration synchronization function (automatic synchronization to slave nodes after the master node modifies metadata) and a flexible arbitration mechanism (arbitration disk or host memory), reducing manual intervention costs. Third, it is compatible with mainstream levels such as RAID0 / 1 / 10 and protocols such as NVMe and SAS, and supports expansion from dual-node to multi-node clusters, adapting to different storage scenarios. Fourth, it optimizes read and write performance through host memory caching (HDD only) and CXL.mem shared caching, while eliminating the need for SSD caching and supercapacitor design, achieving a balance between reliability and cost.

[0034] The optional embodiments of the present invention have been described in detail above with reference to the accompanying drawings. However, the embodiments of the present invention are not limited to the specific details in the above embodiments. Within the scope of the technical concept of the embodiments of the present invention, various simple modifications can be made to the technical solutions of the embodiments of the present invention, and these simple modifications all fall within the protection scope of the embodiments of the present invention.

[0035] It should also be noted that the various specific technical features described in the above embodiments can be combined in any suitable manner without contradiction. To avoid unnecessary repetition, the embodiments of the present invention will not describe the various possible combinations separately.

[0036] Those skilled in the art will understand that all or part of the steps in the methods of the above embodiments can be implemented by a program instructing related hardware. This program is stored in a storage medium and includes several instructions to cause a microcontroller, chip, or processor to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as a USB flash drive, a portable hard drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.

[0037] Furthermore, various different implementations of the present invention can be combined arbitrarily, as long as they do not violate the spirit of the present invention, they should also be regarded as the content disclosed in the present invention.

Claims

1. A high-availability device for RAID cards based on multi-path clustering, characterized in that, include: At least two RAID cards are connected to the server motherboard via the PCIe bus; The RAID engine is used to handle RAID read and write data, configure RAID levels, and handle hard drive hot-swapping. The front-end multipath module is used to establish multiple data channels with the host, transmit read and write data and commands, expose virtual disks with the same NGUID to the host, and enable the host multipath module to aggregate access to the virtual disks. The backend multipath module is used to establish multiple data channels with the dual-port hard drive, transmit the read and write data and commands of the dual-port hard drive, and realize disk access link redundancy and automatic fault switching. The system management module is used to handle management and control events, state transitions, and internal processes of the RAID card, both internal and external. The cluster management module is used to establish communication channels between RAID cards, maintain heartbeat information between RAID cards, and manage the status of cluster nodes.

2. The RAID card high-availability device according to claim 1, characterized in that, The backend multipath module isolates I / O conflicts, specifies the working controller node attributes of the virtual disk, and supports automatic switching in case of link failure. When a link failure is detected, it automatically switches to the backup access path.

3. The RAID card high-availability device according to claim 1, characterized in that, The cluster management module is used to establish communication channels between RAID cards, maintain heartbeat information between RAID cards, and manage the status of cluster nodes, specifically including: A communication channel is established between RAID cards via the PCIe P2P protocol. This communication channel transmits cluster heartbeat messages and cluster configuration messages between the RAID cards. The primary and backup node election is carried out through a primary and backup contention process to determine a unique primary node for accessing DDF metadata, serializing event processing, and synchronizing configuration. In response to a split-brain event, a dual-port hard drive is designated as the arbitrator disk. The RAID card competes for the arbitrator disk lock using the SCSI Reserve command. The winning RAID card survives, while the losing RAID card leaves the cluster.

4. The RAID card high-availability device according to claim 1, characterized in that, It also includes a cache area, which is only effective for HDD and is used to accelerate read and write operations; For SSDs, the cache region is disabled.

5. A high-availability method for RAID cards based on multi-path clustering, characterized in that, The high-availability device for a RAID card according to any one of claims 1-4 includes the following steps: S1: Multiple RAID cards are combined into a cluster. The cluster management module performs the election of master and slave nodes to determine the master node and slave node. The master node is responsible for metadata management and event processing. S2: The front-end multipath module exposes virtual disks with the same NGUID to the host, enabling the host to access the virtual disks through the multipath module. S3: The backend multipath module connects to a dual-port hard drive and specifies the working controller of the virtual disk through the ALUA mechanism, restricting the host to access the corresponding virtual disk only through the working controller. S4: The cluster management module maintains heartbeat information between RAID cards through communication channels and monitors node status in real time; S5: When a RAID card failure is detected, a failover operation is triggered, switching the virtual disk belonging to the failed RAID card to the normal RAID card for management; when the failed RAID card recovers, a failback operation is triggered, and the virtual disk returns to its original RAID card.

6. The method according to claim 5, characterized in that, S1 specifically includes: The primary / standby node election is initiated by the cluster management module through a primary / standby competition process. The election result is determined by an election algorithm. The election result is used to allocate the work controllers of the virtual disks to achieve IO load sharing. In S3, the ALUA mechanism avoids access conflicts from multiple RAID cards by statically assigning ownership of virtual disks.

7. The method according to claim 5, characterized in that, In step S4, the heartbeat information is achieved by periodically sending status identifiers and timestamps. If no heartbeat information is received for a preset number of consecutive times, the corresponding node is determined to be faulty. In S5, during the fault switching and fault recovery process, the virtual disk status and configuration data are synchronized through the system management module.

8. The method according to claim 5, characterized in that, It also includes a cluster configuration synchronization step. When a user initiates a configuration operation through the management tool, the configuration command is sent to the master node. After the master node modifies the physical disk metadata, it notifies the slave nodes through the cluster management module. The slave nodes read and synchronize the configuration data from the physical disk metadata area.

9. The method according to claim 5, characterized in that, The communication channel between RAID cards can be replaced by host shared memory based on CXL.mem technology. The host shared memory includes a cluster message queue, an IO forwarding queue, an arbitration data area, and a shared cache area. It supports the Active-Active access mode of virtual disks and allocates IO requests to the corresponding RAID cards through a hash algorithm.

Citation Information

Cited By

  • Task-level redundancy fault-tolerant system based on RTOS kernel

    CN122111766A