Systems and methods for improving data center availability using rack-to-rack storage link cables
By using rack-to-rack storage link cables within the data center and combining RAID parity and system-level erase codes, the availability and durability issues of data center storage architecture under various failure scenarios are solved, achieving higher data storage robustness and cost efficiency.
Patent Information
- Application Number
- CN202111638606.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2020-12-31
- Filing Date
- 2021-12-29
- Publication Date
- 2025-10-17
- Estimated Expiration
- 2041-12-29
AI Technical Summary
Existing data center storage architectures struggle to guarantee high availability and durability when faced with various failure scenarios, especially sudden space failures and sudden space-free failures. Existing redundancy and parity solutions are insufficient to prevent data loss in certain situations.
By employing rack-to-rack storage link cables, point-to-point connections are established between computing units to provide fault backup, ensuring that storage drives remain accessible even in the event of a computing unit failure. Combined with RAID parity checking and system-level erase codes, this enhances the availability and durability of data storage.
It improves the availability and durability of data center storage systems under various failure scenarios, reduces the risk of data loss, lowers hardware costs and complexity, and enhances system robustness.
Smart Images

Figure CN114691432B_ABST
Abstract
Description
SUMMARY
[0001] The present disclosure relates to a method, system, and apparatus for improving data center availability using rack-to-rack storage link cables. In one embodiment, a first data storage rack has a first compute unit coupled to a first plurality of storage drives via a first storage controller. A second data storage rack has a second compute unit coupled to a second plurality of storage drives via a second storage controller. A first rack-to-rack storage link cable couples the first compute unit to the second storage controller such that the first compute unit is able to provide access to the second plurality of drives in response to a failure that prevents the second compute unit from providing access to the second plurality of drives via a system network.
[0002] In another embodiment, a method involves coupling first and second compute units of first and second data storage racks to a system network. The first and second compute units provide access to respective first and second pluralities of drives in the first and second data storage racks via the system network. A first failure of the second compute unit is detected that prevents the second compute unit from providing access to the second plurality of drives via the system network. In response to detecting the first failure, the first compute unit is coupled to the second plurality of drives via a first rack-to-rack storage link cable. Access to the second plurality of drives via the system network is provided via the first compute unit after the first failure. These and other features and aspects of various embodiments can be appreciated from the following detailed discussion and from the drawings.
[0003] BRIEF DESCRIPTION OF DRAWINGS
[0004] The following discussion discusses reference to the following drawings, wherein like numerals can be used to identify like / components in multiple drawings.
[0005] Figure 1 and Figure 2 is a block diagram of a data center system in accordance with example embodiments;
[0006] Figure 3 , Figure 4 and Figure 5 illustrates data system failures that a data center system in accordance with example embodiments can encounter;
[0007] Figure 6 and Figure 7 is a schematic diagram of a data center system in accordance with example embodiments;
[0008] Figure 7 is a diagram illustrating use of spare storage with reduced areal density disk surfaces in accordance with example embodiments;
[0009] Figure 8is a block diagram illustrating functional components of a system and device according to example embodiments; and
[0010] Figure 9 is a flowchart of a method according to example embodiments. DETAILED DESCRIPTION
[0011] This disclosure generally relates to data centers. A data center is a facility (e.g., a building) that houses large numbers of computer systems and related components such as network infrastructure and data storage systems. Many modern data centers, also known as cloud data centers, are large computing facilities connected to the public Internet and used to serve a wide variety of applications such as cloud storage, cloud computing, website hosting, e-commerce, messaging, etc. In this disclosure, embodiments relate to data storage services within large-scale data centers.
[0012] The storage capacity of a modern data center can reach hundreds of petabytes. This is often provided as a cloud storage service on the Internet. One advantage of using a data center for cloud storage is that the economies of scale can make storage on a data center much less expensive than maintaining one’s own data storage facility. In addition, a data center can employ state-of-the-art protection of storage media, ensuring the availability and durability of data even in the event of equipment failure.
[0013] Generally, availability relates to redundancy in storage nodes and compute nodes, such that backup computers and / or storage devices can quickly replace a failed unit, often without human intervention. Durability relates to the ability to recover from lost portions of stored data, e.g., due to storage device failure, data corruption, etc. Durability can be improved through the use of redundant data such as parity data and erasure codes. The concepts of availability and durability are somewhat related but can be independent in certain scenarios. For example, if a central processing unit (CPU) of a single-CPU data storage rack fails, all of the storage provided by the rack can be unavailable. However, assuming the CPU failure did not damage the storage devices, the data in this scenario can still be safe and thus does not negatively impact durability. For the purposes of this disclosure, the term “reliability” can be used to describe both availability and durability.
[0014] In the embodiments described below, strategies are described that can improve the availability of data center storage beyond what is provided by existing architectures. These strategies can be used with enhanced durability schemes such that data center storage becomes more reliable in the face of a number of different failure scenarios. These strategies can be used with known storage architectures such as Lustre, PVFS, BeeGFS, Cloudian, ActiveScale, SwiftStack, Ceph, HDFS, etc.
[0015] In Figure 1 and Figure 2 example server architectures that can utilize features according to example embodiments are shown. In Figure 1 a block diagram illustrates a number of server storage racks 100 according to a first example embodiment. Rack 101 will be described in further detail, and Figure 1 all racks 100 in can have the same or similar features. Rack 101 is configured as a high-availability (HA) storage server, and utilizes a first computing unit 102 and a second computing unit 103, which are indicated as servers in the figures. For the purposes of this discussion, the term "computing unit" or "server" means an integrated computing device (e.g., board, chassis) having at least one CPU, memory, and an input / output (I / O) bus. Computing units / servers 102, 103 will also have one or more network interfaces that couple the rack to a data center network.
[0016] A number of storage drives 104 (e.g., hard disk drives, solid state drives (SSDs)) are located in rack 100. Note that the term "drive" in this disclosure does not imply a restriction on the type or form factor of the storage medium, nor on the enclosure and interface circuitry of the drive. In some embodiments, the number of storage drives 104 can exceed one hundred per rack, and the storage drives 104 are coupled to one or more storage controllers 106. In this example, storage controller 106 is shown as a Redundant Array of Independent Disks (RAID) controller that uses dedicated hardware and / or firmware to arrange the storage devices into one or more RAID virtual devices. Generally, this involves selecting a number of drives and / or a number of drive partitions to assemble into a larger virtual storage device (sometimes referred to as a volume). Depending on the type of RAID volume (e.g., RAID 1, RAID 5, RAID 6), redundancy and / or parity can be introduced to improve the durability of the volume. Typically, large data centers can use RAID 6, and / or can use proprietary or non-standard schemes (e.g., non-aggregated parity).
[0017] The HA compute units 102, 103 are each connected to drives, with an arrangement that allows the rack to continue to operate in the event of a failure of one of the compute units 102, 103. For example, the storage drives 104 can be divided into two groups, each coupled to a different storage controller 106. The first HA compute unit 102 is designated as the primary unit for the first storage controller 106, and as the secondary unit for the second storage controller 106. The second HA compute unit 103 is designated as the primary unit for the second storage controller 106 and as the secondary unit for the first storage controller 106. The HA compute units 102, 103 monitor each other’s activities to detect signs of failure. If one of the HA compute units 102, 103 is detected as failing, the other HA compute unit 102, 103 will take over as the primary for the failed unit, thus maintaining availability of the entire rack 101. Other ways of coupling the compute units 102, 103 in an HA arrangement can exist, for example, using a single storage controller 106, more than two storage controllers 106, etc., and this example is presented for illustration and not limitation.
[0018] At the system level, data 108 for the storage rack 100 can be distributed across racks, for example, using a round-robin segment storage scheme 110. In this scheme 110, data units (e.g., files, storage objects) are divided into portions that are stored on different racks. This type of arrangement is used in storage architectures such as Lustre, PVFS, BeeGFS, etc., and can be applicable to both file-based and object-based storage systems.
[0019] In Figure 2 The block diagram illustrates a number of server storage racks 200 according to a second example embodiment. The rack 201 will be described in further detail, and Figure 2 All racks 200 in
[0020] In contrast to the arrangement in Figure 1 The arrangement in Figure 2 The arrangement in Figure 2The system in
[0045] can utilize what is referred to herein as software reliability, where reliability is provided by a network component (e.g., storage middleware) that stores data 208 between racks 200 using erasure coding and replication 210. Typically, this involves the storage middleware partitioning a data object into portions distributed across multiple racks. The storage middleware also uses erasure coding to calculate and separately store redundant portions in other parts of the rack. The erasure code data can be used to recover lost portions of data due to drive or rack failures. This type of arrangement is used in storage architectures such as Cloudian, ActiveScale, SwiftStack, Ceph, and HDFS.
[0021] Figure 1 and Figure 2 The two different arrangements shown in
[15] may have their particular strengths related to reliability, as neither approach can protect against all failure events. In particular, two large-scale failure modes can be considered when analyzing the reliability of these architectures. The first failure mode is spatial failure bursts, which involve multiple simultaneous drive failures within a single rack. Figure 3 An example of this pattern is shown, where an X represents a failed drive within chassis 300 . Figure 4 Another example that can also be considered as a spatial failure burst is shown in FIG, where a single server in rack 400 fails. In this example, data is not necessarily lost, but it is unavailable. Figure 4 The situation shown could also be caused by some other single point of failure, such as power delivery, cooling, storage controller, etc. This can be solved by using Figure 2 The wiping between the shells shown prevents Figure 3 and Figure 4 As shown in the spatial fault burst. Figure 1 The parity within the enclosure shown may not be sufficient to protect against sudden bursts of space failures, for example, if enough drives within the RAID volume fail, or if all compute units / servers within a rack fail.
[0022] The second form of failure mode is a burst of unspaced failures, where multiple drive failures occur simultaneously across multiple racks. Figure 5 An example of this failure mode is shown in FIG, where x represents a failed drive in storage rack 500-502. Figure 1 As shown, erase / parity is used within the storage frame to prevent sudden failures without space, but as Figure 2 Using scrub on a storage rack as shown may not be sufficient. For example, Figure 2 Some software reliability configuration recommendations are as follows Figure 2Each compute unit / server in the illustrated setup does not exceed 12 drives. However, by today's standards, 12 drives per enclosure is a small number and this small number reduces the cost efficiency of the data storage unit, which is an advantage of this type of architecture. For inexpensive mass storage, to achieve cost targets, it can be desirable to have more than 100 drives per compute unit / server.
[0023] In Figure 6 In the illustrated setup, the first compute unit 604 is coupled to the second plurality of storage drives 609 via a first compute-to-drive storage link cable 615. The second compute unit 605 is coupled to the first plurality of storage drives 608 via a second compute-to-drive storage link cable 617. The third compute unit 606 is coupled to the third plurality of storage drives 610 via a third compute-to-drive storage link cable 619. The first, second, and third compute-to-drive storage link cables 615, 617, and 619 are coupled to the first, second, and third storage controllers 612, 613, and 614, respectively. The first, second, and third storage controllers 612, 613, and 614 are coupled to the first, second, and third pluralities of storage drives 608, 609, and 610, respectively. The first, second, and third compute units 604, 605, and 606 are coupled to the network 631.
[0024] The first rack-to-rack storage link cable 616 couples the first compute unit 604 to the second plurality of drives 609 such that the first compute unit 604 can provide access to the second plurality of drives 609 in response to a first failure that prevents the second compute unit 605 from accessing the second plurality of drives 609. The second rack-to-rack storage link cable 618 couples the third compute unit 606 to the first plurality of drives 608 such that the third compute unit 606 can provide access to the first plurality of drives 608 in response to a second failure that prevents the first compute unit 604 from accessing the first plurality of drives 608. Note that the failures described above for the first and second compute units 604 and 605 can include failures of the compute units themselves (e.g., CPU, memory, I / O), failures of the links between the compute units and the drives and / or storage controllers, power supply failures that affect the drives and / or compute units, etc.
[0025] The storage link cables 616, 618 can be any type of cable that is compliant with a point-to-point storage protocol, and that can operate over rack-to-rack distances, including SATA, SaS, SCSI, Fibre Channel, Ethernet, etc. Note that Ethernet cables can be configured to run point-to-point, e.g., without intervention of a switch, in which case the cables can be configured as crossover Ethernet cables. Generally, these cables 616, 618 can be distinguished from the network cabling that typically couples the compute units 604-606 to the system network 631. Since the storage controllers 612-614 can present each drive array 608-610 as a single storage device, only one cable 616, 618 can be needed, although more cables can be used for redundancy, performance, and / or to account for multiple storage controllers, as discussed in more detail below.
[0026] Figure 6 Also shown in FIG. 6 is an additional data storage rack 621-623 coupled to the network 631, each rack having a compute unit 624-626, respectively, coupled to and providing access to a plurality of storage drives 628-630 via a storage controller 632-634, respectively. Rack-to-rack storage link cables 640-642 couple the compute units 624-626 to the drive arrays 633, 634, and 614. Thus, in this illustrative example of six data storage racks 601-603, 621-623, each compute unit 604-606, 624-626 has a backup unit while still requiring only one compute unit / server per rack. When a compute unit fails, the disks controlled by that unit can remain available via the backup. Note that the primary and backup compute units can not be strictly considered "paired" since here a circular order HA is enabled. In this arrangement, each compute unit backs up the adjacent compute unit, with the "last" compute unit 626 backing up the "first" compute unit 606. This circular order arrangement allows an odd group of compute units to provide HA backup to each other.
[0027] In other embodiments, the data storage racks 601-603, 621-623 can be coupled in pairs, similar to the intra-rack HA servers. This is indicated by the dashed lines 619, which represent optional storage link cables that couple the compute units 604 to the disks 610. The optional storage link cables 619 can be used in addition to the storage link cables 618 that provide backup to the disks 608, can be used instead of the storage link cables 616, or as an addition to the storage link cables 616. In other embodiments, multiple intra-rack backup groups can be formed, such that the racks within each group back up independently of the other groups, which helps limit the length of the storage link cables. For example, a set of 24 racks can be divided into four groups of six racks, with the six racks in each group arranged in a circular order backup arrangement, as shown in Figure 6
[0028] In Figure 7 the figure shows a data center system 700 according to an example embodiment. The system 700 includes a first data storage rack 701, a second data storage rack 702, and a third data storage rack 703, each having a first compute unit 704, a second compute unit 705, and a third compute unit 706, respectively. The compute units 704, 705, and 706 are generally coupled to a network 720. The first compute unit 701 is coupled to and provides access to portions 708a, 708b of storage drives 708, which are coupled to the first compute unit 704 through storage controllers 712a, 712b. The second compute unit 705 is coupled to and provides access to portions 709a, 709b of storage drives 709, which are coupled to the first compute unit 705 through storage controllers 713a, 713b. The third compute unit 706 is coupled to and provides access to portions 710a, 710b of storage drives 710, which are coupled to the first compute unit 706 through storage controllers 714a, 714b.
[0029] Each compute unit 704-706 has two rack-to-rack storage link cables that extend to drive array portions of two different racks, with the compute unit acting as a backup. For example, the compute unit 704 has a rack-to-rack storage link cable 716a that is coupled to the drive portion 714b of the rack 703 and a rack-to-rack storage link cable 716b that is coupled to the drive portion 713a of the rack 702. This pattern is repeated for the other compute units 705, 706, and also forms a circular order coupling, such that this backup arrangement is available for an odd number of racks. Note that the storage link cables 716a, 716b are optional, and the compute units 704, 705, 706 can be coupled to the storage drives 708-710 through the storage controllers 712-714, as shown by the solid lines 717-719. Figure 6 In contrast to the arrangement in the'1 1 1 patent, this arrangement can distribute the load of a failed compute unit to multiple other compute units while doubling the number of rack-to-rack cables per rack. For example, if the first compute unit 704 fails, portion 712a will be available on network 720 via the third compute unit 706, and portion 712b will be available on network 720 via the second compute unit 706.
[0030] In Figure 8 In the'1 1 1 patent, a block diagram illustrates internal components of a rack 802, 804 in a system 800 according to example embodiments. The components of rack 802 will be described in greater detail, and other racks in system 800 can be similarly or identically configured. Rack 802 includes a server / compute unit 806 having at least one CPU 807, random access memory (RAM) 808, a network interface card (NIC) 809, and an I / O interface 810. These components of compute unit 806 can be coupled via a motherboard or other circuitry known in the art.
[0031] I / O interface 810 is coupled to a storage controller 812, which includes its own controller circuitry 814, e.g., a system on a chip (SoC) that governs the operation of the controller. These operations include control of a drive array 816, which includes multiple persistent storage devices (e.g., HDDs, SSDs) that can be coupled to one or more circuit boards (e.g., backplanes) and can be arranged into storage cartridges. Rack 802 can include multiple instances of drive array 816 and / or storage controller 812.
[0032] Compute unit 806 and / or storage controller 812 can include HA control modules 818, 819 that enable the components of rack 802 to act as a storage control backup for one or more other racks via a storage link cable 820. One or both of HA control modules 818, 819 can further enable one or more other racks to act as a storage control backup for rack 802 via a storage link cable 824.
[0033] Another backup cable 823 is also shown, which can provide similar backup functionality for other racks not shown in this figure. For example, additional data storage racks can be serially coupled by respective rack-to-rack storage link cables. The compute unit of each rack can provide access to a next plurality of disk drives of a next rack in response to a failure of the next compute unit of the next rack. In this case, cable 823 provides backup for the drives of the first of these additional racks, and the last compute unit of these additional data storage racks provides backup for drives 816 via data link cable 824. In another arrangement, if data link cables 823, 824 are connected together, this would be a paired backup arrangement.
[0034] HA control modules 818, 819 can allow for self-healing of the storage rack, such that a centralized entity (e.g., storage middleware 832) need not be required to detect failures and assign backup servers. For example, HA module 818 can communicate with associated HA module 821 on rack 804, e.g., via system network 830. These communications can include keep-alive type messages that determine whether compute unit 826 of rack 804 is operational. Similar communications can be performed by the HA module on storage controller 812, 822 to see whether storage controller 822 is operational. If compute unit 826 is determined to be non-responsive but storage controller 822 is responsive, then compute unit 806 can take control of controller 822 and its associated disk array 825. This can also involve system controller 822 cutting any links with compute unit 826 to prevent any conflicts from occurring if compute unit 826 comes back online later.
[0035] Note that in order for compute unit 806 to take over for failed compute unit 826, it can be configured to act as a network proxy for the failed unit. For example, if compute units 806, 826 have network hostnames of "rack802" and "rack804," respectively, they can be configured to respond to network file system (NFS) uniform resource locators (URLs) "nfs: / / rack802: / path / to / data" and "nfs: / / rack804: / path / to / data," respectively. In this case, compute unit 806 can be configured to respond to NFS URLs of the form "nfs: / / rack804: / path / to / data" by responding with the contents of the file at "nfs: / / rack802: / path / to / data." Similarly, compute unit 822 can be configured to respond to NFS URLs of the form "nfs: / / rack802: / path / to / data" by responding with the contents of the file at "nfs: / / rack804: / path / to / data." <uuid1>" and " nfs: / / rack804: / <uuid2>” provides access to the arrays 816, 825, where UUID1 and UUID2 represent universally unique identifiers (UUIDs) provided by the respective storage controllers 812, 822 to identify their storage arrays 816, 825.
[0036] If computing unit 826 fails, the hostname "rack804" will most likely not respond to network requests. Therefore, computing unit 806 can be configured to also use the hostname "rack804," for example, by reconfiguring the domain name server (DNS) of network 830 to point the "rack804" hostname to the Internet Protocol (IP) address of computing unit 806. Because UUIDs are used to identify the respective arrays 816, 825 in URLs, the backup computing unit 806 can seamlessly take over network requests on behalf of the failed unit 826. Note that if computing unit 826 later comes back online in this scenario, there will be no confusion on the network due to NFS remapping, because computing unit 826 does not typically rely on its network hostname for internal network operations, but rather on "localhost" or the loopback IP address.
[0037] In other embodiments, aspects related to detecting failed computing units and allocating backups can be coordinated by a network entity, such as storage middleware 832. Generally, storage middleware 832 acts as a single, universal storage interface for clients 834 to access storage services in a data center. Storage middleware 832 can run on any number of computing nodes within a data center (including one or more dedicated computing nodes), as a distributed service, on clients 834, on one or more storage rack computing units in a storage rack, and so on. Storage middleware 832 can be optimized for different types of storage access scenarios encountered in large data centers, for example, optimizing for aspects such as data throughput, latency, and reliability. In this case, storage middleware 832 can include its own HA module 833 that communicates with HA modules on the rack (e.g., one or both of HA modules 818 and 819).
[0038] The activities of the middleware HA module 833 may be similar to those described in the self-healing example described above, e.g., keep-alive messages, remapping of network shares. Similar to the example of backup of NFS volumes described above, the middleware HA module 833 may detect the failure of the hostname "rack804". However, since the storage middleware 832 may abstract all access to storage on behalf of the client, it may change its internal mapping to account for the switch in the backup unit, e.g., nfs: / / rack804: / <uuid2>changed to nfs: / / rack802 / <uuid2>This can also be accompanied by a message to the HA module 818 of the rack 802 to assume control of the array 825 via the cable 820.
[0039] Note that management of the backup operation by the middleware HA module 833 can still involve some peer-to-peer communication via the storage rack components. For example, even if the middleware HA module 833 coordinates the remapping of network requests to the backup compute unit 806, the HA module 818 of the backup compute unit 806 can still communicate with the controller card 822 to take over the host storage interface for subsequent storage operations, and further to cut off the host interface to the failed compute unit 826 in case the latter comes back online. Note that if the failed compute unit 826 does come back online and appears to be fully operational again, the backup operation can be reversed to switch control of the drive array 825 back from the compute unit 806 to the compute unit 826.
[0040] In the embodiments described above, any type of local or system-level data durability scheme can be used to improve availability between data storage racks using storage link cables. As mentioned previously, one scheme involves using RAID parity such that a RAID volume can be protected against failure within the drives that make up the volume. Another scheme involves dividing data into multiple parts, computing erasure code data for each part, and distributing each part and erasure code data among different storage units (e.g., storage racks).
[0041] Another durability scheme that can be used with the illustrated scheme to improve availability between data storage racks is a hybrid scheme that uses both RAID parity and system-level erasure. For example, a system can be designed with n% of data overhead for redundancy, with a first amount n1% of the overhead dedicated to RAID parity and a second amount n2% of the overhead dedicated to erasure, where n1+ n2= n. This can make the system more robust against data loss in some failure scenarios shown in FIG. 3. More details of this hybrid durability scheme are described in U.S. Patent Application 16 / 794,951, filed February 19, 2020, entitled “Multi-Level Erasure System with Cooperative Optimization,” which is hereby incorporated by reference in its entirety. Figures 3-5
[0042] In Figure 9 In the middle, a flowchart illustrates a method according to an example embodiment. The method involves coupling first and second compute units of first and second data storage racks to a system network (step 900). The first and second compute units provide access to respective first and second pluralities of drives in the first and second data storage racks via the system network. For example, each plurality of drives can be coupled to a respective storage controller that presents the drives as a reduced number of virtual devices (e.g., RAID or JBOD volumes) to the compute units.
[0043] A failure of the second compute unit is detected (step 901), which prevents the second compute unit from providing access to the second plurality of drives via the system network. In response to detecting the failure, the first compute unit is coupled to the second plurality of drives via the first rack-to-rack storage link cable (step 902). Note that for this and other coupling instances, the storage link cable can already be physically connected between the respective compute units and the drives of the first and second racks. Thus, the coupling shown in this figure generally involves electrical and logical coupling between units that are already physically connected by the storage link cable. Access to the second plurality of drives is provided via the first compute unit after the first failure (step 903).
[0044] Figure 9 An optional step that can be performed by the system to protect the disks of the first data storage rack is also shown in the middle. As indicated by steps 904-906, one approach is to use a round-robin type of connection as described above. This can involve detecting a second failure of the first compute unit (step 904), which prevents the first compute unit from providing access to the first plurality of drives via the system network. In response to detecting the second failure, a third compute unit is coupled to the first plurality of drives via the second rack-to-rack storage link cable (step 905). The third compute unit provides access to a third plurality of drives in a third data storage rack via the system network. Access to the first plurality of drives is provided via the third compute unit after the second failure (step 906).
[0045] Another approach to protecting the disks of the first data storage rack is pairwise backup, as indicated by steps 907-909. This can involve detecting a second failure of the first compute unit (step 907), which prevents the first compute unit from providing access to the first plurality of drives via the system network. In response to detecting the second failure, the second compute unit is coupled to the first plurality of drives via the second rack-to-rack storage link cable (step 908). Access to the second plurality of drives is provided via the second compute unit after the second failure (step 909).
[0046] The various embodiments described above can be implemented using circuitry, firmware and / or software modules that interact with one another. Those skilled in the art will readily recognize the interchangeability of hardware and software modules and the fact that the described functionality can be implemented in either hardware or software, or a combination thereof. Given the teachings of the present disclosure, software and hardware modules particular to various embodiments are readily determined. For example, the flowcharts and control diagrams shown in this document can be used to create computer-readable instructions / code to be executed by a processor. Such instructions can be stored on a non-transitory computer-readable medium and conveyed to a processor to be executed as is known in the art. The structures and processes illustrated above are merely representative examples of embodiments that can be used to provide the functionality described above.
[0047] Unless otherwise indicated, all numbers expressing feature sizes, amounts, and other physical characteristics disclosed herein should be understood as approximations within a range of tolerances, and are intended to be read in a manner that includes the term "about." Unless otherwise indicated, the numerical parameters set forth in the foregoing description and associated drawings are approximations. Any numerical value, however, can inherently contain certain errors necessarily resulting from the standard deviation found in the respective testing measurements and methodologies used in producing the numerical values. The recitation of numerical ranges by endpoints includes all numbers subsumed within that range (e.g., 1 to 5 includes 1, 1.5, 2, 2.75, 3, 3.80, 4, and 5) and any range within that range.
[0048] The foregoing description of example embodiments has been presented for the purposes of illustration and description. It is not intended to be exhaustive or to limit the embodiments to the precise form disclosed. Many modifications and variations are possible in light of the above teaching. Any or all features of the disclosed embodiments can be applied individually or in any combination, without intending to limit the embodiments to the precise forms disclosed. The scope of the disclosure is not intended to be limited to the particular embodiments described herein, but rather, is intended to cover all modifications and variations falling within the scope of the appended claims.
[0049] Further examples:
[0050] Example 1. A system comprising:
[0051] a first data storage rack comprising a first compute unit coupled to a first plurality of storage drives via a first storage controller;
[0052] a second data storage rack comprising a second compute unit coupled to a second plurality of storage drives via a second storage controller; and
[0053] a first rack-to-rack storage link cable coupling the first compute unit to the second storage controller such that the first compute unit can provide access to the second plurality of drives in response to a first failure that prevents the second compute unit from providing access to the second plurality of drives via a system network.
[0054] Example 2. The system of example 1, further comprising:
[0055] a third data storage rack comprising a third compute unit coupled to a third plurality of storage drives via a third storage controller; and
[0056] a second rack-to-rack storage link cable coupling the third compute unit to the first storage controller such that the third compute unit can provide access to the first plurality of drives in response to a second failure that prevents the first compute unit from accessing the first plurality of drives.
[0057] Example 3. The system of example 2, further comprising a third rack-to-rack storage link cable coupling the second compute unit to the third storage controller such that the second compute unit can provide access to the third plurality of drives in response to a third failure that prevents the third compute unit from accessing the third plurality of drives.
[0058] Example 4. The system of example 1, further comprising:
[0059] two or more additional data storage racks, each of the two or more additional data storage racks serially coupled by a respective rack-to-rack storage link cable such that for each rack, the compute unit of each rack can provide access to a next plurality of drives of a next rack in response to a failure that prevents the next compute unit of the next rack from providing access to the next plurality of drives; and
[0060] a second rack-to-rack storage link cable coupling a last compute unit of a last of the two or more additional storage racks to the first storage controller such that the last compute unit can provide access to the first plurality of drives in response to a second failure that prevents the first compute unit from accessing the first plurality of drives.
[0061] Example 5. The system of example 1, further comprising a second rack-to-rack storage link cable coupling the second compute unit to the first storage controller such that the second compute unit can provide access to the first plurality of drives in response to a second failure that prevents the first compute unit from accessing the first plurality of drives.
[0062] Example 6. The system of example 1, wherein the first compute unit is operable to detect the first failure and provide access to the second plurality of drives via the system network in response to the first failure.
[0063] Example 7. The system of example 1, further comprising storage middleware coupled to the first compute unit and the second compute unit via the system network, the storage middleware operable to detect the first failure and instruct the first compute unit to provide access to the second plurality of drives via the system network in response to the first failure.
[0064] Example 8. The system of example 1, wherein the first rack-to-rack storage link cable comprises one of a SAS cable, a SATA cable, a SCSI cable, a fiber channel cable, or a point-to-point Ethernet cable.
[0065] Example 9. A method comprising:
[0066] coupling first and second compute units of first and second data storage racks to a system network, the first and second compute units providing access to respective first and second pluralities of drives in the first and second data storage racks via the system network;
[0067] detecting a first failure of the second compute unit, the first failure preventing the second compute unit from providing access to the second plurality of drives via the system network;
[0068] in response to detecting the first failure, coupling the first compute unit to the second plurality of drives via a first rack-to-rack storage link cable; and
[0069] providing access to the second plurality of drives via the first compute unit after the first failure.
[0070] Example 10. The method of example 9, further comprising:
[0071] detecting a second failure of the first compute unit, the second failure preventing the first compute unit from providing access to the first plurality of drives via the system network;
[0072] in response to detecting the second failure, coupling a third compute unit to the first plurality of drives via a second rack-to-rack storage link cable, the third compute unit providing access to a third plurality of drives in a third data storage rack via the system network; and
[0073] after the third failure, providing access to the third plurality of drives via the second computing unit.
[0074] Example 11. The method of example 10, further comprising:
[0075] detecting a third failure of the third computing unit that prevents the third computing unit from providing access to the third plurality of drives via the system network;
[0076] in response to detecting the third failure, coupling the second computing unit to the third plurality of drives via a third rack-to-rack storage link cable; and
[0077] after the third failure, providing access to the third plurality of drives via the second computing unit.
[0078] Example 12. The method of example 9, further comprising:
[0079] serially coupling two or more additional data storage racks by respective rack-to-rack storage link cables such that for each rack, the computing unit of each rack is capable of providing access to a next plurality of drives of a next rack in response to a failure that prevents a next computing unit of the next rack from providing access to the next plurality of drives; and coupling a second rack-to-rack storage link cable from a last computing unit of a last storage rack of the two or more additional storage racks to the first plurality of drives such that the last computing unit is capable of providing access to the first plurality of drives in response to a second failure that prevents the first computing unit from accessing the first plurality of drives.
[0080] Example 13. The method of example 9, further comprising:
[0081] detecting a second failure of the first computing unit that prevents the first computing unit from providing access to the first plurality of drives via the system network;
[0082] in response to detecting the second failure, coupling the second computing unit to the first plurality of drives via a second rack-to-rack storage link cable; and
[0083] after the second failure, providing access to the first plurality of drives via the second computing unit.
[0084] Example 14. The method of example 9, wherein the first failure is detected by the first computing unit and the first computing unit provides access to the second plurality of drives via the system network in response to the first failure.
[0085] Example 15. The method of example 9, wherein a storage middleware coupled to the first computing unit and the second computing unit via the system network detects the first failure and instructs the first computing unit to provide access to the second plurality of drives via the system network in response to the first failure.
[0086] Example 16. A system comprising:
[0087] a first data storage rack, a second data storage rack, and a third data storage rack, each of the first data storage rack, the second data storage rack, and the third data storage rack including a first computing unit, a second computing unit, and a third computing unit, respectively, the first computing unit, the second computing unit, and the third computing unit coupled to a first plurality of storage drives, a second plurality of storage drives, and a third plurality of storage drives, respectively, via a system network and providing access to the first plurality of storage drives, the second plurality of storage drives, and the third plurality of storage drives;
[0088] a first rack-to-rack storage link cable coupling the first computing unit to the second plurality of drives such that the first computing unit can provide access to the second plurality of drives via the system network in response to a first failure disabling the second computing unit; and
[0089] a second rack-to-rack storage link cable coupling the third computing unit to the first plurality of drives such that the third computing unit can provide access to the first plurality of drives via the system network in response to a second failure disabling the first computing unit.
[0090] Example 17. The system of example 16, further comprising a storage middleware coupled to the first computing unit, the second computing unit, and the third computing unit, the storage middleware operable to perform one or both of:
[0091] detecting the first failure and instructing the first computing unit to provide access to the second plurality of drives via the system network in response to the first failure; and
[0092] detecting the second failure and instructing the third computing unit to provide access to the first plurality of drives via the system network in response to the first failure.
[0093] Example 18. The system of example 16, wherein the first computing unit is operable to detect the first failure and provide access to the second plurality of drives via the system network in response to the first failure, and the third computing unit is operable to detect the second failure and provide access to the first plurality of drives via the system network in response to the second failure.
[0094] Example 19. The system of example 16, wherein the first rack-to-rack storage link cable couples the first computing unit to a first portion of the second plurality of drives such that the first computing unit is able to provide access to the first portion of the second plurality of drives in response to the first failure, the system further comprising:
[0095] a fourth data storage rack having a fourth computing unit; and a third rack-to-rack storage link cable that couples the fourth computing unit to a second portion of the second plurality of drives such that the fourth computing unit is able to provide access to the second portion of the second plurality of drives in response to the first failure.
[0096] Example 20. The system of example 16, wherein each of the first rack-to-rack storage link cable and the second rack-to-rack storage link cable comprises one of a SAS cable, a SATA cable, a SCSI cable, a fiber channel cable, or a point-to-point Ethernet cable.
Claims
1. A system comprising: a first data storage rack comprising a first computing unit coupled to a first plurality of storage drives via a first storage controller; a second data storage rack comprising a second computing unit coupled to a second plurality of storage drives via a second storage controller; as well as A first rack-to-rack storage link cable couples the first computing unit to the second storage controller, enabling the first computing unit to provide access to the second plurality of drives via the second storage controller in response to a first failure, the first failure preventing the second computing unit from providing access to the second plurality of drives via a system network.
2. The system of claim 1, further comprising: a third data storage rack comprising a third computing unit coupled to a third plurality of storage drives via a third storage controller; as well as A second rack-to-rack storage link cable couples the third computing unit to the first storage controller, enabling the third computing unit to provide access to the first plurality of drives in response to a second failure that prevents the first computing unit from accessing the first plurality of drives.
3. The system of claim 2 , further comprising a third rack-to-rack storage link cable coupling the second computing unit to the third storage controller so that the second computing unit can provide access to the third plurality of drives in response to a third failure that prevents the third computing unit from accessing the third plurality of drives.
4. The system of claim 1 , further comprising: two or more additional data storage racks, each of the two or more additional data storage racks coupled in series by a respective rack-to-rack storage link cable such that, for each rack, a computing unit of each rack is capable of providing access to a next plurality of drives of a next rack in response to a failure that prevents a next computing unit of a next rack from providing access to the next plurality of drives; as well as a second rack-to-rack storage link cable coupling a last compute unit of a last storage rack of the two or more additional storage racks to the first storage controller, enabling the last compute unit to provide access to the first plurality of drives in response to a second failure that prevents the first compute unit from accessing the first plurality of drives.
5. The system of claim 1 , further comprising a second rack-to-rack storage link cable coupling the second computing unit to the first storage controller so that the second computing unit can provide access to the first plurality of drives in response to a second failure that prevents the first computing unit from accessing the first plurality of drives.
6. The system of claim 1, wherein: The first computing unit is operable to detect the first failure and, in response to the first failure, provide access to the second plurality of drives via the system network.
7. The system of claim 1 further comprises a storage middleware coupled to the first computing unit and the second computing unit via the system network, the storage middleware being operable to detect the first failure and, in response to the first failure, instruct the first computing unit to provide access to the second plurality of drives via the system network.
8. The system of claim 1, wherein: The first rack-to-rack storage link cable includes one of a SaS cable, a SATA cable, a SCSI cable, a Fibre Channel cable, or a point-to-point Ethernet cable.
9. A method comprising: coupling first and second computing units of a first and second data storage racks to a system network, the first and second computing units providing access to first and second pluralities of drives, respectively, in the first and second data storage racks via the system network; detecting a first failure of the second computing unit, the first failure preventing the second computing unit from providing access to the second plurality of drives via the system network; responsive to detecting the first failure, coupling the first computing unit to a second storage controller of the second plurality of drives via a first rack-to-rack storage link cable; as well as After the first failure, access to the second plurality of drives is provided via the first computing unit and the second storage controller.
10. The method of claim 9, further comprising: detecting a second failure of the first computing unit, the second failure preventing the first computing unit from providing access to the first plurality of drives via the system network; responsive to detecting the second failure, coupling a third computing unit to the first plurality of drives via a second rack-to-rack storage link cable, the third computing unit providing access to a third plurality of drives in a third data storage rack via the system network; as well as After the second failure, access to the first plurality of drives is provided via the third computing unit.
11. The method of claim 10, further comprising: detecting a third failure of the third computing unit, the third failure preventing the third computing unit from providing access to the third plurality of drives via the system network; responsive to detecting the third failure, coupling the second computing unit to the third plurality of drives via a third rack-to-rack storage link cable; as well as After the third failure, access to the third plurality of drives is provided via the second computing unit.
12. The method of claim 9, further comprising: coupling two or more additional data storage racks in series via respective rack-to-rack storage link cables such that, for each rack, a computing unit of each rack is capable of providing access to a next plurality of drives of a next rack in response to a failure that prevents a next computing unit of a next rack from providing access to a next plurality of drives; and coupling a second rack-to-rack storage link cable from a last computing unit of a last storage rack of the two or more additional storage racks to the first plurality of drives so that the last computing unit can provide access to the first plurality of drives in response to a second failure that prevents the first computing unit from accessing the first plurality of drives.
13. The method of claim 9, further comprising: detecting a second failure of the first computing unit, the second failure preventing the first computing unit from providing access to the first plurality of drives via the system network; responsive to detecting the second failure, coupling the second computing unit to the first plurality of drives via a second rack-to-rack storage link cable; as well as After the second failure, access to the first plurality of drives is provided via the second computing unit.
14. The method of claim 9, wherein: The first failure is detected by the first computing unit, and the first computing unit provides access to the second plurality of drives via the system network in response to the first failure.
15. The method of claim 9, wherein: Storage middleware coupled to the first computing unit and the second computing unit via the system network detects the first failure and, in response to the first failure, instructs the first computing unit to provide access to the second plurality of drives via the system network.
16. A system comprising: a first data storage rack, a second data storage rack, and a third data storage rack, each of the first data storage rack, the second data storage rack, and the third data storage rack comprising a first computing unit, a second computing unit, and a third computing unit, respectively, the first computing unit, the second computing unit, and the third computing unit being coupled to a first plurality of storage drives, a second plurality of storage drives, and a third plurality of storage drives, respectively, via a system network and providing access to the first plurality of storage drives, the second plurality of storage drives, and the third plurality of storage drives; a first rack-to-rack storage link cable coupling the first computing unit to a second storage controller, enabling the first computing unit to provide access to the second plurality of drives via the system network in response to a first failure disabling the second computing unit; as well as A second rack-to-rack storage link cable couples the third computing unit to the first storage controller, enabling the third computing unit to provide access to the first plurality of drives via the system network in response to a second fault disabling the first computing unit.
17. The system of claim 16, further comprising a storage middleware coupled to the first computing unit, the second computing unit, and the third computing unit, the storage middleware being operable to perform one or both of the following: detecting the first fault and, in response to the first fault, instructing the first computing unit to provide access to the second plurality of drives via the system network; and The second failure is detected, and in response to the second failure, the third computing unit is instructing to provide access to the first plurality of drives via the system network.
18. The system of claim 16, wherein: The first computing unit is operable to detect the first fault and, in response to the first fault, provide access to the second plurality of drives via the system network, and the third computing unit is operable to detect the second fault and, in response to the second fault, provide access to the first plurality of drives via the system network.
19. The system of claim 16, wherein: The first rack-to-rack storage link cable couples the first computing unit to a first portion of the second plurality of drives such that the first computing unit can provide access to the first portion of the second plurality of drives in response to the first failure, the system further comprising: a fourth data storage rack having a fourth computing unit; and A third rack-to-rack storage link cable couples the fourth computing unit to a second portion of the second plurality of drives, enabling the fourth computing unit to provide access to the second portion of the second plurality of drives in response to the first failure.
20. The system of claim 16, wherein: Each of the first rack-to-rack storage link cable and the second rack-to-rack storage link cable comprises one of a SaS cable, a SATA cable, a SCSI cable, a Fibre Channel cable, or a point-to-point Ethernet cable.
Citation Information
Patent Citations
Multi-level erasure system with cooperative optimization
US20210255925A1
Fault tolerant multiple network servers
US5696895A
Dual access pathways to serially-connected mass data storage units
US7594134B1