Synchronously replicating datasets and other managed objects to cloud-based storage systems
By employing NVRAM as a data buffer and centralizing storage management, the storage system addresses inefficiencies in conventional systems, enhancing reliability and reducing latency through direct data mapping and centralized control, ensuring efficient data replication and storage operations in cloud environments.
Patent Information
- Application Number
- JP2025134796
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2017-12-14
- Filing Date
- 2025-08-13
- Publication Date
- 2025-11-20
AI Technical Summary
Conventional storage systems face inefficiencies in managing data replication and storage operations, particularly in cloud-based environments, leading to increased latency and reduced reliability due to unnecessary write operations and complex management of storage drives.
Implementing a storage system architecture that utilizes non-volatile RAM (NVRAM) as a buffer for data writes, with a centralized storage array controller managing device operations and direct mapping of data blocks, reducing reliance on individual storage drive controllers and optimizing data management through higher-level system processes.
This approach enhances data storage reliability and efficiency by minimizing redundant writes and improving latency, while enabling seamless failover and high availability through centralized control and direct data mapping, thus optimizing storage operations in cloud-based systems.
Smart Images

Figure 2025172070000001_ABST
Abstract
Description
[Brief explanation of the drawings]
[0001] [Figure 1A]
[0001] FIG. 1 illustrates a first example system for data storage in accordance with some embodiments. [Figure 1B]
[0002] FIG. 1 illustrates a second example system for data storage according to some embodiments. [Figure 1C]
[0003] FIG. 1 illustrates a third example system for data storage according to some embodiments. [Figure 1D]
[0004] FIG. 10 illustrates a fourth example system for data storage according to some embodiments. [Figure 2A]
[0005] 1 is a perspective view of a storage cluster having multiple storage nodes and internal storage coupled to each storage node to provide network-attached storage, according to some embodiments. [Figure 2B]
[0006] FIG. 2 is a block diagram illustrating an interconnect switch coupling multiple storage nodes according to some embodiments. [Figure 2C]
[0007] FIG. 2 is a multi-level block diagram illustrating the contents of a storage node and the contents of one of the non-volatile solid-state storage units according to some embodiments. [Figure 2D]
[0008] FIG. 2 illustrates a storage server environment using some of the storage node and storage unit embodiments of the above figures, according to some embodiments. [Figure 2E]
[0009] FIG. 1 is a blade hardware block diagram illustrating the control, compute, and storage planes and authorities interacting with the underlying physical resources according to some embodiments. [Figure 2F]
[0010] FIG. 1 illustrates an elasticity software layer on a blade of a storage cluster according to some embodiments. [Figure 2G]
[0011] FIG. 1 illustrates an authority and storage resources at blades of a storage cluster, according to some embodiments. [Figure 3A]
[0012] FIG. 1 illustrates a diagram of a storage system coupled for data communication with a cloud service provider in accordance with some embodiments of the present disclosure. [Figure 3B]
[0013] FIG. 1 illustrates a diagram of a storage system according to some embodiments of the present disclosure. [Figure 4]
[0014] FIG. 1 illustrates a block diagram showing multiple storage systems supporting a pod according to some embodiments of the present disclosure. [Figure 5]
[0015] FIG. 1 illustrates a block diagram showing multiple storage systems supporting a pod according to some embodiments of the present disclosure. [Figure 6]
[0016] FIG. 1 illustrates a block diagram showing multiple storage systems supporting a pod according to some embodiments of the present disclosure. [Figure 7]
[0017] FIG. 2 illustrates a flowchart illustrating an example method for establishing a synchronous replication relationship between two or more storage systems in accordance with some embodiments of the present disclosure. [Figure 8]
[0018] FIG. 10 illustrates a flowchart illustrating a further example method for establishing a synchronous replication relationship between two or more storage systems in accordance with some embodiments of the present disclosure. [Figure 9]
[0019] FIG. 10 illustrates a flowchart illustrating a further example method for establishing a synchronous replication relationship between two or more storage systems in accordance with some embodiments of the present disclosure. [Figure 10]
[0020] FIG. 10 illustrates a flowchart illustrating a further example method for establishing a synchronous replication relationship between two or more storage systems in accordance with some embodiments of the present disclosure. [Figure 11]
[0021] FIG. 1 illustrates a flowchart showing an example method for servicing I / O operations directed to data sets that are synchronized across multiple storage systems, in accordance with some embodiments of the present disclosure. [Figure 12]
[0022] FIG. 10 illustrates a flowchart depicting a further example method for servicing I / O operations directed to data sets synchronized across multiple storage systems, in accordance with some embodiments of the present disclosure. [Figure 13]
[0023] FIG. 10 illustrates a flowchart depicting a further example method for servicing I / O operations directed to data sets synchronized across multiple storage systems, in accordance with some embodiments of the present disclosure. [Figure 14]
[0024] FIG. 10 illustrates a flowchart depicting a further example method for servicing I / O operations directed to data sets synchronized across multiple storage systems, in accordance with some embodiments of the present disclosure. [Figure 15]
[0025] FIG. 1 illustrates a flowchart showing an example method for mediating between storage systems that synchronously replicate data sets in accordance with some embodiments of the present disclosure. [Figure 16]
[0026] FIG. 1 illustrates a flowchart showing an example method for mediating between storage systems that synchronously replicate data sets in accordance with some embodiments of the present disclosure. [Figure 17]
[0027] FIG. 1 illustrates a flowchart showing an example method for mediating between storage systems that synchronously replicate data sets in accordance with some embodiments of the present disclosure. [Figure 18]
[0028] FIG. 2 illustrates a flowchart showing an example method for recovery for a storage system that synchronously replicates data sets, in accordance with some embodiments of the present disclosure. [Figure 19]
[0029] FIG. 2 illustrates a flowchart showing an example method for recovery for a storage system that synchronously replicates data sets, in accordance with some embodiments of the present disclosure. [Figure 20]
[0030] FIG. 2 illustrates a flowchart showing an example method for recovery for a storage system that synchronously replicates data sets, in accordance with some embodiments of the present disclosure. [Figure 21]
[0031] FIG. 2 illustrates a flowchart showing an example method for resynchronization for a storage system that synchronously replicates data sets in accordance with some embodiments of the present disclosure. [Figure 22]
[0032] FIG. 10 illustrates a flowchart showing an additional example method for resynchronization for a storage system that synchronously replicates data sets, in accordance with some embodiments of the present disclosure. [Figure 23]
[0033] FIG. 10 illustrates a flowchart showing an additional example method for resynchronization for a storage system that synchronously replicates data sets, in accordance with some embodiments of the present disclosure. [Figure 24]
[0034] FIG. 10 illustrates a flowchart showing an additional example method for resynchronization for a storage system that synchronously replicates data sets, in accordance with some embodiments of the present disclosure. [Figure 25]
[0035] FIG. 10 illustrates a flowchart showing an additional example method for resynchronization for a storage system that synchronously replicates data sets, in accordance with some embodiments of the present disclosure. [Figure 26]
[0036] FIG. 10 illustrates a flowchart showing an additional example method for resynchronization for a storage system that synchronously replicates data sets, in accordance with some embodiments of the present disclosure. [Figure 27]
[0037] FIG. 10 illustrates a flowchart showing an additional example method for resynchronization for a storage system that synchronously replicates data sets, in accordance with some embodiments of the present disclosure. [Figure 28]
[0038] FIG. 10 illustrates a flowchart showing an additional example method for resynchronization for a storage system that synchronously replicates data sets, in accordance with some embodiments of the present disclosure. [Figure 29]
[0039] FIG. 10 illustrates a flowchart showing an additional example method for resynchronization for a storage system that synchronously replicates data sets, in accordance with some embodiments of the present disclosure. [Figure 30]
[0040] FIG. 2 illustrates a flowchart showing an example method for managing connectivity to a synchronously replicated storage system in accordance with some embodiments of the present disclosure. [Figure 31]
[0041] FIG. 10 illustrates a flowchart depicting a further example method for managing connectivity to a synchronously replicated storage system in accordance with some embodiments of the present disclosure. [Figure 32]
[0042] FIG. 10 illustrates a flowchart depicting a further example method for managing connectivity to a synchronously replicated storage system in accordance with some embodiments of the present disclosure. [Figure 33]
[0043] FIG. 10 illustrates a flowchart depicting a further example method for managing connectivity to a synchronously replicated storage system in accordance with some embodiments of the present disclosure. [Figure 34]
[0044] FIG. 10 illustrates a flowchart depicting a further example method for managing connectivity to a synchronously replicated storage system in accordance with some embodiments of the present disclosure. [Figure 35]
[0045] FIG. 10 illustrates a flowchart depicting a further example method for managing connectivity to a synchronously replicated storage system in accordance with some embodiments of the present disclosure. [Figure 36]
[0046] FIG. 10 illustrates a flowchart depicting a further example method for managing connectivity to a synchronously replicated storage system in accordance with some embodiments of the present disclosure. [Figure 37]
[0047] FIG. 1 illustrates a flowchart illustrating an example method for automated storage system configuration for an intermediary service according to some embodiments of the present disclosure. [Figure 38]
[0048] FIG. 1 illustrates a flowchart illustrating an example method for automated storage system configuration for an intermediary service according to some embodiments of the present disclosure. [Figure 39]
[0049] FIG. 1 illustrates a flowchart illustrating an example method for automated storage system configuration for an intermediary service according to some embodiments of the present disclosure. [Figure 40]
[0050] FIG. 1 illustrates a flowchart illustrating an example method for automated storage system configuration for an intermediary service according to some embodiments of the present disclosure. [Figure 41]
[0051] FIG. 1 illustrates a diagram of a metadata representation that may be implemented as a structured collection of metadata objects that together may represent a logical volume or portion of a logical volume of storage data, in accordance with some embodiments of the present disclosure. [Figure 42A]
[0052] FIG. 2 illustrates a flowchart showing an example method for synchronizing metadata between storage systems that synchronously replicate datasets in accordance with some embodiments of the present disclosure. [Figure 42B]
[0053] FIG. 2 illustrates a flowchart showing an example method for synchronizing metadata between storage systems that synchronously replicate datasets in accordance with some embodiments of the present disclosure. [Figure 43]
[0054] FIG. 1 illustrates a flowchart showing an example method for determining active membership among storage systems that synchronously replicate data sets in accordance with some embodiments of the present disclosure. [Figure 44]
[0055] FIG. 1 illustrates a flowchart showing an example method for determining active membership among storage systems that synchronously replicate data sets in accordance with some embodiments of the present disclosure. [Figure 45]
[0056] FIG. 1 illustrates a flowchart showing an example method for determining active membership among storage systems that synchronously replicate data sets in accordance with some embodiments of the present disclosure. [Figure 46]
[0057] FIG. 1 illustrates a flowchart showing an example method for determining active membership among storage systems that synchronously replicate data sets in accordance with some embodiments of the present disclosure. [Figure 47]
[0058] FIG. 1 illustrates a flowchart showing an example method for determining active membership among storage systems that synchronously replicate data sets in accordance with some embodiments of the present disclosure. [Figure 48]
[0059] FIG. 2 illustrates a flowchart showing an example method for synchronizing metadata between storage systems that synchronously replicate datasets in accordance with some embodiments of the present disclosure. DETAILED DESCRIPTION OF THE INVENTION
[0002]
[0060] Example methods, apparatus, and articles of manufacture for synchronously replicating data sets and other managed objects to a cloud-based storage system according to embodiments of the present disclosure are described with reference to the accompanying drawings, beginning with FIG. 1A. FIG. 1A illustrates an example system for data storage according to some embodiments. System 100 (also referred to herein as a "storage system") includes a number of elements for purposes of explanation rather than limitation. It should be noted that system 100 may include the same elements, more elements, or fewer elements, configured in the same or different ways in other embodiments.
[0003]
[0061] System 100 includes several computing devices 164A-164B, which may be implemented, for example, as servers, workstations, personal computers, notebooks, etc. (also referred to herein as "client devices") in a data center. Computing devices 164A-164B may be coupled for data communication to one or more storage arrays 102A-102B through a storage area network ("SAN") 158 or a local area network ("LAN") 160.
[0004]
[0062] SAN 158 may be implemented with a variety of data communication fabrics, devices, and protocols. For example, fabrics for SAN 158 may include Fibre Channel, Ethernet, InfiniBand, Serial Attached Small Computer System Interface ("SAS"), etc. Data communication protocols for use with SAN 158 may include Advanced Technology Attachment ("ATA"), Fibre Channel Protocol, Small Computer System Interface ("SCSI"), Internet Small Computer System Interface ("iSCSI"), HyperSCSI, Non-Volatile Memory Express ("NVMe") over Fabrics, etc. Note that SAN 158 is provided for purposes of illustration and not limitation. Other data communication couplings may be implemented between computing devices 164A-164B and storage arrays 102A-102B.
[0005]
[0063] Additionally, LAN 160 may be implemented with a variety of fabrics, devices, and protocols. For example, fabrics for LAN 160 may include Ethernet (802.3), wireless (802.11), etc. Data communication protocols for use in LAN 160 may include Transmission Control Protocol ("TCP"), User Datagram Protocol ("UDP"), Internet Protocol ("IP"), Hypertext Transfer Protocol ("HTTP"), Wireless Access Protocol ("WAP"), Handheld Device Transfer Protocol ("HDTP"), Session Initiation Protocol ("SIP"), Real Time Protocol ("RTP"), etc.
[0006]
[0064] Storage arrays 102A-102B may provide persistent data storage for computing devices 164A-164B. In embodiments, storage array 102A may be included in a chassis (not shown) and storage array 102B may be included in another chassis (not shown). Storage arrays 102A and 102B may include one or more storage array controllers 110 (also referred to herein as "controllers"). Storage array controller 110 may be implemented as an automated computing module including computer hardware, computer software, or a combination of computer hardware and software. In some embodiments, storage array controller 110 may be configured to perform a variety of storage tasks. Storage tasks may include writing data received from computing devices 164A-164B to storage arrays 102A-102B, erasing data from storage arrays 102A-102B, retrieving data from storage arrays 102A-102B and providing the data to computing devices 164A-164B, monitoring and reporting disk utilization and performance, performing redundancy operations such as RAID or RAID-like data redundancy operations, compressing data, encrypting data, etc.
[0007]
[0065] Storage array controller 110 may be implemented in a variety of ways, including as a field programmable gate array ("FPGA"), a programmable logic chip ("PLC"), an application specific integrated circuit ("ASIC"), a system on a chip ("SOC"), or any computing device including discrete components such as a processing unit, a central processing unit, computer memory, or various adapters. Storage array controller 110 may include a data communications adapter configured to support communications over, for example, SAN 158 or LAN 160. In some embodiments, storage array controller 110 may be independently coupled to LAN 160. In embodiments, storage array controller 110 may include an I / O controller or the like that couples storage array controller 110 for data communications to persistent storage resources 170A-170B (also referred to herein as "storage resources") through a midplane (not shown). Persistent storage resources 170A-170B may include any number of storage drives 171A-171F (also referred to herein as "storage devices") and any number of non-volatile random access memory (NVRAM) devices (not shown).
[0008]
[0066] In some embodiments, the NVRAM devices of persistent storage resources 170A-170B may be configured to receive data to be stored on storage drives 171A-171F from storage array controller 110. In some examples, the data may originate from computing devices 164A-164B. In some examples, writing data to the NVRAM devices may be performed more quickly than writing data directly to storage drives 171A-171F. In embodiments, storage array controller 110 may be configured to utilize the NVRAM devices as a quickly accessible buffer for data known to be written to storage drives 171A-171F. The latency of write requests using the NVRAM devices as a buffer may be improved compared to systems in which storage array controller 110 writes data directly to storage drives 171A-171F. In some embodiments, the NVRAM devices may be implemented with computer memory in the form of high-bandwidth, low-latency RAM. NVRAM devices are referred to as "non-volatile" because they may receive or include an intrinsic power source that maintains the state of the RAM after a loss of main power to the NVRAM device. Such a power source may be a battery, one or more capacitors, etc. In response to a power loss, the NVRAM device may be configured to write the contents of the RAM to persistent storage, such as storage drives 171A-171F.
[0009]
[0067] In embodiments, storage drives 171A-171F may refer to any device configured to persistently record data, where "persistently" or "persistent" refers to the device's ability to maintain recorded data after a loss of power. In some embodiments, storage drives 171A-171F may correspond to non-disk storage media. For example, storage drives 171A-171F may be one or more solid-state drives ("SSDs"), flash memory-based storage, any type of solid-state non-volatile memory, or any other type of non-mechanical storage device. In other embodiments, storage drives 171A-171F may include mechanical or rotating hard disks, such as hard disk drives ("HDDs").
[0010]
[0068] In some embodiments, the storage array controller 110 may be configured to remove device management responsibilities from the storage drives 171A-171F in the storage arrays 102A-102B. For example, the storage array controller 110 may manage control information that may describe the state of one or more memory blocks of the storage drives 171A-171F. The control information may indicate, for example, that a particular memory block has failed and should no longer be written to, that a particular memory block contains boot code for the storage array controller 110, the number of program-erase ("P / E") cycles that have been performed on a particular memory block, the age of the data stored in a particular memory block, the type of data stored in a particular memory block, etc. In some embodiments, the control information may be stored with the associated memory block as metadata. In other embodiments, the control information for the storage drives 171A-171F may be stored in one or more specific memory blocks of the storage drives 171A-171F selected by the storage array controller 110. The selected memory block may be tagged with an identifier indicating that the selected memory block contains control information. The identifiers may be utilized by storage array controller 110 in conjunction with storage drives 171A-171F to quickly identify memory blocks containing the control information. For example, storage controller 110 may issue a command to locate memory blocks containing the control information. Note that because the control information is very large, portions of the control information may be stored in multiple locations; the control information may be stored in multiple locations, for example, for redundancy, or the control information may otherwise be distributed across multiple memory blocks of storage drives 171A-171F.
[0011]
[0069] In an embodiment, storage array controller 110 may remove device management responsibility from storage drives 171A-171F of storage arrays 102A-102B by retrieving, from storage drives 171A-171F, control information that describes the state of one or more memory blocks of storage drives 171A-171F. Retrieving the control information from storage drives 171A-171F may be performed, for example, by storage array controller 110 querying storage drives 171A-171F about the location of the control information for a particular storage drive 171A-171F. Storage drives 171A-171F may be configured to execute instructions that enable storage drives 171A-171F to identify the location of the control information. The instructions may be executed by a controller (not shown) associated with or otherwise located at storage drives 171A-171F and may cause storage drives 171A-171F to scan a portion of each memory block to identify memory blocks that store control information for storage drives 171A-171F. Storage drives 171A-171F may respond by sending a response message to storage array controller 110 that includes the location of the control information for storage drives 171A-171F. In response to receiving the response message, storage array controller 110 may issue a request to read the data stored at the address associated with the location of the control information for storage drives 171A-171F.
[0012]
[0070] In other embodiments, storage array controller 110 may further remove device management responsibilities from storage drives 171A-171F by performing storage drive management operations in response to receiving the control information. Storage drive management operations may include, for example, operations normally performed by storage drives 171A-171F (e.g., a controller (not shown) associated with a particular storage drive 171A-171F). Storage drive management operations may include, for example, ensuring that data is not written to failed memory blocks within storage drives 171A-171F, ensuring that data is written to memory blocks within storage drives 171A-171F so that an appropriate wear rate is achieved, etc.
[0013]
[0071] In embodiments, storage arrays 102A-102B may implement two or more storage array controllers 110. For example, storage array 102A may include storage array controller 110A and storage array controller 110B. At a given instance, a single storage array controller 110 (e.g., storage array controller 110A) of storage system 100 may be designated with primary status (also referred to herein as the “primary controller”), and another storage array controller 110 (e.g., storage array controller 110A) may be designated with secondary status (also referred to herein as the “secondary controller”). A primary controller may have certain rights, such as permission to modify data on persistent storage resources 170A-170B (e.g., writing data to persistent storage resources 170A-170B). At least some of the rights of the primary controller may supersede the rights of a secondary controller. For example, a secondary controller may not have permission to modify data on persistent storage resources 170A-170B when the primary controller has the right. The status of storage array controller 110 may change. For example, storage array controller 110A may be designated with secondary status, and storage array controller 110B may be designated with primary status.
[0014]
[0072] In some embodiments, a primary controller, such as storage array controller 110A, may serve as the primary controller for one or more storage arrays 102A-102B, and a second controller, such as storage array controller 110B, may serve as a secondary controller for one or more storage arrays 102A-102B. For example, storage array controller 110A may be the primary controller for storage array 102A and storage array 102B, and storage array controller 110B may be the secondary controller for storage arrays 102A and 102B. In some embodiments, storage controllers 110C and 110D (also referred to as “storage processing modules”) may not have primary or secondary status. Storage array controllers 110C and 110D, implemented as storage processing modules, may serve as a communication interface between the primary and secondary controllers (e.g., storage array controllers 110A and 110B, respectively) and storage array 102B. For example, storage array controller 110A of storage array 102A may send a write request to storage array 102B via SAN 158. The write request may be received by both storage array controllers 110C and 110D of storage array 102B. Storage array controllers 110C and 110D facilitate the communication, e.g., sending the write request to the appropriate storage drives 171A-171F. Note that in some embodiments, storage processing modules may be used to increase the number of storage drives controlled by the primary and secondary controllers.
[0015]
[0073] In an embodiment, storage array controller 110 is communicatively coupled to one or more storage drives 171A-171F and to one or more NVRAM devices (not shown) included as part of storage arrays 102A-102B via a midplane (not shown). Storage array controller 110 may be coupled to the midplane via one or more data communication links, and the midplane may be coupled to storage drives 171A-171F and the NVRAM devices via one or more data communication links. The data communication links described herein are collectively represented by data communication links 108A-108D and may include, for example, Peripheral Component Interconnect Express ("PCIe") buses.
[0016]
[0074] FIG. 1B illustrates an example system for data storage according to some embodiments. The storage array controller 101 illustrated in FIG. 1B may be similar to the storage array controller 110 described with respect to FIG. 1A. In one example, the storage array controller 101 may be similar to the storage array controller 110A or the storage array controller 110B. The storage array controller 101 includes multiple elements for purposes of explanation rather than limitation. Note that the storage array controller 101 may include the same elements, more elements, or fewer elements configured similarly or differently in other embodiments. Note that elements of FIG. 1A may be included below to help illustrate features of the storage array controller 101.
[0017]
[0075] Storage array controller 101 may include one or more processing units 104 and random access memory ("RAM") 111. Processing unit 104 (or controller 101) represents one or more general-purpose processing units, such as, for example, a microprocessor, a central processing unit, or the like. More specifically, processing unit 104 (or controller 101) may be a multiple instruction set computing ("CISC") microprocessor, a reduced instruction set computing ("RISC") microprocessor, a very long instruction word ("VLIW") microprocessor, or a processor implementing other instruction sets or a processor implementing a combination of instruction sets. Processing unit 104 (or controller 101) may also be one or more special-purpose processing units, such as, for example, an application specific integrated circuit ("ASIC"), a field programmable gate array ("FPGA"), a digital signal processor ("DSP"), a network processor, or the like.
[0018]
[0076] The processing unit 104 may be connected to RAM 111 via a data communication link 106, which may be implemented as a high-speed memory bus such as a double data rate 4 ("DDR4") bus. Stored in RAM 111 is an operating system 112. In some implementations, instructions 113 are stored in RAM 111. The instructions 113 may include computer program instructions for performing operations on a direct-mapped flash storage system. In one embodiment, a direct-mapped flash storage system is a flash storage system that directly addresses data blocks within a flash drive, without address translation, performed by a storage controller of the flash drive.
[0019]
[0077] In an embodiment, storage array controller 101 includes one or more host bus adapters 103A-103C coupled to processing unit 104 via data communication links 105A-105C. In an embodiment, host bus adapters 103A-103C may be computer hardware that connects a host system (e.g., a storage array controller) to other networks and storage arrays. In some examples, host bus adapters 103A-103C may be Fibre Channel adapters that allow storage array controller 101 to connect to a SAN, Ethernet adapters that allow storage array controller 101 to connect to a LAN, etc. Host bus adapters 103A-103C may be coupled to processing unit 104 via data communication links 105A-105C, such as a PCIe bus.
[0020]
[0078] In an embodiment, storage array controller 101 may include a host bus adapter 114 coupled to an expander 115. Expander 115 may be used to attach a host system to a larger number of storage drives. Expander 115 may be, for example, a SAS expander utilized to allow host adapter 114 to attach to storage drives in an embodiment in which host bus adapter 114 is implemented as a SAS controller.
[0021]
[0079] In an embodiment, storage array controller 101 may include a switch 116 coupled to processing unit 104 via data communication link 109. Switch 116 may be a computer hardware device that creates multiple endpoints from a single endpoint, thereby allowing multiple devices to share a single endpoint. Switch 116 may be, for example, a PCIe switch coupled to a PCIe bus (e.g., data bus link 109) and presenting multiple PCIe connection points at a midplane.
[0022]
[0080] In an embodiment, storage array controller 101 includes data communication link 107 for coupling storage array controller 101 to other storage array controllers. In some examples, data communication link 107 may be a Quick Path Interconnect (QPI) interconnect.
[0023]
[0081] A conventional storage system using conventional flash drives may implement processes throughout the flash drives that are part of the conventional storage system. For example, higher-level processes in the storage system may initiate and control processes throughout the flash drives. However, flash drives in a conventional storage system may include their own storage controller that also executes processes. Thus, in the case of a conventional storage system, both higher-level processes (e.g., initiated by the storage system) and lower-level processes (e.g., initiated by the storage controller of the storage system) may be executed.
[0024]
[0082] To address various deficiencies of conventional storage systems, operations may be performed by higher-level processes rather than by lower-level processes. For example, a flash storage system may include a flash drive that does not include a storage controller to provide the processes. Thus, the flash storage system's operating system itself may initiate and control the processes. This is accomplished by direct mapping flash storage systems that address data blocks directly within the flash drive, which may be performed by the flash drive's storage controller without address translation.
[0025]
[0083] The operating system of the flash storage system may identify and maintain a list of allocation units across multiple flash drives of the flash storage system. An allocation unit may be a whole erase block or multiple erase blocks. The operating system may maintain a map or address ranges that directly map addresses to erase blocks on the flash drives of the flash storage system.
[0026]
[0084] Direct mapping to erase blocks of a flash drive may be used to rewrite data and erase data. For example, operations may be performed on one or more allocation units containing first and second data, where the first data is retained and the second data is no longer in use by the flash storage system. The operating system may initiate a process to write the first data to a new location in another allocation unit, erase the second data, and mark the allocation unit as available for use for subsequent data. In this manner, the process may be performed solely by the higher-level operating system of the flash storage system, without additional lower-level processes being performed by the flash drive's controller.
[0027]
[0085] Advantages of processes being performed solely by the flash storage system's operating system include increased reliability of the flash drives in the flash storage system, since unnecessary or redundant write operations are not performed during the process. One possible point of novelty here is the concept of initiating and controlling the process in the flash storage system's operating system. Furthermore, the process can be controlled by the operating system across multiple flash drives. This is in contrast to processes being performed by the flash drive's storage controller.
[0028]
[0086] A storage system may consist of two storage array controllers that share a set of drives for failover purposes, or the storage system may consist of a single storage array controller that provides storage services utilizing multiple drives, or the storage system may consist of a distributed network of storage array controllers, each with some number of drives or some amount of flash storage, that cooperate to provide a complete storage service and cooperate with respect to various aspects of the storage service, including storage allocation and garbage collection.
[0029]
[0087] 1C illustrates a third example system 117 for data storage according to some embodiments. System 117 (also referred to herein as a "storage system") includes numerous elements for purposes of illustration rather than limitation. Note that system 117 may include the same elements, more elements, or fewer elements configured similarly or differently in other embodiments.
[0030]
[0088] In one embodiment, system 117 includes dual Peripheral Component Interconnect ("PCI") flash storage devices 118 with separately addressable high-speed write storage. System 117 may include storage controller 119. In one embodiment, storage controller 119 may be a CPU, ASIC, FPGA, or any other circuitry that may implement the necessary control structures in accordance with the present disclosure. In one embodiment, system 117 includes flash memory devices (e.g., including flash memory devices 120a-120n) operatively coupled to various channels of storage device controller 119. Flash memory devices 120a-120n may be presented to controller 119 as addressable collections of flash pages, erase blocks, and / or control elements sufficient to allow storage device controller 119 to program and retrieve various aspects of the flash. In one embodiment, storage device controller 119 may perform operations on flash memory devices 120A-120N, including storing and retrieving data content of pages, allocating and erasing any blocks, tracking statistics related to the use and reuse of flash memory pages, erase blocks, and cells, tracking and predicting error codes and failures in flash memory, controlling voltage levels associated with programming, and retrieving content such as flash cells.
[0031]
[0089] In one embodiment, system 117 may include RAM 121 for storing separately addressable high-speed write data. In one embodiment, RAM 121 may be one or more separate, discrete components. In another embodiment, RAM 121 may be integrated into storage device controller 119 or multiple storage device controllers. RAM 121 may be utilized for other purposes, such as temporary program memory for a processing unit (e.g., a CPU) within storage device controller 119.
[0032]
[0090] In one embodiment, system 119 may include a stored energy device 122, such as a rechargeable battery or a capacitor. Stored energy device 122 may store enough energy to power some amount of RAM (e.g., RAM 121), some amount of flash memory (e.g., flash memories 120a through 120n), and storage device controller 119 for a time sufficient to write the contents of the RAM to flash memory. In one embodiment, storage device controller 119 may write the contents of RAM to flash memory when the storage device controller detects a loss of external power.
[0033]
[0091] In one embodiment, system 117 includes two data communication links 123a, 123b. In one embodiment, data communication links 123a, 123b may be PCI interfaces. In other embodiments, data communication links 123a, 123b may be based on other communication standards (e.g., HyperTransport, InfiniBand, etc.). Data communication links 123a, 123b may be based on the Non-Volatile Memory Express ("NVMe") or NVMe over Fabric ("NVMf") specifications, which allow external connections from other components of storage system 117 to storage device controller 119. Note that for convenience, data communication links may be referred to herein synonymously as PCI buses.
[0034]
[0092] System 117 may also include an external power source (not shown), which may be provided via one or both data communication links 123a, 123b, or may be provided separately. An alternative embodiment includes a separate flash memory (not shown) dedicated for use in storing the contents of RAM 121. Storage device controller 119 may present a logical device on the PCI bus, which may include an addressable fast-write logical device, or a separate portion of the logical address space of storage device 118, which may be presented as PCI memory or as persistent storage. In one embodiment, operations to store into the device are directed into RAM 121. In the event of a power outage, storage device controller 119 may write stored content associated with the addressable fast-write logical storage to a flash memory (e.g., flash memories 120a through 120n) for long-term persistent storage.
[0035]
[0093] In one embodiment, a logical device may include some representation of some or all of the contents of flash memory devices 120a-120n that allows a storage system (e.g., storage system 117) including storage device 118 to directly address flash memory pages and reprogram erase blocks directly from storage system components external to the storage device through a PCI bus. The representation may also allow one or more of the external components to control and retrieve other aspects of flash memory, including some or all of: tracking statistics related to the use and reclamation of flash memory pages, erase blocks, and cells across all flash memory devices; tracking and predicting error codes and failures within and across flash memory devices; controlling voltage levels associated with programming; and retrieving content such as flash cells.
[0036]
[0094] In one embodiment, stored energy device 122 may be sufficient to ensure completion of ongoing operations on flash memory devices 107a-120n, and stored energy device 122 may provide power to storage device controller 119 and associated flash memory devices (e.g., 120a-120n) for those operations and for fast write RAM storage to flash memory. Stored energy device 122 may be used to store accumulated statistics and other parameters kept and tracked by flash memory devices 120a-120n and / or storage device controller 119. A separate capacitor or stored energy device (e.g., a smaller capacitor near or embedded in the flash memory device itself) may be used for some or all of the operations described herein.
[0037]
[0095] Various techniques may be used to track and optimize the life of the stored energy components, such as adjusting the voltage level over time, partially discharging the stored energy device 122 and measuring the corresponding discharge characteristics, etc. If the available energy decreases over time, the effective available capacity of the addressable fast-write storage may be reduced to ensure that it can be safely written based on the currently available stored energy.
[0038]
[0096] 1D illustrates a third example system 124 for data storage in accordance with some embodiments. In one embodiment, system 124 includes storage controllers 125a and 125b. In one embodiment, storage controllers 125a and 125b are operatively coupled to dual PCI storage devices 119a, 119b, and 119c, 119d, respectively. Storage controllers 125a and 125b may be operatively coupled to any number of host computers 127a-127n (e.g., via storage network 130).
[0039]
[0097] In one embodiment, two storage controllers (e.g., 125a and 125b) provide storage services, such as a block storage array (e.g., SCS), a file server, an object server, a database, or a data analytics service. Storage controllers 125a, 125b may provide services to host computers 127a-127n outside of storage system 124 through any number of network interfaces (e.g., 126a-126d). Storage controllers 125a, 125b may provide integrated services or applications entirely within storage system 124, forming a converged storage and computing system. Storage controllers 125a, 125b may utilize fast write memory within or across storage devices 119a-119d to record ongoing operations to ensure that operations are not lost in the event of a power outage, removal of a storage controller, an outage of a storage controller or storage system, or any failure of one or more software or hardware components within storage system 124.
[0040]
[0098] In one embodiment, controllers 125a, 125b act as PCI masters for one or the other PCI bus 128a, 128b. In another embodiment, 128a and 128b may be based on other communication standards (e.g., HyperTransport, InfiniBand, etc.). Other storage system embodiments may operate storage controllers 125a, 125b as multi-masters for PCI buses 128a, 128b. Alternatively, a PCI / NVMe / NVMf switching infrastructure or fabric may connect multiple storage controllers. Some storage system embodiments may allow storage devices to communicate directly with each other rather than communicating only with the storage controller. In one embodiment, storage device controller 119a may be operable under direction from storage controller 125a to combine data stored in flash memory devices and transfer data from data stored in RAM (e.g., RAM 121 of FIG. 1C ). For example, a recomputed version of the RAM contents may be transferred after the storage controller determines that operations are fully committed across the storage system, or when fast write memory for a device reaches a certain usage capacity, or after a certain amount of time, to improve data safety or to free up addressable fast write capacity for reuse. This mechanism may be used, for example, to avoid a second transfer on a bus (e.g., 128a, 128b) from storage controller 125a, 125b. In one embodiment, the recomputation may include compressing data, connecting indexing or other metadata, combining multiple data segments together, performing erasure code calculations, etc.
[0041]
[0099] In one embodiment, under instructions from storage controller 125a, 125b, storage device controller 119a, 119b may operate to calculate and transfer data from data stored in RAM (e.g., RAM 121 in FIG. 1C) to other storage devices without the involvement of storage controller 125a, 125b. This operation may be used to mirror data stored from one controller 125a to another controller 125b, or it may be used to offload compression, data consolidation, and / or erasure coding calculations and transfers to storage devices to reduce the load on PCI bus 128a, 128b from storage controller or storage controller interface 129a, 129b.
[0042]
[0100] The storage device controller 119 may include mechanisms for implementing high availability primitives for use by other parts of the storage system external to the dual PCI storage device 118. For example, reservation or exclusion primitives may be provided so that in a storage system having two storage controllers providing highly available storage services, one storage controller may prevent the other storage controller from accessing or continuing to access the storage device. This could be used, for example, when one controller detects that the other controller is not functioning properly, or when the interconnect between the two storage controllers may itself not be functioning properly.
[0043]
[0101] In one embodiment, a storage system for use with dual PCI direct mapped storage devices with separately addressable fast write storage includes a system that manages erase blocks or groups of erase blocks as allocation units for storing data on behalf of a storage service, for storing metadata associated with a storage service (e.g., indexes, logs, etc.), or for proper management of the storage system itself. Flash pages, which may be several kilobytes in size, may be written as data arrives or once the storage system has persisted the data for a long interval of time (e.g., beyond a defined threshold of time). To commit data more quickly or to reduce the number of writes to the flash memory device, the storage controller may first write the data to separately addressable fast write storage on one or more storage devices.
[0044]
[0102] In one embodiment, storage controllers 125a, 125b may initiate the use of erase blocks within and across storage devices (e.g., 118) according to the age and expected remaining life of the storage device, or based on other statistics. Storage controllers 125a, 125b may initiate garbage collection and data migration data between storage devices according to pages that are no longer needed, as well as to manage the lifespan of flash pages and erase blocks and to manage overall system performance.
[0045]
[0103] In one embodiment, storage system 124 may utilize mirroring and / or erasure coding schemes as part of storing data to addressable fast-write storage and / or as part of writing data to allocation units associated with erase blocks. Erasure codes may be used within erase blocks or allocation units, or across flash memory devices on a single storage device, as well as across storage devices, to provide redundancy against single or multiple storage device failures or to protect against internal corruption of flash memory pages resulting from flash memory operations or from degradation of flash memory cells. Various levels of mirroring and erasure coding may be used to recover from multiple types of failures, occurring separately or in combination.
[0046]
[0104] The embodiments illustrated with respect to FIGS. 2A through 2G depict a storage cluster that stores user data, such as user data originating from one or more user or client systems or other sources external to the storage cluster. The storage cluster uses erasure coding and redundant copies of metadata to distribute user data across storage nodes housed within a chassis or across multiple chassis. Erasure coding refers to a method of protecting or reconstructing data where the data is stored across a collection of different locations, such as disks, storage nodes, or geographic locations. Flash memory is one type of solid-state memory that may be integrated with embodiments, although embodiments may be extended to other storage media, including other types of solid-state memory or non-solid-state memory. Storage location and workload control are distributed across storage locations in a clustered peer-to-peer system. Tasks such as mediating communications between various storage nodes, detecting when a storage node becomes unavailable, and balancing I / O (input / output) across various storage nodes are all handled on a distributed basis. Data is laid out or distributed across multiple storage nodes in data fragments or stripes, which in some embodiments support data recovery. Data ownership is reassigned within the cluster regardless of input and output patterns. This architecture, described in more detail below, allows storage nodes within a cluster to fail while the system remains operational, since the data can be reconstructed from other storage nodes and therefore remains available for input and output operations. In various embodiments, the storage nodes may be referred to as cluster nodes, blades, or servers.
[0047]
[0105] A storage cluster may be contained within a chassis, i.e., an enclosure that houses one or more storage nodes. Contained within the chassis are mechanisms for providing power to each storage node, such as a power distribution bus, and communication mechanisms, such as a communication bus, that enable communication between the storage nodes. According to some embodiments, the storage cluster can run as an independent system in one location. In one embodiment, the chassis includes at least two instances of both power distribution and communication buses, which can be independently enabled or disabled. The internal communication bus may be an Ethernet bus, although other technologies, such as PCIe, InfiniBand, and others, are equally suitable. The chassis provides ports for an external communication bus to enable communication between multiple chassis and with client systems, either directly or through a switch. External communication may use technologies such as Ethernet, InfiniBand, Fibre Channel, and the like. In some embodiments, the external communication bus uses different communication bus technologies for inter-chassis communication and client communication. When a switch is deployed within or between chassis, the switch may perform the function of translating between multiple protocols or technologies. When multiple chassis are connected to define a storage cluster, the storage cluster may be accessed by clients using a proprietary interface or a standard interface, such as Network File System ("NFS"), Common Internet File System ("CIFS"), Small Computer System Interface ("SCSI"), or Hypertext Transfer Protocol ("HTTP"). Translation from the client protocol may occur at the switch, the chassis external communication bus, or within each storage node. In some embodiments, multiple chassis may be coupled or connected to each other through an aggregator switch. Some and / or all of the coupled or connected chassis may be designated as a storage cluster.As mentioned above, each chassis may have multiple blades, and each blade has a media access control (MAC) address, but the storage cluster, in some embodiments, is presented to the external network as having a single cluster IP address and a single MAC address.
[0048]
[0106] Each storage node may be one or more storage servers, each connected to one or more non-volatile solid-state memory units, sometimes referred to as storage units or storage devices. One embodiment includes a single storage server for each storage node and between one and eight non-volatile solid-state memory units, although this example is not intended to be limiting. The storage server may include a processor, DRAM, and interfaces for internal communication and power distribution, respectively. In some embodiments, inside the storage node, the interfaces and storage units share a communication bus, such as PCI Express. The non-volatile solid-state memory units may directly access the internal communication bus interface through the storage node communication bus or may require the storage node to access the bus interface. The non-volatile solid-state memory units include an embedded CPU, a solid-state storage controller, and, in some embodiments, a large amount of solid-state mass storage, e.g., between 2 and 32 terabytes ("TB"). An embedded volatile storage medium, e.g., DRAM, and an energy storage device are included in the non-volatile solid-state memory units. In some embodiments, the energy storage device is a capacitor, supercapacitor, or battery that allows for the transfer of a subset of the DRAM contents to a stable storage medium upon power loss. In some embodiments, the non-volatile solid-state memory unit is constructed with storage class memory, such as phase change memory or magnetoresistive random access memory ("MRAM"), which replaces DRAM and allows for a reduced power hold-up apparatus.
[0049]
[0107] One of the many features of the storage nodes and non-volatile solid-state storage is the ability to proactively rebuild data within a storage cluster. The storage nodes and non-volatile solid-state storage can determine when a storage node or non-volatile solid-state storage within a storage cluster is unreachable, regardless of whether there is an attempt to read the data associated with that storage node or non-volatile solid-state storage. The storage nodes and non-volatile solid-state storage then cooperate to at least partially recover and rebuild the data in a new location. This constitutes proactive rebuilding in that the system rebuilds data without waiting until the data is needed for a read access initiated by a client system utilizing the storage cluster. These and additional details of the storage memory and its operation are described below.
[0050]
[0108] FIG. 2A is a perspective view of a storage cluster 161 having multiple storage nodes 150 and internal solid-state memory coupled to each storage node to provide network-attached storage or a storage area network, according to some embodiments. A network-attached storage, storage area network, storage cluster, or other storage memory may include one or more storage clusters 161, each having one or more storage nodes 150, with a flexible and reconfigurable configuration of both physical components and the amount of storage memory provided thereby. The storage cluster 161 is designed to fit into a rack, and one or more racks may be set up and populated as desired for the storage memory. The storage cluster 161 includes a chassis 138 having multiple slots 142. It should be understood that the chassis 138 may also be referred to as a housing, enclosure, or rack unit. While other slot counts are readily devised, in one embodiment, the chassis 138 has 14 slots 142. For example, some embodiments have four slots, eight slots, 16 slots, 32 slots, or other suitable numbers of slots. In some embodiments, each slot 142 may house one storage node 150. The chassis 138 includes a flap 148 that can be utilized to mount the chassis 138 on a rack. The fans 144 provide air circulation for cooling the storage nodes 150 and their components, although other cooling components could be used, or embodiments without cooling components could be devised. The switch fabric 146 couples the storage nodes 150 in the chassis 138 to each other and to a network for communication to memory. In the embodiment shown herein, the slots 142 to the left of the switch fabric 146 and fans 144 are shown occupied by a storage node 150, while the slots 142 to the right of the switch fabric 146 and fans 144 are empty and available for insertion of a storage node 150 for illustrative purposes.This configuration is an example, and one or more storage nodes 150 could occupy slots 142 in a variety of additional configurations. The configuration of storage nodes need not be contiguous or adjacent, in some embodiments. Storage nodes 150 are hot-pluggable, meaning that storage nodes 150 can be inserted into or removed from slots 142 in chassis 138 without shutting down or powering down the system. Upon insertion or removal of storage nodes 150 from slots 142, the system recognizes the change and automatically reconfigures to adapt to the change. In some embodiments, the reconfiguration includes restoring redundancy and / or rebalancing data or load.
[0051]
[0109] Each storage node 150 may include multiple components. Although other attachments and / or components may be used in additional embodiments, in the embodiment shown, storage node 150 includes a CPU 156, i.e., a printed circuit board 159 mounted with a processor, memory 154 coupled to CPU 156, and non-volatile solid-state storage 152 coupled to CU 156. Memory 154 contains instructions executed by and / or data affected by CPU 156. As described further below, non-volatile solid-state storage 152 may include flash, or in additional embodiments, other types of solid-state memory.
[0052]
[0110] Referring to FIG. 2A , storage cluster 161 is scalable, meaning that storage capacity having non-uniform storage sizes is easily added as described above. One or more storage nodes 150 can be plugged into or removed from each chassis, and the storage cluster is self-configuring in some embodiments. Plug-in storage nodes 150, whether installed in a chassis as delivered or as a later add-on, may have different sizes. For example, in one embodiment, storage nodes 150 may have any multiple of 4TB, such as 8TB, 12TB, 16TB, 32TB, etc. In additional embodiments, storage nodes 150 could have other storage amounts or any multiple of storage capacity. The storage capacity of each storage node 150 is broadcast and influences the determination of how data is striped. For maximum storage efficiency, embodiments may self-configure as widely as possible in stripes, subject to given requirements for continued operation with the loss of at most one or at most two non-volatile solid-state storage units 152 or storage nodes 150 in the chassis.
[0053]
[0111] FIG. 2B is a block diagram illustrating communication interconnects 171A-171F and power distribution bus 172 coupling multiple storage nodes 150. Referring back to FIG. 2A, in some embodiments, communication interconnects 171A-171F may be included in or implemented in switch fabric 146. When multiple storage clusters 161 occupy a rack, in some embodiments, communication interconnects 171A-171F may be included in or implemented in an upper portion of the rack switch. As shown in FIG. 2B, storage cluster 161 is enclosed within a single chassis 138. External port 176 is coupled to storage nodes 150 through communication interconnects 171A-171F, while external port 174 is coupled directly to the storage nodes. External power port 178 is coupled to power distribution bus 172. Storage nodes 150 may include varying amounts and capacities of non-volatile solid-state storage 152, as described with respect to FIG. 2A. Additionally, one or more storage nodes may be compute-only storage nodes, as shown in FIG. 2B . Authority 168 is implemented on non-volatile solid-state storage 152, such as a list or other data structure stored in memory. In some embodiments, authority 168 is stored in non-volatile solid-state storage 152 and supported by software running on a controller or other processor of non-volatile solid-state storage 152. In additional embodiments, authority 168 is implemented on storage node 150, such as a list or other data structure stored in memory 154, and supported by software running on CPU 156 of storage node 150. Authority 168, in some embodiments, controls how and where data is stored on non-volatile solid-state storage 152. This control helps determine what type of erasure coding scheme is applied to the data and which portions of the data storage node 150 owns. Each authority 168 may be assigned to non-volatile solid-state storage 152.Each authority may control a set of inode numbers, segment numbers, or other data identifiers assigned to data by the file system, by storage node 150, or by non-volatile solid-state storage 152, in various embodiments.
[0054]
[0112] In some embodiments, every data and every metadata has redundancy within the system. Furthermore, every data and every metadata has an owner, sometimes called an authority. If the authority is unreachable, for example, due to a storage node failure, there is a succession plan for how to find the data or its metadata. In various embodiments, there are redundant copies of authority 168. Authority 168, in some embodiments, has a relationship to storage nodes 150 and non-volatile solid-state storage 152. Each authority 168 covering a range of data segment numbers or other identifiers of data may be assigned to a particular non-volatile solid-state storage 152. In some embodiments, authorities 168 for all such ranges are distributed across the non-volatile solid-state storage 152 of the storage cluster. Each storage node 150 has a network port that provides access to the non-volatile solid-state storage(s) 152 of that storage node 150. Data may be stored in segments that are associated with a segment number, which in some embodiments is an indirection for configuring a RAID (Redundant Array of Independent Disks) stripe. The assignment and use of authority 168 thus establishes indirection to the data. Indirection may, according to some embodiments, be referred to as the ability to reference data indirectly, in this case through authority 168. A segment identifies a collection of non-volatile solid-state storage 152 and a local identifier into that collection of non-volatile solid-state storage 152 that may contain the data. In some embodiments, the local identifier is an offset into the device and may be reused consecutively by multiple segments. In other embodiments, the local identifier is unique to a particular segment and is never reused.The offsets in the non-volatile solid-state storage 152 are applied to locating data for writing to or reading from the non-volatile solid-state storage 152 (in the form of a RAID stripe). The data is striped across multiple units of the non-volatile solid-state storage 152, which may include or may be different from the non-volatile solid-state storage 152 with authority 168 for the particular data segment.
[0055]
[0113] For example, during data migration or data reconstruction, if there is a change in the location of a particular segment of data, the authority 168 for that data segment should be consulted in that non-volatile solid-state storage 152 or storage node 150 that has that authority 168. To locate a particular piece of data, embodiments calculate a hash value of the data segment or apply an inode number or data segment number. The output of this operation points to the non-volatile solid-state storage 152 that has the authority 168 for that piece of data. In some embodiments, there are two stages to this operation. The first stage is mapping an entity identifier (ID), such as a segment number, inode number, or directory number, to an authority identifier. This mapping may include calculating a hash or bitmask, etc. The second stage is mapping the authority identifier to a particular non-volatile solid-state storage 152, which may be done through explicit mapping. The operation is repeatable, so that when the calculation is performed, the result of the calculation repeatably and reliably points to the particular non-volatile solid-state storage 152 that has the authority 168. The operation may include a set of reachable storage nodes as input. If the set of reachable non-volatile solid-state units changes, the optimal set changes. In some embodiments, the persistence value is the current assignment (which is always true), and the calculated value is the target assignment toward which the cluster attempts to reconfigure. This calculation may be used to determine the optimal non-volatile solid-state storage 152 for the authority given a set of non-volatile solid-state storages 152 that are reachable and that constitute the same cluster. The calculation also determines an ordered set of peer non-volatile solid-state storages 152, which also records the authority to non-volatile solid-state storage mapping so that the authority can be determined even if the assigned non-volatile solid-state storage is unreachable.If a particular authority 168 is unavailable in some embodiments, a duplicate or alternative authority 168 may be consulted.
[0056]
[0114] 2A and 2B, two of the many tasks of the CPU 156 on a storage node 150 are splitting write data and reassembling read data. When the system determines that data is to be written, the authority 168 for that data is located as described above. When the segment ID of the data has already been determined, the write request is forwarded to the non-volatile solid-state storage 152 currently determined to host the determined authority 168 from the segment. The host CPU 156 of the storage node 150 on which the non-volatile solid-state storage 152 and corresponding authority 168 reside then splits or shards the data and sends the data to the various non-volatile solid-state storages 152. The transmitted data is written as data stripes according to an erasure coding scheme. In some embodiments, data is requested to be pulled, while in other embodiments, data is pushed. Conversely, when data is read, the authority 168 for the segment ID containing the data is located as described above. The host CPU 156 of the storage node 150 on which the non-volatile solid-state storage 152 and corresponding authority 168 reside requests data from the non-volatile solid-state storage and corresponding storage node pointed to by the authority. In some embodiments, the data is read from flash storage as a data stripe. The host CPU 156 of the storage node 150 then reassembles the read data, corrects any errors (if present) according to an appropriate erasure coding scheme, and transfers the reassembled data to the network. In additional embodiments, some or all of these tasks may be handled within the non-volatile solid-state storage 152. In some embodiments, a segment host requests data to be sent to the storage node 150 by requesting a page from storage and then sending the data to the storage node that made the original request.
[0057]
[0115] In some systems, such as UNIX-style file systems, data is organized in index nodes or inodes, which specify data structures that represent objects in the file system. An object might be, for example, a file or a directory. Metadata may be associated with an object as attributes such as permission data and creation timestamps, among other attributes. A segment number might be assigned to all or a portion of such an object in the file system. In other systems, data segments are organized with segment numbers assigned elsewhere. For purposes of illustration, the unit of distribution is an entity, which might be a file, directory, or segment. That is, an entity is a unit of data or metadata stored by the storage system. Entities are grouped into collections called authorities. Each authority has an authority owner, which is a storage node that has exclusive rights to update the authority's entities. In other words, storage nodes contain authorities, which in turn contain entities.
[0058]
[0116] According to some embodiments, a segment is a logical container of data. A segment is an address space between the media address space and a physical flash location. That is, data segment numbers reside within this address space. A segment may also contain metadata that allows data redundancy to be restored (rewritten to a different flash location or device) without the involvement of high-level software. In one embodiment, the internal format of a segment includes client data and intermediate mappings to determine the location of data. Each data segment is protected against, for example, memory and other failures, if applicable, by dividing the segment into several data and parity shards. The data and parity shards are distributed, or striped, across the non-volatile solid-state storage 152 coupled to the host CPU 156 (see FIGS. 2E and 2G) according to an erasure coding scheme. The use of the term segment, in some embodiments, refers to the container and its location within the segment's address space. The use of the term stripe refers to the same collection of shards as a segment, including how the shards are distributed along with redundancy or parity information according to some embodiments.
[0059]
[0117] A series of address space translations occurs throughout the storage system. At the top are directory entries (filenames) that link to inodes. Inodes point into the media address space where data is logically stored. Media addresses may be mapped through a series of indirection to spread the load of large files or implement data services such as deduplication or snapshots. Media addresses may be mapped through a series of indirection to spread the load of large files or implement data services such as deduplication or snapshots. Segment addresses are then translated into physical flash locations. Physical flash locations, according to some embodiments, have address ranges limited by the amount of flash in the system. Media addresses and segment addresses are logical containers, and in some embodiments, use identifiers of 128 bits or more to be effectively infinite, with a likelihood of reuse calculated to be longer than the expected life of the system. Addresses from the logical containers are assigned hierarchically in some embodiments. Initially, each non-volatile solid-state storage unit 152 may be assigned a range of address space. Within this assigned range, non-volatile solid-state storage 152 can assign addresses without synchronizing with other non-volatile solid-state storage 152.
[0060]
[0118] Data and metadata are stored per storage device and in a set of basic storage layouts optimized for varying workload patterns. These layouts incorporate multiple redundancy schemes, compression formats, and indexing algorithms. Some of these layouts store information about authorities and authority masters, while others store file metadata and file data. Redundancy schemes include error correction codes to tolerate corrupted bits within a single storage device (e.g., a NAND flash chip), erasure codes to tolerate failures of multiple storage nodes, and replication schemes to tolerate data center or regional failures. In some embodiments, low-density parity check (LDPC) codes are used within a single storage unit. Reed-Solomon coding is used in storage clusters, and mirroring is used in storage grids in some embodiments. Metadata may be stored using an ordered log-structured index (e.g., a log-structured merge tree), and large data may not be stored in a log-structured layout.
[0061]
[0119] To maintain consistency across multiple copies of an entity, storage nodes implicitly agree through computation on two things: (1) the authority that contains the entity, and (2) the storage node that contains the authority. The assignment of entities to authorities may be done by pseudo-randomly assigning entities to authorities, by dividing entities into ranges based on externally generated keys, or by placing a single entity into each authority. Examples of pseudo-random schemes are linear hashing and the Replication Under Scalable Hashing ("RUSH") family of hashes, which includes Controlled Replication Under Scalable Hashing ("CRUSH"). In some embodiments, pseudo-random assignment is utilized solely to assign authorities to nodes, since the set of nodes may change. The set of authorities cannot change, and therefore any subjective function may be applied in these embodiments. Some placement schemes automatically place authorities on storage nodes. While other placement schemes rely on explicit mapping of authorities to storage nodes. In some embodiments, a pseudo-random scheme is utilized to map from each authority to a set of candidate authority owners. A pseudo-random data distribution function associated with CRUSH may assign authorities to storage nodes and create a list of locations where authorities are assigned. Each storage node may have a copy of the pseudo-random data distribution function and arrive at the same calculation for distributing and later discovering or locating authorities. Each pseudo-random scheme requires a reachable set of storage nodes as input in some embodiments to infer the same target node. Once an entity is placed into an authority, it may be stored on a physical device so that expected failures do not result in unexpected data loss. In some embodiments, a rebalancing algorithm attempts to store copies of all entities within the authority in the same layout and on the same set of machines.
[0062]
[0120] Examples of expected failures include device failure, machine theft, data center fire, and regional disasters such as nuclear or geological events. Different failures lead to different levels of acceptable data loss. In some embodiments, theft of a storage node does not affect the security or reliability of the system. On the other hand, depending on the system configuration, a regional event could lead to no data loss, loss of a few seconds or minutes of updates, or even complete data loss.
[0063]
[0121] In embodiments, data placement for storage redundancy is independent of authority placement for data consistency. In some embodiments, storage nodes that include an authority do not include any persistent storage devices. Instead, the storage nodes are connected to non-volatile solid-state storage units that do not include an authority. The communication interconnect between the storage nodes and the non-volatile solid-state storage units is comprised of multiple communication technologies and has uneven performance and fault tolerance characteristics. In some embodiments, as described above, the non-volatile solid-state storage units are connected to the storage nodes via PCI Express, and the storage nodes are connected to each other within a single chassis using an Ethernet backplane, and the chassis are connected to each other to form a storage cluster. The storage cluster, in some embodiments, is connected to clients using Ethernet or Fibre Channel. When multiple storage clusters are configured in a storage grid, the multiple storage clusters are connected using the Internet or other long-distance networking links, such as “macro-scale” links or private links that do not traverse the Internet.
[0064]
[0122] The authority owner has the exclusive right to modify entities, move entities from one non-volatile solid-state storage unit to another, and add and delete copies of entities. This allows basic data redundancy to be maintained. When an authority owner fails, is decommissioned, or becomes overloaded, the authority is transferred to a new storage node. Transient failures make it important to ensure that all non-faulty machines agree on the new authority location. Ambiguity caused by transient failures can be resolved manually by a remote system administrator or by a local hardware administrator (e.g., by physically removing the failed machine from the cluster or pressing a button on the failed machine), or automatically by a consensus protocol such as Paxos or a hot-warm failover scheme. In some embodiments, a consensus protocol is used and failover is automatic. If too many failures or replication events occur in too short a period of time, the system enters a self-preservation mode, suspending replication and data movement activity until an administrator intervenes according to some embodiments.
[0065]
[0123] As authorities are transferred between storage nodes and authority owners update their authority entities, the system transfers messages between storage nodes and non-volatile solid-state storage units. Regarding persistent messages, messages with different purposes are of different types. Depending on the message type, the system maintains different ordering and durability guarantees. As persistent messages are processed, they are temporarily stored in multiple persistent and non-persistent storage hardware technologies. In some embodiments, messages are stored in RAM, NVRAM, and NAND flash devices, and various protocols are used to efficiently use each storage medium. Latency-sensitive client requests may be persisted in replicated NVRAM and then later in NAND. Meanwhile, background rebalancing operations are persisted directly to NAND.
[0066]
[0124] Persistent messages are persistently stored before being sent. This allows the system to continue servicing client requests despite failures and component replacement. Many hardware components contain unique identifiers that are visible to system administrators, manufacturers, the hardware supply chain, and ongoing monitoring and quality control infrastructure, but applications running on top of the infrastructure addresses virtualize the addresses. These virtualized addresses do not change over the life of the storage system despite component failures and replacements. This allows components of the storage system to be replaced over time without reconfiguration or interruption to the processing of client requests. In other words, the system supports non-disruptive updates.
[0067]
[0125] In some embodiments, the virtualized addresses are stored with full redundancy. A continuous monitoring system correlates hardware and software status and hardware identifiers. This allows for the detection and prediction of failures due to defective components and manufacturing details. The monitoring system also, in some embodiments, allows for the proactive transfer of authorities and entities away from affected devices before a failure occurs by removing components from critical paths.
[0068]
[0126] 2C is a multi-level block diagram illustrating the contents of a storage node 150 and the contents of the non-volatile solid-state storage 152 of the storage node 150. Data is communicated to and from the storage node 150 by a network interface controller (“NIC”) 202, in some embodiments. Each storage node 150 includes a CPU 156 and one or more non-volatile solid-state storage devices 152, as described above. Moving down one level in FIG. 2C, each non-volatile solid-state storage device 152 includes a non-volatile random access memory (“NVRAM”) 204 and a relatively fast non-volatile solid-state memory such as flash memory 206. In some embodiments, the NVRAM 204 may be a component that does not require program / erase cycles (DRAM, MRAM, PCM), and may be memory that can support being written to much more frequently than the memory is read. 2C , NVRAM 204 is implemented in one embodiment as a high-speed volatile memory, such as dynamic random access memory (DRAM) 216, backed up by energy storage 218. Energy storage 218 provides sufficient power to keep DRAM 216 powered long enough for its contents to be transferred to flash memory 206 in the event of a power outage. In some embodiments, energy storage 218 is a capacitor, supercapacitor, electron, or other device that provides an adequate supply of energy sufficient to enable the transfer of the contents of DRAM 216 to a stable storage medium in the event of a power loss. Flash memory 206 is implemented as multiple flash dies 222, which may also be referred to as a package of flash dies 222 or an array of flash dies 222. It should be understood that flash dies 222 could be packaged in any of several ways, such as a single die per package, multiple dies per package (i.e., a multi-chip package), in a hybrid package, as bare dies on a printed circuit board or other substrate, as encapsulated dies, etc.In the illustrated embodiment, non-volatile solid-state storage 152 includes a controller 212 or other processor and an input / output (I / O) port 210 coupled to controller 212. I / O port 210 is coupled to CPU 156 and / or network interface controller 202 of flash storage node 150. Flash input / output (I / O) port 220 is coupled to flash die 222, and direct memory access unit (DMA) 214 is coupled to controller 212, DRAM 216, and flash die 222. In the illustrated embodiment, I / O port 210, controller 212, DMA unit 214, and flash I / O port 220 are implemented on a programmable logic device ("PLD") 208, such as, for example, a field programmable gate array (FPGA). In this embodiment, each flash die 222 has pages organized as 16 kB (kilobyte) pages 224 and registers 226 through which data can be written to or read from the flash die 222. In additional embodiments, other types of solid-state memory are used instead of or in addition to the flash memory shown in the flash die 222.
[0069]
[0127] In various embodiments disclosed herein, a storage cluster 161 may generally be contrasted with a storage array. Storage nodes 150 are part of a collection that creates the storage cluster 161. Each storage node 150 owns a slice of data and the computing required to provide the data. Multiple storage nodes 150 cooperate to store and retrieve data. Storage memory or storage devices typically used in storage arrays are not primarily involved in processing or manipulating data. Storage memory or storage devices in a storage array receive commands to read, write, or erase data. Storage memory or storage devices in a storage array are unaware of the larger system in which they are embedded or what the data represents. Storage memory or storage devices in a storage array may include various types of storage memory, such as RAM, solid-state drives, and hard disk drives. The storage units 152 described herein have multiple interfaces that are active simultaneously and serve multiple purposes. In some embodiments, some of the functionality of a storage node 150 is shifted into the storage unit 152, transforming the storage unit 152 into a combination of storage unit 152 and storage node 150. Placing the computation (on the stored data) within the storage units 152 places this computation closer to the data itself. Various system embodiments have a hierarchy of storage node layers with different capabilities. In contrast, in a storage array, the controller owns and knows everything about all of the data it manages within the shelves or storage devices. In a storage cluster 161, multiple controllers of multiple storage units 152 and / or storage nodes 150 cooperate in various ways (e.g., for erasure coding, data sharding, metadata communication and redundancy, storage capacity expansion or contraction, data recovery, etc.) as described herein.
[0070]
[0128] FIG. 2D illustrates a storage server environment using an embodiment of the storage node 150 and storage unit 152 of FIGS. 2A-2C. In this version, each storage unit 152 includes a processor, such as a controller 212 (see FIG. 2C), an FPGA (Field Programmable Gate Array), flash memory 206, and NVRAM 204 (which is DRAM 216 backed up by supercapacitors, see FIGS. 2B and 2C) on a PCIe (Peripheral Component Interconnect Express) board within the chassis 138 (see FIG. 2A). The storage unit 152 may be implemented as a single board containing storage and may be the largest tolerable failure domain within the chassis. In some embodiments, up to two storage units 152 can fail, and the device continues without data loss.
[0071]
[0129] Physical storage is divided into named regions based on application usage in some embodiments. NVRAM 204 is a contiguous block of reserved memory within storage unit 152's DRAM 216 and is backed by NAND flash. NVRAM 204 is logically divided into multiple memory regions, each written to as a spool (e.g., spool_region). Space within the NVRAM 204 spool is managed independently by each authority 168. Each device provides an amount of storage space to each authority 168, which in turn manages the lifetime and allocation of that space. Examples of spools include distributed transactions or concepts. When primary power to storage unit 152 fails, an onboard supercapacitor provides short-term power holdup. During this holdup interval, the contents of NVRAM 204 are flushed to flash memory 206. Upon the next power-up, the contents of NVRAM 204 are retrieved from flash memory 206.
[0072]
[0130] With respect to storage unit controllers, logical "controller" responsibilities are distributed across each of the blades containing an authority 168. This distribution of logical control is illustrated in FIG. 2D as host controller 242, mid-tier controller 244, and storage unit controller(s) 246. While the parts may be physically located on the same blade, management of the control plane and storage plane is handled independently. Each authority 168 effectively acts as an independent controller. Each authority 168 provides its own data and metadata structures, its own background workers, and maintains its own lifecycle.
[0073]
[0131] Figure 2E is a blade 252 hardware block diagram showing the control plane 254, compute plane, and storage planes 256, 258, and authorities 168 that interact with the underlying physical resources using an embodiment of the storage node 150 and storage unit 152 of Figures 2A-2C in the storage server environment of Figure 2D. The control plane 254 is partitioned into several authorities 168 that can use the computational resources in the compute plane 256 to run on any of the blades 252. The storage plane 258 is partitioned into a collection of devices, each of which provides access to flash 206 and NVRAM 204 resources.
[0074]
[0132] In the compute and storage planes 256, 258 of FIG. 2E, authorities 168 interact with underlying physical resources (i.e., devices). From the perspective of an authority 168, its resources are striped across all of the physical devices. From the perspective of a device, it provides resources to all authorities 168, regardless of where the authority happens to run. Each authority 168 has assigned or is assigned one or more partitions 260 of storage memory within a storage unit 152, such as partitions 260 of flash memory 206 and NVRAM 204. Each authority 168 uses its assigned partitions 260 to write or read user data. Authorities may be associated with different amounts of physical storage in the system. For example, one authority 168 might have more partitions 260 or larger-sized partitions 260 on one or more storage units 152 than one or more other authorities 168.
[0075]
[0133] FIG. 2F illustrates elasticity software layers for a blade 252 of a storage cluster, according to some embodiments. In an elasticity architecture, the elasticity software is symmetric. That is, each blade's compute module 270 executes the three identical layers of the process shown in FIG. 2F. A storage manager 274 executes read and write requests from other blades 252 to data and metadata stored in the local storage unit 152, NVRAM 204, and flash 206. An authority 168 fulfills client requests by issuing the necessary reads and writes to the blade 252 on whose storage unit 152 the corresponding data or metadata resides. An endpoint 272 analyzes client connection requests received from the switch fabric 146 oversight software, relays the client connection request to the authority 168 responsible for fulfillment, and relays the authority 168's response to the client. The symmetric three-tier architecture enables a high degree of concurrency in the storage system. Elasticity scales out efficiently and reliably in these embodiments. Additionally, Elasticity implements a unique scale-out approach that maximizes concurrency by balancing work evenly across all resources, regardless of client access patterns, eliminating much of the need for cross-blade coordination that typically occurs with traditional distributed locking.
[0076]
[0134] Still referring to FIG. 2F , authority 168, executing on compute module 270 of blade 252, performs the internal operations required to service client requests. One feature of elasticity is that authority 168 is stateless; that is, while the authority caches active data and metadata in its own blade 252's DRAM for fast access, the authority stores any updates to its NVRAM 204 partitions on three separate blades 252 until the updates are written to flash 206. All storage system writes to NVRAM 204 are, in some embodiments, triplicate for partitions on three separate blades 252. With triple-mirrored NVRAM 204 and persistent storage protected by parity and Reed-Solomon RAID checksums, the storage system can survive the simultaneous failure of two blades 252 without losing data, metadata, or access to either.
[0077]
[0135] Because authorities 168 are stateless, authorities can migrate between blades 252. Each authority 168 has a unique identifier. NVRAM 204 and flash 206 partitions are associated with the identifier of the authority 168, not with the blade 252 on which some authority is running. Thus, when an authority 168 migrates, the authority 168 continues to manage the same storage partitions from its new location. When a new blade 252 is installed in an embodiment of a storage cluster, the system partitions the new blade's 252's storage for use by authorities 168 in the system, migrates selected authorities 168 to the new blade 252, and automatically rebalances the load by starting endpoints 272 on the new blade 252 and including them in the client connection distribution algorithm of the switch fabric 146.
[0078]
[0136] From its new location, the migrated authority 168 persists the contents of its NVRAM 204 partition on flash 206, processes read and write requests from other authorities 168, and serves client requests directed to it by endpoints 272. Similarly, if a blade 252 fails or is removed, the system redistributes its authorities 168 among the system's remaining blades 252. The redistributed authorities 168 continue to perform their original functions from their new locations.
[0079]
[0137] FIG. 2G illustrates an authority 168 and storage resources of a blade 252 of a storage cluster, according to some embodiments. Each authority 168 is exclusively responsible for a flash 206 and NVRAM 204 partition on each blade 252. The authority 168 manages the contents and integrity of its partition independently of other authorities 168. The authority 168 compresses incoming data, temporarily stores it in its NVRAM 204 partition, and then hardens, RAID-protects, and persists the data in a segment of storage within its flash 206 partition. As the authority 168 writes data to the flash 206, the storage manager 274 performs the necessary flash transformations to optimize write performance and maximize media lifespan. In the background, the authority 168 “garbage collects,” or reclaims, space occupied by data that clients make obsolete by overwriting the data. It should be understood that because the partitions of authority 168 have no common elements, there is no need for distributed locking among clients and writes or to perform background functions.
[0080]
[0138] The embodiments described herein may utilize a variety of software protocols, communication protocols, and / or networking protocols. Furthermore, hardware and / or software configurations may be adjusted to accommodate the various protocols. For example, embodiments may utilize Active Directory, a database-based system that provides authentication, directory, policy, and other services in a Windows environment. In these embodiments, LDAP (Lightweight Directory Access Protocol) is an example application protocol for querying and modifying items in a directory service provider such as Active Directory. In some embodiments, a Network Lock Manager ("NLM") is utilized as a mechanism to provide System V-style advisory files and work in conjunction with the Network File System ("NFS") to record locks over the network. The Server Message Block ("SMB") protocol, one version of which is also known as the Common Internet File System ("CIFS"), may be integrated with the storage systems described herein. SMP operates as an application-layer network protocol typically used to provide shared access to files, printers, and serial ports, as well as various communications between nodes on a network. SMB also provides an authenticated inter-process communication mechanism. AMAZON™ S3 (Simple Storage Service) is a web service provided by Amazon Web Services, and the system described herein may interface with Amazon S3 through web service interfaces (REST (Representable State Transfer), SOAP (Simple Object Access Protocol), and BitTorrent). A RESTful API (Application Programming Interface) breaks down a transaction to create a series of small modules, each of which addresses a specific fundamental part of the transaction.The control or permissions provided in these embodiments, particularly for object data, may include the use of access control lists ("ACLs"). An ACL is a list of permissions granted to an object; an ACL specifies not only which actions are allowed on a given object, but also which user or system processes are allowed to access the object. The system provides an identification and location system for computers on the network and utilizes Internet Protocol version 6 ("IPV6"), in addition to IPv4, as the communications protocol for sending traffic across the Internet. Routing of packets between networked systems may include equal-cost multipath routing ("ECMP"). ECMP is a routing strategy in which next-hop packets forwarded to a single destination may occur on multiple "best paths" that rank highest in routing metric calculations. Because multipath routing is a hop-by-hop decision limited to a single router, multipath routing can be used with most routing protocols. The software may support multitenancy, an architecture in which a single instance of a software application serves multiple customers. Each customer may be referred to as a tenant. Tenants may be given the ability to customize some parts of the application, but in some embodiments may not customize the application's code. Embodiments may maintain an audit log. An audit log is a document that records events in a computing system. In addition to documenting which resources were accessed, audit log entries typically include destination and source addresses, a timestamp, and user login information for compliance with various regulations. Embodiments may support various key management policies, such as encryption key rotation. Additionally, the system may support a dynamic root password or some variant of a dynamically changing password.
[0081]
[0139] 3A illustrates a diagram of a storage system 306 coupled for data communication with a cloud service provider 302 in accordance with some embodiments of the present disclosure. While not shown in greater detail, the storage system 306 shown in FIG. 3A may be similar to the storage systems described above with respect to FIGS. 1A-1D and 2A-2G. In some embodiments, the storage system 306 shown in FIG. 3A may be implemented as a storage system including unbalanced active / active controllers, a storage system including balanced active / active controllers, a storage system including active / active controllers in which less than all of the resources of each controller are utilized such that each controller has spare resources that can be used to support failover, a storage system including fully active / active controllers, a storage system including data set separation controllers, a storage system including a dual-tier architecture with a front-end controller and a back-end unified storage controller, a storage system including a scale-out cluster of dual-controller arrays, and combinations of such embodiments.
[0082]
[0140] 3A , storage system 306 is coupled to cloud service provider 302 via data communications link 304. Data communications link 304 may be implemented as a dedicated data communications link, as a data communications path provided through the use of one of a data communications network such as a wide area network (“WAN”) or a local area network (“LAN”), or some other mechanism capable of transporting digital information between storage system 306 and cloud service provider 302. Such data communications link 304 may be entirely wired, entirely wireless, or some aggregation of wired and wireless data communications paths. In such an example, digital information may be exchanged between storage system 306 and cloud service provider 302 via data communications link 304 using one or more data communications protocols. For example, digital information may be exchanged between storage system 306 and cloud service provider 302 via data communications link 304 using Handheld Device Transfer Protocol ("HDTP"), Hypertext Transfer Protocol ("HTTP"), Internet Protocol ("IP"), Real-Time Transport Protocol ("RTP"), Transmission Control Protocol ("TCP"), User Datagram Protocol ("UDP"), Wireless Application Protocol ("WAP"), or other protocols.
[0083]
[0141] The cloud service provider 302 shown in FIG. 3A may be implemented as a system and computing environment that provides services to users of the cloud service provider 302 through the sharing of computing resources, for example, via a data communication link 304. The cloud service provider 302 may provide on-demand access to a shared pool of configurable computing resources, such as computer networks, servers, storage, applications, and services. The shared pool of configurable resources may be quickly set up and released to users of the cloud service provider 302 with minimal administrative effort. Generally, users of the cloud service provider 302 are unaware of the exact computing resources utilized by the cloud service provider 302 to provide services. While in many cases, such cloud service providers 302 may be accessible via the Internet, readers skilled in the art will recognize that any system that conceptualizes the use of shared resources to provide services to users over any data communication link may be considered a cloud service provider 302.
[0084]
[0142] 3A , cloud service provider 302 may be configured to provide various services to storage system 306 and users of storage system 306 through various service model implementations. For example, cloud service provider 302 may be configured to provide services to storage system 306 and users of storage system 306 through an infrastructure-as-a-service ("IaaS") service model implementation in which cloud service provider 302 provides computing infrastructure, such as virtual machines and other resources, as a service to subscribers. Furthermore, cloud service provider 302 may be configured to provide services to storage system 306 and users of storage system 306 through a platform-as-a-service ("PaaS") service model implementation in which cloud service provider 302 provides a development environment to application developers. Such a development environment may include, for example, an operating system, a programming language execution environment, a database, a web server, or other components that may be used by application developers to develop and run software solutions on a cloud platform. Additionally, cloud service provider 302 may be configured to provide services to storage system 306 and users of storage system 306 through an implementation of a software-as-a-service ("SaaS") service model in which cloud service provider 302 provides application software, databases, as well as the platform used to run applications to storage system 306 and users of storage system 306, providing on-demand software to storage system 306 and users of storage system 306, eliminating the need to install and run applications on local computers, and simplifying application maintenance and support.Cloud service provider 302 may further be configured to provide services to storage system 306 and users of storage system 306 through an implementation of an authentication as a service ("AaaS") service model, in which cloud service provider 302 provides authentication services that can be used to secure access to applications, data sources, or other resources. Cloud service provider 302 may also be configured to provide services to storage system 306 and users of storage system 306 through an implementation of a storage as a service model, in which cloud service provider 302 provides access to its storage infrastructure for use by storage system 306 and users of storage system 306. The reader will understand that the above-described service models are included for illustrative purposes only and do not represent any limitations on the services that may be provided by cloud service provider 302 or the service models that may be implemented by cloud service provider 302, and that cloud service provider 302 may be configured to provide additional services to storage system 306 and users of storage system 306 through implementation of additional service models.
[0085]
[0143] In the example shown in FIG. 3A , cloud service provider 302 may be implemented, for example, as a private cloud, a public cloud, or a combination of private and public clouds. In embodiments in which cloud service provider 302 is implemented as a private cloud, cloud service provider 302 may focus on providing services to a single organization rather than serving multiple organizations. In embodiments in which cloud service provider 302 is implemented as a public cloud, cloud service provider 302 may provide services to multiple organizations. Public and private cloud deployment models may differ and may involve various advantages and disadvantages. For example, because public cloud deployments require the sharing of computing infrastructure across different organizations, such deployments may not be ideal for organizations with security concerns, mission-critical workloads, uptime requirements, etc. While private cloud deployments can address some of these issues, private cloud deployments may require on-premise staff to manage the private cloud. In yet another alternative embodiment, cloud service provider 302 may be implemented as a hybrid cloud deployment of private and public cloud services.
[0086]
[0144] Although not explicitly shown in FIG. 3A , the reader will understand that additional hardware and software components may be necessary to facilitate delivery of cloud services to the storage system 306 and users of the storage system 306. For example, the storage system 306 may be coupled to (or include) a cloud storage gateway. Such a cloud storage gateway may be implemented, for example, as a hardware- or software-based appliance located at the facility containing the storage system 306. Such a cloud storage gateway may act as a bridge between local applications running on the storage array 306 and remote, cloud-based storage leveraged by the storage array 306. Using a cloud storage gateway, an organization can move temporary iSCSI or NAS storage to the cloud service provider 302, allowing the organization to conserve space in its on-premises storage system. Such a cloud storage gateway may be configured to emulate a disk array, block-based device, file server, or other storage system that can translate SCSI commands, file server commands, or other appropriate commands into a RESTful namespace protocol that facilitates communication with the cloud service provider 302.
[0087]
[0145] A cloud migration process may occur during which data, applications, or other elements are moved from an organization's local system (or from another cloud environment) to cloud service provider 302 to enable storage system 306 and users of storage system 306 to take advantage of services offered by cloud service provider 302. To successfully migrate the data, applications, or other elements to the cloud service provider's 302 environment, middleware such as a cloud migration tool may be utilized to bridge the gap between the cloud service provider's 302 environment and the organization's environment. Such cloud migration tools may also be configured to address security concerns associated with transmitting sensitive data over a data communications network to cloud service provider 302, as well as potentially high network costs and long migration times associated with migrating large amounts of data to cloud service provider 302. To enable storage system 306 and users of storage system 306 to further take advantage of services offered by cloud service provider 302, a cloud orchestrator may be used to prepare and coordinate automated tasks to create an integrated process or workflow. Such a cloud orchestrator may perform tasks such as configuring various components, whether those components are cloud or on-premise components, as well as managing the interconnections between such components. The cloud orchestrator can simplify communication and connections between components to ensure that links are properly configured and maintained.
[0088]
[0146] 3A , and as briefly described above, cloud service provider 302 may be configured to provide services to storage system 306 and users of storage system 306 through the use of a SaaS service model in which cloud service provider 302 provides application software, databases, as well as the platform used to run the applications to storage system 306 and users of storage system 306, providing on-demand software to storage system 306 and users of storage system 306, eliminating the need to install and run applications on local computers, which may simplify application maintenance and support. Such applications may take many forms in accordance with various embodiments of the present disclosure. For example, cloud service provider 302 may be configured to provide storage system 306 and users of storage system 306 with access to a data analysis application. Such a data analysis application may be configured, for example, to receive telemetry data communicated by storage system 306 to its servers. Such telemetry data may describe various operating characteristics of storage system 306 and may be analyzed, for example, to determine the health of storage system 306, to identify the workload running on storage system 306, to predict when storage system 306 will run out of various resources, and to recommend configuration changes, hardware or software upgrades, workflow transitions, or other actions that may improve the operation of storage system 306.
[0089]
[0147] Cloud service provider 302 may also be configured to provide storage system 306 and users of storage system 306 with access to virtualized computing environments. Such virtualized computing environments may be implemented, for example, as virtual machines or other virtualized computer hardware platforms, virtualized storage devices, virtualized computer network resources, etc. Examples of such virtualized environments may include virtual machines created to emulate actual computers, virtualized desktop environments that separate logical desktops from physical machines, virtualized file systems that enable uniform access to different types of concrete file systems, and many others.
[0090]
[0148] For further explanation, Figure 3B illustrates a diagram of a storage system 306 according to some embodiments of the present disclosure. Although not described in great detail, the storage system 306 shown in Figure 3B may be similar to the storage systems described above with reference to Figures 1A-1D and 2A-2G, as the storage system may include many of the components described above.
[0091]
[0149] The storage system 306 shown in FIG. 3B may include storage resources 308, which may be implemented in many ways. For example, in some embodiments, the storage resources 308 may include nanoRAM or another form of nonvolatile random access memory that utilizes carbon nanotubes deposited on a substrate. In some embodiments, the storage resources 308 may include 3D cross-point nonvolatile memory in which bit storage is based on bulk resistance changes, along with stackable cross-gridded data access arrays. In some embodiments, the storage resources 308 may include flash memory, including single-level cell ("SLC") NAND flash, multi-level cell ("MLC") NAND flash, triple-level cell ("TLC") NAND flash, quad-level cell ("QLC") NAND flash, and others. In some embodiments, the storage resources 308 may include nonvolatile magnetoresistive random access memory ("MRAM"), including spin-transfer torque ("STT") MRAM, in which data is stored using magnetic storage elements. In some embodiments, example storage resources 308 may include non-volatile phase change memory (“PCM”), which may have the ability to hold multiple bits in a single cell because the cell can achieve several distinct intermediate states. In some embodiments, storage resources 308 may include quantum memory, which allows for the storage and retrieval of optical quantum information. In some embodiments, example storage resources 308 may include resistive random access memory (“ReRAM”), in which data is stored by changing the resistance across a dielectric solid-state material. In some embodiments, storage resources 308 may include storage class memory (“SCM”), in which solid-state non-volatile memory may be fabricated at high density using some combination of sublithographic patterning techniques, multiple bits per cell, multiple layers of devices, etc.The reader will understand that other forms of computer memory and storage devices may be utilized by the storage system described above, including DRAM, SRAM, EEPROM, universal memory, and many others. The storage resources 308 shown in Figure 3A may be implemented in a variety of form factors, including, but not limited to, dual in-line memory modules ("DIMMs"), non-volatile dual in-line memory modules ("NVDIMMs"), M.2, U.2, and others.
[0092]
[0150] The example storage system 306 shown in FIG. 3B may implement a variety of storage architectures. For example, storage systems according to some embodiments of the present disclosure may utilize block storage, where data is stored in blocks, with each block essentially performing the function of an individual hard drive. Storage systems according to some embodiments of the present disclosure may utilize object storage, where data is managed as objects. Each object may include the data itself, a variable amount of metadata, and a globally unique identifier, and object storage can be implemented at multiple levels (e.g., device level, system level, interface level). Storage systems according to some embodiments of the present disclosure utilize file storage, where data is stored in a hierarchical structure. Such data is stored in files and folders and presented in the same format to both the system that stores it and the system that retrieves it.
[0093]
[0151] 3B may be implemented as a storage system in which additional storage resources can be added using a scale-up model, additional storage resources can be added using a scale-out model, or some combination thereof. In a scale-up model, additional storage may be added by adding additional storage devices. However, in a scale-out model, additional storage nodes may be added to a cluster of storage nodes, and such storage nodes may include additional processing resources, additional networking resources, etc.
[0094]
[0152] The storage system 306 shown in FIG. 3B also includes communication resources 310 that may be useful in facilitating data communication between components within the storage system 306 as well as data communication between the storage system 306 and computing devices external to the storage system 306. The communication resources 310 may be configured to utilize a variety of different protocols and data communication fabrics to facilitate data communication between components within the storage system as well as computing devices external to the storage system. For example, the communication resources 310 may include Fibre Channel ("FC") technology, such as an FC fabric and FC protocol, which may transport SCSI commands over an FC network. The communication resources 310 may also include FC over Ethernet ("FCoE") technology, in which FC frames are encapsulated and transmitted over an Ethernet network. The communication resources 310 may also include InfiniBand ("IB") technology, in which switched fabric technology is utilized to facilitate transmission between channel adapters. Communication resources 310 may also include NVM Express ("NVMe") technology and NVMe over Fabric ("NVMeoF") technology through which non-volatile storage media attached via a PCI Express ("PCIe") bus may be accessed.Communications resources 310 may also include mechanisms for accessing storage resources 308 within storage system 306 using Serial Attached SCSI ("SAS"), Serial ATA ("SATA") bus interfaces for connecting storage resources 308 within storage system 306 to host bus adapters within storage system 306, Internet Small Computer System Interface ("iScSI") technology to provide block-level access to storage resources 308 within storage system 306 and other communications resources that may be useful for facilitating data communication between components within storage system 306, as well as data communication between storage system 306 and computing devices external to storage system 306.
[0095]
[0153] 3B also includes processing resources 312 that may be useful in executing computer program instructions and performing other computational tasks within storage system 306. Processing resources 312 may include one or more central processing units (“CPUs”) as well as one or more application-specific integrated circuits (“ASICs”) customized for some particular purpose. Processing resources 312 may also include one or more digital signal processors (“DSPs”), one or more field programmable gate arrays (“FPGAs”), one or more systems-on-chips (“SoCs”), or other forms of processing resources 312. Storage system 306 may utilize storage resources 312 to perform various tasks, including, but not limited to, supporting the execution of software resources 314, which are described in more detail below.
[0096]
[0154] 3B also includes software resources 314 that, when executed by processing resources 312 within storage system 306, may perform various tasks. Software resources 314 may include, for example, one or more modules of computer program instructions that, when executed by processing resources 312 within storage system 306, are useful in implementing various data protection techniques to preserve the integrity of data stored within the storage system. The reader will understand that such data protection techniques may be implemented, for example, by system software running on computer hardware within the storage system, by a cloud service provider, or in other ways. Such data protection techniques may include, for example, data archiving techniques, in which data that is no longer in active use is moved to a separate storage device or separate storage system for long-term retention, data backup techniques, in which data stored in a storage system may be copied and stored elsewhere to avoid data loss in the event of equipment failure or some other form of catastrophic event with the storage system, data replication techniques, in which data stored in a storage system is replicated to another storage system so that the data is accessible via multiple storage systems, data snapshot techniques, in which the state of data in a storage system is captured at various points in time, data and database cloning techniques, through which copies of replicas of data and databases may be created, and other data protection techniques. Through the use of such data protection techniques, business continuity and disaster recovery objectives may be met, as failure of a storage system may not result in loss of data stored in the storage system.
[0097]
[0155] Software resources 314 may also include software useful in implementing software-defined storage ("SDS"). In such an example, software resources 314 may include one or more modules of computer program instructions that, when executed, are useful in policy-based provisioning and management of data storage independent of the underlying hardware. Such software resources 314 may be useful in implementing storage visualization to separate storage hardware from the software that manages the storage hardware.
[0098]
[0156] Software resources 314 may also include software useful in facilitating and optimizing I / O operations directed to storage resources 308 within storage system 306. For example, software resources 314 may include software modules that perform various data reduction techniques, such as data compression, data deduplication, and others. Software resources 314 may also include software modules that intelligently group I / O operations together to facilitate better use of underlying storage resources 308, software modules that perform data migration operations to migrate data out of the storage system, and software modules that perform other operations. Such software resources 314 may be implemented as one or more software containers or in many other ways.
[0099]
[0157] The reader will understand that the various components shown in FIG. 3B may be grouped into one or more optimized computing packages as a centralized infrastructure. Such a centralized infrastructure may include a pool of computing, storage, and networking resources that may be shared by multiple applications and collectively managed using policy-driven processes. Such a centralized infrastructure may minimize compatibility issues between the various components in the storage system 306 while also reducing various costs associated with establishing and operating the storage system 306. Such a centralized infrastructure may be implemented in a centralized infrastructure reference architecture, in standalone equipment, in a software-driven hyper-centralized approach (e.g., a hyper-centralized infrastructure), or in other ways.
[0100]
[0158] The reader will appreciate that the storage system 306 shown in Figure 3B may be useful for supporting various types of software applications. For example, the storage system 306 may be useful in supporting artificial intelligence ("AI") applications, database applications, DevOps projects, electronic design automation tools, event-driven software applications, high-performance computing applications, simulation applications, high-speed data capture and analysis applications, machine learning applications, media generation applications, media provisioning applications, picture archiving and communication systems ("PACS") applications, software development applications, virtual reality applications, augmented reality applications, and many other types of applications by providing storage resources for such applications.
[0101]
[0159] The storage system described above may operate to support a wide variety of applications. Given the fact that the storage system includes computational resources, storage resources, and a wide variety of other resources, the storage system may be well-suited to support resource-intensive applications, such as AI applications. Such AI applications may enable a device to recognize its environment and take actions that maximize its probability of success toward some goal. Examples of such AI applications may include IBM Watson, Microsoft Oxford, Google DeepMind, Baidu Minwa, and others. The storage system described above may also be well-suited to support other types of resource-intensive applications, such as machine learning applications. Machine learning applications may perform various types of data analysis to automate analytical model building. Machine learning applications use algorithms that iteratively learn from data, allowing computers to learn without being explicitly programmed.
[0102]
[0160] In addition to the resources already described, the storage systems described above may also include a graphics processing unit (“GPU”), sometimes referred to as a visual processing unit (“VPU”). Such a GPU may be implemented as specialized electronic circuitry that rapidly manipulates and modifies memory to accelerate the creation of images in a frame buffer intended for output to a display device. Such a GPU may be included within any of the computing devices that are part of the storage systems described above, including as one of many individually scalable components of the storage system; other examples of individually scalable components of such storage systems may include storage components, memory components, computational components (e.g., CPUs, FPGAs, ASICs), networking components, software components, and others. In addition to the GPU, the storage systems described above may also include a neural network processor (“NNP”) for use in various aspects of neural network processing. Such NNPs may be used instead of (or in addition to) a GPU and may be individually scalable.
[0103]
[0161] As mentioned above, the storage systems described herein may be configured to support artificial intelligence applications, machine learning applications, big data analytics applications, and many other types of applications. The rapid growth of these types of applications is driven by three technologies: deep learning (DL), GPU processors, and big data. Deep learning is a computing model that utilizes massively parallel neural networks inspired by the human brain. Instead of experts handcrafting software, deep learning models write their own software by learning from many examples. GPUs are modern processors with thousands of cores that are well-suited to run algorithms that loosely represent the parallel nature of the human brain.
[0104]
[0162] Advances in deep neural networks have ignited a new wave of algorithms and tools for data scientists to utilize that data with artificial intelligence (AI). With improved algorithms, larger datasets, and diverse frameworks (including open-source software libraries for machine learning across a range of tasks), data scientists are tackling new use cases such as autonomous vehicles, natural language processing, and many others. However, training deep neural networks requires both high-quality input data and large amounts of computation. GPUs are massively parallel processors that can operate on large amounts of data simultaneously. When coupled with multi-GPU clusters, high-throughput pipelines may be required to send input data from storage to the computation engine. Deep learning is more than simply building and training a model. There is also an entire data pipeline that must be designed for the scale, iteration, and experimentation required for data science teams to succeed.
[0105]
[0163] Data is at the heart of modern AI and deep learning algorithms. Before training can begin, one problem that must be addressed revolves around collecting labeled data, which is critical for training accurate AI models. Full-scale AI deployments may be required to continuously collect, organize, transform, label, and store large amounts of data. Adding additional high-quality data points directly translates into more accurate models and better insights. Data samples may undergo a series of processing steps, including, but not limited to, 1) ingesting data from external sources into the training system and storing the data in a raw format; 2) organizing the data and converting it into a format convenient for training, including linking data samples to appropriate labels; 3) iteratively exploring parameters and models, rapidly testing them on smaller datasets, and focusing on the most promising models for pushing into the production cluster; 4) running a training phase to select random batches of input data containing both new and older samples and feeding the random batches to a production CPU server for calculations to update the model parameters; and 5) evaluating, including using a holdback portion of the data not used in training to evaluate model accuracy on holdout data. This lifecycle may apply for any type of parallelized machine learning, not just neural networks or deep learning. For example, a standard machine learning framework may rely on a CPU instead of a GPU, but the data ingestion and training workflow may be the same. Readers will understand that a single shared storage data hub creates a coordination point throughout the lifecycle without requiring extra data copies during the ingestion, preprocessing, and training stages. Ingested data is rarely used for just one purpose, and shared storage provides the flexibility to train multiple different models or apply traditional analytics to the data.
[0106]
[0164] The reader understands that each stage of an AI data pipeline may have varying requirements from a data hub (e.g., a storage system or collection of storage systems). A scale-out storage system must deliver uncompromising performance for all types of access types and patterns—from small metadata to heavy, large files, random access patterns to sequential access patterns, and low concurrency to high concurrency. Because the system may serve unstructured workloads, the storage system described above may serve as an ideal AI data hub. In the first stage, data is ideally ingested and stored on the same data hub used by subsequent stages to avoid excessive data copying. The next two steps can be run on standard compute servers, optionally including GPUs, and then in the fourth and final stage, full training production jobs are run on powerful CPU-accelerated servers. Often, there is a production pipeline alongside an experimental pipeline running on the same dataset. Furthermore, GPU-accelerated servers can be used independently for different models or can be combined together to train on one larger model, even spanning multiple systems for distributed training. If the shared storage layer is slow, then the data must be copied to local storage for each stage, resulting in wasted time staging the data on different servers. An ideal data hub for an AI training pipeline would perform similarly to data stored locally on the server nodes, while also possessing the simplicity and performance to allow all pipeline stages to operate simultaneously.
[0107]
[0165] Data scientists work to improve the usability of trained models through a variety of methods: more data, better data, smarter training, and deeper models. Often, there are teams of data scientists sharing the same dataset and working in parallel to create new and improved training models. Often, there are teams of data scientists working through these stages simultaneously on the same shared dataset. Multiple concurrent workloads of data processing, experimentation, and full-scale training layer the demands of multiple access patterns onto the storage layer. In other words, rather than being able to satisfy large file reads, storage must handle a mixture of large and small file reads and writes. Finally, with multiple data scientists exploring datasets and models, storing data in its native format can be crucial to provide each user with the flexibility to transform, organize, and use the data in their own way. The storage systems described above can provide a natural shared storage home for datasets, data protection redundancy (e.g., by using RAID 6), and the performance necessary to be a common access point for multiple developers and multiple experiments. Using the storage system described above can avoid the need to carefully copy subsets of data for local work, saving both engineering and CPU-accelerated server time. These copies become an ever-growing burden as raw data sets persist and desired transformations constantly update and change.
[0108]
[0166] Readers will understand that the underlying reason for the rapid growth in deep learning success is the continuous improvement of models with larger dataset sizes. In contrast, older machine learning algorithms, such as logistic regression, stop improving accuracy at smaller dataset sizes. In this way, the separation of compute and storage resources also allows for independent scaling of each tier, avoiding much of the complexity inherent in managing both together. As dataset sizes grow or new datasets are considered, the scale-out storage system must be able to easily expand. Similarly, if more concurrent training is required, additional GPUs or other compute resources can be added without concern for their internal storage. Additionally, the above-described storage system may make it easier to build, operate, and grow AI systems because of the random read bandwidth provided by the storage system, the storage system's ability to randomly read small files (50 KB) at high speed (meaning that no extra effort is required to consolidate individual data points to create larger storage-friendly files), the storage system's ability to scale capacity and performance as data sets grow or throughput requirements increase, the storage system's ability to support files or objects, the storage system's ability to tune performance for large or small files (i.e., the user does not need to set up a file system), the storage system's ability to support non-destructive hardware and software updates even during production model training, and for many other reasons.
[0109]
[0167] Many types of input, including text, audio, or images, are natively stored as small files, so small file performance at the storage tier can be important. If the storage tier does not handle small files well, extra steps are required to preprocess and group samples into larger files. Storage built on top of a spinning platter that relies on SSDs as a caching tier may fall short of the required performance. Because training with random input batches yields more accurate models, the entire dataset must be accessible with full performance. SSD caches only provide high performance for a small subset of data and are ineffective at hiding the latency of the spinning drives.
[0110]
[0168] The reader will understand that the storage systems described above may be configured to support storage of blockchains (among other types of data). Such blockchains may be implemented as a continuously growing list of records, called blocks, that are linked and secured using cryptography. Each block in a blockchain may include a hash pointer as a link to the previous block, a timestamp, transaction data, etc. Blockchains may be designed to be resistant to data modification and can serve as an open distributed ledger that can record transactions between two parties efficiently, in a verifiable, and permanent manner. This makes blockchains potentially suitable for recording events, medical records, and other record-keeping activities, such as identity management, transaction processing, and others.
[0111]
[0169] The reader will further understand that in some embodiments, the storage systems described above may be paired with other resources to support the applications described above. For example, one infrastructure might include primary computing in the form of servers and workstations specialized in using general-purpose computing with graphics processing units ("GPGPUs") to accelerate deep learning applications, interconnected to a compute engine to train parameters for deep neural networks. Each system may have Ethernet external connectivity, InfiniBand external connectivity, some other form of external connectivity, or some combination thereof. In such examples, GPUs might be grouped for a single large training run or used independently to train multiple models. The infrastructure might also include a storage system, such as those described above to provide a scale-out all-flash file or object store, with data accessible through high-performance protocols such as NFS, S3, etc. The infrastructure might also include redundant top-of-rack Ethernet switches connected to the storage and computing via ports in MLAG port channels for redundancy. The infrastructure could also include additional computation in the form of white-box servers, optionally using GPUs, for data ingestion, pre-processing, and model debugging. The reader understands that additional infrastructure is also possible.
[0112]
[0170] The reader understands that the above-described system may be more suitable for the above-described applications than other systems, which may include distributed direct-attached storage (DDAS) solutions deployed on server nodes. Such DDAS solutions may be built to handle larger, less sequential accesses, but may not be able to handle smaller, random accesses as well. The reader understands that the above-described storage system may be utilized to provide the above-described applications with a platform that is preferable to leveraging cloud-based resources, or otherwise as part of a platform to support the above-described applications, because the storage system may be included in an on-site or in-house infrastructure that is more secure, more locally and internally managed, and more robust in feature set and performance. For example, services built on platforms such as IBM's Watson may require companies to distribute personal user information, such as financial transaction information or identifiable patient records, to other institutions. Thus, cloud-based offerings of AI as a service may be less desirable for various technical as well as business reasons than AI that is internally managed and provided as a service supported by a storage system, such as the above-described storage system.
[0113]
[0171] The reader will understand that the storage system described above may be configured to support other AI-related tools, either alone or in conjunction with other computing devices. For example, the storage system may utilize tools such as ONXX or other open neural network exchange formats that facilitate the transfer of models created in different AI frameworks. Similarly, the storage system may be configured to support tools such as Amazon's Gluon, which enable developers to prototype, build, and train deep learning models.
[0114]
[0172] The reader will further understand that the storage systems described above may be deployed as edge solutions. Such edge solutions may be implemented to optimize cloud computing systems by performing data processing at the edge of the network, closer to the source of the data. Edge computing may push applications, data, and computing power (i.e., services) from a centralized point to the logical edge of the network. Using an edge solution, such as the storage systems described above, computational tasks may be performed using the computational resources provided by such storage systems, data may be stored using the storage resources of the storage systems, and cloud-based services may be accessed using various resources of the storage systems (including networking resources). By performing computational tasks at edge solutions, storing data at edge solutions, and generally utilizing edge solutions, consumption of expensive cloud-based resources may be avoided, and in fact, performance improvements may be experienced with greater reliance on cloud-based resources.
[0115]
[0173] While many tasks may benefit from leveraging edge solutions, some specific applications may be particularly suited for deployment in such environments. For example, devices such as drones, autonomous vehicles, robots, and others may require high-speed processing—so fast, in fact, that uploading data back to a cloud environment for data processing support may simply be too slow. Similarly, machines such as locomotives and gas turbines, which generate large amounts of information through the use of abundant data-generating sensors, may benefit from the high-speed data processing capabilities of edge solutions. As an additional example, some IoT devices, such as connected video cameras, may not be suitable for leveraging cloud-based resources because sending data to the cloud may be impractical (not just from a privacy, security, or financial perspective) simply due to the sheer volume of data involved. Thus, many tasks that rely on data processing, storage, or communication may be better suited by platforms that include edge solutions, such as the storage systems described above.
[0116]
[0174] Consider the specific example of inventory management in a warehouse, distribution center, or similar location. A large inventory operation, freight storage operation, shipping operation, order fulfillment operation, manufacturing operation, or other operation has high-resolution digital cameras that produce large amounts of inventory on inventory shelves and a firehose of large amounts of data. All of this data may be captured into an image processing system that can reduce the amount of data into a firehose of smaller data. All of the smaller data may be stored on-premise in storage. The on-premise storage at the edge of the facility may be coupled to the cloud for external reporting, real-time control, and cloud storage. Inventory management may be performed using the results of the image processing, whereby inventory is tracked on the shelves, replenished, moved, shipped, revised with new products, or discontinued / obsolete products removed, etc. The above situation is a prime candidate for the configurable processing and storage system embodiments described above. A combination of compute-specific blades and offload blades suitable for image processing, perhaps with deep learning on offload FPGAs or offload custom blade(s), could take the firehose of massive amounts of data from all of the digital cameras and create a firehose of smaller data. All of the smaller data would then be stored by the storage node, and whichever combination of storage blade types best handles the data flow would work with the storage unit. This is an example of storage and function acceleration and consolidation. Depending on the need for external communication with and processing within the cloud, as well as the reliability of network connections and cloud resources, the system would be sized for managing bursty workloads and storage and compute with variable conductivity reliability. Also, depending on other inventory management aspects, the system could be configured for scheduling and resource management in a hybrid edge / cloud environment.
[0117]
[0175] The storage systems described above may also be optimized for use in big data analytics. Big data analytics may be broadly described as the process of examining large and varied data sets to uncover hidden patterns, unknown correlations, market trends, customer preferences, and other useful information that can help organizations make more informed business decisions. Big data analytics applications enable data scientists, predictive modelers, statisticians, and other analytics professionals to analyze ever-increasing volumes of structured transactional data, along with other forms of data often left untapped by traditional business intelligence (BI) and analytics programs. As part of that process, semi-structured and unstructured data, such as internet clickstream data, web server logs, social media content, text from customer emails and survey responses, mobile phone call detail records, IoT sensor data, and other data, may be converted into a structured form. Big data analytics is a form of advanced analytics that requires complex applications involving elements such as predictive models, statistical algorithms, and what-if analysis powered by high-performance analytics systems.
[0118]
[0176] The storage system described above may also support applications (including being implemented as a system interface) that perform tasks in response to human speech. For example, the storage system may support running intelligent personal assistant applications such as Amazon's Alexa, Apple's Siri, Google Voice, Samsung Bixby, Microsoft Cortana, and others. While the examples described above utilize voice as input, the storage system described above may also support chatbots, talkbots, chatterbots, or artificial conversational entities configured to conduct conversations via auditory or textual methods. Similarly, the storage system may actually execute such applications to allow users, such as system administrators, to interact with the storage system via speech. In embodiments according to the present disclosure, such applications may be utilized as interfaces for various system management operations, but generally, such applications may provide voice interaction, music playback, creating to-do lists, setting alarms, streaming podcasts, playing audiobooks, and other real-time information such as weather, traffic, and news.
[0119]
[0177] The storage systems described above may also implement an AI platform to achieve the vision of self-driving storage. Such an AI platform may be configured to achieve global predictive intelligence by collecting and analyzing large volumes of storage system telemetry data points to enable effortless management, analysis, and support. Indeed, such storage systems may be able to predict both capacity and performance, as well as generate intelligent advice for workload deployment, interaction, and optimization. Such an AI platform may be configured to scan all incoming storage system telemetry data against a library of problem fingerprints to capture hundreds of performance-related variables that are used to predict and resolve incidents and forecast performance loads in real time, before the incidents impact customer environments.
[0120]
[0178] For additional explanation, Figure 4 illustrates a block diagram showing multiple storage systems (402, 404, 406) supporting pods in accordance with some embodiments of the present disclosure. While not shown in great detail, the storage systems (402, 404, 406) shown in Figure 4 may be similar to the storage systems described above with respect to Figures 1A-1D, 2A-2G, 3A-3B, or any combination thereof. In practice, the storage systems (402, 404, 406) shown in Figure 4 may include the same components as the storage systems described above, fewer components, or additional components.
[0121]
[0179] 4, each of the storage systems (402, 404, 406) is shown as having at least one computer processor (408, 410, 412), computer memory (414, 416, 418), and computer storage (420, 422, 424). In some embodiments, the computer memory (414, 416, 418) and the computer storage (420, 422, 424) may be part of the same hardware device, while in other embodiments, the computer memory (414, 416, 418) and the computer storage (420, 422, 424) may be part of different hardware devices. The distinction between computer memory (414, 416, 418) and computer storage (420, 422, 424) in this particular example is that computer memory (414, 416, 418) may be physically proximate to computer processors (408, 410, 412) and may store computer program instructions executed by computer processors (408, 410, 412), while computer storage (420, 422, 424) may be implemented as non-volatile storage for storing user data, metadata describing the user data, etc. For example, with reference to the above example of FIG. 1A, the computer processors (408, 410, 412) and computer memory (414, 416, 418) for particular storage systems (402, 404, 406) may reside within one or more of the controllers (110A-110D). Meanwhile, the attached storage devices (171A-171F) may act as computer storage (420, 422, 424) within a particular storage system (402, 404, 406).
[0122]
[0180] In the example shown in FIG. 4 , the illustrated storage systems (402, 404, 406) may be attached to one or more pods (430, 432) in accordance with some embodiments of the present disclosure. Each of the pods (430, 432) shown in FIG. 4 may include a dataset (426, 428). For example, a first pod (430) with three attached storage systems (402, 404, 406) includes a first dataset (426), and a second pod (432) with two attached storage systems (404, 406) includes a second dataset (428). In such an example, when a particular storage system attaches to a pod, the pod's dataset is copied to the particular storage system and then kept up to date as the dataset is modified. A storage system can be removed from a pod, resulting in the dataset no longer being kept up to date on the removed storage system. In the example shown in FIG. 4, any storage system that is active for a pod (that is, the most recent, operational, non-faulted member of a non-faulted pod) can receive and process requests to modify or read the pod's dataset.
[0123]
[0181] In the example shown in FIG. 4 , each pod (430, 432) may include a collection of managed objects and management operations, as well as a collection of access operations to modify or read the dataset (426, 428) associated with the particular pod (430, 432). In such an example, a management operation may equally modify or query a managed object through either of the storage systems. Similarly, an access operation to read or modify a dataset may equally operate through either of the storage systems. In such an example, each storage system stores a separate copy of the dataset as an appropriate subset of the dataset stored and advertised for use by the storage system, but a managed object or operation to modify a dataset performed and completed through any one storage system is reflected in a subsequent managed object to query the pod or a subsequent access operation to read the dataset.
[0124]
[0182] The reader will understand that pods may implement more functionality than simply clustered, synchronously replicated data sets. For example, pods can be used to implement tenants whereby data sets are securely isolated from one another in some way. Pods can also be used to implement virtual arrays or virtual storage systems where each pod presents as a unique storage entity on a network (e.g., a storage area network or Internet Protocol network) with a separate address. In the case of a multi-storage system pod implementing a virtual storage system, all physical storage systems associated with the pod may present themselves in some way as the same storage system (e.g., as if multiple physical storage systems were no different than multiple network ports into a single storage system).
[0125]
[0183] The reader understands that a pod may also be a unit of management, representing a collection of volumes, file systems, object / analytic stores, snapshots, and other management entities, and that management changes in any one storage system (e.g., renaming, property changes, exporting, or managing permissions for any portion of the pod's dataset) are automatically reflected in all active storage systems associated with the pod. Additionally, a pod may also be a unit of data collection and data analysis, where performance and capacity metrics may be presented aggregated across all active storage systems for the pod, or presented to challenge data collection and analysis per pod separately, or perhaps presenting the contribution of each attached storage system to incoming content and performance per pod.
[0126]
[0184] One model of pod membership may be defined as a list of storage systems and a subset of that list of storage systems that are considered in sync for a pod. A storage system may be considered in sync for a pod if the pod is at least in recovery with identical idle content for the last read copy of the dataset associated with the pod. Idle content is the content after any ongoing modifications have completed without processing new modifications. This is sometimes referred to as "crash-recoverable" consistency. Pod recovery implements a process to reconcile differences in applying concurrent updates to synchronized storage systems within a pod. Recovery may resolve any inconsistencies between storage systems upon completion of concurrent modifications that were requested for various members of the pod but not signaled to either requestor as successfully completed. Storage systems that are listed as pod members but not listed as in sync for a pod can be described as "detached" from the pod. Storage systems listed as pod members are in sync for a pod and are currently available to actively serve data for the pod, being "online" for the pod.
[0127]
[0185] Each storage system member of a pod may have its own copy of the membership, including which storage systems it last knew were in sync and which storage systems it last knew contained the complete set of pod members. To be online for a pod, a storage system must consider itself in sync for the pod and be in communication with all other storage systems it considers in sync for the pod. If a storage system cannot be sure that it is in sync and in communication with all other storage systems it considers in sync, then it must stop processing new incoming requests for the pod (or complete them with an error or exception) until it is sure that it is in sync and in communication with all other storage systems it is in sync with. A first storage system may determine that a second paired storage system should be detached, which would allow the first storage system to continue because it is currently in sync with all systems in its list. Alternatively, however, the second storage system must be prevented from determining that the first storage system should be detached while the second storage system continues to operate. This would create a "split-brain" situation that could lead to incoherent datasets, dataset corruption, or application corruption, among other dangers.
[0128]
[0186] Situations requiring a determination of how to proceed when not communicating with a paired storage system may occur while a storage system is running normally and then realizes that communication has been lost, while it is currently recovering from some previous failure, while it is rebooting or resuming from a temporary power loss or restored communication outage, while it is switching operation from one set of storage system controllers to another for whatever reason, or during or after any combination of these or other types of events. In practice, whenever a storage system associated with a pod cannot communicate with all known non-detached members, the storage system either waits a short time until communication can be established, goes offline, and continues to wait, or the storage system may determine that it is safe to detach the non-communicating storage system by some means and then proceed, without risk of the non-communicating storage system inferring otherwise and causing a split-brain situation. If the safe detachment occurs quickly enough, the storage system can remain online for the pod with little more than a short delay and without causing application outages for applications that may issue requests to the remaining online storage system.
[0129]
[0187] An example of this situation is when a storage system may know it is out of date. This may occur, for example, when a first storage system is first added to a pod that is already associated with one or more storage systems, or when the first storage system reconnects to another storage system and finds that the other storage system has already marked the first storage system as detached. In this case, the first storage system simply waits until it connects to some other set of storage systems that it is in sync with for the pod.
[0130]
[0188] This model requires some consideration of how storage systems are added to or removed from a pod or from the syncing pod member list. Because each storage system has its own copy of the list, and because two independent storage systems cannot update their local copies at exactly the same time, and because the local copies are all that is available across reboots or various failure situations, care must be taken to ensure that transient inconsistencies do not cause problems. For example, if one storage system is syncing for a pod and a second storage system is added, and then the second storage system is updated to show both storage systems as syncing initially, and then there is a failure and restart of both storage systems, the second may start up and wait to connect to the first storage system. Meanwhile, the first may not be aware that it should or will wait for the second storage system. If the second storage system then responds to its inability to connect to the first storage system by going through a process to detach it, it may then succeed in completing a process that the first storage system is unaware of, resulting in a split-brain. Therefore, it may be necessary to ensure that storage systems do not inappropriately disagree as to whether they may choose to undergo a detachment process when they are not communicating.
[0131]
[0189] One way to ensure that storage systems do not inappropriately disagree regarding whether they may choose to undergo the detachment process if they are not communicating is to ensure that when adding a new storage system to the in-sync member list for a pod, the new storage system first remembers that it is a detached member (and possibly that it has been added as an in-sync member). Existing in-sync storage systems can then locally remember that the new storage system is an in-sync pod member before the new storage system remembers that same fact locally. If there is a reboot or network outage set up before the new storage system remembers its in-sync status, then the original storage system may detach the new storage system due to lack of communication, while the new storage system waits. To remove a communicating storage system from a pod, the reverse version of this change may be required. First, the removing storage system remembers that it is no longer in-sync, then the remaining storage systems remember that the removing storage system is no longer in-sync, and then all storage systems remove the removing storage system from their pod membership lists. Depending on the implementation, the intermediate persistent detach state may not be necessary. Whether or not local copies of membership lists require attention may depend on the model storage system to monitor each other or to validate its membership. If a consensus model is used for both, or an external system (or external distributed or clustered system) is used to store and validate pod membership, then inconsistencies in locally stored membership lists may not be an issue.
[0132]
[0190] When communication fails or one or more storage systems in a pod fail, or a storage system starts (or fails over to a secondary controller) and is unable to communicate with its paired storage systems for a pod, and one or more storage systems decide to detach one or more paired storage systems, some algorithm or mechanism must be utilized to determine that it is safe to do so and follow through on the detachment. One means of resolving detachment is to use a majority (or quorum) model for membership. In the case of three storage systems, as long as two are communicating, they can agree to detach the non-communicating third storage system, but that third storage system cannot independently choose to detach either of the other two. Chaos can occur when storage system communication is inconsistent. For example, storage system A may be communicating with storage system B but not with C. Meanwhile, storage system B may be communicating with both A and C. So A and B will detach C, or B and C will detach A, but more communication between pod members may be required to figure this out.
[0133]
[0191] The quorum membership model requires care when adding and removing storage systems. For example, if a fourth storage system is added, then the "majority" of storage systems is now three. Transitioning from three storage systems (where two is required for majority) to a pod containing a fourth storage system (where three is required for majority) may require something similar to the model described above to carefully add the storage system to the synchronization list. For example, the fourth storage system starts in an attaching state but is not yet attached, which would never trigger a quorum vote. Once in that state, each of the original three pod members is updated to recognize the fourth member and the new requirement for a majority of three storage systems to detach it. Removing a storage system from a pod may similarly move that storage system to a locally stored "detached" state before updating other pod members. A variant of this is to use a distributed consensus mechanism, such as PAXOS or RAFT, to implement any membership changes or process detach requests.
[0134]
[0192] Another means of managing membership transitions is to use an external system outside of the storage system itself to handle pod membership. To come online for a pod, the storage system must first contact the external pod membership system to verify that it is in sync for the pod. Any storage system that is online for a pod should then remain in communication with the pod membership system and wait or go offline if the storage system loses communication. An external pod membership manager could be implemented as a highly available cluster using a variety of cluster tools, such as Oracle RAC, Linux HA, VERITAS Cluster Server, IBM's HACMP, or others. An external pod membership manager could also use a distributed configuration tool such as Etcd or Zookeeper, or a reliable distributed database such as Amazon's DynamoDB.
[0135]
[0193] In the example shown in FIG. 4 , in accordance with some embodiments of the present disclosure, the illustrated storage systems (402, 404, 406) may receive a request to read a portion of a dataset (426, 428) and process the request to read the portion of the dataset locally. The reader will understand that because the dataset (426, 428) should be consistent across all storage systems (402, 404, 406) of the pod, a request (e.g., a write operation) to modify the dataset (426, 428) requires coordination among the storage systems (402, 404, 406) of the pod, but fulfilling a request to read the portion of the dataset (426, 428) does not require similar coordination among the storage systems (402, 404, 406). Thus, a particular storage system receiving a read request may service the read request locally by reading the portion of the dataset (426, 428) stored in its storage device, without synchronous communication with other storage systems in the pod. A read request received by one storage system for a replicated data set in a replicated cluster is expected to avoid any communication in the majority of cases, at least when received by a running storage system in a cluster that is also nominally running. Such reads should typically be processed normally by simply reading from the local copy of the clustered data, with no additional interaction with other storage systems in the cluster required.
[0136]
[0194] The reader understands that storage systems may take measures to ensure read consistency, so that read requests return the same result regardless of which storage system processes the read request. For example, the resulting clustered dataset contents for any set of updates received by any set of storage systems in a cluster should be consistent across the cluster, at least whenever the updates are idle (all previous modification operations are marked as complete, and no new update requests are received or processed). More specifically, instances of a clustered dataset across a set of storage systems may differ only as a result of updates that have not yet completed. That is, for example, any two write requests that overlap in their volume block ranges, or any combination of snapshots, compare-and-writes, or virtual block range copies that overlap with write requests, must produce consistent results across all copies of the dataset. Two operations should not produce results as if they occurred in one order on one storage system and in a different order on another storage system of a replicated cluster.
[0137]
[0195] Furthermore, read requests may be made time-ordered consistent. For example, if one read request is received and completed at a replicated cluster, and that read is then followed by another read request for an overlapping address range, and one or both reads overlap in time and volume address range with a modify request received by the replicated cluster in any way (whether the read or the modify is received by the same storage system or a different storage system in the replicated cluster), then if the first read reflects the results of the update, then the second read should also reflect the results of that update, rather than possibly returning data that preceded the update. If the first read does not reflect the update, then the second read will either reflect the update or not. This ensures that "time" for the data segment cannot progress backward between the two read requests.
[0138]
[0196] In the example shown in FIG. 4, the illustrated storage systems (402, 404, 406) may detect an interruption in data communication with one or more of the other storage systems and determine whether a particular storage system should remain in the pod. The interruption in data communication with one or more of the other storage systems may occur for a variety of reasons. For example, the interruption in data communication with one or more of the other storage systems may occur because one of the storage systems has failed, because a network interconnect has failed, or for some other reason. An important aspect of a synchronous replicated cluster is ensuring that any failure handling does not result in irrecoverable discrepancies or any discrepancies in responses. For example, if the network fails between two storage systems, at most one of the storage systems can continue to process new incoming I / O requests for the pod. And, if one storage system continues processing, the other storage system will not be able to process any new requests, including read requests, to completion.
[0139]
[0197] In the example shown in FIG. 4, the illustrated storage systems (402, 404, 406) may determine whether a particular storage system should remain in a pod in response to detecting an interruption in data communication with one or more of the other storage systems. As described above, to be "online" as part of a pod, a storage system must consider itself in sync for the pod and be in communication with all other storage systems that it considers in sync for the pod. If a storage system cannot be confident that it is in sync and in communication with all other storage systems that are in sync, then the storage system stops processing new incoming requests to access the datasets (426, 428). Thus, a storage system may determine whether a particular storage system should remain online as part of a pod through, for example, a combination of both steps where the storage system must verify that it can communicate with all other storage systems it considers to be in sync for the pod and that all other storage systems it considers to be in sync for the pod also consider the storage system to be attached to the pod, by determining (via one or more test messages) whether the storage system can communicate with all other storage systems it considers to be in sync for the pod, and that all other storage systems it considers to be in sync for the pod also consider the storage system to be attached to the pod, or through some other mechanism.
[0140]
[0198] 4, the illustrated storage systems (402, 404, 406) may keep the datasets on the particular storage system accessible for management and dataset operations in response to determining that the particular storage system should remain within the pod. The storage systems may keep the datasets (426, 428) on the particular storage system accessible for management and dataset operations, for example, by accepting and processing requests to access versions of the datasets (426, 428) stored on the storage systems, by accepting and processing management operations associated with the datasets (426, 428) issued by a host or authorized administrator, by accepting and processing management operations associated with the datasets (426, 428) issued by one of the other storage systems, or in some other manner.
[0141]
[0199] 4, however, the illustrated storage systems (402, 404, 406) may render datasets on the particular storage systems inaccessible for management and dataset operations in response to determining that the particular storage system should not remain in the pod. The storage systems may render datasets (426, 428) on the particular storage systems inaccessible for management and dataset operations, for example, by rejecting requests to access versions of the datasets (426, 428) stored on the storage systems, by rejecting management operations associated with the datasets (426, 428) issued by a host or other authorized administrator, by rejecting management operations associated with the datasets (426, 428) issued by one of the other storage systems in the pod, or in some other manner.
[0142]
[0200] In the example shown in Figure 4, the illustrated storage systems (402, 404, 406) may detect that an interruption in data communication with one or more of the other storage systems has been repaired and may make the datasets on the particular storage system accessible for management and dataset operations. The storage system may detect that the interruption in data communication with one or more of the other storage systems has been repaired, for example, by receiving a message from one or more of the other storage systems. In response to detecting that the interruption in data communication with one or more of the other storage systems has been repaired, the storage system may make the datasets (426, 428) on the particular storage system accessible for management and dataset operations once the previously detached storage system has resynchronized with the storage systems that remain attached to the pod.
[0143]
[0201] In the example shown in FIG. 4 , the illustrated storage systems (402, 404, 406) may go offline from the pod such that the particular storage system no longer enables management and dataset operations. The illustrated storage systems (402, 404, 406) may go offline from the pod such that the particular storage system no longer enables management and dataset operations for a variety of reasons. For example, the illustrated storage systems (402, 404, 406) may go offline from the pod because of some failure at the storage system itself, because an update or some other maintenance is occurring at the storage system, because of a communication failure, or for many other reasons. In such an example, as described in more detail in the resynchronization section included below, since the particular storage system went offline and came back online with the pod such that the particular storage system enables management and dataset operations, the illustrated storage systems (402, 404, 406) may subsequently update the dataset on the particular storage system to include all updates to the dataset.
[0144]
[0202] In the example shown in FIG. 4, the illustrated storage systems (402, 404, 406) may identify a target storage system to asynchronously receive a dataset, where the target storage system is not one of multiple storage systems to which the dataset is synchronously replicated. Such a target storage system may represent a backup storage system, for example, as any storage system that utilizes a synchronously replicated dataset. In practice, synchronous replication may be used to distribute copies of a dataset closer to some rack of servers for better local read performance. One such case is a smaller top-of-rack storage system that is symmetrically replicated to a larger storage system centrally located in a data center or campus, which is more carefully managed for reliability or connected to an external network for asynchronous replication or backup services.
[0145]
[0203] 4, the illustrated storage systems (402, 404, 406) may identify a portion of a dataset that has not been asynchronously replicated to the target storage system by any of the other storage systems and asynchronously replicate to the target storage system the portion of the dataset that has not been asynchronously replicated to the target storage system by any of the other storage systems, with the two or more storage systems collectively replicating the entire dataset to the target storage system. In this manner, the work associated with asynchronously replicating a particular dataset may be divided among the members of the pod, such that each storage system of the pod is only responsible for asynchronously replicating a subset of the dataset to the target storage system.
[0146]
[0204] In the example shown in FIG. 4 , the illustrated storage systems (402, 404, 406) may be detached from the pod, such that the particular storage system detached from the pod is no longer included in the set of storage systems across which a dataset is synchronously replicated. For example, if storage system (404) in FIG. 4 detaches from pod (430) shown in FIG. 4 , pod (430) would only include storage systems (402, 406) as storage systems across which dataset (426) included in pod (430) would be synchronously replicated. In such an example, detaching a storage system from a pod might also include deleting the dataset from the particular storage system detached from the pod. Continuing with the example in which storage system (404) in FIG. 4 detaches from pod (430) shown in FIG. 4 , dataset (426) included in pod (430) might be deleted or otherwise removed from storage system (404).
[0147]
[0205] The reader will also understand that there are several unique management features enabled by the pod model that can be supported. The pod model itself also presents several issues that may be addressed by an implementation. For example, when a storage system is offline for a pod but is otherwise running because an interconnect failed and another storage system for the pod won arbitration, there may be a desire or need to access the offline pod's dataset on the offline storage system. One solution may be to simply enable the pod in some detached mode so that the dataset can be accessed. However, that solution is risky, and it may make it much more difficult to reconcile the pod's metadata and data when the storage systems regain communication. Furthermore, a host may still have separate paths to access the offline storage system as well as the online storage system. In that case, the host may issue I / O to both storage systems even though they are no longer kept in sync. This is because the host sees target ports reporting a volume with the same identifier, and the host I / O driver assumes it sees additional paths to the same volume. Even if the host makes a naive guess, this can result in rather detrimental data corruption, as the reads and writes issued to both storage systems are no longer in agreement. A variant of this case is in a clustered application, such as a shared storage clustered database, where a clustered application running on one host may be reading from or writing to one storage system, and the same clustered application running on another host may be reading from or writing to a "detached" storage system, but the two instances of the clustered application are communicating between each other under the assumption that the data sets they each see are in perfect agreement with respect to completed writes.Because they do not match, the assumption is violated and the application's dataset (e.g., database) can quickly become corrupted.
[0148]
[0206] One way to solve both of these problems is to allow offline pods, or perhaps snapshots of offline pods, to be copied to a new pod with a new volume that has a sufficiently new identity that host I / O drivers and clustered applications do not confuse the copied volume with a still-online volume on another storage system. Because each pod maintains a complete copy of the dataset that is crash-consistent but perhaps slightly different from the copy of the pod dataset on another storage system, and because each pod has an independent copy of all data and metadata needed to operate on the pod contents, making a virtual copy of some or all of the pod's volumes or snapshots to a new volume in the new pod is a straightforward problem. For example, in a logical extent graph implementation, all that is required is to define a new volume in the new pod that references the pod's volumes or snapshots and the logical extent graph from the copied pod associated with the logical extent graph being marked as copied on write. Similar to how a volume snapshot copied to a new volume might be implemented, the new volume should be treated as a new volume. The volume may have the same administrative name, albeit in the new pod namespace, but it should have a different base identifier and a different logical unit identifier than the original volume.
[0149]
[0207] In some cases, it may be possible to use virtual network isolation techniques (e.g., by creating a virtual LAN in the case of an IP network or a virtual SAN in the case of a Fibre Channel network) so that isolation of volumes presented to some interfaces can be ensured to be inaccessible from host network interfaces or host SCSI initiation ports that might also see the original volume. In such cases, it may be safe to provide the copy of the volume with the same SCSI or other storage identifiers as the original volume. This could be used, for example, if an application expects to see a particular set of storage identifiers in order to function without undue burden on reconfiguration.
[0150]
[0208] Some of the techniques described herein may also be used outside of active disaster situations to test readiness to handle disasters. Readiness testing (sometimes called "fire drills") is typically required for disaster recovery configurations, and frequent and repeated testing is deemed necessary to ensure that most or all aspects of the disaster recovery plan are correct and account for any recent changes to the application, data set, or equipment changes. Readiness testing should be non-disruptive to current production operations, including replicas. While in many cases, the actual operations cannot actually be invoked in the active configuration, a good approach is to use storage operations to make copies of the production data set, and then perhaps combine this with the use of virtual networking to create an isolated environment containing all the data deemed necessary for critical applications that must start successfully in the event of a disaster. Having such copies of synchronously replicated (or even asynchronously replicated) data sets available within the site (or collection of sites) that is expected to perform disaster recovery readiness testing procedures and then start critical applications on that data set to ensure it starts and functions is an excellent tool because it helps ensure that critical portions of the application data set were not left out in the disaster recovery plan. Where necessary and practical, this could be combined with a virtual isolation network, perhaps coupled with an isolated collection of physical or virtual machines, to approximate as closely as possible a real-world disaster recovery takeover situation. Virtually copying a pod (or collection of pods) to another pod as a point-in-time image of the pod dataset not only allows isolation to a single site (or several sites) separate from the original pod, but also immediately creates an isolated dataset that contains all the copied elements and can be operated essentially the same as the original pod. Furthermore, these are fast operations, and they can be easily dismantled and tested repeatedly as often as desired.
[0151]
[0209] Several enhancements could be made to move even closer toward a complete disaster recovery test. For example, in conjunction with isolated networks, SCSI logical unit identities or other types of identities would be copied to the target pod so that test servers, virtual machines, and applications see the same identifiers. Furthermore, the server management environment would be configured to respond to requests from a specific virtual set of virtual networks and respond to requests or operations on the original pod name so that scripts do not require the use of test variants along with alternate "test" versions of object names. Additional enhancements can be used when the hosting server infrastructure that takes over in the event of a disaster is available for use during testing. This includes cases where a disaster recovery datacenter is generally fully equipped with alternate server infrastructure that is not used until dictated to do so by a disaster. It also includes cases where the infrastructure may be used for non-critical operations (e.g., running analytics on production data, or simply supporting application development or other functions that may be critical but can be paused if needed for more critical functions). Specifically, host definitions and configurations, and the server infrastructure that uses them, can be set up as they are for an actual disaster recovery takeover event, tested as part of the disaster recovery takeover test, with the tested volumes connected to these host definitions from the virtual pod copies used to provide snapshots of the datasets. From the perspective of the storage systems involved, these host definitions and configurations used for testing, and the volume-to-host connection configurations used during testing, can then be reused when an actual disaster takeover event is triggered, significantly minimizing the configuration differences between the test configuration and the actual configuration that will be used in the event of a disaster recovery takeover.
[0152]
[0210] In some cases, it may make sense to move volumes from a first pod to a new, second pod containing only those volumes. Pod membership and high availability and recovery features can then be adjusted separately, and management of the two resulting pod data sets can then be decoupled from each other. Also, operations that can be performed in one direction should also be possible in the other direction. At some point, it may make sense to take two pods and merge them into one, such that the volumes in each of the original two pods now track each other for storage system membership and high availability and recovery features and events. Both operations can be achieved safely and with reasonably minimal or no disruption to running applications by relying on the proposed features for changing brokering or quorum characteristics for pods, as described in the previous sections. For example, with brokering, a pod's mediator can be changed using a sequence of steps in which each storage system in the pod is changed to depend on both the first mediator and the second mediator, and then each is changed to depend only on the second mediator. If a failure occurs mid-sequence, some storage systems may rely on both the first mediator and the second mediator, but recovery and failure handling never results in some storage systems relying solely on the first mediator and others relying solely on the second mediator. Quorum can be handled similarly by temporarily relying on winning both the first and second quorum models to proceed to recovery. This may result in a very brief period where the availability of a pod facing a failure depends on additional resources, thus reducing potential availability, but this period is so brief that the reduction in availability is often negligible.With intermediation, if the change in mediator parameters is merely a change in the key used for intermediation and the intermediary service used is the same, then the potential reduction in availability is also less, since it now depends on two calls to the same service versus one call to that service, and only on two separate services rather than separate calls.
[0153]
[0211] The reader should note that changing the quorum model can be quite complex. An additional step may be required in which the storage system joins the second quorum model but does not rely on the second quorum model to prevail, followed by a step that also relies on the second quorum model. This may be necessary to account for the fact that if only one system processes the change to rely on the quorum model, then it will never achieve quorum because there is never a majority. With this model in place to change the high availability parameters (brokerage relationship, quorum model, takeover priority), safety procedures can be created for these operations to split a pod into two or join two pods into one. This may require adding one other feature: linking a second pod to a first pod for high availability, so that if the two pods contain compatible high availability parameters, the second pod linked to the first pod can rely on the first pod to determine and drive detach relationship processing and behavior, offline and sync states, and recovery and resync actions.
[0154]
[0212] To split one pod into two, a distributed operation may be formed to move some volumes into a newly created pod, which can be described as follows: create a second pod that moves a collection of volumes that were previously in the first pod; copy high availability parameters from the first pod to the second pod to ensure they are compatible for linking; and link the second pod to the first pod for high availability. This operation may be encoded as a message and should be implemented by each storage system in the pod, such that the storage system ensures that the operation occurs entirely on that storage system, or not at all if the operation is interrupted by a failure. Once all synchronized storage systems for the two pods have processed this operation, the storage system may then process a subsequent operation that modifies the second pod so that it is no longer linked to the first pod. As with other changes to high availability features for pods, this involves first having each in-sync storage system rely on both the previous model (where high availability is linked to the first pod) and the new model (where it is now highly available independently). In the case of brokering or quorum, this means that the storage system that processed the change first relies on the brokering or quorum achieved accordingly for the first pod, and then relies on the new, separate brokering (e.g., new broker key) or quorum being achieved for the second pod before proceeding following the failure that required the test for brokering or quorum. As with the above description of changing the quorum model, an intermediate step may be configuring the storage system to participate in a quorum for the second pod before the step in which the storage system joins and relies on the quorum for the second pod. Once all in-sync storage systems have processed the change to rely on the new parameters for brokering or quorum for both the first pod and the second pod, the split is complete.
[0155]
[0213] Joining a second pod to a first pod essentially works in reverse. First, the second pod must be adjusted to be compatible with the first pod by having an identical list of storage systems and a compatible high availability model. This may include any set of steps, such as steps for adding or removing storage systems or for modifying mediators and quorum models, as described elsewhere herein. Depending on the implementation, arriving at an identical list of storage systems may be all that is necessary. Joining proceeds by processing an operation at each in-sync storage system to link the second pod to the first pod for high availability. Each storage system that processes the operation then depends on the first pod for high availability and then on the second pod for high availability. Once all in-sync storage systems for the second pod have processed the operation, the storage systems then remove the link between the second pod and the first pod, migrate volumes from the second pod to the first pod, and remove the second pod, respectively. Host or application data set access can be preserved through these operations as long as the implementation allows appropriate direction of host or application data set modification or read operations to volumes by identity, and as long as identity is preserved appropriately for the storage protocol or storage model (e.g., as long as logical unit identifiers for volumes and use of target ports to access volumes are preserved in the case of SCSI).
[0156]
[0214] Migrating volumes between pods can present problems. If the pods have the same set of in-sync membership storage systems, then it may be simple to temporarily suspend operations on the migrating volumes, switch control over operations on those volumes to the control software and structures for the new pod, and then resume operations. If networks and ports migrate appropriately between pods, this allows for a seamless migration with continuous uptime for applications, apart from a very brief interruption in operations. Depending on the implementation, the suspend operation may not be necessary, or may be so internal to the system that the interruption in operations has no impact. Copying volumes between pods with different in-sync membership sets is increasingly problematic. This is less of an issue if the target pod of the copy has a subset of in-sync members from the source pod. Member storage systems can be removed safely enough without the need for more work. However, if the target pod adds an in-sync member storage system to a volume on the source pod, then the added storage system must be synchronized to include the volume's contents before they can be used. Until synchronized, this leaves the copied volumes distinctly different from volumes that are already synchronized in that failure handling is different and requests from member storage systems that are not yet synchronized may not work or may have to be forwarded or may not be fast because reads must traverse the interconnect. Also, the internal implementation must handle some volumes that are in synchronization and ready for failure handling while others are not synchronized.
[0157]
[0215] There are other issues related to the reliability of operations in the face of failures. Coordinating the migration of volumes across multiple storage system pods is a distributed operation. If pods are the unit of failure handling and recovery, and if intermediaries or quorums or whatever means are used to avoid split-brain situations, then switching a volume from one pod to another pod, and then to the pod's storage system, with a particular set of state and configurations and relationships for failure handling, recovery, intermediaries, and quorums must be careful to coordinate the changes involved in that operation for every volume. Operations cannot be atomically distributed across storage systems, but must be staged in some way. The intermediary and quorum models inherently provide tools for implementing distributed transaction atomicity within pods, but this may not extend to inter-pod operations without increasing implementation complexity.
[0158]
[0216] Consider also the simple migration of a volume from a first pod to a second pod in the case of two pods sharing the same first and second storage systems. At some point, the storage systems coordinate to determine that the volume is now in the second pod and no longer in the first pod. Without an inherent mechanism for transaction atomicity across storage systems for the two pods, a native implementation might leave the volume in the first pod on the first storage system and the second pod on the second storage system in the event of a network failure that causes a failure process to detach the storage systems from the two pods. If the pods independently determine which storage system succeeds in detaching the other, the result would be either the same storage system detaching the other for both pods, in which case the volume migration recovery results must be consistent, or different storage systems would detach the other for the two pods. If a first storage system detaches a second storage system for a first pod and the second storage system detaches the first storage system for the second pod, then recovery will recover the volume to the first pod on the first storage system and to the second pod on the second storage system, and the volume will then be running and exported to hosts and storage applications on both storage systems. Alternatively, if a second storage system detaches the first storage system for the first pod and the first storage system detaches the second storage system for the second pod, then recovery may result in the volume being discarded from the second pod by the first storage system and the volume being discarded from the first pod by the second storage system, causing the volume to disappear entirely. If the pods between which the volume is being migrated are on different sets of storage systems, then things may become even more complicated.
[0159]
[0217] A solution to these problems may be to use an intermediate pod in conjunction with the techniques described above for splitting and joining pods. This intermediate pod must never be presented as a visible management object associated with the storage system. In this model, a volume to be moved from a first pod to a second pod is first split from the first pod into the new intermediate pod using the split operation described above. The storage system membership for the intermediate pod can then be adjusted to match the membership of the storage system by adding or removing storage systems from the pod as needed. The intermediate pod may then be joined to the second pod.
[0160]
[0218] For additional explanation, Figure 5 sets forth a flowchart illustrating steps that may be performed by storage systems (402, 404, 406) supporting pods in accordance with some embodiments of the present disclosure. While not shown in great detail, the storage systems (402, 404, 406) illustrated in Figure 5 may be similar to the storage systems described above with respect to Figures 1A-1D, 2A-2G, 3A-3B, 4, or any combination thereof. In practice, the storage systems (402, 404, 406) illustrated in Figure 5 may include the same components as the storage systems described above, fewer components, or additional components.
[0161]
[0219] In the example method shown in FIG. 5, a storage system (402) may attach to a pod (508). The model for pod membership may include a list of storage systems and a subset of that list of storage systems that are presumed to be in sync for the pod. A storage system is in sync for a pod if it is in recovery with the same idle content as at least the last created copy of the dataset associated with the pod. Idle content is the content after no new modifications are processed and any ongoing modifications have completed. This is sometimes referred to as "crash recoverable" consistency. A storage system that is listed as a pod member but not listed as in sync for the pod may be described as "detached" from the pod. A storage system that is listed as a pod member is in sync for the pod and is currently available to actively serve data for the pod, being "online" for the pod.
[0162]
[0220] 5, a storage system (402) may attach (508) to a pod by synchronizing its locally stored version of the dataset (426) with the latest versions of the dataset (426) stored in the pod's other storage systems (404, 406) that are online, e.g., as the term is described above. In such an example, in order for the storage system (402) to attach (508) to the pod, the pod definition stored locally in each of the storage systems (402, 404, 406) in the pod needs to be updated in order for the storage system (402) to attach (508) to the pod. In such an example, each storage system member of the pod may have its own copy of the membership, including which storage systems it last knew were in sync and which storage systems it last knew contained the entire collection of pod members.
[0163]
[0221] 5, a storage system (402) may receive (510) a request to read a portion of a dataset (426), and the storage system (402) may locally process (512) the request to read the portion of the dataset (426). The reader will understand that while a request to modify the dataset (426) (e.g., a write operation) requires coordination among the storage systems (402, 404, 406) in the pod, fulfilling a request to read the portion of the dataset (426) does not require similar coordination among the storage systems (402, 404, 406) because the dataset (426) must be consistent across all storage systems (402, 404, 406) in the pod. Thus, a particular storage system (402) receiving a read request may service the read request locally by reading a portion of the dataset (426) stored in its (402) storage device, without synchronous communication with other storage systems (404, 406) in the pod. Read requests received by one storage system for a replicated dataset in a replicated cluster are expected to avoid any communication in the vast majority of cases, at least when received by a storage system that is also running in a nominally running cluster. Such reads should typically be successfully processed by simply reading from the local copy of the clustered dataset, with no additional interaction with other storage systems in the cluster required.
[0164]
[0222] The reader understands that storage systems may take steps to ensure read consistency, so that a read request returns the same result regardless of which storage system processes the read request. For example, the resulting clustered dataset contents for any set of updates received by any set of storage systems in a cluster should be consistent across the cluster, at least at any point in time when the updates are idle (all previous modification operations are marked as completed, and no new update requests have been received or processed). That is, instances of a clustered dataset across a set of storage systems may differ only as a result of updates that have not yet completed. That is, for example, any two write requests that overlap on that volume block range, or any combination of snapshot, compare-and-write, or virtual block range copy that overlap with a write request, must produce consistent results for all copies of the dataset. Two operations cannot produce results as if they occurred in one order on one storage system and in a different order on another storage system of a replication cluster.
[0165]
[0223] Furthermore, read requests may be time-consistent. For example, if one read request is received and completed at a replicated cluster, and then that read is followed by another read request for an overlapping address range that is received by the replicated cluster, and one or both reads overlap in time and volume address range to any extent with the modify request received by the replicated cluster (regardless of whether the read or the modify is received by the same storage system in the replicated cluster or a different storage system), then if the first read reflects the results of the update, then the second read should also reflect the results of that update, perhaps rather than returning data that preceded the update. If the first read does not reflect the update, then the second read may or may not reflect the update. This ensures that the data segment's "time" never goes back in time between the two read requests.
[0166]
[0224] In the example method shown in FIG. 5, the storage system (402) may detect (514) an interruption in data communication with one or more of the other storage systems (404, 406). The interruption in data communication with one or more of the other storage systems (404, 406) may occur for a variety of reasons. For example, the interruption in data communication with one or more of the other storage systems (404, 406) may occur because one of the storage systems (402, 404, 406) has failed, because the network interconnect has failed, or for some other reason. An important aspect of synchronous replication clustering is ensuring that any failure handling does not result in an irrecoverable inconsistency or correspondence inconsistency. For example, if the network fails between two storage systems, at most one of the storage systems may continue to process new incoming I / O requests for the pod. And, if one storage system continues processing, the other storage system will not be able to process new requests, including read requests, to completion.
[0167]
[0225] In the example method shown in FIG. 5, a storage system (402) may determine (516) whether a particular storage system (402) should remain online as part of a pod. As discussed above, to be “online” as part of a pod, a storage system must consider itself in sync for the pod and be in communication with all other storage systems that it considers in sync for the pod. If a storage system cannot be sure that it is in sync and in communication with all other storage systems that are in sync, then the storage system may stop processing new incoming requests to access the dataset (426). Thus, a storage system (402) may determine (516) whether a particular storage system (402) should remain online as part of a pod by determining (e.g., via one or more test messages) whether the storage system can communicate with all other storage systems (404, 406) that it considers to be in sync for the pod, by determining whether all other storage systems (404, 406) that it considers to be in sync for the pod also consider the storage system (402) to be attached to the pod, by determining whether the particular storage system (402) must verify that it can communicate with all other storage systems (404, 406) that it considers to be in sync for the pod and that all other storage systems (404, 406) that it considers to be in sync for the pod also consider the storage system (402) to be attached to the pod, or by some other mechanism.
[0168]
[0226] 5 , in response to affirmatively determining (518) that a particular storage system (402) should remain online as part of a pod, a storage system (402) may keep (522) a dataset (426) on the particular storage system (402) accessible for management and dataset operations. The storage system (402) may make (522) a dataset (426) on the particular storage system (402) accessible for management and dataset operations by, for example, accepting and processing requests to access versions of the dataset (426) stored on the storage system (402), accepting and processing management operations associated with the dataset (426) issued by a host or authorized administrator, accepting and processing management operations associated with the dataset (426) issued by one of the other storage systems (404, 406) in the pod or in some other manner.
[0169]
[0227] 5, the storage system 402 may also make the dataset 426 on the particular storage system 402 inaccessible for management and dataset operations in response to determining 520 that the particular storage system should not remain online as part of the pod. The storage system 402 may make 524 the dataset 426 on the particular storage system 402 inaccessible for management and dataset operations, for example, by rejecting requests to access the version of the dataset 426 stored on the storage system 402, by rejecting management operations associated with the dataset 426 issued by a host or other authorized administrator, by rejecting management operations associated with the dataset 426 issued by one of the other storage systems 404, 406 in the pod, or in some other manner.
[0170]
[0228] 5, the storage system (402) may detect (526) that a disruption in data communication with one or more of the other storage systems (404, 406) has been repaired. The storage system (402) may detect (526) that a disruption in data communication with one or more of the other storage systems (404, 406) has been repaired, for example, by receiving a message from one or more of the other storage systems (404, 406). In response to detecting (526) that a disruption in data communication with one or more of the other storage systems (404, 406) has been repaired, the storage system (402) may make (528) the dataset (426) on the particular storage system (402) accessible for management and dataset operations.
[0171]
[0229] The reader will understand that the example shown in Figure 5 describes an embodiment in which various actions are shown as occurring in the same order, although no ordering is required. Additionally, other embodiments may exist in which the storage system (402) only performs a subset of the described actions. For example, the storage system (402) may perform the steps of detecting (514) an interruption in data communication with one or more of the other storage systems (404, 406), determining (516) whether the particular storage system (402) should remain in the pod, keeping (522) a dataset (426) on the particular storage system (402) accessible for management and dataset operations, or making (524) a dataset (426) on the particular storage system (402) inaccessible for management and dataset operations without first receiving (510) a request to read a portion of the dataset (426) and processing (512) the request to read the portion of the dataset (426) locally. Additionally, a storage system (402) may detect (526) that an interruption in data communication with one or more of the other storage systems (404, 406) has been repaired and may make (528) the dataset (426) on the particular storage system (402) accessible for management and dataset operations without first receiving (510) a request to read a portion of the dataset (426) and processing (512) the request to read the portion of the dataset (426) locally. Indeed, none of the steps described herein are explicitly required in all embodiments as a prerequisite for performing any other step described herein.
[0172]
[0230] For additional explanation, Figure 6 sets forth a flowchart illustrating steps that may be performed by storage systems (402, 404, 406) supporting pods in accordance with some embodiments of the present disclosure. While not shown in great detail, the storage systems (402, 404, 406) illustrated in Figure 6 may be similar to the storage systems described above with respect to Figures 1A-1D, 2A-2G, 3A-3B, 4, or any combination thereof. In practice, the storage systems (402, 404, 406) illustrated in Figure 6 may include the same components as the storage systems described above, fewer components, or additional components.
[0173]
[0231] In the example method shown in Figure 6, two or more of the storage systems (402, 404) may each identify (608) a target storage system (618) for asynchronously receiving the dataset (426). The target storage system (618) for asynchronously receiving the dataset (426) may be implemented as cloud storage provided by a cloud service provider or in many other ways, such as a backup storage system located in a different data center than either of the storage systems (402, 404) that are members of a particular pod. The reader will understand that the target storage system (618) is not one of multiple storage systems (402, 404) across which the dataset (426) is synchronously replicated, and therefore the target storage system (618) does not initially contain an up-to-date local copy of the dataset (426).
[0174]
[0232] 6, two or more of the storage systems (402, 404) may each identify (610) a portion of the dataset (426) that has not been asynchronously replicated to a target storage (618) system by any of the other storage systems that are members of the pod that includes the dataset (426). In such an example, the storage systems (402, 404) may each asynchronously replicate (612) to the target storage system (618) a portion of the dataset (426) that has not been asynchronously replicated to the target storage system by any of the other storage systems. Consider an example in which a first storage system (402) is responsible for asynchronously replicating a first portion of the dataset (426) (e.g., a first half of its address space) to the target storage system (618). In such an example, the second storage system (404) would be responsible for asynchronously replicating a second portion of the dataset (426) (e.g., a second half of the address space) to the target storage system (618) such that the two or more storage systems (402, 404) collectively replicate the entire dataset (426) to the target storage system (618).
[0175]
[0233] The reader will understand that, as described above, by using pods, the replication relationship between two storage systems can be switched from one in which data is asynchronously replicated to one in which data is synchronously replicated. For example, if storage system A is configured to asynchronously replicate a dataset to storage system B, creating a pod that includes the dataset, storage system A as a member, and storage system B as a member can switch the relationship from one in which data is asynchronously replicated to one in which data is synchronously replicated. Similarly, by using pods, the replication relationship between two storage systems can be switched from one in which data is synchronously replicated to one in which data is asynchronously replicated. For example, if a pod that includes a dataset, storage system A as a member, and storage system B as a member is created by simply not stretching the pod (to remove storage system A as a member or to remove storage system B as a member), the relationship in which data is synchronously replicated between the storage systems can immediately be switched to one in which data is asynchronously replicated. In this manner, storage systems can be switched back and forth between asynchronous replication and synchronous replication as needed.
[0176]
[0234] This switch may be facilitated by an implementation relying on similar techniques for both synchronous and asynchronous replication. For example, if resynchronization of a synchronously replicated dataset relies on the same or compatible mechanism as used for asynchronous replication, then switching to asynchronous replication is conceptually identical to dropping the in-sync state and leaving the relationship in a state similar to "durable recovery" mode. Similarly, switching from asynchronous to synchronous replication may conceptually work by completing resynchronization and "catching up" as the switching system becomes an in-sync pod member, just as it does.
[0177]
[0235] Alternatively, or in addition, if both synchronous and asynchronous replication rely on similar or identical common metadata, or a common model for representing and identifying logical extents or stored block identities, or a common model for representing content-addressable stored blocks, then these aspects of commonality may be exploited to dramatically reduce the content that may need to be transferred when switching to and from synchronous and asynchronous replication. Furthermore, if a data set is asynchronously replicated from storage system A to storage system B, and system B further asynchronously replicates that data set to storage system C, then the common metadata model, common logical extents or block identities, or common representation of content-addressable storage blocks can dramatically reduce the data transfers required to enable synchronous replication between storage system A and storage system C.
[0178]
[0236] The reader will further understand that by using pods, as described above, replication techniques can be used to perform tasks other than data replication. Indeed, since a pod may include a collection of managed objects, tasks such as migrating a virtual machine can be implemented using the pod and replication techniques described herein. For example, if virtual machine A is running on storage system A, by creating a pod that includes virtual machine A as a managed object, storage system A as a member, and storage system B as a member, virtual machine A and any associated images and definitions can be migrated to storage system B, at which time the pod may simply be destroyed, membership may be updated, or other action may be taken as necessary.
[0179]
[0237] For additional explanation, Figure 7 sets forth a flowchart illustrating an example method for establishing a synchronous replication relationship between two or more storage systems (714, 724, 728) according to some embodiments of the present disclosure. While not shown in great detail, the storage systems (714, 724, 728) illustrated in Figure 7 may be similar to the storage systems described above with respect to Figures 1A-1D, 2A-2G, 3A-3B, or any combination thereof. In practice, the storage systems (714, 724, 728) illustrated in Figure 7 may include the same components as the storage systems described above, fewer components, or additional components.
[0180]
[0238] The example method shown in Figure 7 includes identifying (702), for a data set (712), multiple storage systems (714, 724, 728) across which the data set (712) is synchronously replicated. The data set (712) shown in Figure 7 may be embodied, for example, as the contents of a particular volume, as the contents of a particular shard of a volume, or as any other collection of one or more data elements. The data set (712) may be synchronized across the multiple storage systems (714, 724, 728) such that each storage system (714, 724, 728) maintains a local copy of the data set (712). In the example described herein, such a dataset (712) is synchronously replicated across storage systems (714, 724, 728) such that the dataset (712) can be accessed through any of the storage systems (714, 724, 728) that have performance characteristics such that any one storage system in the cluster operates substantially less optimally than any other storage system in the cluster, at least as long as the cluster and the particular storage system being accessed are nominally operating. In such a system, modifications to the dataset (712) should be made to copies of the dataset resident on each storage system (714, 724, 728) so that accessing the dataset (712) on any of the storage systems (714, 724, 728) produces consistent results. For example, a write request issued to the dataset must be serviced by all storage systems (714, 724, 728) or by none of the storage systems (714, 724, 728). Similarly, some groups of operations (e.g., two write operations directed to the same location in a dataset) must be executed in the same order on all storage systems (714, 724, 728) so that the copies of the dataset resident on each storage system (714, 724, 728) are ultimately identical.Modifications to the dataset (712) need not occur exactly simultaneously, although some actions (e.g., issuing an acknowledgment that a write request is directed to the dataset and enabling read access to locations within the dataset targeted by write requests that have not yet completed on all storage systems) may be deferred until copies of the dataset (712) on each storage system (714, 724, 728) have been modified.
[0181]
[0239] In the example method shown in FIG. 7 , identifying (702), for a dataset (712), multiple storage systems (714, 724, 728) across which the dataset (712) is synchronously replicated may be performed, for example, by examining a pod definition or similar data structure that associates the dataset (712) with one or more storage systems (714, 724, 728) that nominally store the dataset (712). A "pod," as the term is used here and throughout the remainder of this application, may be implemented as a management entity representing a dataset, a collection of management objects and management operations, a collection of access operations to modify or read the dataset, and multiple storage systems. Such management operations may modify or query management objects equally through any of the storage systems, and access operations to read or modify the dataset equally through any of the storage systems. Each storage system may store a separate copy of the dataset as an appropriate subset of the dataset stored and advertised for use by the storage system, and operations that modify a management object or dataset performed and completed through any one storage system are reflected in subsequent management objects to query the pod or subsequent access operations to read the dataset. Additional details regarding "pods" may be found in previously filed provisional patent application Ser. No. 62 / 518,071, which is incorporated herein by reference. In such an example, a pod definition may include at least an identification of a dataset (712) and a collection of storage systems (714, 724, 728) across which the dataset (712) is synchronously replicated. Such a pod may encapsulate some of many (possibly optional) characteristics, including symmetric access, flexible addition / removal of replicas, highly available data consistency, centralized user management across storage systems in relation to the dataset, management host access, application clustering, etc.A storage system may be added to a pod, resulting in the pod's dataset (712) being copied to that storage system and then kept up to date as the dataset (712) is modified. A storage system may also be removed from a pod, resulting in the dataset (712) being no longer kept up to date with the removed storage system. In such an example, the pod definition or similar data structure may be updated as storage systems are added to and removed from a particular pod.
[0182]
[0240] The example method shown in FIG. 7 also includes configuring (704) one or more data communication links (716, 718, 720) between each of the multiple storage systems (714, 724, 728) used to synchronously replicate the data set (712). In the example method shown in FIG. 6, the storage systems (714, 724, 728) within a pod must communicate with each other both for high-bandwidth data transfers and for cluster, status, and management communications. These distinct types of communications may be over the same data communication links (716, 718, 720), or in alternative embodiments, these distinct types of communications may be over separate data communication links (716, 718, 720). In a cluster of dual-controller storage systems, both controllers of each storage system should have the nominal ability to communicate with both controllers for any paired storage system (i.e., any other storage system within the pod).
[0183]
[0241] In a primary / secondary controller design, all cluster communications for active replication may occur between primary controllers until a failure occurs. In such systems, some communication may occur between the primary and secondary controllers, or between secondary controllers of separate storage systems, to verify that the data communication links between such entities are operational. In other cases, virtual network addresses may be used to limit the configuration required for network links between data centers or to simplify the design of clustered aspects of storage systems. In an active / active controller design, cluster communications may continue from all active controllers of one storage system to some or all active controllers of any paired storage system, or the cluster communications may be filtered through a common switch, or the cluster communications may use virtual network addresses to simplify configuration, or the cluster communications may use some combination. In a scale-out design, two or more common network switches may be used to connect all scale-out storage controllers in the storage system to the network switch for handling data traffic. The switch may or may not use techniques to limit the number of network addresses exposed, so that the paired storage system does not need to be configured with the network addresses of all storage controllers.
[0184]
[0242] 7 , configuring 704 one or more data communication links (716, 718, 720) between each of the multiple storage systems (714, 724, 728) used to synchronously replicate the data set (712) may be performed, for example, by configuring the storage systems (716, 718, 720) to communicate over defined ports over a data communication network, by configuring the storage systems (716, 718, 720) to communicate over a point-to-point data communication link between two of the storage systems (716, 724, 728), or in a variety of other ways. If secure communication is required, some form of key exchange may be required, or communication could occur or be initiated through some service, such as SSH (Secure SHell), SSL, or some other service or protocol built around public key or Diffie-Hellman key sharing, or a suitable alternative. Secure communications could also be mediated through some vendor-provided cloud service that is tied in some way to the customer identity. Alternatively, a service configured to run on the customer premises, e.g., running on a virtual machine or virtual container, could be used to broker the key exchange necessary for secure communications between the replicating storage systems (716, 718, 720). The reader will understand that a pod containing more than two storage systems may require communication links between most or all of the individual storage systems. In the example shown in FIG. 6, three data communication links (716, 718, 720) are shown, although additional data communication links may be present in other embodiments.
[0185]
[0243] The reader will understand that communication between the storage systems (714, 724, 728) across which the dataset (712) is synchronously replicated serves several purposes. For example, one purpose is to deliver data from one storage system (714, 724, 728) to another storage system (714, 724, 728) as part of an I / O operation. For example, processing a write typically requires delivering the write content and some description of the write to any paired storage system for a pod. Another purpose served by data communication between the storage systems (714, 724, 728) may be to communicate configuration changes and analytical data for processing the creation, expansion, deletion, or renaming of volumes, files, object buckets, etc. Another purpose served by data communication between the storage systems (714, 724, 728) may be to facilitate communications involved in detecting and handling storage system and interconnect failures. This type of communication can be time-critical and may need to be prioritized to ensure it does not get stuck behind long network queuing delays when large bursts of write traffic are suddenly dumped on the data center interconnect.
[0186]
[0244] The reader will further understand that different types of communications may use the same or different connections, and may use the same or different networks, in various combinations. Furthermore, some communications may be encrypted and protected, while other communications may not be encrypted. In some cases, a data communications link could be used to forward I / O requests from one storage system to another (either directly as the request itself or as a logical description of the operation the I / O request represents). This could be used, for example, when one storage system has up-to-date, in-sync content for a pod and another storage system does not currently have up-to-date, in-sync content for the pod. In such a case, as long as the data communications link is running, a request can be forwarded from a storage system that is not up-to-date and in-sync to a storage system that is up-to-date and in-sync.
[0187]
[0245] The example method shown in Figure 7 also includes exchanging (706) timing information (710, 722, 726) between the plurality of storage systems (714, 724, 728) for at least one of the plurality of storage systems (714, 724, 728). In the example method shown in Figure 6, the timing information (710, 722, 726) for a particular storage system (714, 724, 728) may be embodied, for example, as a value of a clock within the storage system (714, 724, 728). In an alternative embodiment, the timing information (710, 722, 726) for a particular storage system (714, 724, 728) may be embodied as a value that serves as a proxy for the clock value. The value that serves as a proxy for the clock value may be included in a token exchanged between the storage systems. For example, a value that acts as a proxy for a clock value may be implemented, such as a sequence number that a particular storage system (714, 724, 728) or storage controller may internally record as having been sent at a particular time. In such an example, if a token (e.g., a sequence number) is also received, the associated clock value may be found and utilized as a criterion for determining whether a valid lease is still in effect. In the example method shown in FIG. 6 , exchanging (706) timing information (710, 722, 726) among the plurality of storage systems (714, 724, 728) for at least one of the plurality of storage systems (714, 724, 728) may be performed, for example, by each storage system (714, 724, 728) sending timing information to each other storage system (714, 724, 728) in the pod periodically, on-demand, within a predetermined amount of time after lease establishment, within a predetermined amount of time before the lease is set to expire, as part of an attempt to initiate or re-establish a synchronous replication relationship, or in some other manner.
[0188]
[0246] The example method shown in Figure 7 also includes establishing (708) a synchronous replication lease according to the timing information (710, 722, 726) for at least one of the plurality of storage systems (714, 724, 728), where the synchronous replication lease identifies a period of time during which the synchronous replication relationship is valid. In the example method shown in Figure 7, the synchronous replication relationship is formed as a collection of storage systems (714, 724, 728) that replicate some data (712) between these largely independent stores, with each storage system (714, 724, 728) having its own copy and its own separate internal management of relevant data structures for defining storage objects, mapping objects to physical storage, for deduplication, defining mapping of content to snapshots, etc. A synchronous replication relationship may be specific to a particular data set, such that a particular storage system (714, 724, 728) may be associated with more than one synchronous replication relationship, each differentiated by the data set being described and may further comprise a different set of additional member storage systems.
[0189]
[0247] In the example method shown in FIG. 7 , a synchronous replication lease may be established (708) depending on timing information (710, 722, 726) for at least one of a plurality of storage systems (714, 724, 728) in a variety of different ways. In one embodiment, the storage system may establish (708) a synchronous replication lease by utilizing the timing information (710, 722, 726) for each of the plurality of storage systems (714, 724, 728) to adjust the clock. In such an example, once the clock is adjusted for each of the storage systems (714, 724, 728), the storage system may establish (708) a synchronous replication lease that extends for a predetermined period of time beyond the adjusted clock value. For example, if the clock for each storage system (714, 724, 728) is adjusted to a value of X, the storage systems (714, 724, 728) may each be configured to establish a synchronous replication lease that is valid for up to X+2 seconds.
[0190]
[0248] In an alternative embodiment, the need to adjust clocks between storage systems (714, 724, 728) may be avoided while still achieving timing guarantees. In such an embodiment, the storage controllers in each storage system (714, 724, 728) may have a local monotonically increasing clock. A synchronous replication lease may be established between storage controllers (e.g., a primary controller of one storage system communicating with a primary controller of a paired storage system) by each controller sending its clock value to the other storage controller along with the last clock value it received from the other storage controller (708). When a particular controller also receives its clock value from another controller, the particular controller adds any agreed-upon lease interval to the received clock value and uses it to establish its local synchronous replication lease (708). In this manner, the synchronous replication lease may be calculated according to the value of the local clock received from the other storage system.
[0191]
[0249] Consider an example in which a storage controller of a first storage system (714) is in communication with a storage controller of a second storage system (724). In such an example, assume that the value of the monotonically increasing clock for the storage controller of the first storage system (714) is 1000 milliseconds. Further assume that the storage controller of the first storage system (714) sends a message to the storage controller of the second storage system (724) indicating that the value of its clock was 1000 milliseconds when the message was generated. In such an example, assume that 500 milliseconds after the storage controller of the first storage system (714) sends a message to the storage controller of the second storage system (724) indicating that its clock value was 1000 milliseconds when the message was generated, the storage controller of the first storage system (714) receives a message from the storage controller of the second storage system (724) indicating that 1) the value of the monotonically increasing clock of the storage controller of the second storage system (724) was 5000 milliseconds when the message was generated, and 2) the last value of the monotonically increasing clock of the storage controller of the first storage system (714) received by the second storage system (724) was 1000 milliseconds. In such an example, if the agreed-upon lease interval is 2000 milliseconds, the first storage system (714) establishes (708) a synchronous replication lease that is valid until the monotonically increasing clock for the storage controller of the first storage system (714) reaches a value of 3000 milliseconds.If the storage controller of the first storage system (714) does not receive a message from the storage controller of the second storage system (724) containing an updated value of the monotonically increasing clock for the storage controller of the first storage system (714) by the time the monotonically increasing clock for the storage controller of the first storage system (714) reaches a value of 3000 milliseconds, the first storage system (714) may treat the synchronous replication lease as expired and take various actions, which are described in more detail below. The reader will understand that the storage controllers in the remaining storage systems (724, 728) in the pod will react similarly and perform similar tracking and updating of the synchronous replication lease. Essentially, the receiving controller may be guaranteed that the network and its paired controller were running somewhere during that time interval, and that its paired controller received the message it sent somewhere during that time interval. Without any adjustment of the clock, the receiving controller cannot know exactly where in that time interval the network and paired controller were performing, and cannot truly know if there were any queuing delays in sending its clock value or in receiving it back.
[0192]
[0250] In a pod consisting of two storage systems, each with a simple primary controller, where the primary controllers exchange clocks as part of their cluster communication, each primary controller can use an activity lease to place limits on when it does not know with certainty what its paired controller was doing. At the point where the primary controller becomes uncertain (when the controller's active lease on the connection has expired), the primary controller can send a message indicating that it is uncertain and that a properly synchronized connection must be re-established before the active lease can be resumed. If the network is functioning in one direction but not properly in the other, these messages may be received and no response may be received. This may be the first indication by the running paired controller that the connection is not performing normally, since its own active lease may not have yet expired due to a different combination of lost messages and queuing delays. As a result, if such a message is received, it should consider its own active lease to have expired as well, and it should begin sending its own messages, attempting to coordinate synchronization of the connection and resumption of the active lease. Until that happens and a new set of clock exchanges is successful, no controller can consider its active lease valid.
[0193]
[0251] In this model, the controller may wait lease interval seconds after it begins sending a re-establishment message, and if it receives no response, it may be assured that either the paired controller is down or the paired controller's own lease for the connection has expired. To handle small amounts of clock drift, the controller may wait slightly longer than the lease interval (i.e., re-establish the lease). When the controller receives the re-establishment message, it will immediately consider the re-establishment lease to have expired rather than waiting (because it knows that sending the controller's active lease has expired), but it often makes sense to try additional messaging before giving up if the message loss was a temporary condition caused, for example, by a congested network switch.
[0194]
[0252] In an alternative embodiment, in addition to establishing a synchronous replication lease, a cluster membership lease may also be established upon receipt of a clock value from a paired storage system or upon receipt of a clock exchanged with a paired storage system. In such an example, each storage system may have its own synchronous replication lease with every paired storage system and its own cluster membership lease. Expiration of a synchronous replication lease with any pair may result in processing outages. However, cluster membership cannot be recalculated until cluster membership leases with all pairs have expired. Therefore, the duration of a cluster membership lease should be set based on the interaction of messages and clock values to ensure that a cluster membership lease with a pair does not expire until after the pair's synchronous replication link for that link has expired. The reader will understand that a cluster membership lease may be established by each storage system in a pod and may be associated with a communication link between any two storage systems that are members of the pod. Furthermore, a cluster membership lease may extend after the expiration of a synchronous replication lease for a period at least as long as the period for the expiration of the synchronous replication lease. The cluster membership lease may be extended upon receipt of a clock value received from a paired storage system as part of a clock exchange, and the cluster membership lease duration from the current clock value may be at least as long as the duration established for the last synchronous replication lease extension based on the exchanged clock values. In additional embodiments, additional cluster membership information may be exchanged over the connection, including when the session was initially negotiated. The reader will understand that in embodiments utilizing cluster membership leases, each storage system (or storage controller) may have its own value for the cluster membership lease.Such leases should not expire until all synchronous replication leases across all pod members are guaranteed to have expired, given that cluster lease expiration allows new membership to be established, e.g., through mediator contention, and synchronous replication lease expiration forces a pause in the processing of new requests. In such instances, quiescing must be guaranteed to occur everywhere before cluster membership actions can be taken.
[0195]
[0253] The reader will now be directed to identifying (702) a plurality of storage systems (714, 724, 728) across which the dataset (712) is synchronously replicated, with only one of the storage systems (714) being a data set (712); configuring (704) one or more data communication links (716, 718, 720) between each of the plurality of storage systems (714, 724, 728) used to synchronously replicate the dataset (712); While the steps are illustrated as exchanging (706) timing information (710, 722, 726) for at least one of the plurality of storage systems (714, 724, 728) between the plurality of storage systems (714, 724, 728) and establishing (708) a synchronous replication lease in accordance with the timing information (710, 722, 726) for at least one of the plurality of storage systems (714, 724, 728), it will be understood that the remaining storage systems (724, 728) may also perform such steps. In practice, establishing a synchronous replication relationship between two or more storage systems (714, 724, 728) may require cooperation and interaction between two or more storage systems (714, 724, 728), so that all three storage systems (714, 724, 728) may perform one or more of the above-described steps simultaneously.
[0196]
[0254] For further explanation, Figure 8 sets forth a flowchart illustrating an additional example method for establishing a synchronous replication relationship between two or more storage systems (714, 724, 728) according to some embodiments of the present disclosure. The example method illustrated in Figure 8 includes identifying (702), for a data set (712), a plurality of storage systems (714, 724, 728) across which the data set (712) is to be synchronously replicated; configuring (704) one or more data communication links (716, 718, 720) between each of the plurality of storage systems (714, 724, 728) used to synchronously replicate the data set (712); and configuring (705) one or more data communication links (716, 718, 720) between the plurality of storage systems (714, 724, 728) to synchronously replicate the data set (712). The example method shown in FIG. 8 is similar to the example method shown in FIG. 46 because it also includes exchanging (706) timing information (710, 722, 726) for at least one of the storage systems (714, 724, 728) and establishing a synchronous replication lease in accordance with the timing information (710, 722, 726) for at least one of the plurality of storage systems (714, 724, 728), where the synchronous replication lease identifies a period of time during which the synchronous replication relationship is valid.
[0197]
[0255] In the example method shown in Figure 8, establishing (708) a synchronous replication lease in accordance with timing information (710, 722, 726) for at least one of the plurality of storage systems (714, 724, 728) may include coordinating (802) clocks between the plurality of storage systems (714, 724, 728). In the example method shown in Figure 8, coordinating (802) clocks between the plurality of storage systems (714, 724, 728) may be performed, for example, by exchanging one or more messages sent between the storage systems (714, 724, 728). The one or more messages sent between the storage systems (714, 724, 728) may include information such as the clock value of the storage system whose clock value is used by all other storage systems, instructions to set the clock value for all storage systems to a predetermined value, and a confirmation message from the storage system that its clock value has been updated. In such examples, the storage systems (714, 724, 728) may be configured such that a clock value for a particular storage system (e.g., a leader storage system) should be used by all other storage systems, clock values from all of the storage systems that meet some particular criteria (e.g., highest clock value) should be used by all other storage systems, etc. In such examples, some predetermined amount of time may be added to a clock value received from another storage system to account for transmission time associated with the exchange of messages.
[0198]
[0256] In the example method shown in Figure 8, establishing (708) a synchronous replication lease in accordance with timing information (710, 722, 726) for at least one of the plurality of storage systems (714, 724, 728) may include exchanging (804) uncoordinated clocks among the plurality of storage systems (714, 724, 728). Exchanging (804) uncoordinated clocks among the plurality of storage systems (714, 724, 728) may be performed, for example, by the storage controllers of each storage system (714, 724, 728) exchanging values of local monotonically increasing clocks, as described in more detail above. In such an example, each storage system (714, 724, 728) may utilize agreed-upon synchronous replication lease intervals and messaging received from the other storage systems (714, 724, 728) to establish (708) a synchronous replication lease.
[0199]
[0257] 8 also includes postponing (806) processing of I / O requests received after a synchronous replication lease has expired. I / O requests received by any of the storage systems after a synchronous replication lease has expired may be postponed for a predetermined amount of time sufficient to attempt to re-establish the synchronous replication relationship, such as until a new synchronous replication lease is established. In such an example, the storage system may postpone processing of the I / O request (806) by failing with some type of "busy" or temporary failure indication, or in some other manner.
[0200]
[0258] For further explanation, Figure 9 sets forth a flowchart illustrating an additional example method for establishing a synchronous replication relationship between two or more storage systems (714, 724, 728) in accordance with some embodiments of the present disclosure. The example method shown in Figure 9 includes identifying (702), for a data set (712), a plurality of storage systems (714, 724, 728) across which the data set (712) is to be synchronously replicated; configuring (704) one or more data communication links (716a, 716b, 718a, 718b, 720a, 720b) between each of the plurality of storage systems (714, 724, 728) used to synchronously replicate the data set (712); and configuring (705) one or more data communication links (716a, 716b, 718a, 718b, 720a, 720b) between the plurality of storage systems (714, 724, 728) used to synchronously replicate the data set (712). 9 is similar to the example method shown in FIG. 46 because it also includes exchanging (706) timing information (710, 722, 726) for at least one of the plurality of storage systems (714, 724, 728) and establishing (708) a synchronous replication lease in accordance with the timing information (710, 722, 726) for at least one of the plurality of storage systems (714, 724, 728), where the synchronous replication lease identifies a period of time during which the synchronous replication relationship is valid.
[0201]
[0259] 9, configuring (704) one or more data communication links (716a, 716b, 718a, 718b, 720a, 720b) between each of the plurality of storage systems (714, 724, 728) used to synchronously replicate the dataset (712) may include configuring (902) a data communication link (716a, 716b, 718a, 718b, 720a, 720b) between each of the plurality of storage systems (714, 724, 728) used to synchronously replicate the dataset (712) for each of a plurality of data communication types. In the example method shown in FIG. 9, each storage system may be configured to generate a plurality of data communication types that the storage system sends to other storage systems in the pod. For example, a storage system may generate a first type of data communication that includes data that is part of an I / O operation (e.g., data written to the storage system as part of a write request issued by a host), a second type of data communication that includes a configuration change (e.g., information generated in response to creating, expanding, deleting, or renaming a volume), and a third type of data communication that includes communications related to detecting and handling storage system and interconnect failures. In such examples, the data communication type may be determined, for example, based on which software module initiated the message, based on which hardware component initiated the message, based on the type of event that initiated the message, and in other manners. In the example method shown in FIG. 9, configuring (902) a data communication link (716a, 716b, 718a, 718b, 720a, 720b) between each of the plurality of storage systems (714, 724, 728) for each of the plurality of data communication types may be performed, for example, by configuring the storage systems to use a separate interconnect for each of the plurality of data communication types, by configuring the storage systems to use a separate network for each of the plurality of data communication types, or in other ways.
[0202]
[0260] 9 also includes detecting 904 that the synchronous replication lease has expired. In the example method shown in FIG. 9, detecting 904 that the synchronous replication lease has expired may be performed, for example, by a particular storage system comparing its current clock value to the period for which the lease was valid. The storage systems (714, 724, 728) are configured to adjust their clocks to set the value of the clock within each storage system (714, 724, 728) to a value of 5000 milliseconds, and each storage system (714, 724, 728) establishes 708 a synchronous replication lease that extends beyond that clock value for a lease interval of 2000 milliseconds, whereby the synchronous replication lease for each storage system (714, 724, 728) has expired when the clock within the particular storage system (714, 724, 728) reaches a value greater than 7000 milliseconds. In such an example, detecting that a synchronous replication lease has expired (904) may be performed by determining that a clock within a particular storage system (714, 724, 728) has reached a value of 7001 milliseconds or greater.
[0203]
[0261] The reader will understand that the occurrence of other events may also cause each storage system (714, 724, 728) to immediately treat the synchronous replication lease as expired. For example, a storage system (714, 724, 728) may immediately treat the synchronous replication lease as expired if it detects a communication failure between the storage system (714, 724, 728) and another storage system (714, 724, 728) in the pod, a storage system (714, 724, 728) may immediately treat the synchronous replication lease as expired if it receives a lease re-establishment message from another storage system (714, 724, 728) in the pod, a storage system (714, 724, 728) may immediately treat the synchronous replication lease as expired if it detects that another storage system (714, 724, 728) in the pod has failed, and so on. In such an example, the occurrence of any of the events described in the preceding paragraphs may cause the storage system to detect (904) that the synchronous replication lease has expired.
[0204]
[0262] The example method shown in Figure 9 also includes reestablishing 906 the synchronous replication relationship. In the example method shown in Figure 9, reestablishing 906 the synchronous replication relationship may be performed, for example, through the use of one or more reestablishment messages. Such reestablishment messages may include, for example, an identification of the pods with which the synchronous replication relationship is to be reestablished, information needed to configure one or more data communication links, updated timing information, etc. In this manner, the storage systems (714, 724, 728) may each identify (702) for a data set (712) a plurality of storage systems (714, 724, 728) across which the data set (712) is to be synchronously replicated; configure (704) one or more data communication links (716a, 716b, 718a, 718b, 720a, 720b) between each of the plurality of storage systems (714, 724, 728) used to synchronously replicate the data set (712); and configure (704) one or more data communication links (716a, 716b, 718a, 718b, 720a, 720b) between the plurality of storage systems (714, 724, 728). and establishing (708) a synchronous replication lease in accordance with the timing information (710, 722, 726) for at least one of the plurality of storage systems (714, 724, 728), wherein the synchronous replication lease identifies a period of time during which the synchronous replication relationship is valid.
[0205]
[0263] In the example method shown in FIG. 9 , expiration of a synchronous replication lease may be followed by some set of events, such as a re-establishment message followed by a new activity event or some other action. Data, configuration, or other communications may be in transition while the synchronous replication lease expires and is re-established. In fact, communications may not be received until, for example, after a new synchronous replication lease is established. In such cases, communications may have been sent based on one understanding of the state of a pod, cluster, or network link, and may now be received by a storage system (714, 724, 728) that has a different understanding of one aspect or another of that state. Thus, generally, there should be some means of ensuring that communications received are discarded if they were sent before some set of clusters or link states change. One way to ensure that communications sent before some convergence of the cluster or received communications are discarded if link conditions change is to establish some session identifier (e.g., a number) associated with establishing or re-establishing the link with the working synchronous replication lease being extended. After cluster communications are re-established, the link gets a new session identifier. This identifier may be included with data messages, configuration messages, or other communication messages. Any message received with an incorrect session identifier is discarded or results in an error response indicating a session identifier mismatch.
[0206]
[0264] The reader will understand that how the storage systems (714, 724, 728) respond to the re-establishment of a synchronous replication lease may vary based on the different implementations of the storage systems and pods. In the case of a simple primary controller with two storage systems, a new request to perform an operation (read, write, file operation, object operation, management operation, etc.) on the storage system that is received after the controller's synchronous replication lease has expired may have its processing postponed, aborted, or fail with some kind of "retry later" error code. Thus, a running primary storage controller may be guaranteed not to process new requests if its paired storage controller's synchronous replication lease has expired, and it may be certain when its own re-established lease has expired. After the re-established lease has expired, it is safe for the controller to consider its paired controller offline and then consider corrective actions, including continuing storage processing without the paired controller. Exactly what those actions may be may vary based on a wide variety of considerations and implementation details.
[0207]
[0265] In the case of storage systems with primary and secondary controllers, the primary controller still running on a storage system may attempt to connect to the former second controller of the paired storage system, assuming that the former second controller of the paired storage system may take over. Alternatively, the primary controller still running on a storage system may wait a certain amount of time, perhaps a maximum secondary takeover time. If the secondary controller connects and establishes a new connection with a new synchronous replication lease within a reasonable amount of time, the pod may then recover itself to a consistent state (described below) and then proceed normally. If the paired secondary controller does not connect quickly enough, then the still-running primary controller may take additional action, such as, for example, considering the paired storage system to have failed and then attempting to determine whether it should continue operating without the paired storage system. The primary controller may instead maintain an active lease connection to a secondary controller on the paired storage system in the pod. In that case, expiration of the primary-to-primary re-established lease could instead result in the surviving primary using that connection to query the secondary for takeover, rather than having to first establish that connection. Two primary storage controllers could be running while the network is not functioning between them, but the network is functioning between one or the other primary controller and the paired secondary controller. In that case, high availability monitoring within the storage system might not by itself detect the condition that triggers a failover from the primary to the secondary controller. Responses to that condition could include triggering a primary-to-secondary failover anyway to simply resume synchronous replication, routing communication traffic from the primary through the secondary, or acting exactly as if communication had completely failed between the two storage systems, resulting in the same failure processing that occurred.
[0208]
[0266] When multiple controllers are active for a pod (including both dual active-active controller storage systems and scale-out storage systems), leases may still be maintained through individual controller cluster communication with any or all controllers in the paired storage systems. In this case, a stale synchronous replication lease may lead to a stall in new request processing for the pod across the storage system. The leasing model may be extended with an exchange of clocks and paired clock responses among all active controllers in the storage system, with additional exchange of their clocks with every paired controller in the paired storage system. If there is an operational path in which a particular local controller's clock is exchanged with any paired controller, the controllers can then use that path for independent synchronous replication leases and possibly for independent re-establishment of leases. In this case, local controllers in a storage system may further exchange clocks among each other for local leases between each other as well. This may already be built into the local storage system's high availability and monitoring mechanisms, but any timing related to the storage system's high availability mechanisms should be taken into account in the duration of the active and re-established leases, or any additional delays between the re-established lease expiration and the action taken to handle the interconnect failure.
[0209]
[0267] Alternatively, storage system-to-storage system cluster communication or leasing protocols alone may be assigned to one primary controller at a time in each multi-controller or scale-out system, at least for a particular pod. This service may migrate from controller to controller as a result of failure, or perhaps as a result of load imbalance. Alternatively, cluster communication or leasing protocols may run on a subset of controllers (e.g., two) to limit the complexity of analyzing clock exchange or failure conditions. Each local controller may need to exchange clocks among controllers that process storage system-to-storage system leases, and the time to respond after lease expiration must be adjusted accordingly, to account for potential cascading delays when individual controllers can be sure that they have achieved processing pauses. Connections that are not currently relied upon for leases related to processing pauses may still be monitored for alerts.
[0210]
[0268] The example method shown in Figure 9 also includes attempting to take over I / O processing for the dataset (908). In the example method shown in Figure 9, attempting to take over I / O processing for the dataset (712) (908) may be performed by a storage system (714, 724, 728) racing to a mediator, for example. If a particular storage system (714, 724, 728) successfully takes over I / O processing for the dataset (712), all accesses of the dataset (712) will be serviced by the particular storage system (714, 724, 728) until a synchronous replication relationship can be re-established, and any changes to the dataset (712) that occurred after the previous synchronous replication relationship expired may then be transferred and persisted on the other storage system (714, 724, 728). In such an example, an attempt 908 to take over I / O processing for the dataset 712 may occur only after the expiration of some period of time after the synchronous replication lease expires. For example, an attempt to determine how to begin after a link failure (including one or more of the storage systems attempting to take over I / O processing for the dataset) may not begin until a period of time after the synchronous replication lease expires, i.e., at least as long as the maximum lease time resulting from a clock exchange.
[0211]
[0269] The reader will understand that while in many of the examples shown above, only one of the storage systems (714) is shown performing the steps described above, establishing a synchronous replication relationship between two or more storage systems may require cooperation and interaction between two or more storage systems, and therefore in practice all of the storage systems (714, 724, 728) in a pod (or in a pod being formed) may simultaneously perform one or more of the steps described above.
[0212]
[0270] For further explanation, Figure 10 sets forth a flowchart illustrating an additional example method for establishing a synchronous replication relationship between two or more storage systems (1024, 1046) according to some embodiments of the present disclosure. While the example method illustrated in Figure 10 illustrates an embodiment in which the dataset (1022) is synchronously replicated across only two storage systems (1024, 1046), the example illustrated in Figure 10 may be extended to embodiments in which the dataset (1022) is synchronously replicated across additional storage systems that may perform steps similar to those performed by the two illustrated storage systems (1024, 1046).
[0213]
[0271] The example method shown in Figure 10 includes configuring (1002), by the storage system (1024), one or more data communication links (1052) between the storage system (1024, 1046) and a second storage system (1046). In the example method shown in Figure 10, the storage system (1024) may configure (1002) the one or more data communication links (1052) between the storage system (1024) and the second storage system (1046), for example, by identifying a port defined over a data communication network used to exchange data communication with the second storage system (1046), by identifying a point-to-point data communication link used to exchange data communication with the second storage system (1046), or in a variety of ways. If secure communications are required, some form of key exchange may be required, or the communications may occur or be initiated through some service, such as SSH (Secure SHell), SSL, or some other service or protocol built around public key or Diffie-Hellman key sharing, or a suitable alternative. Secure communications may also be mediated through some vendor-provided cloud service that is tied in some way to a customer identity. Alternatively, a service configured to run on the customer premises, e.g., running in a virtual machine or virtual container, could be used to broker the key exchange necessary for secure communications between the replicating storage systems (1024, 1046). The reader will understand that a pod containing more than two storage systems may require communication links between most or all of the individual storage systems. In the example method shown in FIG. 10, the second storage system (1046) may similarly configure (1026) one or more data communication links (1052) between the storage system (1024) and the second storage system (1046).
[0214]
[0272] 10 also includes transmitting (1004) timing information (1048) for the storage system (1024) from the storage system (1024) to the second storage system (1046). The timing information (1048) for the storage system (1024) may be embodied, for example, as the value of a clock within the storage system (1024), as a representation of a clock value (e.g., a sequence number that the storage system (1024) can first record), as the most recently received value of the clock within the second storage system (1046), etc. In the example method shown in Figure 10, the storage system (1024) may send (1004) timing information (1048) for the storage system (1024) to the second storage system (1046), for example, via one or more messages transmitted from the storage system (1024) to the second storage system (1046) over a data communication link (1052) between the two storage systems (1024, 1046). In the example shown in Figure 10, the second storage system (1046) may similarly send (1030) timing information (1050) for the second storage system (1046) from the second storage system (1046) to the storage system (1024).
[0215]
[0273] In the example method shown in Figure 10, transmitting (1004) timing information (1048) for the storage system (1024) from the storage system (1024) to the second storage system (1046) may include transmitting (1006) the value of the clock of the storage system (1024). In the example shown in Figure 10, the storage system (1024) may transmit (1006) the value of the clock of the storage system (1024) to the second storage system (1046) as part of an effort to coordinate the clocks between the storage systems (1024, 1046). In such an example, the storage system (1024) may include a local monotonically increasing clock whose value is transmitted (1006) via one or more messages sent to the second storage system (1046) over a data communications link (1052) between the two storage systems (1024, 1046). In the example method shown in FIG. 10, transmitting (1030) timing information (1050) for the second storage system (1046) from the second storage system (1046) to the storage system (1024) may also include transmitting (1032) the value of a clock of the second storage system (1046).
[0216]
[0274] In the example method shown in Figure 10, transmitting (1004) timing information (1048) for the storage system (1024) from the storage system (1024) to the second storage system (1046) may also include transmitting (1008) the most recently received value of the clock of the second storage system (1046). In the example method shown in Figure 10, transmitting (1008) the most recently received value of the clock of the second storage system (1046) may be performed, for example, as part of an effort to eliminate the need to adjust clocks between the storage systems (1024, 1046) while still achieving timing guarantees. In such an embodiment, each storage system (1024, 1046) may have a local monotonically increasing clock. A synchronous replication lease may be established between storage systems (1024, 1048) by each storage system (1024, 1048) sending its clock value to the other storage system (1024, 1048) along with the last clock value it received from the other storage system (1024, 1048). When a particular storage system (1024, 1048) also receives a clock value from another storage system (1024, 1048), the particular storage system (1024, 1048) may add any agreed-upon lease interval to the received clock value and use it to establish a synchronous replication lease. In the example method shown in FIG. 10, transmitting (1030) timing information (1050) for the second storage system (1046) from the second storage system (1046) to the storage system (1024) may also include transmitting (1034) the most recently received value of the clock of the storage system (1024).
[0217]
[0275] The example method shown in Figure 10 also includes receiving (1010) by the storage system (1024) from the second storage system (1046) timing information (1050) for the second storage system (1046). In the example method shown in Figure 10, the storage system (1024) may receive (1010) the timing information (1050) for the second storage system (1046) from the second storage system (1046) via one or more messages transmitted from the second storage system (1046) over a data communication link (1052) between the two storage arrays (1024, 1046). In the example shown in Figure 10, the second storage system (1046) may similarly receive (1028) the timing information for the storage system (1024) from the storage system (1024).
[0218]
[0276] The example method shown in Figure 10 also includes setting (1012) a clock value at the storage system (1024) according to the timing information (1050) for the second storage system (1046). In the example shown in Figure 10, setting (1012) a clock value at the storage system (1024) according to the timing information (1050) for the second storage system (1046) may be performed, for example, as part of an effort to coordinate clocks between the two storage systems (1024, 1046). In such an example, the two storage systems (1024, 1046) may be configured, for example, to set their respective clock values to a value that is some predetermined amount higher than the highest clock value between the storage systems (1024, 1046), to set their respective clock values to a value equal to the highest clock value between the pair of storage systems (1024, 1046), to set their respective clock values to a value generated by applying some function to the respective clock values of each storage system (1024, 1046), or in some other manner. In the example method shown in Figure 10, the second storage system (1046) may similarly set (1036) its clock value in accordance with the timing information (1048) for storage system (1024).
[0219]
[0277] The example method shown in Figure 10 also includes establishing (1014) a synchronous replication lease. In the example method shown in Figure 10, establishing (1014) a synchronous replication lease may be performed, for example, by establishing a synchronous replication lease that extends for some predetermined lease interval beyond the equivalent clock values between the two storage systems (1024, 1046), by establishing a synchronous replication lease that extends for some predetermined lease interval beyond the unadjusted clock value associated with one of the storage systems (1024, 1046), or in some other manner. In the example method shown in Figure 19, the second storage system (1046) may similarly set (1036) its clock value in accordance with the timing information (1048) for storage system (1024).
[0220]
[0278] The example method shown in Figure 10 also includes detecting (1016) by the storage system (1024) that the synchronous replication lease has expired. In the example method shown in Figure 10, detecting (1016) that the synchronous replication lease has expired may be performed, for example, by the storage system (1024) comparing the current clock value to the period for which the lease was valid. Consider an example in which the storage systems (1024, 1046) are configured to adjust the value of the clock within each storage system (1024, 1046) to a value of 5000 milliseconds, and each storage system (1024, 1046) establishes (1038) an extended synchronous replication lease for a lease interval of 2000 milliseconds beyond that clock value, thereby causing the synchronous replication lease for each storage system (1024, 1046) to expire when the clock within a particular storage system (1024, 1046) reaches a value greater than 10,000 milliseconds. In such an example, detecting that the synchronous replication lease has expired (1006) may be performed by determining that a clock within the storage system (1024) has reached a value of 10001 milliseconds or greater. In the example method shown in Figure 10, the second storage system (1046) may similarly detect that the synchronous replication lease has expired (1040).
[0221]
[0279] The example method illustrated in Figure 10 also includes attempting (1020) to take over I / O processing for the dataset (1022) by the storage system (1024). In the example method illustrated in Figure 10, attempting (1020) to take over I / O processing for the dataset (1022) may be performed by the storage system (1024) dispatching to a mediator, for example. If the storage system (1024) successfully takes over I / O processing for the dataset (1022), all accesses of the dataset (1022) will be serviced by the storage system (1024) until the synchronous replication relationship is re-established, and any changes to the dataset (1022) that occurred after the previous synchronous replication relationship expired may be transferred to and persisted in the second storage system (1046). In the example method shown in FIG. 10, the second storage system (1046) may also attempt to take over I / O processing (1044) for the dataset (1022).
[0222]
[0280] The example method shown in FIG. 10 also includes, by the storage system (1024), attempting (1018) to reestablish the synchronous replication relationship. In the example method shown in FIG. 10, attempting (1018) to reestablish the synchronous replication relationship may be performed, for example, by using one or more reestablishment messages. Such reestablishment messages may include, for example, an identification of the pod with which the synchronous replication relationship is to be reestablished, information needed to configure one or more data communication links, updated timing information, etc. In this manner, the storage system (1024) may reestablish the synchronous replication relationship substantially in the same manner as the synchronous replication relationship was originally created. In the example shown in FIG. 10, the second storage system (1046) may similarly attempt (1042) to reestablish the synchronous replication relationship.
[0223]
[0281] For additional explanation, Figure 11 sets forth a flow chart illustrating an example method for servicing I / O operations directed to a dataset (1142) that is synchronized across multiple storage systems (1138, 1140) in accordance with some embodiments of the present disclosure. While not shown in great detail, the storage systems (1138, 1140) illustrated in Figure 11 may be similar to the storage systems described above with respect to Figures 1A-1D, 2A-2G, 3A-3B, or any combination thereof. In practice, the storage system illustrated in Figure 11 may include the same components as the storage systems described above, fewer components, or additional components.
[0224]
[0282] The data set (1142) shown in Figure 11 may be embodied, for example, as the contents of a particular volume, the contents of a particular share of a volume, or any other collection of one or more data elements. The data set (1142) may be synchronized across multiple storage systems (1138, 1140) such that each storage system (1138, 1140) maintains a local copy of the data set. In the example described herein, such data set (1142) is synchronously replicated across storage systems (1138, 1140) such that the data set (1142) can be accessed through any of the storage systems (1138, 1140) that have performance characteristics such that any one storage system in the cluster operates substantially less optimally than any other storage system in the cluster, at least as long as the cluster and the particular storage system being accessed are nominally performing. In such a system, modifications to a dataset (1142) should be made to copies of the dataset resident on each storage system (1138, 1140) so that accessing the dataset (1142) on any storage system (1138, 1140) produces consistent results. For example, a write request to a dataset must be serviced on all storage systems (1138, 1140), or on none of the storage systems (1138, 1140) that were nominally running at the beginning of the write and remained nominally running until the write was completed. Similarly, some groups of operations (e.g., two write operations directed to the same location in a dataset) must be performed in the same order, or other steps must be taken on all storage systems (1138, 1140), so that the dataset is ultimately identical on all storage systems (1138, 1140).Modifications to the dataset (1142) do not need to occur at exactly the same time, although some actions (e.g., issuing an acknowledgment that a write request is directed to the dataset and allowing read access to locations in the dataset targeted by write requests that have not yet completed on both storage systems) may be deferred until the copies of the dataset on each storage system (1138, 1140) have been modified.
[0225]
[0283] In the example method shown in Figure 11, designating one storage system (1140) as a "leader" and another storage system (1138) as a "follower" may refer to the respective relationship of each storage system for purposes of synchronously replicating a particular data set across the storage systems. In such an example, and as described in more detail below, the leader storage system (1140) may be responsible for performing some processing of incoming I / O operations, passing relevant information to the follower storage system (1138), or performing other tasks not required of the follower storage system (1140). The leader storage system (1140) may be responsible for performing tasks not required of the follower storage system (1138) for all incoming I / O operations, or alternatively, the leader-follower relationship may be specific to only a subset of the I / O operations received by either storage system. For example, a leader-follower relationship may be specific to I / O operations directed to a first volume, a first group of volumes, a first group of logical addresses, a first group of physical addresses, or some other logical or physical delineator. In this manner, a first storage system may act as a leader storage system for I / O operations directed to a first set of volumes (or other delineators), while a second storage system may act as a leader storage system for I / O operations directed to a second set of volumes (or other delineators). The example method shown in FIG. 11 illustrates an embodiment in which synchronizing multiple storage systems (1138, 1140) occurs in response to receiving a request (1104) to modify a data set (1142) by a leader storage system (1140), although synchronizing multiple storage systems (1138, 1140) may be performed in response to receiving a request (1104) to modify a data set (1142) by a follower storage system (1138), as described in more detail below.
[0226]
[0284] 11 includes receiving 1106 a request 1104 by a leader storage system 1140 to modify a dataset 1142. The request 1104 to modify the dataset 1142 may be implemented, for example, as a request to write data to a location in the storage system 1140 that contains data included in the dataset 1142, as a request to write data to a volume that contains data included in the dataset 1142, as a request to take a snapshot of the dataset 1142, essentially as a virtual range copy, as an UNMAP operation that represents the deletion of some portion of the data in the dataset 1142, as a modifying transformation of the dataset 1142 (rather than a change to a portion of the data within the dataset), or as some other operation that causes a change to some portion of the data included in the dataset 1142. In the example method shown in FIG. 11, a request (1104) to modify a dataset (1142) is issued by a host (1102), which may be embodied, for example, as an application running on a virtual machine, as an application running on a computing device connected to the storage system (1140), or as some other entity configured to access the storage system (1140).
[0227]
[0285] 11 also includes generating 1108, by the leader storage system 1140, information 1110 describing modifications to the dataset 1142. The leader storage system 1140 may generate 1108 the information 1110 describing modifications to the dataset 1142 by determining the proper outcome of overlapping modifications (e.g., the proper outcome of two requests modifying the same storage location), for example, by determining ordering versus any other operations that are in progress, and computing any distributed state changes, for example, to common elements of metadata, across all members of a pod (e.g., all storage systems to which the dataset is synchronously replicated). The information 1110 describing modifications to the dataset 1142 may be implemented, for example, as system-level information used to describe I / O operations performed by the storage system. The leader storage system 1140 may generate 1108 information describing modifications to the dataset 1142 by processing the request 1104 to modify the dataset 1142 sufficiently to understand what must occur in order to service the request 1104 to modify the dataset 1142. For example, the leader storage system 1140 may determine whether some ordering of execution of the request 1104 to modify the dataset 1142 relative to other requests to modify the dataset 1142 is required, or some other step must be taken, as described in more detail below, to produce equivalent results at each storage system 1138, 1140.
[0228]
[0286] Consider an example in which a request (1104) to modify a dataset (1142) is implemented as a request to copy blocks from a first address range in the dataset (1142) to a second address range in the dataset (1142). In such an example, assume that three other write operations (Write A, Write B, and Write C) are directed to the first address range in the dataset (1142). In such an example, if the leader storage system (1140) services Write A and Write B (but does not service Write C) before copying blocks from the first address range in the dataset (1142) to the second address range in the dataset (1142), then the follower storage system (1138) must also service Write A and Write B (but does not service Write C) before copying blocks from the first address range in the dataset (1142) to the second address range in the dataset (1142) to produce consistent results. Thus, when the leader storage system (1140) generates (1108) information (1110) describing modifications to the data set (1142), in this example, the leader storage system (1140) could generate information (e.g., sequence numbers for Write A and Write B) that identifies other operations that must be completed before the follower storage system (1138) can process the request (1104) to modify the data set (1142).
[0229]
[0287] Consider an additional example in which two requests (e.g., Write A and Write B) are directed to overlapping portions of a data set (1142). In such an example, if the follower storage system (1138) services Write B and later services Write A, but the leader storage system (1140) services Write A and later services Write B, the data set (1142) will not be consistent across both storage systems (1138, 1140). Thus, when the leader storage system (1140) generates (1108) information (1110) describing modifications to the data set (1142), in this example, the leader storage system (1140) could generate information (e.g., sequence numbers for Write A and Write B) that identifies the order in which the requests should be executed. Alternatively, rather than generating information (1110) describing modifications to a dataset (1142) that requires intermediate operations from each storage system (1138, 1140), the leader storage system (1140) may generate information (1110) describing modifications to a dataset (1142) that includes information identifying the appropriate outcome of the two requests (1108). For example, if write B logically follows (and overlaps with) write A, the end result should be that dataset (1142) includes portions of write B that overlap with write A, rather than including portions of write A that overlap with write B. Such an outcome would be facilitated by merging the results in memory and writing the results of such a merge to dataset (1142), rather than strictly requiring a particular storage system (1138, 11410) to perform write A and then subsequently perform write B. The reader will understand that more subtle cases pertain to snapshots and virtual address range copies.
[0230]
[0288] The reader further understands that the correct outcome for any operation must be committed to a point where it is recoverable before receipt of the operation can be acknowledged. However, in some cases, multiple operations can be committed together, and in other cases, operations can be partially committed if recovery ensures correctness. For example, a snapshot could be committed locally with dependencies recorded on expected write A and write B, but A or B itself might not be committed. If receipt of the snapshot cannot be acknowledged and lost I / O cannot be recovered from another array, recovery could end up canceling the snapshot. Also, if write B overlaps with write A, the leader storage system may "instruct" B to be behind A, but A is actually discarded, and then the operation that writes A will simply wait for B. Write A, Write B, Write C, and Write D combined with the snapshot between A, B and C, D will be able to commit and / or acknowledge receipt of some or all parts together as long as recovery does not cause a snapshot inconsistency across the array and as long as the acknowledgment does not complete the later operation before the earlier operation has persisted to the point where it is guaranteed to be recoverable.
[0231]
[0289] 11 also includes transmitting 1112 information 1110 from the leader storage system 1140 to the follower storage system 1138 describing the modifications to the data set 1142. Transmitting 1112 information 1110 from the leader storage system 1140 to the follower storage system 1138 describing the modifications to the data set 1142 may be performed, for example, by the leader storage system 1140 sending one or more messages to the follower storage system 1138. The leader storage system 1140 may also send an I / O payload 1114 for the request 1104 to modify the data set 1142 in the same message or in one or more different messages. The I / O payload 1114 may be implemented as data written to storage in the follower storage system 1138, for example, when a request 1104 to modify a dataset 1142 is implemented as a request to write data to the dataset 1142. In such an example, because the request 1104 to modify the dataset 1142 was received 1106 by the leader storage system 1140, the follower storage system 1138 has not yet received the I / O payload 1114 associated with the request 1104 to modify the dataset 1142.In the example shown in FIG. 11 , information (1110) describing modifications to the dataset (1142) and an I / O payload (1142) associated with a request (1104) to modify the dataset (1142) may be transmitted (1112) from the leader storage system (1140) to the follower storage system (1138) via one or more data communications networks coupling the leader storage system (1140) to the follower storage system (1138), via one or more dedicated data communications links (e.g., a first link for transmitting the I / O payload and a second link for transmitting the information describing the modifications to the dataset) coupling the leader storage system (1140) to the follower storage system (1138), or via some other mechanism.
[0232]
[0290] 11 also includes receiving 1116, by the follower storage system 1138, information 1110 describing modifications to the data set 1142. The follower storage system 1138 may receive 1116 the information 1110 describing modifications to the data set 1142 and the I / O payload 1114 from the leader storage system 1140, for example, via one or more messages sent from the leader storage system 1140 to the follower storage system 1138. The one or more messages may be sent from the leader storage system 1140 to the follower storage system 1138 via one or more dedicated data communication links between the two storage systems 1138, 1140, by the leader storage system 1140 writing the messages to a predetermined storage location (e.g., a queue location) on the follower storage system 1138 using RDMA or a similar mechanism, or by other methods.
[0233]
[0291] In one embodiment, a follower storage system (1138) may receive (1116) information (1110) and I / O payload (1114) describing modifications to a dataset (1142) from a leader storage system (1140) using SCSI requests (write from sender to receiver, or read from receiver to sender) as the communication mechanism. In such an embodiment, a SCSI write request is used to encode the information intended to be sent (including whatever data and metadata) and may be delivered to a special pseudo-device, or over a specially configured SCSI network, or through any other agreed-upon addressing mechanism. Alternatively, the model may use a special device, a specially configured SCSI network, or other agreed-upon mechanism to issue a set of open SCSI read requests from the receiver to the sender. The encoded information, including the data and metadata, is delivered to the receiver in response to one or more of these open SCSI requests. Such a model may be implemented over a Fibre Channel SCSI network, often deployed as a “dark fibre” storage network between datacenters. Such a model also allows for multipathing from the host to the remote array and use of the same network lines for bulk array-to-array communication.
[0234]
[0292] The example method illustrated in Figure 11 also includes processing (1118), by the follower storage system (1138), the request (1104) to modify the dataset (1142). In the example method illustrated in Figure 11, the follower storage system (1138) may process (1118) the request (1104) to modify the dataset (1142) by modifying the contents of one or more storage devices (e.g., NVRAM devices, SSDs, HDDs) included in the follower storage system (1138) according to not only the I / O payload (1114) received from the leader storage system (1140) but also the information (1110) describing the modifications to the dataset (1142). Consider an example in which a request (1104) to modify a data set (1142) is implemented as a write operation directed to a volume included in the data set (1142), and the information (1110) describing the modification to the data set (1142) indicates that the write operation can only be performed after a previously issued write operation has been processed. In such an example, processing (1118) the request (1104) to modify the data set (1142) may be implemented by the follower storage system (1138) first verifying that the previously issued write operation has been processed at the follower storage system (1138), and then writing the I / O payload (1114) associated with the write operation to one or more storage devices included in the follower storage system (1138). In such an example, a request (1104) to modify a dataset (1142) may be considered completed and successfully processed, for example, when the I / O payload (1114) is committed to persistent storage in the follower storage system (1138).
[0235]
[0293] The example method illustrated in Figure 11 also includes acknowledging (1120) by the follower storage system (1138) to the leader storage system (1140) receipt of completion of the request (1104) to modify the data set (1142). In the example method illustrated in Figure 11, acknowledging (1120) by the follower storage system (1138) to the leader storage system (1140) receipt of completion of the request (1104) to modify the data set (1142) may be performed by the follower storage system (1138) sending an acknowledgement (1122) message to the leader storage system (1140). Such a message may include, for example, information identifying the particular request (1104) to modify the completed data set (1142), as well as any additional information useful in acknowledging (1120) receipt of completion of the request (1104) to modify the data set (1142) by the follower storage system (1138). In the example method shown in FIG. 11, acknowledging (1120) receipt of completion of the request (1104) to modify the data set (1142) to the leader storage system (1140) is indicated by the follower storage system (1138) issuing an acknowledgment (1122) message to the leader storage system (1140).
[0236]
[0294] The example method illustrated in Figure 11 also includes processing (1124), by the leader storage system (1140), the request (1104) to modify the dataset (1142). In the example method illustrated in Figure 11, the leader storage system (1140) may process (1124) the request (1104) to modify the dataset (1142) by modifying not only the I / O payload (1114) received as part of the request (1104) to modify the dataset (1142), but also the contents of one or more storage devices (e.g., NVRAM devices, SSDs, HDDs) included in the leader storage system (1140) according to information (1110) describing the modifications to the dataset (1142). Consider an example in which a request (1104) to modify a dataset (1142) is implemented as a write operation directed to a volume included in the dataset (1142), and the information (1110) describing the modification to the dataset (1142) indicates that the write operation can only be performed after a previously issued write operation has been processed. In such an example, processing (1124) the request (1104) to modify the dataset (1142) may be implemented by the leader storage system (1140) first verifying that the previously issued write operation has been processed by the leader storage system (1140), and then writing the I / O payload (1114) associated with the write operation to one or more storage devices included in the leader storage system (1140). In such an example, a request (1104) to modify a dataset (1142) may be considered completed and successfully processed, for example, when the I / O payload (1114) is committed to persistent storage in the reader storage system (1140).
[0237]
[0295] 11 also includes receiving 1126 an indication from the follower storage system 1138 that the follower storage system 1138 has processed 1104 the request to modify the data set 1136. In this example, the indication that the follower storage system 1138 has processed 1104 the request to modify the data set 1136 is embodied as an acknowledgment 1122 message sent from the follower storage system 1138 to the leader storage system 1140. The reader will understand that although many of the steps described above are shown and described as occurring in a particular order, no particular order is actually required. In practice, the follower storage system 1138 and the leader storage system 1140 are independent storage systems, and therefore each storage system may be performing some of the steps described above in parallel. For example, the follower storage system (1138) may receive (1116) information (1110) describing modifications to the dataset (1142), may process (1118) the request (1104) to modify the dataset (1142), or may acknowledge (1120) receipt of completion of the request (1104) to modify the dataset (1142) before the leader storage system (1140) processes (1124) the request (1104) to modify the dataset (1142). Alternatively, the leader storage system (1140) may have processed (1124) the request (1104) to modify the dataset (1142), processed (1118) the request (1104) to modify the dataset (1142), or acknowledged (1120) receipt of completion of the request (1104) to modify the dataset (1142) before the follower storage system (1138) received (1116) the information (1110) describing the modifications to the dataset (1142).
[0238]
[0296] The example method shown in Figure 11 also includes acknowledging (1134) receipt, by the leader storage system (1140), of completion of the request (1104) to modify the data set (1142). In the example method shown in Figure 11, acknowledging (1134) receipt of completion of the request (1104) to modify the data set (1142) may be performed using one or more acknowledgment (1136) messages sent from the leader storage system (1140) to the host (1102) or via some other suitable mechanism. In the example method shown in Figure 11, the leader storage system (1140) may determine (1128) whether the request (1104) to modify the data set (1142) has been processed (1118) by the follower storage system (1138) before acknowledging (1134) receipt of completion of the request (1104) to modify the data set (1142). The leader storage system (1140) may determine (1128) whether the request (1104) to modify the data set (1142) has been processed (1118) by the follower storage system (1138), for example, by determining whether the leader storage system (1140) has received an acknowledgment message or other message from the follower storage system (1138) indicating that the request (1104) to modify the data set (1142) has been processed (1118) by the follower storage system (1138). In such an example, if the leader storage system (1140) affirmatively determines (1130) that the request (1104) to modify the dataset (1142) has been processed (1118) by the follower storage system (1138) and also processed (1118) by the leader storage system (1138), the leader storage system (1140) may proceed by acknowledging (1134) receipt of completion of the request (1104) to modify the dataset (1142) to the host (1102) that initiated the request (1104) to modify the dataset (1142).However, if the leader storage system (1140) determines that the request (1104) to modify the dataset (1142) has not yet been processed (1118) (1132) by the follower storage system (1138) or has not yet been processed (1124) by the leader storage system (1138), the leader storage system (1140) may acknowledge (1134) receipt of completion of the request (1104) to modify the dataset (1142) only when the request (1104) to modify the dataset (1142) has been successfully processed on all storage systems (1138, 1140) across which the dataset (1142) is synchronously replicated, and therefore the leader storage system (1140) may not yet acknowledge (1134) receipt of completion of the request (1104) to modify the dataset (1142) from the host (1102) that initiated the request (1104) to modify the dataset (1142).
[0239]
[0297] The reader will note that in the example method shown in FIG. 11 , sending (1112) information describing modifications to a data set (1142) from a leader storage system (1140) to a follower storage system (1138) and acknowledging (1120) receipt of completion of a request (1104) to modify the data set (1142) by the follower storage system (1138) to the leader storage system (1140) may be accomplished using a single round-trip message. Single round-trip messaging may be used, for example, with Fibre Channel as the data interconnect. Typically, the SCSI protocol is used with Fibre Channel. Because some older replication technologies can be built to essentially replicate data as SCSI transactions over a Fibre Channel network, such interconnects are commonly set up between data centers. Furthermore, traditional Fibre Channel SCSI infrastructures have had less overhead and lower latency than Ethernet and TCP / IP-based networks. Additionally, when data centers use Fibre Channel and are internally connected to disconnect storage arrays, the Fibre Channel network may be extended to other data centers so that hosts in one data center can switch to accessing a storage array at a remote data center if the local storage array fails.
[0240]
[0298] SCSI could be used as a general communication mechanism, even though it is typically designed for use with block storage protocols for storing and retrieving data on block-oriented volumes (or for tape). For example, a SCSI READ or SCSI WRITE could be used to deliver or retrieve message data between storage controllers in paired storage systems. A typical implementation of a SCSI WRITE requires two message round-trips: the SCSI initiator sends a SCSI CDB describing the SCSI WRITE operation, the SCSI target receives the CDB, and the SCSI target sends a "ready to receive" message to the SCSI initiator. The SCSI initiator then sends the data to the SCSI target, and once the SCSI WRITE is complete, the SCSI target responds with a successful completion to the SCSI initiator. On the other hand, a SCSI READ request requires only one round-trip: the SCSI initiator sends a SCSI CDB describing the SCSI READ operation, the SCSI target receives the CDB, and responds with the data and then a successful completion. As a result, in terms of distance, a SCSI READ incurs half the distance-related latency as a SCSI WRITE. Therefore, it may be faster for a data communications receiver to receive a message using a SCSI READ request than for the message sender to send data using a SCSI WRITE request. Using a SCSI READ simply requires that the message sender act as a SCSI target and the message receiver act as a SCSI initiator. The message receiver may send several SCSI CDB READ requests to any message sender, and the message sender will respond to one of the pending CDB READ requests when message data is available. The SCSI subsystem may time out if a READ request is pending for too long (e.g., 10 seconds), so a READ request should be responded to within a few seconds, even if there is no message to send.
[0241]
[0299] SCSI tape requests support variable response data, which may be more flexible for returning variable-sized message data, as described in the SCSI Stream Commands standard of the T10 Technical Committee of the InterNational Committee on Information Technology Standards. The SCSI standard also supports an immediate mode for SCSI WRITE requests, which could allow for a single round-trip SCSI WRITE command. The reader will understand that many of the embodiments described below also utilize single-round-trip messaging.
[0242]
[0300] For additional explanation, FIG. 12 sets forth a flowchart illustrating an additional example method for servicing I / O operations directed to a dataset (1142) synchronized across multiple storage systems (1138, 1140, 1150) in accordance with some embodiments of the present disclosure. While not shown in great detail, the storage systems (1138, 1140, 1150) illustrated in FIG. 11 may be similar to the storage systems described above with respect to FIGS. 1A-1D, 2A-2G, 3A-3B, or any combination thereof. In practice, the storage system illustrated in FIG. 11 may include the same components as the storage systems described above, fewer components, or additional components. The example method illustrated in FIG. 12 is similar to the example method illustrated in FIG. 11. The example method shown in FIG. 12 also includes receiving (1106) a request to modify (1142) a data set (1142) by a leader storage system (1140), generating (1108) information (1110) by the storage system (1140) describing the modifications to the data set (1142), transmitting (1112) the information (1110) describing the modifications to the data set (1142) from the leader storage system (1140) to a follower storage system (1138), receiving (1116) the information (1110) describing the modifications to the data set (1142) by the follower storage system (1138), and transmitting (1116) the information (1110) describing the modifications to the data set (1142) to the follower storage system (1138). 11, as it also includes processing (1118) the request (1104) to modify the data set (1142) by the follower storage system (1138), acknowledging (1120) receipt of completion of the request (1104) to modify the data set (1142) by the leader storage system (1140), processing (1124) the request (1104) to modify the data set (1142) by the leader storage system (1140), and acknowledging (1134) receipt of completion of the request (1104) to modify the data set (1142) by the leader storage system (1140).
[0243]
[0301] However, the example method shown in FIG. 12 differs from the example method shown in FIG. 11 because it illustrates an embodiment in which a data set (1142) is synchronously replicated across three storage systems, one of which is a leader storage system (1140) and the remaining storage systems are follower storage systems (1138, 1150). In such an example, the additional follower storage system (1150) performs many of the same steps as the follower storage system (1138) shown in FIG. 11, such as receiving (1142) information (1110) from the leader storage system (1140) describing modifications to the dataset (1142), processing (1142) a request (1104) to modify the dataset (1142) in accordance with the information (1110) describing the modifications to the dataset (1142), and confirming (1146) receipt of completion of the request (1104) to modify the dataset (1142) to the leader storage system (1140), such as by using an acknowledgement (1148) message or other suitable mechanism.
[0244]
[0302] In the example method shown in Figure 12, the information (1110) describing modifications to a dataset (1142) may include ordering information (1152) for a request (1104) to modify the dataset (1142). In the example method shown in Figure 12, the ordering information (1152) for a request (1104) to modify the dataset (1142) may express a description ...
Claims
1. 1. A plurality of storage systems across which a data set is synchronously replicated, each storage system including a computer memory and a computer processor, the computer memory within each of the storage systems, when executed by the computer processor of a particular storage system, providing to the particular storage system: Attaching to a pod, the pod having the dataset, a set of management objects and operations, and a set of access operations for modifying or reading the dataset; Management operations can modify or query managed objects equally well through any of said storage systems; access operations to read or modify said data set operate equally well through any of said storage systems; each storage system storing a separate copy of the dataset as an appropriate subset of the dataset stored and advertised for use by the storage system; A completed management object or operation to modify the dataset performed through any one storage system is reflected in a subsequent management object to query the pod or a subsequent access operation to read the dataset. and a plurality of storage systems. comprising computer program instructions for causing the computer to perform Multiple storage systems.
2. One or more of the storage systems, when executed by the computer processor of a particular storage system, receiving a request to read a portion of the data set; locally processing the request to read the portion of the dataset; comprising computer program instructions for causing the computer to perform 10. A plurality of storage systems according to claim 1.
3. One or more of the storage systems, when executed by the computer processor of a particular storage system, detecting an interruption in data communication with one or more of the other storage systems; determining whether the particular storage system should remain in the pod; In response to determining that the particular storage system should remain in the pod, keeping the dataset on the particular storage system accessible for management and dataset operations; in response to determining that the particular storage system should not remain in the pod, making the dataset on the particular storage system inaccessible for management and dataset operations; comprising computer program instructions for causing the computer to perform 10. A plurality of storage systems according to claim 1.
4. One or more of the storage systems, when executed by the computer processor of a particular storage system, detecting that the interruption in data communication with one or more of the other storage systems has been repaired; making the dataset on the particular storage system accessible for management and dataset operations; comprising computer program instructions for causing the computer to perform 4. A plurality of storage systems according to claim 3.
5. Two or more of the storage systems, when executed by the computer processor of each storage system, identifying a target storage system for asynchronously receiving the dataset, the target storage system being not one of the plurality of storage systems across which the dataset is synchronously replicated; identifying a portion of the data set that is not asynchronously replicated to the target storage system by any of the other storage systems; asynchronously replicating to the target storage system the portion of the dataset that has not been asynchronously replicated to the target storage system by any of the other storage systems, wherein the two or more storage systems collectively replicate the entire dataset to the target storage system; comprising computer program instructions for causing the computer to perform 10. A plurality of storage systems according to claim 1.
6. The plurality of storage systems of claim 1 , wherein at least one of the storage systems is implemented as cloud storage provided by a cloud service provider.
7. 1. A method for synchronously replicating a database across multiple storage systems, comprising: Attaching a pod to the plurality of storage systems, the pod receiving the dataset, a set of management objects and management operations, and a set of access operations for modifying or reading the dataset; Management operations can modify or query managed objects equally well through any of said storage systems; access operations to read or modify said data set operate equally well through any of said storage systems; each storage system storing a separate copy of the dataset as an appropriate subset of the dataset stored and advertised for use by the storage system; A completed management object or operation to modify the dataset performed through any one storage system is reflected in a subsequent management object to query the pod or a subsequent access operation to read the dataset. Attaching multiple storage systems, including A method comprising:
8. receiving a request to read a portion of the data set by a particular storage system, the particular storage system being one of the plurality of storage systems; locally processing the request to read the portion of the dataset by the particular storage system; and The method of claim 7 further comprising:
9. detecting, by a particular storage system of the plurality of storage systems, an interruption in data communication with one or more of the other storage systems of the plurality of storage systems; determining whether the particular storage system should remain in the pod; In response to determining that the particular storage system should remain in the pod, keeping the dataset on the particular storage system accessible for management and dataset operations; In response to determining that the particular storage system should not remain in the pod, making the dataset on the particular storage system inaccessible for management and dataset operations; The method of claim 7 further comprising:
10. detecting that the interruption in data communication with one or more of the other storage systems has been repaired; making the dataset on the particular storage system accessible for management and dataset operations; 10. The method of claim 9, further comprising:
11. identifying a target storage system for asynchronously receiving the dataset, the target storage system not being one of the plurality of storage systems across which the dataset is synchronously replicated; identifying a portion of the data set that is not asynchronously replicated to the target storage system by any of the other storage systems; asynchronously replicating to the target storage system the portion of the dataset that has not been asynchronously replicated to the target storage system by any of the other storage systems, wherein the two or more storage systems collectively replicate the entire dataset to the target storage system; The method of claim 1 further comprising:
12. The method of claim 7 , wherein at least one of the storage systems is implemented as cloud storage provided by a cloud service provider.
13. 1. An apparatus for synchronously replicating a data set across multiple storage systems, the apparatus comprising: a computer processor; and computer memory operatively coupled to the computer processor, the computer memory comprising: When executed by the computer processor, the apparatus: Attaching to a pod, the pod having the dataset, a set of management objects and operations, and a set of access operations for modifying or reading the dataset; Management operations can modify or query managed objects equally well through any of said storage systems; access operations to read or modify said data set operate equally well through any of said storage systems; each storage system storing a separate copy of the dataset as an appropriate subset of the dataset stored and advertised for use by the storage system; A completed management object or operation to modify the dataset performed through any one storage system is reflected in a subsequent management object to query the pod or a subsequent access operation to read the dataset. Attaching steps involving multiple storage systems and the computer memory having computer program instructions disposed therein that cause the computer to execute the
14. When executed by the computer processor, the apparatus: receiving a request to read a portion of the data set; locally processing the request to read the portion of the dataset; 14. The apparatus of claim 13, further comprising computer program instructions to cause the apparatus to perform:
15. When executed by the computer processor, the apparatus: detecting an interruption in data communication with one or more of the other storage systems; determining whether the particular storage system should remain in the pod; In response to determining that the particular storage system should remain in the pod, keeping the dataset on the particular storage system accessible for management and dataset operations; in response to determining that the particular storage system should not remain in the pod, making the dataset on the particular storage system inaccessible for management and dataset operations; The apparatus of claim 8 , further comprising computer program instructions to perform:
16. When executed by the computer processor, the apparatus: detecting that the interruption in data communication with one or more of the other storage systems has been repaired; making the dataset on the particular storage system accessible for management and dataset operations; 16. The apparatus of claim 15, further comprising computer program instructions to cause the apparatus to perform:
17. When executed by the computer processor, the apparatus: identifying a target storage system for asynchronously receiving the dataset, the target storage system being not one of the plurality of storage systems across which the dataset is synchronously replicated; identifying a portion of the data set that is not asynchronously replicated to the target storage system by any of the other storage systems; asynchronously replicating to the target storage system the portion of the dataset that has not been asynchronously replicated to the target storage system by any of the other storage systems, wherein the two or more storage systems collectively replicate the entire dataset to the target storage system; The apparatus of claim 8 , further comprising computer program instructions to perform:
Citation Information
Patent Citations
How to prevent "split brain" in computer clustering systems
JP2004516575A
Cluster database with remote data mirroring
JP2007518195A
Post-cluster brain split quorum processing method and quorum storage device and system
WO2016107173A1