Role Enforcement for Storage as a Service
By shifting data management responsibilities to the operating system in a direct-mapped flash storage system, the inefficiencies and reliability issues in conventional storage systems are addressed, resulting in improved performance and reduced redundant operations.
Patent Information
- Application Number
- JP2023564236
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2021-07-01
- Filing Date
- 2022-05-11
- Publication Date
- 2025-12-01
- Estimated Expiration
- 2042-05-11
AI Technical Summary
Conventional storage systems face inefficiencies due to lower-level processes being performed by storage controllers, leading to unnecessary write operations and reduced reliability, especially in flash storage systems.
Implementing a direct-mapped flash storage system where higher-level processes, such as data block addressing and management, are handled by the operating system rather than the storage controller, optimizing zone management and reducing redundant operations.
This approach enhances the reliability and efficiency of flash storage systems by minimizing unnecessary write operations and improving overall system performance.
Smart Images

Figure 0007778158000001 
Figure 0007778158000002 
Figure 0007778158000003
Abstract
Description
[Brief explanation of the drawings]
[0001] [Figure 1A] 1 illustrates a first exemplary system for data storage, according to some implementations. [Figure 1B] 1 illustrates a second exemplary system for data storage, according to some implementations. [Figure 1C] 1 illustrates a third exemplary system for data storage, according to some implementations. [Figure 1D] 1 illustrates a fourth exemplary system for data storage, according to some implementations. [Figure 2A] FIG. 1 is a perspective view of a storage cluster having multiple storage nodes and internal storage coupled to each storage node to provide network-attached storage, according to some embodiments. [Figure 2B] FIG. 2 is a block diagram illustrating an interconnect switch coupling multiple storage nodes, according to some embodiments. [Figure 2C] FIG. 2 is a multi-level block diagram illustrating the contents of a storage node and the contents of one of the non-volatile solid-state storage units, according to some embodiments. [Figure 2D] 1 illustrates a storage server environment that uses embodiments of the storage nodes and storage units of some of the previous figures, according to some embodiments. [Figure 2E] FIG. 1 is a blade hardware block diagram illustrating the control plane, the compute and storage plane, and authorities interacting with the underlying physical resources, according to some embodiments. [Figure 2F] 1 depicts an elasticity software layer within a blade of a storage cluster, according to some embodiments. [Figure 2G] 1 depicts permissions and storage resources within blades of a storage cluster, according to some embodiments. [Figure 3A]1 illustrates an illustration of a storage system coupled for data communication with a cloud service provider, according to some embodiments of the present disclosure. [Figure 3B] 1 illustrates a diagram of a storage system in accordance with some embodiments of the present disclosure. [Figure 3C] An example of a cloud-based storage system according to some embodiments of the present disclosure is described. [Figure 3D] 1 illustrates an exemplary computing device that may be specifically configured to perform one or more of the processes described herein. [Figure 3E] 1 illustrates an example fleet of storage systems for providing storage services, according to some embodiments of the present disclosure. [Figure 4] 1 illustrates a block diagram including an edge management service for delivering storage services, according to some embodiments of the present disclosure. [Figure 5] 10 depicts a flowchart illustrating an exemplary method for providing data management as a service, according to some embodiments of the present disclosure. [Figure 6] 1 illustrates a block diagram that includes a system for role enforcement for storage as a service, according to some embodiments of the present disclosure. [Figure 7] 10 depicts a flowchart illustrating an exemplary method for role enforcement for storage as a service, according to some embodiments of the present disclosure. [Figure 8] 10 depicts a flowchart illustrating an example method for adding role enforcement for storage as a service, according to some embodiments of the present disclosure. [Figure 9] 10 depicts a flowchart illustrating an example method for adding role enforcement for storage as a service, according to some embodiments of the present disclosure. [Figure 10] 10 depicts a flowchart illustrating an example method for adding role enforcement for storage as a service, according to some embodiments of the present disclosure. [Figure 11] 10 depicts a flowchart illustrating an example method for adding role enforcement for storage as a service, according to some embodiments of the present disclosure. [Figure 12] 10 depicts a flowchart illustrating an example method for adding role enforcement for storage as a service, according to some embodiments of the present disclosure. [Figure 13] 10 depicts a flowchart illustrating an example method for adding role enforcement for storage as a service, according to some embodiments of the present disclosure. DETAILED DESCRIPTION OF THE INVENTION
[0002] Exemplary methods, apparatus, and products for storage-as-a-service role enforcement according to embodiments of the present disclosure are described with reference to the accompanying drawings, beginning with FIG. 1A. FIG. 1A illustrates an exemplary system for data storage according to some implementations. System 100 (also referred to herein as a "storage system") includes a number of elements, for purposes of illustration and not limitation. It should be noted that system 100 may include the same, more, or fewer elements, configured in the same or different ways, in other implementations.
[0003] System 100 includes multiple computing devices 164A-B. The computing devices (also referred to herein as "client devices") may be embodied as, for example, servers in a data center, workstations, personal computers, notebooks, etc. The computing devices 164A-B may be coupled for data communication to one or more storage arrays 102A-B via a storage area network ("SAN") 158 or a local area network ("LAN") 160.
[0004] SAN 158 may be implemented using a variety of data communication fabrics, devices, and protocols. For example, fabrics for SAN 158 may include Fibre Channel, Ethernet, InfiniBand, Serial Attached Small Computer System Interface ("SAS"), etc. Data communication protocols used with SAN 158 may include Advanced Technology Attachment ("ATA"), Fibre Channel Protocol, Small Computer System Interface ("SCSI"), Internet Small Computer System Interface ("iSCSI"), HyperSCSI, Non-Volatile Memory Express ("NVMe") over fabric, etc. Note that SAN 158 is provided for purposes of illustration and not limitation. Other data communication couplings may be implemented between computing devices 164A-B and storage arrays 102A-B.
[0005] LAN 160 may also be implemented using a variety of fabrics, devices, and protocols. For example, fabrics for LAN 160 may include Ethernet (802.3), wireless (802.11), etc. Data communication protocols used in LAN 160 may include Transmission Control Protocol ("TCP"), User Datagram Protocol ("UDP"), Internet Protocol ("IP"), HyperText Transfer Protocol ("HTTP"), Wireless Access Protocol ("WAP"), Handheld Device Transport Protocol ("HDTP"), Session Initiation Protocol ("SIP"), Real Time Protocol ("RTP"), etc. LAN 160 may also be connected to the Internet 162.
[0006] Storage arrays 102A-B can provide persistent data storage for computing devices 164A-B. In implementations, storage array 102A can be housed in a chassis (not shown) and storage array 102B can be housed in another chassis (not shown). Storage arrays 102A and 102B can include one or more storage array controllers 110A-D (also referred to herein as "controllers"). Storage array controllers 110A-D can be embodied as modules of an automated computing machine including computer hardware, computer software, or a combination of computer hardware and software. In some implementations, storage array controllers 110A-D can be configured to perform various storage tasks. Storage tasks may include writing data received from computing devices 164A-B to storage arrays 102A-B, erasing data from storage arrays 102A-B, retrieving data from storage arrays 102A-B and providing the data to computing devices 164A-B, monitoring and reporting disk usage and performance, performing redundancy operations such as a Redundant Array of Independent Drives ("RAID") or RAID-like data redundancy operations, compressing data, encrypting data, etc.
[0007] The storage array controllers 110A-D may be implemented in a variety of ways, including as a Field Programmable Gate Array ("FPGA"), a Programmable Logic Chip ("PLC"), an Application Specific Integrated Circuit ("ASIC"), a System-on-Chip ("SOC"), or any computing device that includes discrete components such as a processing device, a central processing unit, computer memory, or various adapters. The storage array controllers 110A-D may include a data communications adapter configured to support communications over, for example, the SAN 158 or the LAN 160. In some implementations, the storage array controllers 110A-D may be independently coupled to the LAN 160. In implementations, the storage array controllers 110A-D may include an I / O controller or the like that couples the storage array controllers 110A-D to persistent storage resources 170A-B (also referred to herein as "storage resources") for data communications over a midplane (not shown). Persistent storage resources 170A-B may include any number of storage drives 171A-F (also referred to herein as "storage devices") and any number of non-volatile random access memory ("NVRAM") devices (not shown).
[0008] In some implementations, the NVRAM devices of persistent storage resources 170A-B may be configured to receive data to be stored on storage drives 171A-F from storage array controllers 110A-D. In some examples, the data may originate from computing devices 164A-B. In some examples, writing data to an NVRAM device may be performed more quickly than writing data directly to storage drives 171A-F. In implementations, storage array controllers 110A-D may be configured to utilize an NVRAM device as a quickly accessible buffer for data to be written to storage drives 171A-F. The latency of write requests using an NVRAM device as a buffer may be improved relative to systems in which storage array controllers 110A-D write data directly to storage drives 171A-F. In some implementations, the NVRAM devices may be implemented using computer memory in the form of high-bandwidth, low-latency RAM. NVRAM devices are referred to as “non-volatile” because they may receive or include their own power source that maintains the RAM state after a main power loss to the NVRAM device. Such a power source may be a battery, one or more capacitors, etc. In response to a power loss, the NVRAM device may be configured to write the contents of the RAM to persistent storage, such as storage drives 171A-F.
[0009] In implementations, storage drives 171A-F may refer to any device configured to persistently record data, where "persistently" or "persistent" refers to the device's ability to maintain recorded data after a loss of power. In some implementations, storage drives 171A-F may correspond to non-disk storage media. For example, storage drives 171A-F may be one or more solid-state drives ("SSDs"), flash memory-based storage, any type of solid-state non-volatile memory, or any other type of non-mechanical storage device. In other implementations, storage drives 171A-F may include mechanical or rotating hard disks, such as hard disk drives ("HDDs").
[0010] In some implementations, the storage array controllers 110A-D may be configured to offload device management responsibilities from the storage drives 171A-F in the storage arrays 102A-B. For example, the storage array controllers 110A-D may manage control information that may describe the state of one or more memory blocks in the storage drives 171A-F. The control information may indicate, for example, that a particular memory block has failed and should no longer be written to, that a particular memory block contains boot code for the storage array controllers 110A-D, the number of program-erase ("P / E") cycles performed on a particular memory block, the age of the data stored in a particular memory block, the type of data stored in a particular memory block, etc. In some implementations, the control information may be stored as metadata with the associated memory block. In other implementations, the control information for the storage drives 171A-F may be stored in one or more specific memory blocks of the storage drives 171A-F selected by the storage array controllers 110A-D. The selected memory block may be tagged with an identifier indicating that the selected memory block contains control information. The identifiers may be utilized by storage array controllers 110A-D in conjunction with storage drives 171A-F to quickly identify memory blocks containing the control information. For example, storage controllers 110A-D may issue commands that specify the locations of memory blocks containing the control information. Note that the control information may be large enough that portions of the control information may be stored in multiple locations, the control information may be stored in multiple locations, for example, for redundancy purposes, or the control information may be otherwise distributed across multiple memory blocks within storage drives 171A-F.
[0011] In an implementation, storage array controllers 110A-D can offload device management responsibilities from storage drives 171A-F of storage arrays 102A-B by retrieving control information from storage drives 171A-F that describes the state of one or more memory blocks within storage drives 171A-F. Retrieving the control information from storage drives 171A-F may be performed, for example, by storage array controllers 110A-D querying storage drives 171A-F for the location of the control information for a particular storage drive 171A-F. The storage drives 171A-F may be configured to execute instructions that enable the storage drives 171A-F to identify the location of the control information. The instructions may be executed by a controller (not shown) associated with or otherwise located on the storage drives 171A-F, and may cause the storage drives 171A-F to scan a portion of each memory block to identify the memory block that stores the control information for the storage drives 171A-F. The storage drives 171A-F may respond by sending a response message to the storage array controllers 110A-D that includes the location of the control information for the storage drives 171A-F. In response to receiving the response message, the storage array controllers 110A-D may issue a request to read the data stored at the address associated with the location of the control information for the storage drives 171A-F.
[0012] In other implementations, storage array controllers 110A-D can further offload device management responsibilities from storage drives 171A-F by performing storage drive management operations in response to receiving the control information. The storage drive management operations may include, for example, operations typically performed by storage drives 171A-F (e.g., a controller (not shown) associated with a particular storage drive 171A-F). The storage drive management operations may include, for example, ensuring that data is not written to failed memory blocks within storage drives 171A-F, ensuring that data is written to memory blocks within storage drives 171A-F such that proper wear leveling is achieved, etc.
[0013] In implementations, storage arrays 102A-B may implement two or more storage array controllers 110A-D. For example, storage array 102A may include storage array controller 110A and storage array controller 110B. At a given instance, a single storage array controller 110A-D (e.g., storage array controller 110A) of storage system 100 may be designated with primary status (also referred to herein as a “primary controller”), and the other storage array controller 110A-D (e.g., storage array controller 110A) may be designated with secondary status (also referred to herein as a “secondary controller”). The primary controller may have certain rights, such as permission to modify data in persistent storage resources 170A-B (e.g., write data to persistent storage resources 170A-B). At least some of the rights of the primary controller may supersede the rights of the secondary controller. For example, if the primary controller has the right, the secondary controller may not have permission to modify the data in the persistent storage resources 170A-B. The states of the storage array controllers 110A-D may change. For example, the storage array controller 110A may be designated with secondary status, and the storage array controller 110B may be designated with primary status.
[0014] In some implementations, a primary controller, such as storage array controller 110A, may serve as the primary controller for one or more storage arrays 102A-B, and a second controller, such as storage array controller 110B, may serve as a secondary controller for one or more storage arrays 102A-B. For example, storage array controller 110A may be the primary controller for storage array 102A and storage array 102B, and storage array controller 110B may be the secondary controller for storage arrays 102A and 102B. In some implementations, storage array controllers 110C and 110D (also referred to as "storage processing modules") may not have primary or secondary status. Storage array controllers 110C and 110D, implemented as storage processing modules, may serve as a communication interface between the primary and secondary controllers (e.g., storage array controllers 110A and 110B, respectively) and storage array 102B. For example, storage array controller 110A of storage array 102A may send a write request to storage array 102B via SAN 158. The write request may be received by both storage array controllers 110C and 110D of storage array 102B. Storage array controllers 110C and 110D facilitate the communication, for example, sending the write request to the appropriate storage drives 171A-F. Note that in some implementations, storage processing modules can be used to increase the number of storage drives controlled by the primary and secondary controllers.
[0015] In an implementation, storage array controllers 110A-D are communicatively coupled to one or more storage drives 171A-F and one or more NVRAM devices (not shown) included as part of storage arrays 102A-B via a midplane (not shown). Storage array controllers 110A-D may be coupled to the midplane via one or more data communication links, and the midplane may be coupled to storage drives 171A-F and the NVRAM devices via one or more data communication links. The data communication links described herein are collectively illustrated by data communication links 108A-D and may include, for example, a Peripheral Component Interconnect Express ("PCIe") bus.
[0016] FIG. 1B illustrates an exemplary system for data storage, according to some implementations. The storage array controller 101 illustrated in FIG. 1B may be similar to storage array controllers 110A-D described with respect to FIG. 1A. In one example, storage array controller 101 may be similar to storage array controller 110A or storage array controller 110B. Storage array controller 101 includes multiple elements for purposes of illustration and not limitation. Note that in other implementations, storage array controller 101 may include the same, more, or fewer elements, configured in the same or different ways. Note that elements of FIG. 1A may be included below to help illustrate features of storage array controller 101.
[0017] Storage array controller 101 may include one or more processing devices 104 and random access memory ("RAM") 111. Processing device 104 (or controller 101) represents one or more general-purpose processing devices, such as a microprocessor, central processing unit, or the like. More specifically, processing device 104 (or controller 101) may be a complex instruction set computing ("CISC") microprocessor, a reduced instruction set computing ("RISC") microprocessor, a very long instruction word ("VLIW") microprocessor, or a processor implementing other instruction sets or a processor implementing a combination of instruction sets. Processing device 104 (or controller 101) may also be one or more special-purpose processing devices, such as an ASIC, an FPGA, a digital signal processor ("DSP"), a network processor, or the like.
[0018] The processing device 104 may be connected to RAM 111 via a data communication link 106, which may be embodied as a high-speed memory bus such as a Double-Data Rate 4 ("DDR4") bus. An operating system 112 is stored in RAM 111. In some implementations, instructions 113 are stored in RAM 111. The instructions 113 may include computer program instructions for performing operations in a direct-mapped flash storage system. In one embodiment, a direct-mapped flash storage system is a system that directly addresses blocks of data within a flash drive without address translation being performed by the flash drive's storage controller.
[0019] In an implementation, storage array controller 101 includes one or more host bus adapters 103A-C coupled to processing device 104 via data communication links 105A-C. In an implementation, host bus adapters 103A-C may be computer hardware that connects a host system (e.g., a storage array controller) to other networks and storage arrays. In some examples, host bus adapters 103A-C may be Fibre Channel adapters that allow storage array controller 101 to connect to a SAN, Ethernet adapters that allow storage array controller 101 to connect to a LAN, etc. Host bus adapters 103A-C may be coupled to processing device 104 via data communication links 105A-C, such as a PCIe bus.
[0020] In an implementation, storage array controller 101 may include a host bus adapter 114 coupled to an expander 115. Expander 115 may be used to attach a host system to a larger number of storage drives. Expander 115 may be, for example, a SAS expander utilized to allow host bus adapter 114 to attach to storage drives in an implementation in which host bus adapter 114 is embodied as a SAS controller.
[0021] In an implementation, storage array controller 101 may include a switch 116 coupled to processing device 104 via data communication link 109. Switch 116 may be a computer hardware device that can create multiple endpoints from a single endpoint, thereby allowing multiple devices to share a single endpoint. Switch 116 may be, for example, a PCIe switch coupled to a PCIe bus (e.g., data communication link 109) and providing multiple PCIe connection points to a midplane.
[0022] In an implementation, storage array controller 101 includes data communication link 107 for coupling storage array controller 101 to other storage array controllers. In some examples, data communication link 107 may be a QuickPath Interconnect (QPI) interconnect.
[0023] A conventional storage system that uses conventional flash drives may implement processes across flash drives that are part of the conventional storage system. For example, higher-level processes in the storage system may initiate and control processes across the flash drives. However, flash drives in a conventional storage system may include their own storage controller that also implements processes. Thus, a conventional storage system may implement both higher-level processes (e.g., initiated by the storage system) and lower-level processes (e.g., initiated by the storage controller of the storage system).
[0024] To address various shortcomings of conventional storage systems, operations can be performed by higher-level processes rather than by lower-level processes. For example, a flash storage system may include a flash drive that does not include a storage controller to provide the processes. Thus, the operating system of the flash storage system itself can initiate and control the processes. This can be achieved by a direct-mapped flash storage system that directly addresses data blocks within the flash drive without the address translation performed by the flash drive's storage controller.
[0025] In implementations, storage drives 171A-F may be one or more zoned storage devices. In some implementations, one or more zoned storage devices may be a single HDD. In implementations, one or more storage devices may be flash-based SSDs. In a zoned storage device, the zoned namespace on the zoned storage device may be grouped by natural size and addressed by aligned groups of blocks to form multiple addressable zones. In implementations utilizing SSDs, the natural size may be based on the erase block size of the SSD. In some implementations, the zones of a zoned storage device may be defined during initialization of the zoned storage device. In implementations, the zones may be dynamically defined as data is written to the zoned storage device.
[0026] In some implementations, zones may be heterogeneous, with some zones each being a page group and other zones being multiple page groups. In implementations, some zones may correspond to an erase block and other zones may correspond to multiple erase blocks. In one implementation, zones may be any combination of different numbers of pages within page groups and / or erase blocks for a heterogeneous mix of programming modes, manufacturers, product types, and / or product generations of storage devices, as applied to heterogeneous assembly, upgrades, distributed storage, etc. In one implementation, zones may be defined as having usage characteristics, such as characteristics that support data with a particular type of lifespan (e.g., very short-lived or very long-lived). These characteristics may be used by the zoned storage device to determine how the zone is managed over the expected lifespan of the zone.
[0027] It should be understood that zones are virtual constructs. Any particular zone may not have a fixed location on the storage device. Until allocated, a zone may not have any location on the storage device. In various implementations, a zone may correspond to a number representing a virtually allocatable chunk of space, the size of an erase block or other block size. When the system allocates or opens a zone, the zone is allocated to flash or other solid-state storage memory, and when the system writes to the zone, the pages are written to the mapped flash or other solid-state storage memory of the zoned storage device. When the system closes a zone, the associated erase block or block of other size is completed. At some point in the future, the system can delete the zone, which frees up the zone's allocated space. During its lifetime, a zone may be moved to a different location on the zoned storage device, for example, when the zoned storage device undergoes internal maintenance.
[0028] In implementations, zones in a zoned storage device can be in different states. A zone can be empty, with no data stored in it. An empty zone can be opened explicitly or implicitly by writing data to the zone. This is the initial state of a zone on a new zoned storage device, but can also be the result of a zone reset. In some implementations, an empty zone can have a specified location within the flash memory of the zoned storage device. In one implementation, the location of an empty zone can be selected when the zone is first opened or first written to (or later, if the write is buffered in memory). Zones can be in the open state either implicitly or explicitly, and a zone in the open state can be written to store data using a write command or an append command. In one implementation, a zone in the open state can be written using a copy command, which copies data from a different zone. In some implementations, a zoned storage device can have a limit on the number of open zones at a particular time.
[0029] A closed zone is a zone that has been partially written to but entered the closed state after issuing an explicit close operation. A closed zone may remain available for future writes, but may reduce some of the runtime overhead consumed by keeping the zone open. In implementations, a zoned storage device may have a limit on the number of closed zones at a particular time. A full zone is a zone that stores data and can no longer be written to. A zone can be in the full state either after a write has written data to the entire zone or as a result of a zone close operation. Before the close operation, the zone may or may not be completely written to. However, after the close operation, the zone will not be open for further writes without first performing a zone reset operation.
[0030] The mapping from zones to erase blocks (or single tracks in an HDD) may be arbitrary, dynamic, or hidden from view. The process of opening a zone may be an operation that allows a new zone to be dynamically mapped to the underlying storage of a zoned storage device, and then allows data to be written to the zone by appending writes until the zone reaches capacity. A zone may be terminated at any point, after which no further data may be written to the zone. When the data stored in a zone is no longer needed, the zone may be reset, thereby effectively deleting the zone's contents from the zoned storage device and making the physical storage held by that zone available for subsequent data storage. Once a zone is written and terminated, the zoned storage device ensures that the data stored in the zone is not lost until the zone is reset. In the time between writing data to a zone and resetting the zone, the zone may be moved between single tracks or erase blocks, such as by copying data, to keep data refreshed as part of a maintenance operation in the zoned storage device, or to handle the aging of memory cells in an SSD.
[0031] In implementations utilizing HDDs, resetting a zone may allow a single track to be allocated to a new opened zone that may be opened at some point in the future. In implementations utilizing SSDs, resetting a zone causes the zone's associated physical erase blocks to be erased and then reused for storage of data. In some implementations, a zoned storage device may have a limit on the number of zones that are open at a given time to reduce the amount of overhead dedicated to keeping zones open.
[0032] The operating system of the flash storage system can identify and maintain a list of allocation units across multiple flash drives of the flash storage system. An allocation unit can be an entire erase block or multiple erase blocks. The operating system can maintain a map or address ranges that directly map addresses to erase blocks on the flash drives of the flash storage system.
[0033] Direct mapping to erase blocks on a flash drive can be used to rewrite data and erase data. For example, operations can be performed on one or more allocation units that include first and second data, where the first data is retained and the second data is no longer in use by the flash storage system. The operating system can initiate a process to write the first data to a new location in another allocation unit, erase the second data, and mark the allocation unit as available for subsequent data. Thus, the process can be performed solely by the higher-level operating system of the flash storage system, without additional lower-level processes performed by the flash drive's controller.
[0034] Advantages of processes performed solely by the flash storage system's operating system include improved reliability of the flash drives in the flash storage system, since unnecessary or redundant write operations are not performed during the process. One potential novelty is the concept of initiating and controlling the process in the flash storage system's operating system. Additionally, the process may be controlled by the operating system across multiple flash drives. This is in contrast to processing performed by the flash drive's storage controller.
[0035] A storage system may consist of two storage array controllers that share a set of drives for failover purposes, or may consist of a single storage array controller that provides storage services utilizing multiple drives, or may consist of a distributed network of storage array controllers, each having some number of drives or some amount of flash storage, where the storage array controllers in the network cooperate to provide a complete storage service and cooperate with respect to various aspects of the storage service, including storage allocation and garbage collection.
[0036] 1C illustrates a third exemplary system 117 for data storage, according to some implementations. System 117 (also referred to herein as a "storage system") includes a number of elements, for purposes of illustration and not limitation. Note that system 117 may include the same, more, or fewer elements, configured in the same or different ways, in other implementations.
[0037] In one embodiment, system 117 includes dual Peripheral Component Interconnect ("PCI") flash storage device 118 with separately addressable high-speed write storage. System 117 may include storage device controller 119. In one embodiment, storage device controllers 119A-D may be CPUs, ASICs, FPGAs, or any other circuitry capable of implementing the necessary control structures in accordance with the present disclosure. In one embodiment, system 117 includes flash memory devices (e.g., including flash memory devices 120a-n) operably coupled to various channels of storage device controller 119. Flash memory devices 120a-n may be presented to storage device controllers 119A-D as addressable collections of flash pages, erase blocks, and / or control elements sufficient to enable controllers 119A-D to program and retrieve various aspects of the flash. In one embodiment, storage device controllers 119A-D can perform operations on flash memory devices 120a-n, including storing and retrieving data contents of pages, allocating and erasing any blocks, tracking statistics regarding the use and reuse of flash memory pages, erase blocks, and cells, tracking and predicting error codes and failures within flash memory, controlling voltage levels associated with programming flash cells and retrieving contents, etc.
[0038] In one embodiment, system 117 may include RAM 121 for storing separately addressable high-speed write data. In one embodiment, RAM 121 may be one or more separate discrete devices. In another embodiment, RAM 121 may be integrated into storage device controllers 119A-D or multiple storage device controllers. RAM 121 may also be utilized for other purposes, such as temporary program memory for a processing device (e.g., a CPU) within storage device controller 119.
[0039] In one embodiment, system 117 may include a stored energy device 122, such as a rechargeable battery or capacitor. Stored energy device 122 may store enough energy to power storage device controller 119, a certain amount of RAM (e.g., RAM 121), and a certain amount of flash memory (e.g., flash memories 120a-120n) for a sufficient time to write the contents of the RAM to flash memory. In one embodiment, storage device controllers 119A-D may write the contents of RAM to flash memory if the storage device controller detects a loss of external power.
[0040] In one embodiment, system 117 includes two data communication links 123a, 123b. In one embodiment, data communication links 123a, 123b may be PCI interfaces. In other embodiments, data communication links 123a, 123b may be based on other communication standards (e.g., HyperTransport, InfiniBand, etc.). Data communication links 123a, 123b may be based on the Non-Volatile Memory Express (“NVMe”) or NVMe over fabric (“NVMf”) specifications, which allow external connections from other components within storage system 117 to storage device controllers 119A-D. Note that the data communication links may be referred to interchangeably as PCI buses for convenience herein.
[0041] System 117 may also include an external power source (not shown), which may be provided via one or both of data communication links 123a, 123b, or may be provided separately. An alternative embodiment includes a separate flash memory (not shown) dedicated for use in storing the contents of RAM 121. Storage device controllers 119A-D may present a logical device on the PCI bus, which may include an addressable fast-write logical device, or a separate portion of the logical address space of storage device 118, which may be presented as PCI memory or persistent storage. In one embodiment, operations to store to the device are directed to RAM 121. During a power outage, storage device controllers 119A-D may write stored content associated with the addressable fast-write logical storage to flash memory (e.g., flash memories 120a-n) for long-term persistent storage.
[0042] In one embodiment, the logical device may include some representation of some or all of the contents of flash memory devices 120a-n that allows a storage system, including storage device 118 (e.g., storage system 117), to directly address flash memory pages and reprogram erase blocks directly from storage system components external to the storage device via a PCI bus. This representation may also allow one or more of the external components to control and retrieve other aspects of the flash memory, including some or all of tracking statistics related to the use and reuse of flash memory pages, erase blocks, and cells across all flash memory devices, tracking and predicting error codes and failures within and across flash memory devices, controlling voltage levels associated with programming and retrieving the contents of flash cells, etc.
[0043] In one embodiment, stored energy device 122 may be sufficient to ensure completion of ongoing operations on flash memory devices 120a-120n, and stored energy device 122 may power storage device controllers 119A-D and associated flash memory devices (e.g., 120a-n) for those operations as well as for fast write RAM storage to flash memory. Stored energy device 122 may be used to store cumulative statistics and other parameters kept and tracked by flash memory devices 120a-n and / or storage device controller 119. A separate capacitor or stored energy device (such as a smaller capacitor near or embedded within the flash memory device itself) may be used for some or all of the operations described herein.
[0044] Various schemes may be used to track and optimize the life of stored energy components, such as adjusting voltage levels over time, partially discharging stored energy device 122 to measure corresponding discharge characteristics, etc. If available energy decreases over time, the effective available capacity of the addressable fast-write storage may be reduced to ensure that it can be safely written based on the currently available stored energy.
[0045] 1D illustrates a third exemplary storage system 124 for data storage, according to some implementations. In one embodiment, storage system 124 includes storage controllers 125a, 125b. In one embodiment, storage controllers 125a, 125b are operably coupled to a dual PCI storage device. Storage controllers 125a, 125b can be operably coupled to a number of host computers 127a-n (e.g., via storage network 130).
[0046] In one embodiment, two storage controllers (e.g., 125a and 125b) provide storage services such as an SCS block storage array, a file server, an object server, a database, or a data analysis service. The storage controllers 125a, 125b can provide services to host computers 127a-n external to the storage system 124 through a number of network interfaces (e.g., 126a-d). The storage controllers 125a, 125b can provide services or applications integrated entirely within the storage system 124, forming an integrated storage and computing system. The storage controllers 125a, 125b can utilize high-speed write memory in or across the storage devices 119a-d to journal ongoing operations, ensuring that operations are not lost due to power outage, removal of a storage controller, storage controller or storage system shutdown, or any failure of one or more software or hardware components within the storage system 124.
[0047] In one embodiment, storage controllers 125a, 125b act as PCI masters for one or the other PCI bus 128a, 128b. In another embodiment, 128a and 128b may be based on other communication standards (e.g., HyperTransport, InfiniBand, etc.). Other storage system embodiments may allow storage controllers 125a, 125b to act as multi-masters for both PCI buses 128a, 128b. Alternatively, a PCI / NVMe / NVMf switching infrastructure or fabric may connect multiple storage controllers. Some storage system embodiments may allow storage devices to communicate directly with each other rather than communicating only with the storage controllers. In one embodiment, storage device controller 119a may be operable under direction from storage controller 125a to combine and transfer data to be stored in a flash memory device from data stored in RAM (e.g., RAM 121 of FIG. 1C ). For example, a recomputed version of the RAM contents can be transferred after the storage controller determines that the operation has been fully committed across the storage system, or when the fast write memory on the device reaches a certain used capacity, or after a certain amount of time to ensure improved data security or to free up addressable fast write capacity for reuse. This mechanism can be used, for example, to avoid a second transfer over a bus (e.g., 128a, 128b) from the storage controller 125a, 125b. In one embodiment, the recomputation can include compressing the data, attaching indexing or other metadata, combining multiple data segments together, performing erasure coding calculations, etc.
[0048] In one embodiment, under direction from storage controller 125a, 125b, storage device controller 119a, 119b may be operable to calculate and transfer data from data stored in RAM (e.g., RAM 121 of FIG. 1C) to other storage devices without the involvement of storage controller 125a, 125b. This operation may be used to mirror data stored in one storage controller 125a to another storage controller 125b, or to offload compression, data aggregation, and / or erasure coding calculations and transfer them to storage devices to reduce the load on storage controller or storage controller interface 129a, 129b to PCI bus 128a, 128b.
[0049] The storage device controllers 119A-D may include mechanisms for implementing high availability primitives used by other parts of the storage system external to the dual PCI storage device 118. For example, in a storage system having two storage controllers providing highly available storage services, reservation or exclusion primitives may be provided so that one storage controller can prevent the other storage controller from accessing or continuing to access the storage device. This may be used, for example, when one controller detects that the other controller is not functioning properly, or when the interconnect between the two storage controllers itself may not be functioning properly.
[0050] In one embodiment, a storage system for use with dual PCI direct-mapped storage devices having separately addressable, fast-write storage includes a system for managing erase blocks or groups of erase blocks as allocation units for storing data on behalf of a storage service, for storing metadata associated with the storage service (e.g., indexes, logs, etc.), or for proper management of the storage system itself. Flash pages, which may be several kilobytes in size, may be written as data arrives or as the storage system persists the data for an extended period of time (e.g., above a defined time threshold). To commit data more quickly or reduce the number of writes to the flash memory device, the storage controller may first write the data to separately addressable, fast-write storage on one or more storage devices.
[0051] In one embodiment, the storage controllers 125a, 125b can initiate the use of erase blocks within and across storage devices (e.g., 118) according to the age and expected remaining life of the storage device, or based on other statistics. The storage controllers 125a, 125b can initiate garbage collection and data migration of data between storage devices according to pages that are no longer needed, manage the lifespan of flash pages and erase blocks, and manage overall system performance.
[0052] In one embodiment, storage system 124 may utilize mirroring and / or erasure coding schemes as part of storing data in addressable, fast-write storage and / or as part of writing data to allocation units associated with erase blocks. Erasure codes may be used across storage devices, as well as within erase blocks or allocation units, or within and across flash memory devices on a single storage device, to provide redundancy against single or multiple storage device failures or to protect against internal corruption of flash memory pages resulting from flash memory operation or degradation of flash memory cells. Mirroring and erasure coding at various levels may be used to recover from multiple types of failures, occurring separately or in combination.
[0053] The embodiments depicted with reference to Figures 2A-2G illustrate a storage cluster that stores user data, such as user data originating from one or more user or client systems or other sources external to the storage cluster. The storage cluster distributes user data across storage nodes housed within a chassis or across multiple chassis using erasure coding and redundant copies of metadata. Erasure coding refers to a method of data protection or reconstruction in which data is stored across a set of distinct locations, such as disks, storage nodes, or geographic locations. Flash memory is one type of solid-state memory that may be integrated with embodiments, but embodiments can be extended to other types of solid-state memory or other storage media, including non-solid-state memory. Control of storage locations and workloads is distributed across storage locations in a clustered peer-to-peer system. Tasks such as mediating communications between various storage nodes, detecting when a storage node becomes unavailable, and balancing I / O (input and output) across various storage nodes are all handled on a distributed basis. Data, in some embodiments, is placed or distributed across multiple storage nodes in data fragments or stripes that support data recovery. Ownership of data can be reassigned within the cluster regardless of input and output patterns. This architecture, described in more detail below, allows storage nodes within a cluster to fail while the system remains operational, as data can be reconstructed from other storage nodes and therefore remain available for input and output operations. In various embodiments, the storage nodes may be referred to as cluster nodes, blades, or servers.
[0054] A storage cluster may be contained within a chassis, i.e., a housing that houses one or more storage nodes. Included within the chassis are mechanisms for providing power to each storage node, such as a power distribution bus, and communication mechanisms, such as a communication bus, that enable communication between the storage nodes. According to some embodiments, the storage cluster may operate as an independent system in one location. In one embodiment, the chassis includes at least two instances of both power distribution and communication buses that can be independently enabled or disabled. The internal communication bus may be an Ethernet bus, although other technologies, such as PCIe, InfiniBand, and others, are equally suitable. The chassis provides ports for an external communication bus to enable communication between multiple chassis and client systems, either directly or via a switch. External communication can use technologies such as Ethernet, InfiniBand, or Fibre Channel. In some embodiments, the external communication bus uses different communication bus technologies for inter-chassis communication and client communication. When a switch is deployed within or between chassis, the switch can act as a translator between multiple protocols or technologies. When multiple chassis are connected to define a storage cluster, the storage cluster may be accessed by clients using either proprietary or standard interfaces, such as network file system ("NFS"), common internet file system ("CIFS"), small computer system interface ("SCSI"), or hypertext transfer protocol ("HTTP"). Translation from the client protocol may occur within a switch, a chassis external communication bus, or each storage node. In some embodiments, multiple chassis may be coupled or connected to each other through an aggregator switch. Some and / or all of the coupled or connected chassis may be designated as a storage cluster.As mentioned above, each chassis may have multiple blades, and each blade has a media access control ("MAC") address, but the storage cluster, in some embodiments, is presented to the external network as having a single cluster MAC address and a single IP.
[0055] Each storage node may be one or more storage servers, each connected to one or more non-volatile solid-state memory units, which may be referred to as storage units or storage devices. One embodiment includes a single storage server and 1 to 8 non-volatile solid-state memory units within each storage node, but this example is not meant to be limiting. The storage server may include a processor, DRAM, an interface for an internal communication bus, and power distribution for each of the power buses. In some embodiments, within the storage node, the interface and storage units share a communication bus, e.g., PCI Express. The non-volatile solid-state memory units may directly access the internal communication bus interface via the storage node communication bus or may require the storage node to access the bus interface. In some embodiments, the non-volatile solid-state memory units include an embedded CPU, a solid-state storage controller, and a solid-state mass storage device, e.g., in amounts of 2 to 32 terabytes ("TB"). The non-volatile solid-state memory units include an internal volatile storage medium, such as DRAM, and an energy storage device. In some embodiments, the energy storage device is a capacitor, supercapacitor, or battery that allows for the transfer of a subset of the DRAM contents to a stable storage medium in the event of power loss. In some embodiments, the non-volatile solid-state memory unit is constructed with storage class memory such as phase change or magnetoresistive random access memory ("MRAM"), which replaces DRAM and allows for reduced power holdup devices.
[0056] One of the many features of the storage nodes and non-volatile solid-state storage is the ability to proactively rebuild data in a storage cluster. The storage nodes and non-volatile solid-state storage can determine when a storage node or non-volatile solid-state storage in a storage cluster becomes unreachable, regardless of whether there is an attempt to read the data associated with that storage node or non-volatile solid-state storage. The storage nodes and non-volatile solid-state storage then cooperate to recover and rebuild the data, at least in part, in a new location. This constitutes proactive rebuilding in that the system rebuilds data without waiting until the data is needed for a read access initiated by a client system using the storage cluster. These and further details of the storage nodes and their operation are discussed below.
[0057] FIG. 2A is a perspective view of a storage cluster 161 having multiple storage nodes 150 and internal solid-state memory coupled to each storage node to provide network-attached storage or a storage area network, according to some embodiments. A network-attached storage, storage area network, or storage cluster or other storage memory may include one or more storage clusters 161, each having one or more storage nodes 150, with a flexible and reconfigurable arrangement of both physical components and the amount of storage memory provided thereby. The storage cluster 161 is designed to fit into a rack, and one or more racks can be set up and populated as desired for storage memory. The storage cluster 161 includes a chassis 138 having multiple slots 142. It should be understood that the chassis 138 may also be referred to as a housing, enclosure, or rack unit. In one embodiment, the chassis 138 has 14 slots 142, although other numbers of slots are readily contemplated. For example, some embodiments have 4 slots, 8 slots, 16 slots, 32 slots, or other suitable numbers of slots. Each slot 142 can accommodate one storage node 150 in some embodiments. The chassis 138 includes a flap 148 that can be utilized to mount the chassis 138 in a rack. The fans 144 provide air circulation to cool the storage nodes 150 and their components, although other cooling components may be used, or an embodiment without cooling components may be devised. The switch fabric 146 couples the storage nodes 150 within the chassis 138 to each other and to a network for communication to memory. In one embodiment depicted herein, the slots 142 to the left of the switch fabric 146 and fans 144 are shown as occupied by a storage node 150, while the slots 142 to the right of the switch fabric 146 and fans 144 are empty and available for inserting a storage node 150 for illustrative purposes.This configuration is an example, and one or more storage nodes 150 can occupy slots 142 in a variety of additional arrangements. The arrangement of storage nodes need not be contiguous or adjacent in some embodiments. Storage nodes 150 are hot-pluggable, meaning that storage nodes 150 can be inserted into or removed from slots 142 in chassis 138 without shutting down or powering down the system. Upon insertion or removal of storage node 150 into or from slot 142, the system recognizes the change and automatically reconfigures to adapt. Reconfiguration, in some embodiments, includes restoring redundancy and / or rebalancing data or load.
[0058] Each storage node 150 may have multiple components. In the embodiment shown herein, storage node 150 includes a CPU 156, i.e., a printed circuit board 159 on which the processor is implemented, memory 154 coupled to CPU 156, and non-volatile solid-state storage 152 coupled to CPU 156, although other implementations and / or components may be used in further embodiments. Memory 154 contains instructions to be executed by CPU 156 and / or data to be operated on by CPU 156. As described further below, non-volatile solid-state storage 152 may include flash, or in further embodiments, other types of solid-state memory.
[0059] Referring to FIG. 2A , the storage cluster 161 is scalable, meaning that storage capacity having non-uniform storage sizes is easily added, as described above. In some embodiments, one or more storage nodes 150 can be plugged in or removed from each chassis, and the storage cluster self-configures. Plug-in storage nodes 150, whether installed in the chassis at delivery or added later, can have different sizes. For example, in one embodiment, the storage nodes 150 can have any multiple of 4 TB, e.g., 8 TB, 12 TB, 16 TB, 32 TB, etc. In further embodiments, the storage nodes 150 can have any multiple of other storage amounts or capacities. The storage capacity of each storage node 150 is broadcast and influences the determination of how data is striped. For maximum storage efficiency, one embodiment can self-configure as widely as possible within a stripe, subject to a given requirement for continuous operation with the loss of up to one or up to two non-volatile solid-state storage 152 units or storage nodes 150 within a chassis.
[0060] FIG. 2B is a block diagram illustrating a communication interconnect 173 and a power distribution bus 172 coupling multiple storage nodes 150. Referring back to FIG. 2A, the communication interconnect 173 may, in some embodiments, be included in or implemented with the switch fabric 146. When multiple storage clusters 161 occupy a rack, in some embodiments, the communication interconnect 173 may be included in or implemented with a top-of-rack switch. As illustrated in FIG. 2B, the storage cluster 161 is enclosed within a single chassis 138. The external port 176 is coupled to the storage node 150 via the communication interconnect 173, and the external port 174 is coupled directly to the storage node. The external power port 178 is coupled to the power distribution bus 172. The storage node 150 may include various amounts and capacities of non-volatile solid-state storage 152, as described with reference to FIG. 2A. Additionally, one or more of the storage nodes 150 may be compute-only storage nodes, as illustrated in FIG. 2B. Authorities 168 are implemented on non-volatile solid-state storage 152, for example, as a list or other data structure stored in-memory. In some embodiments, authorities are stored within non-volatile solid-state storage 152 and supported by software executing on a controller or other processor of non-volatile solid-state storage 152. In further embodiments, authorities 168 are implemented on storage node 150, for example, as a list or other data structure stored in memory 154 and supported by software executing on CPU 156 of storage node 150. In some embodiments, authorities 168 control how and where data is stored in non-volatile solid-state storage 152. This control helps determine what type of erasure coding scheme is applied to the data and which storage node 150 has which portion of the data. Each authority 168 can be assigned to non-volatile solid-state storage 152.Each authority, in various embodiments, can control a range of inode numbers, segment numbers, or other data identifiers assigned to data by the file system, by storage node 150, or by non-volatile solid-state storage 152.
[0061] In some embodiments, all data and all metadata are redundant within the system. Furthermore, all data and all metadata have an owner, which may be referred to as an authority. If the authority is unreachable, for example due to a storage node failure, there is a succession plan for how to find the data or its metadata. In various embodiments, there are redundant copies of authority 168. In some embodiments, authority 168 has a relationship to storage nodes 150 and non-volatile solid-state storage 152. Each authority 168 covering a range of data segment numbers or other identifiers of data may be assigned to a particular non-volatile solid-state storage 152. In some embodiments, authorities 168 for all such ranges are distributed across the non-volatile solid-state storage 152 of the storage cluster. Each storage node 150 has a network port that provides access to the non-volatile solid-state storage 152 of that storage node 150. Data may be stored in segments associated with a segment number, which in some embodiments is an indirect reference to the configuration of a RAID (Redundant Array of Independent Disks) stripe. Thus, the assignment and use of authority 168 establishes an indirect reference to the data. Indirect referencing may, according to some embodiments, be referred to as the ability to reference data indirectly, in this case via authority 168. A segment identifies a set of non-volatile solid-state storage 152 and a local identifier to the set of non-volatile solid-state storage 152 that may contain the data. In some embodiments, the local identifier is an offset into the device and may be reused sequentially by multiple segments. In other embodiments, the local identifier is unique to a particular segment and is never reused. The offset within non-volatile solid-state storage 152 is applied to locating data for writing to or reading from non-volatile solid-state storage 152 (in the form of a RAID stripe).Data is striped across multiple units of non-volatile solid-state storage 152, which may or may not include non-volatile solid-state storage 152 with authority 168 for particular data segments.
[0062] For example, during data migration or data reconstruction, if there is a change in where a particular segment of data is located, the authority 168 for that data segment should be referenced in the non-volatile solid-state storage 152 or storage node 150 that has that authority 168. To locate particular data, embodiments calculate a hash value for the data segment or apply an inode number or data segment number. The output of this operation points to the non-volatile solid-state storage 152 that has the authority 168 for that particular data. In some embodiments, this operation has two stages. The first stage is mapping an entity identifier (ID), such as a segment number, inode number, or directory number, to an authority identifier. This mapping may include a calculation such as a hash or bit mask. The second stage is mapping the authority identifier to a particular non-volatile solid-state storage 152, which can be done through explicit mapping. This operation is repeatable, so once the calculation is performed, the result of the calculation repeatably and reliably points to the particular non-volatile solid-state storage 152 that has that authority 168. The operation may include as input a set of reachable storage nodes. If the set of reachable non-volatile solid-state storage units changes, the optimal set changes. In some embodiments, the persisted value is the current allocation (which is always true) and the calculated value is the target allocation to which the cluster attempts to reconfigure. This calculation may be used to determine the optimal non-volatile solid-state storage 152 for authority, given a set of non-volatile solid-state storages 152 that are reachable and that constitute the same cluster. The calculation also determines an ordered set of peer non-volatile solid-state storages 152 that also record a mapping of authority to non-volatile solid-state storage so that authority can be determined even if the assigned non-volatile solid-state storage is unreachable. In some embodiments, if a particular authority 168 is unavailable, a replica or substitute authority 168 may be referenced.
[0063] 2A and 2B, two of the many tasks of the CPU 156 on a storage node 150 are to split write data and reassemble read data. When the system determines that data is to be written, the authority 168 for that data is located as described above. If the segment ID of the data has already been determined, the write request is forwarded from the segment to the non-volatile solid-state storage 152 currently determined to be the host of the determined authority 168. The host CPU 156 of the storage node 150 where the non-volatile solid-state storage 152 and corresponding authority 168 reside then splits or shards the data and transmits the data to the various non-volatile solid-state storages 152. The transmitted data is written as data stripes according to an erasure coding scheme. In some embodiments, data is requested to be pulled, while in other embodiments, data is pushed. Conversely, when data is read, the authority 168 for the segment ID containing the data is located as described above. The host CPU 156 of the storage node 150 where the non-volatile solid-state storage 152 and corresponding authority 168 reside requests data from the non-volatile solid-state storage and corresponding storage node pointed to by the authority. In some embodiments, the data is read from flash storage as a data stripe. The host CPU 156 of the storage node 150 then reassembles the read data, corrects any errors (if any) according to an appropriate erasure coding scheme, and transfers the reassembled data to the network. In further embodiments, some or all of these tasks may be processed in the non-volatile solid-state storage 152. In some embodiments, a segment host requests data to be sent to the storage node 150 by requesting a page from storage and then sending the data to the storage node that made the original request.
[0064] In an embodiment, authority 168 operates to determine how an operation proceeds for a particular logical element. Each logical element may be operated through a particular authority across multiple storage controllers of a storage system. Authority 168 may communicate with multiple storage controllers to cause the multiple storage controllers to collectively perform the operation for those particular logical elements.
[0065] In embodiments, a logical element may be, for example, a file, a directory, an object bucket, an individual object, a delimited portion of a file or object, some other form of key-value pair database, or a table. In embodiments, performing an operation may involve, for example, ensuring consistency, structural integrity, and / or recoverability with other operations on the same logical element, reading metadata and data associated with the logical element, determining what data should be durably written to the storage system to persist any changes due to the operation, or where it may be determined that metadata and data are stored across modular storage devices attached to multiple storage controllers in the storage system.
[0066] In some embodiments, operations are token-based transactions for efficient communication within a distributed system. Each transaction may be accompanied by or associated with a token that grants permission to perform the transaction. Authority 168, in some embodiments, may maintain the pre-transaction state of the system until the completion of the operation. Token-based communication can be achieved without global locks across the system and also allows for the resumption of operations in the event of an interruption or other failure.
[0067] In some systems, e.g., UNIX-style file systems, data is handled in index nodes or inodes, which specify data structures that represent objects within the file system. An object may be, for example, a file or a directory. Metadata may be associated with an object as attributes such as permission data and creation timestamps, among other attributes. A segment number may be assigned to all or part of such an object within the file system. In other systems, data segments are handled with segment numbers assigned elsewhere. For purposes of explanation, the unit of distribution is an entity, which may be a file, directory, or segment. That is, an entity is a unit of data or metadata stored by the storage system. Entities are grouped into sets called authorities. Each authority has an authority owner, which is a storage node that has exclusive rights to update the entities within the authority. In other words, storage nodes contain authorities, which in turn contain entities.
[0068] A segment is a logical container of data according to some embodiments. A segment is an address space between a media address space and a physical flash location; i.e., data segment numbers reside in this address space. A segment may also contain metadata that allows data redundancy to be restored (rewritten to a different flash location or device) without the involvement of higher-level software. In one embodiment, the internal format of a segment includes client data and a media mapping for determining the location of that data. Each data segment is protected against, for example, memory and other failures, by dividing the segment into multiple data and parity shards, if applicable. The data and parity shards are distributed, or striped, across the non-volatile solid-state storage 152 coupled to the host CPU 156 (see FIGS. 2E and 2G) according to an erasure coding scheme. The use of the term segment, in some embodiments, refers to the container and its location in the segment's address space. The use of the term stripe refers to the same set of shards as a segment, and, according to some embodiments, includes how the shards are distributed along with the redundancy or parity information.
[0069] A series of address space translations occurs throughout the storage system. At the top are directory entries (file names) that link to inodes. The inodes point to the media address space where data is logically stored. Media addresses can be mapped through a series of indirection media to distribute large file loads or implement data services such as deduplication or snapshots. Media addresses can be mapped through a series of indirection media to distribute large file loads or implement data services such as deduplication or snapshots. Next, segment addresses are translated to physical flash locations. According to some embodiments, physical flash locations have an address range that is limited by the amount of flash in the system. Media addresses and segment addresses are logical containers, and in some embodiments, use 128-bit or larger identifiers to be effectively infinite, with the potential for reuse calculated to be longer than the expected life of the system. In some embodiments, addresses from the logical containers are allocated hierarchically. Initially, each non-volatile solid-state storage 152 unit can be assigned a range of address space. Within this allocated range, non-volatile solid-state storage 152 can allocate addresses without synchronization with other non-volatile solid-state storage 152 .
[0070] Data and metadata are stored via a set of underlying storage layouts optimized for different workload patterns and storage devices. These layouts incorporate multiple redundancy schemes, compression formats, and indexing algorithms. Some of these layouts store information about authority and authority masters, while others store file metadata and file data. Redundancy schemes include error correction codes that tolerate corrupted bits within a single storage device (e.g., a NAND flash chip), erasure codes that tolerate failures of multiple storage nodes, and replication schemes that tolerate data center or regional failures. In some embodiments, low-density parity check ("LDPC") codes are used within a single storage unit. In some embodiments, Reed-Solomon coding is used within a storage cluster, and mirroring is used within a storage grid. Metadata may be stored using an ordered log-structured index (e.g., a log-structured merge tree), and large data may not be stored in a log-structured layout.
[0071] To maintain consistency across multiple copies of an entity, storage nodes implicitly agree on two things through computation: (1) the authorities that contain the entity, and (2) the storage nodes that contain the authorities. The assignment of entities to authorities can be done by pseudo-randomly assigning entities to authorities, by dividing entities into ranges based on an externally generated key, or by placing a single entity in each authority. Examples of pseudo-random schemes are the hash family of linear hashing and Replication Under Scalable Hashing ("RUSH"), including Controlled Replication Under Scalable Hashing ("CRUSH"). In some embodiments, pseudo-random assignment is utilized solely to assign authorities to nodes, since the set of nodes may change. Because the set of authorities cannot change, any subjective function may be applied in these embodiments. Some placement schemes automatically place authorities on storage nodes, while others rely on explicit mapping of authorities to storage nodes. In some embodiments, a pseudo-random scheme is utilized to map from each authority to a set of candidate authority holders. A pseudo-random data distribution function associated with CRUSH can assign authorities to storage nodes and create a list of where authorities are assigned. Each storage node has a copy of the pseudo-random data distribution function and can arrive at the same calculation for distribution and later discover or locate authorities. Each pseudo-random scheme, in some embodiments, requires a reachable set of storage nodes as input to conclude the same target node. Once entities are placed within authorities, they can be stored on physical devices such that expected failures do not lead to unexpected data loss. In some embodiments, the rebalancing algorithm attempts to store copies of all entities within an authority in the same layout on the same set of machines.
[0072] Examples of expected failures include device failure, stolen machinery, data center fire, and regional disasters such as nuclear or geological events. Different failures result in different levels of tolerable data loss. In some embodiments, a stolen storage node does not affect the security or reliability of the system, but depending on the system configuration, a regional event may result in no data loss, loss of a few seconds or minutes of updates, or even complete data loss.
[0073] In embodiments, the placement of data for storage redundancy is independent of the placement of authority for data consistency. In some embodiments, the storage nodes containing the authority do not contain any persistent storage. Instead, the storage nodes are connected to non-volatile solid-state storage units that do not contain authority. The communication interconnect between the storage nodes and the non-volatile solid-state storage units is comprised of multiple communication technologies and has non-uniform performance and fault-tolerance characteristics. In some embodiments, as described above, the non-volatile solid-state storage units are connected to the storage nodes via PCI Express, and the storage nodes are connected together within a single chassis using an Ethernet backplane, and the chassis are connected together to form a storage cluster. The storage cluster, in some embodiments, is connected to clients using Ethernet or Fibre Channel. When multiple storage clusters are configured into a storage grid, the multiple storage clusters are connected using the Internet or other long-distance networking links, such as "metro-scale" links or private links that do not traverse the Internet.
[0074] Authorities have exclusive rights to modify entities, migrate entities from one non-volatile solid-state storage unit to another, and add and remove copies of entities. This allows redundancy of the underlying data to be maintained. If an authority fails, is scheduled for decommissioning, or is overloaded, authority is transferred to a new storage node. In transient failures, it is important to ensure that all surviving machines agree on the new authority location. Ambiguity caused by transient failures can be resolved automatically through consensus protocols such as Paxos, hot-warm failover methods, manual intervention by a remote system administrator, or by a local hardware administrator (such as by physically removing the failed machine from the cluster or pressing a button on the failed machine). In some embodiments, a consensus protocol is used and failover is automatic. According to some embodiments, if too many failures or replication events occur within too short a period of time, the system enters a self-preservation mode, halting replication and data movement activities until an administrator intervenes.
[0075] As authorities are transferred between storage nodes and authority owners update entities within those authorities, the system transfers messages between the storage nodes and non-volatile solid-state storage units. Regarding persistent messages, messages with different purposes are of different types. Depending on the message type, the system maintains different ordering and durability guarantees. When persistent messages are being processed, the messages are temporarily stored in multiple durable and non-durable storage hardware technologies. In some embodiments, messages are stored in RAM, NVRAM, and NAND flash devices, and different protocols are used to efficiently use each storage medium. Latency-sensitive client requests may be persisted to replicated NVRAM and then later to NAND, while background rebalancing operations are persisted directly to NAND.
[0076] Persistent messages are stored persistently before being transmitted. This allows the system to continue servicing client requests despite failures and component replacement. Many hardware components contain unique identifiers that are visible to system administrators, manufacturers, the hardware supply chain, and an ongoing monitoring and quality control infrastructure; however, applications running on top of the infrastructure addresses virtualize the addresses. These virtualized addresses do not change over the life of the storage system, despite component failures and replacements. This allows components of the storage system to be replaced over time without reconfiguration or interruption of client request processing; i.e., the system supports non-disruptive upgrades.
[0077] In some embodiments, the virtualized addresses are stored with full redundancy. A continuous monitoring system correlates hardware and software status with hardware identifiers, allowing for the detection and prediction of failures due to defective components and manufacturing details. The monitoring system also, in some embodiments, allows for the proactive transfer of privileges and entities from affected devices before failures occur by removing components from the critical path.
[0078] FIG. 2C is a multilevel block diagram illustrating the contents of a storage node 150 and the contents of its non-volatile solid-state storage 152. Data is communicated to and from the storage node 150 by a network interface controller ("NIC") 202, in some embodiments. Each storage node 150 includes a CPU 156 and one or more non-volatile solid-state storage devices 152, as described above. Moving down one level in FIG. 2C, each non-volatile solid-state storage device 152 includes a non-volatile random access memory ("NVRAM") 204 and a relatively fast non-volatile solid-state memory such as flash memory 206. In some embodiments, the NVRAM 204 may be a component that does not require program / erase cycles (DRAM, MRAM, PCM) and may be memory that can support being written to much more frequently than the memory is read. Moving to another level in FIG. 2C, the NVRAM 204, in one embodiment, is implemented as a fast volatile memory such as dynamic random access memory (DRAM) 216 backed up by energy storage 218. Energy storage 218 provides sufficient power to continue powering DRAM 216 long enough for the contents to be transferred to flash memory 206 in the event of a power failure. In some embodiments, energy storage 218 is a capacitor, supercapacitor, battery, or other device that provides a suitable supply of energy sufficient to allow the transfer of the contents of DRAM 216 to a stable storage medium in the event of a power loss. Flash memory 206 is implemented as multiple flash dies 222, sometimes referred to as a package of flash dies 222 or an array of flash dies 222. It should be understood that flash dies 222 may be packaged in any number of ways, such as a single die per package, multiple dies per package (i.e., a multi-chip package), a hybrid package, bare dies on a circuit printed board or other substrate, encapsulated dies, etc.In the illustrated embodiment, non-volatile solid-state storage 152 includes a controller 212 or other processor and input / output (I / O) ports 210 coupled to controller 212. I / O ports 210 are coupled to CPU 156 and / or network interface controller 202 of flash storage node 150. Flash input / output ports 220 are coupled to flash dies 222, and direct memory access (DMA) units 214 are coupled to controller 212, DRAM 216, and flash dies 222. In the illustrated embodiment, I / O ports 210, controller 212, DMA units 214, and flash I / O ports 220 are implemented on a programmable logic device (“PLD”) 208, e.g., an FPGA. In this embodiment, each flash die 222 includes pages organized as 16 kB (kilobyte) pages 224 and registers 226 that can write data to or read data from flash dies 222. In further embodiments, other types of solid-state memory are used in place of or in addition to the flash memory illustrated in flash die 222.
[0079] A storage cluster 161, in various embodiments as disclosed herein, can be generally contrasted with a storage array. Storage nodes 150 are part of a collection that makes up the storage cluster 161. Each storage node 150 owns a slice of data and the computing necessary to serve the data. Multiple storage nodes 150 cooperate to store and retrieve data. Storage memory or devices, as generally used in storage arrays, are not significantly involved in processing and manipulating data. Storage memory or devices in a storage array receive commands to read, write, or erase data. Storage memory or devices in a storage array are unaware of the larger system in which they are embedded or what the data represents. Storage memory or devices in a storage array may include various types of storage memory, such as RAM, solid-state drives, and hard disk drives. The non-volatile solid-state storage 152 units described herein have multiple interfaces that are simultaneously active and serve multiple purposes. In some embodiments, some of the functionality of a storage node 150 is shifted to the storage unit 152, transforming the storage unit 152 into a combination of storage unit 152 and storage node 150. Placing computing (over storage data) in the storage unit 152 places this computing closer to the data itself. Various system embodiments have a hierarchy of storage node tiers with different capabilities. In contrast, in a storage array, the controller owns and knows everything about all the data it manages in the shelves or storage devices. In a storage cluster 161, multiple non-volatile solid-state storage 152 units and / or multiple controllers in storage nodes 150 cooperate in various ways (e.g., for erasure coding, data sharding, metadata communication and redundancy, expanding or shrinking storage capacity, data recovery, etc.) as described herein.
[0080] FIG. 2D illustrates a storage server environment using the storage node 150 and storage 152 unit embodiments of FIGS. 2A-2C. In this version, each non-volatile solid-state storage 152 unit includes a processor, such as a controller 212 (see FIG. 2C), an FPGA, flash memory 206, and NVRAM 204 (supercapacitor-backed DRAM 216, see FIGS. 2B and 2C) on a PCIe (Peripheral Component Interconnect Express) board within the chassis 138 (see FIG. 2A). The non-volatile solid-state storage 152 units may be implemented as a single board containing storage, which may be the maximum tolerable failure domain within the chassis. In some embodiments, up to two non-volatile solid-state storage 152 units may fail, and the device continues without data loss.
[0081] In some embodiments, physical storage is divided into named regions based on application usage. NVRAM 204 is a contiguous block of reserved memory within nonvolatile solid-state storage 152, DRAM 216, and backed by NAND flash. NVRAM 204 is logically divided into multiple memory regions, two of which are written as spools (e.g., spool_regions). Space within the NVRAM 204 spool is managed independently by each authority 168. Each device provides a certain amount of storage space to each authority 168, which then manages the lifetime and allocation within that space. Examples of spools include distributed transactions or concepts. When primary power to the nonvolatile solid-state storage 152 unit fails, an onboard supercapacitor provides a short duration of power holdup. During this holdup interval, the contents of NVRAM 204 are flushed to flash memory 206. Upon the next power-up, the contents of NVRAM 204 are restored from flash memory 206.
[0082] With respect to the storage unit controller, the logical "controller" responsibilities are distributed across each of the blades, including authority 168. This distribution of logical control is illustrated in FIG. 2D as host controller 242, mid-tier controller 244, and storage unit controller 246. Control plane and storage plane management are handled independently, although some may be physically co-located on the same blade. Each authority 168 effectively functions as an independent controller. Each authority 168 provides its own data and metadata structures, its own background workers, and maintains its own lifecycle.
[0083] FIG. 2E is a hardware block diagram of a blade 252, using the embodiment of the storage node 150 and storage unit 152 of FIGS. 2A-2C in the storage server environment of FIG. 2D , illustrating a control plane 254, a compute plane 256, and a storage plane 258 that interact with the underlying physical resources, and an authority 168. The control plane 254 is divided into multiple authorities 168 that can run on any of the blades 252 using the computational resources in the compute plane 256. The storage plane 258 is divided into a set of devices that each provide access to the flash 206 and NVRAM 204 resources. In one embodiment, the compute plane 256 can perform the operations of a storage array controller on one or more devices of the storage plane 258 (e.g., a storage array) as described herein.
[0084] In the compute plane 256 and storage plane 258 of FIG. 2E, authorities 168 interact with the underlying physical resources (i.e., devices). From the perspective of an authority 168, its resources are striped across all of the physical devices. From the device's perspective, the device provides resources to all authorities 168, regardless of where the authorities happen to run. Each authority 168 has allocated or is allocated one or more partitions 260 of storage memory in the storage units 152, e.g., partitions 260 in the flash memory 206 and NVRAM 204. Each authority 168 uses its allocated partitions 260 to write or read user data. Authorities can be associated with different amounts of physical storage in the system. For example, one authority 168 can have more partitions 260 or larger-sized partitions 260 in one or more storage units 152 than one or more other authorities 168.
[0085] FIG. 2F illustrates elasticity software layers within a blade 252 of a storage cluster, according to some embodiments. In an elastic architecture, the elasticity software is symmetric; that is, each blade's compute module 270 executes the three identical layers of processes depicted in FIG. 2F. A storage manager 274 executes read and write requests from other blades 252 to data and metadata stored in the local storage unit 152, NVRAM 204, and flash 206. An authority 168 fulfills client requests by issuing the necessary reads and writes to the blade 252 on the storage unit 152 where the corresponding data or metadata resides. An endpoint 272 analyzes client connection requests received from the monitoring software of the switch fabric 146, relays the client connection request to the authority 168 responsible for fulfillment, and relays the authority's 168 response to the client. The symmetric three-tier architecture enables a high degree of concurrency in the storage system. Elasticity scales out efficiently and reliably in these embodiments. Additionally, Elasticity implements inherent scale-out techniques that maximize concurrency by balancing work evenly across all resources regardless of client access patterns, eliminating much of the need for cross-blade coordination that typically occurs with traditional distributed locking.
[0086] 2F , authorities 168 executing within compute modules 270 of blades 252 perform the internal operations necessary to fulfill client requests. One feature of resiliency is that authorities 168 are stateless, i.e., they cache active data and metadata in their own blade's 252 DRAM for fast access, but they store all updates in their NVRAM 204 partitions on three separate blades 252 until the updates are written to flash 206. In some embodiments, all storage system writes to NVRAM 204 are triplicate across partitions on three separate blades 252. With triple-mirrored NVRAM 204 and persistent storage protected by parity and Reed-Solomon RAID checksums, the storage system can tolerate the simultaneous failure of two blades 252 without losing data, metadata, or access to either.
[0087] Because authorities 168 are stateless, they can migrate between blades 252. Each authority 168 has a unique identifier. NVRAM 204 and flash 206 partitions are associated with the identifier of the authority 168, not the blade 252 on which they are running. Thus, when an authority 168 migrates, the authority 168 continues to manage the same storage partitions from its new location. When a new blade 252 is installed in one embodiment of a storage cluster, the system automatically rebalances the load by partitioning the new blade's 252's storage for use by authorities 168 in the system, migrating selected authorities 168 to the new blade 252, and starting endpoints 272 on the new blade 252 and including them in the switch fabric 146's client connection distribution algorithm.
[0088] From their new locations, the migrated authorities 168 persist the contents of their NVRAM 204 partitions on flash 206, process read and write requests from other authorities 168, and fulfill client requests that endpoints 272 direct to them. Similarly, if a blade 252 fails or is removed, the system redistributes its authorities 168 among the remaining blades 252 in the system. The redistributed authorities 168 continue to perform their original functions from their new locations.
[0089] FIG. 2G depicts authorities 168 and storage resources within blades 252 of a storage cluster, according to some embodiments. Each authority 168 is exclusively responsible for a partition of flash 206 and NVRAM 204 on each blade 252. Authorities 168 manage the contents and integrity of their partitions independently of other authorities 168. Authorities 168 compress incoming data, temporarily store it in their NVRAM 204 partitions, and then consolidate, RAID-protect, and persist the data in segments of storage in their flash 206 partitions. As authorities 168 write data to flash 206, storage manager 274 performs the necessary flash transformations to optimize write performance and maximize media lifespan. In the background, authorities 168 “garbage collect,” i.e., reclaim space occupied by data no longer needed by clients overwriting data. It should be appreciated that because the partitions of authorities 168 are disjoint, there is no need for distributed locking to execute clients and writes or to execute background functions.
[0090] The embodiments described herein may utilize various software, communication, and / or networking protocols. In addition, hardware and / or software configurations may be adjusted to accommodate various protocols. For example, embodiments may utilize Active Directory, a database-based system that provides authentication, directory, policy, and other services in a WINDOWS™ environment. In these embodiments, the Lightweight Directory Access Protocol (LDAP) is an example of an application protocol for querying and modifying entries in a directory service provider such as Active Directory. In some embodiments, a network lock manager ("NLM") is utilized in conjunction with the Network File System ("NFS") to provide System V-style advisory file and record locking over a network. The Server Message Block ("SMB") protocol, one version of which is also known as the Common Internet File System ("CIFS"), may be integrated with the storage systems described herein. SMP operates as an application-layer network protocol typically used to provide shared access to files, printers, and serial ports, as well as various communications between nodes on a network. SMB also provides an authenticated inter-process communication mechanism. AMAZON™ S3 (Simple Storage Service) is a web service provided by Amazon Web Services, and the system described herein can interface with Amazon S3 via web service interfaces (REST (Representational State Transfer), SOAP (Simple Object Access Protocol), and BitTorrent).A RESTful API (application programming interface) breaks down transactions into a series of small modules, each addressing a specific underlying part of the transaction. The control or permissions provided in these embodiments, particularly for object data, may include the use of access control lists ("ACLs"). An ACL is a list of permissions attached to an object, specifying which users or system processes are allowed to access the object and which actions are allowed for a given object. The system provides an identification and location system for computers on the network and may utilize Internet Protocol version 6 ("IPv6") as well as IPv4 for communication protocols that route traffic through the Internet. Routing of packets between networked systems may include equal-cost multi-path routing ("ECMP"), a routing strategy in which next-hop packet forwarding to a single destination may occur over multiple "best paths" that combine on top of a routing metric calculation. Multi-path routing can be used with most routing protocols because it is a hop-by-hop decision limited to a single router. The software may support multitenancy, an architecture in which a single instance of a software application serves multiple customers. Each customer may be referred to as a tenant. Tenants may be given the ability to customize some parts of the application, but in some embodiments may not customize the application's code. Embodiments may maintain audit logs. An audit log is a document that records events in a computing system.In addition to documenting which resources were accessed, audit log entries typically include destination and source addresses, timestamps, and user login information for compliance with various regulations. Embodiments can support various key management policies, such as encryption key rotation. Additionally, the system can support dynamic root passwords or some variation that dynamically changes passwords.
[0091] FIG. 3A illustrates a diagram of a storage system 306 coupled for data communication with a cloud service provider 302, according to some embodiments of the present disclosure. While not depicted in greater detail, the storage system 306 depicted in FIG. 3A may be similar to the storage systems described above with reference to FIGS. 1A-1D and 2A-2G. In some embodiments, the storage system 306 depicted in FIG. 3A may be embodied as a storage system including unbalanced active / active controllers, a storage system including balanced active / active controllers, a storage system including active / active controllers in which fewer than all of each controller's resources are utilized such that each controller has spare resources that can be used to support failover, a storage system including fully active / active controllers, a storage system including controllers with separated data sets, a storage system including a dual-tier architecture with a front-end controller and a back-end unified storage controller, a storage system including a scale-out cluster of dual-controller arrays, and combinations of such embodiments.
[0092] 3A , storage system 306 is coupled to cloud service provider 302 via data communications link 304. Data communications link 304 may be embodied as a dedicated data communications link, as a data communications path provided through the use of one or more data communications networks, such as a wide area network ("WAN") or LAN, or as some other mechanism capable of transferring digital information between storage system 306 and cloud service provider 302. Such data communications link 304 may be entirely wired, entirely wireless, or some collection of wired and wireless data communications paths. In such an example, digital information may be exchanged between storage system 306 and cloud service provider 302 via data communications link 304 using one or more data communications protocols. For example, digital information may be exchanged between storage system 306 and cloud service provider 302 via data communications link 304 using Handheld Device Transfer Protocol ("HDTP"), Hypertext Transfer Protocol ("HTTP"), Internet Protocol ("IP"), Real-Time Transport Protocol ("RTP"), Transmission Control Protocol ("TCP"), User Datagram Protocol ("UDP"), Wireless Application Protocol ("WAP"), or other protocols.
[0093] The cloud service provider 302 depicted in FIG. 3A may be embodied as a system and computing environment that provides a vast array of services to users of the cloud service provider 302, for example, through the sharing of computing resources over a data communication link 304. The cloud service provider 302 can provide on-demand access to a shared pool of configurable computing resources, such as computer networks, servers, storage, applications, and services. The shared pool of configurable resources can be rapidly provisioned and released to users of the cloud service provider 302 with minimal administrative effort. Generally, users of the cloud service provider 302 are unaware of the exact computing resources utilized by the cloud service provider 302 to provide services. While such cloud service providers 302 may often be accessible via the Internet, readers skilled in the art will recognize that any system that abstracts the use of shared resources to provide services to users over any data communication link may be considered a cloud service provider 302.
[0094] 3A, cloud service provider 302 may be configured to provide various services to storage system 306 and users of storage system 306 through the implementation of various service models. For example, cloud service provider 302 may be configured to provide services through an implementation of an infrastructure as a service ("IaaS") service model, through an implementation of a platform as a service ("PaaS") service model, through an implementation of a software as a service ("SaaS") service model, through an implementation of an authentication as a service ("AaaS") service model, or through an implementation of a storage as a service model in which cloud service provider 302 provides access to its storage infrastructure for use by storage system 306 and users of storage system 306. The reader will understand that the above-described service models are included for illustrative purposes only and do not represent limitations on the services that may be provided by cloud service provider 302 or on the service models that may be implemented by cloud service provider 302, and that cloud service provider 302 may be configured to provide additional services to storage system 306 and users of storage system 306 through the implementation of additional service models.
[0095] 3A , cloud service provider 302 may be embodied as, for example, a private cloud, a public cloud, or a combination of private and public clouds. In an embodiment in which cloud service provider 302 is embodied as a private cloud, cloud service provider 302 may be dedicated to serving a single organization rather than serving multiple organizations. In an embodiment in which cloud service provider 302 is embodied as a public cloud, cloud service provider 302 may provide services to multiple organizations. In yet another embodiment, cloud service provider 302 may be embodied as a mix of private and public cloud services, comprising a hybrid cloud deployment.
[0096] Although not explicitly depicted in FIG. 3A , the reader will understand that a significant amount of additional hardware and software components may be required to facilitate delivery of cloud services to storage system 306 and users of storage system 306. For example, storage system 306 may be coupled to (or include) a cloud storage gateway. Such a cloud storage gateway may be embodied, for example, as a hardware- or software-based appliance located on-premises with storage system 306. Such a cloud storage gateway can act as a bridge between local applications running on storage system 306 and remote, cloud-based storage utilized by storage system 306. Through the use of a cloud storage gateway, an organization may be able to move its primary iSCSI or NAS storage to cloud service provider 302, thereby enabling the organization to conserve space on its on-premises storage system. Such a cloud storage gateway may be configured to emulate a disk array, block-based device, file server, or other storage system that can translate SCSI commands, file server commands, or other appropriate commands into a RESTful space protocol that facilitates communication with cloud service provider 302.
[0097] To enable storage system 306 and users of storage system 306 to utilize services offered by cloud service provider 302, a cloud migration process may be performed, during which data, applications, or other elements from an organization's local system (or from another cloud environment) are moved to cloud service provider 302. To successfully migrate data, applications, or other elements to the cloud service provider's 302 environment, middleware such as a cloud migration tool may be utilized to bridge the gap between the cloud service provider's 302 environment and the organization's environment. Such cloud migration tools may also be configured to address potentially high network costs and long transfer times associated with migrating large amounts of data to cloud service provider 302, as well as security issues associated with transmitting sensitive data over a data communications network to cloud service provider 302. To further enable storage system 306 and users of storage system 306 to utilize services offered by cloud service provider 302, a cloud orchestrator may also be used to arrange and coordinate automated tasks in pursuit of creating an integrated process or workflow. Such a cloud orchestrator can perform tasks such as configuring various components, whether they are cloud or on-premise components, and managing the interconnections between such components. The cloud orchestrator can simplify inter-component communication and connections to ensure that links are properly configured and maintained.
[0098] In the example depicted in FIG. 3A , as briefly described above, cloud service provider 302 may be configured to provide services to storage system 306 and users of storage system 306 through the use of a SaaS service model, eliminating the need to install and run applications on local computers and simplifying application maintenance and support. Such applications can take many forms in accordance with various embodiments of the present disclosure. For example, cloud service provider 302 may be configured to provide storage system 306 and users of storage system 306 with access to a data analysis application. Such data analysis application may be configured to receive, for example, vast amounts of telemetry data transmitted by storage system 306 to their homes. Such telemetry data may describe various operational characteristics of storage system 306 and can be analyzed for a myriad of purposes, including, for example, determining the health of storage system 306, identifying workloads running on storage system 306, predicting when storage system 306 will run out of various resources, and recommending configuration changes, hardware or software upgrades, workflow transitions, or other actions that can improve the operation of storage system 306.
[0099] Cloud service provider 302 may also be configured to provide access to virtualized computing environments to storage system 306 and users of storage system 306. Such virtualized computing environments may be embodied, for example, as virtual machines or other virtualized computer hardware platforms, virtual storage devices, virtualized computer network resources, etc. Examples of such virtualized environments may include virtual machines created to emulate real computers, virtualized desktop environments that separate logical desktops from physical machines, virtualized file systems that allow uniform access to different types of concrete file systems, etc.
[0100] While the example depicted in FIG. 3A illustrates storage system 306 being coupled for data communication with cloud service provider 302, in other embodiments, storage system 306 may be part of a public cloud deployment in which hybrid cloud elements (e.g., private cloud services, on-premises infrastructure, etc.) and public cloud elements (e.g., private cloud services, infrastructure, etc., that may be provided by one or more cloud service providers) are combined to form a single solution through orchestration between various platforms. Such hybrid cloud deployments may leverage hybrid cloud management software, such as Microsoft™'s Azure™ Arc, which centralizes management of the hybrid cloud deployment to any infrastructure and enables deployment of services anywhere. In such an example, the hybrid cloud management software may be configured to create, update, and delete resources (both physical and virtual) that form the hybrid cloud deployment, allocate compute and storage to specific workloads, monitor workloads and resources for performance, policy compliance, updates and patches, security status, or perform various other tasks.
[0101] The reader will understand that pairing the storage systems described herein with one or more cloud service providers can enable a variety of offerings. For example, disaster recovery as a service ("DRaaS") can be provided, in which cloud resources are utilized to protect applications and data from disruptions caused by disasters, including embodiments in which the storage system can serve as a primary data store. In such embodiments, full system backups can be taken to enable business continuity in the event of a system failure. In such embodiments, cloud data backup technology (by itself or as part of a larger DRaaS solution) can also be integrated into an overall solution that includes the storage systems and cloud service providers described herein.
[0102] The storage systems described herein, as well as cloud service providers, can be utilized to provide a variety of security features. For example, the storage system can encrypt data at rest (encrypted data can be sent to and from the storage system) and can utilize Key Management-as-a-Service ("KMaaS") to manage encryption keys, keys for locking and unlocking storage devices, and the like. Similarly, a cloud data security gateway or similar mechanism can be utilized to ensure that data stored within the storage system is not improperly stored in the cloud as part of a cloud data backup operation. Furthermore, microsegmentation or identity-based segmentation can be utilized in data centers that include the storage systems or within cloud service providers to create secure zones that allow workloads to be isolated from one another in data center and cloud deployments.
[0103] For further explanation, Figure 3B sets forth a diagram of a storage system 306 according to some embodiments of the present disclosure. Although not depicted in greater detail, the storage system 306 depicted in Figure 3B may be similar to the storage systems described above with reference to Figures 1A-1D and 2A-2G, as the storage systems may include many of the components described above.
[0104] 3B may include a vast amount of storage resources 308, which may be embodied in many forms. For example, the storage resources 308 may include nanoRAM or another form of nonvolatile random-access memory utilizing carbon nanotubes deposited on a substrate, 3D cross-point nonvolatile memory, flash memory, including single-level cell ("SLC") NAND flash, multi-level cell ("MLC") NAND flash, triple-level cell ("TLC") NAND flash, quad-level cell ("QLC") NAND flash, or others. Similarly, the storage resources 308 may include nonvolatile magnetoresistive random-access memory ("MRAM"), including spin transfer torque ("STT") MRAM. Exemplary storage resource 308 may alternatively include other forms of storage resources, including non-volatile phase-change memory ("PCM"), quantum memory that enables the storage and retrieval of photonic quantum information, resistive random-access memory ("ReRAM"), storage class memory ("SCM"), or any combination of the resources described herein. The reader will understand that other forms of computer memory and storage devices may be utilized by the above-described storage system, including DRAM, SRAM, EEPROM, universal memory, etc.The storage resources 308 depicted in FIG. 3A may be embodied in a variety of form factors, including, but not limited to, dual in-line memory modules ("DIMMs"), non-volatile dual in-line memory modules ("NVDIMMs"), M.2, U.2, and others.
[0105] The storage resources 308 depicted in FIG. 3B may include various forms of SCM. SCM can effectively treat high-speed, non-volatile memory (e.g., NAND flash) as an extension of DRAM, so that the entire data set can be treated as an in-memory data set residing entirely within DRAM. SCM can include, for example, non-volatile media such as NAND flash. Such NAND flash may be accessed using NVMe, which can use the PCIe bus as its transport, offering relatively low access latency compared to older protocols. Indeed, network protocols used for SSDs in all-flash arrays include NVMe over Ethernet (ROCE, NVMe TCP), Fibre Channel (NVMe FC), InfiniBand (iWARP), and others that enable treating high-speed, non-volatile memory as an extension of DRAM. Given the fact that DRAM is often byte-addressable and high-speed, non-volatile memory such as NAND flash is block-addressable, a controller software / hardware stack may be required to convert block data into bytes stored on the media. Examples of media and software that can be used as SCM can include, for example, 3D XPoint, Intel Memory Drive Technology, Samsung's Z-SSD, and others.
[0106] The storage resource 308 depicted in FIG. 3B can also include racetrack memory (also referred to as domain wall memory). Such racetrack memory can be embodied in a solid-state device as a form of nonvolatile solid-state memory that relies on the charge of electrons as well as the unique strength and orientation of the magnetic field generated by electrons as they spin. By using spin-coherent currents to move magnetic domains along nanoscale Permalloy wires, the magnetic domains can pass by a magnetic read / write head positioned near the wire as the current passes through the wire, thereby altering the magnetic domains and recording a pattern of bits. Many such wires and read / write elements can be packaged together to create a racetrack memory device.
[0107] The exemplary storage system 306 depicted in FIG. 3B can implement a variety of storage architectures. For example, a storage system according to some embodiments of the present disclosure may utilize block storage, where data is stored in blocks, with each block essentially acting as an individual hard drive. A storage system according to some embodiments of the present disclosure may utilize object storage, where data is managed as objects. Each object may include the data itself, a variable amount of metadata, and a globally unique identifier, and object storage may be implemented at multiple levels (e.g., device level, system level, interface level). A storage system according to some embodiments of the present disclosure utilizes file storage, where data is stored in a hierarchical structure. Such data is stored in files and folders and can be presented in the same format to both the system that stores it and the system that retrieves it.
[0108] 3B may be embodied as a storage system in which additional storage resources can be added through the use of a scale-up model, a scale-out model, or some combination thereof. In a scale-up model, additional storage may be added by adding additional storage devices. However, in a scale-out model, additional storage nodes may be added to a cluster of storage nodes, and such storage nodes may include additional processing resources, additional networking resources, etc.
[0109] 3B may utilize the storage resources described above in a variety of different ways. For example, portions of the storage resources may be utilized to function as a write cache, storage resources within the storage system may be utilized as a read cache, or tiering may be achieved within the storage system by placing data within the storage system according to one or more tiering policies.
[0110] 3B also includes communication resources 310 that may be useful for facilitating data communication between components within the storage system 306, as well as between the storage system 306 and computing devices external to the storage system 306, including embodiments in which these resources are separated by relatively large expanses. The communication resources 310 may be configured to utilize a variety of different protocols and data communication fabrics to facilitate data communication between components within the storage system and computing devices external to the storage system. For example, communication resources 310 may include Fibre Channel ("FC") technology, such as an FC fabric and FC protocol capable of transporting SCSI commands over an FC network, FC over Ethernet ("FCoE") technology, in which FC frames are encapsulated and transmitted over an Ethernet network, InfiniBand ("IB") technology, in which a switched fabric topology is utilized to facilitate transmission between channel adapters, NVM Express ("NVMe") technology and NVMe over fabric ("NVMeoF") technology, in which non-volatile storage media attached via a PCI Express ("PCIe") bus can be accessed, and others. Indeed, the storage systems described above may directly or indirectly utilize neutrino communication technologies and devices in which information (including binary information) is transmitted using beams of neutrinos.
[0111] The communication resources 310 may also include mechanisms for accessing the storage resources 308 in the storage system 306 using serial attached SCSI ("SAS"), serial ATA ("SATA") bus interfaces for connecting the storage resources 308 in the storage system 306 to host bus adapters in the storage system 306, Internet Small Computer System Interface ("iSCSI") technology for providing block-level access to the storage resources 308 in the storage system 306, and other communication resources that may be useful in facilitating data communication between components within the storage system 306, as well as data communication between the storage system 306 and computing devices external to the storage system 306.
[0112] 3B also includes processing resources 312 that may be useful for executing computer program instructions and performing other computational tasks within the storage system 306. The processing resources 312 may include one or more ASICs customized for any particular purpose, as well as one or more CPUs. The processing resources 312 may also include one or more DSPs, one or more FPGAs, one or more systems on a chip ("SoC"), or other forms of processing resources 312. The storage system 306 may utilize the storage resources 312 to perform various tasks, including, but not limited to, supporting the execution of software resources 314, which are described in more detail below.
[0113] 3B also includes software resources 314 that, when executed by processing resources 312 within storage system 306, can perform a vast number of tasks. Software resources 314 may include, for example, one or more modules of computer program instructions that, when executed by processing resources 312 within storage system 306, are useful for implementing various data protection techniques. Such data protection techniques may be implemented, for example, by system software running on computer hardware within the storage system, by a cloud service provider, or otherwise. Such data protection techniques may include data archiving, data backup, data replication, data snapshots, data and database cloning, and other data protection techniques.
[0114] Software resources 314 may also include software that is useful in implementing software-defined storage ("SDS"). In such an example, software resources 314 may include one or more modules of computer program instructions that, when executed, are useful in policy-based provisioning and management of data storage independent of the underlying hardware. Such software resources 314 may be useful in implementing storage virtualization to separate storage hardware from the software that manages the storage hardware.
[0115] The software resources 314 may also include software useful for facilitating and optimizing I / O operations directed to the storage system 306. For example, the software resources 314 may include software modules that implement various data reduction techniques, such as data compression, data deduplication, and others. The software resources 314 may include software modules that intelligently group I / O operations to facilitate better use of the underlying storage resources 308, software modules that perform data migration operations for migrating data from within the storage system, and software modules that perform other functions. Such software resources 314 may be embodied as one or more software containers or in many other ways.
[0116] For further explanation, Figure 3C sets forth an example of a cloud-based storage system 318 according to some embodiments of the present disclosure. In the example depicted in Figure 3C, the cloud-based storage system 318 is created entirely within a cloud computing environment 316, such as, for example, Amazon Web Services ("AWS")™, Microsoft Azure™, Google Cloud Platform™, IBM Cloud™, Oracle Cloud™, and others. The cloud-based storage system 318 may be used to provide services similar to those that may be provided by the storage systems described above.
[0117] The cloud-based storage system 318 depicted in FIG. 3C includes two cloud computing instances 320, 322, each used to support the execution of storage controller applications 324, 326. The cloud computing instances 320, 322 may be embodied as instances of cloud computing resources (e.g., virtual machines) that may be provided by the cloud computing environment 316 to support the execution of software applications such as the storage controller applications 324, 326. For example, each of the cloud computing instances 320, 322 may run on an Azure VM, and each Azure VM may include high-speed temporary storage that can be utilized as a cache (e.g., as a read cache). In one embodiment, the cloud computing instances 320, 322 may be embodied as Amazon Elastic Compute Cloud ("EC2") instances. In such an example, an Amazon Machine Image ("AMI") that includes the storage controller applications 324, 326 may be booted to create and configure a virtual machine capable of running the storage controller applications 324, 326.
[0118] 3C , the storage controller applications 324, 326 may be embodied as modules of computer program instructions that, when executed, perform various storage tasks. For example, the storage controller applications 324, 326 may be embodied as modules of computer program instructions that, when executed, perform the same tasks as the controllers 110A, 110B of FIG. 1A described above, such as writing data to the cloud-based storage system 318, erasing data from the cloud-based storage system 318, retrieving data from the cloud-based storage system 318, monitoring and reporting disk usage and performance, performing redundancy operations such as RAID or RAID-like data redundancy operations, compressing data, encrypting data, deduplication data, etc. The reader will understand that, because there are two cloud computing instances 320, 322, each including a storage controller application 324, 326, in some embodiments, one cloud computing instance 320 can operate as a primary controller as described above, and the other cloud computing instance 322 can operate as a secondary controller as described above. The reader will understand that the storage controller applications 324, 326 depicted in FIG. 3C may comprise the same source code running within different cloud computing instances 320, 322, such as separate EC2 instances.
[0119] The reader will understand that other embodiments that do not include primary and secondary controllers are within the scope of this disclosure. For example, each cloud computing instance 320, 322 can act as a primary controller for some portion of the address space supported by the cloud-based storage system 318, each cloud computing instance 320, 322 can act as a primary controller where servicing of I / O operations directed to the cloud-based storage system 318 is divided in some other manner, and so on. Indeed, in other embodiments where cost savings may take priority over performance requirements, there may be only a single cloud computing instance that includes the storage controller application.
[0120] The cloud-based storage system 318 depicted in Figure 3C includes cloud computing instances 340a, 340b, 340n having local storage 330, 334, 338. The cloud computing instances 340a, 340b, 340n may be embodied as instances of cloud computing resources that may be provided by the cloud computing environment 316 to support the execution of software applications, for example. The cloud computing instances 340a, 340b, 340n of Figure 3C may differ from the cloud computing instances 320, 322 described above because the cloud computing instances 340a, 340b, 340n of Figure 3C have local storage 330, 334, 338 resources, whereas the cloud computing instances 320, 322 that support the execution of storage controller applications 324, 326 need not have local storage resources. Cloud computing instances 340a, 340b, 340n having local storage 330, 334, 338 may be embodied, for example, as an EC2 M5 instance including one or more SSDs, as an EC2 R5 instance including one or more SSDs, as an EC2 I3 instance including one or more SSDs, etc. In some embodiments, local storage 330, 334, 338 must be embodied as solid-state storage (e.g., SSDs) rather than storage that utilizes hard disk drives.
[0121] 3C , each of the cloud computing instances 340 a, 340 b, 340 n with local storage 330, 334, 338 may include a software daemon 328, 332, 336 that, when executed by the cloud computing instances 340 a, 340 b, 340 n, can present itself to the storage controller application 324, 326 as if the cloud computing instance 340 a, 340 b, 340 n were a physical storage device (e.g., one or more SSDs). In such an example, the software daemon 328, 332, 336 may include computer program instructions similar to those that would typically be included on a storage device, such that the storage controller application 324, 326 can send and receive the same commands that a storage controller would send to a storage device. In this manner, the storage controller application 324, 326 may include code that is the same (or substantially the same) as the code executed by the controller in the storage system described above. In these and similar embodiments, communication between the storage controller applications 324, 326 and the cloud computing instances 340a, 340b, 340n with the local storage 330, 334, 338 may utilize iSCSI, NVMe over TCP, messaging, a custom protocol, or some other mechanism.
[0122] 3C , each of the cloud computing instances 340a, 340b, 340n with local storage 330, 334, 338 may also be coupled to block storage 342, 344, 346 provided by the cloud computing environment 316, such as, for example, Amazon Elastic Bookstore (Elastic Block Store, “EBS”) volumes. In such an example, the block storage 342, 344, 346 provided by the cloud computing environment 316 may be utilized in a manner similar to the way NVRAM devices described above are utilized, such that when a software daemon 328, 332, 336 (or some other module) running within a particular cloud computing instance 340a, 340b, 340n receives a request to write data, it may initiate writing data to its attached EBS volume as well as to its local storage 330, 334, 338 resource. In some alternative embodiments, data may be written only to local storage 330, 334, 338 resources within the particular cloud containing the instance 340 a, 340 b, 340 n. In alternative embodiments, rather than using block storage 342, 344, 346 provided by the cloud computing environment 316 as NVRAM, actual RAM on each of the cloud computing instances 340 a, 340 b, 340 n having local storage 330, 334, 338 may be used as NVRAM, thereby reducing network utilization costs associated with using EBS volumes as NVRAM. In yet another embodiment, high-performance block storage resources such as one or more Azure Ultra Disks may be utilized as NVRAM.
[0123] Storage controller applications 324, 326 may be used to perform various tasks such as deduplicating the data included in the request, compressing the data included in the request, determining where to write the data included in the request, and then ultimately sending a request to write a deduplicated, encrypted, or possibly updated version of the data to one or more of cloud computing instances 340a, 340b, 340n having local storage 330, 334, 338. Either of cloud computing instances 320, 322, in some embodiments, may receive a request to read data from cloud-based storage system 318 and ultimately send the request to read the data to one or more of cloud computing instances 340a, 340b, 340n having local storage 330, 334, 338.
[0124] When a request to write data is received by a particular cloud computing instance 340a, 340b, 340n having local storage 330, 334, 338, the software daemons 328, 332, 336 may be configured not only to write the data to their own local storage 330, 334, 338 resources and any suitable block storage 342, 344, 346 resources, but the software daemons 328, 332, 336 may also be configured to write the data to cloud-based object storage 348 attached to the particular cloud computing instance 340a, 340b, 340n. The cloud-based object storage 348 attached to the particular cloud computing instance 340a, 340b, 340n may be embodied as, for example, Amazon Simple Storage Service ("S3"). In other embodiments, cloud computing instances 320, 322, including storage controller applications 324, 326, respectively, can initiate storage of data to local storage 330, 334, 338 of cloud computing instances 340a, 340b, 340n and cloud-based object storage 348. In other embodiments, rather than storing data using both cloud computing instances 340a, 340b, 340n with local storage 330, 334, 338 (also referred to herein as "virtual drives") and cloud-based object storage 348, the persistent storage tier may be implemented in other ways. For example, one or more Azure Ultra disks may be used to persistently store data (e.g., after the data is written to the NVRAM tier).
[0125] While the local storage 330, 334, 338 and block storage 342, 344, 346 resources utilized by cloud computing instances 340a, 340b, 340n may support block-level access, the cloud-based object storage 348 attached to a particular cloud computing instance 340a, 340b, 340n only supports object-based access. Thus, software daemons 328, 332, 336 may be configured to retrieve blocks of data, package those blocks into objects, and write the objects to the cloud-based object storage 348 attached to a particular cloud computing instance 340a, 340b, 340n.
[0126] Consider an example in which data is written in 1 MB blocks to the local storage 330, 334, 338 and block storage 342, 344, 346 resources utilized by cloud computing instances 340 a, 340 b, 340 n. In such an example, assume that a user of the cloud-based storage system 318 issues a request to write data that, after being compressed and deduplicated by the storage controller applications 324, 326, will require writing 5 MB of data. In such an example, writing the data to the local storage 330, 334, 338 and block storage 342, 344, 346 resources utilized by the cloud computing instances 340 a, 340 b, 340 n is relatively straightforward because five blocks, each 1 MB in size, are written to the local storage 330, 334, 338 and block storage 342, 344, 346 resources utilized by the cloud computing instances 340 a, 340 b, 340 n. In such an example, software daemons 328, 332, 336 may also be configured to create five objects containing separate 1 MB chunks of data. Thus, in some embodiments, each object written to cloud-based object storage 348 may be identical (or nearly identical) in size. Readers will understand that in such an example, each object may include metadata associated with the data itself (e.g., the first 1 MB of the object is the data, and the remainder is metadata associated with the data). Readers will understand that cloud-based object storage 348 can be incorporated into cloud-based storage system 318 to increase the durability of cloud-based storage system 318.
[0127] In some embodiments, all data stored by cloud-based storage system 318 may be stored in both 1) cloud-based object storage 348 and 2) at least one of local storage 330, 334, 338 or block storage 342, 344, 346 resources utilized by cloud computing instances 340a, 340b, 340n. In such embodiments, the local storage 330, 334, 338 and block storage 342, 344, 346 resources utilized by cloud computing instances 340a, 340b, 340n may effectively operate as a cache that generally contains all data that is also stored in S3, such that all reads of data may be serviced by cloud computing instances 340a, 340b, 340n without requiring cloud computing instances 340a, 340b, 340n to access cloud-based object storage 348. However, the reader will understand that in other embodiments, all data stored by cloud-based storage system 318 may be stored in cloud-based object storage 348, but less than all data stored by cloud-based storage system 318 may be stored in at least one of the local storage 330, 334, 338 resources or block storage 342, 344, 346 resources utilized by cloud computing instances 340a, 340b, 340n. In such examples, various policies may be utilized to determine which subsets of data stored by cloud-based storage system 318 should reside in both 1) cloud-based object storage 348 and 2) at least one of the local storage 330, 334, 338 resources or block storage 342, 344, 346 resources utilized by cloud computing instances 340a, 340b, 340n.
[0128] One or more modules of computer program instructions executing within the cloud-based storage system 318 (e.g., a monitoring module running on its own EC2 instance) may be designed to handle the failure of one or more of the cloud computing instances 340a, 340b, 340n having local storage 330, 334, 338. In such an example, the monitoring module may handle the failure of one or more of the cloud computing instances 340a, 340b, 340n having local storage 330, 334, 338 by creating one or more new cloud computing instances having local storage, retrieving data stored in the failed cloud computing instances 340a, 340b, 340n from cloud-based object storage 348, and storing the data retrieved from cloud-based object storage 348 in local storage in the newly created cloud computing instances. The reader will understand that many variations of this process may be implemented.
[0129] The reader will understand that various performance aspects of the cloud-based storage system 318 can be monitored (e.g., by a monitoring module running on an EC2 instance) so that the cloud-based storage system 318 can be scaled up or out as needed. For example, if the cloud computing instances 320, 322 used to support the execution of the storage controller applications 324, 326 are undersized and are not adequately servicing the I / O requests issued by users of the cloud-based storage system 318, the monitoring module may create a new, more powerful cloud computing instance (e.g., a type of cloud computing instance that includes more processing power, more memory, etc.) that includes the storage controller application so that the new, more powerful cloud computing instance can begin operating as the primary controller. Similarly, if the monitoring module determines that the cloud computing instances 320, 322 used to support the execution of the storage controller applications 324, 326 are oversized and cost savings can be achieved by switching to smaller, less powerful cloud computing instances, the monitoring module can create a new, less powerful (and cheaper) cloud computing instance that includes the storage controller application so that the new, less powerful cloud computing instance can begin operating as the primary controller.
[0130] The storage system described above may implement intelligent data backup techniques, which allow data stored in the storage system to be copied and stored in a different location to avoid data loss in the event of equipment failure or other forms of catastrophic disaster. For example, the storage system described above may be configured to inspect each backup to avoid restoring the storage system to an undesirable state. Consider an example in which malware infects a storage system. In such an example, the storage system may include software resource 314 that can scan each backup to distinguish between backups captured before the malware infected the storage system and backups captured after the malware infected the storage system. In such an example, the storage system may restore itself from a backup that does not contain the malware, or at least may not restore the portion of the backup that contained the malware. In such an example, the storage system may include software resource 314 that may scan each backup to identify the presence of malware (or viruses, or anything else undesirable), for example, by identifying write operations served by the storage system that originate from network subnets served by the storage system that are suspected of delivering malware, by identifying write operations served by the storage system that originate from users that are suspected of delivering malware, by identifying write operations served by the storage system, by inspecting the content of the write operations against malware fingerprints, and in many other ways.
[0131] The reader will further appreciate that backups (often in the form of one or more snapshots) may also be utilized to facilitate rapid recovery of the storage system. Consider an example where a storage system is infected with ransomware that locks users out of the storage system. In such an example, software resources 314 within the storage system may be configured to detect the presence of the ransomware and may further be configured to restore the storage system to a point in time using retained backups prior to the time the ransomware infected the storage system. In such an example, the presence of ransomware may be detected explicitly through the use of software tools utilized by the system, through the use of a key (e.g., a USB drive) inserted into the storage system, or similar methods. Similarly, the presence of ransomware may be inferred in response to system activity that meets a predetermined fingerprint, such as, for example, no reads or writes being input to the system for a predetermined period of time.
[0132] The reader will understand that the various components described above can be grouped into one or more optimized computing packages as an integrated infrastructure. Such an integrated infrastructure can include a pool of computer, storage, and networking resources that can be shared by multiple applications and collectively managed using policy-driven processes. Such an integrated infrastructure can be implemented using an integrated infrastructure reference architecture, using standalone equipment, using a software-driven hyper-integrated approach (e.g., a hyper-integrated infrastructure), or in other ways.
[0133] The reader will understand that the storage systems described in this disclosure may be useful for supporting various types of software applications. Indeed, a storage system may be "application-aware," in the sense that the storage system may acquire, maintain, or otherwise access information describing connected applications (e.g., applications that utilize the storage system) and optimize the operation of the storage system based on intelligence about the applications and their usage patterns. For example, the storage system may optimize data layout, optimize caching behavior, optimize "QoS" levels, or perform some other optimization designed to improve the storage performance experienced by the application.
[0134] As an example of one type of application that may be supported by the storage system described herein, the storage system 306 may be useful in supporting artificial intelligence ("AI") applications, database applications, XOps projects (e.g., DevOps projects, DataOps projects, MLOps projects, ModelOps projects, PlatformOps projects), electronic design automation tools, event-driven software applications, high-performance computing applications, simulation applications, high-speed data capture and analysis applications, machine learning applications, media production applications, media serving applications, picture archiving and communication system ("PACS") applications, software development applications, virtual reality applications, augmented reality applications, and many other types of applications by providing storage resources to such applications.
[0135] Given the fact that storage systems include computational resources, storage resources, and a wide variety of other resources, storage systems may be well suited to supporting resource-intensive applications such as, for example, AI applications, which may be deployed in a variety of fields, including predictive maintenance in manufacturing and related fields, healthcare applications such as patient data and risk analytics, retail and marketing deployments (e.g., search advertising, social media advertising), supply chain solutions, fintech solutions such as business analytics and reporting tools, operational deployments such as real-time analytics tools, application performance management tools, IT infrastructure management tools, and the like.
[0136] Such AI applications may enable devices to perceive their environment and take actions that maximize their chances of success for some purpose. Examples of such AI applications may include IBM Watson™, Microsoft Oxford™, Google DeepMind™, Baidu Minwa™, and others.
[0137] The storage systems described above are also well suited to supporting other types of resource-intensive applications, such as machine learning applications. Machine learning applications can perform various types of data analysis and automate the construction of analytical models. Using algorithms that iteratively learn from data, machine learning applications can enable computers to learn without being explicitly programmed. One particular area of machine learning is called reinforcement learning, which involves taking appropriate actions to maximize rewards in specific situations.
[0138] In addition to the resources already described, the storage systems described above may also include a graphics processing unit (GPU), sometimes referred to as a visual processing unit (VPU). Such a GPU may be embodied as dedicated electronic circuitry that rapidly manipulates and modifies memory to accelerate the creation of images in a frame buffer intended for output to a display device. Such a GPU may be included within any of the computing devices that are part of the storage systems described above, including as one of many individually scalable components of the storage system; other examples of individually scalable components of such storage systems may include storage components, memory components, computational components (e.g., CPUs, FPGAs, ASICs), networking components, software components, and others. In addition to the GPU, the storage systems described above may also include a neural network processor (NNP) for use in various aspects of neural network processing. Such NNPs may be used instead of (or in addition to) a GPU and may be independently scalable.
[0139] As mentioned above, the storage systems described herein can be configured to support artificial intelligence applications, machine learning applications, big data analytics applications, and many other types of applications. The rapid growth in these types of applications is driven by three technologies: deep learning (DL), GPU processors, and big data. Deep learning is a computing model that utilizes massively parallel neural networks inspired by the human brain. Instead of experts handcrafting software, deep learning models write their own software by learning from many examples. Such GPUs can contain thousands of cores that are well suited to running algorithms that roughly represent the parallelism of the human brain.
[0140] Advances in deep neural networks, including the development of multi-layer neural networks, have sparked a new wave of algorithms and tools for data scientists to harness their data with artificial intelligence (AI). With improved algorithms, larger datasets, and a variety of frameworks (including open-source software libraries for machine learning across a range of tasks), data scientists are addressing new use cases such as autonomous vehicles, natural language processing and understanding, computer vision, machine reasoning, strong AI, and more. Applications of such technologies may include machine and vehicle object detection, identification, and avoidance; visual recognition, classification, and tagging; algorithmic financial trading strategy performance management; simultaneous localization and mapping; predictive maintenance of high-value machinery; cybersecurity threat prevention; automated expertise; image recognition and classification; question answering; robotics; text analysis (extraction, classification), and text generation and translation. Applications of AI technologies are embodied in a wide range of products, including, for example, Amazon Echo's speech recognition technology, which enables users to talk to their machines; Google Translate™, which enables machine-based language translation; Spotify's Discover Weekly, which provides recommendations about new songs and artists that users may like based on user usage and traffic analysis; Quill's text generation offering, which takes structured data and turns it into a narrative story; and Chatbots, which provide real-time, context-specific answers to questions in a dialogue format.
[0141] Data is at the heart of modern AI and deep learning algorithms. Before training can begin, one issue that must be addressed is collecting labeled data, which is critical for training accurate AI models. Full-scale AI deployments may require continuously collecting, cleaning, transforming, labeling, and storing large amounts of data. Adding additional high-quality data points directly leads to more accurate models and better insights. Data samples may be subjected to a series of processing steps, including, but not limited to: 1) ingesting data from external sources into the training system and storing the data in raw form; 2) cleaning and transforming the data in a format convenient for training, including linking data samples to appropriate labels; 3) iterating to explore parameters and models, rapidly testing them with smaller datasets, and converging on the most promising model to push to a production cluster; 4) running a training phase to select random batches of input data, including both new and older samples, and feeding them to a production GPU server for computation to update model parameters; and 5) evaluation, which involves using a holdback portion of the data not used in training to evaluate model accuracy against holdout data. This lifecycle can be applied to any type of parallelized machine learning, not just neural networks or deep learning. For example, a standard machine learning framework may rely on a CPU instead of a GPU, but the data ingestion and training workflow may be the same. Readers will understand that a single shared storage data hub creates a coordination point across the entire lifecycle without requiring extra data copies between the ingestion, preprocessing, and training stages. Ingested data is rarely used for only one purpose, and shared storage provides the flexibility to train multiple different models or apply traditional analytics to the data.
[0142] The reader will understand that each stage in an AI data pipeline can have different requirements from the data hub (e.g., a storage system or collection of storage systems). A scale-out storage system must provide uncompromising performance for all access types and patterns, from small, metadata-heavy files to large files, from random access patterns to sequential access patterns, and from low concurrency to high concurrency. The storage system described above can serve as an ideal AI data hub because the system can service unstructured workloads. In the first stage, data is ideally ingested and stored on the same data hub used by subsequent stages to avoid excessive data copying. The next two steps can be performed on standard compute servers, optionally including GPUs, and then in the fourth and final stage, the complete training production job is run on powerful GPU-accelerated servers. Often, a production pipeline exists alongside an experimental pipeline running on the same dataset. Furthermore, GPU-accelerated servers can be used independently for different models or can be combined together to train one larger model, even across multiple systems for distributed training. If the shared storage tier is slow, data must be copied to local storage for each phase, resulting in wasted time staging data to different servers. An ideal data hub for an AI training pipeline would provide similar performance to data stored locally on server nodes, while also possessing the simplicity and performance to allow all pipeline stages to operate simultaneously.
[0143] In order for the above-described storage system to function as a data hub or as part of an AI deployment, in some embodiments, the storage system may be configured to provide DMA between storage devices included in the storage system and one or more GPUs used in an AI or big data analytics pipeline. One or more GPUs may be coupled to the storage system via, for example, NVMe-over-Fabric ("NVMe-oF"), bypassing bottlenecks such as the host CPU and allowing the storage system (or one of the components included therein) to directly access the GPU memory. In such an example, the storage system may leverage API hooks to the GPU to transfer data directly to the GPU. For example, the GPU may be embodied as an Nvidia™ GPU, and the storage system may support GPUDirect Storage ("GDS") software or have similar proprietary software that enables the storage system to transfer data to the GPU via RDMA or a similar mechanism.
[0144] While the preceding paragraphs discuss deep learning applications, the reader will understand that the storage systems described herein may also be part of a distributed deep learning ("DDL") platform to support the execution of DDL algorithms. The storage systems described above may also be paired with other technologies, such as TensorFlow, an open-source software library for dataflow programming across a range of tasks that may be used in machine learning applications, such as neural networks, to facilitate the development of such machine learning models, applications, and the like.
[0145] The storage systems described above can also be used in neuromorphic computing environments. Neuromorphic computing is a form of computing that mimics brain cells. To support neuromorphic computing, an architecture of interconnected "neurons" replaces traditional computing models with low-power signals traveling directly between neurons for more efficient computation. Neuromorphic computing can utilize very-large-scale integration (VLSI) systems that include electronic analog circuits to mimic the neurobiological architecture present in the nervous system, as well as analog, digital, and mixed-mode analog / digital VLSI and software systems that implement models of the nervous system for perception, motor control, or multisensory integration.
[0146] The reader will understand that the storage systems described above may be configured to support the storage or use of blockchains and derived items (among other types of data), such as, for example, open source blockchains and related tools that are part of the IBM™ Hyperledger project, permissioned blockchains in which a certain number of trusted parties are permitted to access the blockchain, blockchain products that allow developers to build their own distributed ledger projects, and others. The blockchains and storage systems described herein may be utilized to support on-chain storage of data as well as off-chain storage of data.
[0147] Off-chain storage of data can be implemented in various ways and can occur when the data itself is not stored within the blockchain. For example, in one embodiment, a hash function can be utilized, and the data itself can be fed into the hash function to generate a hash value. In such an example, a hash of a large piece of data may be embedded within a transaction instead of the data itself. The reader will understand that in other embodiments, alternatives to blockchain can be used to facilitate decentralized storage of information. For example, one alternative to blockchain that can be used is blockweave. While traditional blockchains store every transaction to achieve validation, blockweave allows for secure decentralization without using the entire chain, thereby enabling low-cost on-chain storage of data. Such blockweaves can utilize consensus mechanisms based on proof of access (PoA) and proof of work (PoW).
[0148] The storage systems described above, alone or in combination with other computing devices, can be used to support in-memory computing applications. In-memory computing involves the storage of information in RAM distributed across a cluster of computers. The reader will understand that the storage systems described above, particularly those configurable with customizable amounts of processing, storage, and memory resources (e.g., systems in which blades include configurable amounts of each type of resource), can be configured to provide an infrastructure capable of supporting in-memory computing. Similarly, the storage systems described above can include component parts (e.g., NVDIMMs, 3D cross-point storage providing persistent, high-speed random-access memory) that can actually provide an improved in-memory computing environment compared to an in-memory computing environment that relies on RAM distributed across dedicated servers.
[0149] In some embodiments, the storage systems described above can be configured to operate as hybrid in-memory computing environments that include universal interfaces to all storage media (e.g., RAM, flash storage, 3D cross-point storage). In such embodiments, users may not have knowledge of the details of where their data is stored, but can still address the data using the same complete, unified API. In such embodiments, the storage system can (in the background) move data to the fastest available tier, including intelligently placing data according to various characteristics of the data or some other heuristic. In such examples, the storage system can even utilize existing products such as Apache Ignite and GridGain to move data between various storage tiers, or the storage system can utilize custom software to move data between various storage tiers. The storage systems described herein can implement various optimizations to improve the performance of in-memory computing, such as, for example, having computations occur as close to the data as possible.
[0150] The reader will further understand that in some embodiments, the storage systems described above can be paired with other resources to support the applications described above. For example, one infrastructure may include primary computing in the form of servers and workstations specialized in using general-purpose computing on graphics processing units (GPGPUs) to accelerate deep learning applications, interconnected to a compute engine to train parameters for deep neural networks. Each system may have Ethernet external connectivity, InfiniBand external connectivity, some other form of external connectivity, or some combination thereof. In such examples, GPUs may be grouped for a single large-scale training run or used independently to train multiple models. The infrastructure may also include storage systems such as those described above to provide a scale-out all-flash file or object store, with data accessible via high-performance protocols such as NFS, S3, etc. The infrastructure may also include redundant top-of-rack Ethernet switches connected to the storage and computers via ports in MLAG port channels for redundancy, for example. The infrastructure may also include additional computers in the form of white-box servers, optionally with GPUs, for data ingestion, preprocessing, and model debugging. The reader will appreciate that additional infrastructure is possible.
[0151] The reader will understand that the storage system described above can be configured to support other AI-related tools, either alone or in cooperation with other computing machines. For example, the storage system can utilize tools such as ONXX or other open neural network exchange formats, which make it easier to transfer models written in different AI frameworks. Similarly, the storage system can be configured to support tools such as Amazon's Gluon, which allows developers to prototype, build, and train deep learning models. In fact, the storage system described above can be part of a larger platform, such as IBM™ Cloud Private for Data, which includes integrated data science, data engineering, and application building services.
[0152] The reader will further understand that the storage systems described above can also be deployed as edge solutions. Such edge solutions may be suitable for optimizing cloud computing systems by performing data processing at the edge of the network, close to the source of the data. Edge computing can push applications, data, and computing power (i.e., services) from centralized points to the logical extremes of the network. Through the use of edge solutions such as the storage systems described above, computational tasks can be performed using the computational resources provided by such storage systems, data can be stored using the storage resources of the storage systems, and cloud-based services can be accessed through the use of various resources (including networking resources) of the storage systems. By performing computational tasks on edge solutions, storing data on edge solutions, and utilizing edge solutions in general, consumption of expensive cloud-based resources can be avoided and, in fact, performance improvements can be experienced relative to a heavier reliance on cloud-based resources.
[0153] While many tasks can benefit from the use of edge solutions, some specific uses may be particularly suited to deployment in such environments. For example, drones, autonomous vehicles, robots, and other devices may require very fast processing; in practice, transmitting data up to a cloud environment and back to receive data processing support may simply be too slow. As an additional example, some IoT devices, such as connected video cameras, may not be well suited to utilizing cloud-based resources simply because the sheer volume of data involved may make transmitting data to the cloud impractical (not just from a privacy, security, or financial perspective). Thus, many tasks involving data processing, storage, or communication may indeed be better suited by platforms that include edge solutions, such as the storage systems described above.
[0154] The storage system described above, alone or in combination with other computing resources, can function as a network edge platform, combining compute resources, storage resources, networking resources, cloud technologies, and network virtualization technologies. As part of the network, the edge can take on similar characteristics as other network facilities, from customer premises and backhaul aggregation facilities to points of presence (PoPs) and regional data centers. Readers will understand that network workloads such as virtual network functions (VNFs) and others reside on the network edge platform. Network edge platforms enabled by the combination of containers and virtual machines may rely on controllers and schedulers that are no longer geographically co-located with data processing resources. Microservice-based functions can be split into control planes, user and data planes, or even state machines, allowing independent optimization and scaling techniques to be applied. Such user and data planes can be enabled both through the increasing number of accelerators present in server platforms, such as FPGAs and smart NICs, and through SDN-enabled merchant silicon and programmable ASICs.
[0155] The above-described storage systems can also be optimized for use in big data analytics, including leveraging containerized analytics architectures, for example, as part of a configurable data analytics pipeline, making analytics capabilities more configurable. Big data analytics can generally be described as the process of examining large and diverse data sets to uncover hidden patterns, unknown correlations, market trends, customer preferences, and other useful information that can help organizations make more informed business decisions. As part of that process, semi-structured and unstructured data, such as internet clickstream data, web server logs, social media content, text from customer emails and survey responses, mobile phone call detail records, IoT sensor data, and other data, may be converted into a structured form.
[0156] The storage system described above may also support (including implementing as a system interface) applications that perform tasks in response to human speech. For example, the storage system may support the execution of intelligent personal assistant applications such as Amazon's Alexa™, Apple Siri™, Google Voice™, Samsung Bixby™, Microsoft Cortana™, and others. While the examples described in the previous sentence utilize voice as input, the storage system described above may also support chatbots, talkbots, chatterbots, or artificial conversational entities, or other applications configured to conduct conversations via auditory or textual methods. Similarly, the storage system may actually execute such applications to enable users, such as system administrators, to interact with the storage system via voice. While such applications generally enable voice interaction, music playback, creating to-do lists, setting alarms, streaming podcasts, playing audiobooks, and providing other real-time information such as weather, traffic, and news, in embodiments according to the present disclosure, such applications may be used as an interface to various system management operations.
[0157] The storage systems described above can also implement AI platforms to deliver on the vision of self-driving storage. Such AI platforms can be configured to provide global predictive intelligence by collecting and analyzing large volumes of storage system telemetry data points, enabling easy management, analysis, and support. Indeed, such storage systems may be capable of predicting both capacity and performance and generating intelligent advice regarding workload deployment, interaction, and optimization. Such AI platforms can be configured to scan all incoming storage system telemetry data against a library of problem fingerprints, capturing hundreds of performance-related variables used to predict performance loads, in order to predict and resolve incidents in real time before they impact customer environments.
[0158] The storage systems described above can support the serialization or concurrent execution of artificial intelligence applications, machine learning applications, data analysis applications, data transformations, and other tasks that may collectively form an AI ladder. Such an AI ladder may be effectively formed by combining such elements to form a complete data science pipeline in which dependencies exist between the elements of the AI ladder. For example, AI may require that some form of machine learning has taken place, which may require that some form of analytics has taken place, which may require that some form of data and information construction has taken place, etc. Thus, each element can be considered a rung in the AI ladder that can collectively form a complete and sophisticated AI solution.
[0159] The above-described storage systems can also be used, alone or in combination with other computing environments, to deliver AI to any experience where AI permeates broad and expansive aspects of business and life. For example, AI may play a key role in the delivery of deep learning solutions, deep reinforcement learning solutions, artificial general intelligence solutions, autonomous vehicles, cognitive computing solutions, commercial UAVs or drones, conversational user interfaces, enterprise taxonomies, ontology management solutions, machine learning solutions, smart dust, smart robots, smart workplaces, etc.
[0160] The above-described storage systems may also be used, alone or in combination with other computing environments, to provide a wide range of transparent and immersive experiences (including those using digital twins of various "things," such as people, places, processes, systems, etc.) where technology can introduce transparency between people, businesses, and things. Such transparent and immersive experiences may be provided as augmented reality technology, connected homes, virtual reality technology, brain-computer interfaces, human augmentation technology, nanotube electronics, volumetric displays, 4D printing technology, or others.
[0161] The above-described storage systems may also be used, alone or in combination with other computing environments, to support a wide variety of digital platforms, including, for example, 5G wireless systems and platforms, digital twin platforms, edge computing platforms, IoT platforms, quantum computing platforms, serverless PaaS, software-defined security, neuromorphic computing platforms, and the like.
[0162] The storage system described above may also be part of a multi-cloud environment in which multiple cloud computing and storage services are deployed in a single heterogeneous architecture. To facilitate the operation of such a multi-cloud environment, DevOps tools may be deployed to enable orchestration across clouds. Similarly, continuous development and continuous integration tools may be deployed to standardize processes for continuous integration and delivery, new feature rollout, and cloud workload provisioning. By standardizing these processes, a multi-cloud strategy may be implemented that enables the utilization of the best provider for each workload.
[0163] The storage systems described above can be used as part of a platform that enables the use of cryptographic anchors that can be used to authenticate the origin and content of a product and ensure that it matches the blockchain record associated with the product. Similarly, as part of a suite of tools for securing data stored on the storage system, the storage systems described above can implement various encryption techniques and schemes, including lattice cryptography. Lattice cryptography can involve the construction of cryptographic primitives that include a lattice in either the construction itself or in security proofs. Unlike public-key schemes such as RSA, Diffie-Hellman, or Elliptic-Curve cryptosystems, which are easily attacked by quantum computers, some lattice-based configurations are believed to be resistant to attacks by both classical and quantum computers.
[0164] A quantum computer is a device that performs quantum computing. Quantum computing is calculation using quantum mechanical phenomena such as superposition and entanglement. Quantum computers differ from conventional computers based on transistors because such conventional computers require data to be encoded into binary digits (bits), each of which is always in one of two distinct states (0 or 1). In contrast to conventional computers, quantum computers use qubits, which can be in superposition of states. A quantum computer maintains a set of qubits, and a single qubit can represent 1, 0, or any quantum superposition of the two qubit states. A pair of qubits can be in any quantum superposition of four states, and three qubits can be in any superposition of eight states. A quantum computer with n qubits can generally be in any superposition of up to 2^n different states simultaneously, while a conventional computer can only be in one of these states at any one time. A quantum Turing machine is a theoretical model of such a computer.
[0165] The storage systems described above can be paired with FPGA acceleration servers as part of a larger AI or ML infrastructure. Such FPGA acceleration servers may reside near the storage systems described above (e.g., in the same data center) or may be incorporated into an appliance that includes one or more storage systems, one or more FPGA acceleration servers, a networking infrastructure supporting communication between the one or more storage systems and the one or more FPGA acceleration servers, and other hardware and software components. Alternatively, the FPGA acceleration servers can reside in a cloud computing environment that can be used to perform computation-related tasks for AI and ML jobs. Any of the above-described embodiments can collectively function as an FPGA-based AI or ML platform. The reader will understand that in some embodiments of an FPGA-based AI or ML platform, the FPGA included in the FPGA acceleration server can be reconfigured for different types of ML models (e.g., LSTM, CNN, GRU). The ability to reconfigure the FPGA included in the FPGA acceleration server can enable the acceleration of ML or AI applications based on the most optimal numerical precision and memory model being used. The reader will understand that by treating a collection of FPGA-accelerated servers as a pool of FPGAs, any CPU in the data center can utilize the pool of FPGAs as a shared hardware microservice, rather than limiting the server to the dedicated accelerators plugged into it.
[0166] The FPGA-accelerated servers and GPU-accelerated servers described above can implement a model of computing in which CPU models and parameters are pinned to high-bandwidth on-chip memory and much data is streamed through the high-bandwidth on-chip memory, rather than holding a small amount of data in machine learning and executing a long stream of instructions on it, as is done in traditional computing models. Because FPGAs can be programmed with only the instructions necessary to execute this type of computing model, FPGAs can be even more efficient than GPUs for this type of computing model.
[0167] The storage systems described above can be configured to provide parallel storage, for example, through the use of a parallel file system such as BeeGFS. Such a parallel file system can include a distributed metadata architecture. For example, a parallel file system may include components including multiple metadata servers across which metadata is distributed, and services for clients and storage servers.
[0168] The above-described system can support the execution of a variety of software applications. Such software applications can be deployed in various ways, including a container-based deployment model. Containerized applications can be managed using various tools. For example, containerized applications may be managed using Docker Swarm, Kubernetes, and others. Containerized applications can be used to facilitate a serverless, cloud-native computing deployment and management model for software applications. In support of a serverless, cloud-native computing deployment and management model for software applications, containers can be used as part of an event handling mechanism (e.g., AWS Lambda) such that various events spin up containerized applications to act as event handlers.
[0169] The above-described systems may be deployed in various ways, including in a manner that supports fifth-generation ("5G") networks. 5G networks may support substantially faster data communications than prior-generation mobile communications networks, potentially leading to the decentralization of data and computing resources, as modern large-scale data centers may become less prominent and may be replaced by more local micro-data centers, for example, closer to mobile network towers. The above-described systems may be included in such local micro-data centers or may be part of or paired with multi-access edge computing ("MEC") systems. Such MEC systems may enable cloud computing capabilities and IT service environments at the edge of cellular networks. By running applications and performing associated processing tasks closer to cellular customers, network congestion may be reduced and applications may perform better.
[0170] The storage system described above can be configured to implement NVMe Zoned Namespaces. Through the use of NVMe Zoned Namespaces, the logical address space of the namespace is divided into zones. Each zone provides a logical block address range that is written sequentially and must be explicitly reset before being rewritten, thereby enabling the creation of a namespace that exposes the natural boundaries of the device and offloading management of internal mapping tables to the host. To implement NVMe Zoned Namespaces ("Zoned Namespaces"), ZNS SSDs or other forms of zoned block devices that expose the namespace logical address space using zones can be utilized. When zones are aligned to the internal physical characteristics of the device, multiple inefficiencies in data placement can be eliminated. In such embodiments, each zone can be mapped to a separate application, such that functions such as wear leveling and garbage collection can be performed per zone or per application, rather than across the entire device. To support ZNS, the storage controllers described herein can be configured to interact with zoned block devices, for example, through the use of the Linux™ kernel zoned block device interface or other tools.
[0171] The storage systems described above may also be configured to implement zoned storage in other ways, such as through the use of shingled magnetic recording (SMR) storage devices. In instances where zoned storage is used, a device-managed embodiment may be deployed, where the storage device hides this complexity by managing it in firmware and presenting an interface like any other storage device. Alternatively, zoned storage may be implemented through a host-managed embodiment that relies on the operating system to know how to handle the drive and writes sequentially only to specific areas of the drive. Zoned storage may similarly be implemented using a host-aware embodiment, where a combination of drive-managed and host-managed implementations is deployed.
[0172] The storage systems described herein can be used to form a data lake. A data lake can act as the first place an organization's data flows, and such data can be in a raw format. Metadata tagging can be implemented to facilitate searching of data elements within the data lake, especially in embodiments where the data lake includes multiple stores of data in formats that are not easily accessible or readable (e.g., unstructured data, semi-structured data, structured data). From the data lake, data can proceed downstream to a data warehouse, where the data can be more processed, packaged, and stored in a consumable format. The storage systems described above can also be used to implement such a data warehouse. In addition, a data mart or data hub can enable even more easily consumed data, and the storage systems described above can be used to provide the underlying storage resources required for the data mart or data hub. In embodiments, queries to the data lake may require a schema-on-read approach, in which the plan or schema is applied as the data is pulled from the stored location, rather than as the plan or schema is entered.
[0173] The storage systems described herein may also be configured to implement a recovery point objective ("RPO"), which may be established by a user, an administrator, a system default, as part of a storage class or service that the storage system participates in delivering, or in some other manner. A "recovery point objective" is a target for the maximum time difference between the last update to a source dataset and the last recoverable replicated dataset update that is correctly recoverable from a continuously or frequently updated copy of the source dataset. An update is correctly recoverable if it properly takes into account all updates processed to the source dataset prior to the last recoverable replicated dataset update.
[0174] In synchronous replication, the RPO is zero, which means that under normal operation, all completed updates on the source data set should be present and correctly recoverable on the copy data set. In best-effort near-synchronous replication, the RPO can be as low as a few seconds. In snapshot-based replication, the RPO can be roughly calculated as the interval between snapshots plus the time to transfer modifications between the previous, already transferred snapshot and the latest replicated snapshot.
[0175] If updates accumulate faster than they can be replicated, the RPO can be missed. In the case of snapshot-based replication, the RPO can be missed if more replicated data accumulates between two snapshots than can be replicated between taking a snapshot and replicating that snapshot's cumulative updates to the copy. Again, in snapshot-based replication, if the data being replicated accumulates at a rate faster than it can be transferred in the time between subsequent snapshots, replication can begin to fall further behind, widening the gap between the expected recovery point objective and the actual recovery point represented by the last successfully replicated update.
[0176] The storage systems described above may be part of a shared-nothing storage cluster. In a shared-nothing storage cluster, each node of the cluster has local storage and communicates with other nodes in the cluster over a network, and the storage used by the cluster is (generally) provided solely by the storage connected to each individual node. A collection of nodes synchronously replicating a data set may be an example of a local storage cluster, since each storage system has shared-nothing and communicates with other storage systems over a network; these storage systems (generally) do not use other storage to which they share access through some interconnect. In contrast, some of the storage systems described above are themselves configured as shared storage clusters, since there are drive shelves shared by paired controllers. However, other storage systems described above are configured as shared-nothing storage clusters, since all storage is local to a particular node (e.g., blade) and all communication is via the network linking the compute nodes to each other.
[0177] In other embodiments, other forms of shared-nothing storage clusters may include embodiments in which any node in the cluster has a local copy of all the storage it needs, and to ensure that data is not lost, or because other nodes are also using that storage, the data is mirrored to other nodes in the cluster via synchronous replication. In such an embodiment, if a new cluster node needs some data, it can be copied to the new node from other nodes that have copies of that data.
[0178] In some embodiments, a mirror copy-based shared storage cluster may store multiple copies of all the cluster's stored data, with each subset of the data replicated to a particular set of nodes and different subsets of the data replicated to different sets of nodes. In some variations, embodiments may store all of the cluster's stored data on all nodes, while in other variations, the nodes may be divided so that a first set of nodes all store the same set of data and a second, different set of nodes all store a different set of data.
[0179] The reader will understand that a RAFT-based database (e.g., etcd) can operate like a shared-nothing cluster, with all RAFT nodes storing all data. However, the amount of data stored in a RAFT cluster can be limited so that redundant copies do not consume too much storage. A container server cluster may also be capable of replicating all data to all cluster nodes, assuming containers do not tend to be too large and their bulk data (data manipulated by applications running within containers) is stored elsewhere, such as an S3 cluster or external file server. In such an example, container storage may be provided by the cluster directly through its shared-nothing storage model, and those containers provide the images that form the execution environment for portions of an application or service.
[0180] For further explanation, FIG. 3D illustrates an exemplary computing device 350 that may be particularly configured to perform one or more of the processes described herein. As shown in FIG. 3D, computing device 350 may include a communication interface 352, a processor 354, a storage device 356, and an input / output (“I / O”) module 358, communicatively coupled to each other via a communication infrastructure 360. While an exemplary computing device 350 is shown in FIG. 3D, the components illustrated in FIG. 3D are not intended to be limiting. In other embodiments, additional or alternative components may be used. The components of computing device 350 shown in FIG. 3D will now be described in further detail.
[0181] The communication interface 352 may be configured to communicate with one or more computing devices. Examples of the communication interface 352 include, but are not limited to, a wired network interface connection (such as a network interface card), a wireless network interface connection (such as a wireless network interface card connection), a modem, an audio / video connection, and any other suitable interface.
[0182] The processor 354 generally represents any type or form of processing unit capable of processing data and / or interpreting, executing, and / or directing the execution of one or more of the instructions, processes, and / or operations described herein. The processor 354 may perform operations by executing computer-executable instructions 362 (e.g., applications, software, code, and / or other executable data instances) stored on the storage device 356.
[0183] Storage device 356 may include one or more data storage media, devices, or configurations, and any type, form, and combination of data storage media and / or devices may be used. For example, storage device 356 may include, but is not limited to, any combination of non-volatile and / or volatile media described herein. Electronic data, including the data described herein, may be stored temporarily and / or persistently in storage device 356. For example, data representing computer-executable instructions 362 configured to instruct processor 354 to perform any of the operations described herein may be stored in storage device 356. In some examples, data may be located in one or more databases residing in storage device 356.
[0184] I / O module 358 may include one or more I / O modules configured to receive user input and provide user output. I / O module 358 may include any hardware, firmware, software, or combination thereof that supports input and output capabilities. For example, I / O module 358 may include hardware and / or software for capturing user input, including, but not limited to, a keyboard or keypad, a touchscreen component (e.g., a touchscreen display), a receiver (e.g., an RF or infrared receiver), a motion sensor, and / or one or more input buttons.
[0185] I / O module 358 may include one or more devices for presenting output to a user, including, but not limited to, a graphics engine, a display (e.g., a display screen), one or more output drivers (e.g., a display driver), one or more audio speakers, and one or more audio drivers. In particular embodiments, I / O module 358 is configured to provide graphical data to a display for presentation to a user. The graphical data may represent one or more graphical user interfaces and / or any other graphical content as may be useful in a particular implementation. In some examples, any of the systems, computing devices, and / or other components described herein may be implemented by computing device 350.
[0186] For further explanation, FIG. 3E illustrates an example of a fleet of storage systems 376 for providing storage services (also referred to herein as “data services”). The fleet of storage systems 376 depicted in FIG. 3 includes multiple storage systems 374a, 374b, 374c, 374d, and 374n, each of which may be similar to the storage systems described herein. The storage systems 374a, 374b, 374c, 374d, and 374n in the fleet of storage systems 376 may be embodied as the same storage system or as different types of storage systems. For example, two of the storage systems 374a, 374n depicted in FIG. 3E are depicted as cloud-based storage systems because the resources collectively forming each of the storage systems 374a, 374n are provided by separate cloud service providers 370, 372. For example, first cloud service provider 370 may be Amazon AWS™ while second cloud service provider 372 is Microsoft Azure™, although in other embodiments, one or more public clouds, private clouds, or a combination thereof may be used to provide the underlying resources used to form a particular storage system within fleet of storage systems 376.
[0187] 3E includes an edge management service 382 for delivering storage services, according to some embodiments of the present disclosure. The delivered storage services (also referred to herein as "data services") may include, for example, services that provide a specific amount of storage to consumers, services that provide storage to consumers subject to certain service level agreements, services that provide storage to consumers subject to certain regulatory requirements, and many other services.
[0188] 3E may be embodied as one or more modules of computer program instructions executing on computer hardware, such as one or more computer processors. Alternatively, edge management service 382 may be embodied as one or more modules of virtualized instructions executing on computer program execution environments, such as one or more virtual machines, within one or more containers or in some other manner. In other embodiments, edge management service 382 may be embodied as a combination of the above-described embodiments, including embodiments in which one or more modules of computer program instructions included in edge management service 382 are distributed across multiple physical or virtual execution environments.
[0189] The edge management service 382 may act as a gateway to provide storage services to storage consumers, where the storage services leverage storage provided by one or more storage systems 374a, 374b, 374c, 374d, 374n. For example, the edge management service 382 may be configured to provide storage services to host devices 378a, 378b, 378c, 378d, 378n running one or more applications that consume the storage services. In such an example, the edge management service 382 can act as a gateway between the host devices 378a, 378b, 378c, 378d, 378n and the storage systems 374a, 374b, 374c, 374d, 374n, rather than requiring the host devices 378a, 378b, 378c, 378d, 378n to directly access the storage systems 374a, 374b, 374c, 374d, 374n.
[0190] While the edge management service 382 of FIG. 3E exposes the storage services module 380 to the host devices 378a, 378b, 378c, 378d, and 378n of FIG. 3E, in other embodiments, the edge management service 382 may expose the storage services module 380 to other consumers of various storage services. The various storage services may be presented to the consumers via one or more user interfaces, via one or more APIs, or through some other mechanism provided by the storage services module 380. Thus, the storage services module 380 depicted in FIG. 3E may be embodied as one or more modules of virtualized instructions executing on physical hardware, on a computer program execution environment, or a combination thereof, where execution of such modules enables consumers of storage services to be offered, select, and access various storage services.
[0191] The edge management services 382 of Figure 3E also includes a system management services module 384. The system management services module 384 of Figure 3E includes one or more modules of computer program instructions that, when executed, perform various operations in coordination with the storage systems 374a, 374b, 374c, 374d, 374n to provide storage services to the host devices 378a, 378b, 378c, 378d, 378n. The system management services module 384 may be configured to perform tasks such as, for example, provisioning storage resources from the storage systems 374a, 374b, 374c, 374d, 374n via one or more APIs exposed by the storage systems 374a, 374b, 374c, 374d, 374n, migrating datasets or workloads between the storage systems 374a, 374b, 374c, 374d, 374n via one or more APIs exposed by the storage systems 374a, 374b, 374c, 374d, 374n, and setting one or more tunable parameters (i.e., one or more configurable settings) on the storage systems 374a, 374b, 374c, 374d, 374n via one or more APIs exposed by the storage systems 374a, 374b, 374c, 374d, 374n. For example, many of the services described below relate to embodiments in which storage systems 374a, 374b, 374c, 374d, 374n are configured to operate in some manner. In such examples, system management services module 384 may be responsible for using APIs (or some other mechanism) provided by storage systems 374a, 374b, 374c, 374d, 374n to configure storage systems 374a, 374b, 374c, 374d, 374n to operate in the manner described below.
[0192] In addition to configuring storage systems 374a, 374b, 374c, 374d, 374n, the edge management service 382 itself may be configured to perform various tasks required to provide various storage services. Consider an example where a storage service includes a service that, when selected and applied, obfuscates personally identifiable information ("PII") contained in a dataset when the dataset is accessed. In such an example, storage systems 374a, 374b, 374c, 374d, 374n may be configured to obfuscate PII when servicing read requests directed to the dataset. Alternatively, the storage systems 374a, 374b, 374c, 374d, 374n may service the read by returning data that includes PII, but the edge management service 382 itself may obfuscate the PII as the data passes through the edge management service 382 on its way from the storage systems 374a, 374b, 374c, 374d, 374n to the host devices 378a, 378b, 378c, 378d, 378n.
[0193] The storage systems 374a, 374b, 374c, 374d, and 374n depicted in FIG. 3E may be embodied as one or more of the storage systems (including variations thereof) described above with reference to FIGS. 1A-3D. In practice, the storage systems 374a, 374b, 374c, 374d, and 374n may function as a pool of storage resources, with individual components within the pool having different performance characteristics, different storage characteristics, and the like. For example, one of the storage systems 374a may be a cloud-based storage system, another storage system 374b may be a storage system providing block storage, another storage system 374c may be a storage system providing file storage, another storage system 374d may be a relatively high-performance storage system, while another storage system 374n may be a relatively low-performance storage system, and so on. In other embodiments, only a single storage system may be present.
[0194] The storage systems 374a, 374b, 374c, 374d, and 374n depicted in FIG. 3E may also be organized into different failure domains such that a failure of one storage system 374a is completely independent of a failure of another storage system 374b. For example, each of the storage systems may receive power from an independent power system, each of the storage systems may be coupled for data communication via an independent data communication network, etc. Furthermore, storage systems in a first failure domain may be accessed through a first gateway, while storage systems in a second failure domain may be accessed through a second gateway. For example, the first gateway may be a first instance of edge management service 382, and the second gateway may be a second instance of edge management service 382, including embodiments in which each instance is separate or each instance is part of a distributed edge management service 382.
[0195] As an illustrative example of available storage services, a user may be presented with storage services associated with different levels of data protection. For example, a user may be presented with storage services that, when selected and implemented, assure the user that data associated with that user is protected such that various recovery point objectives (“RPOs”) can be guaranteed. A first available storage service, for example, may ensure that a subset of data sets associated with the user are protected such that any data older than five seconds can be recovered in the event of a failure of the primary data store, while a second available storage service may ensure that a subset of data sets associated with the user are protected such that any data older than five minutes can be recovered in the event of a failure of the primary data store.
[0196] Additional examples of storage services that may be presented to a user, selected by the user, and ultimately applied to a dataset associated with the user may include one or more data compliance services. Such data compliance services may be embodied as services that may be provided to a consumer of the data compliance services (i.e., a user) to ensure, for example, that the user's dataset is managed to comply with various regulatory requirements. For example, one or more data compliance services may be provided to a user to ensure that the user's dataset is managed in a manner that complies with the General Data Protection Regulation ("GDPR"); one or more data compliance services may be provided to a user to ensure that the user's dataset is managed in a manner that complies with the Sarbanes-Oxley Act of 2002 ("SOX"); or one or more data compliance services may be provided to a user to ensure that the user's dataset is managed in a manner that complies with some other regulatory act. Additionally, one or more data compliance services may be provided to a user to ensure that their datasets are managed in adherence to some non-governmental guidance (e.g., in adherence to best practices for audit purposes), one or more data compliance services may be provided to a user to ensure that their datasets are managed in adherence to the requirements of a particular client or organization, etc.
[0197] Consider an example in which a specific data compliance service is designed to ensure that a user's data sets are managed in a manner that complies with the requirements set forth in the GDPR. While a complete list of the GDPR's requirements can be found in the regulation itself, for illustrative purposes, an example of a requirement set forth in the GDPR requires that a pseudonymization process must be applied to stored data to transform personal data such that the resulting data cannot be attributed to a specific data subject without the use of additional information. For example, data encryption techniques can be applied to make the original data unintelligible, and such data encryption techniques cannot be reversed without access to the correct decryption key. Thus, the GDPR may require that the decryption key be kept separate from the pseudonymized data. One specific data compliance service may be offered to ensure compliance with the requirements set forth in this paragraph.
[0198] To provide this particular data compliance service, the data compliance service may be presented to the user (e.g., via a GUI) and selected by the user. In response to receiving a selection of a particular data compliance service, one or more storage service policies may be applied to the dataset associated with the user to perform the particular data compliance service. For example, a storage service policy may be applied that requires the dataset to be encrypted before being stored in a storage system, a cloud environment, or elsewhere. To enforce this policy, not only may a requirement be enforced that the dataset be encrypted when stored, but a requirement may also be introduced that requires the dataset to be encrypted before transmitting the dataset (e.g., sending the dataset to another party). In such an example, a storage service policy may also be introduced that requires that any encryption key used to encrypt the dataset not be stored on the same system that stores the dataset itself. The reader will understand that many other forms of data compliance services may be provided and implemented by embodiments of the present disclosure.
[0199] Another example of a storage service that may be presented to a user, selected by the user, and ultimately applied to a dataset associated with the user may include one or more high-availability storage services. Such high-availability storage services may be embodied as services that may be provided to consumers (i.e., users) of the high-availability storage services to ensure, for example, that the user's dataset is guaranteed to have a particular level of uptime (i.e., be available for a predetermined amount of time). For example, a first high-availability storage service may be provided to a user to ensure that the user's dataset has three-nines availability, meaning that the dataset is available 99.9% of the time. However, a second high-availability storage service may be provided to a user to ensure that the user's dataset has five-nines availability, meaning that the dataset is available 99.999% of the time. Other high-availability storage services that ensure other levels of availability may be provided. Similarly, high-availability storage services may also be delivered in a manner that ensures various levels of uptime for entities other than datasets. For example, a particular highly available storage service may ensure that one or more virtual machines, one or more containers, one or more data access endpoints, or some other entity is available for a particular amount of time.
[0200] Consider an example where a particular highly available storage service is designed to ensure that a user's dataset has five 9s availability, i.e., the dataset is available 99.999% of the time. In such an example, one or more storage service policies associated with such highly available storage service may be applied and enforced to provide this uptime guarantee. For example, a storage service policy may be applied that requires the dataset to be mirrored across a predetermined number of storage systems, a storage service policy may be applied that requires the dataset to be mirrored across a predetermined number of availability zones in a cloud environment, a storage service policy may be applied that requires the dataset to be replicated in a particular manner (e.g., synchronously replicated so that multiple up-to-date copies of the dataset exist), etc.
[0201] The reader will understand that many other forms of highly available storage services may be provided and implemented by embodiments of the present disclosure. For example, not only may such a highly available storage service be associated with a storage service policy that ensures that a certain number of copies of a dataset exist to protect against failures in the storage infrastructure (e.g., failure of the storage system itself), but the highly available storage service may also be associated with storage service policies designed to prevent the dataset from becoming unavailable for other reasons. For example, a highly available storage service may be associated with one or more storage service policies that require a certain level of redundancy in networking paths, so that availability to meet uptime requirements associated with a dataset is not constituted by the lack of a data communication path for accessing the storage resource containing the dataset. Other embodiments of the present disclosure may provide and implement additional forms of highly available storage services.
[0202] Another example of a storage service that may be presented to a user, selected by the user, and ultimately applied to a dataset associated with the user may include one or more disaster recovery services. Such disaster recovery services may be embodied as services that may be offered to consumers of the disaster recovery services (i.e., users) to ensure, for example, that the user's dataset can be recovered according to certain parameters in the event of a disaster. For example, a first disaster recovery service may be offered to a user to ensure that the user's dataset can be recovered according to a first RPO and a first recovery time objective ("RTO") if the storage system storing the dataset fails or some other form of disaster occurs that makes the dataset unavailable. However, a second disaster recovery service may be offered to the user to ensure that the user's dataset can be recovered according to a second RPO and a second RTO if the storage system storing the dataset fails or some other form of disaster occurs that makes the dataset unavailable. Other disaster recovery services may be offered that ensure other levels of recoverability outside of the RPO and RTO. Similarly, disaster recovery services may be delivered in a manner that ensures various levels of recoverability for entities other than datasets. For example, a particular disaster recovery service may ensure that one or more virtual machines, one or more containers, one or more data access endpoints, or some other entity can be recovered in a certain amount of time or according to some other metric. In this example, disaster recovery services can be presented to a user, a selection of one or more selected disaster recovery services can be received, and the selected disaster recovery services can be applied to a dataset (or other entity) associated with the user.
[0203] Consider an example where a particular disaster recovery service is designed to ensure that a user's dataset has a zero RPO and RTO. This means that if a particular storage system that stores the dataset fails or some other form of disaster occurs, the dataset must be immediately available without data loss. In such an example, one or more storage service policies associated with such disaster recovery service may be applied and enforced to deliver these RPO / RTO guarantees. For example, a storage service policy may be applied that requires that a dataset be synchronously replicated across a predetermined number of storage systems, meaning that a request to modify a dataset may be acknowledged as completed only when all of the storage systems containing copies of the dataset have modified the dataset in accordance with the request. In other words, the modification operation (e.g., a write) may be applied to all copies of the dataset present on storage systems or to none of the copies of the dataset present on storage systems, so that the failure of one storage system does not prevent a user from accessing the same up-to-date copy of the dataset.
[0204] The reader will understand that many other forms of disaster recovery services can be provided and implemented by embodiments of the present disclosure. For example, such disaster recovery services can be associated with storage service policies that do more than just ensure that a disaster can be recovered according to specific RPO or RTO requirements; disaster recovery services may also be associated with storage service policies designed to ensure that a dataset (or other entity) can be restored to a predetermined location or system; disaster recovery services may also be associated with storage service policies designed to ensure that a dataset (or other entity) can be restored at or below a specified maximum cost; disaster recovery services may also be associated with storage service policies designed to ensure that a particular action is taken in response to detecting a disaster (e.g., spinning up a clone of a failed system), etc. Other embodiments of the present disclosure can provide and implement additional forms of disaster recovery services.
[0205] Another example of a storage service that may be presented to a user, selected by the user, and ultimately applied to a dataset associated with the user may include one or more data archiving services (including data offloading services). Such data archiving services may be embodied, for example, as services that are provided to consumers (i.e., users) of the data archiving services and that may ensure that the user's dataset is archived in a particular manner, such as according to a particular set of preferences, parameters, etc. For example, one or more data archiving services may be provided to a user to ensure that the user's dataset is archived if the data has not been accessed within a predetermined period of time, if the data has been invalid for a particular period of time, after the dataset has reached a particular size, when the storage system storing the dataset has reached a predetermined utilization level, etc.
[0206] Consider an example in which a particular data archiving service is designed to ensure that portions of a user's dataset that are invalidated (e.g., the portion is replaced with an updated portion, the portion is deleted) are archived within 24 hours of the data being invalidated. To provide this particular data archiving service, the data archiving service may be presented to the user (e.g., via a GUI) and selected by the user. In response to receiving the selection of the particular data archiving service, one or more storage service policies may be applied to the dataset associated with the user to perform the particular data archiving service. For example, a storage service policy may be applied that requires that any actions that cause data to be invalidated (e.g., overwrite, delete) be cataloged, and that every 12 hours a process checks for the cataloged actions and migrates invalidated data to the archive. Similarly, a storage service policy may impose restrictions on processes such as garbage collection, if appropriate, to prevent such processes from deleting data that has not yet been archived. The reader will understand that in this example, the data archiving service may operate in coordination with one or more data compliance services, as requirements for archiving data may be created by one or more regulations. For example, a user who selects a particular data compliance service that requires that certain data be retained for a predetermined period of time may automatically trigger a particular data archiving service that can provide the requested level of data archiving and retention. Similarly, some data archiving services may be incompatible with some data compliance services, and as a result, a user may be prevented from selecting two competing or otherwise incompatible services.
[0207] Another example of a storage service that may be presented to a user, selected by the user, and ultimately applied to a dataset associated with the user may include one or more quality-of-service ("QoS") storage services. Such a QoS storage service may be embodied as a service that may be provided to a consumer (i.e., a user) of a data compliance service to ensure, for example, that the user's dataset may be accessed according to a predetermined performance metric. For example, a particular QoS storage service may guarantee that reads directed to the user's dataset may be serviced within a particular amount of time, that writes directed to the user's dataset may be serviced within a particular amount of time, that the user may be guaranteed a certain number of IOPS directed to the user's dataset, etc.
[0208] Consider an example in which a particular QoS storage service is designed to ensure that a user's dataset can be accessed in a manner that guarantees that read and write latencies will be lower than a predetermined amount of time. To provide this particular QoS storage service, the QoS storage service may be presented to the user (e.g., via a GUI) and selected by the user. In response to receiving a selection of the particular QoS storage service, one or more storage service policies may be applied to the dataset associated with the user to implement the particular QoS storage service. For example, a storage service policy may be applied that requires the dataset to be maintained in storage that can be used to provide the required read and write latencies. For example, if the particular QoS storage service requires relatively low read and write latencies, then the storage service policy associated with this particular QoS storage service may require that the user's dataset be stored on relatively high-performance storage that is located relatively close to any hosts issuing I / O operations targeting that dataset.
[0209] The reader will understand that many other forms of QoS storage services may be provided and implemented by embodiments of the present disclosure. Indeed, a QoS storage service may operate in coordination with one or more other storage services. For example, a particular QoS storage service may operate in coordination with a particular high-availability storage service because the QoS storage service may be associated with particular availability requirements that may be implemented through the application of the particular high-availability storage service. Similarly, a particular QoS storage service may operate in coordination with a particular data replication service because the QoS storage service may be associated with particular performance guarantees that can only be achieved by replicating (via a particular data replication service) a data set to storage relatively close to the source of I / O operations directed at the data set. Similarly, some QoS storage services may be incompatible with other storage services, such that a user may be prevented from selecting between two conflicting or otherwise incompatible services.
[0210] Another example of a storage service that may be presented to a user, selected by the user, and ultimately applied to a dataset associated with the user may include one or more data protection services. Such data protection services may be embodied as services that may be provided to a consumer of the data protection service (i.e., a user) to ensure, for example, that the user's dataset is protected and disseminated in a particular manner. For example, one or more data protection services may be provided to a user to ensure that the user's dataset is managed in a manner that limits how personal data can be used or disseminated.
[0211] Consider an example in which a particular data protection service is designed to ensure that a user's dataset is not disseminated outside of a particular organization associated with the user. To provide this particular data protection service, the data protection service may be presented to the user (e.g., via a GUI) and selected by the user. In response to receiving the selection of the particular data protection service, one or more storage service policies may be applied to the dataset associated with the user to perform the particular data protection service. For example, a storage service policy may be applied that requires that a dataset not be replicated, backed up, or otherwise stored in a public cloud such as Amazon AWS™. Similarly, a storage service policy may be applied that requires that specific credentials be provided to access the dataset, such as credentials that can be used to verify that the requester is an authorized member of the particular organization associated with the user.
[0212] The reader will understand that many other forms of data protection services may be provided and implemented by embodiments of the present disclosure. Indeed, a data protection service may operate in concert with one or more other storage services. For example, a particular data protection service may operate in concert with one or more data compliance services, as one or more regulations may create requirements for restricting access to or sharing of data. For example, a user selecting a particular data compliance service that requires that certain data cannot be shared may automatically trigger a particular data protection service that can deliver the required level of data privacy. Similarly, some data protection services may be incompatible with some other storage services, such that a user may be prevented from selecting two competing or otherwise incompatible services.
[0213] Another example of a storage service that may be presented to a user, selected by the user, and ultimately applied to a dataset associated with the user may include one or more virtualization management services (including container orchestration). Such virtualization management services may be embodied as services that may be provided to consumers of the virtualization management services (i.e., users) to, for example, ensure that the user's virtualized resources are managed according to predetermined policies. For example, one or more virtualization management services may be provided to a user to ensure that the user's datasets may be presented in a manner such that the datasets are accessible by one or more virtual machines or containers associated with the user, even if the underlying datasets reside in storage that may not normally be available to the virtual machines or containers. For example, a virtualization management service may be configured to present the datasets as part of a virtual volume made available to a virtual machine, container, or some other form of virtual execution environment. Additionally, one or more virtualization management services may be provided to a user to ensure that the user's virtualized execution environment can be backed up and restored in the event of a failure. For example, a virtualization management service may restore an image of a virtual machine or capture state information associated with a virtual machine so that the virtual machine can be restored to its previous state in the event of a virtual machine failure. In other embodiments, a virtualization management service may be used to manage virtualized execution environments associated with users and provide storage resources to such virtualized execution environments.
[0214] Consider an example where a particular virtualization management service is designed to provide persistent storage to containerized applications that otherwise cannot retain data beyond the life of the container itself, and data associated with a particular container is retained for 24 hours after the container is destroyed. To provide this particular virtualization management service, one or more storage service policies can be applied to a dataset or virtualization environment associated with a user to perform the particular virtualization management service. For example, a storage service policy may be applied that creates and configures virtual volumes that can be accessed by the container, and the virtual volumes are backed by physical storage on a physical storage system. In such an example, when a container is destroyed, the storage service policy may also include a rule that prevents a garbage collection process or some other process from deleting the contents of the physical storage on the physical storage system used to back up the virtual volumes for at least 24 hours.
[0215] The reader will understand that many other forms of virtualization management services may be provided and implemented by embodiments of the present disclosure. Indeed, a virtualization management service may operate in coordination with one or more other storage services. For example, a particular virtualization management service may operate in coordination with one or more QoS storage services because a virtualized environment (e.g., virtual machine, container) that is given access to persistent storage via a virtualization management service may also have performance requirements that can be met through the enforcement of a particular QoS storage service. For example, a user selecting a particular QoS storage service that requires a particular entity to receive access to storage resources that can meet specific performance requirements may trigger a particular virtualization management service if the entity is a virtualized entity. Similarly, some virtualization management services may be incompatible with some other storage services, and as a result, a user may be prevented from selecting two conflicting or otherwise incompatible services.
[0216] Another example of a storage service that may be presented to a user, selected by the user, and ultimately applied to a dataset associated with the user may include one or more fleet management services. Such fleet management services may be embodied, for example, as services that may be provided to a fleet management service consumer (i.e., a user) to ensure that the user's storage resources (physical, cloud-based, and combinations thereof) and associated resources are managed in a particular manner. For example, one or more fleet management services may be provided to a user to ensure that the user's dataset is distributed across a fleet of storage systems to best suit the performance needs associated with the dataset. For example, a production version of a dataset may be placed on a relatively high-performance storage system to ensure that the dataset can be accessed using relatively low-latency operations (e.g., read, write), while a test / development version of the dataset may be placed on a relatively low-performance storage system because accessing the dataset using relatively high-latency operations (e.g., read, write) may be acceptable in a test / development environment. In other examples, one or more fleet management services may be provided to a user to ensure that the user's datasets are distributed in a manner that achieves load balancing goals such that some storage systems are not overloaded and other storage systems are not underutilized, ensure that datasets are distributed in a manner that achieves a high level of data reduction (e.g., grouping similar datasets together in hopes of achieving better data deduplication than occurs with a random distribution of datasets), ensure that datasets are distributed in a manner that complies with data compliance regulations, etc.
[0217] Consider an example in which a particular fleet management service is designed to ensure that a user's datasets are distributed in such a way that the datasets are placed on storage systems that are physically closest to the hosts that most frequently access the datasets. To provide this particular fleet management service, one or more storage service policies may be applied that require the location of the hosts that most frequently access the particular datasets to be considered when placing the datasets. Furthermore, the one or more storage service policies that may be applied may further require that if a different host happens to be the host that most frequently accesses the particular dataset, the particular dataset should be replicated to the storage system that is most physically closest to the different hosts.
[0218] The reader will understand that many other forms of fleet management services may be provided and implemented by embodiments of the present disclosure. Indeed, a fleet management service may operate in cooperation with one or more other storage services. For example, a particular fleet management service may operate in cooperation with one or more data compliance services because its ability to move datasets in such a manner may be limited by regulatory requirements enforced by one or more data compliance services. For example, if a user selects a particular data compliance service that limits its ability to move datasets, the particular fleet management service may consider only target storage systems to which it can move datasets without violating policies enforced by the selected data compliance service when the fleet management system is evaluating where datasets residing on source storage systems should be moved in pursuit of some fleet management objective. Similarly, some fleet management services may be incompatible with some other storage services, such that a user may be prevented from selecting two conflicting or otherwise incompatible services.
[0219] Another example of a storage service that may be presented to a user, selected by the user, and ultimately applied to a dataset associated with the user may include one or more cost optimization services. Such cost optimization services may be embodied as services that may be provided to consumers of the cost optimization service (i.e., users) to ensure, for example, that the user's datasets, storage systems, and other resources are managed in a manner that minimizes costs to the users. For example, one or more cost optimization services may be provided to a user to ensure that the user's datasets are replicated to minimize costs (e.g., in dollar terms) associated with replicating data from a source storage system to any of multiple available target storage systems, ensure that the user's datasets are managed to minimize costs associated with storing the datasets, ensure that the user's storage systems or other resources are managed in a manner that reduces power consumption costs associated with operation of the storage systems or other resources, or otherwise, including ensuring that the user's datasets, storage systems, or other resources are managed to minimize or reduce the cumulative cost of multiple expenses associated with the datasets, storage systems, or other resources. Additionally, the one or more cost optimization services may further take into account contractual costs, such as, for example, financial penalties associated with violating a service level agreement associated with a particular user. Indeed, costs associated with implementing an upgrade or performing some other action may also be taken into account, along with many other forms of costs that may be associated with providing storage services and data solutions to customers.
[0220] Consider an example in which a particular cost-optimization service is designed to ensure that a user's dataset is managed such that the cost of storing the dataset is minimized, despite the fact that the user has a separate requirement that the dataset be stored in a local on-premises storage system and mirrored to at least one other storage resource. For example, the dataset may be mirrored to the cloud or to an off-site storage system. To provide this particular cost-optimization service, a storage service policy may be applied that requires that, for each possible replication target, the cost associated with transmitting the dataset to the replication target and the cost associated with storing the dataset on the replication target must be taken into account. In such an example, enforcing such a storage service policy may result in the dataset being mirrored to the replication target at the lowest expected cost.
[0221] The reader will understand that many other forms of cost optimization services may be provided and implemented by embodiments of the present disclosure. Indeed, a cost optimization service may operate in cooperation with one or more other storage services. For example, a particular cost optimization service may operate in cooperation with one or more QoS storage services, because the ability to store a dataset within a particular storage resource may be limited by performance requirements associated with the QoS storage service. For example, a user selecting a particular QoS storage service that creates a requirement that a dataset must be accessible within a certain maximum latency may limit the ability of the cost optimization service to consider all possible storage resources as possible locations where the dataset may be stored, because some storage resources (or combinations of the storage resource and other resources, such as networking resources required to access the storage resource) may not be capable of providing the level of performance required by the QoS storage service. Thus, a particular cost optimization service may only consider target storage systems in which the dataset resides and that can still be accessed according to the requirements of the QoS storage service selected for the dataset. Similarly, some cost optimization services may be incompatible with some other storage services, such that a user may be prevented from selecting two conflicting or otherwise incompatible services.
[0222] Another example of a storage service that may be presented to a user, selected by the user, and ultimately applied to a dataset associated with the user may include one or more workload placement services. Such workload placement services may be embodied as services that may be provided to consumers of workload placement services (i.e., users) to ensure, for example, that the user's dataset is managed to comply with various requirements regarding where data is stored within a system that includes different storage resources. For example, one or more workload placement services may be provided to a user to ensure that the user's dataset is managed in a manner that load-balances access to the data across different storage resources, ensure that the user's dataset is managed in a manner that optimizes a particular performance metric (e.g., read latency, write latency, data reduction) for a selected dataset, ensure that the user's dataset is managed in a manner such that mission-critical datasets are unlikely to be unavailable or suffer relatively long access times, etc.
[0223] Consider an example in which a particular workload placement service is designed to achieve a predetermined load balancing goal across three on-premises storage systems associated with a particular user. To provide this particular workload placement service, the three storage systems may be periodically monitored to apply storage service policies that require ensuring that each storage system serves a relatively similar number of IOPS and stores a relatively similar amount of data. In such an example, if a first storage system stores a relatively large amount of data but serves a relatively small number of IOPS on a third storage system, enforcing the storage service policies associated with the particular workload placement service may result in moving a relatively large (e.g., GB) but infrequently accessed data set on the first storage system to the third storage system, as well as moving a relatively small (e.g., GB) but frequently accessed data set stored on the third storage system to the first storage system, all in the pursuit of ensuring that each storage system serves a relatively similar number of IOPS and stores a relatively similar amount of data.
[0224] The reader will understand that many other forms of workload placement services may be provided and implemented by embodiments of the present disclosure. Indeed, a workload placement service may operate in cooperation with one or more other storage services. For example, a particular workload placement service may operate in cooperation with one or more QoS storage services, because the ability to store a dataset within a particular storage resource may be limited by performance requirements associated with the QoS storage service. For example, a user selecting a particular QoS storage service that creates a requirement that a dataset must be accessible within a certain maximum latency may limit the ability of the workload placement service to load balance across storage resources, because some storage resources (or combinations of storage resources with other resources, such as networking resources, required to access the storage resources) may not be able to provide the performance level required by the QoS storage service. Thus, a particular workload placement service may only consider target storage systems in which the dataset resides and that can still be accessed according to the requirements of the QoS storage service selected for the dataset. Similarly, some workload placement services may be incompatible with some other storage services, such that a user may be prevented from selecting two conflicting or otherwise incompatible services.
[0225] Another example of a storage service that may be presented to a user, selected by the user, and ultimately applied to a dataset associated with the user may include one or more dynamic scaling services. Such dynamic scaling services may be embodied, for example, as services that are provided to consumers (i.e., users) of the dynamic scaling services to ensure that storage resources associated with the user's dataset are scaled up and down as needed. For example, one or more dynamic scaling services may be provided to a user to ensure that the user's dataset and storage resources are managed in a manner that satisfies various objectives that may be achieved by the scaling means.
[0226] Consider an example in which a particular dynamic scaling service is designed to ensure that a user's mission-critical dataset is managed in such a way that no dataset resides on any storage resource that is utilized above 85% in terms of storage capacity or IOPS capacity. To provide this particular dynamic scaling service, a storage service policy is applied that may require: 1) that the storage resource be scaled up (if possible) when this utilization threshold is reached; or 2) that the workload be relocated when this threshold is reached to bring the storage resource storing the mission-critical dataset below 75% in terms of storage capacity and IOPS capacity. For example, if the dataset resides on a cloud-based storage system as described above, the cloud-based storage system may be scaled by adding additional virtual drives (i.e., cloud computing instances with local storage), the cloud-based storage system may be scaled by using more powerful cloud computing instances to run the storage controller application, etc. Alternatively, if the dataset resides on a physical storage system that cannot be immediately scaled, some of the dataset may be migrated off the physical storage system until utilization levels become acceptable.
[0227] The reader will understand that many other forms of dynamic scaling services can be provided and implemented by embodiments of the present disclosure. Indeed, a dynamic scaling service can work in concert with one or more other storage services. For example, a particular dynamic scaling service can work in concert with one or more QoS storage services, because the ability to provide a particular level of performance, such as requested by a QoS storage service, may depend on having appropriately scaled storage resources (or other resources). For example, a user selecting a particular QoS storage service that creates a requirement that a data set must be accessible within a particular latency maximum may immediately trigger one or more dynamic scaling services needed to scale resources so that the latency target can be met. Similarly, some dynamic scaling services may be incompatible with some other storage services, such that a user may be prevented from selecting two conflicting or otherwise incompatible services.
[0228] Another example of a storage service that may be presented to a user, selected by the user, and ultimately applied to a dataset associated with the user may include one or more performance optimization services. Such performance optimization services may be embodied as services that may be provided to a consumer (i.e., a user) of the performance optimization service to ensure, for example, that the user's dataset, storage resources, and other resources are managed to maximize performance as measured by various possible metrics. For example, one or more performance optimization services may be provided to a user to ensure that the user's storage resources are maximized in terms of the aggregate amount of IOPS that can be served, to ensure that the lifespan of different storage resources is maximized through the application of wear-leveling policies, by ensuring that the user's storage resources are managed to minimize total power consumption, by ensuring that the user's storage resources are managed to ensure resource uptime, or in other ways. Similarly, one or more performance optimization services may be provided to a user to ensure that the user's dataset is accessible according to various performance goals. For example, a user's dataset may be managed to provide the best performance in terms of IOPS for a particular dataset, a user's dataset may be managed to provide the best performance in terms of data reduction for a particular dataset, a user's dataset may be managed to provide the best performance in terms of availability for a particular dataset, or may be managed in some other manner.
[0229] Consider an example in which a particular performance optimization service is designed to ensure that users' storage resources are managed in a manner that collectively maximizes the amount of data that can be collectively stored on the storage resources. For example, a storage service policy may be applied that requires that a particular type of data be stored on a storage resource that implements a compression method likely to achieve the best data compression results. For example, if one storage system utilizes a compression algorithm that is more effective at compressing text data and a second storage system utilizes a compression algorithm that is more effective at compressing video data, implementing the storage service policy may result in the video data being stored on the second storage system and the text data being stored on the first storage system. Similarly, a storage service policy may be applied that requires data from similar host applications to be stored on the same storage resource to improve the level of data deduplication that can be achieved. For example, if data from a database application may be more effectively deduplicated when deduplicated against data from other database applications, and data from an image processing application may be more effectively deduplicated when deduplicated against data from other image processing applications, implementing a storage service policy may result in all data from the database application being stored on a first storage system and all data from the image processing application being stored on a second storage system in pursuit of a better deduplication rate than would be achieved by storing data from each type of application on the same storage system. By achieving a better data reduction rate in the backend storage system, more data can be stored on the storage resources from a user's perspective.
[0230] For example, using the compression example above, if a first storage system can compress some data from an uncompressed size of 1 TB to a compressed size of 300 GB, while a second storage system can only compress the same 1 TB of data to a compressed size of 600 GB (because the storage systems use different compression algorithms), then storing the data on the second storage system will require consuming an additional 300 GB of storage from the backend pool of storage systems with fixed physical capacity, and therefore the amount of available storage from the user's perspective will appear different (although their logical capacity can be improved through intelligent placement of data).
[0231] The reader will understand that many other forms of performance optimization services may be provided and implemented by embodiments of the present disclosure. Indeed, a performance optimization service may operate in coordination with one or more other storage services. For example, a particular performance optimization service may operate in coordination with one or more QoS storage services, because the ability to provide a particular level of performance required by a QoS storage service may limit the ability to place a dataset on a particular storage resource. For example, if a user selects a particular QoS storage service that creates a requirement that a dataset must be accessible within a particular latency maximum, but also selects a particular performance optimization service designed to maximize logical storage capacity, only storage resources that can meet both requirements may be candidates for receiving the dataset, even if the other storage resource may be able to provide better results with respect to one service (while not being able to meet the requirements of another service). For example, a particular storage system that can only provide relatively high I / O latency may be able to perform excellent data compression of the dataset by supporting a highly efficient compression algorithm for that dataset, but the dataset may not be able to be placed on that particular storage system (because placing the dataset in such a manner would violate QoS policy). Thus, some performance optimization services may be incompatible with some other storage services, and as a result, a user may be prevented from choosing between two competing or otherwise incompatible services.
[0232] Another example of a storage service that may be presented to a user, selected by the user, and ultimately applied to a dataset associated with the user may include one or more network connectivity services. Such network connectivity services may be embodied as services that may be provided to consumers of network connectivity services (i.e., users) to ensure, for example, that the user's datasets, storage resources, networking resources, and other resources are managed to adhere to various connectivity requirements. For example, one or more network connectivity services may be provided to a user to ensure that the user's datasets are reachable via a predetermined number of networking paths that are independent of any shared hardware or software components. Similarly, one or more network connectivity services may be provided to a user to ensure that storage resources that store the user's datasets can be communicated via data communication paths that meet specific requirements, as well as many other requirements, by ensuring that the user's datasets are managed in a manner that informs host applications of the optimal data communication path for accessing the datasets, to ensure that the user's datasets are managed in a manner that is reachable only via secure data communication channels, and so on.
[0233] Consider an example in which a particular network connectivity service is designed to ensure that a user's dataset is reachable via a predetermined number of networking paths that are independent of any shared hardware or software components. For example, a storage service policy may be applied that requires that the dataset reside on at least two separate storage resources (e.g., two separate storage systems) that may be reachable from application hosts or other devices that access the dataset via separate data communications networks. To that end, applying the storage service policy may cause the dataset to be replicated from one storage resource to another, applying the storage service policy may activate a mirroring mechanism to ensure that the dataset resides on both storage resources, or some other mechanism may be used to enforce the policy.
[0234] The reader will understand that many other forms of network connectivity services can be provided and implemented by embodiments of the present disclosure. Indeed, a network connectivity service can operate in coordination with one or more other storage services. For example, a particular network connectivity service can operate in coordination with one or more replication services, QoS services, data compliance services, and other services, because the ability to position a dataset to comply with a particular network connectivity service can be limited based on the ability to replicate the dataset in accordance with the replication service, further based on the ability to meet performance requirements guaranteed by the QoS service, and further based on the ability to position the dataset to comply with one or more data compliance services. For example, in the example above where the network connectivity service required that a dataset reside on at least two separate storage resources (e.g., two separate storage systems) that could be reachable from an application host or other device accessing the dataset via separate data communications networks, a combination of storage resources could be selected only if other services could also be provided. In some embodiments, if two storage resources that could not deliver all of the selected services were not available, the user may be prompted to eliminate some of the selected services, may select two storage systems that come closest to being able to deliver the selected services using a best-fit scheme, or may take some other action. Thus, some network connectivity services may be incompatible with some other storage services, and as a result, the user may be prevented from selecting two conflicting or otherwise incompatible services.
[0235] Another example of a storage service that may be presented to a user, selected by the user, and ultimately applied to a dataset associated with the user may include one or more data analysis services. Such data analysis services may be embodied as services that may be offered to consumers of the data analysis services (i.e., users) to provide data analysis capabilities on the user's dataset, storage resources, and other resources. For example, one or more data analysis services may be provided to analyze the content of the user's dataset, cleanse the content of the user's dataset, perform data collection operations to create or augment the user's dataset, etc.
[0236] The reader will understand that many other forms of data analysis services may be provided and implemented by embodiments of the present disclosure. Indeed, a data analysis service may operate in cooperation with one or more other storage services. For example, a particular data analysis service may operate in cooperation with one or more QoS storage services, such that access to a data set for purposes of performing data analysis may be given lower priority than more traditional access of the data (e.g., user-initiated reads and writes) to avoid violating performance guarantees set forth in the QoS storage services. Thus, some data analysis services may be incompatible with some other storage services, and as a result, a user may be prevented from choosing between two competing or otherwise incompatible services.
[0237] Another example of a storage service that may be presented to a user, selected by the user, and ultimately applied to a dataset associated with the user may include one or more data portability services. Such data portability services may be embodied as services that may be offered to a consumer of the data portability service (i.e., a user) to, for example, enable the user to perform various data movement, data transformation, or similar processes on the user's dataset. For example, one or more data portability services may be provided to a user to enable the user to migrate their dataset from one storage resource to another, enable the user to convert their dataset from one format (e.g., block data) to another format (e.g., object data), enable the user to consolidate data, enable the user to transfer their dataset from one data controller (e.g., a first cloud service vendor) to another data controller (e.g., a second cloud service vendor), enable the user to convert their dataset from complying with a first set of regulations to complying with a second set of regulations, etc.
[0238] Consider an example in which a particular data portability service is designed to allow a user to transfer their dataset from a first data controller (e.g., a first cloud service vendor) to a second data controller (e.g., a second cloud service vendor). To provide this particular data portability service, a storage service policy may be applied that periodically converts the dataset to be compatible with the second data controller's infrastructure. The reader will understand that many other forms of data portability services may be provided and implemented by embodiments of the present disclosure. Indeed, a data portability service may operate in coordination with one or more other storage services. For example, a particular data portability service may operate in coordination with one or more data compliance services, such that the migration of a dataset from a first data controller to a second data controller may be restricted to prevent the user from violating regulatory compliance set forth in one or more data compliance services. Thus, some data portability services may be incompatible with some other storage services, which may prevent a user from choosing between two competing or otherwise incompatible services.
[0239] Another example of a storage service that may be presented to a user, selected by the user, and ultimately applied to a dataset associated with the user may include one or more upgrade management services. Such upgrade management services may be embodied, for example, as services that may be provided to consumers (i.e., users) of data compliance services to ensure that the user's datasets, storage resources, and other resources can be upgraded or kept up-to-date as various updates become available. For example, one or more upgrade management services may be provided to a user to ensure that the user's storage resources are upgraded upon the occurrence of certain thresholds (e.g., age, utilization), ensure that system software is upgraded as patches, new releases become available, upgrade cloud components, ensure that new cloud service offerings become available, ensure that storage-related resources such as file systems are upgraded when upgrades or updates become available, etc.
[0240] Consider an example in which a particular upgrade management service is designed to ensure that a user's storage resources are managed in a manner that ensures that new software updates to the user's storage system are applied when new software updates become available. To provide this particular upgrade management service, a storage service policy may be applied that requires the storage system to periodically check for updates, download any updates, and install the updates within 24 hours of their availability. The reader will understand that many other forms of upgrade management services may be provided and implemented by embodiments of the present disclosure. Indeed, an upgrade management service may operate in coordination with one or more other storage services. For example, a particular upgrade management service may operate in coordination with one or more QoS services, such that updates or upgrades are applied only when QoS requirements can be maintained by the resource being upgraded or by some other resource. Thus, some upgrade management services may be incompatible with some other storage services, which may prevent a user from choosing between two conflicting or otherwise incompatible services.
[0241] Another example of a storage service that may be presented to a user, selected by the user, and ultimately applied to a dataset associated with the user may include one or more data security services. Such data security services may be embodied as services that may be provided to a consumer of the data security service (i.e., a user) to ensure, for example, that the user's datasets, storage resources, and other resources are managed to adhere to various security requirements. For example, one or more data security services may be provided to a user to ensure that the user's datasets are encrypted according to particular standards both at rest (when stored on the storage resources) and in motion so that end-to-end encryption is achieved. In practice, the one or more data security services may include guarantees that describe how data is protected at rest, guarantees that describe how data is protected in motion, guarantees that describe the private / public key system used, guarantees that describe how access to the datasets or resources is restricted, etc.
[0242] Consider an example in which a particular data security service is designed to ensure that a data set stored on a particular storage resource is encrypted using a key maintained on a resource other than the storage resource, such as a key server. To provide this particular data security service, a storage service policy may be applied that requires the storage resource to request a key from a key server, encrypt the data set (or any unencrypted portions thereof), and delete the encryption key whenever the data set is modified (e.g., via a write). Similarly, to service a read, the storage resource may need to request a key from a key server, decrypt the data set, and delete the encryption key. The reader will understand that many other forms of data security services can be provided and implemented by embodiments of the present disclosure. Indeed, a data security service can operate in coordination with one or more other storage services. For example, a particular data security service can operate in coordination with one or more QoS services, such that only certain QoS services may be available when a particular data security service is selected because requirements for implementing various security functions may limit the extent to which high performance guarantees can be made. Thus, some data security services may be incompatible with some other storage services, and as a result, a user may be prevented from choosing between two competing or otherwise incompatible services.
[0243] Another example of a storage service that may be presented to a user, selected by the user, and ultimately applied to a dataset associated with the user may include one or more integrated system management services. Such integrated system management services may be embodied as services that may be provided to consumers (i.e., users) of the integrated system management services, for example, to ensure that the user's datasets, storage resources, and other resources within the integrated system are managed in a manner that complies with certain policies. For example, one or more integrated system management services may be provided to a user to ensure that an integrated infrastructure including storage resources and one or more GPU servers designed for AI / ML applications can be managed in a particular manner. Similarly, one or more integrated system management services may be provided to a user to ensure that an integrated infrastructure including storage resources and on-premises cloud infrastructure (e.g., Amazon's warehouse) can be managed in a particular manner. For example, one or more integrated system management services may ensure that I / O operations directed to storage resources and initiated by GPU servers within the aforementioned integrated infrastructure are prioritized over I / O operations initiated by devices external to the integrated infrastructure. The reader will understand that many other forms of integrated system management services may be provided and implemented by embodiments of the present disclosure. In fact, an integrated systems management service may operate in concert with one or more other storage services. Thus, some integrated systems management services may be incompatible with some other storage services, and as a result, a user may be prevented from choosing between two competing or otherwise incompatible services.
[0244] Another example of a storage service that may be presented to a user, selected by the user, and ultimately applied to a dataset associated with the user may include one or more application development services. Such application development services may be embodied as services that may be offered to consumers (i.e., users) of data compliance services, for example, to facilitate application development and testing, as well as to perform any other aspects of the application development and testing cycle. For example, one or more application development services may be offered to a user to enable the user to quickly clone a production dataset for development purposes, clone a production dataset with obfuscated personally identifiable information, spin up additional virtual machines or containers for testing, manage all of the connectivity required between test execution environments and the datasets utilized by such environments, and the like. To provide such application development services, one or more storage service policies may be applied to a dataset associated with the user to perform the particular application development service. For example, a storage service policy may be applied that creates a clone of a production dataset upon user request, where personally identifiable information in the dataset is obfuscated in the clone, and then stores the clone on storage resources available for development and testing operations.
[0245] The reader will understand that many other forms of application development services may be provided and implemented by embodiments of the present disclosure. Indeed, an application development service may operate in concert with one or more other storage services. For example, a particular application development service may operate in concert with one or more replication policies such that clones of a production data set may only be sent to non-production environments (e.g., development and test environments). Thus, some application development services may be incompatible with some other storage services, and as a result, a user may be prevented from choosing between two competing or otherwise incompatible services.
[0246] While examples are given above in which a user may select multiple services and compatibility may need to be established between the selected services, the reader will understand that there are many other combinations of services (as well as individual services) that may be presented to a user, selected by the user (when such selections are received by one or more edge management services 406 or similar mechanisms), and ultimately applied to a dataset, storage resource, or some other resource associated with the user. The reader will further understand that many of the example storage services (and other services) described above may include some level of overlap and may be associated with similar, related, or even identical storage service policies.
[0247] The reader will further appreciate that various mechanisms may be used to attach one or more storage service policies to a particular dataset. For example, metadata identifying the particular storage service policy to which the dataset is subject may be attached to the dataset. Alternatively, a centralized repository may be maintained that associates an identifier for each dataset with the storage service policy to which that dataset is subject. Similarly, various devices may maintain information describing the datasets they process. For example, a storage system may maintain information describing the storage service policy to which each dataset stored within the storage system is subject, networking equipment may maintain information describing the storage service policy to which each dataset traversing the networking equipment is subject, etc. Alternatively, such information may be maintained elsewhere and accessible to various devices. In other embodiments, other mechanisms may be used to attach one or more storage service policies to a particular dataset.
[0248] The storage systems 374a, 374b, 374c, 374d, 374n in the fleet of storage systems 376 may be collectively managed, for example, by one or more fleet management modules. The fleet management modules may be part of or separate from the system management services module 384 depicted in FIG. 3E. The fleet management modules may perform tasks such as monitoring the health of each storage system in the fleet, initiating updates or upgrades on one or more storage systems in the fleet, migrating workloads for load balancing or other performance purposes, and many other tasks. Accordingly, and for many other reasons, the storage systems 374a, 374b, 374c, 374d, 374n may be coupled to one another via one or more data communication links to exchange data between the storage systems 374a, 374b, 374c, 374d, 374n.
[0249] The storage systems described herein may support various forms of data replication. For example, two or more of the storage systems may synchronously replicate a dataset between each other. In synchronous replication, separate copies of a particular dataset may be maintained by multiple storage systems, but all accesses (e.g., reads) of the dataset should yield consistent results regardless of which storage system the access is directed to. For example, reads directed to any of the storage systems synchronously replicating the dataset must return identical results. Thus, updates to versions of a dataset need not occur at exactly the same time, but precautions must be taken to ensure consistent access to the dataset. For example, if an update (e.g., a write) directed to a dataset is received by a first storage system, the update may be acknowledged as complete only if all storage systems synchronously replicating the dataset have applied the update to their copies of the dataset. In such examples, synchronous replication may be performed through the use of I / O forwarding (e.g., a write received at a first storage system is forwarded to a second storage system), communication between storage systems (e.g., each storage system indicates it has completed the update), or in other ways.
[0250] In other embodiments, datasets may be replicated through the use of checkpoints. In checkpoint-based replication (also referred to as "near-synchronous replication"), a set of updates to a dataset (e.g., one or more write operations directed to a dataset) may occur between different checkpoints such that the dataset is updated to a particular checkpoint only if all updates to the dataset prior to the particular checkpoint have been completed. Consider an example in which a first storage system stores a live copy of a dataset being accessed by a user of the dataset. In this example, assume that the dataset is being replicated from the first storage system to a second storage system using checkpoint-based replication. For example, the first storage system may send a first checkpoint (time t=0) to the second storage system, then send a first set of updates to the dataset, then send a second checkpoint (time t=1), then send a second set of updates to the dataset, then send a third checkpoint (time t=2). In such an example, if the second storage system has implemented all updates in the first set of updates, but has not yet implemented all updates in the second set of updates, the copy of the dataset stored on the second storage system may be up to the second checkpoint. Alternatively, if the second storage system has implemented all updates in both the first set of updates and the second set of updates, the copy of the dataset stored on the second storage system may be up to the third checkpoint. The reader will understand that various types of checkpoints (e.g., metadata-only checkpoints) may be used, and that checkpoints may be distributed based on various factors (e.g., time, number of operations, RPO settings), etc.
[0251] In other embodiments, datasets may be replicated through snapshot-based replication (also referred to as "asynchronous replication"). In snapshot-based replication, a snapshot of a dataset may be sent from a replication source, such as a first storage system, to a replication target, such as a second storage system. In such an embodiment, each snapshot may include the entire dataset or a subset of the dataset, e.g., only the portions of the dataset that have changed since the last snapshot was sent from the replication source to the replication target. The reader will understand that snapshots may be sent on-demand, based on a policy that takes into account various factors (e.g., time, number of operations, RPO settings), or in some other manner.
[0252] The storage systems described above, alone or in combination, can be configured to function as continuous data protection stores. Continuous data protection stores are features of storage systems that record updates to a dataset, allowing a consistent image of the dataset's previous contents to be accessed at a low time granularity (often on the order of seconds or even less), stretching back a reasonable period of time (often hours or days). They allow access to very recent consistent points in time for a dataset, and also allow access to points in time of a dataset immediately prior to an event, such as when a portion of the dataset is corrupted or otherwise lost, while retaining a number of updates close to the maximum number of updates immediately prior to that event. Conceptually, they are like a sequence of snapshots of a dataset taken very frequently and retained over an extended period of time, although continuous data protection stores are often implemented quite differently from snapshots. Storage systems that implement continuous data protection stores can further provide means to access these points in time, to access one or more of these points in time as snapshots or as clone copies, or to revert a dataset to one of these recorded points in time.
[0253] Over time, to reduce overhead, some time points held in the continuous data protection store can be merged with other nearby points in time, essentially removing some of these time points from the store. This can reduce the capacity required to store updates. It may also be possible to convert these limited number of time points into snapshots of longer duration. For example, such a store may keep a low-granularity sequence of time points going back a few hours from the present, and merge or remove some time points to reduce overhead up to additional days. Going further back than that, some of these time points can be converted into snapshots that represent a consistent point-in-time picture from just every few hours.
[0254] While some embodiments are described primarily in the context of a storage system, those skilled in the art will recognize that embodiments of the present disclosure may also take the form of a computer program product disposed on a computer-readable storage medium for use with any suitable processing system. Such a computer-readable storage medium may be any storage medium for machine-readable information, including magnetic, optical, solid-state, or other suitable media. Examples of such media include magnetic disks in hard drives or diskettes, compact discs for optical drives, magnetic tape, and other media that will occur to those skilled in the art. Those skilled in the art will readily recognize that any computer system with suitable programming means is capable of performing the steps described herein as embodied in a computer program product. Additionally, those skilled in the art will recognize that while some of the embodiments described herein are directed to software installed and executed on computer hardware, alternative embodiments implemented as firmware or hardware are well within the scope of the present disclosure.
[0255] In some examples, a non-transitory computer-readable medium storing computer-readable instructions may be provided in accordance with the principles described herein. The instructions, when executed by a processor of a computing device, may direct the processor and / or computing device to perform one or more operations, including one or more of the operations described herein. Such instructions may be stored and / or transmitted using any of a variety of known computer-readable media.
[0256] Non-transitory computer-readable media referred to herein may include any non-transitory storage medium that participates in providing data (e.g., instructions) that can be read and / or executed by a computing device (e.g., by a processor of a computing device). For example, non-transitory computer-readable storage media may include, but are not limited to, any combination of non-volatile storage media and / or volatile media. Exemplary non-volatile storage media include, but are not limited to, read-only memory, flash memory, solid-state drives, magnetic storage devices (e.g., hard disks, floppy disks, magnetic, etc.), ferroelectric random-access memory ("RAM"), and optical disks (e.g., compact disks, digital video disks, Blu-ray disks, etc.). Exemplary volatile storage media include, but are not limited to, RAM (e.g., dynamic RAM).
[0257] For further explanation, Figure 4 sets forth a block diagram including an edge management service 382 for delivering storage services, according to some embodiments of the present disclosure. The example depicted in Figure 4 is similar to the example depicted in Figure 4, which also includes one or more storage systems 374a, 374n and one or more host devices 378a, 378n.
[0258] 4 also includes an administrator 404 who can access the edge management service 382 via the management interface 402. The management interface 402 may be embodied, for example, as a user interface that allows the administrator 404 to select particular services provided by the edge management service 382 and the underlying storage systems 374a, 374n. In such an example, the administrator 404 can request services from the edge management service 382 through the management interface 402 via a gateway 406 (e.g., a virtual private network) between the management interface 402 and the edge management service 382. In this particular example, the management interface 402 is provided by resources in a cloud 408 computing environment (e.g., public cloud, private cloud), while the edge management service 382 and the storage systems 374a, 374b are located on-premises 410, such as in a particular customer's data center.
[0259] For further explanation, FIG. 5 sets forth a flowchart illustrating an example method for delivering storage services according to some embodiments of the present disclosure. The method illustrated in FIG. 5 may be performed, at least in part, by edge management service 382. Edge management service 382 may be embodied, for example, as computer program instructions running on virtualized computer hardware such as a virtual machine, as computer program instructions running with a container, or in some other manner. In such examples, one or more data service modules may be running in a public cloud environment such as Amazon AWS™, Microsoft Azure™, or the like. Alternatively, edge management service 382 may be running in a private cloud environment, in a hybrid cloud environment, on dedicated hardware and software such as found in a data center, or in some other environment.
[0260] Edge management service 382 may be configured to at least assist in the process of presenting one or more available data services to a user, receiving a selection of one or more selected data services, and applying one or more data service policies to a data set associated with the user in response to the one or more selected data services, as described in more detail below. Additionally, edge management service 382 may be configured to perform other steps, as described in more detail below. In this manner, edge management service 382 can essentially act as a gateway to physical devices, such as one or more storage systems (e.g., including the storage systems described above and variations thereof), one or more networking devices, one or more processing devices, and other devices that can drive the operation of such devices to provide various storage services.
[0261] The exemplary method depicted in FIG. 5 includes presenting one or more available data services to a user (502). The one or more available data services may be embodied as services that may be provided to a data service consumer (i.e., a user) to manage data associated with the data service consumer. The data services may be applied to various forms of data, including, for example, one or more files in a file system, one or more objects, one or more data blocks collectively forming a dataset, one or more data blocks collectively forming a volume, multiple volumes, etc. The data services may be applied to a dataset selected by the user or, alternatively, to some data selected based on one or more policies or heuristics. For example, some data in a dataset (e.g., data having personally identifiable information) may have one set of data services applied to it, while other data in the same dataset (e.g., data without personally identifiable information) may have a different set of data services applied to it.
[0262] As illustrative examples of available data services that may be presented 502 to a user, data services associated with different levels of data protection may be presented 502 to a user. For example, a user may be presented with data services that, when selected and implemented, assure the user that data associated with that user is protected such that various recovery point objectives (“RPOs”) can be guaranteed. A first available data service, for example, may ensure that a subset of data sets associated with the user are protected such that any data older than five seconds can be recovered in the event of a failure of the primary data store, while a second available data service may ensure that a subset of data sets associated with the user are protected such that any data older than five minutes can be recovered in the event of a failure of the primary data store. The reader will understand that this is just one example of available date services that may be presented 502 to a user. Additional available date services that may be presented 502 to a user are described in more detail below.
[0263] The exemplary method depicted in FIG. 5 also includes receiving (504) a selection of one or more selected data services. The one or more selected data services may represent data services chosen from a set of all available data services that a user desires to apply to a particular data set associated with the user. Continuing with the example above, a first available data service may ensure that a portion of a data set associated with the user is protected such that any data older than five seconds can be restored in the event of a failure of the primary data store, while a second available data service may ensure that a data set associated with the user is protected such that any data older than five minutes can be restored in the event of a failure of the primary data store. The reader will understand, however, that the monetary cost of the first available data service may be higher than the monetary cost of the second available data service, and thus the user may weigh these cost differences and select the data service most appropriate for the user's application, data, etc. The user's selection of one or more selected data services may be received (504), for example, via a GUI presented to the user, as part of a performance tier selected by the user, or in some other manner.
[0264] The exemplary method depicted in FIG. 5 also includes applying 506 one or more data service policies to a dataset associated with the user, depending on the one or more selected data services. The one or more data service policies may be embodied as one or more rules that, for example, when enforced or implemented, cause the associated data service to be provided. Continuing with the above example in which a user selects a first available data service to ensure that a portion of a dataset associated with the user is protected so that any data older than five seconds can be restored in the event of a failure of the primary data store, the one or more data service policies that may be applied 506 may include, for example, a policy that causes a snapshot of the user's dataset to be taken every five seconds and sent to one or more backup storage systems. In this way, even if the primary data store fails or otherwise becomes unavailable, a snapshot was taken just five seconds before the failure and can be restored from the one or more backup storage systems. In other examples, another data service policy may be applied 506 to replay the I / O event that modified the dataset to provide the associated selected data service, or some other data service policy that causes the associated selected data service to be provided 506.
[0265] The reader will understand that in some embodiments, while one or more data service modules may be configured to present one or more available data services to a user (502) and receive one or more selected data service selections (504), the one or more data service modules may not be responsible for applying one or more data service policies to a dataset associated with the user (506). Instead, the one or more data service modules may instruct some other entity (via one or more APIs, via one or more messages, or in some other manner) to apply one or more data service policies to a dataset associated with the user (506). For example, the one or more data service modules may utilize an API provided by the storage system to cause the storage system to perform a backup operation as described above, providing the first available data service described above, such that the dataset associated with the user is protected and any data older than five seconds can be restored in the event of a failure of the primary data store.
[0266] The reader will understand that many data services may be presented to the user (502) and selected by the user, ultimately resulting in one or more data service policies being applied to a data set associated with the user (506) in response to the one or more data services selected by the user. While a non-exhaustive list of data services that may be made available is included above, the reader will understand that additional data services may be made available according to some embodiments of the present disclosure.
[0267] In the exemplary method depicted in FIG. 5 , applying one or more data service policies to a dataset associated with a user (506) may include identifying an initial placement for the dataset (508). The initial placement for the dataset may be embodied, for example, as one or more storage systems that are initially identified as the storage systems on which the dataset should be stored. Identifying the initial placement for the dataset (508) may be performed, for example, by identifying all storage systems (or a combination thereof) that have the necessary resources needed to support the dataset. For example, only storage systems with sufficient capacity (e.g., in terms of MB, GB, TB, etc.) to store the dataset may be selected as candidates for storing the dataset, only storage systems with sufficient performance capacity (e.g., in terms of IOPS for the dataset that can be supported) to service the dataset may be selected as candidates for storing the dataset, or only storage systems with some other form of functionality sufficient for storing the dataset (e.g., the ability to encrypt data, file system support, block storage support, the ability to meet certain availability requirements) may be selected as candidates for storing the dataset. Indeed, the determination of whether a particular storage system is a candidate for storing the dataset may include many criteria. In such an embodiment, each dataset may be paired with metadata describing various requirements associated with the dataset that may be used to evaluate whether a particular storage system (or a particular combination of storage systems) is a candidate for storing the dataset.
[0268] In such an example, once one or more storage systems have been identified as candidates for storing the dataset, each of the candidates may be evaluated to identify the best match based on some criteria. For example, the storage system capable of storing the dataset at the lowest cost may be identified as the storage system for storing the dataset (508). In practice, the determination as to which storage system should be identified as the storage system for storing the dataset (508) may be based on multiple criteria. Such criteria may include, for example, cost, the amount of resources available at each candidate storage system, performance experienced by entities accessing the dataset (e.g., read latency, write latency), etc. In some embodiments, a score may be generated for each candidate storage system that identifies how well the candidate storage system can meet all of the needs or requirements associated with storing the dataset, such that the best match storage system may be identified as the storage system for storing the dataset (508).
[0269] In the exemplary method depicted in FIG. 5 , applying one or more data service policies to a dataset associated with a user (506) may include migrating the dataset (510). In contrast to the example above in which an initial placement of the dataset was identified (508), migrating the dataset (510) may be performed in response to detecting that some change to the overall storage environment has caused the current placement of the dataset to no longer be a sufficient or optimal placement of the dataset. For example, changes to the overall storage environment may have occurred as a result of new workloads being added to the storage environment (e.g., additional datasets being stored in a storage system), as a result of existing workloads growing or shrinking, as a result of a storage resource failing, as a storage resource being added to the storage environment, as a result of one or more storage resources being upgraded, as a result of changes in networking resources utilized to access the storage environment (e.g., a switch being added, a switch being removed), and for many other reasons. Generally, changes to the overall storage environment may have been made as a result of changes to the underlying hardware or to the way the storage environment is utilized.
[0270] In such examples, migrating (510) the dataset may be performed by determining that the current placement of the dataset is no longer sufficient or optimal. Such a determination may be made, for example, by periodically identifying one or more candidate storage systems for storing the dataset, evaluating each candidate, and identifying the best match among the candidates. Thus, in such examples, the determination may be made periodically in a scheduled manner. In alternative embodiments, the occurrence of some event may trigger this process. For example, if the capacity utilization of a particular storage system storing the dataset reaches a threshold, one or more candidate storage systems for storing the dataset may be identified and evaluated to identify the best match among the candidates. Similarly, if one or more storage resources are added to or removed from the overall storage environment, one or more candidate storage systems for storing the dataset may be identified and evaluated to identify the best match among the candidates.
[0271] In the examples described herein, costs or impacts associated with migrating a dataset may be taken into consideration when determining whether to migrate 510 a dataset. For example, if migrating 510 a particular dataset would result in a relatively small benefit, but the costs associated with migrating 510 the dataset were substantial, the migration may not be permitted to occur. Similarly, if migrating 510 a particular dataset would result in a better fit for the dataset, but would cause substantial harm to the ability to service other datasets, the migration may not be permitted to occur.
[0272] 5 , applying 506 one or more data service policies to a dataset associated with a user may include provisioning 512 resources on one or more storage systems. Provisioning 512 resources on one or more storage systems may be performed, for example, by initiating the creation of one or more volumes for storing the dataset through an API exposed by one or more of the storage resources, by initiating the creation of one or more file systems for storing the dataset through an API exposed by one or more of the storage resources, etc. In such an example, provisioning 512 resources on one or more storage systems may also include reserving capacity, I / O bandwidth, or some other aspect of the resources of the storage systems to enable enforcement of a level of service guarantee or similar performance requirement.
[0273] While the above examples relate to embodiments in which resources are provisioned 512 (or reserved) from one or more storage systems, in other embodiments, other forms of resources needed by a user to access their data may also be provisioned or reserved. For example, networking resources may be provisioned or reserved, monitoring resources may be provisioned or reserved, computing resources may be provisioned or reserved, management resources may be provisioned or reserved, etc.
[0274] In the exemplary method depicted in FIG. 5, applying one or more data service policies to a dataset associated with a user (506) may include setting one or more tunables for one or more storage systems (514). Each tunable may be embodied, for example, as a configuration parameter, an operational parameter, or other parameter that affects the operation of one or more storage systems. For example, if a tunable parameter is used to determine how frequently snapshots are taken of a particular dataset, the tunable parameter may be set to enforce a desired RPO for the dataset (514). Similarly, if a tunable parameter is used to determine whether end-to-end encryption should be enforced for access to a particular dataset, the tunable parameter may be set to enable or disable end-to-end encryption for the dataset (514). (End-to-end encryption means that data is encrypted while stored on a storage system and is transferred encrypted over a network as part of accessing the dataset, typically using different encryption keys at rest and in motion.) In such an example, setting 514 one or more tunables for one or more storage systems may be performed through the use of one or more APIs exposed from the storage systems and to the edge management service 382.
[0275] 5 , applying 506 one or more data service policies to a dataset associated with a user may include modifying 516 one or more tunables for one or more storage systems. While the above example relates to one embodiment in which one or more tunables are set 514, in some embodiments, changes may be made that require modifying 516 the tunables. Such changes may include, for example, changes to the entire storage environment (e.g., hardware changes, workload changes), changes to requirements associated with a particular dataset, changes to a set of services selected by a user, etc. In such an example, modifying 516 the tunables may be performed by the edge management service 382, which determines that some tunables should be adjusted and modifies such tunables through one or more APIs available to the edge management service 382.
[0276] In the exemplary method depicted in FIG. 5 , applying one or more data service policies to datasets associated with a user (506) may include performing one or more fleet management operations on one or more storage systems (516). Fleet management operations may include, for example, rebalancing datasets and workloads among storage systems, updating or upgrading software on one or more of the storage systems, etc. For example, permutations of how the workload can be distributed may be evaluated to find different permutations that satisfy different requirements associated with each dataset in an overall optimized manner (e.g., dataset 1 is stored on storage system 1 and dataset 2 is stored on storage system 2, versus dataset 1 is stored on storage system 2 and dataset 2 is stored on storage system 1). For example, if a first arrangement of the datasets results in less capacity consumption than a second arrangement of the datasets (e.g., through improved data deduplication), in one embodiment where the distribution is optimized to conserve capacity, performing fleet management operations (516) may include arranging the datasets according to the first arrangement. In other embodiments, the datasets may be arranged to optimize for other consideration(s).
[0277] 5 , applying 506 one or more data service policies to a dataset associated with a user may include modifying 518 the data transmitted to the host device in response to a request to access the dataset. Modifying 518 the data transmitted to the host device in response to a request to access the dataset can be performed, for example, by the edge management service 382 by masking or removing PII to adhere to a requirement that PII not be shared with the dataset accessor, encrypting / decrypting the data, enforcing a requirement that the dataset be accessed and stored using end-to-end encryption, etc. In such an example, modifying 518 the data transmitted to the host device in response to a request to access the dataset can be performed by the edge management service 382 to offload some requirements from the underlying storage system and instead leverage the edge management service 382 to perform some of the tasks required to deliver a particular service.
[0278] 5, applying 506 one or more data service policies to a dataset associated with a user may include attaching 520 one or more policies to the dataset. In such an example, attaching 520 one or more policies to the dataset may be performed by the edge management service 382, which adds or modifies metadata associated with the dataset, such metadata being used by the storage system to determine which policies apply to the dataset. For example, a first value in a particular metadata field may indicate whether the dataset should be replicated synchronously or asynchronously.
[0279] 5 , applying 506 one or more data service policies to a dataset associated with a user may include enforcing 522 the one or more policies associated with the dataset. While the above-described embodiment relates to an implementation in which the edge management service 382 only informs a storage system (or other resource) about policies to be attached 520 to the dataset, in this example, the edge management service 382 may itself be responsible (at least in part) for enforcing 522 one or more policies associated with the dataset. For example, if a policy defines that only encrypted data should be passed from the storage system to a host device when accessing the dataset, the edge management service 382 may actually encrypt the data to enforce 522 such policy.
[0280] While the above-described embodiments relate to embodiments in which one or more available storage services are presented to a user, a selection of one or more selected storage services is received, and one or more storage service policies for a dataset associated with the user are applied depending on the one or more selected storage services, the reader will understand that in other embodiments, storage classes (rather than services) may be presented and ultimately applied in a similar manner. For example, a "high performance" storage class may be presented that is ultimately delivered through the application of a predetermined set of services associated with the "high performance" storage class. Similarly, a "highly resilient" storage class may be presented that is ultimately delivered through the application of a predetermined set of services associated with the "highly resilient" storage class. In additional embodiments, the exact nature of what is "presented" and made available for selection by a user may take other forms, all of which may be implemented and provided as described herein.
[0281] For further explanation, Figure 6 sets forth a block diagram including an edge management service 382 for role enforcement for storage as a service, according to some embodiments of the present disclosure. The example depicted in Figure 6 is similar to the example depicted in Figure 4, and the example depicted in Figure 6 also includes a gateway 406 and one or more storage systems 374a, 374n. The host devices of Figure 4 are referred to as client devices 678a, 678n in Figure 6. Figure 6 also includes an additional client device 602 communicatively coupled to the edge management service 382 via the gateway 406.
[0282] Each client device (client devices 678a, 678n coupled to edge management service 382 and client device 602 coupled to edge management service 382 via gateway 406) is a computing system used by a storage system client. Each client device 678a, 678n, 602 may be used by or otherwise under the control of a client assigned to a storage consumer role or a storage provider role. A storage consumer is a client that utilizes (i.e., consumes) data storage on storage system 374a, 374n. A storage provider is a client that manages (i.e., provides) storage and storage services for storage consumers. Edge management service 382 of FIG. 6 enforces the roles of storage consumer and storage provider on the storage system as a service. Specifically, edge management service 382 of FIG. 6 manages the roles of storage consumer and storage provider, services data management instructions from storage consumer clients, and services storage management instructions from storage provider clients. As described above, the edge management service 382 provides storage services to clients utilizing client devices 678a, 678n, 602 from storage systems 374a, 374n communicatively coupled to the edge management service 382.
[0283] In this particular example, the edge management service 382 and the fleet 376 of storage systems 374a, 374b are located on-premise 410, such as at a particular customer's data center. In another example, the edge management service 382 and gateway 406 may be provided by a remote cloud computing environment for the storage systems. As used herein, the term "cloud-based" refers to a system or collection of systems that provide services remotely over a data communications link (such as the data communications link described above in FIG. 3A).
[0284] The gateway 406 (also referred to as a network gateway) is hardware, software, or an aggregation of hardware and software through which the client device 602 exchanges information with the edge management service 382. The gateway 406 can provide the client device 602 with secure access to the edge management service 382 and the fleet 376 of storage systems 374 a, 374 n. The gateway 406 can require authorization and authentication of the client device 602 before granting access to the edge management service 382 and the storage systems 374 a, 374 n. Examples of the gateway 406 can include a VPN, a network node, a server, a network bridge, and a network appliance.
[0285] Clients assigned different roles can communicate with the edge management service 382 through different communication paths. For example, as shown in FIG. 6, client devices 678a, 678n communicate directly with the edge management service 382. Clients using client devices 678a, 678n may be assigned the role of storage consumer and may communicate with the edge management service 382 through a storage service module, as shown in FIG. 3E. Also shown in FIG. 6, client device 602 communicates with the edge management service 382 through a gateway 406. Client using client device 602 is assigned the role of storage provider and can access the edge management service 382 through the gateway, allowing the client 602 to execute storage management instructions.
[0286] As used herein, the term “local to” refers to entities that are physically proximate to one another, such as entities within the same data center. The edge management service 382 and gateway 406 may be local to the storage systems 374a, 374b in that the edge management service 382 may be hosted on a system within the same physical building as the hardware components of the storage systems 374a, 374b. Similarly, the terms “remote to” and “remotely” refer to entities that are not physically proximate to one another, such as entities that communicate over a wide area network (e.g., the Internet). The edge management service 382 and gateway 406 may be provided by a cloud environment that is remote to the storage systems 374a, 374b in that the edge management service 382 and gateway 406 may be hosted on one or more systems outside the physical building that houses the hardware components of the storage systems 374a, 374b and may be under the control of entities that are separate and distinct from the entities that control the hardware components of the storage systems 374a, 374b.
[0287] Although the following example methods are depicted as being performed by edge management service 382, role enforcement for storage as a service methods may be performed by other entities in a system separate and distinct from edge management service 382. For example, the system providing storage as a service enforcement may be cloud-based and remote to storage systems 374a, 374b. Alternatively, the system providing storage as a service enforcement may be on-premise and local to storage systems 374a, 374b.
[0288] For further explanation, FIG. 7 sets forth a flowchart illustrating an exemplary method of role enforcement for storage as a service according to some embodiments of the present disclosure. The method of FIG. 7 includes managing (702) a plurality of roles for a storage system 374, including a storage consumer role and a storage provider role, where the storage consumer role is associated with data management instructions that are enabled for the storage consumer role and disabled for the storage provider role, and the storage provider role is associated with storage management instructions that are enabled for the storage provider role and disabled for the storage consumer role. Clients assigned to the storage consumer role can successfully execute some instructions that are unavailable to clients assigned to the storage provider role. Similarly, clients assigned to the storage provider role can successfully execute some instructions that are unavailable to clients assigned to the storage consumer role. The edge management service 382 enforces these restrictions and capabilities for the storage consumer and storage provider roles. A client associated with a particular role refers to a client that has been assigned that role by a management entity. Furthermore, roles can be mutually exclusive, meaning that a client can only be assigned a single role at any given time.
[0289] Managing 702 multiple roles, including storage consumer and storage provider roles, for storage system 374 can be performed by enforcing the roles assigned to each client and disallowing or preventing a client associated with one role from successfully executing instructions specific to a different role. For example, edge management service 382 can disallow or prevent a client assigned the role of storage provider from successfully executing data management instructions, such as instructions to delete data. Similarly, edge management service 382 can disallow or prevent a client assigned the role of storage consumer from successfully executing storage management instructions, such as instructions to modify protection policies on the storage system.
[0290] 7 further includes servicing 704 a data management instruction from a first client associated with a storage consumer role, the data management instruction being an instruction to manipulate data on storage system 374. Servicing 704 a data management instruction from a first client associated with a storage consumer role, the data management instruction being an instruction to manipulate data on storage system 374, may be performed by receiving the data management instruction, verifying that the role assigned to the client sending the data management instruction is the storage consumer role, and executing the data management instruction using storage system 374. As used herein, the term "servicing" may refer to implementing an instruction or instructing another entity in a system to implement the received instruction. Examples of data management instructions include a data write instruction, a data read instruction, a data delete instruction, and a storage class instantiation instruction.
[0291] 7 further includes servicing 706 a storage management command from a second client associated with the storage provider role, the storage management command being a command to manage storage system 374. Servicing 706 a storage management command from a second client associated with the storage provider role, the storage management command being a command to manage storage system 374, may be performed by receiving the storage management command, verifying that the role assigned to the client sending the storage management command is the storage provider role, and executing the storage management command using storage system 374. Examples of storage management commands include a create region command, a create availability zone command, a define storage class command, and a protection policy command.
[0292] For further explanation, Figure 8 sets forth a flowchart illustrating an additional exemplary method of role enforcement for storage as a service according to some embodiments of the present disclosure. The exemplary method depicted in Figure 8 is similar to the exemplary method depicted in Figure 7, and also includes managing (702) a plurality of roles, including a storage consumer role and a storage provider role, for a storage system 374, where the storage consumer role is associated with data management instructions that are enabled for the storage consumer role and disabled for the storage provider role, and the storage provider role is associated with storage management instructions that are enabled for the storage provider role and disabled for the storage consumer role; servicing (704) data management instructions from a first client associated with the storage consumer role, the data management instructions being instructions for manipulating data on the storage system 374; and servicing (706) storage management instructions from a second client associated with the storage provider role, the storage management instructions being instructions for managing the storage system 374.
[0293] In the exemplary method depicted in FIG. 8 , servicing 704 a data management instruction from a first client associated with the storage consumer role, the data management instruction being an instruction to manipulate data on the storage system 374, includes servicing 802 a delete instruction from the first client associated with the storage consumer role. The instruction to manipulate data may specifically be an instruction to delete data from the storage system 374. Servicing 802 a delete instruction from the first client associated with the storage consumer role may be performed by locating data targeted by the delete instruction and marking the data for garbage collection. Clients assigned to the role of storage provider may be prevented or not permitted to delete data from the storage system.
[0294] For further explanation, Figure 9 sets forth a flowchart illustrating an additional exemplary method of role enforcement for storage as a service according to some embodiments of the present disclosure. The exemplary method depicted in Figure 9 is similar to the exemplary method depicted in Figure 7, and also includes managing (702) a plurality of roles, including a storage consumer role and a storage provider role, for a storage system 374, where the storage consumer role is associated with data management instructions that are enabled for the storage consumer role and disabled for the storage provider role, and the storage provider role is associated with storage management instructions that are enabled for the storage provider role and disabled for the storage consumer role; servicing (704) data management instructions from a first client associated with the storage consumer role, the data management instructions being instructions for manipulating data on the storage system 374; and servicing (706) storage management instructions from a second client associated with the storage provider role, the storage management instructions being instructions for managing the storage system 374.
[0295] In the exemplary method depicted in FIG. 9 , servicing 706 a storage management instruction from a second client associated with the storage provider role, the storage management instruction being an instruction to manage the storage system 374, includes servicing 902 ...
Claims
1. 1. A method comprising: managing, by an edge management service for a storage system, a plurality of roles including a storage consumer role and a storage provider role, wherein the edge management service enforces role exclusivity such that data management commands associated with the storage consumer role are enabled for the storage consumer role and disabled for the storage provider role, and storage management commands associated with the storage provider role are enabled for the storage provider role and disabled for the storage consumer role, and the storage consumer role is associated with data management commands and the storage provider role is associated with storage management commands; servicing a data management instruction from a first client associated with the storage consumer role, the data management instruction being an instruction to manipulate data on the storage system; and servicing storage management instructions from a second client associated with the storage provider role, the storage management instructions being instructions for managing the storage system without affecting the data.
2. 2. The method of claim 1, wherein servicing the data management instruction from the first client associated with the storage consumer role, the data management instruction being an instruction to manipulate data on the storage system, comprises servicing a delete instruction from the first client associated with the storage consumer role.
3. 2. The method of claim 1, wherein servicing the storage management instruction from the second client associated with the storage provider role, the storage management instruction being an instruction to manage the storage system, comprises servicing an instruction to modify a protection policy on the storage system from the second client associated with the storage provider role.
4. The method of claim 1 , wherein the storage system is communicatively coupled to the edge management service that provides storage services from the storage system to clients.
5. servicing the data management instructions from the first client associated with the storage consumer role, the data management instructions being instructions for manipulating data on the storage system; receiving, by the edge management service, the data management instructions from the first client, the edge management service providing storage services to the client from the storage system; and verifying, by the edge management service, that the role associated with the first client is the storage consumer role.
6. servicing the storage management instructions from the second client associated with the storage provider role, the storage management instructions being instructions for managing the storage system; receiving, by the edge management service via a gateway, the storage management instructions from the second client, wherein the edge management service provides storage services from the storage system to clients, and the gateway provides the second client with access to the edge management service; and verifying, by the edge management service, that the role associated with the second client is the storage provider role.
7. The method of claim 1 , wherein the data management instructions include a data write instruction, a data read instruction, a data delete instruction, and a storage class instantiation instruction.
8. The method of claim 1 , wherein the storage management instructions include a region creation instruction, an availability zone creation instruction, a storage class definition instruction, and a protection policy instruction.
9. 2. The method of claim 1, wherein servicing the storage management instructions from the second client associated with the storage provider role, the storage management instructions being instructions for managing the storage system, comprises servicing instructions from the second client associated with the storage provider role to implement a storage service that is applied to a dataset on the storage system.
10. 2. The method of claim 1, wherein servicing the storage management instructions from the second client associated with the storage provider role, the storage management instructions being instructions for managing the storage system, comprises servicing a request for metrics describing storage system performance from the second client associated with the storage provider role.
11. 1. An apparatus comprising: a computer processor; and a computer memory operatively coupled to the computer processor, the computer memory, when executed by the computer processor, causing the apparatus to: managing, by an edge management service for a storage system, a plurality of roles including a storage consumer role and a storage provider role, wherein the edge management service enforces role exclusivity such that data management commands associated with the storage consumer role are enabled for the storage consumer role and disabled for the storage provider role, and storage management commands associated with the storage provider role are enabled for the storage provider role and disabled for the storage consumer role, and the storage consumer role is associated with data management commands and the storage provider role is associated with storage management commands; servicing a data management instruction from a first client associated with the storage consumer role, the data management instruction being an instruction to manipulate data on the storage system; servicing storage management instructions from a second client associated with the storage provider role, the storage management instructions being instructions for managing the storage system without affecting the data.
12. 12. The apparatus of claim 11, wherein servicing the data management instruction from the first client associated with the storage consumer role, the data management instruction being an instruction to manipulate data on the storage system, comprises servicing a delete instruction from the first client associated with the storage consumer role.
13. 12. The apparatus of claim 11, wherein servicing the storage management instructions from the second client associated with the storage provider role, the storage management instructions being instructions for managing the storage system, comprises servicing instructions from the second client associated with the storage provider role to modify a protection policy on the storage system.
14. The apparatus of claim 11 , wherein the storage system is communicatively coupled to the edge management service that provides storage services from the storage system to clients.
15. servicing the data management instructions from the first client associated with the storage consumer role, the data management instructions being instructions for manipulating data on the storage system; receiving, by the edge management service, the data management instructions from the first client, the edge management service providing storage services to the client from the storage system; and verifying, by the edge management service, that the role associated with the first client is the storage consumer role.
16. servicing the storage management instructions from the second client associated with the storage provider role, the storage management instructions being instructions for managing the storage system; receiving, by the edge management service via a gateway, the storage management instructions from the second client, wherein the edge management service provides storage services from the storage system to clients, and the gateway provides the second client with access to the edge management service; and verifying, by the edge management service, that the role associated with the second client is the storage provider role.
17. The apparatus of claim 11 , wherein the data management instructions include a data write instruction, a data read instruction, a data delete instruction, and a storage class instantiation instruction.
18. The apparatus of claim 11 , wherein the storage management instructions include region creation instructions, availability zone creation instructions, storage class definition instructions, and protection policy instructions.
19. 12. The apparatus of claim 11, wherein servicing the storage management instructions from the second client associated with the storage provider role, the storage management instructions being instructions for managing the storage system, comprises servicing instructions from the second client associated with the storage provider role to implement a storage service applied to a dataset on the storage system.
20. A computer program product disposed on a computer-readable medium, the computer program product, when executed, causing a computer to: managing, by an edge management service for a storage system, a plurality of roles including a storage consumer role and a storage provider role, wherein the edge management service enforces role exclusivity such that data management commands associated with the storage consumer role are enabled for the storage consumer role and disabled for the storage provider role, and storage management commands associated with the storage provider role are enabled for the storage provider role and disabled for the storage consumer role, and the storage consumer role is associated with data management commands and the storage provider role is associated with storage management commands; servicing a data management instruction from a first client associated with the storage consumer role, the data management instruction being an instruction to manipulate data on the storage system; servicing storage management instructions from a second client associated with the storage provider role, the storage management instructions being instructions for managing the storage system without affecting the data.
Citation Information
Patent Citations
Storage system and storage managing method
JP2006185386A
A trusted extended markup language for trusted computing and data services.
JP2013513834A