Context-driven user interface for a storage system

A direct-mapped flash storage system with non-volatile RAM buffering and zone management within the operating system addresses inefficiencies in conventional storage systems, enhancing reliability and reducing unnecessary write operations.

JP7815428B2Active Publication Date: 2026-02-17PURE STORAGE INC
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
JP2024523170
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2021-10-18
Filing Date
2022-10-18
Publication Date
2026-02-17
Estimated Expiration
2042-10-18

AI Technical Summary

Technical Problem

Conventional storage systems face inefficiencies due to lower-level processes being performed by storage controllers, leading to unnecessary write operations and reduced reliability, especially in flash storage systems.

Method used

Implementing a direct-mapped flash storage system where higher-level processes initiate and control operations without address translation by the storage controller, utilizing non-volatile RAM as a buffer for improved latency and reliability, and managing zones and allocation units within the operating system.

Benefits of technology

Enhances reliability and reduces unnecessary write operations by allowing the operating system to manage storage processes directly, improving efficiency and reducing overhead in flash storage systems.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007815428000001
    Figure 0007815428000001
  • Figure 0007815428000002
    Figure 0007815428000002
  • Figure 0007815428000003
    Figure 0007815428000003
Patent Text Reader

Abstract

1. A context-driven user interface for a storage system, comprising: receiving a request from a user account to access a system interface for the system; identifying at least one significant system characteristic that describes a current aspect of the system; reconfiguring the system interface based on the at least one significant system characteristic; and presenting the reconfigured system interface to a user of the user account.
Need to check novelty before this filing date? Find Prior Art

Description

[Brief explanation of the drawings]

[0001] [Figure 1A] 1 illustrates a first exemplary system for data storage, according to some implementations. [Figure 1B] 1 illustrates a second exemplary system for data storage, according to some implementations. [Figure 1C] 1 illustrates a third exemplary system for data storage, according to some implementations. [Figure 1D] 1 illustrates a fourth exemplary system for data storage, according to some implementations. [Figure 2A] FIG. 1 is a perspective view of a storage cluster having multiple storage nodes and internal storage coupled to each storage node to provide network-attached storage, according to some embodiments. [Figure 2B] FIG. 2 is a block diagram illustrating an interconnect switch coupling multiple storage nodes, according to some embodiments. [Figure 2C] FIG. 2 is a multi-level block diagram illustrating the contents of a storage node and the contents of one of the non-volatile solid-state storage units, according to some embodiments. [Figure 2D] 1 illustrates a storage server environment that uses embodiments of the storage nodes and storage units of some of the previous figures, according to some embodiments. [Figure 2E] FIG. 1 is a blade hardware block diagram illustrating the control plane, compute and storage plane, and authorities interacting with underlying physical resources, according to some embodiments. [Figure 2F] 1 depicts an elasticity software layer within a blade of a storage cluster, according to some embodiments. [Figure 2G] 1 depicts permissions and storage resources within blades of a storage cluster, according to some embodiments. [Figure 3A]1 illustrates an illustration of a storage system coupled for data communication with a cloud service provider, according to some embodiments of the present disclosure. [Figure 3B] 1 illustrates a diagram of a storage system in accordance with some embodiments of the present disclosure. [Figure 3C] An example of a cloud-based storage system according to some embodiments of the present disclosure is described. [Figure 3D] 1 illustrates an exemplary computing device that may be specifically configured to perform one or more of the processes described herein. [Figure 3E] Illustrates an example fleet of storage systems 376 for providing storage services (also referred to herein as "data services"). [Figure 4] 10 depicts a flowchart illustrating an exemplary method for a context-driven user interface for a storage system, according to some embodiments of the present disclosure. [Figure 5] 10A-10C are flowcharts illustrating additional exemplary methods of a context-driven user interface for a storage system, according to some embodiments of the present disclosure. [Figure 6] 10A-10C are flowcharts illustrating additional exemplary methods of a context-driven user interface for a storage system, according to some embodiments of the present disclosure. [Figure 7] 10A-10C are flowcharts illustrating additional exemplary methods of a context-driven user interface for a storage system, according to some embodiments of the present disclosure. [Figure 8] 10A-10C are flowcharts illustrating additional exemplary methods of a context-driven user interface for a storage system, according to some embodiments of the present disclosure. DETAILED DESCRIPTION OF THE INVENTION

[0002] Exemplary methods, apparatus, and products for a context-driven user interface for a storage system according to embodiments of the present disclosure are described with reference to the accompanying drawings, beginning with FIG. 1A. FIG. 1A illustrates an exemplary system for data storage according to some implementations. System 100 (also referred to herein as a "storage system") includes a number of elements for purposes of illustration, not limitation. It should be noted that system 100 may include the same, more, or fewer elements, configured in the same or different ways, in other implementations.

[0003] System 100 includes multiple computing devices 164A-B. The computing devices (also referred to herein as "client devices") may be embodied as, for example, servers in a data center, workstations, personal computers, notebooks, etc. The computing devices 164A-B may be coupled for data communication to one or more storage arrays 102A-B via a storage area network ("SAN") 158 or a local area network ("LAN") 160.

[0004] SAN 158 may be implemented using a variety of data communication fabrics, devices, and protocols. For example, fabrics for SAN 158 may include Fibre Channel, Ethernet, InfiniBand, Serial Attached Small Computer System Interface ("SAS"), etc. Data communication protocols used with SAN 158 may include Advanced Technology Attachment ("ATA"), Fibre Channel Protocol, Small Computer System Interface ("SCSI"), Internet Small Computer System Interface ("iSCSI"), HyperSCSI, Non-Volatile Memory Express ("NVMe") over fabric, etc. It may be noted that SAN 158 is provided for illustration and not limitation. Other data communication couplings may be implemented between computing devices 164A-B and storage arrays 102A-B.

[0005] LAN 160 may also be implemented using a variety of fabrics, devices, and protocols. For example, fabrics for LAN 160 may include Ethernet (802.3), wireless (802.11), etc. Data communication protocols used in LAN 160 may include Transmission Control Protocol ("TCP"), User Datagram Protocol ("UDP"), Internet Protocol ("IP"), HyperText Transfer Protocol ("HTTP"), Wireless Access Protocol ("WAP"), Handheld Device Transport Protocol ("HDTP"), Session Initiation Protocol ("SIP"), Real Time Protocol ("RTP"), etc. LAN 160 may also be connected to the Internet 162.

[0006] Storage arrays 102A-B can provide persistent data storage for computing devices 164A-B. In implementations, storage array 102A can be housed in a chassis (not shown) and storage array 102B can be housed in another chassis (not shown). Storage arrays 102A and 102B can include one or more storage array controllers 110A-D (also referred to herein as "controllers"). Storage array controllers 110A-D can be embodied as modules of an automated computing machine including computer hardware, computer software, or a combination of computer hardware and software. In some implementations, storage array controllers 110A-D can be configured to perform various storage tasks. Storage tasks may include writing data received from computing devices 164A-B to storage arrays 102A-B, erasing data from storage arrays 102A-B, retrieving data from storage arrays 102A-B and providing the data to computing devices 164A-B, monitoring and reporting disk usage and performance, performing redundancy operations such as a Redundant Array of Independent Drives ("RAID") or RAID-like data redundancy operations, compressing data, encrypting data, etc.

[0007] The storage array controllers 110A-D may be implemented in a variety of ways, including as a Field Programmable Gate Array ("FPGA"), a Programmable Logic Chip ("PLC"), an Application Specific Integrated Circuit ("ASIC"), a System-on-Chip ("SOC"), or any computing device that includes discrete components such as a processing device, a central processing unit, computer memory, or various adapters. The storage array controllers 110A-D may include a data communications adapter configured to support communications over, for example, the SAN 158 or the LAN 160. In some implementations, the storage array controllers 110A-D may be independently coupled to the LAN 160. In implementations, the storage array controllers 110A-D may include an I / O controller or the like that couples the storage array controllers 110A-D to persistent storage resources 170A-B (also referred to herein as "storage resources") for data communications over a midplane (not shown). Persistent storage resources 170A-B may include any number of storage drives 171A-F (also referred to herein as "storage devices") and any number of non-volatile random access memory ("NVRAM") devices (not shown).

[0008] In some implementations, the NVRAM devices of persistent storage resources 170A-B may be configured to receive data to be stored on storage drives 171A-F from storage array controllers 110A-D. In some examples, the data may originate from computing devices 164A-B. In some examples, writing data to an NVRAM device may be performed more quickly than writing data directly to storage drives 171A-F. In implementations, storage array controllers 110A-D may be configured to utilize an NVRAM device as a quickly accessible buffer for data to be written to storage drives 171A-F. The latency of write requests using an NVRAM device as a buffer may be improved relative to systems in which storage array controllers 110A-D write data directly to storage drives 171A-F. In some implementations, the NVRAM devices may be implemented using computer memory in the form of high-bandwidth, low-latency RAM. NVRAM devices are referred to as “non-volatile” because they may receive or include their own power source that maintains the RAM state after a main power loss to the NVRAM device. Such a power source may be a battery, one or more capacitors, etc. In response to a power loss, the NVRAM device may be configured to write the contents of the RAM to persistent storage, such as storage drives 171A-F.

[0009] In implementations, storage drives 171A-F may refer to any device configured to persistently record data, where "persistently" or "persistent" refers to the device's ability to maintain recorded data after a loss of power. In some implementations, storage drives 171A-F may correspond to non-disk storage media. For example, storage drives 171A-F may be one or more solid-state drives ("SSDs"), flash memory-based storage, any type of solid-state non-volatile memory, or any other type of non-mechanical storage device. In other implementations, storage drives 171A-F may include mechanical or rotating hard disks, such as hard disk drives ("HDDs").

[0010] In some implementations, the storage array controllers 110A-D may be configured to offload device management responsibilities from the storage drives 171A-F in the storage arrays 102A-B. For example, the storage array controllers 110A-D may manage control information that may describe the state of one or more memory blocks in the storage drives 171A-F. The control information may indicate, for example, that a particular memory block has failed and should no longer be written to, that a particular memory block contains boot code for the storage array controllers 110A-D, the number of program-erase ("P / E") cycles performed on a particular memory block, the age of the data stored in a particular memory block, the type of data stored in a particular memory block, etc. In some implementations, the control information may be stored as metadata with the associated memory block. In other implementations, the control information for the storage drives 171A-F may be stored in one or more specific memory blocks of the storage drives 171A-F selected by the storage array controllers 110A-D. The selected memory block may be tagged with an identifier indicating that the selected memory block contains control information. The identifiers may be utilized by storage array controllers 110A-D in conjunction with storage drives 171A-F to quickly identify memory blocks containing the control information. For example, storage controllers 110A-D may issue commands that specify the locations of memory blocks containing the control information. Note that the control information may be large enough that portions of the control information may be stored in multiple locations, the control information may be stored in multiple locations, for example, for redundancy purposes, or the control information may be otherwise distributed across multiple memory blocks within storage drives 171A-F.

[0011] In an implementation, storage array controllers 110A-D can offload device management responsibilities from storage drives 171A-F of storage arrays 102A-B by retrieving control information from storage drives 171A-F that describes the state of one or more memory blocks within storage drives 171A-F. Retrieving the control information from storage drives 171A-F may be performed, for example, by storage array controllers 110A-D querying storage drives 171A-F for the location of the control information for a particular storage drive 171A-F. The storage drives 171A-F may be configured to execute instructions that enable the storage drives 171A-F to identify the location of the control information. The instructions may be executed by a controller (not shown) associated with or otherwise located on the storage drives 171A-F, and may cause the storage drives 171A-F to scan a portion of each memory block to identify the memory block that stores the control information for the storage drives 171A-F. The storage drives 171A-F may respond by sending a response message to the storage array controllers 110A-D that includes the location of the control information for the storage drives 171A-F. In response to receiving the response message, the storage array controllers 110A-D may issue a request to read the data stored at the address associated with the location of the control information for the storage drives 171A-F.

[0012] In other implementations, storage array controllers 110A-D can further offload device management responsibilities from storage drives 171A-F by performing storage drive management operations in response to receiving the control information. The storage drive management operations may include, for example, operations typically performed by storage drives 171A-F (e.g., a controller (not shown) associated with a particular storage drive 171A-F). The storage drive management operations may include, for example, ensuring that data is not written to failed memory blocks within storage drives 171A-F, ensuring that data is written to memory blocks within storage drives 171A-F such that proper wear leveling is achieved, etc.

[0013] In implementations, storage arrays 102A-B may implement two or more storage array controllers 110A-D. For example, storage array 102A may include storage array controller 110A and storage array controller 110B. At a given instance, a single storage array controller 110A-D (e.g., storage array controller 110A) of storage system 100 may be designated with primary status (also referred to herein as a “primary controller”), and the other storage array controller 110A-D (e.g., storage array controller 110A) may be designated with secondary status (also referred to herein as a “secondary controller”). The primary controller may have certain rights, such as permission to modify data in persistent storage resources 170A-B (e.g., write data to persistent storage resources 170A-B). At least some of the rights of the primary controller may supersede the rights of the secondary controller. For example, if the primary controller has the rights, the secondary controller may not have permission to modify the data in the persistent storage resources 170A-B. The states of the storage array controllers 110A-D may change. For example, the storage array controller 110A may be designated with secondary status, and the storage array controller 110B may be designated with primary status.

[0014] In some implementations, a primary controller, such as storage array controller 110A, may serve as the primary controller for one or more storage arrays 102A-B, and a second controller, such as storage array controller 110B, may serve as a secondary controller for one or more storage arrays 102A-B. For example, storage array controller 110A may be the primary controller for storage array 102A and storage array 102B, and storage array controller 110B may be the secondary controller for storage arrays 102A and 102B. In some implementations, storage array controllers 110C and 110D (also referred to as "storage processing modules") may not have primary or secondary status. Storage array controllers 110C and 110D, implemented as storage processing modules, may serve as a communication interface between the primary and secondary controllers (e.g., storage array controllers 110A and 110B, respectively) and storage array 102B. For example, storage array controller 110A of storage array 102A may send a write request to storage array 102B via SAN 158. The write request may be received by both storage array controllers 110C and 110D of storage array 102B. Storage array controllers 110C and 110D facilitate the communication, for example, sending the write request to the appropriate storage drives 171A-F. Note that in some implementations, storage processing modules can be used to increase the number of storage drives controlled by the primary and secondary controllers.

[0015] In an implementation, storage array controllers 110A-D are communicatively coupled to one or more storage drives 171A-F and one or more NVRAM devices (not shown) included as part of storage arrays 102A-B via a midplane (not shown). Storage array controllers 110A-D may be coupled to the midplane via one or more data communication links, and the midplane may be coupled to storage drives 171A-F and the NVRAM devices via one or more data communication links. The data communication links described herein are collectively illustrated by data communication links 108A-D and may include, for example, a Peripheral Component Interconnect Express ("PCIe") bus.

[0016] FIG. 1B illustrates an exemplary system for data storage, according to some implementations. The storage array controller 101 illustrated in FIG. 1B may be similar to storage array controllers 110A-D described with respect to FIG. 1A. In one example, storage array controller 101 may be similar to storage array controller 110A or storage array controller 110B. Storage array controller 101 includes multiple elements for purposes of illustration and not limitation. Note that in other implementations, storage array controller 101 may include the same, more, or fewer elements, configured in the same or different ways. Note that elements of FIG. 1A may be included below to help illustrate features of storage array controller 101.

[0017] Storage array controller 101 may include one or more processing devices 104 and random access memory ("RAM") 111. Processing device 104 (or controller 101) represents one or more general-purpose processing devices, such as a microprocessor, central processing unit, or the like. More specifically, processing device 104 (or controller 101) may be a complex instruction set computing ("CISC") microprocessor, a reduced instruction set computing ("RISC") microprocessor, a very long instruction word ("VLIW") microprocessor, or a processor implementing other instruction sets or a processor implementing a combination of instruction sets. Processing device 104 (or controller 101) may also be one or more special-purpose processing devices, such as an ASIC, an FPGA, a digital signal processor ("DSP"), a network processor, or the like.

[0018] The processing device 104 may be connected to RAM 111 via a data communication link 106, which may be embodied as a high-speed memory bus such as a Double-Data Rate 4 ("DDR4") bus. An operating system 112 is stored in RAM 111. In some implementations, instructions 113 are stored in RAM 111. The instructions 113 may include computer program instructions for performing operations in a direct-mapped flash storage system. In one embodiment, a direct-mapped flash storage system is a system that directly addresses blocks of data within a flash drive without address translation being performed by the flash drive's storage controller.

[0019] In an implementation, storage array controller 101 includes one or more host bus adapters 103A-C coupled to processing device 104 via data communication links 105A-C. In an implementation, host bus adapters 103A-C may be computer hardware that connects a host system (e.g., a storage array controller) to other networks and storage arrays. In some examples, host bus adapters 103A-C may be Fibre Channel adapters that allow storage array controller 101 to connect to a SAN, Ethernet adapters that allow storage array controller 101 to connect to a LAN, etc. Host bus adapters 103A-C may be coupled to processing device 104 via data communication links 105A-C, such as a PCIe bus.

[0020] In an implementation, storage array controller 101 may include a host bus adapter 114 coupled to an expander 115. Expander 115 may be used to attach a host system to a larger number of storage drives. Expander 115 may be, for example, a SAS expander utilized to allow host bus adapter 114 to attach to storage drives in an implementation in which host bus adapter 114 is embodied as a SAS controller.

[0021] In an implementation, storage array controller 101 may include a switch 116 coupled to processing device 104 via data communication link 109. Switch 116 may be a computer hardware device that can create multiple endpoints from a single endpoint, thereby allowing multiple devices to share a single endpoint. Switch 116 may be, for example, a PCIe switch coupled to a PCIe bus (e.g., data communication link 109) and providing multiple PCIe connection points to a midplane.

[0022] In an implementation, storage array controller 101 includes data communication link 107 for coupling storage array controller 101 to other storage array controllers. In some examples, data communication link 107 may be a QuickPath Interconnect (QPI) interconnect.

[0023] A conventional storage system that uses conventional flash drives may implement processes across flash drives that are part of the conventional storage system. For example, higher-level processes in the storage system may initiate and control processes across the flash drives. However, flash drives in a conventional storage system may include their own storage controller that also implements processes. Thus, a conventional storage system may implement both higher-level processes (e.g., initiated by the storage system) and lower-level processes (e.g., initiated by the storage controller of the storage system).

[0024] To address various shortcomings of conventional storage systems, operations can be performed by higher-level processes rather than by lower-level processes. For example, a flash storage system may include a flash drive that does not include a storage controller to provide the processes. Thus, the operating system of the flash storage system itself can initiate and control the processes. This can be achieved by a direct-mapped flash storage system that directly addresses data blocks within the flash drive without the address translation performed by the flash drive's storage controller.

[0025] In implementations, storage drives 171A-F may be one or more zoned storage devices. In some implementations, one or more zoned storage devices may be a single HDD. In implementations, one or more storage devices may be flash-based SSDs. In a zoned storage device, the zoned namespace on the zoned storage device may be grouped by natural size and addressed by aligned groups of blocks to form multiple addressable zones. In implementations utilizing SSDs, the natural size may be based on the erase block size of the SSD. In some implementations, the zones of a zoned storage device may be defined during initialization of the zoned storage device. In implementations, the zones may be dynamically defined as data is written to the zoned storage device.

[0026] In some implementations, zones may be heterogeneous, with some zones each being a page group and other zones being multiple page groups. In implementations, some zones may correspond to an erase block and other zones may correspond to multiple erase blocks. In one implementation, zones may be any combination of different numbers of pages within page groups and / or erase blocks for a heterogeneous mix of programming modes, manufacturers, product types, and / or product generations of storage devices, as applied to heterogeneous assembly, upgrades, distributed storage, etc. In one implementation, zones may be defined as having usage characteristics, such as characteristics that support data with a particular type of lifespan (e.g., very short-lived or very long-lived). These characteristics may be used by the zoned storage device to determine how the zone is managed over the expected lifespan of the zone.

[0027] It should be understood that zones are virtual constructs. Any particular zone may not have a fixed location on the storage device. Until allocated, a zone may not have any location on the storage device. A zone may, in various implementations, correspond to a number representing a virtually allocatable chunk of space, the size of an erase block or other block size. When the system allocates or opens a zone, the zone is allocated in flash or other solid-state storage memory, and when the system writes to the zone, the pages are written to the mapped flash or other solid-state storage memory of the zoned storage device. When the system closes a zone, the associated erase block or block of other size is completed. At some point in the future, the system can delete the zone, which frees up the zone's allocated space. During its lifetime, a zone may be moved to a different location on the zoned storage device, for example, when the zoned storage device undergoes internal maintenance.

[0028] In implementations, zones in a zoned storage device can be in different states. A zone can be empty, with no data stored in it. An empty zone can be opened explicitly or implicitly by writing data to the zone. This is the initial state of a zone on a new zoned storage device, but can also be the result of a zone reset. In some implementations, an empty zone can have a specified location within the flash memory of the zoned storage device. In one implementation, the location of an empty zone can be selected when the zone is first opened or first written to (or later, if the write is buffered in memory). Zones can be in the open state either implicitly or explicitly, and a zone in the open state can be written to store data using a write command or an append command. In one implementation, a zone in the open state can be written using a copy command, which copies data from a different zone. In some implementations, a zoned storage device can have a limit on the number of open zones at a particular time.

[0029] A closed zone is a zone that has been partially written to but entered the closed state after issuing an explicit close operation. A closed zone may remain available for future writes, but may reduce some of the runtime overhead consumed by keeping the zone open. In implementations, a zoned storage device may have a limit on the number of closed zones at a particular time. A full zone is a zone that stores data and can no longer be written to. A zone may be in the full state either after a write has written data to the entire zone or as a result of a zone close operation. Before the close operation, the zone may or may not be completely written to. However, after the close operation, the zone may not be open for further writes without first performing a zone reset operation.

[0030] The mapping from zones to erase blocks (or single tracks in an HDD) may be arbitrary, dynamic, or hidden from view. The process of opening a zone may be an operation that allows a new zone to be dynamically mapped to the underlying storage of a zoned storage device, and then allows data to be written to the zone by appending writes until the zone reaches capacity. A zone may be terminated at any point, after which no further data may be written to the zone. When the data stored in a zone is no longer needed, the zone may be reset, thereby effectively deleting the zone's contents from the zoned storage device and making the physical storage held by that zone available for subsequent data storage. Once a zone is written and terminated, the zoned storage device ensures that the data stored in the zone is not lost until the zone is reset. In the time between writing data to a zone and resetting the zone, the zone may be moved between single tracks or erase blocks, such as by copying data, to keep data refreshed as part of a maintenance operation in the zoned storage device, or to handle the aging of memory cells in an SSD.

[0031] In implementations utilizing HDDs, resetting a zone may allow a single track to be allocated to a new opened zone that may be opened at some point in the future. In implementations utilizing SSDs, resetting a zone causes the zone's associated physical erase blocks to be erased and then reused for storage of data. In some implementations, a zoned storage device may have a limit on the number of zones that are open at a given time to reduce the amount of overhead dedicated to keeping zones open.

[0032] The operating system of the flash storage system can identify and maintain a list of allocation units across multiple flash drives of the flash storage system. An allocation unit can be an entire erase block or multiple erase blocks. The operating system can maintain a map or address ranges that directly map addresses to erase blocks on the flash drives of the flash storage system.

[0033] Direct mapping to erase blocks on a flash drive can be used to rewrite data and erase data. For example, operations can be performed on one or more allocation units that include first and second data, where the first data is retained and the second data is no longer in use by the flash storage system. The operating system can initiate a process to write the first data to a new location in another allocation unit, erase the second data, and mark the allocation unit as available for subsequent data. Thus, the process can be performed solely by the higher-level operating system of the flash storage system, without additional lower-level processes performed by the flash drive's controller.

[0034] Advantages of processes performed solely by the flash storage system's operating system include improved reliability of the flash drives in the flash storage system, since unnecessary or redundant write operations are not performed during the process. One potential novelty is the concept of initiating and controlling the process in the flash storage system's operating system. Additionally, the process may be controlled by the operating system across multiple flash drives. This is in contrast to processing performed by the flash drive's storage controller.

[0035] A storage system may consist of two storage array controllers that share a set of drives for failover purposes, or may consist of a single storage array controller that provides storage services utilizing multiple drives, or may consist of a distributed network of storage array controllers, each having some number of drives or some amount of flash storage, where the storage array controllers in the network cooperate to provide a complete storage service and cooperate with respect to various aspects of the storage service, including storage allocation and garbage collection.

[0036] 1C illustrates a third exemplary system 117 for data storage, according to some implementations. System 117 (also referred to herein as a "storage system") includes a number of elements for purposes of illustration and not limitation. Note that system 117 may include the same, more, or fewer elements, configured in the same or different ways, in other implementations.

[0037] In one embodiment, system 117 includes dual Peripheral Component Interconnect ("PCI") flash storage device 118 with separately addressable high-speed write storage. System 117 may include storage device controller 119. In one embodiment, storage device controllers 119A-D may be CPUs, ASICs, FPGAs, or any other circuitry capable of implementing the necessary control structures in accordance with this disclosure. In one embodiment, system 117 includes flash memory devices (e.g., including flash memory devices 120a-n) operably coupled to various channels of storage device controller 119. Flash memory devices 120a-n may be presented to controller 119A-D as addressable collections of flash pages, erase blocks, and / or control elements sufficient to enable storage device controller 119A-D to program and retrieve various aspects of the flash. In one embodiment, storage device controllers 119A-D can perform operations on flash memory devices 120a-n, including storing and retrieving data contents of pages, allocating and erasing any blocks, tracking statistics regarding the use and reuse of flash memory pages, erase blocks, and cells, tracking and predicting error codes and failures within flash memory, controlling voltage levels associated with programming flash cells and retrieving contents, etc.

[0038] In one embodiment, system 117 may include RAM 121 for storing separately addressable high-speed write data. In one embodiment, RAM 121 may be one or more separate discrete devices. In another embodiment, RAM 121 may be integrated into storage device controllers 119A-D or multiple storage device controllers. RAM 121 may also be utilized for other purposes, such as temporary program memory for a processing device (e.g., a CPU) within storage device controller 119.

[0039] In one embodiment, system 117 may include a stored energy device 122, such as a rechargeable battery or capacitor. Stored energy device 122 may store enough energy to power storage device controller 119, a certain amount of RAM (e.g., RAM 121), and a certain amount of flash memory (e.g., flash memories 120a-120n) for a sufficient time to write the contents of the RAM to flash memory. In one embodiment, storage device controllers 119A-D may write the contents of RAM to flash memory if the storage device controller detects a loss of external power.

[0040] In one embodiment, system 117 includes two data communication links 123a, 123b. In one embodiment, data communication links 123a, 123b may be PCI interfaces. In other embodiments, data communication links 123a, 123b may be based on other communication standards (e.g., HyperTransport, InfiniBand, etc.). Data communication links 123a, 123b may be based on the Non-Volatile Memory Express (“NVMe”) or NVMe over fabric (“NVMf”) specifications, which allow external connections from other components within storage system 117 to storage device controllers 119A-D. Note that the data communication links may be referred to interchangeably as PCI buses for convenience herein.

[0041] System 117 may also include an external power source (not shown), which may be provided via one or both of data communication links 123a, 123b, or may be provided separately. An alternative embodiment includes a separate flash memory (not shown) dedicated for use in storing the contents of RAM 121. Storage device controllers 119A-D may present a logical device on the PCI bus, which may include an addressable fast-write logical device, or a separate portion of the logical address space of storage device 118, which may be presented as PCI memory or persistent storage. In one embodiment, operations to store to the device are directed to RAM 121. During a power outage, storage device controllers 119A-D may write stored content associated with the addressable fast-write logical storage to flash memory (e.g., flash memories 120a-n) for long-term persistent storage.

[0042] In one embodiment, a logical device may include some representation of some or all of the contents of flash memory devices 120a-n that allows a storage system, including storage device 118 (e.g., storage system 117), to directly address flash memory pages and reprogram erase blocks directly from storage system components external to the storage device via a PCI bus. This representation may also allow one or more of the external components to control and retrieve other aspects of the flash memory, including some or all of tracking statistics related to the use and reuse of flash memory pages, erase blocks, and cells across all flash memory devices, tracking and predicting error codes and failures within and across flash memory devices, controlling voltage levels associated with programming and retrieving the contents of flash cells, etc.

[0043] In one embodiment, stored energy device 122 may be sufficient to ensure completion of ongoing operations on flash memory devices 120a-120n, and stored energy device 122 may power storage device controllers 119A-D and associated flash memory devices (e.g., 120a-n) for those operations as well as for fast write RAM storage to flash memory. Stored energy device 122 may be used to store cumulative statistics and other parameters kept and tracked by flash memory devices 120a-n and / or storage device controller 119. A separate capacitor or stored energy device (such as a smaller capacitor near or embedded within the flash memory device itself) may be used for some or all of the operations described herein.

[0044] Various schemes may be used to track and optimize the life of stored energy components, such as adjusting voltage levels over time, partially discharging stored energy device 122 to measure corresponding discharge characteristics, etc. If available energy decreases over time, the effective available capacity of the addressable fast-write storage may be reduced to ensure that it can be safely written based on the currently available stored energy.

[0045] 1D illustrates a third exemplary storage system 124 for data storage, according to some implementations. In one embodiment, storage system 124 includes storage controllers 125a, 125b. In one embodiment, storage controllers 125a, 125b are operably coupled to a dual PCI storage device. Storage controllers 125a, 125b can be operably coupled to a number of host computers 127a-n (e.g., via storage network 130).

[0046] In one embodiment, two storage controllers (e.g., 125a and 125b) provide storage services such as an SCS block storage array, a file server, an object server, a database, or a data analysis service. The storage controllers 125a, 125b can provide services to host computers 127a-n external to the storage system 124 through a number of network interfaces (e.g., 126a-d). The storage controllers 125a, 125b can provide services or applications integrated entirely within the storage system 124, forming an integrated storage and computing system. The storage controllers 125a, 125b can utilize high-speed write memory in or across the storage devices 119a-d to journal ongoing operations, ensuring that operations are not lost due to power outage, removal of a storage controller, storage controller or storage system shutdown, or any failure of one or more software or hardware components within the storage system 124.

[0047] In one embodiment, storage controllers 125a, 125b act as PCI masters for one or the other PCI bus 128a, 128b. In another embodiment, 128a and 128b may be based on other communication standards (e.g., HyperTransport, InfiniBand, etc.). Other storage system embodiments may allow storage controllers 125a, 125b to act as multi-masters for both PCI buses 128a, 128b. Alternatively, a PCI / NVMe / NVMf switching infrastructure or fabric may connect multiple storage controllers. Some storage system embodiments may allow storage devices to communicate directly with each other rather than communicating only with the storage controllers. In one embodiment, storage device controller 119a may be operable under direction from storage controller 125a to combine and transfer data to be stored in a flash memory device from data stored in RAM (e.g., RAM 121 of FIG. 1C ). For example, a recomputed version of the RAM contents can be transferred after the storage controller determines that the operation has been fully committed across the storage system, or when the fast write memory on the device reaches a certain used capacity, or after a certain amount of time to ensure improved data security or to free up addressable fast write capacity for reuse. This mechanism can be used, for example, to avoid a second transfer over a bus (e.g., 128a, 128b) from the storage controller 125a, 125b. In one embodiment, the recomputation can include compressing the data, attaching indexing or other metadata, combining multiple data segments together, performing erasure coding calculations, etc.

[0048] In one embodiment, under direction from storage controller 125a, 125b, storage device controller 119a, 119b may be operable to calculate and transfer data from data stored in RAM (e.g., RAM 121 of FIG. 1C) to other storage devices without the involvement of storage controller 125a, 125b. This operation may be used to mirror data stored in one storage controller 125a to another storage controller 125b, or to offload compression, data aggregation, and / or erasure coding calculations and transfer them to storage devices to reduce the load on storage controller or storage controller interface 129a, 129b to PCI bus 128a, 128b.

[0049] The storage device controllers 119A-D may include mechanisms for implementing high availability primitives used by other parts of the storage system external to the dual PCI storage device 118. For example, in a storage system having two storage controllers providing highly available storage services, reservation or exclusion primitives may be provided so that one storage controller can prevent the other storage controller from accessing or continuing to access the storage device. This may be used, for example, when one controller detects that the other controller is not functioning properly, or when the interconnect between the two storage controllers itself may not be functioning properly.

[0050] In one embodiment, a storage system for use with dual PCI direct-mapped storage devices having separately addressable fast-write storage includes a system for managing erase blocks or groups of erase blocks as allocation units for storing data on behalf of a storage service, for storing metadata associated with the storage service (e.g., indexes, logs, etc.), or for proper management of the storage system itself. Flash pages, which may be several kilobytes in size, may be written as data arrives or as the storage system persists the data for a long time interval (e.g., above a defined time threshold). To commit data more quickly or reduce the number of writes to the flash memory device, the storage controller may first write the data to separately addressable fast-write storage on the other storage device.

[0051] In one embodiment, the storage controllers 125a, 125b can initiate the use of erase blocks within and across storage devices (e.g., 118) according to the age and expected remaining life of the storage device, or based on other statistics. The storage controllers 125a, 125b can initiate garbage collection and data migration of data between storage devices according to pages that are no longer needed, manage the lifespan of flash pages and erase blocks, and manage overall system performance.

[0052] In one embodiment, storage system 124 may utilize mirroring and / or erasure coding schemes as part of storing data in addressable, fast-write storage and / or as part of writing data to allocation units associated with erase blocks. Erasure codes may be used across storage devices, as well as within erase blocks or allocation units, or within and across flash memory devices on a single storage device, to provide redundancy against single or multiple storage device failures or to protect against internal corruption of flash memory pages resulting from flash memory operation or degradation of flash memory cells. Mirroring and erasure coding at various levels may be used to recover from multiple types of failures, occurring separately or in combination.

[0053] The embodiments depicted with reference to Figures 2A-2G illustrate a storage cluster that stores user data, such as user data originating from one or more user or client systems or other sources external to the storage cluster. The storage cluster distributes user data across storage nodes housed within a chassis or across multiple chassis using erasure coding and redundant copies of metadata. Erasure coding refers to a method of data protection or reconstruction in which data is stored across a set of distinct locations, such as disks, storage nodes, or geographic locations. Flash memory is one type of solid-state memory that may be integrated with embodiments, but embodiments can be extended to other types of solid-state memory or other storage media, including non-solid-state memory. Control of storage locations and workloads is distributed across storage locations in a clustered peer-to-peer system. Tasks such as mediating communications between various storage nodes, detecting when a storage node becomes unavailable, and balancing I / O (input and output) across various storage nodes are all handled on a distributed basis. Data, in some embodiments, is placed or distributed across multiple storage nodes in data fragments or stripes that support data recovery. Ownership of data can be reassigned within the cluster regardless of input and output patterns. This architecture, described in more detail below, allows storage nodes within a cluster to fail while the system remains operational, as data can be reconstructed from other storage nodes and therefore remain available for input and output operations. In various embodiments, the storage nodes may be referred to as cluster nodes, blades, or servers.

[0054] A storage cluster may be contained within a chassis, i.e., a housing that houses one or more storage nodes. Included within the chassis are mechanisms for providing power to each storage node, such as a power distribution bus, and communication mechanisms, such as a communication bus, that enable communication between the storage nodes. According to some embodiments, the storage cluster may operate as an independent system in one location. In one embodiment, the chassis includes at least two instances of both power distribution and communication buses that can be independently enabled or disabled. The internal communication bus may be an Ethernet bus, although other technologies, such as PCIe, InfiniBand, and others, are equally suitable. The chassis provides ports for an external communication bus to enable communication between multiple chassis and client systems, either directly or via a switch. External communication can use technologies such as Ethernet, InfiniBand, or Fibre Channel. In some embodiments, the external communication bus uses different communication bus technologies for inter-chassis communication and client communication. When a switch is deployed within or between chassis, the switch can act as a translator between multiple protocols or technologies. When multiple chassis are connected to define a storage cluster, the storage cluster may be accessed by clients using either proprietary or standard interfaces, such as network file system ("NFS"), common internet file system ("CIFS"), small computer system interface ("SCSI"), or hypertext transfer protocol ("HTTP"). Translation from the client protocol may occur within a switch, a chassis external communication bus, or each storage node. In some embodiments, multiple chassis may be coupled or connected to each other through an aggregator switch. Some and / or all of the coupled or connected chassis may be designated as a storage cluster.As mentioned above, each chassis may have multiple blades, and each blade has a media access control ("MAC") address, but the storage cluster, in some embodiments, is presented to the external network as having a single cluster MAC address and a single IP.

[0055] Each storage node may be one or more storage servers, each connected to one or more non-volatile solid-state memory units, which may be referred to as storage units or storage devices. One embodiment includes a single storage server and 1 to 8 non-volatile solid-state memory units within each storage node, but this example is not meant to be limiting. The storage server may include a processor, DRAM, an interface for an internal communication bus, and power distribution for each of the power buses. In some embodiments, within the storage node, the interface and storage units share a communication bus, e.g., PCI Express. The non-volatile solid-state memory units may directly access the internal communication bus interface via the storage node communication bus or may require the storage node to access the bus interface. In some embodiments, the non-volatile solid-state memory units include an embedded CPU, a solid-state storage controller, and a solid-state mass storage device, e.g., in amounts of 2 to 32 terabytes ("TB"). The non-volatile solid-state memory units include an internal volatile storage medium, such as DRAM, and an energy storage device. In some embodiments, the energy storage device is a capacitor, supercapacitor, or battery that allows for the transfer of a subset of the DRAM contents to a stable storage medium in the event of power loss. In some embodiments, the non-volatile solid-state memory unit is constructed with storage class memory such as phase change or magnetoresistive random access memory ("MRAM"), which replaces DRAM and allows for reduced power holdup devices.

[0056] One of the many features of the storage nodes and non-volatile solid-state storage is the ability to proactively rebuild data in a storage cluster. The storage nodes and non-volatile solid-state storage can determine when a storage node or non-volatile solid-state storage in a storage cluster becomes unreachable, regardless of whether there is an attempt to read the data associated with that storage node or non-volatile solid-state storage. The storage nodes and non-volatile solid-state storage then cooperate to recover and rebuild the data, at least in part, in a new location. This constitutes proactive rebuilding in that the system rebuilds data without waiting until the data is needed for a read access initiated by a client system using the storage cluster. These and further details of the storage nodes and their operation are discussed below.

[0057] FIG. 2A is a perspective view of a storage cluster 161 having multiple storage nodes 150 and internal solid-state memory coupled to each storage node to provide network-attached storage or a storage area network, according to some embodiments. A network-attached storage, storage area network, or storage cluster or other storage memory may include one or more storage clusters 161, each having one or more storage nodes 150, with a flexible and reconfigurable arrangement of both physical components and the amount of storage memory provided thereby. The storage cluster 161 is designed to fit into a rack, and one or more racks can be set up and populated as desired for storage memory. The storage cluster 161 includes a chassis 138 having multiple slots 142. It should be understood that the chassis 138 may also be referred to as a housing, enclosure, or rack unit. In one embodiment, the chassis 138 has 14 slots 142, although other numbers of slots are readily contemplated. For example, some embodiments have 4 slots, 8 slots, 16 slots, 32 slots, or other suitable numbers of slots. Each slot 142 can accommodate one storage node 150 in some embodiments. The chassis 138 includes a flap 148 that can be utilized to mount the chassis 138 in a rack. The fans 144 provide air circulation to cool the storage nodes 150 and their components, although other cooling components may be used, or an embodiment without cooling components may be devised. The switch fabric 146 couples the storage nodes 150 within the chassis 138 to each other and to a network for communication to memory. In one embodiment depicted herein, for illustrative purposes, the slots 142 to the left of the switch fabric 146 and fans 144 are shown as occupied by a storage node 150, while the slots 142 to the right of the switch fabric 146 and fans 144 are empty and available for inserting a storage node 150.This configuration is an example, and one or more storage nodes 150 can occupy slots 142 in a variety of additional arrangements. The arrangement of storage nodes need not be contiguous or adjacent in some embodiments. Storage nodes 150 are hot-pluggable, meaning that storage nodes 150 can be inserted into or removed from slots 142 in chassis 138 without shutting down or powering down the system. Upon insertion or removal of storage node 150 into or from slot 142, the system recognizes the change and automatically reconfigures to adapt. Reconfiguration, in some embodiments, includes restoring redundancy and / or rebalancing data or load.

[0058] Each storage node 150 may have multiple components. In the embodiment shown herein, storage node 150 includes a CPU 156, i.e., a printed circuit board 159 on which the processor is implemented, memory 154 coupled to CPU 156, and non-volatile solid-state storage 152 coupled to CPU 156, although other implementations and / or components may be used in further embodiments. Memory 154 contains instructions to be executed by CPU 156 and / or data to be operated on by CPU 156. As described further below, non-volatile solid-state storage 152 may include flash, or in further embodiments, other types of solid-state memory.

[0059] Referring to FIG. 2A , the storage cluster 161 is scalable, meaning that storage capacity having non-uniform storage sizes is easily added, as described above. In some embodiments, one or more storage nodes 150 can be plugged in or removed from each chassis, and the storage cluster self-configures. Plug-in storage nodes 150, whether installed in the chassis at delivery or added later, can have different sizes. For example, in one embodiment, the storage nodes 150 can have any multiple of 4 TB, e.g., 8 TB, 12 TB, 16 TB, 32 TB, etc. In further embodiments, the storage nodes 150 can have any multiple of other storage amounts or capacities. The storage capacity of each storage node 150 is broadcast and influences the determination of how data is striped. For maximum storage efficiency, one embodiment can self-configure as widely as possible within a stripe, subject to a given requirement for continuous operation with the loss of up to one or up to two non-volatile solid-state storage 152 units or storage nodes 150 within a chassis.

[0060] FIG. 2B is a block diagram illustrating a communication interconnect 173 and a power distribution bus 172 coupling multiple storage nodes 150. Referring back to FIG. 2A, the communication interconnect 173 may, in some embodiments, be included in or implemented with the switch fabric 146. When multiple storage clusters 161 occupy a rack, in some embodiments, the communication interconnect 173 may be included in or implemented with a top-of-rack switch. As illustrated in FIG. 2B, the storage cluster 161 is enclosed within a single chassis 138. The external port 176 is coupled to the storage node 150 via the communication interconnect 173, and the external port 174 is coupled directly to the storage node. The external power port 178 is coupled to the power distribution bus 172. The storage node 150 may include various amounts and capacities of non-volatile solid-state storage 152, as described with reference to FIG. 2A. Additionally, one or more of the storage nodes 150 may be compute-only storage nodes, as illustrated in FIG. 2B. Authorities 168 are implemented on non-volatile solid-state storage 152, for example, as a list or other data structure stored in-memory. In some embodiments, authorities are stored within non-volatile solid-state storage 152 and supported by software executing on a controller or other processor of non-volatile solid-state storage 152. In further embodiments, authorities 168 are implemented on storage node 150, for example, as a list or other data structure stored in memory 154 and supported by software executing on CPU 156 of storage node 150. In some embodiments, authorities 168 control how and where data is stored in non-volatile solid-state storage 152. This control helps determine what type of erasure coding scheme is applied to the data and which storage node 150 has which portion of the data. Each authority 168 can be assigned to non-volatile solid-state storage 152.Each authority, in various embodiments, can control a range of inode numbers, segment numbers, or other data identifiers assigned to data by the file system, by storage node 150, or by non-volatile solid-state storage 152.

[0061] In some embodiments, all data and all metadata are redundant within the system. Furthermore, all data and all metadata have an owner, which may be referred to as an authority. If the authority is unreachable, for example due to a storage node failure, there is a succession plan for how to find the data or its metadata. In various embodiments, there are redundant copies of authority 168. In some embodiments, authority 168 has a relationship to storage nodes 150 and non-volatile solid-state storage 152. Each authority 168 covering a range of data segment numbers or other identifiers of data may be assigned to a particular non-volatile solid-state storage 152. In some embodiments, authorities 168 for all such ranges are distributed across the non-volatile solid-state storage 152 of the storage cluster. Each storage node 150 has a network port that provides access to the non-volatile solid-state storage 152 of that storage node 150. Data may be stored in segments associated with a segment number, which in some embodiments is an indirect reference to the configuration of a RAID (Redundant Array of Independent Disks) stripe. Thus, the assignment and use of authority 168 establishes an indirect reference to the data. Indirect referencing may, according to some embodiments, be referred to as the ability to indirectly reference data, in this case via authority 168. A segment identifies a set of non-volatile solid-state storage 152 and a local identifier to the set of non-volatile solid-state storage 152 that may contain the data. In some embodiments, the local identifier is an offset into the device and may be reused sequentially by multiple segments. In other embodiments, the local identifier is unique to a particular segment and is never reused. The offset within non-volatile solid-state storage 152 is applied to locating data for writing to or reading from non-volatile solid-state storage 152 (in the form of a RAID stripe).Data is striped across multiple units of non-volatile solid-state storage 152, which may or may not include non-volatile solid-state storage 152 with authority 168 for particular data segments.

[0062] For example, during data migration or data reconstruction, if there is a change in where a particular segment of data is located, the authority 168 for that data segment should be referenced in the non-volatile solid-state storage 152 or storage node 150 that has that authority 168. To locate particular data, embodiments calculate a hash value for the data segment or apply an inode number or data segment number. The output of this operation points to the non-volatile solid-state storage 152 that has the authority 168 for that particular data. In some embodiments, this operation has two stages. The first stage is mapping an entity identifier (ID), such as a segment number, inode number, or directory number, to an authority identifier. This mapping may include a calculation such as a hash or bit mask. The second stage is mapping the authority identifier to a particular non-volatile solid-state storage 152, which can be done through explicit mapping. This operation is repeatable, so once the calculation is performed, the result of the calculation repeatably and reliably points to the particular non-volatile solid-state storage 152 that has that authority 168. The operation may include as input a set of reachable storage nodes. If the set of reachable non-volatile solid-state storage units changes, the optimal set changes. In some embodiments, the persisted value is the current allocation (which is always true) and the calculated value is the target allocation to which the cluster attempts to reconfigure. This calculation may be used to determine the optimal non-volatile solid-state storage 152 for authority, given a set of non-volatile solid-state storages 152 that are reachable and that constitute the same cluster. The calculation also determines an ordered set of peer non-volatile solid-state storages 152 that also record a mapping of authority to non-volatile solid-state storage so that authority can be determined even if the assigned non-volatile solid-state storage is unreachable. In some embodiments, if a particular authority 168 is unavailable, a replica or substitute authority 168 may be referenced.

[0063] 2A and 2B, two of the many tasks of the CPU 156 on a storage node 150 are to split write data and reassemble read data. When the system determines that data is to be written, the authority 168 for that data is located as described above. If the segment ID of the data has already been determined, the write request is forwarded from the segment to the non-volatile solid-state storage 152 currently determined to be the host of the determined authority 168. The host CPU 156 of the storage node 150 where the non-volatile solid-state storage 152 and corresponding authority 168 reside then splits or shards the data and transmits the data to the various non-volatile solid-state storages 152. The transmitted data is written as data stripes according to an erasure coding scheme. In some embodiments, data is requested to be pulled, while in other embodiments, data is pushed. Conversely, when data is read, the authority 168 for the segment ID containing the data is located as described above. The host CPU 156 of the storage node 150 where the non-volatile solid-state storage 152 and corresponding authority 168 reside requests data from the non-volatile solid-state storage and corresponding storage node pointed to by the authority. In some embodiments, the data is read from flash storage as a data stripe. The host CPU 156 of the storage node 150 then reassembles the read data, corrects any errors (if any) according to an appropriate erasure coding scheme, and transfers the reassembled data to the network. In further embodiments, some or all of these tasks may be processed in the non-volatile solid-state storage 152. In some embodiments, a segment host requests data to be sent to the storage node 150 by requesting a page from storage and then sending the data to the storage node that made the original request.

[0064] In an embodiment, authority 168 operates to determine how an operation proceeds for a particular logical element. Each logical element may be operated through a particular authority across multiple storage controllers of a storage system. Authority 168 may communicate with multiple storage controllers to cause the multiple storage controllers to collectively perform the operation for those particular logical elements.

[0065] In embodiments, a logical element may be, for example, a file, a directory, an object bucket, an individual object, a delimited portion of a file or object, some other form of key-value pair database, or a table. In embodiments, performing an operation may involve, for example, ensuring consistency, structural integrity, and / or recoverability with other operations on the same logical element, reading metadata and data associated with the logical element, determining what data should be durably written to the storage system to persist any changes due to the operation, or where it may be determined that metadata and data are stored across modular storage devices attached to multiple storage controllers in the storage system.

[0066] In some embodiments, operations are token-based transactions for efficient communication within a distributed system. Each transaction may be accompanied by or associated with a token that grants permission to perform the transaction. Authority 168, in some embodiments, may maintain the pre-transaction state of the system until the completion of the operation. Token-based communication can be achieved without global locks across the system and also allows for the resumption of operations in the event of an interruption or other failure.

[0067] In some systems, e.g., UNIX-style file systems, data is handled in index nodes or inodes, which specify data structures that represent objects within the file system. An object may be, for example, a file or a directory. Metadata may be associated with an object as attributes such as permission data and creation timestamps, among other attributes. A segment number may be assigned to all or part of such an object within the file system. In other systems, data segments are handled with segment numbers assigned elsewhere. For purposes of explanation, the unit of distribution is an entity, which may be a file, directory, or segment. That is, an entity is a unit of data or metadata stored by the storage system. Entities are grouped into sets called authorities. Each authority has an authority owner, which is a storage node that has exclusive rights to update the entities within the authority. In other words, storage nodes contain authorities, which in turn contain entities.

[0068] A segment is a logical container of data according to some embodiments. A segment is an address space between a media address space and a physical flash location; i.e., data segment numbers reside in this address space. A segment may also contain metadata that allows data redundancy to be restored (rewritten to a different flash location or device) without the involvement of higher-level software. In one embodiment, the internal format of a segment includes client data and a media mapping for determining the location of that data. Each data segment is protected against, for example, memory and other failures, by dividing the segment into multiple data and parity shards, if applicable. The data and parity shards are distributed, or striped, across the non-volatile solid-state storage 152 coupled to the host CPU 156 (see FIGS. 2E and 2G) according to an erasure coding scheme. The use of the term segment, in some embodiments, refers to the container and its location in the segment's address space. The use of the term stripe refers to the same set of shards as a segment, and, according to some embodiments, includes how the shards are distributed along with the redundancy or parity information.

[0069] A series of address space translations occurs throughout the storage system. At the top are directory entries (file names) that link to inodes. The inodes point to the media address space where data is logically stored. Media addresses can be mapped through a series of indirection media to distribute large file loads or implement data services such as deduplication or snapshots. Media addresses can be mapped through a series of indirection media to distribute large file loads or implement data services such as deduplication or snapshots. Next, segment addresses are translated to physical flash locations. According to some embodiments, physical flash locations have an address range that is limited by the amount of flash in the system. Media addresses and segment addresses are logical containers, and in some embodiments, use 128-bit or larger identifiers to be effectively infinite, with the potential for reuse calculated to be longer than the expected life of the system. In some embodiments, addresses from the logical containers are allocated hierarchically. Initially, each non-volatile solid-state storage 152 unit can be assigned a range of address space. Within this allocated range, non-volatile solid-state storage 152 can allocate addresses without synchronization with other non-volatile solid-state storage 152 .

[0070] Data and metadata are stored via a set of underlying storage layouts optimized for different workload patterns and storage devices. These layouts incorporate multiple redundancy schemes, compression formats, and indexing algorithms. Some of these layouts store information about authority and authority masters, while others store file metadata and file data. Redundancy schemes include error correction codes that tolerate corrupted bits within a single storage device (e.g., a NAND flash chip), erasure codes that tolerate failures of multiple storage nodes, and replication schemes that tolerate data center or regional failures. In some embodiments, low-density parity check ("LDPC") codes are used within a single storage unit. In some embodiments, Reed-Solomon coding is used within a storage cluster, and mirroring is used within a storage grid. Metadata may be stored using an ordered log-structured index (e.g., a log-structured merge tree), and large data may not be stored in a log-structured layout.

[0071] To maintain consistency across multiple copies of an entity, storage nodes implicitly agree on two things through computation: (1) the authorities that contain the entity, and (2) the storage nodes that contain the authorities. The assignment of entities to authorities can be done by pseudo-randomly assigning entities to authorities, by dividing entities into ranges based on an externally generated key, or by placing a single entity in each authority. Examples of pseudo-random schemes are the hash family of linear hashing and Replication Under Scalable Hashing ("RUSH"), including Controlled Replication Under Scalable Hashing ("CRUSH"). In some embodiments, pseudo-random assignment is utilized solely to assign authorities to nodes, since the set of nodes may change. Because the set of authorities cannot change, any subjective function may be applied in these embodiments. Some placement schemes automatically place authorities on storage nodes, while others rely on explicit mapping of authorities to storage nodes. In some embodiments, a pseudo-random scheme is utilized to map from each authority to a set of candidate authority holders. A pseudo-random data distribution function associated with CRUSH can assign authorities to storage nodes and create a list of where authorities are assigned. Each storage node has a copy of the pseudo-random data distribution function and can arrive at the same calculation for distribution and later discover or locate authorities. Each pseudo-random scheme, in some embodiments, requires a reachable set of storage nodes as input to conclude the same target node. Once entities are placed within authorities, they can be stored on physical devices such that expected failures do not lead to unexpected data loss. In some embodiments, the rebalancing algorithm attempts to store copies of all entities within an authority in the same layout on the same set of machines.

[0072] Examples of expected failures include device failure, stolen machinery, data center fire, and regional disasters such as nuclear or geological events. Different failures result in different levels of tolerable data loss. In some embodiments, a stolen storage node does not affect the security or reliability of the system, but depending on the system configuration, a regional event may result in no data loss, loss of a few seconds or minutes of updates, or even complete data loss.

[0073] In embodiments, the placement of data for storage redundancy is independent of the placement of authority for data consistency. In some embodiments, the storage nodes containing the authority do not contain any persistent storage. Instead, the storage nodes are connected to non-volatile solid-state storage units that do not contain authority. The communication interconnect between the storage nodes and the non-volatile solid-state storage units is comprised of multiple communication technologies and has non-uniform performance and fault-tolerance characteristics. In some embodiments, as described above, the non-volatile solid-state storage units are connected to the storage nodes via PCI Express, and the storage nodes are connected together within a single chassis using an Ethernet backplane, and the chassis are connected together to form a storage cluster. The storage cluster, in some embodiments, is connected to clients using Ethernet or Fibre Channel. When multiple storage clusters are configured into a storage grid, the multiple storage clusters are connected using the Internet or other long-distance networking links, such as "metro-scale" links or private links that do not traverse the Internet.

[0074] Authorities have exclusive rights to modify entities, migrate entities from one non-volatile solid-state storage unit to another, and add and remove copies of entities. This allows redundancy of the underlying data to be maintained. If an authority fails, is scheduled for decommissioning, or is overloaded, authority is transferred to a new storage node. In transient failures, it is important to ensure that all surviving machines agree on the new authority location. Ambiguity caused by transient failures can be resolved automatically through consensus protocols such as Paxos, hot-warm failover methods, manual intervention by a remote system administrator, or by a local hardware administrator (such as by physically removing the failed machine from the cluster or pressing a button on the failed machine). In some embodiments, a consensus protocol is used and failover is automatic. According to some embodiments, if too many failures or replication events occur within too short a period of time, the system enters a self-preservation mode, halting replication and data movement activities until an administrator intervenes.

[0075] As authorities are transferred between storage nodes and authority owners update entities within those authorities, the system transfers messages between the storage nodes and non-volatile solid-state storage units. Regarding persistent messages, messages with different purposes are of different types. Depending on the message type, the system maintains different ordering and durability guarantees. When persistent messages are being processed, the messages are temporarily stored in multiple durable and non-durable storage hardware technologies. In some embodiments, messages are stored in RAM, NVRAM, and NAND flash devices, and different protocols are used to efficiently use each storage medium. Latency-sensitive client requests may be persisted to replicated NVRAM and then later to NAND, while background rebalancing operations are persisted directly to NAND.

[0076] Persistent messages are stored persistently before being transmitted. This allows the system to continue servicing client requests despite failures and component replacement. Many hardware components contain unique identifiers that are visible to system administrators, manufacturers, the hardware supply chain, and an ongoing monitoring and quality control infrastructure; however, applications running on top of the infrastructure addresses virtualize the addresses. These virtualized addresses do not change over the life of the storage system, despite component failures and replacements. This allows components of the storage system to be replaced over time without reconfiguration or interruption of client request processing; i.e., the system supports non-disruptive upgrades.

[0077] In some embodiments, the virtualized addresses are stored with full redundancy. A continuous monitoring system correlates hardware and software status with hardware identifiers, allowing for the detection and prediction of failures due to defective components and manufacturing details. The monitoring system also, in some embodiments, allows for the proactive transfer of privileges and entities from affected devices before failures occur by removing components from the critical path.

[0078] FIG. 2C is a multilevel block diagram illustrating the contents of a storage node 150 and the contents of its non-volatile solid-state storage 152. Data is communicated to and from the storage node 150 by a network interface controller ("NIC") 202, in some embodiments. Each storage node 150 includes a CPU 156 and one or more non-volatile solid-state storage devices 152, as described above. Moving down one level in FIG. 2C, each non-volatile solid-state storage device 152 includes a non-volatile random access memory ("NVRAM") 204 and a relatively fast non-volatile solid-state memory such as flash memory 206. In some embodiments, the NVRAM 204 may be a component that does not require program / erase cycles (DRAM, MRAM, PCM) and may be memory that can support being written to much more frequently than the memory is read. Moving to another level in FIG. 2C, the NVRAM 204, in one embodiment, is implemented as a fast volatile memory such as dynamic random access memory (DRAM) 216 backed up by energy storage 218. Energy storage 218 provides sufficient power to continue powering DRAM 216 long enough for the contents to be transferred to flash memory 206 in the event of a power failure. In some embodiments, energy storage 218 is a capacitor, supercapacitor, battery, or other device that provides a suitable supply of energy sufficient to allow the transfer of the contents of DRAM 216 to a stable storage medium in the event of a power loss. Flash memory 206 is implemented as multiple flash dies 222, sometimes referred to as a package of flash dies 222 or an array of flash dies 222. It should be understood that flash dies 222 may be packaged in any number of ways, such as a single die per package, multiple dies per package (i.e., a multi-chip package), a hybrid package, bare dies on a circuit printed board or other substrate, encapsulated dies, etc.In the illustrated embodiment, non-volatile solid-state storage 152 includes a controller 212 or other processor and an input / output (I / O) port 210 coupled to controller 212. I / O port 210 is coupled to CPU 156 and / or network interface controller 202 of flash storage node 150. Flash input / output (I / O) port 220 is coupled to flash die 222, and direct memory access (DMA) unit 214 is coupled to controller 212, DRAM 216, and flash die 222. In the illustrated embodiment, I / O port 210, controller 212, DMA unit 214, and flash I / O port 220 are implemented on a programmable logic device (“PLD”) 208, e.g., an FPGA. In this embodiment, each flash die 222 has pages organized as 16 kB (kilobyte) pages 224 and registers 226 that can write data to or read data from the flash die 222. In further embodiments, other types of solid-state memory are used instead of or in addition to the flash memory illustrated in the flash die 222.

[0079] A storage cluster 161, in various embodiments as disclosed herein, can be generally contrasted with a storage array. Storage nodes 150 are part of a collection that makes up the storage cluster 161. Each storage node 150 owns a slice of data and the computing necessary to serve the data. Multiple storage nodes 150 cooperate to store and retrieve data. Storage memory or devices, as generally used in storage arrays, are not significantly involved in processing and manipulating data. Storage memory or devices in a storage array receive commands to read, write, or erase data. Storage memory or devices in a storage array are unaware of the larger system in which they are embedded or what the data represents. Storage memory or devices in a storage array may include various types of storage memory, such as RAM, solid-state drives, and hard disk drives. The non-volatile solid-state storage 152 units described herein have multiple interfaces that are simultaneously active and serve multiple purposes. In some embodiments, some of the functionality of a storage node 150 is shifted to the storage unit 152, transforming the storage unit 152 into a combination of storage unit 152 and storage node 150. Placing computing (over storage data) in the storage unit 152 places this computing closer to the data itself. Various system embodiments have a hierarchy of storage node tiers with different capabilities. In contrast, in a storage array, the controller owns and knows everything about all the data it manages in the shelves or storage devices. In a storage cluster 161, multiple non-volatile solid-state storage 152 units and / or multiple controllers in storage nodes 150 cooperate in various ways (e.g., for erasure coding, data sharding, metadata communication and redundancy, expanding or shrinking storage capacity, data recovery, etc.) as described herein.

[0080] FIG. 2D illustrates a storage server environment using the storage node 150 and storage 152 unit embodiments of FIGS. 2A-2C. In this version, each non-volatile solid-state storage 152 unit includes a processor, such as a controller 212 (see FIG. 2C), an FPGA, flash memory 206, and NVRAM 204 (supercapacitor-backed DRAM 216, see FIGS. 2B and 2C) on a PCIe (Peripheral Component Interconnect Express) board within the chassis 138 (see FIG. 2A). The non-volatile solid-state storage 152 units may be implemented as a single board containing storage, which may be the maximum tolerable failure domain within the chassis. In some embodiments, up to two non-volatile solid-state storage 152 units may fail, and the device continues without data loss.

[0081] In some embodiments, physical storage is divided into named regions based on application usage. NVRAM 204 is a contiguous block of reserved memory within nonvolatile solid-state storage 152, DRAM 216, and backed by NAND flash. NVRAM 204 is logically divided into multiple memory regions, two of which are written as spools (e.g., spool_regions). Space within the NVRAM 204 spool is managed independently by each authority 168. Each device provides a certain amount of storage space to each authority 168, which then manages the lifetime and allocation within that space. Examples of spools include distributed transactions or concepts. When primary power to the nonvolatile solid-state storage 152 unit fails, an onboard supercapacitor provides a short duration of power holdup. During this holdup interval, the contents of NVRAM 204 are flushed to flash memory 206. Upon the next power-up, the contents of NVRAM 204 are restored from flash memory 206.

[0082] With respect to the storage unit controller, the logical "controller" responsibilities are distributed across each of the blades, including authority 168. This distribution of logical control is illustrated in FIG. 2D as host controller 242, mid-tier controller 244, and storage unit controller 246. Control plane and storage plane management are handled independently, although some may be physically co-located on the same blade. Each authority 168 effectively functions as an independent controller. Each authority 168 provides its own data and metadata structures, its own background workers, and maintains its own lifecycle.

[0083] FIG. 2E is a hardware block diagram of a blade 252, using the embodiment of the storage node 150 and storage unit 152 of FIGS. 2A-2C in the storage server environment of FIG. 2D , illustrating a control plane 254, a compute plane 256, and a storage plane 258 that interact with the underlying physical resources, and an authority 168. The control plane 254 is divided into multiple authorities 168 that can run on any of the blades 252 using the computational resources in the compute plane 256. The storage plane 258 is divided into a set of devices that each provide access to the flash 206 and NVRAM 204 resources. In one embodiment, the compute plane 256 can perform the operations of a storage array controller on one or more devices of the storage plane 258 (e.g., a storage array) as described herein.

[0084] In the compute plane 256 and storage plane 258 of FIG. 2E, authorities 168 interact with the underlying physical resources (i.e., devices). From the perspective of an authority 168, its resources are striped across all of the physical devices. From the device's perspective, the device provides resources to all authorities 168, regardless of where the authorities happen to run. Each authority 168 allocates or is allocated one or more partitions 260 of storage memory in a storage unit 152, e.g., partitions 260 in flash memory 206 and NVRAM 204. Each authority 168 uses its allocated partitions 260 to write or read user data. Authorities can be associated with different amounts of physical storage in the system. For example, one authority 168 can have more partitions 260 or larger-sized partitions 260 in one or more storage units 152 than one or more other authorities 168.

[0085] FIG. 2F illustrates elasticity software layers within a blade 252 of a storage cluster, according to some embodiments. In an elastic architecture, the elasticity software is symmetric; that is, each blade's compute module 270 executes the three identical layers of processes depicted in FIG. 2F. A storage manager 274 executes read and write requests from other blades 252 to data and metadata stored in the local storage unit 152, NVRAM 204, and flash 206. An authority 168 fulfills client requests by issuing the necessary reads and writes to the blade 252 on the storage unit 152 where the corresponding data or metadata resides. An endpoint 272 analyzes client connection requests received from the monitoring software of the switch fabric 146, relays the client connection request to the authority 168 responsible for fulfillment, and relays the authority's 168 response to the client. The symmetric three-tier architecture enables a high degree of concurrency in the storage system. Elasticity scales out efficiently and reliably in these embodiments. Additionally, Elasticity implements inherent scale-out techniques that maximize concurrency by balancing work evenly across all resources regardless of client access patterns, eliminating much of the need for cross-blade coordination that typically occurs with traditional distributed locking.

[0086] 2F , authorities 168 executing within compute modules 270 of blades 252 perform the internal operations necessary to fulfill client requests. One feature of resiliency is that authorities 168 are stateless, i.e., they cache active data and metadata in their own blade's 252 DRAM for fast access, but they store all updates in their NVRAM 204 partitions on three separate blades 252 until the updates are written to flash 206. In some embodiments, all storage system writes to NVRAM 204 are triplicate across partitions on three separate blades 252. With triple-mirrored NVRAM 204 and persistent storage protected by parity and Reed-Solomon RAID checksums, the storage system can tolerate the simultaneous failure of two blades 252 without losing data, metadata, or access to either.

[0087] Because authorities 168 are stateless, they can migrate between blades 252. Each authority 168 has a unique identifier. NVRAM 204 and flash 206 partitions are associated with the identifier of the authority 168, not the blade 252 on which they are running. Thus, when an authority 168 migrates, the authority 168 continues to manage the same storage partitions from its new location. When a new blade 252 is installed in one embodiment of a storage cluster, the system automatically rebalances the load by partitioning the new blade's 252's storage for use by authorities 168 in the system, migrating selected authorities 168 to the new blade 252, and starting endpoints 272 on the new blade 252 and including them in the switch fabric 146's client connection distribution algorithm.

[0088] From their new locations, the migrated authorities 168 persist the contents of their NVRAM 204 partitions on flash 206, process read and write requests from other authorities 168, and fulfill client requests that endpoints 272 direct to them. Similarly, if a blade 252 fails or is removed, the system redistributes its authorities 168 among the remaining blades 252 in the system. The redistributed authorities 168 continue to perform their original functions from their new locations.

[0089] FIG. 2G depicts authorities 168 and storage resources within blades 252 of a storage cluster, according to some embodiments. Each authority 168 is exclusively responsible for a partition of flash 206 and NVRAM 204 on each blade 252. An authority 168 manages the contents and integrity of its partition independently of other authorities 168. Authorities 168 compress incoming data, temporarily store it in their NVRAM 204 partitions, and then consolidate, RAID-protect, and persist the data in segments of storage in their flash 206 partitions. As authorities 168 write data to flash 206, storage manager 274 performs the necessary flash transformations to optimize write performance and maximize media lifespan. In the background, authorities 168 “garbage collect,” i.e., reclaim space occupied by data no longer needed by clients overwriting data. It should be appreciated that because the partitions of authorities 168 are disjoint, there is no need for distributed locking to execute clients and writes or to execute background functions.

[0090] The embodiments described herein may utilize various software, communication, and / or networking protocols. In addition, hardware and / or software configurations may be adjusted to accommodate various protocols. For example, embodiments may utilize Active Directory, a database-based system that provides authentication, directory, policy, and other services in a WINDOWS™ environment. In these embodiments, the Lightweight Directory Access Protocol (LDAP) is an example of an application protocol for querying and modifying entries in a directory service provider such as Active Directory. In some embodiments, a network lock manager ("NLM") is utilized in conjunction with the Network File System ("NFS") to provide System V-style advisory file and record locking over a network. The Server Message Block ("SMB") protocol, one version of which is also known as the Common Internet File System ("CIFS"), may be integrated with the storage systems described herein. SMP operates as an application-layer network protocol typically used to provide shared access to files, printers, and serial ports, as well as various communications between nodes on a network. SMB also provides an authenticated inter-process communication mechanism. AMAZON™ S3 (Simple Storage Service) is a web service provided by Amazon Web Services, and the system described herein can interface with Amazon S3 through web service interfaces (REST (Representational State Transfer), SOAP (Simple Object Access Protocol), and BitTorrent). A RESTful API (Application Programming Interface) breaks down transactions into a series of small modules.Each module addresses a specific underlying part of a transaction. Control or permission provided in these embodiments, particularly for object data, may include the use of access control lists ("ACLs"). An ACL is a list of permissions attached to an object, specifying which users or system processes are allowed to access the object and which actions are permitted for a given object. The system provides an identification and location system for computers on the network and may utilize Internet Protocol version 6 ("IPv6") as well as IPv4 for communication protocols that route traffic through the Internet. Routing of packets between networked systems may include equal-cost multi-path routing ("ECMP"), a routing strategy in which next-hop packet forwarding to a single destination may occur over multiple "best paths" that combine at the top of a routing metric calculation. Multi-path routing can be used with most routing protocols because it is a hop-by-hop decision limited to a single router. The software may support multi-tenancy, an architecture in which a single instance of a software application serves multiple customers. Each customer may be referred to as a tenant. Tenants may be given the ability to customize some portions of the application, although in some embodiments, they may not customize the application's code. Embodiments may maintain audit logs. An audit log is a document that records events in a computing system. In addition to documenting which resources were accessed, audit log entries typically include destination and source addresses, timestamps, and user login information for compliance with various regulations. Embodiments may support various key management policies, such as encryption key rotation.Additionally, the system may support dynamic root passwords or some variation that allows passwords to change dynamically.

[0091] FIG. 3A illustrates a diagram of a storage system 306 coupled for data communication with a cloud service provider 302, according to some embodiments of the present disclosure. While not depicted in greater detail, the storage system 306 depicted in FIG. 3A may be similar to the storage systems described above with reference to FIGS. 1A-1D and 2A-2G. In some embodiments, the storage system 306 depicted in FIG. 3A may be embodied as a storage system including unbalanced active / active controllers, a storage system including balanced active / active controllers, a storage system including active / active controllers in which fewer than all of each controller's resources are utilized such that each controller has spare resources that can be used to support failover, a storage system including fully active / active controllers, a storage system including controllers with separated data sets, a storage system including a dual-tier architecture with a front-end controller and a back-end unified storage controller, a storage system including a scale-out cluster of dual-controller arrays, and combinations of such embodiments.

[0092] 3A , storage system 306 is coupled to cloud service provider 302 via data communications link 304. Data communications link 304 may be embodied as a dedicated data communications link, as a data communications path provided through the use of one or more data communications networks, such as a wide area network ("WAN") or LAN, or as some other mechanism capable of transferring digital information between storage system 306 and cloud service provider 302. Such data communications link 304 may be entirely wired, entirely wireless, or some collection of wired and wireless data communications paths. In such an example, digital information may be exchanged between storage system 306 and cloud service provider 302 via data communications link 304 using one or more data communications protocols. For example, digital information may be exchanged between storage system 306 and cloud service provider 302 via data communications link 304 using Handheld Device Transfer Protocol ("HDTP"), Hypertext Transfer Protocol ("HTTP"), Internet Protocol ("IP"), Real-Time Transport Protocol ("RTP"), Transmission Control Protocol ("TCP"), User Datagram Protocol ("UDP"), Wireless Application Protocol ("WAP"), or other protocols.

[0093] The cloud service provider 302 depicted in FIG. 3A may be embodied as a system and computing environment that provides a vast array of services to users of the cloud service provider 302, for example, through the sharing of computing resources over a data communication link 304. The cloud service provider 302 can provide on-demand access to a shared pool of configurable computing resources, such as computer networks, servers, storage, applications, and services. The shared pool of configurable resources can be rapidly provisioned and released to users of the cloud service provider 302 with minimal administrative effort. Generally, users of the cloud service provider 302 are unaware of the exact computing resources utilized by the cloud service provider 302 to provide services. While such cloud service providers 302 may often be accessible via the Internet, readers skilled in the art will recognize that any system that abstracts the use of shared resources to provide services to users over any data communication link may be considered a cloud service provider 302.

[0094] 3A, cloud service provider 302 may be configured to provide various services to storage system 306 and users of storage system 306 through the implementation of various service models. For example, cloud service provider 302 may be configured to provide services through an implementation of an infrastructure as a service ("IaaS") service model, through an implementation of a platform as a service ("PaaS") service model, through an implementation of a software as a service ("SaaS") service model, through an implementation of an authentication as a service ("AaaS") service model, or through an implementation of a storage as a service model in which cloud service provider 302 provides access to its storage infrastructure for use by storage system 306 and users of storage system 306. The reader will understand that the above-described service models are included for illustrative purposes only and do not represent limitations on the services that may be provided by cloud service provider 302 or on the service models that may be implemented by cloud service provider 302, and that cloud service provider 302 may be configured to provide additional services to storage system 306 and users of storage system 306 through the implementation of additional service models.

[0095] 3A , cloud service provider 302 may be embodied as, for example, a private cloud, a public cloud, or a combination of private and public clouds. In an embodiment in which cloud service provider 302 is embodied as a private cloud, cloud service provider 302 may be dedicated to serving a single organization rather than serving multiple organizations. In an embodiment in which cloud service provider 302 is embodied as a public cloud, cloud service provider 302 may provide services to multiple organizations. In yet an alternative embodiment, cloud service provider 302 may be embodied as a mix of private and public cloud services, comprising a hybrid cloud deployment.

[0096] Although not explicitly depicted in FIG. 3A , the reader will understand that a significant amount of additional hardware and software components may be required to facilitate the delivery of cloud services to the storage system 306 and users of the storage system 306. For example, the storage system 306 may be coupled to (or include) a cloud storage gateway. Such a cloud storage gateway may be embodied, for example, as a hardware- or software-based appliance located on-premises with the storage system 306. Such a cloud storage gateway can act as a bridge between local applications running on the storage system 306 and remote, cloud-based storage utilized by the storage system 306. Through the use of a cloud storage gateway, an organization may be able to move its primary iSCSI or NAS storage to the cloud service provider 302, thereby enabling the organization to conserve space on its on-premises storage system. Such a cloud storage gateway may be configured to emulate a disk array, block-based device, file server, or other storage system that can translate SCSI commands, file server commands, or other appropriate commands into a RESTful space protocol that facilitates communication with the cloud service provider 302.

[0097] To enable storage system 306 and users of storage system 306 to utilize services offered by cloud service provider 302, a cloud migration process may be performed, during which data, applications, or other elements from an organization's local system (or from another cloud environment) are moved to cloud service provider 302. To successfully migrate data, applications, or other elements to the cloud service provider's 302 environment, middleware such as a cloud migration tool may be utilized to bridge the gap between the cloud service provider's 302 environment and the organization's environment. Such cloud migration tools may also be configured to address potentially high network costs and long transfer times associated with migrating large amounts of data to cloud service provider 302, as well as security issues associated with transmitting sensitive data over a data communications network to cloud service provider 302. To further enable storage system 306 and users of storage system 306 to utilize services offered by cloud service provider 302, a cloud orchestrator may also be used to arrange and coordinate automated tasks in pursuit of creating an integrated process or workflow. Such a cloud orchestrator can perform tasks such as configuring various components, whether they are cloud or on-premise components, and managing the interconnections between such components. The cloud orchestrator can simplify inter-component communication and connections to ensure that links are properly configured and maintained.

[0098] In the example depicted in FIG. 3A , as briefly described above, cloud service provider 302 may be configured to provide services to storage system 306 and users of storage system 306 through the use of a SaaS service model, eliminating the need to install and run applications on local computers and simplifying application maintenance and support. Such applications may take many forms in accordance with various embodiments of the present disclosure. For example, cloud service provider 302 may be configured to provide storage system 306 and users of storage system 306 with access to a data analysis application. Such data analysis application may be configured to receive, for example, vast amounts of telemetry data transmitted by storage system 306 to their homes. Such telemetry data may describe various operational characteristics of storage system 306 and can be analyzed for a myriad of purposes, including, for example, determining the health of storage system 306, identifying workloads running on storage system 306, predicting when storage system 306 will run out of various resources, and recommending configuration changes, hardware or software upgrades, workflow transitions, or other actions that can improve the operation of storage system 306.

[0099] Cloud service provider 302 may also be configured to provide access to virtualized computing environments to storage system 306 and users of storage system 306. Such virtualized computing environments may be embodied, for example, as virtual machines or other virtualized computer hardware platforms, virtual storage devices, virtualized computer network resources, etc. Examples of such virtualized environments may include virtual machines created to emulate real computers, virtualized desktop environments that separate logical desktops from physical machines, virtualized file systems that allow uniform access to different types of concrete file systems, etc.

[0100] While the example depicted in FIG. 3A illustrates storage system 306 being coupled for data communication with cloud service provider 302, in other embodiments, storage system 306 may be part of a hybrid cloud deployment in which private cloud elements (e.g., private cloud services, on-premises infrastructure, etc.) and public cloud elements (e.g., private cloud services, infrastructure, etc., that may be provided by one or more cloud service providers) are combined to form a single solution through orchestration between various platforms. Such hybrid cloud deployments may leverage hybrid cloud management software, such as, for example, Microsoft™'s Azure™ Arc, which centralizes management of the hybrid cloud deployment to any infrastructure and enables deployment of services anywhere. In such an example, the hybrid cloud management software may be configured to create, update, and delete resources (both physical and virtual) that form the hybrid cloud deployment, allocate compute and storage to specific workloads, monitor workloads and resources for performance, policy compliance, updates and patches, security status, or perform various other tasks.

[0101] The reader will understand that pairing the storage systems described herein with one or more cloud service providers can enable a variety of offerings. For example, disaster recovery as a service ("DRaaS") can be provided, in which cloud resources are utilized to protect applications and data from disruptions caused by disasters, including embodiments in which the storage system can serve as a primary data store. In such embodiments, full system backups can be taken to enable business continuity in the event of a system failure. In such embodiments, cloud data backup technology (by itself or as part of a larger DRaaS solution) can also be integrated into an overall solution that includes the storage systems and cloud service providers described herein.

[0102] The storage systems described herein, as well as cloud service providers, can be utilized to provide a variety of security features. For example, the storage system can encrypt data at rest (encrypted data can be sent to and from the storage system) and can utilize Key Management-as-a-Service ("KMaaS") to manage encryption keys, keys for locking and unlocking storage devices, and the like. Similarly, a cloud data security gateway or similar mechanism can be utilized to ensure that data stored within the storage system is not improperly stored in the cloud as part of a cloud data backup operation. Furthermore, microsegmentation or identity-based segmentation can be utilized in data centers that include the storage systems or within cloud service providers to create secure zones that allow workloads to be isolated from one another in data center and cloud deployments.

[0103] For further explanation, Figure 3B sets forth a diagram of a storage system 306 according to some embodiments of the present disclosure. Although not depicted in greater detail, the storage system 306 depicted in Figure 3B may be similar to the storage systems described above with reference to Figures 1A-1D and 2A-2G, as the storage systems may include many of the components described above.

[0104] 3B may include a vast amount of storage resources 308, which may be embodied in many forms. For example, the storage resources 308 may include nanoRAM or another form of nonvolatile random-access memory utilizing carbon nanotubes deposited on a substrate, 3D cross-point nonvolatile memory, flash memory, including single-level cell ("SLC") NAND flash, multi-level cell ("MLC") NAND flash, triple-level cell ("TLC") NAND flash, quad-level cell ("QLC") NAND flash, or others. Similarly, the storage resources 308 may include nonvolatile magnetoresistive random-access memory ("MRAM"), including spin transfer torque ("STT") MRAM. Exemplary storage resource 308 may alternatively include other forms of storage resources, including non-volatile phase-change memory ("PCM"), quantum memory that enables the storage and retrieval of photonic quantum information, resistive random-access memory ("ReRAM"), storage class memory ("SCM"), or any combination of the resources described herein. The reader will understand that other forms of computer memory and storage devices may be utilized by the above-described storage system, including DRAM, SRAM, EEPROM, universal memory, etc.The storage resources 308 depicted in FIG. 3A may be embodied in a variety of form factors, including, but not limited to, dual in-line memory modules ("DIMMs"), non-volatile dual in-line memory modules ("NVDIMMs"), M.2, U.2, and others.

[0105] The storage resources 308 depicted in FIG. 3B may include various forms of SCM. SCM can effectively treat high-speed, non-volatile memory (e.g., NAND flash) as an extension of DRAM, so that the entire data set can be treated as an in-memory data set residing entirely within DRAM. SCM can include, for example, non-volatile media such as NAND flash. Such NAND flash may be accessed using NVMe, which can use the PCIe bus as its transport, offering relatively low access latency compared to older protocols. Indeed, network protocols used for SSDs in all-flash arrays include NVMe over Ethernet (ROCE, NVMe TCP), Fibre Channel (NVMe FC), InfiniBand (iWARP), and others that enable treating high-speed, non-volatile memory as an extension of DRAM. Given the fact that DRAM is often byte-addressable and high-speed, non-volatile memory such as NAND flash is block-addressable, a controller software / hardware stack may be required to convert block data into bytes stored on the media. Examples of media and software that can be used as SCM can include, for example, 3D XPoint, Intel Memory Drive Technology, Samsung's Z-SSD, and others.

[0106] The storage resource 308 depicted in FIG. 3B can also include racetrack memory (also referred to as domain wall memory). Such racetrack memory can be embodied in a solid-state device as a form of nonvolatile solid-state memory that relies on the charge of electrons as well as the unique strength and orientation of the magnetic field generated by electrons as they spin. By using spin-coherent currents to move magnetic domains along nanoscale Permalloy wires, the magnetic domains can pass by a magnetic read / write head positioned near the wire as the current passes through the wire, thereby altering the magnetic domains and recording a pattern of bits. Many such wires and read / write elements can be packaged together to create a racetrack memory device.

[0107] The exemplary storage system 306 depicted in FIG. 3B may implement a variety of storage architectures. For example, a storage system according to some embodiments of the present disclosure may utilize block storage, where data is stored in blocks, with each block essentially acting as an individual hard drive. A storage system according to some embodiments of the present disclosure may utilize object storage, where data is managed as objects. Each object may include the data itself, a variable amount of metadata, and a globally unique identifier, and object storage may be implemented at multiple levels (e.g., device level, system level, interface level). A storage system according to some embodiments of the present disclosure utilizes file storage, where data is stored in a hierarchical structure. Such data is stored in files and folders and can be presented in the same format to both the system that stores it and the system that retrieves it.

[0108] 3B may be embodied as a storage system in which additional storage resources can be added through the use of a scale-up model, a scale-out model, or some combination thereof. In a scale-up model, additional storage may be added by adding additional storage devices. However, in a scale-out model, additional storage nodes may be added to a cluster of storage nodes, and such storage nodes may include additional processing resources, additional networking resources, etc.

[0109] 3B may utilize the storage resources described above in a variety of different ways. For example, portions of the storage resources may be utilized to function as a write cache, storage resources within the storage system may be utilized as a read cache, or tiering may be achieved within the storage system by placing data within the storage system according to one or more tiering policies.

[0110] 3B also includes communication resources 310 that may be useful for facilitating data communication between components within the storage system 306, as well as between the storage system 306 and computing devices external to the storage system 306, including embodiments in which these resources are separated by relatively large expanses. The communication resources 310 may be configured to utilize a variety of different protocols and data communication fabrics to facilitate data communication between components within the storage system and computing devices external to the storage system. For example, communication resources 310 may include Fibre Channel ("FC") technology, such as an FC fabric and FC protocol capable of transporting SCSI commands over an FC network, FC over Ethernet ("FCoE") technology, in which FC frames are encapsulated and transmitted over an Ethernet network, InfiniBand ("IB") technology, in which a switched fabric topology is utilized to facilitate transmission between channel adapters, NVM Express ("NVMe") technology and NVMe over fabric ("NVMeoF") technology, in which non-volatile storage media attached via a PCI Express ("PCIe") bus can be accessed, and others. Indeed, the storage systems described above may directly or indirectly utilize neutrino communication technologies and devices in which information (including binary information) is transmitted using beams of neutrinos.

[0111] The communication resources 310 may also include mechanisms for accessing the storage resources 308 in the storage system 306 using serial attached SCSI ("SAS"), serial ATA ("SATA") bus interfaces for connecting the storage resources 308 in the storage system 306 to host bus adapters in the storage system 306, Internet Small Computer System Interface ("iSCSI") technology for providing block-level access to the storage resources 308 in the storage system 306, and other communication resources that may be useful in facilitating data communication between components within the storage system 306, as well as data communication between the storage system 306 and computing devices external to the storage system 306.

[0112] 3B also includes processing resources 312 that may be useful for executing computer program instructions and performing other computational tasks within the storage system 306. The processing resources 312 may include one or more ASICs customized for any particular purpose, as well as one or more CPUs. The processing resources 312 may also include one or more DSPs, one or more FPGAs, one or more systems on a chip ("SoC"), or other forms of processing resources 312. The storage system 306 may utilize the storage resources 312 to perform various tasks, including, but not limited to, supporting the execution of software resources 314, which are described in more detail below.

[0113] 3B also includes software resources 314 that, when executed by processing resources 312 within storage system 306, can perform a vast number of tasks. Software resources 314 may include, for example, one or more modules of computer program instructions that, when executed by processing resources 312 within storage system 306, are useful for implementing various data protection techniques. Such data protection techniques may be implemented, for example, by system software running on computer hardware within the storage system, by a cloud service provider, or otherwise. Such data protection techniques may include data archiving, data backup, data replication, data snapshots, data and database cloning, and other data protection techniques.

[0114] Software resources 314 may also include software that is useful in implementing software-defined storage ("SDS"). In such an example, software resources 314 may include one or more modules of computer program instructions that, when executed, are useful in policy-based provisioning and management of data storage independent of the underlying hardware. Such software resources 314 may be useful in implementing storage virtualization to separate storage hardware from the software that manages the storage hardware.

[0115] The software resources 314 may also include software useful for facilitating and optimizing I / O operations directed to the storage system 306. For example, the software resources 314 may include software modules that implement various data reduction techniques, such as data compression, data deduplication, and others. The software resources 314 may include software modules that intelligently group I / O operations to facilitate better use of the underlying storage resources 308, software modules that perform data migration operations for migrating data from within the storage system, and software modules that perform other functions. Such software resources 314 may be embodied as one or more software containers or in many other ways.

[0116] For further explanation, Figure 3C sets forth an example of a cloud-based storage system 318 according to some embodiments of the present disclosure. In the example depicted in Figure 3C, the cloud-based storage system 318 is created entirely within a cloud computing environment 316, such as, for example, Amazon Web Services ("AWS")™, Microsoft Azure™, Google Cloud Platform™, IBM Cloud™, Oracle Cloud™, and others. The cloud-based storage system 318 can be used to provide services similar to those that can be provided by the storage systems described above.

[0117] The cloud-based storage system 318 depicted in FIG. 3C includes two cloud computing instances 320, 322, each used to support the execution of storage controller applications 324, 326. The cloud computing instances 320, 322 may be embodied as instances of cloud computing resources (e.g., virtual machines) that may be provided by the cloud computing environment 316 to support the execution of software applications such as the storage controller applications 324, 326. For example, each of the cloud computing instances 320, 322 may run on an Azure VM, and each Azure VM may include high-speed temporary storage that can be utilized as a cache (e.g., as a read cache). In one embodiment, the cloud computing instances 320, 322 may be embodied as Amazon Elastic Compute Cloud ("EC2") instances. In such an example, an Amazon Machine Image ("AMI") that includes the storage controller applications 324, 326 may be booted to create and configure a virtual machine capable of running the storage controller applications 324, 326.

[0118] 3C , the storage controller applications 324, 326 may be embodied as modules of computer program instructions that, when executed, perform various storage tasks. For example, the storage controller applications 324, 326 may be embodied as modules of computer program instructions that, when executed, perform the same tasks as the controllers 110A, 110B of FIG. 1A described above, such as writing data to the cloud-based storage system 318, erasing data from the cloud-based storage system 318, retrieving data from the cloud-based storage system 318, monitoring and reporting disk usage and performance, performing redundancy operations such as RAID or RAID-like data redundancy operations, compressing data, encrypting data, deduplication data, etc. The reader will understand that, because there are two cloud computing instances 320, 322, each including a storage controller application 324, 326, in some embodiments, one cloud computing instance 320 can operate as a primary controller as described above, and the other cloud computing instance 322 can operate as a secondary controller as described above. The reader will understand that the storage controller applications 324, 326 depicted in FIG. 3C may comprise the same source code running within different cloud computing instances 320, 322, such as separate EC2 instances.

[0119] The reader will understand that other embodiments that do not include primary and secondary controllers are within the scope of this disclosure. For example, each cloud computing instance 320, 322 can act as a primary controller for some portion of the address space supported by the cloud-based storage system 318, each cloud computing instance 320, 322 can act as a primary controller where servicing of I / O operations directed to the cloud-based storage system 318 is divided in some other manner, and so on. Indeed, in other embodiments where cost savings may take priority over performance requirements, there may be only a single cloud computing instance that includes the storage controller application.

[0120] The cloud-based storage system 318 depicted in Figure 3C includes cloud computing instances 340a, 340b, 340n having local storage 330, 334, 338. The cloud computing instances 340a, 340b, 340n may be embodied as instances of cloud computing resources that may be provided by the cloud computing environment 316 to support the execution of software applications, for example. The cloud computing instances 340a, 340b, 340n of Figure 3C may differ from the cloud computing instances 320, 322 described above because the cloud computing instances 340a, 340b, 340n of Figure 3C have local storage 330, 334, 338 resources, whereas the cloud computing instances 320, 322 that support the execution of storage controller applications 324, 326 need not have local storage resources. Cloud computing instances 340a, 340b, 340n having local storage 330, 334, 338 may be embodied, for example, as an EC2 M5 instance including one or more SSDs, as an EC2 R5 instance including one or more SSDs, as an EC2 I3 instance including one or more SSDs, etc. In some embodiments, local storage 330, 334, 338 must be embodied as solid-state storage (e.g., SSDs) rather than storage that utilizes hard disk drives.

[0121] 3C , each of the cloud computing instances 340 a, 340 b, 340 n having local storage 330, 334, 338 may include a software daemon 328, 332, 336 that, when executed by the cloud computing instances 340 a, 340 b, 340 n, can present itself to the storage controller application 324, 326 as if the cloud computing instance 340 a, 340 b, 340 n were a physical storage device (e.g., one or more SSDs). In such an example, the software daemon 328, 332, 336 may include computer program instructions similar to those that would typically be included on a storage device, such that the storage controller application 324, 326 can send and receive the same commands that a storage controller would send to a storage device. In this manner, the storage controller application 324, 326 may include code that is the same (or substantially the same) as the code executed by the controller in the storage system described above. In these and similar embodiments, communication between the storage controller applications 324, 326 and the cloud computing instances 340a, 340b, 340n with the local storage 330, 334, 338 may utilize iSCSI, NVMe over TCP, messaging, a custom protocol, or some other mechanism.

[0122] 3C , each of the cloud computing instances 340a, 340b, 340n having local storage 330, 334, 338 may also be coupled to block storage 342, 344, 346 provided by the cloud computing environment 316, such as, for example, Amazon Elastic Bookstore (Elastic Block Store, “EBS”) volumes. In such an example, the block storage 342, 344, 346 provided by the cloud computing environment 316 may be utilized in a manner similar to the way NVRAM devices described above are utilized, in that when a software daemon 328, 332, 336 (or some other module) running within a particular cloud computing instance 340a, 340b, 340n receives a request to write data, it may initiate writing data to its attached EBS volume as well as to its local storage 330, 334, 338 resource. In some alternative embodiments, data may be written only to local storage 330, 334, 338 resources within the particular cloud containing the instance 340 a, 340 b, 340 n. In alternative embodiments, rather than using block storage 342, 344, 346 provided by the cloud computing environment 316 as NVRAM, actual RAM on each of the cloud computing instances 340 a, 340 b, 340 n having local storage 330, 334, 338 may be used as NVRAM, thereby reducing network utilization costs associated with using EBS volumes as NVRAM. In yet another embodiment, high-performance block storage resources such as one or more Azure Ultra Disks may be utilized as NVRAM.

[0123] Storage controller applications 324, 326 may be used to perform various tasks such as deduplicating the data included in the request, compressing the data included in the request, determining where to write the data included in the request, and then ultimately sending a request to write a deduplicated, encrypted, or possibly updated version of the data to one or more of cloud computing instances 340a, 340b, 340n having local storage 330, 334, 338. Either of cloud computing instances 320, 322, in some embodiments, may receive a request to read data from cloud-based storage system 318 and ultimately send the request to read the data to one or more of cloud computing instances 340a, 340b, 340n having local storage 330, 334, 338.

[0124] When a request to write data is received by a particular cloud computing instance 340a, 340b, 340n having local storage 330, 334, 338, the software daemons 328, 332, 336 may be configured not only to write the data to their own local storage 330, 334, 338 resources and any suitable block storage 342, 344, 346 resources, but the software daemons 328, 332, 336 may also be configured to write the data to cloud-based object storage 348 attached to the particular cloud computing instance 340a, 340b, 340n. The cloud-based object storage 348 attached to the particular cloud computing instance 340a, 340b, 340n may be embodied as, for example, Amazon Simple Storage Service ("S3"). In other embodiments, cloud computing instances 320, 322, including storage controller applications 324, 326, respectively, can initiate storage of data to local storage 330, 334, 338 of cloud computing instances 340a, 340b, 340n and cloud-based object storage 348. In other embodiments, rather than storing data using both cloud computing instances 340a, 340b, 340n with local storage 330, 334, 338 (also referred to herein as "virtual drives") and cloud-based object storage 348, the persistent storage tier may be implemented in other ways. For example, one or more Azure Ultra disks may be used to persistently store data (e.g., after the data is written to the NVRAM tier).

[0125] While the local storage 330, 334, 338 and block storage 342, 344, 346 resources utilized by cloud computing instances 340a, 340b, 340n may support block-level access, the cloud-based object storage 348 attached to a particular cloud computing instance 340a, 340b, 340n only supports object-based access. Thus, software daemons 328, 332, 336 may be configured to retrieve blocks of data, package those blocks into objects, and write the objects to the cloud-based object storage 348 attached to a particular cloud computing instance 340a, 340b, 340n.

[0126] Consider an example in which data is written in 1 MB blocks to the local storage 330, 334, 338 and block storage 342, 344, 346 resources utilized by cloud computing instances 340 a, 340 b, 340 n. In such an example, assume that a user of the cloud-based storage system 318 issues a request to write data that, after being compressed and deduplicated by the storage controller applications 324, 326, will require writing 5 MB of data. In such an example, writing the data to the local storage 330, 334, 338 and block storage 342, 344, 346 resources utilized by the cloud computing instances 340 a, 340 b, 340 n is relatively straightforward because five blocks, each 1 MB in size, are written to the local storage 330, 334, 338 and block storage 342, 344, 346 resources utilized by the cloud computing instances 340 a, 340 b, 340 n. In such an example, software daemons 328, 332, 336 may also be configured to create five objects containing separate 1 MB chunks of data. Thus, in some embodiments, each object written to cloud-based object storage 348 may be identical (or nearly identical) in size. Readers will understand that in such an example, each object may include metadata associated with the data itself (e.g., the first 1 MB of the object is the data, and the remainder is metadata associated with the data). Readers will understand that cloud-based object storage 348 can be incorporated into cloud-based storage system 318 to increase the durability of cloud-based storage system 318.

[0127] In some embodiments, all data stored by cloud-based storage system 318 may be stored in both 1) cloud-based object storage 348 and 2) at least one of local storage 330, 334, 338 or block storage 342, 344, 346 resources utilized by cloud computing instances 340a, 340b, 340n. In such embodiments, the local storage 330, 334, 338 and block storage 342, 344, 346 resources utilized by cloud computing instances 340a, 340b, 340n may effectively operate as a cache that generally contains all data that is also stored in S3, such that all reads of data may be serviced by cloud computing instances 340a, 340b, 340n without requiring cloud computing instances 340a, 340b, 340n to access cloud-based object storage 348. However, the reader will understand that in other embodiments, all data stored by cloud-based storage system 318 may be stored in cloud-based object storage 348, but less than all data stored by cloud-based storage system 318 may be stored in at least one of the local storage 330, 334, 338 resources or block storage 342, 344, 346 resources utilized by cloud computing instances 340a, 340b, 340n. In such examples, various policies may be utilized to determine which subsets of data stored by cloud-based storage system 318 should reside in both 1) cloud-based object storage 348 and 2) at least one of the local storage 330, 334, 338 resources or block storage 342, 344, 346 resources utilized by cloud computing instances 340a, 340b, 340n.

[0128] One or more modules of computer program instructions executing within the cloud-based storage system 318 (e.g., a monitoring module running on its own EC2 instance) may be designed to handle the failure of one or more of the cloud computing instances 340a, 340b, 340n having local storage 330, 334, 338. In such an example, the monitoring module may handle the failure of one or more of the cloud computing instances 340a, 340b, 340n having local storage 330, 334, 338 by creating one or more new cloud computing instances having local storage, retrieving data stored in the failed cloud computing instances 340a, 340b, 340n from cloud-based object storage 348, and storing the data retrieved from cloud-based object storage 348 in local storage in the newly created cloud computing instances. The reader will understand that many variations of this process may be implemented.

[0129] The reader will understand that various performance aspects of the cloud-based storage system 318 can be monitored (e.g., by a monitoring module running on an EC2 instance) so that the cloud-based storage system 318 can be scaled up or out as needed. For example, if the cloud computing instances 320, 322 used to support the execution of the storage controller applications 324, 326 are undersized and are not adequately servicing the I / O requests issued by users of the cloud-based storage system 318, the monitoring module may create a new, more powerful cloud computing instance (e.g., a type of cloud computing instance that includes more processing power, more memory, etc.) that includes the storage controller application so that the new, more powerful cloud computing instance can begin operating as the primary controller. Similarly, if the monitoring module determines that the cloud computing instances 320, 322 used to support the execution of the storage controller applications 324, 326 are oversized and cost savings can be achieved by switching to smaller, less powerful cloud computing instances, the monitoring module can create a new, less powerful (and cheaper) cloud computing instance that includes the storage controller application so that the new, less powerful cloud computing instance can begin operating as the primary controller.

[0130] The storage system described above may implement intelligent data backup techniques, which allow data stored in the storage system to be copied and stored in a different location to avoid data loss in the event of equipment failure or other forms of catastrophic disaster. For example, the storage system described above may be configured to inspect each backup to avoid restoring the storage system to an undesirable state. Consider an example in which malware infects a storage system. In such an example, the storage system may include software resource 314 that can scan each backup to distinguish between backups captured before the malware infected the storage system and backups captured after the malware infected the storage system. In such an example, the storage system may restore itself from a backup that does not contain the malware, or at least may not restore the portion of the backup that contained the malware. In such an example, the storage system may include software resource 314 that may scan each backup to identify the presence of malware (or viruses, or anything else undesirable), for example, by identifying write operations served by the storage system that originate from network subnets served by the storage system that are suspected of delivering malware, by identifying write operations served by the storage system that originate from users that are suspected of delivering malware, by identifying write operations served by the storage system, by inspecting the content of the write operations against malware fingerprints, and in many other ways.

[0131] The reader will further appreciate that backups (often in the form of one or more snapshots) may also be utilized to facilitate rapid recovery of the storage system. Consider an example where a storage system is infected with ransomware that locks users out of the storage system. In such an example, software resources 314 within the storage system may be configured to detect the presence of the ransomware and may further be configured to restore the storage system to a point in time using retained backups prior to the time the ransomware infected the storage system. In such an example, the presence of ransomware may be detected explicitly through the use of software tools utilized by the system, through the use of a key (e.g., a USB drive) inserted into the storage system, or similar methods. Similarly, the presence of ransomware may be inferred in response to system activity that meets a predetermined fingerprint, such as, for example, no reads or writes being input to the system for a predetermined period of time.

[0132] The reader will understand that the various components described above can be grouped into one or more optimized computing packages as an integrated infrastructure. Such an integrated infrastructure can include a pool of computer, storage, and networking resources that can be shared by multiple applications and collectively managed using policy-driven processes. Such an integrated infrastructure can be implemented using an integrated infrastructure reference architecture, using standalone equipment, using a software-driven hyper-integrated approach (e.g., a hyper-integrated infrastructure), or in other ways.

[0133] The reader will understand that the storage systems described in this disclosure may be useful for supporting various types of software applications. Indeed, the storage systems may be "application-aware," in the sense that the storage systems may acquire, maintain, or otherwise access information describing connected applications (e.g., applications that utilize the storage system) and optimize their operation based on intelligence about the applications and their usage patterns. For example, the storage systems may optimize data layout, optimize caching behavior, optimize "QoS" levels, or perform some other optimization designed to improve the storage performance experienced by the applications.

[0134] As an example of one type of application that may be supported by the storage system described herein, the storage system 306 may be useful in supporting artificial intelligence ("AI") applications, database applications, XOps projects (e.g., DevOps projects, DataOps projects, MLOps projects, ModelOps projects, PlatformOps projects), electronic design automation tools, event-driven software applications, high-performance computing applications, simulation applications, high-speed data capture and analysis applications, machine learning applications, media production applications, media serving applications, picture archiving and communication system ("PACS") applications, software development applications, virtual reality applications, augmented reality applications, and many other types of applications by providing storage resources to such applications.

[0135] Given the fact that storage systems include computational resources, storage resources, and a wide variety of other resources, storage systems may be well suited to supporting resource-intensive applications such as, for example, AI applications, which may be deployed in a variety of fields, including predictive maintenance in manufacturing and related fields, healthcare applications such as patient data and risk analytics, retail and marketing deployments (e.g., search advertising, social media advertising), supply chain solutions, fintech solutions such as business analytics and reporting tools, operational deployments such as real-time analytics tools, application performance management tools, IT infrastructure management tools, and the like.

[0136] Such AI applications may enable devices to perceive their environment and take actions that maximize their chances of success for some purpose. Examples of such AI applications may include IBM Watson™, Microsoft Oxford™, Google DeepMind™, Baidu Minwa™, and others.

[0137] The storage systems described above are also well suited to supporting other types of resource-intensive applications, such as machine learning applications. Machine learning applications can perform various types of data analysis and automate the construction of analytical models. Using algorithms that iteratively learn from data, machine learning applications can enable computers to learn without being explicitly programmed. One particular area of ​​machine learning is called reinforcement learning, which involves taking appropriate actions to maximize rewards in specific situations.

[0138] In addition to the resources already described, the storage systems described above may also include a graphics processing unit (GPU), sometimes referred to as a visual processing unit (VPU). Such a GPU may be embodied as dedicated electronic circuitry that rapidly manipulates and modifies memory to accelerate the creation of images in a frame buffer intended for output to a display device. Such a GPU may be included within any of the computing devices that are part of the storage systems described above, including as one of many individually scalable components of the storage system; other examples of individually scalable components of such storage systems may include storage components, memory components, computational components (e.g., CPUs, FPGAs, ASICs), networking components, software components, and others. In addition to the GPU, the storage systems described above may also include a neural network processor (NNP) for use in various aspects of neural network processing. Such NNPs may be used instead of (or in addition to) a GPU and may be independently scalable.

[0139] As mentioned above, the storage systems described herein can be configured to support artificial intelligence applications, machine learning applications, big data analytics applications, and many other types of applications. The rapid growth in these types of applications is driven by three technologies: deep learning (DL), GPU processors, and big data. Deep learning is a computing model that utilizes massively parallel neural networks inspired by the human brain. Instead of experts handcrafting software, deep learning models write their own software by learning from many examples. Such GPUs can contain thousands of cores that are well suited to running algorithms that roughly represent the parallelism of the human brain.

[0140] Advances in deep neural networks, including the development of multi-layer neural networks, have sparked a new wave of algorithms and tools for data scientists to harness their data with artificial intelligence (AI). With improved algorithms, larger datasets, and a variety of frameworks (including open-source software libraries for machine learning across a range of tasks), data scientists are addressing new use cases such as autonomous vehicles, natural language processing and understanding, computer vision, machine reasoning, strong AI, and more. Applications of such technologies may include machine and vehicle object detection, identification, and avoidance; visual recognition, classification, and tagging; algorithmic financial trading strategy performance management; simultaneous localization and mapping; predictive maintenance of high-value machinery; cybersecurity threat prevention; automated expertise; image recognition and classification; question answering; robotics; text analysis (extraction, classification), and text generation and translation. Applications of AI technologies are embodied in a wide range of products, including, for example, Amazon Echo's speech recognition technology, which enables users to talk to their machines; Google Translate™, which enables machine-based language translation; Spotify's Discover Weekly, which provides recommendations for new songs and artists that users may like based on user usage and traffic analysis; Quill's text generation offering, which takes structured data and turns it into narrative stories; and Chatbots, which provide real-time, context-specific answers to questions in a dialogue format.

[0141] Data is at the heart of modern AI and deep learning algorithms. Before training can begin, one issue that must be addressed is collecting labeled data, which is critical for training accurate AI models. Full-scale AI deployments may require continuously collecting, cleaning, transforming, labeling, and storing large amounts of data. Adding additional high-quality data points directly leads to more accurate models and better insights. Data samples may be subjected to a series of processing steps, including, but not limited to: 1) ingesting data from external sources into the training system and storing the data in raw form; 2) cleaning and transforming the data in a format convenient for training, including linking data samples to appropriate labels; 3) iterating to explore parameters and models, rapidly testing them with smaller datasets, and converging on the most promising model to push to a production cluster; 4) running a training phase to select random batches of input data, including both new and older samples, and feeding them to a production GPU server for computation to update model parameters; and 5) evaluation, which involves using a holdback portion of the data not used in training to evaluate model accuracy against holdout data. This lifecycle can be applied to any type of parallelized machine learning, not just neural networks or deep learning. For example, a standard machine learning framework may rely on a CPU instead of a GPU, but the data ingestion and training workflow may be the same. Readers will understand that a single shared storage data hub creates a coordination point across the entire lifecycle without requiring extra data copies between the ingestion, preprocessing, and training stages. Ingested data is rarely used for only one purpose, and shared storage provides the flexibility to train multiple different models or apply traditional analytics to the data.

[0142] The reader will understand that each stage in an AI data pipeline can have different requirements from the data hub (e.g., a storage system or collection of storage systems). A scale-out storage system must provide uncompromising performance for all access types and patterns, from small, metadata-heavy files to large files, from random access patterns to sequential access patterns, and from low concurrency to high concurrency. The storage system described above can serve as an ideal AI data hub because the system can service unstructured workloads. In the first stage, data is ideally ingested and stored on the same data hub used by subsequent stages to avoid excessive data copying. The next two steps can be performed on standard compute servers, optionally including GPUs, and then in the fourth and final stage, the complete training production job is run on powerful GPU-accelerated servers. Often, a production pipeline exists alongside an experimental pipeline running on the same dataset. Furthermore, GPU-accelerated servers can be used independently for different models or can be combined together to train one larger model, even across multiple systems for distributed training. If the shared storage tier is slow, data must be copied to local storage for each phase, resulting in wasted time staging data on different servers. An ideal data hub for an AI training pipeline would provide similar performance to data stored locally on server nodes, while also possessing the simplicity and performance to allow all pipeline stages to operate simultaneously.

[0143] In order for the above-described storage system to function as a data hub or as part of an AI deployment, in some embodiments, the storage system may be configured to provide DMA between storage devices included in the storage system and one or more GPUs used in an AI or big data analytics pipeline. One or more GPUs may be coupled to the storage system via, for example, NVMe-over-Fabric ("NVMe-oF"), bypassing bottlenecks such as the host CPU and allowing the storage system (or one of the components included therein) to directly access the GPU memory. In such an example, the storage system may leverage API hooks to the GPU to transfer data directly to the GPU. For example, the GPU may be embodied as an Nvidia™ GPU, and the storage system may support GPUDirect Storage ("GDS") software or have similar proprietary software that enables the storage system to transfer data to the GPU via RDMA or a similar mechanism.

[0144] While the preceding paragraphs discuss deep learning applications, the reader will understand that the storage systems described herein may also be part of a distributed deep learning ("DDL") platform to support the execution of DDL algorithms. The storage systems described above may also be paired with other technologies, such as TensorFlow, an open-source software library for dataflow programming across a range of tasks that may be used in machine learning applications, such as neural networks, to facilitate the development of such machine learning models, applications, and the like.

[0145] The storage systems described above can also be used in neuromorphic computing environments. Neuromorphic computing is a form of computing that mimics brain cells. To support neuromorphic computing, an architecture of interconnected "neurons" replaces traditional computing models with low-power signals traveling directly between neurons for more efficient computation. Neuromorphic computing can utilize very-large-scale integration (VLSI) systems that include electronic analog circuits to mimic the neurobiological architecture present in the nervous system, as well as analog, digital, and mixed-mode analog / digital VLSI and software systems that implement models of the nervous system for perception, motor control, or multisensory integration.

[0146] The reader will understand that the storage systems described above may be configured to support the storage or use of blockchains and derived items (among other types of data), such as, for example, open source blockchains and related tools that are part of the IBM™ Hyperledger project, permissioned blockchains in which a certain number of trusted parties are permitted to access the blockchain, blockchain products that allow developers to build their own distributed ledger projects, and others. The blockchains and storage systems described herein may be utilized to support on-chain storage of data as well as off-chain storage of data.

[0147] Off-chain storage of data can be implemented in various ways and can occur when the data itself is not stored within the blockchain. For example, in one embodiment, a hash function can be utilized, and the data itself can be fed into the hash function to generate a hash value. In such an example, a hash of a large piece of data may be embedded within a transaction instead of the data itself. The reader will understand that in other embodiments, alternatives to blockchain can be used to facilitate decentralized storage of information. For example, one alternative to blockchain that can be used is blockweave. While traditional blockchains store every transaction to achieve validation, blockweave allows for secure decentralization without using the entire chain, thereby enabling low-cost on-chain storage of data. Such blockweaves can utilize consensus mechanisms based on proof of access (PoA) and proof of work (PoW).

[0148] The storage systems described above, alone or in combination with other computing devices, can be used to support in-memory computing applications. In-memory computing involves the storage of information in RAM distributed across a cluster of computers. The reader will understand that the storage systems described above, particularly those configurable with customizable amounts of processing, storage, and memory resources (e.g., systems in which blades include configurable amounts of each type of resource), can be configured to provide an infrastructure capable of supporting in-memory computing. Similarly, the storage systems described above can include component parts (e.g., NVDIMMs, 3D cross-point storage providing persistent, high-speed random-access memory) that can actually provide an improved in-memory computing environment compared to an in-memory computing environment that relies on RAM distributed across dedicated servers.

[0149] In some embodiments, the storage systems described above can be configured to operate as hybrid in-memory computing environments that include a universal interface to all storage media (e.g., RAM, flash storage, 3D cross-point storage). In such embodiments, users may not have knowledge of the details of where their data is stored, but can still address the data using the same complete, unified API. In such embodiments, the storage system can (in the background) move data to the fastest available tier, including intelligently placing data according to various characteristics of the data or some other heuristic. In such examples, the storage system can even utilize existing products such as Apache Ignite and GridGain to move data between various storage tiers, or the storage system can utilize custom software to move data between various storage tiers. The storage systems described herein can implement various optimizations to improve the performance of in-memory computing, such as, for example, having computations occur as close to the data as possible.

[0150] The reader will further understand that in some embodiments, the storage systems described above can be paired with other resources to support the applications described above. For example, one infrastructure may include primary computing in the form of servers and workstations specialized in using general-purpose computing on graphics processing units (GPGPUs) to accelerate deep learning applications, interconnected to a compute engine to train parameters for deep neural networks. Each system may have Ethernet external connectivity, InfiniBand external connectivity, some other form of external connectivity, or some combination thereof. In such examples, GPUs may be grouped for a single large-scale training run or used independently to train multiple models. The infrastructure may also include storage systems such as those described above to provide a scale-out all-flash file or object store where data can be accessed via high-performance protocols such as NFS, S3, etc. The infrastructure may also include redundant top-of-rack Ethernet switches connected to the storage and computers via ports in MLAG port channels for redundancy, for example. The infrastructure may also include additional computers in the form of white-box servers, optionally with GPUs, for data ingestion, preprocessing, and model debugging. The reader will appreciate that additional infrastructure is possible.

[0151] The reader will understand that the storage system described above can be configured to support other AI-related tools, either alone or in cooperation with other computing machines. For example, the storage system can utilize tools such as ONXX or other open neural network exchange formats, which make it easier to transfer models written in different AI frameworks. Similarly, the storage system can be configured to support tools such as Amazon's Gluon, which allows developers to prototype, build, and train deep learning models. In fact, the storage system described above can be part of a larger platform, such as IBM™ Cloud Private for Data, which includes integrated data science, data engineering, and application building services.

[0152] The reader will further understand that the storage systems described above can also be deployed as edge solutions. Such edge solutions may be suitable for optimizing cloud computing systems by performing data processing at the edge of the network, close to the source of the data. Edge computing can push applications, data, and computing power (i.e., services) from centralized points to the logical extremes of the network. Through the use of edge solutions such as the storage systems described above, computational tasks can be performed using the computational resources provided by such storage systems, data can be stored using the storage resources of the storage systems, and cloud-based services can be accessed through the use of various resources (including networking resources) of the storage systems. By performing computational tasks on edge solutions, storing data on edge solutions, and utilizing edge solutions in general, consumption of expensive cloud-based resources can be avoided and, in fact, performance improvements can be experienced relative to a heavier reliance on cloud-based resources.

[0153] While many tasks can benefit from the use of edge solutions, some specific uses may be particularly suited to deployment in such environments. For example, drones, autonomous vehicles, robots, and other devices may require very fast processing; in practice, transmitting data up to a cloud environment and back to receive data processing support may simply be too slow. As an additional example, some IoT devices, such as connected video cameras, may not be well suited to utilizing cloud-based resources simply because the sheer volume of data involved may make transmitting data to the cloud impractical (not just from a privacy, security, or financial perspective). Thus, many tasks involving data processing, storage, or communication may indeed be better suited by platforms that include edge solutions, such as the storage systems described above.

[0154] The storage system described above, alone or in combination with other computing resources, can function as a network edge platform, combining compute resources, storage resources, networking resources, cloud technologies, and network virtualization technologies. As part of the network, the edge can take on similar characteristics as other network facilities, from customer premises and backhaul aggregation facilities to points of presence (PoPs) and regional data centers. Readers will understand that network workloads such as virtual network functions (VNFs) and others reside on the network edge platform. Network edge platforms enabled by the combination of containers and virtual machines may rely on controllers and schedulers that are no longer geographically co-located with data processing resources. Microservice-based functions can be split into control planes, user and data planes, or even state machines, allowing independent optimization and scaling techniques to be applied. Such user and data planes can be enabled both through the increasing number of accelerators present in server platforms, such as FPGAs and smart NICs, and through SDN-enabled merchant silicon and programmable ASICs.

[0155] The above-described storage systems can also be optimized for use in big data analytics, including leveraging containerized analytics architectures, for example, as part of a configurable data analytics pipeline, making analytics capabilities more configurable. Big data analytics can generally be described as the process of examining large and diverse data sets to uncover hidden patterns, unknown correlations, market trends, customer preferences, and other useful information that can help organizations make more informed business decisions. As part of that process, semi-structured and unstructured data, such as internet clickstream data, web server logs, social media content, text from customer emails and survey responses, mobile phone call detail records, IoT sensor data, and other data, may be converted into a structured form.

[0156] The storage system described above may also support (including implementing as a system interface) applications that perform tasks in response to human speech. For example, the storage system may support the execution of intelligent personal assistant applications such as Amazon's Alexa™, Apple Siri™, Google Voice™, Samsung Bixby™, Microsoft Cortana™, and others. While the examples described in the previous sentence utilize voice as input, the storage system described above may also support chatbots, talkbots, chatterbots, or artificial conversational entities, or other applications configured to conduct conversations via auditory or textual methods. Similarly, the storage system may actually execute such applications to enable users, such as system administrators, to interact with the storage system via voice. While such applications generally enable voice interaction, music playback, creating to-do lists, setting alarms, streaming podcasts, playing audiobooks, and providing other real-time information such as weather, traffic, and news, in embodiments according to the present disclosure, such applications may also be used as interfaces to various system management operations.

[0157] The storage systems described above can also implement AI platforms to deliver on the vision of self-driving storage. Such AI platforms can be configured to provide global predictive intelligence by collecting and analyzing large volumes of storage system telemetry data points, enabling easy management, analysis, and support. Indeed, such storage systems may be capable of predicting both capacity and performance and generating intelligent advice regarding workload deployment, interaction, and optimization. Such AI platforms can be configured to scan all incoming storage system telemetry data against a library of problem fingerprints, capturing hundreds of performance-related variables used to predict performance loads, in order to predict and resolve incidents in real time before they impact customer environments.

[0158] The storage systems described above can support the serialization or concurrent execution of artificial intelligence applications, machine learning applications, data analysis applications, data transformations, and other tasks that may collectively form an AI ladder. Such an AI ladder may be effectively formed by combining such elements to form a complete data science pipeline in which dependencies exist between the elements of the AI ​​ladder. For example, AI may require that some form of machine learning has taken place, which may require that some form of analytics has taken place, which may require that some form of data and information construction has taken place, etc. Thus, each element can be considered a rung in the AI ​​ladder that can collectively form a complete and sophisticated AI solution.

[0159] The above-described storage systems can also be used, alone or in combination with other computing environments, to deliver AI to any experience where AI permeates broad and expansive aspects of business and life. For example, AI may play a key role in the delivery of deep learning solutions, deep reinforcement learning solutions, artificial general intelligence solutions, autonomous vehicles, cognitive computing solutions, commercial UAVs or drones, conversational user interfaces, enterprise taxonomies, ontology management solutions, machine learning solutions, smart dust, smart robots, smart workplaces, etc.

[0160] The above-described storage systems may also be used, alone or in combination with other computing environments, to provide a wide range of transparent and immersive experiences (including those using digital twins of various "things," such as people, places, processes, systems, etc.) where technology can introduce transparency between people, businesses, and things. Such transparent and immersive experiences may be provided as augmented reality technology, connected homes, virtual reality technology, brain-computer interfaces, human augmentation technology, nanotube electronics, volumetric displays, 4D printing technology, or others.

[0161] The above-described storage systems may also be used, alone or in combination with other computing environments, to support a wide variety of digital platforms, including, for example, 5G wireless systems and platforms, digital twin platforms, edge computing platforms, IoT platforms, quantum computing platforms, serverless PaaS, software-defined security, neuromorphic computing platforms, and the like.

[0162] The storage system described above may also be part of a multi-cloud environment in which multiple cloud computing and storage services are deployed in a single heterogeneous architecture. To facilitate the operation of such a multi-cloud environment, DevOps tools may be deployed to enable orchestration across clouds. Similarly, continuous development and continuous integration tools may be deployed to standardize processes for continuous integration and delivery, new feature rollout, and cloud workload provisioning. By standardizing these processes, a multi-cloud strategy may be implemented that enables the utilization of the best provider for each workload.

[0163] The storage systems described above can be used as part of a platform that enables the use of cryptographic anchors that can be used to authenticate the origin and content of a product and ensure that it matches the blockchain record associated with the product. Similarly, as part of a suite of tools for securing data stored on the storage system, the storage systems described above can implement various encryption techniques and schemes, including lattice cryptography. Lattice cryptography can involve the construction of cryptographic primitives that include a lattice in either the construction itself or in security proofs. Unlike public-key schemes such as RSA, Diffie-Hellman, or Elliptic-Curve cryptosystems, which are easily attacked by quantum computers, some lattice-based configurations are believed to be resistant to attacks by both classical and quantum computers.

[0164] A quantum computer is a device that performs quantum computing. Quantum computing is calculation using quantum mechanical phenomena such as superposition and entanglement. Quantum computers differ from conventional computers based on transistors because such conventional computers require data to be encoded into binary digits (bits), each of which is always in one of two distinct states (0 or 1). In contrast to conventional computers, quantum computers use qubits, which can be in superposition of states. Quantum computers maintain a sequence of qubits, and a single qubit can represent 1, 0, or any quantum superposition of the two qubit states. A pair of qubits can be in any quantum superposition of four states, and three qubits can be in any superposition of eight states. A quantum computer with n qubits can generally be in any superposition of up to 2^n different states simultaneously, while a conventional computer can only be in one of these states at any one time. A quantum Turing machine is a theoretical model of such a computer.

[0165] The storage systems described above can be paired with FPGA acceleration servers as part of a larger AI or ML infrastructure. Such FPGA acceleration servers may reside near the storage systems described above (e.g., in the same data center) or may be incorporated into an appliance that includes one or more storage systems, one or more FPGA acceleration servers, a networking infrastructure supporting communication between the one or more storage systems and the one or more FPGA acceleration servers, and other hardware and software components. Alternatively, the FPGA acceleration servers can reside in a cloud computing environment that can be used to perform computation-related tasks for AI and ML jobs. Any of the above-described embodiments can collectively function as an FPGA-based AI or ML platform. The reader will understand that in some embodiments of an FPGA-based AI or ML platform, the FPGA included in the FPGA acceleration server can be reconfigured for different types of ML models (e.g., LSTM, CNN, GRU). The ability to reconfigure the FPGA included in the FPGA acceleration server can enable the acceleration of ML or AI applications based on the most optimal numerical precision and memory model being used. The reader will understand that by treating a collection of FPGA-accelerated servers as a pool of FPGAs, any CPU in the data center can utilize the pool of FPGAs as a shared hardware microservice, rather than limiting the server to the dedicated accelerators plugged into it.

[0166] The FPGA-accelerated servers and GPU-accelerated servers described above can implement a model of computing in which CPU models and parameters are pinned to high-bandwidth on-chip memory and much data is streamed through the high-bandwidth on-chip memory, rather than holding a small amount of data in machine learning and executing a long stream of instructions on it, as is done in traditional computing models. Because FPGAs can be programmed with only the instructions necessary to execute this type of computing model, FPGAs can be even more efficient than GPUs for this type of computing model.

[0167] The storage systems described above can be configured to provide parallel storage, for example, through the use of a parallel file system such as BeeGFS. Such a parallel file system can include a distributed metadata architecture. For example, a parallel file system may include components including multiple metadata servers across which metadata is distributed, and services for clients and storage servers.

[0168] The above-described system can support the execution of a variety of software applications. Such software applications can be deployed in various ways, including a container-based deployment model. Containerized applications can be managed using various tools. For example, containerized applications may be managed using Docker Swarm, Kubernetes, and others. Containerized applications can be used to facilitate a serverless, cloud-native computing deployment and management model for software applications. In support of a serverless, cloud-native computing deployment and management model for software applications, containers can be used as part of an event handling mechanism (e.g., AWS Lambda) such that various events spin up containerized applications to act as event handlers.

[0169] The above-described systems may be deployed in various ways, including in a manner that supports fifth-generation ("5G") networks. 5G networks may support substantially faster data communications than prior-generation mobile communications networks, potentially leading to the decentralization of data and computing resources, as modern large-scale data centers may become less prominent and may be replaced by more local micro-data centers, for example, closer to mobile network towers. The above-described systems may be included in such local micro-data centers or may be part of or paired with multi-access edge computing ("MEC") systems. Such MEC systems may enable cloud computing capabilities and IT service environments at the edge of cellular networks. By running applications and performing associated processing tasks closer to cellular customers, network congestion may be reduced and applications may perform better.

[0170] The storage system described above can be configured to implement NVMe Zoned Namespaces. Through the use of NVMe Zoned Namespaces, the logical address space of the namespace is divided into zones. Each zone provides a logical block address range that is written sequentially and must be explicitly reset before being rewritten, thereby enabling the creation of a namespace that exposes the natural boundaries of the device and offloading management of internal mapping tables to the host. To implement NVMe Zoned Namespaces ("Zoned Namespaces"), ZNS SSDs or other forms of zoned block devices that expose the namespace logical address space using zones can be utilized. When zones are aligned to the internal physical characteristics of the device, multiple inefficiencies in data placement can be eliminated. In such embodiments, each zone can be mapped to a separate application, such that functions such as wear leveling and garbage collection can be performed per zone or per application, rather than across the entire device. To support ZNS, the storage controllers described herein can be configured to interact with zoned block devices, for example, through the use of the Linux™ kernel zoned block device interface or other tools.

[0171] The storage systems described above may also be configured to implement zoned storage in other ways, such as through the use of shingled magnetic recording (SMR) storage devices. In instances where zoned storage is used, a device-managed embodiment may be deployed, where the storage device hides this complexity by managing it in firmware and presenting an interface like any other storage device. Alternatively, zoned storage may be implemented through a host-managed embodiment that relies on the operating system to know how to handle the drive and writes sequentially only to specific areas of the drive. Zoned storage may similarly be implemented using a host-aware embodiment, where a combination of drive-managed and host-managed implementations is deployed.

[0172] The storage systems described herein can be used to form a data lake. A data lake can act as the first place an organization's data flows, and such data can be in a raw format. Metadata tagging can be implemented to facilitate searching of data elements within the data lake, especially in embodiments where the data lake includes multiple stores of data in formats that are not easily accessible or readable (e.g., unstructured data, semi-structured data, structured data). From the data lake, data can proceed downstream to a data warehouse, where the data can be more processed, packaged, and stored in a consumable format. The storage systems described above can also be used to implement such a data warehouse. In addition, a data mart or data hub can enable even more easily consumed data, and the storage systems described above can be used to provide the underlying storage resources required for the data mart or data hub. In embodiments, queries to the data lake may require a schema-on-read approach, in which the plan or schema is applied as the data is pulled from the stored location, rather than as the plan or schema is entered.

[0173] Storage systems described herein may also be configured to implement a recovery point objective ("RPO"), which may be established by a user, an administrator, a system default, as part of a storage class or service that the storage system participates in delivering, or in some other manner. A "recovery point objective" is a target for the maximum time difference between the last update to a source dataset and the last recoverable replicated dataset update that is correctly recoverable, given a reason to do so, from a continuously or frequently updated copy of the source dataset. An update is correctly recoverable if it properly takes into account all updates processed to the source dataset prior to the last recoverable replicated dataset update.

[0174] In synchronous replication, the RPO is zero, which means that under normal operation, all completed updates on the source data set should be present and correctly recoverable on the copy data set. In best-effort near-synchronous replication, the RPO can be as low as a few seconds. In snapshot-based replication, the RPO can be roughly calculated as the interval between snapshots plus the time to transfer modifications between the previous, already transferred snapshot and the latest replicated snapshot.

[0175] If updates accumulate faster than they can be replicated, the RPO can be missed. In the case of snapshot-based replication, the RPO can be missed if more replicated data accumulates between two snapshots than can be replicated between taking a snapshot and replicating that snapshot's cumulative updates to the copy. Again, in snapshot-based replication, if the data being replicated accumulates at a rate faster than it can be transferred in the time between subsequent snapshots, replication can begin to fall further behind, widening the gap between the expected recovery point objective and the actual recovery point represented by the last successfully replicated update.

[0176] The storage systems described above may be part of a shared-nothing storage cluster. In a shared-nothing storage cluster, each node of the cluster has local storage and communicates with other nodes in the cluster over a network, and the storage used by the cluster is (generally) provided solely by the storage connected to each individual node. A collection of nodes synchronously replicating a data set may be an example of a local storage cluster, since each storage system has shared-nothing and communicates with other storage systems over a network; these storage systems (generally) do not use other storage to which they share access through some interconnect. In contrast, some of the storage systems described above are themselves configured as shared storage clusters, since there are drive shelves shared by paired controllers. However, other storage systems described above are configured as shared-nothing storage clusters, since all storage is local to a particular node (e.g., blade) and all communication is via the network linking the compute nodes to each other.

[0177] In other embodiments, other forms of shared-nothing storage clusters may include embodiments in which any node in the cluster has a local copy of all the storage it needs, and the data is mirrored to other nodes in the cluster via synchronous replication to ensure that data is not lost or because other nodes are also using that storage. In such an embodiment, if a new cluster node needs some data, it can be copied to the new node from other nodes that have copies of that data.

[0178] In some embodiments, a mirror copy-based shared storage cluster may store multiple copies of all the cluster's stored data, with each subset of the data replicated to a particular set of nodes and different subsets of the data replicated to different sets of nodes. In some variations, embodiments may store all of the cluster's stored data on all nodes, while in other variations, the nodes may be divided so that a first set of nodes all store the same set of data and a second, different set of nodes all store a different set of data.

[0179] The reader will understand that a RAFT-based database (e.g., etcd) can operate like a shared-nothing cluster, with all RAFT nodes storing all data. However, the amount of data stored in a RAFT cluster can be limited so that redundant copies do not consume too much storage. A container server cluster may also be capable of replicating all data to all cluster nodes, assuming containers do not tend to be too large and their bulk data (data manipulated by applications running within containers) is stored elsewhere, such as an S3 cluster or external file server. In such an example, container storage may be provided by the cluster directly through its shared-nothing storage model, and those containers provide the images that form the execution environment for portions of an application or service.

[0180] For further explanation, FIG. 3D illustrates an exemplary computing device 350 that may be particularly configured to perform one or more of the processes described herein. As shown in FIG. 3D, computing device 350 may include a communication interface 352, a processor 354, a storage device 356, and an input / output (“I / O”) module 358, communicatively coupled to each other via a communication infrastructure 360. While an exemplary computing device 350 is shown in FIG. 3D, the components illustrated in FIG. 3D are not intended to be limiting. In other embodiments, additional or alternative components may be used. The components of computing device 350 shown in FIG. 3D will now be described in further detail.

[0181] The communication interface 352 may be configured to communicate with one or more computing devices. Examples of the communication interface 352 include, but are not limited to, a wired network interface connection (such as a network interface card), a wireless network interface connection (such as a wireless network interface card connection), a modem, an audio / video connection, and any other suitable interface.

[0182] The processor 354 generally represents any type or form of processing unit capable of processing data and / or interpreting, executing, and / or directing the execution of one or more of the instructions, processes, and / or operations described herein. The processor 354 may perform operations by executing computer-executable instructions 362 (e.g., applications, software, code, and / or other executable data instances) stored on the storage device 356.

[0183] Storage device 356 may include one or more data storage media, devices, or configurations, and any type, form, and combination of data storage media and / or devices may be used. For example, storage device 356 may include, but is not limited to, any combination of non-volatile and / or volatile media described herein. Electronic data, including the data described herein, may be stored temporarily and / or persistently in storage device 356. For example, data representing computer-executable instructions 362 configured to instruct processor 354 to perform any of the operations described herein may be stored in storage device 356. In some examples, data may be located in one or more databases residing in storage device 356.

[0184] I / O module 358 may include one or more I / O modules configured to receive user input and provide user output. I / O module 358 may include any hardware, firmware, software, or combination thereof that supports input and output capabilities. For example, I / O module 358 may include hardware and / or software for capturing user input, including, but not limited to, a keyboard or keypad, a touchscreen component (e.g., a touchscreen display), a receiver (e.g., an RF or infrared receiver), a motion sensor, and / or one or more input buttons.

[0185] I / O module 358 may include one or more devices for presenting output to a user, including, but not limited to, a graphics engine, a display (e.g., a display screen), one or more output drivers (e.g., a display driver), one or more audio speakers, and one or more audio drivers. In particular embodiments, I / O module 358 is configured to provide graphical data to a display for presentation to a user. The graphical data may represent one or more graphical user interfaces and / or any other graphical content as may be useful in a particular implementation. In some examples, any of the systems, computing devices, and / or other components described herein may be implemented by computing device 350.

[0186] For further explanation, FIG. 3E illustrates an example of a fleet of storage systems 376 for providing storage services (also referred to herein as “data services”). The fleet of storage systems 376 depicted in FIG. 3 includes multiple storage systems 374a, 374b, 374c, 374d, and 374n, each of which may be similar to the storage systems described herein. The storage systems 374a, 374b, 374c, 374d, and 374n in the fleet of storage systems 376 may be embodied as the same storage system or as different types of storage systems. For example, two of the storage systems 374a, 374n depicted in FIG. 3E are depicted as cloud-based storage systems because the resources collectively forming each of the storage systems 374a, 374n are provided by separate cloud service providers 370, 372. For example, first cloud service provider 370 may be Amazon AWS™ while second cloud service provider 372 is Microsoft Azure™, although in other embodiments, one or more public clouds, private clouds, or a combination thereof may be used to provide the underlying resources used to form a particular storage system within fleet of storage systems 376.

[0187] 3E includes an edge management service 382 for delivering storage services, according to some embodiments of the present disclosure. The delivered storage services (also referred to herein as "data services") may include, for example, services that provide a certain amount of storage to consumers, services that provide storage to consumers subject to certain service level agreements, services that provide storage to consumers subject to certain regulatory requirements, and many other services.

[0188] 3E may be embodied as one or more modules of computer program instructions executing on computer hardware, such as one or more computer processors. Alternatively, edge management service 382 may be embodied as one or more modules of computer program instructions executing on a virtualized execution environment, such as one or more virtual machines, within one or more containers, or in some other manner. In other embodiments, edge management service 382 may be embodied as a combination of the above-described embodiments, including embodiments in which one or more modules of computer program instructions included in edge management service 382 are distributed across multiple physical or virtual execution environments.

[0189] The edge management service 382 may act as a gateway to provide storage services to storage consumers, where the storage services leverage storage provided by one or more storage systems 374a, 374b, 374c, 374d, 374n. For example, the edge management service 382 may be configured to provide storage services to host devices 378a, 378b, 378c, 378d, 378n running one or more applications that consume the storage services. In such an example, the edge management service 382 can act as a gateway between the host devices 378a, 378b, 378c, 378d, 378n and the storage systems 374a, 374b, 374c, 374d, 374n, rather than requiring the host devices 378a, 378b, 378c, 378d, 378n to directly access the storage systems 374a, 374b, 374c, 374d, 374n.

[0190] While the edge management service 382 of FIG. 3E exposes the storage services module 380 to the host devices 378a, 378b, 378c, 378d, and 378n of FIG. 3E, in other embodiments, the edge management service 382 may expose the storage services module 380 to other consumers of various storage services. The various storage services may be presented to the consumer via one or more user interfaces, via one or more APIs, or through some other mechanism provided by the storage services module 380. Thus, the storage services module 380 depicted in FIG. 3E may be embodied as one or more modules of computer program instructions executing on physical hardware, on a virtualized execution environment, or a combination thereof, where execution of such modules enables consumers of storage services to be offered, select from, and access various storage services.

[0191] The edge management services 382 of Figure 3E also includes a system management services module 384. The system management services module 384 of Figure 3E includes one or more modules of computer program instructions that, when executed, perform various operations in coordination with the storage systems 374a, 374b, 374c, 374d, 374n to provide storage services to the host devices 378a, 378b, 378c, 378d, 378n. The system management services module 384 may be configured to perform tasks such as, for example, provisioning storage resources from the storage systems 374a, 374b, 374c, 374d, 374n via one or more APIs exposed by the storage systems 374a, 374b, 374c, 374d, 374n, migrating datasets or workloads between the storage systems 374a, 374b, 374c, 374d, 374n via one or more APIs exposed by the storage systems 374a, 374b, 374c, 374d, 374n, and setting one or more tunable parameters (i.e., one or more configurable settings) on the storage systems 374a, 374b, 374c, 374d, 374n via one or more APIs exposed by the storage systems 374a, 374b, 374c, 374d, 374n. For example, many of the services described below relate to embodiments in which storage systems 374a, 374b, 374c, 374d, 374n are configured to operate in some manner. In such examples, system management services module 384 may be responsible for using APIs (or some other mechanism) provided by storage systems 374a, 374b, 374c, 374d, 374n to configure storage systems 374a, 374b, 374c, 374d, 374n to operate in the manner described below.

[0192] In addition to configuring storage systems 374a, 374b, 374c, 374d, 374n, the edge management service 382 itself may be configured to perform various tasks required to provide various storage services. Consider an example where a storage service includes a service that, when selected and applied, obfuscates personally identifiable information ("PII") contained in a dataset when the dataset is accessed. In such an example, storage systems 374a, 374b, 374c, 374d, 374n may be configured to obfuscate PII when servicing read requests directed to the dataset. Alternatively, the storage systems 374a, 374b, 374c, 374d, 374n may service the read by returning data that includes PII, but the edge management service 382 itself may obfuscate the PII as the data passes through the edge management service 382 on its way from the storage systems 374a, 374b, 374c, 374d, 374n to the host devices 378a, 378b, 378c, 378d, 378n.

[0193] The storage systems 374a, 374b, 374c, 374d, and 374n depicted in FIG. 3E may be embodied as one or more of the storage systems (including variations thereof) described above with reference to FIGS. 1A-3D. In practice, the storage systems 374a, 374b, 374c, 374d, and 374n may function as a pool of storage resources, with individual components within the pool having different performance characteristics, different storage characteristics, and the like. For example, one of the storage systems 374a may be a cloud-based storage system, another storage system 374b may be a storage system providing block storage, another storage system 374c may be a storage system providing file storage, another storage system 374d may be a relatively high-performance storage system, while another storage system 374n may be a relatively low-performance storage system, and so on. In alternative embodiments, there may be only a single storage system.

[0194] The storage systems 374a, 374b, 374c, 374d, and 374n depicted in FIG. 3E may also be organized into different failure domains such that a failure of one storage system 374a is completely independent of a failure of another storage system 374b. For example, each of the storage systems may receive power from an independent power system, each of the storage systems may be coupled for data communication via an independent data communication network, etc. Furthermore, storage systems in a first failure domain may be accessed through a first gateway, while storage systems in a second failure domain may be accessed through a second gateway. For example, the first gateway may be a first instance of edge management service 382, ​​and the second gateway may be a second instance of edge management service 382, ​​including embodiments in which each instance is separate or each instance is part of a distributed edge management service 382.

[0195] As an illustrative example of available storage services, a user may be presented with storage services associated with different levels of data protection. For example, a user may be presented with storage services that, when selected and implemented, assure the user that data associated with that user is protected such that various recovery point objectives (“RPOs”) can be guaranteed. A first available storage service, for example, may ensure that a subset of data sets associated with the user are protected such that any data older than five seconds can be recovered in the event of a failure of the primary data store, while a second available storage service may ensure that a subset of data sets associated with the user are protected such that any data older than five minutes can be recovered in the event of a failure of the primary data store.

[0196] Additional examples of storage services that may be presented to a user, selected by the user, and ultimately applied to a dataset associated with the user may include one or more data compliance services. Such data compliance services may be embodied as services that may be provided to a consumer of the data compliance services (i.e., a user) to ensure, for example, that the user's dataset is managed to comply with various regulatory requirements. For example, one or more data compliance services may be provided to a user to ensure that the user's dataset is managed in a manner that complies with the General Data Protection Regulation ("GDPR"); one or more data compliance services may be provided to a user to ensure that the user's dataset is managed in a manner that complies with the Sarbanes-Oxley Act of 2002 ("SOX"); or one or more data compliance services may be provided to a user to ensure that the user's dataset is managed in a manner that complies with some other regulatory act. Additionally, one or more data compliance services may be provided to a user to ensure that their datasets are managed in adherence to some non-governmental guidance (e.g., in adherence to best practices for audit purposes), one or more data compliance services may be provided to a user to ensure that their datasets are managed in adherence to the requirements of a particular client or organization, etc.

[0197] Consider an example in which a specific data compliance service is designed to ensure that a user's data sets are managed in a manner that complies with the requirements set forth in the GDPR. While a complete list of the GDPR's requirements can be found in the regulation itself, for illustrative purposes, an example of a requirement set forth in the GDPR requires that a pseudonymization process must be applied to stored data to transform personal data such that the resulting data cannot be attributed to a specific data subject without the use of additional information. For example, data encryption techniques can be applied to make the original data unintelligible, and such data encryption techniques cannot be reversed without access to the correct decryption key. Thus, the GDPR may require that the decryption key be kept separate from the pseudonymized data. One specific data compliance service may be offered to ensure compliance with the requirements set forth in this paragraph.

[0198] To provide this particular data compliance service, the data compliance service may be presented to the user (e.g., via a GUI) and selected by the user. In response to receiving a selection of a particular data compliance service, one or more storage service policies may be applied to the dataset associated with the user to perform the particular data compliance service. For example, a storage service policy may be applied that requires the dataset to be encrypted before being stored in a storage system, a cloud environment, or elsewhere. To enforce this policy, not only may a requirement be enforced that the dataset be encrypted when stored, but a requirement may also be introduced that requires the dataset to be encrypted before transmitting the dataset (e.g., sending the dataset to another party). In such an example, a storage service policy may also be introduced that requires that any encryption key used to encrypt the dataset not be stored on the same system that stores the dataset itself. The reader will understand that many other forms of data compliance services may be provided and implemented by embodiments of the present disclosure.

[0199] The storage systems 374a, 374b, 374c, 374d, 374n in the fleet of storage systems 376 may be collectively managed, for example, by one or more fleet management modules. The fleet management modules may be part of or separate from the system management services module 384 depicted in FIG. 3E. The fleet management modules may perform tasks such as monitoring the health of each storage system in the fleet, initiating updates or upgrades on one or more storage systems in the fleet, migrating workloads for load balancing or other performance purposes, and many other tasks. Accordingly, and for many other reasons, the storage systems 374a, 374b, 374c, 374d, 374n may be coupled to one another via one or more data communication links to exchange data between the storage systems 374a, 374b, 374c, 374d, 374n.

[0200] The storage systems described herein may support various forms of data replication. For example, two or more of the storage systems may synchronously replicate a dataset between each other. In synchronous replication, separate copies of a particular dataset may be maintained by multiple storage systems, but all accesses (e.g., reads) of the dataset should yield consistent results regardless of which storage system the access is directed to. For example, reads directed to any of the storage systems synchronously replicating the dataset must return identical results. Thus, updates to versions of a dataset need not occur at exactly the same time, but precautions must be taken to ensure consistent access to the dataset. For example, if an update (e.g., a write) directed to a dataset is received by a first storage system, the update may be acknowledged as complete only if all storage systems synchronously replicating the dataset have applied the update to their copies of the dataset. In such examples, synchronous replication may be performed through the use of I / O forwarding (e.g., a write received at a first storage system is forwarded to a second storage system), communication between storage systems (e.g., each storage system indicates it has completed the update), or in other ways.

[0201] In other embodiments, datasets may be replicated through the use of checkpoints. In checkpoint-based replication (also referred to as "near-synchronous replication"), a set of updates to a dataset (e.g., one or more write operations directed to a dataset) may occur between different checkpoints such that the dataset is updated to a particular checkpoint only if all updates to the dataset prior to the particular checkpoint have been completed. Consider an example in which a first storage system stores a live copy of a dataset being accessed by a user of the dataset. In this example, assume that the dataset is being replicated from the first storage system to a second storage system using checkpoint-based replication. For example, the first storage system may send a first checkpoint (time t=0) to the second storage system, then send a first set of updates to the dataset, then send a second checkpoint (time t=1), then send a second set of updates to the dataset, then send a third checkpoint (time t=2). In such an example, if the second storage system has implemented all updates in the first set of updates but has not yet implemented all updates in the second set of updates, the copy of the dataset stored on the second storage system may be up to the second checkpoint. Alternatively, if the second storage system has implemented all updates in both the first set of updates and the second set of updates, the copy of the dataset stored on the second storage system may be up to the third checkpoint. The reader will understand that various types of checkpoints (e.g., metadata-only checkpoints) may be used, and that checkpoints may be distributed based on various factors (e.g., time, number of operations, RPO settings), etc.

[0202] In other embodiments, datasets may be replicated through snapshot-based replication (also referred to as "asynchronous replication"). In snapshot-based replication, a snapshot of a dataset may be sent from a replication source, such as a first storage system, to a replication target, such as a second storage system. In such an embodiment, each snapshot may include the entire dataset or a subset of the dataset, e.g., only the portions of the dataset that have changed since the last snapshot was sent from the replication source to the replication target. The reader will understand that snapshots may be sent on-demand, based on a policy that takes into account various factors (e.g., time, number of operations, RPO settings), or in some other manner.

[0203] The storage systems described above, alone or in combination, can be configured to function as continuous data protection stores. Continuous data protection stores are features of storage systems that record updates to a dataset, allowing a consistent image of the dataset's previous contents to be accessed at a low time granularity (often on the order of seconds or even less), stretching back a reasonable period of time (often hours or days). They allow access to very recent consistent points in time for a dataset, and also allow access to points in time of a dataset immediately prior to an event, such as when a portion of the dataset is corrupted or otherwise lost, while retaining a number of updates close to the maximum number of updates immediately prior to that event. Conceptually, they are like a sequence of snapshots of a dataset taken very frequently and retained over an extended period of time, although continuous data protection stores are often implemented quite differently from snapshots. Storage systems that implement continuous data protection stores can further provide means to access these points in time, to access one or more of these points in time as snapshots or as clone copies, or to revert a dataset to one of these recorded points in time.

[0204] Over time, to reduce overhead, some time points held in the continuous data protection store can be merged with other nearby points in time, essentially removing some of these time points from the store. This can reduce the capacity required to store updates. It may also be possible to convert these limited number of time points into snapshots of longer duration. For example, such a store may keep a low-granularity sequence of time points going back a few hours from the present, and merge or remove some time points to reduce overhead up to additional days. Going further back than that, some of these time points can be converted into snapshots that represent a consistent point-in-time picture from just every few hours.

[0205] For further explanation, FIG. 4 sets forth a flowchart illustrating an exemplary method for a context-driven user interface for a storage system, according to some embodiments of the present disclosure. The exemplary method depicted in FIG. 4 includes receiving 410 a request to access a system interface 404 of the system 400 from a user account 406. The system interface 404 is a software mechanism for presenting visual elements to a user via a user computing device 408 and the user account 404. The system interface 404 is configured and reconfigured by a user interface (UI) engine 402. The UI engine 402 is hardware, software, or a combination of hardware and software that determines important system characteristics and reconfigures the system interface 404 based on the important system characteristics. The UI engine 402 also services requests to the system interface 404 from users of the user account 404. The user account 406 is the user's identity in the storage system 400. The user account 406 is under the user's control and is managed by a security administrator. The user account 406 can interact with and translate commands from a user on a user computing device 408. The system 400 can be any of the storage systems described above. Although the following is described with reference to the storage system 400, the described methods can be implemented using any computing system having a UI engine 402 and a system interface 404.

[0206] Receiving 410 a request to access the system interface 404 of the system 400 from the user account 406 may be performed by detecting that a user (via the user account 406) has logged into the system 400. The request to access the system interface 404 may be the initialization of a user account session, where the user account session refers to the activity between the time the user of the user account 406 logs into the system 400 and the time the user of the user account 406 logs out of the system 400.

[0207] 4 includes identifying 412 at least one significant system characteristic that describes a current aspect of the system 400. The system characteristic is information about the state or context of the storage system 400. The system characteristic may include metrics about the storage system (e.g., uptime level, RPO, etc.), as well as information about the storage system itself (e.g., software version, hardware description, etc.). The system characteristic may also incorporate information retrieved from systems other than the storage system 400. For example, the system characteristic may include comparative information such as a metric percentile within similar deployments (e.g., bottom 25% of reliability compared to similar companies).

[0208] A critical system characteristic is a system characteristic identified as being important enough to reconfigure the system interface 404 based on the system characteristic. Identifying 412 at least one critical system characteristic can be performed by evaluating system characteristics to determine the presence of one or more critical system characteristics. A system characteristic may be identified as being important based on the system characteristic indicating that the system is operating outside of ideal operating parameters. For example, a system characteristic may be identified as being important if a metric associated with the system characteristic exceeds a threshold tolerance (e.g., a decreased uptime level, an increased RPO, etc.). In a storage system, such a critical system characteristic may be the storage utilization level of the storage system. A system characteristic may be identified as being important if the storage utilization level exceeds a particular threshold (i.e., is over-utilized or under-utilized). The storage utilization level may include the space available for storage and / or the resources available to service storage requests.

[0209] As another example, a system characteristic may be identified as important based on the system characteristic being indicative of a change to the system (e.g., a user needing access to a resource, a feature added to the system, etc.) In a storage system, such an important system characteristic may be the detection of additional resources added to or released from within the storage system and the requirement that such resources be allocated.

[0210] As another example, a system characteristic may be identified as critical because it relates to external systems. Specifically, a system characteristic may be identified as critical if conditions external to the system 400 change such that the system characteristic must be addressed (e.g., a vulnerability is discovered in the system software, an update is available for the system software, government compliance requirements have changed, the system no longer operates according to industry standards, etc.). Such a critical system characteristic in a storage system may be the level or type of encryption used compared to industry benchmarks. If the encryption used by the storage system is not keeping up with the industry benchmarks, the encryption level or type may be identified as a critical system characteristic.

[0211] As another example, a system characteristic may be identified as critical if the system characteristic indicates imminent damage to the system or a user (e.g., data corruption, system elements offline, etc.). Such a critical system characteristic in a storage system may be the presence of a ransomware attack. If a ransomware attack is detected, system security and the attack may be identified as a critical system characteristic. System characteristics such as those described above may be explicitly identified as critical by a user of user account 406 or by an administrator of storage system 400. Alternatively, a system characteristic may be identified as critical based on an automatic trigger, such as a metric exceeding a threshold or detecting a malware attack based on a sequence of requests.

[0212] A request for access to the system interface 404 may itself identify a critical system characteristic. Specifically, the request may include an indication that the user desires the system interface 404 to be reconfigured based on a particular critical system characteristic or group of critical system characteristics. For example, a request for access to the system interface 404 may identify a software upgrade that the user plans to install during the user account session. The software upgrade is then identified as a critical system characteristic.

[0213] Identifying 412 the at least one critical system characteristic may also be performed by tracking a system characteristic or a subset of system characteristics and determining an importance score for each tracked system characteristic over time. The importance score is a value indicating the severity of the state of the system characteristic relative to the states of other system characteristics. The importance score may be based, for example, on the urgency or importance of the system characteristic. The importance score may also be based on the number of (potential) affected users. The at least one critical system characteristic may be identified based on the applied importance scores (e.g., the system characteristic with the highest importance score is identified as the at least one critical system characteristic). Similarly, the UI engine 402 may identify a hierarchy of critical system characteristics, such as a primary group of critical system characteristics, a secondary group of critical system characteristics, etc. The primary and secondary nature of the identified critical system characteristics may be reflected in the reconfigured system interface. Identifying 412 the at least one critical system characteristic may also be performed by determining one or more critical system characteristics that are a priority for resolution. In such a determination, the importance score for each important system characteristic may be based on system priorities (e.g., as specified by a user or administrator) to select one or more important system characteristics to prioritize.

[0214] 4 also includes reconfiguring 414 the system interface 404 based on at least one important system characteristic. Reconfiguring 414 the system interface 404 based on at least one important system characteristic may be performed by adding, removing, and / or rearranging visual elements within the system interface 404. The system interface 404 includes several visual elements, each representing a system characteristic, and interaction elements used to perform tasks related to one or more system characteristics. The visual elements may be grouped together to address a particular important system characteristic or group of important system characteristics.

[0215] Reconfiguring 414 the system interface 404 based on the at least one important system characteristic may include rendering within the system interface 404 a group of visual elements used to perform tasks related to the identified at least one important system characteristic. The reconfigured system interface 404 may exclusively present the group of visual elements used to perform tasks related to the identified at least one important system characteristic. Alternatively, the reconfigured system interface 404 may include groups of visual elements associated with different system characteristics, with the identified at least one important system characteristic being presented more prominently than other visual elements. The system interface 404 is “reconfigured” in that the system interface 404 is modified from a default configuration (based on a system without the identified important system characteristic).

[0216] A group of visual elements associated with the same system characteristic or type of system characteristic may share similar visual characteristics regardless of the reconfigured and identified important system characteristics. Such visual characteristics may include, for example, the same color, the same general interface location (e.g., anchored near the upper left corner), or the same brightness or shade. For example, a group of visual elements may each be rendered in a different hue of blue, and another group may each be rendered in a different hue of red, regardless of the identified important system characteristics for the current user account session.

[0217] Reconfiguring 414 the system interface 404 based on the identified important system characteristics may be performed during the course of a user account session. Specifically, the UI engine 402 may identify new important system characteristics in response to changes to the system 400. Based on the changes to the system 400, the UI engine 402 may reconfigure the system interface 404 to present visual elements associated with the identified important system characteristics.

[0218] 4 includes presenting 416 the reconfigured system interface 404 to a user of the user account 406. Presenting 416 the reconfigured system interface 404 to a user of the user account 406 may be performed by modifying permissions associated with the user account 406 and / or the reconfigured system interface 404 to allow the user account 406 to access the reconfigured system interface 404. Presenting the reconfigured system interface 404 to the user of the user account 406 may also be performed by providing the user account 406 with a path to the reconfigured system interface 404 that can be used to retrieve code representing the reconfigured system interface 404.

[0219] For further explanation, Figure 5 sets forth a flowchart illustrating an additional exemplary method of a context-driven user interface for a storage system, according to some embodiments of the present disclosure. The exemplary method depicted in Figure 5 is similar to the exemplary method depicted in Figure 4, as it includes receiving 410 a request from a user account 406 to access a system interface 404 of a system 400, identifying 412 at least one important system characteristic that describes a current aspect of the system 400, reconfiguring 414 the system interface 404 based on the at least one important system characteristic, and presenting 416 the reconfigured system interface 404 to a user of the user account 406.

[0220] 5, reconfiguring 414 the system interface 404 based on at least one critical system characteristic includes identifying 502 a user account personality for the user account session. As used herein, a user account personality refers to one of multiple roles that a single user account can perform with respect to a system (e.g., a storage system). For example, a user account may be associated with a personality primarily concerned with security and protection of the storage system (i.e., a security personality), a personality primarily concerned with resolving errors in the system (i.e., a troubleshooting personality), a personality primarily concerned with procuring resources for the storage system (i.e., a procurement personality), a personality primarily concerned with deploying resources in the storage system (i.e., a deployment personality), and a personality primarily concerned with government compliance (i.e., a compliance personality).

[0221] Each user account personality may be associated with a collection of visual elements that represent system characteristics and interaction elements used to perform tasks associated with the personality. Visual elements associated with a personality represent system characteristics and interaction elements that support activities performed under that personality. Visual elements associated with the same personality may be grouped together within the system interface. For example, one visual element group may be associated with a security personality and include elements such as a menu for authorizing or deauthorizing storage client accounts, system characteristics related to data utilization, and system characteristics related to snapshots created from existing data. Another visual element group may be associated with a procurement personality and include elements such as suggestions for devices or services to add to the storage system based on system characteristics. Another visual element group may be associated with a troubleshooting personality and include elements such as system characteristics describing current errors present in the system and proposed actions to address the errors. Another visual element group may be associated with a deployment personality and include elements such as procedural instructions for adding resources or services to the storage system and the status of the deployment process. Finally, another group of visual elements may be associated with a compliance personality and include elements such as rules to be followed and details of each rule.

[0222] Identifying 502 a user account personality for a user account session may be performed by determining a user account personality for the identified critical system characteristics. Each critical system characteristic may be associated with a particular user account personality through which the critical system characteristic may be addressed. For example, a critical system characteristic that indicates imminent damage to the system or user may be associated with a security personality. As another example, a critical system characteristic that indicates the system is operating outside ideal operating parameters may be associated with a troubleshooting personality. Once the user account personality is identified, the system interface 404 may be reconfigured by presenting within the system interface visual elements associated with the identified user account personality.

[0223] For further explanation, Figure 6 sets forth a flowchart illustrating an additional exemplary method of a context-driven user interface for a storage system, according to some embodiments of the present disclosure. The exemplary method depicted in Figure 6 is similar to the exemplary method depicted in Figure 4, as it includes receiving 410 a request from a user account 406 to access a system interface 404 of a system 400, identifying 412 at least one important system characteristic that describes a current aspect of the system 400, reconfiguring 414 the system interface 404 based on the at least one important system characteristic, and presenting 416 the reconfigured system interface 404 to a user of the user account 406.

[0224] 6, reconfiguring 414 the system interface 404 based on the at least one important system characteristic includes arranging 602 visual elements within the system interface 402 such that the visual element associated with the at least one important system characteristic is the primary element of the system interface 404. Arranging 602 visual elements within the system interface 402 such that the visual element associated with the at least one important system characteristic is the primary element of the system interface 404 may be performed by applying a visual style to the visual elements associated with the identified important system characteristic to distinguish these visual elements from other groups of visual elements associated with other system characteristics.

[0225] "Primary" refers to a visual element or group of visual elements that are visually emphasized relative to other visual elements in system interface 404. Making a visual element associated with an identified important system characteristic a primary element of system interface 404 can be implemented in various ways. For example, the visual element associated with an identified important system characteristic may be centered in system interface 404. As another example, the visual element associated with an identified important system characteristic may be rendered larger than other groups of visual elements associated with other system characteristics. As another example, the visual element associated with an identified important system characteristic may be rendered in a lighter hue than other groups of visual elements associated with other system characteristics. As another example, the visual element associated with an identified important system characteristic may be presented on top of a layer of windows in system interface 404.

[0226] Reconfiguring 414 the system interface 404 based on the identified important system characteristics may include arranging a group of visual elements associated with any identified secondary important system characteristics such that the visual elements associated with the identified secondary important system characteristics become secondary elements of the system interface 404. Similar to the description above, visual styles may be applied to the visual elements associated with the identified secondary important system characteristics to distinguish them from and make them less prominent than the group of visual elements associated with the primary important system characteristics.

[0227] For further explanation, Figure 7 sets forth a flowchart illustrating an additional exemplary method for a context-driven user interface for a storage system, according to some embodiments of the present disclosure. The exemplary method depicted in Figure 7 is similar to the exemplary method depicted in Figure 4, as it includes receiving 410 a request from a user account 406 to access a system interface 404 of a system 400, identifying 412 at least one important system characteristic that describes a current aspect of the system 400, reconfiguring 414 the system interface 404 based on the at least one important system characteristic, and presenting 416 the reconfigured system interface 404 to a user of the user account 406.

[0228] In the exemplary method depicted in FIG. 7, identifying 412 at least one important system characteristic describing a current aspect of the system 400 includes selecting 702 at least one important system characteristic from a ranked list of system characteristics based on importance. Selecting 702 at least one important system characteristic from the ranked list of system characteristics based on importance may be performed by generating an importance score for each system characteristic. As described above, the importance score may be based on, for example, the urgency or importance of the system characteristic and the number of (potential) users affected. When importance scores are applied to the system characteristics, the system characteristics may be ranked from highest importance score to lowest importance score. The at least one important system characteristic may be selected from among the system characteristics having the highest importance score. Similarly, secondary and tertiary important system characteristics may also be selected from the ranked list.

[0229] 7 , reconfiguring 414 the system interface 404 based on the at least one important system characteristic includes arranging 704 visual elements within the system interface 404 based on a ranking of the at least one important system characteristic within the ranked list of system characteristics. Arranging 704 visual elements within the system interface 404 based on a ranking of the at least one important system characteristic within the ranked list of system characteristics may be performed by rendering visual elements associated with the important system characteristic having the highest importance score from the ranked list and making those visual elements primary elements of the system interface 404. Similarly, visual elements associated with the important system characteristic having the second highest importance score may be made secondary elements of the system interface 404.

[0230] For further explanation, Figure 8 sets forth a flowchart illustrating an additional exemplary method for a context-driven user interface for a storage system, according to some embodiments of the present disclosure. The exemplary method depicted in Figure 8 is similar to the exemplary method depicted in Figure 4, as it includes receiving 410 a request from a user account 406 to access a system interface 404 of a system 400, identifying 412 at least one important system characteristic that describes a current aspect of the system 400, reconfiguring 414 the system interface 404 based on the at least one important system characteristic, and presenting 416 the reconfigured system interface 404 to a user of the user account 406.

[0231] 8 , reconfiguring 414 the system interface 404 based on the at least one important system characteristic includes populating 802 static objects in the system interface 404 with visual elements associated with the at least one important system characteristic. Populating 802 static objects in the system interface 404 with visual elements associated with the at least one important system characteristic may be performed by matching visual elements associated with the important system characteristic with static objects in the system interface 404.

[0232] For each important system characteristic, the system interface 404 can maintain several consistent static objects. Such static objects can include different windows anchored to specific locations on the screen and populated with different visual elements depending on the identified important system characteristic. For example, the system interface 404 may include a primary window anchored to the center of the system interface 404 and using 60% of the height and 60% of the width of the system interface 404, and secondary windows anchored to each corner and occupying the remaining space. The primary window can be populated using visual elements associated with at least one important system characteristic. If secondary important system characteristics are identified, the secondary windows of the system interface 404 can be populated using visual elements associated with those secondary important system characteristics.

[0233] Reconfiguring 414 the system interface 404 based on at least one important system characteristic may further be based on an object relational model. An object relational model is a graph that defines objects and the relationships between the objects. The object relational model of the personality-based system interface 404 described above may include objects that define different personalities for a user account and visual elements for each personality. The object relational model may also define relationships between visual elements associated with each personality. The object relational model may then be used to populate static objects within the system interface 404 or to create dynamic objectification within the system interface 404.

[0234] While some embodiments are described primarily in the context of a storage system, those skilled in the art will recognize that embodiments of the present disclosure may also take the form of a computer program product disposed on a computer-readable storage medium for use with any suitable processing system. Such a computer-readable storage medium may be any storage medium for machine-readable information, including magnetic, optical, solid-state, or other suitable media. Examples of such media include magnetic disks in hard drives or diskettes, compact discs for optical drives, magnetic tape, and other media that will occur to those skilled in the art. Those skilled in the art will readily recognize that any computer system with suitable programming means is capable of performing the steps described herein as embodied in a computer program product. Additionally, those skilled in the art will recognize that while some of the embodiments described herein are directed to software installed and executed on computer hardware, alternative embodiments implemented as firmware or hardware are well within the scope of the present disclosure.

[0235] In some examples, a non-transitory computer-readable medium storing computer-readable instructions may be provided in accordance with the principles described herein. The instructions, when executed by a processor of a computing device, may direct the processor and / or computing device to perform one or more operations, including one or more of the operations described herein. Such instructions may be stored and / or transmitted using any of a variety of known computer-readable media.

[0236] Non-transitory computer-readable media referred to herein may include any non-transitory storage medium that participates in providing data (e.g., instructions) that can be read and / or executed by a computing device (e.g., by a processor of a computing device). For example, non-transitory computer-readable media may include, but are not limited to, any combination of non-volatile storage media and / or volatile media. Exemplary non-volatile storage media include, but are not limited to, read-only memory, flash memory, solid-state drives, magnetic storage devices (e.g., hard disks, floppy disks, magnetic, etc.), ferroelectric random-access memory ("RAM"), and optical disks (e.g., compact disks, digital video disks, Blu-ray disks, etc.). Exemplary volatile storage media include, but are not limited to, RAM (e.g., dynamic RAM).

[0237] One or more embodiments may be described herein with the aid of method steps that illustrate the performance of specified functions and their relationships. The boundaries and sequences of these functional building blocks and method steps have been arbitrarily defined herein for convenience of description. Alternative boundaries and sequences may be defined so long as the specified functions and relationships are properly performed. Accordingly, any such alternative boundaries or sequences are within the scope and spirit of the claims. Furthermore, the boundaries of these functional building blocks have been arbitrarily defined herein for convenience of description. Alternative boundaries may be defined so long as certain key functions are properly performed. Similarly, flow diagram blocks may also be arbitrarily defined herein to illustrate certain key functions.

[0238] To the extent used, flow diagram block boundaries and sequences may be defined otherwise and still perform certain significant functions. Accordingly, such alternative definitions of both the functional building blocks and the flow diagram blocks and sequences are within the scope and spirit of the claims. Those skilled in the art will also recognize that the functional building blocks and other illustrative blocks, modules, and components herein may be implemented as illustrated, or by discrete components, application-specific integrated circuits, processors executing appropriate software, etc., or any combination thereof.

[0239] While particular combinations of various features and characteristics of one or more embodiments are expressly described herein, other combinations of these features and functions are possible as well, and the present disclosure is not limited by the specific examples disclosed herein, but expressly incorporates these other combinations.

[0240] The advantages and features of the present disclosure can be further explained by the following statements.

[0241] 1. A method for receiving a request to access a system interface for a system from a user account, the method including identifying at least one important system characteristic that describes a current aspect of the system, reconfiguring the system interface based on the at least one important system characteristic, and presenting the reconfigured system interface to a user of the user account.

[0242] 2. The method of statement 1, wherein reconfiguring the system interface based on at least one important system characteristic includes identifying a user account personality for the user account session.

[0243] 3. The method described in statement 2 or statement 1, wherein reconfiguring the system interface based on at least one important system characteristic includes arranging visual elements within the system interface such that a visual element associated with the at least one important system characteristic is a primary element of the system interface.

[0244] 4. The method of statement 3, statement 2, or statement 1, wherein identifying at least one important system characteristic that describes a current aspect of the system includes selecting at least one important system characteristic from a ranked list of system characteristics based on importance.

[0245] 5. The method of statement 4, statement 3, statement 2, or statement 1, wherein reconfiguring the system interface based on at least one important system characteristic includes arranging visual elements in the system interface based on a ranking of the at least one important system characteristic in the ranked list of system characteristics.

[0246] 6. The method of statement 5, statement 4, statement 3, statement 2, or statement 1, wherein reconfiguring the system interface based on at least one important system characteristic includes populating the system interface with static objects using visual elements associated with the at least one important system characteristic.

[0247] 7. The method of statement 6, statement 5, statement 4, statement 3, statement 2, or statement 1, wherein the at least one important system characteristic includes a storage utilization level of the system.

[0248] 8. The method of statement 7, statement 6, statement 5, statement 4, statement 3, statement 2, or statement 1, wherein at least one important system characteristic describes a current aspect of the system in relation to an external system.

[0249] 9. The method of statement 8, statement 7, statement 6, statement 5, statement 4, statement 3, statement 2, or statement 1, wherein at least one important system characteristic is identified based on a request to access a system interface.

[0250] 10. The method of statement 9, statement 8, statement 7, statement 6, statement 5, statement 4, statement 3, statement 2, or statement 1, wherein reconfiguring the system interface is further based on an object-relational model.

[0251] One or more embodiments may be described herein with the aid of method steps that illustrate the performance of specified functions and their relationships. The boundaries and sequences of these functional building blocks and method steps have been arbitrarily defined herein for convenience of description. Alternative boundaries and sequences may be defined so long as the specified functions and relationships are properly performed. Accordingly, any such alternative boundaries or sequences are within the scope and spirit of the claims. Furthermore, the boundaries of these functional building blocks have been arbitrarily defined herein for convenience of description. Alternative boundaries may be defined so long as certain key functions are properly performed. Similarly, flow diagram blocks may also be arbitrarily defined herein to illustrate certain key functions.

[0252] To the extent used, flow diagram block boundaries and sequences may be defined otherwise and still perform certain significant functions. Accordingly, such alternative definitions of both the functional building blocks and the flow diagram blocks and sequences are within the scope and spirit of the claims. Those skilled in the art will also recognize that the functional building blocks and other illustrative blocks, modules, and components herein may be implemented as illustrated, or by discrete components, application-specific integrated circuits, processors executing appropriate software, etc., or any combination thereof.

[0253] While particular combinations of various features and characteristics of one or more embodiments are expressly described herein, other combinations of these features and functions are possible as well, and the present disclosure is not limited by the specific examples disclosed herein, but expressly incorporates these other combinations.

Claims

1. 1. A method comprising: In response to a request to access a system interface for a system from a user account, determining a user role associated with the user account; selecting at least one important system characteristic associated with the determined user role based on the evaluation of one or more system characteristics; reconfiguring the system interface based on the at least one important system characteristic, including rendering one or more visual elements associated with the selected at least one important system characteristic; and and presenting the reconfigured system interface to a user of the user account.

2. The method of claim 1, wherein reconfiguring the system interface includes identifying a role for the user based on the request.

3. The method of claim 1, wherein reconfiguring the system interface includes arranging visual elements within the system interface such that a visual element associated with the at least one important system characteristic is visually emphasized more than at least one other visual element of the system interface.

4. The method of claim 1, wherein selecting the at least one important system characteristic includes selecting the at least one important system characteristic from a ranked list of system characteristics based on importance.

5. The method of claim 4, wherein reconfiguring the system interface includes arranging visual elements within the system interface based on a ranking of the at least one important system characteristic within the ranked list of system characteristics.

6. The method of claim 1, wherein reconfiguring the system interface includes populating the system interface with static objects using visual elements associated with the at least one important system characteristic.

7. The method of claim 1 , wherein the at least one important system characteristic comprises a storage utilization level of the system.

8. The method of claim 1 , wherein the at least one important system characteristic describes a current behavior of the system relative to external systems.

9. The method of claim 1 , wherein the at least one important system characteristic is identified based on the request.

10. The method of claim 1 , wherein the reconfiguring of the system interface is further based on an object-relational model.

11. An apparatus comprising: a memory; and a processor operatively coupled to the memory, the processor: In response to a request to access a system interface for the system from a user account, determine a user role associated with the user account; selecting at least one important system characteristic associated with the determined user role based on the evaluation of one or more system characteristics; rendering one or more visual elements associated with the selected at least one important system characteristic; and reconfiguring the system interface based on the at least one important system characteristic; a device configured to present the reconfigured system interface to a user of the user account; 12. The apparatus of claim 11, wherein the processor is further configured to identify a user role based on the request.

13. The device of claim 11, wherein the processor is further configured to arrange visual elements within the system interface such that a visual element associated with the at least one important system characteristic is visually emphasized more than at least one other visual element of the system interface.

14. The apparatus of claim 11, wherein the processor is further configured to select the at least one important system characteristic from a ranked list of system characteristics based on importance.

15. The device described in claim 14, wherein the processor is further configured to arrange visual elements within the system interface based on a ranking of the at least one important system characteristic within the ranked list of system characteristics.

16. The device of claim 11, wherein the processor is further configured to populate the system interface with a static object using a visual element associated with the at least one important system characteristic.

17. The apparatus of claim 11 , wherein the at least one important system characteristic comprises a storage utilization level of the system.

18. The apparatus of claim 11 , wherein the at least one important system characteristic describes a current behavior of the system relative to an external system.

19. The apparatus of claim 11 , wherein the at least one important system characteristic is identified based on the request.

20. A computer program product disposed on a non-transitory computer-readable medium, the medium including instructions that, when executed, cause a processor to: determining a user role associated with the user account in response to a request to access a system interface for the system from the user account; selecting at least one important system characteristic associated with the determined user role based on the evaluation of one or more system characteristics; rendering one or more visual elements associated with the selected at least one important system characteristic; and reconfiguring the system interface based on the at least one important system characteristic; A computer program product having stored thereon instructions for causing the reconfigured system interface to be presented to a user of the user account.

Citation Information

Patent Citations

  • Customization tool for dashboards

    US11144336B1

  • Method, electronic device, and computer program product for monitoring storage system

    US20210342240A1

  • Heuristics for determining the layout of a procedurally generated user interface

    US9300554B1