Profiling user activity to achieve social and governance objectives.

A direct-mapped flash storage system with operating system-controlled processes and zoned storage devices addresses inefficiencies in conventional systems, enhancing reliability and efficiency by eliminating redundant operations.

JP2026086642APending Publication Date: 2026-05-26PURE STORAGE INC

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
PURE STORAGE INC
Filing Date
2026-02-09
Publication Date
2026-05-26

AI Technical Summary

Technical Problem

Conventional storage systems face inefficiencies due to the involvement of both higher-level and lower-level processes, particularly in flash storage systems where flash drives have their own controllers, leading to redundant operations and reduced reliability.

Method used

Implementing a direct-mapped flash storage system where the operating system directly addresses data blocks without address translation by the flash drive's controller, offloading device management responsibilities, and utilizing zones in zoned storage devices for dynamic allocation and management.

Benefits of technology

This approach enhances reliability and reduces redundant operations by allowing the operating system to control processes across multiple flash drives, improving efficiency and reducing wear on storage components.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026086642000001_ABST
    Figure 2026086642000001_ABST
Patent Text Reader

Abstract

To provide methods and apparatus for dynamic and personalized user experiences. [Solution] A method for a dynamic and personalized user experience includes a UI engine receiving a request from a user account to access a system interface for the system, and identifying a user account personality from a plurality of user account personalities based on personality indicators for the user account. Each personality indicator is associated with at least one of the plurality of user account personalities. The method also includes reconfiguring the system interface based on the identified user account personality, presenting the reconfigured system interface to the user of the user account, and granting the user account access to the reconfigured system interface.
Need to check novelty before this filing date? Find Prior Art

Description

Brief Description of the Drawings

[0001] [Figure 1A] Illustrates a first exemplary system for data storage according to some implementation forms. [Figure 1B] Illustrates a second exemplary system for data storage according to some implementation forms. [Figure 1C] Illustrates a third exemplary system for data storage according to some implementation forms. [Figure 1D] Illustrates a fourth exemplary system for data storage according to some implementation forms. [Figure 2A] Perspective view of a storage cluster having a plurality of storage nodes and internal storage coupled to each storage node to provide network-attached storage, according to some embodiments. [Figure 2B] Block diagram showing an interconnect switch that couples a plurality of storage nodes, according to some embodiments. [Figure 2C] Multi-level block diagram showing the content of a storage node and the content of one of the non-volatile solid state storage units, according to some embodiments. [Figure 2D] Illustrates a storage server environment that uses embodiments of the storage nodes and storage units of some of the previous drawings, according to some embodiments. [Figure 2E] Block diagram of blade hardware showing the control plane, the compute and storage planes, and the authorities interacting with the underlying physical resources, according to some embodiments. [Figure 2F] Depicts an elastic software layer within a blade of a storage cluster, according to some embodiments. [Figure 2G] Depicts the authorities and storage resources within a blade of a storage cluster, according to some embodiments. [Figure 3A]A diagram of a storage system coupled with a cloud service provider for data communication, according to some embodiments of this disclosure, is provided. [Figure 3B] A diagram of a storage system according to some embodiments of this disclosure is provided below. [Figure 3C] Examples of cloud-based storage systems according to some embodiments of this disclosure are described below. [Figure 3D] This specification illustrates an exemplary computing device that may be specifically configured to perform one or more of the processes described herein. [Figure 3E] An example of a fleet of storage systems for providing storage services (also referred to herein as "data services") is illustrated. [Figure 3F] This example illustrates a typical container system. [Figure 4] A flowchart illustrating exemplary methods for dynamic, personalized user experiences according to certain embodiments of this disclosure is provided. [Figure 5] A flowchart illustrating additional exemplary methods for dynamic, personalized user experiences according to some embodiments of this disclosure is provided. [Figure 6] A flowchart illustrating additional exemplary methods for dynamic, personalized user experiences according to some embodiments of this disclosure is provided. [Figure 7] A flowchart illustrating additional exemplary methods for dynamic, personalized user experiences according to some embodiments of this disclosure is provided. [Figure 8] A flowchart illustrating additional exemplary methods for dynamic, personalized user experiences according to some embodiments of this disclosure is provided. [Figure 9] A flowchart illustrating exemplary methods for profiling user activity to achieve social and governance objectives, as described in some embodiments of this disclosure, is provided. [Figure 10]A flowchart illustrating additional exemplary methods for profiling user activity to achieve social and governance objectives, as described in some embodiments of this disclosure, is provided. [Modes for carrying out the invention]

[0002] Exemplary methods, apparatuses, and products for dynamic, personalized user experiences according to embodiments of this disclosure will be described with reference to the accompanying drawings beginning with Figure 1A. Figure 1A illustrates an exemplary system for data storage in one of its implementations. System 100 (also referred to herein as the “Storage System”) includes a number of elements, for illustrative purposes only and not limiting. Note that System 100 may include the same, more, or fewer elements, configured in the same or different ways, in other implementations.

[0003] System 100 includes a plurality of computing devices 164A-B. Computing devices (also referred to herein as “client devices”) may be embodied, for example, as servers, workstations, personal computers, notebooks, etc., in a data center. Computing devices 164A-B may be coupled to one or more storage arrays 102A-B for data communication via a storage area network ("SAN") 158 or a local area network ("LAN") 160.

[0004] SAN158 can be implemented using various data communication fabrics, devices, and protocols. For example, fabrics for SAN158 may include Fibre Channel, Ethernet, InfiniBand, and Serial Attached Small Computer System Interface ("SAS"). Data communication protocols used with SAN158 may include Advanced Technology Attachment ("ATA"), Fibre Channel protocol, Small Computer System Interface ("SCSI"), Internet Small Computer System Interface ("iSCSI"), HyperSCSI, and Non-Volatile Memory Express ("NVMe") overfabric. It should be noted that SAN158 is provided for illustrative purposes only, not as an extension. Other data communication couplings may be implemented between computing devices 164A-B and storage arrays 102A-B.

[0005] LAN160 can also be implemented using various fabrics, devices, and protocols. For example, the fabric for LAN160 may include Ethernet (802.3), wireless (802.11), etc. Data communication protocols used in LAN160 may include Transmission Control Protocol ("TCP"), User Datagram Protocol ("UDP"), Internet Protocol ("IP"), HyperText Transfer Protocol ("HTTP"), Wireless Access Protocol ("WAP"), Handheld Device Transport Protocol ("HDTP"), Session Initiation Protocol ("SIP"), Real Time Protocol ("RTP"), etc.

[0006] The storage arrays 102A-B can provide persistent data storage for computing devices 164A-B. In some implementations, storage array 102A can be housed in a chassis (not shown), and storage array 102B can be housed in another chassis (not shown). The storage arrays 102A and 102B may include one or more storage array controllers 110A-D (also referred to herein as “controllers”). The storage array controllers 110A-D can be embodied as modules of an automated computing machine including computer hardware, computer software, or a combination of computer hardware and software. In some implementations, the storage array controllers 110A-D may be configured to perform various storage tasks. Storage tasks may include writing data received from computing devices 164A-B to storage arrays 102A-B, erasing data from storage arrays 102A-B, retrieving data from storage arrays 102A-B and providing it to computing devices 164A-B, monitoring and reporting storage device usage and performance, performing redundant operations such as a Redundant Array of Independent Drive ("RAID") or RAID-like data redundancy, compressing data, and encrypting data.

[0007] The storage array controllers 110A-D may be implemented in various ways, including field programmable gate arrays (FPGAs), programmable logic chips (PLCs), application-specific integrated circuits (ASICs), system-on-chips (SOCs), or any computing device including individual components such as processing devices, central processing units, computer memory, or various adapters. The storage array controllers 110A-D may include, for example, data communication adapters configured to support communication via SAN 158 or LAN 160. In some implementations, the storage array controllers 110A-D may be coupled independently to LAN 160. In some implementations, the storage array controllers 110A-D may include I / O controllers that couple the storage array controllers 110A-D to persistent storage resources 170A-B (also referred to herein as “storage resources”) for data communication via a midplane (not shown). The persistent storage resources 170A-B may include any number of storage drives 171A-F (also referred to herein as “storage devices”) and any number of non-volatile random access memory ("NVRAM") devices (not shown).

[0008] In some implementations, the NVRAM devices of persistent storage resources 170A-B may be configured to receive data stored in storage drives 171A-F from storage array controllers 110A-D. In some examples, the data may originate from computing devices 164A-B. In some examples, writing data to the NVRAM devices may be faster than directly writing data to storage drives 171A-F. In some implementations, the storage array controllers 110A-D may be configured to use the NVRAM devices as a quickly accessible buffer for data that is to be written to storage drives 171A-F. The latency of write requests using the NVRAM devices as a buffer may be improved compared to systems where the storage array controllers 110A-D directly write data to storage drives 171A-F. In some implementations, the NVRAM devices may be implemented using computer memory in the form of high-bandwidth, low-latency RAM. NVRAM devices are referred to as "non-volatile" because they may receive or contain their own power supply to maintain the RAM state after a main power loss to the NVRAM device. Such a power source may be a battery, one or more capacitors, etc. In response to power loss, the NVRAM device may be configured to write the contents of the RAM to persistent storage such as storage drives 171A-F.

[0009] In some implementations, storage drives 171A-F can refer to any device configured to permanently record data, where “permanently” or “persistent” refers to the device’s ability to retain recorded data after power loss. In some implementations, storage drives 171A-F can correspond to non-disk storage media. For example, storage drives 171A-F may be one or more solid-state drives ("SSD"), flash memory-based storage, any type of solid-state non-volatile memory, or any other type of non-mechanical storage device. In other implementations, storage drives 171A-F may include mechanical or rotating hard disks such as hard-disk drives ("HDD").

[0010] In some implementations, storage array controllers 110A-D may be configured to offload device management responsibilities from storage drives 171A-F within storage arrays 102A-B. For example, storage array controllers 110A-D may manage control information that can describe the state of one or more memory blocks within storage drives 171A-F. This control information may indicate, for example, that a particular memory block has failed and should no longer be written to, that a particular memory block contains boot code for storage array controllers 110A-D, the number of program-erase ("P / E") cycles performed on a particular memory block, the age of the data stored in a particular memory block, or the type of data stored in a particular memory block. In some implementations, the control information may be stored as metadata along with the associated memory block. In other implementations, the control information for storage drives 171A-F may be stored in one or more specific memory blocks of storage drives 171A-F selected by storage array controllers 110A-D. The selected memory blocks may be tagged with identifiers indicating that the selected memory blocks contain control information. The identifier may be used by the storage array controllers 110A-D in conjunction with the storage drives 171A-F to quickly identify memory blocks containing control information. For example, the storage controllers 110A-D may issue commands to locate the memory blocks containing control information. Note that the control information may be large enough that parts of it may be stored in multiple locations, or that it may be stored in multiple locations for redundancy purposes, or that it may be distributed across multiple memory blocks within the storage drives 171A-F.

[0011] In some implementations, storage array controllers 110A-D can offload device management responsibilities from storage drives 171A-F of storage arrays 102A-B by retrieving control information from storage drives 171A-F that describes the state of one or more memory blocks within those drives. Retrieving control information from storage drives 171A-F can be performed, for example, by storage array controllers 110A-D querying storage drives 171A-F for the location of the control information for a particular storage drive 171A-F. Storage drives 171A-F may be configured to execute instructions that enable them to identify the location of the control information. These instructions may be executed by a controller (not shown) associated with or otherwise located on storage drives 171A-F, causing storage drives 171A-F to scan a portion of each memory block to identify the memory block that stores the control information for storage drives 171A-F. The storage drives 171A to F may respond by sending a response message to the storage array controllers 110A to D that includes the location of the control information for the storage drives 171A to F. In response to receiving the response message, the storage array controllers 110A to D may issue a request to read the data stored at the address associated with the location of the control information for the storage drives 171A to F.

[0012] In other implementation forms, the storage array controllers 110A to D can further offload the device management responsibility from the storage drives 171A to F by performing storage drive management operations in response to receiving control information. The storage drive management operations can include, for example, operations typically performed by the storage drives 171A to F (e.g., a controller (not shown) associated with a specific storage drive 171A to F). The storage drive management operations can include, for example, ensuring that data is not written to a failed memory block within the storage drives 171A to F, ensuring that data is written to memory blocks within the storage drives 171A to F such that appropriate wear leveling is achieved, and the like.

[0013] In some implementations, storage arrays 102A-B can implement two or more storage array controllers 110A-D. For example, storage array 102A may include storage array controllers 110A and 110B. In a given instance, a single storage array controller 110A-D of the storage system 100 (e.g., storage array controller 110A) may be designated as primary (also referred to herein as "primary controller"), and other storage array controllers 110A-D (e.g., storage array controller 110A) may be designated as secondary (also referred to herein as "secondary controller"). The primary controller may have certain rights, such as permission to modify data in persistent storage resources 170A-B (e.g., to write data to persistent storage resources 170A-B). At least some of the rights of the primary controller may supersede the rights of the secondary controllers. For example, if the primary controller has the right, the secondary controller may not have permission to modify the data in the persistent storage resources 170A-B. The status of storage array controllers 110A-D may change. For example, storage array controller 110A may be designated as secondary, and storage array controller 110B may be designated as primary.

[0014] In some implementations, a primary controller, such as storage array controller 110A, may function as the primary controller for one or more storage arrays 102A-B, and a second controller, such as storage array controller 110B, may function as a secondary controller for one or more storage arrays 102A-B. For example, storage array controller 110A may be the primary controller for storage arrays 102A and 102B, and storage array controller 110B may be a secondary controller for storage arrays 102A and 102B. In some implementations, storage array controllers 110C and 110D (also referred to as "storage processing modules") may not have either primary or secondary status. Storage array controllers 110C and 110D implemented as storage processing modules can function as a communication interface between primary and secondary controllers (e.g., storage array controllers 110A and 110B, respectively) and storage array 102B. For example, the storage array controller 110A of storage array 102A may send a write request to storage array 102B via SAN 158. The write request may be received by both storage array controllers 110C and 110D of storage array 102B. Storage array controllers 110C and 110D facilitate communication and, for example, send the write request to the appropriate storage drives 171A-F. Note that in some implementations, a storage processing module can be used to increase the number of storage drives controlled by the primary and secondary controllers.

[0015] In some implementations, the storage array controllers 110A - D are communicatively coupled via a midplane (not shown) to one or more storage drives 171A - F and one or more non-volatile random access memory (NVRAM) devices (not shown) that are part of the storage arrays 102A - B. The storage array controllers 110A - D may be coupled to the midplane via one or more data communication links, and the midplane may be coupled to the storage drives 171A - F and the NVRAM devices via one or more data communication links. The data communication links described herein are collectively illustrated by the data communication links 108A - D and may include, for example, a Peripheral Component Interconnect Express (PCIe) bus.

[0016] FIG. 1B illustrates an exemplary system for data storage according to some implementations. The storage array controller 101 illustrated in FIG. 1B may be similar to the storage array controllers 110A - D described with respect to FIG. 1A. In one example, the storage array controller 101 may be similar to the storage array controller 110A or the storage array controller 110B. The storage array controller 101 includes a number of elements for purposes of illustration and not limitation. Note that in other implementations, the storage array controller 101 may include the same, more, or fewer elements configured in the same or different ways. Note also that the elements of FIG. 1A may be included below to help illustrate the features of the storage array controller 101.

[0017] The storage array controller 101 may include one or more processing devices 104 and random access memory ("RAM") 111. The processing device 104 (or controller 101) represents one or more general-purpose processing devices, such as a microprocessor or a central processing unit. More specifically, the processing device 104 (or controller 101) may be a complex instruction set computing ("CISC") microprocessor, a reduced instruction set computing ("RISC") microprocessor, a very long instruction word ("VLIW") microprocessor, or a processor implementing another instruction set or a combination of instruction sets. The processing device 104 (or controller 101) may also be one or more dedicated processing devices, such as an ASIC, FPGA, digital signal processor ("DSP"), or network processor.

[0018] The processing device 104 may be connected to the RAM 111 via a data communication link 106, which can be embodied as a high-speed memory bus such as a Double-Data Rate 4 ("DDR4") bus. The RAM 111 stores the operating system 112. In some implementations, instructions 113 are stored in the RAM 111. Instructions 113 may include computer program instructions for performing operations in a direct-mapped flash storage system. In one embodiment, the direct-mapped flash storage system is a system that directly addresses data blocks in a flash drive without address translation performed by the flash drive's storage controller.

[0019] In some implementations, the storage array controller 101 includes one or more host bus adapters 103A-C connected to the processing device 104 via data communication links 105A-C. In some implementations, the host bus adapters 103A-C may be computer hardware connecting the host system (e.g., the storage array controller) to other networks and storage arrays. In some examples, the host bus adapters 103A-C may be a Fibre Channel adapter enabling the storage array controller 101 to connect to a SAN, an Ethernet adapter enabling the storage array controller 101 to connect to a LAN, and so on. The host bus adapters 103A-C may be connected to the processing device 104 via data communication links 105A-C, such as a PCIe bus.

[0020] In some implementations, the storage array controller 101 may include a host bus adapter 114 coupled to an expander 115. The expander 115 can be used to attach the host system to a larger number of storage drives. The expander 115 may be a SAS expander used to enable the host bus adapter 114 to attach to storage drives, for example, in an implementation where the host bus adapter 114 is embodied as a SAS controller.

[0021] In some implementations, the storage array controller 101 may include a switch 116 coupled to the processing device 104 via a data communication link 109. The switch 116 may be a computer hardware device capable of creating multiple endpoints from a single endpoint, thereby enabling multiple devices to share a single endpoint. The switch 116 may be, for example, a PCIe switch coupled to a PCIe bus (e.g., the data communication link 109) and providing multiple PCIe connection points to the midplane.

[0022] In some implementations, the storage array controller 101 includes a data communication link 107 for connecting the storage array controller 101 to other storage array controllers. In some examples, the data communication link 107 may be a QuickPath Interconnect (QPI) interconnect.

[0023] Conventional storage systems using conventional flash drives can implement processes across the flash drives, which are part of the conventional storage system. For example, higher-level processes in the storage system may initiate and control processes across the flash drives. However, flash drives in conventional storage systems may also include their own storage controllers that also perform processes. Therefore, conventional storage systems can implement both higher-level processes (e.g., initiated by the storage system) and lower-level processes (e.g., initiated by the storage controller of the storage system).

[0024] To overcome the various shortcomings of conventional storage systems, operations can be carried out by higher-level processes rather than lower-level processes. For example, a flash storage system may include flash drives that do not include a storage controller that provides the processes. Therefore, the operating system of the flash storage system itself can start and control the processes. This can be achieved by a direct-mapped flash storage system that directly addresses data blocks within the flash drive without address translation performed by the flash drive's storage controller.

[0025] In some implementations, storage drives 171A-F may be one or more zoned storage devices. In some implementations, one or more zoned storage devices may be single HDDs. In some implementations, one or more storage devices may be flash-based SSDs. In zoned storage devices, zoned namespaces on the zoned storage device are grouped by natural size and can be addressed by groups of aligned blocks, forming multiple addressable zones. In some implementations using SSDs, the natural size may be based on the erase block size of the SSD. In some implementations, zones of a zoned storage device may be defined during the initialization of the zoned storage device. In some implementations, zones may be defined dynamically when data is written to the zoned storage device.

[0026] In some implementations, zones may be heterogeneous, with some zones being page groups and others being multiple page groups. In some implementations, some zones may correspond to erase blocks and others to multiple erase blocks. In one implementation, a zone can be any combination of different numbers of pages within page groups and / or erase blocks for heterogeneous combinations of storage device programming modes, manufacturers, product types, and / or product generations, such as those applied to heterogeneous assemblies, upgrades, and distributed storage. In one implementation, a zone may be defined as having usage characteristics such as supporting data with a specific type of lifetime (e.g., very short-lived or very long-lived). These characteristics may be used by the zoned storage device to determine how the zone will be managed over its expected lifetime.

[0027] It is important to understand that zones are virtual structures. No particular zone may have a fixed location on a storage device. Until allocated, a zone may have no location on a storage device. In various implementations, a zone can correspond to a number representing a chunk of virtually allocatable space, which may be the size of an erase block or other block size. When the system allocates or opens a zone, the zone is allocated to flash or other solid-state storage memory, and when the system writes to a zone, pages are written to the mapped flash or other solid-state storage memory of the zoned storage device. When the system closes a zone, the associated erase block or other block size is completed. At some point in the future, the system may delete a zone to free up the space it allocated. During its lifetime, a zone may be moved to a different location on the zoned storage device, for example, when the zoned storage device performs internal maintenance.

[0028] In some implementations, zones in a zoned storage device can be in different states. A zone can be empty, with no data stored in it. An empty zone can be opened explicitly or implicitly by writing data to it. This is the initial state of a zone on a new zoned storage device, but it may also be the result of a zone reset. In some implementations, an empty zone can have a specified location in the flash memory of the zoned storage device. In one implementation, the location of an empty zone can be selected when the zone is first opened or first written to (or later, if the write is buffered in memory). A zone can be open implicitly or explicitly, and an open zone can be written to store data using write or append commands. In one implementation, an open zone can be written to using copy commands to copy data from a different zone. In some implementations, a zoned storage device may have a limit on the number of open zones at any given time.

[0029] A closed zone is a zone that has been partially written to but entered the closed state after an explicit close operation was issued. A closed zone may be left available for future writes, but some of the runtime overhead consumed by keeping the zone open can be reduced. In some implementations, zoned storage devices may have a limit on the number of closed zones at any given time. A full zone is a zone that has stored data and can no longer be written to. A zone can be full either after a write has written data to the entire zone, or as a result of a zone termination operation. Prior to a termination operation, a zone may or may not be fully written to. However, after a termination operation, a zone may not be opened for further writing without first performing a zone reset operation.

[0030] The mapping from a zone to an erase block (or single track in an HDD) may be arbitrary, dynamic, and hidden from view. The process of opening a zone may be an operation that allows a new zone to be dynamically mapped to the underlying storage of the zoned storage device, and then allows data to be written to the zone by adding writes until the zone reaches its capacity. A zone can be terminated at any point, after which no further data can be written to the zone. When the data stored in a zone is no longer needed, the zone can be reset, thereby effectively removing the zone's contents from the zoned storage device and making the physical storage held by that zone available for subsequent data storage. Once a zone is written to and terminated, the zoned storage device ensures that the data stored in the zone is not lost until the zone is reset. In the time between writing data to a zone and resetting a zone, the zone may be moved between single tracks or erase blocks, for example, by copying the data to keep the data refreshed or to handle the aging of memory cells in an SSD, as part of maintenance operations within the zoned storage device.

[0031] In some implementations using HDDs, a zone reset may allow a single track to be allocated to a newly opened zone that may be opened at some point in the future. In some implementations using SSDs, resetting a zone erases the zone's associated physical erase block(s), which may then be reused for data storage. In some implementations, zoned storage devices may have a limit on the number of zones that are open at any given time to reduce the amount of overhead dedicated to keeping zones open.

[0032] The operating system of a flash storage system can identify and maintain a list of allocation units across multiple flash drives of the flash storage system. An allocation unit may be an entire erase block or multiple erase blocks. The operating system may maintain a map or address range that directly maps addresses to erase blocks on the flash drives of the flash storage system.

[0033] Data can be rewritten and erased by using direct mapping to the erase block of a flash drive. For example, the operation may be performed on one or more allocation units containing first data and second data, where the first data is retained and the second data is no longer used by the flash storage system. The operating system can then initiate a process to write the first data to a new location in the other allocation unit, erase the second data, and mark the allocation unit as available for subsequent data. Thus, the process can be performed solely by the higher-level operating system of the flash storage system, without any additional lower-level processes performed by the flash drive controller.

[0034] The advantages of processes performed solely by the operating system of a flash storage system include improved reliability of the flash drives in the flash storage system because unnecessary or redundant write operations are not performed during the process. One possible novelty here is the concept of initiating and controlling the process with the operating system of the flash storage system. In addition, the process can be controlled by the operating system across multiple flash drives. This is in contrast to the processing performed by the storage controller of the flash drives.

[0035] A storage system may consist of two storage array controllers sharing a set of drives for failover purposes, or a single storage array controller providing storage services utilizing multiple drives, or a distributed network of storage array controllers, each having a number of drives or a certain amount of flash storage, which cooperate to provide a complete storage service and cooperate with respect to various aspects of storage services, including storage allocation and garbage collection.

[0036] Figure 1C illustrates a third exemplary system 117 for data storage in one of its implementations. System 117 (also referred to herein as the “storage system”) includes numerous elements for illustrative purposes only, and not limiting. Note that system 117 may include the same, more, or fewer elements, configured in the same or different ways in other implementations.

[0037] In one embodiment, system 117 includes a dual Peripheral Component Interconnect ("PCI") flash storage device 118 having separately addressable high-speed write storage. System 117 may include a storage device controller 119. In one embodiment, storage device controllers 119A-D may be a CPU, ASIC, FPGA, or any other circuit capable of implementing the control structures required in accordance with this disclosure. In one embodiment, system 117 includes flash memory devices (e.g., including flash memory devices 120a-n) operably coupled to various channels of the storage device controller 119. Flash memory devices 120a-n may be presented to controllers 119A-D as an addressable set of flash pages, erase blocks, and / or control elements sufficient to enable the storage device controllers 119A-D to program and retrieve various aspects of the flash. In one embodiment, storage device controllers 119A to D can perform operations on flash memory devices 120a to n, including storing and retrieving data content of pages, placing and erasing arbitrary blocks, tracking statistics on the use and reuse of flash memory pages, erased blocks, and cells, tracking and predicting error codes and failures in flash memory, and controlling voltage levels related to programming flash cells and retrieving content.

[0038] In one embodiment, the system 117 may include RAM 121 for storing separately addressable high-speed write data. In one embodiment, RAM 121 may be one or more separate discrete devices. In another embodiment, RAM 121 may be integrated into storage device controllers 119A-D or a plurality of storage device controllers. RAM 121 may also be used for other purposes, such as temporary program memory for processing devices (e.g., CPU) within the storage device controller 119.

[0039] In one embodiment, the system 117 may include an energy storage device 122, such as a rechargeable battery or capacitor. The energy storage device 122 can store enough energy to power the storage device controller 119, a certain amount of RAM (e.g., RAM 121), and a certain amount of flash memory (e.g., flash memory 120a-120n) for a sufficient amount of time to write the contents of the RAM to the flash memory. In one embodiment, the storage device controllers 119A-D may write the contents of the RAM to the flash memory if the storage device controller detects a loss of external power.

[0040] In one embodiment, the system 117 includes two data communication links 123a and 123b. In one embodiment, the data communication links 123a and 123b may be PCI interfaces. In another embodiment, the data communication links 123a and 123b may be based on other communication standards (e.g., HyperTransport, InfiniBand, etc.). The data communication links 123a and 123b may be based on the Non-Volatile Memory Express ("NVMe") or NVMe over fabric ("NVMf") specification, which enables external connections from other components within the storage system 117 to the storage device controllers 119A-D. Note that, for convenience, the data communication links may be referred to herein interchangeably with PCI buses.

[0041] System 117 may also include an external power supply (not shown), which may be provided via one or both of the data communication links 123a, 123b, or separately. An alternative embodiment includes a separate flash memory (not shown) dedicated for use when storing the contents of RAM 121. Storage device controllers 119A-D may present a logical device on the PCI bus, which may include an addressable high-speed write logical device, or a separate portion of the logical address space of storage device 118, which may be presented as PCI memory or persistent storage. In one embodiment, the operation to store in the device is directed to RAM 121. In the event of a power failure, storage device controllers 119A-D may write the stored contents associated with the addressable high-speed write logical storage to flash memory (e.g., flash memories 120a-n) for long-term persistent storage.

[0042] In one embodiment, the logical device may include some presentation of some or all of the contents of flash memory devices 120a~n, such presentation enabling a storage system, including a storage device 118 (e.g., storage system 117), to directly address flash memory pages and directly reprogram erase blocks from storage system components located outside the storage device via the PCI bus. This presentation may also enable one or more external components to control and retrieve other aspects of the flash memory, such other aspects may include some or all of tracking statistics related to the use and reuse of flash memory pages, erase blocks, and cells across all flash memory devices, tracking and predicting error codes and failures within and across flash memory devices, and controlling voltage levels related to programming and retrieving the contents of flash cells.

[0043] In one embodiment, the energy storage device 122 may be sufficient to ensure the completion of ongoing operations for the flash memory devices 120a-120n, and the energy storage device 122 may power the storage device controllers 119A-D and associated flash memory devices (e.g., 120a-n) for their operations and for storing high-speed write RAM to flash memory. The energy storage device 122 may be used to store cumulative statistics and other parameters held and tracked by the flash memory devices 120a-n and / or the storage device controller 119. A separate capacitor or energy storage device (such as a smaller capacitor near or embedded in the flash memory device itself) may be used for some or all of the operations described herein.

[0044] Various methods can be used to track and optimize the lifespan of the stored energy components, such as adjusting the voltage level over time or partially discharging the stored energy device 122 to measure the corresponding discharge characteristics. If the available energy decreases over time, the effective available capacity of the addressable high-speed write storage may be reduced based on the currently available stored energy to ensure that data can be safely written to.

[0045] Figure 1D illustrates a third exemplary storage system 124 for data storage in some implementation configurations. In one embodiment, the storage system 124 includes storage controllers 125a, 125b. In one embodiment, the storage controllers 125a, 125b are operably coupled to a dual PCI storage device. The storage controllers 125a, 125b can be operably coupled to several host computers 127a~n (for example, via a storage network 130).

[0046] In one embodiment, two storage controllers (e.g., 125a and 125b) provide storage services such as an SCS block storage array, a file server, an object server, a database, or data analysis services. Storage controllers 125a and 125b can provide services to host computers 127a to n outside the storage system 124 through several network interfaces (e.g., 126a to d). Storage controllers 125a and 125b can provide services or applications fully integrated within the storage system 124, forming an integrated storage and computing system. Storage controllers 125a and 125b utilize high-speed write memory in or across storage devices 119a to d to journal ongoing operations, ensuring that operations are not lost due to power failure of one or more software or hardware components within the storage system 124, removal of a storage controller, shutdown of a storage controller or storage system, or any other failure.

[0047] In one embodiment, storage controllers 125a and 125b operate as PCI masters for one or the other PCI bus 128a and 128b. In another embodiment, 128a and 128b may be based on other communication standards (e.g., HyperTransport, InfiniBand, etc.). In other embodiments of the storage system, storage controllers 125a and 125b can operate as multi-masters for both PCI buses 128a and 128b. Alternatively, a PCI / NVMe / NVMf switching infrastructure or fabric may connect multiple storage controllers. Some embodiments of the storage system may allow storage devices to communicate directly with each other rather than only with the storage controllers. In one embodiment, a storage device controller 119a may be able to operate under instructions from storage controller 125a to synthesize and transfer data stored in a flash memory device from data stored in RAM (e.g., RAM 121 in Figure 1C). For example, a recalculated version of the RAM content can be transferred after the storage controller determines that the operation is fully committed across the storage system, or when the high-speed write memory on the device reaches a certain used capacity, or after a certain amount of time, to ensure improved data security or to free up addressable high-speed write capacity for reuse. This mechanism can be used, for example, to avoid a second transfer via a bus (e.g., 128a, 128b) from the storage controllers 125a, 125b. In one embodiment, recalculation may include compressing data, attaching indexing or other metadata, combining multiple data segments, performing loss correction code calculations, and so on.

[0048] In one embodiment, under instructions from storage controllers 125a, 125b, storage device controllers 119a, 119b may be able to operate to compute data from data stored in RAM (e.g., RAM 121 in Figure 1C) and transfer it to another storage device without the involvement of storage controllers 125a, 125b. This operation may be used to mirror data stored in one storage controller 125a to another storage controller 125b, or to offload compression, data aggregation, and / or loss correction coding computations and transfer them to a storage device to reduce the load on the storage controller interfaces 129a, 129b to the storage controllers or PCI buses 128a, 128b.

[0049] The storage device controllers 119A-D may include mechanisms for implementing high-availability primitives used by other parts of the storage system outside of the dual PCI storage device 118. For example, in a storage system having two storage controllers providing highly available storage services, one storage controller may provide reservation or exclusion primitives to prevent the other storage controller from accessing or continuing to access a storage device. This can be used, for example, if one controller detects that the other controller is not functioning properly, or if the interconnection between the two storage controllers itself may not be functioning properly.

[0050] In one embodiment, a storage system for use with a dual PCI direct-mapped storage device having separately addressable high-speed write storage includes a system for managing erase blocks or groups of erase blocks as allocation units for storing data on behalf of a storage service, or for storing metadata associated with a storage service (e.g., indexes, logs, etc.), or for proper management of the storage system itself. Flash pages, which may be several kilobytes in size, can be written as data arrives or as the storage system will persist data over long time intervals (e.g., beyond a defined time threshold). To commit data more quickly or to reduce the number of writes to the flash memory device, the storage controller may first write the data to separately addressable high-speed write storage on the other storage device.

[0051] In one embodiment, storage controllers 125a and 125b can initiate the use of erase blocks within and across storage devices (e.g., 118) according to the years of use and expected remaining lifespan of the storage devices, or based on other statistics. Storage controllers 125a and 125b can initiate garbage collection and data migration between storage devices according to pages that are no longer needed, manage the lifespan of flash pages and erase blocks, and manage the overall system performance.

[0052] In one embodiment, the storage system 124 may utilize mirroring and / or loss correction coding schemes as part of storing data in addressable high-speed write storage and / or as part of writing data to allocation units associated with erase blocks. The erase code may be used across storage devices, within erase blocks or allocation units, or within and across flash memory devices on a single storage device, to provide redundancy against single or multiple storage device failures, or to protect against internal corruption of flash memory pages resulting from flash memory operation or degradation of flash memory cells. Various levels of mirroring and loss correction coding can be used to recover from multiple types of failures occurring separately or in combination.

[0053] The embodiments described with reference to Figures 2A to 2G illustrate a storage cluster that stores user data, such as user data originating from one or more user or client systems, or other sources outside the storage cluster. The storage cluster distributes user data across storage nodes housed in a chassis, or across multiple chassis, using loss correction coding and redundant copies of metadata. Lost correction coding refers to a data protection or reconstruction method in which data is stored across a set of different locations, such as disks, storage nodes, or geographical locations. Flash memory is one type of solid-state memory that can be integrated with the embodiments, but the embodiments can be extended to other types of solid-state memory or other storage media, including non-solid-state memory. Control of storage locations and workloads is distributed across storage locations in a clustered peer-to-peer system. Tasks such as mediating communication between various storage nodes, detecting when a storage node becomes unavailable, and balancing I / O (inputs and outputs) across various storage nodes are all handled on a distributed basis. In some embodiments, data is located or distributed across multiple storage nodes in data fragments or stripes that support data recovery. Data ownership can be reallocated within the cluster, regardless of input and output patterns. This architecture, described in more detail below, allows a storage node in the cluster to fail while the system remains operational, because data can be reconstructed from other storage nodes and therefore remain available for input and output operations. In various embodiments, storage nodes may be referred to as cluster nodes, blades, or servers.

[0054] A storage cluster may be housed in a chassis, i.e., a housing that accommodates one or more storage nodes. Within the chassis are mechanisms for supplying power to each storage node, such as a power distribution bus, and communication mechanisms, such as a communication bus, that enable communication between storage nodes. According to some embodiments, a storage cluster can operate as an independent system in one location. In one embodiment, the chassis includes at least two instances of both the power distribution and communication buses, which can be independently enabled or disabled. The internal communication bus may be an Ethernet bus, but other technologies such as PCIe, InfiniBand, and others are equally suitable. The chassis provides ports for an external communication bus to enable communication between multiple chassis and client systems, either directly or via a switch. External communication can use technologies such as Ethernet, InfiniBand, or Fibre Channel. In some embodiments, the external communication bus uses different communication bus technologies for inter-chassis communication and client communication. If switches are deployed within or between chassis, the switches can function as converters between multiple protocols or technologies. When multiple chassis are connected to define a storage cluster, the storage cluster can be accessed by clients using a proprietary or standard interface such as a network file system ("NFS"), a common internet file system ("CIFS"), a small computer system interface ("SCSI"), or a hypertext transfer protocol ("HTTP"). Conversion from client protocols can be performed within a switch, a chassis external communication bus, or within each storage node. In some embodiments, multiple chassis may be coupled or connected to each other via an aggregator switch. Some and / or all of the coupled or connected chassis may be designated as a storage cluster.As described above, each chassis may have multiple blades, and each blade may have a media access control (MAC) address, but in some embodiments, the storage cluster is presented to the external network as having a single cluster MAC address and a single IP address.

[0055] Each storage node may be one or more storage servers, each storage server connected to one or more non-volatile solid-state memory units, which may be referred to as storage units or storage devices. One embodiment includes a single storage server and 1 to 8 non-volatile solid-state memory units in each storage node, but this example is not limiting. A storage server may include a processor, DRAM, an interface for an internal communication bus, and power distribution for each of the power buses. In some embodiments, within the storage node, the interface and storage units share a communication bus, such as PCI Express. Non-volatile solid-state memory units may have direct access to the internal communication bus interface via the storage node communication bus, or may be required by the storage node to access the bus interface. In some embodiments, a non-volatile solid-state memory unit includes an embedded CPU, a solid-state storage controller, and a solid-state mass storage device of, for example, 2 to 32 terabytes ("TB"). A non-volatile solid-state memory unit includes an internal volatile storage medium such as DRAM and an energy storage device. In some embodiments, the energy storage device is a capacitor, supercapacitor, or battery that allows a subset of the DRAM content to be transferred to a stable storage medium in the event of power loss. In some embodiments, the non-volatile solid-state memory unit is constructed with storage-class memory such as phase-change or magnetoresistive random access memory ("MRAM") that replaces DRAM and enables a reduced power hold-up device.

[0056] One of the many characteristics of storage nodes and non-volatile solid-state storage is their ability to proactively reconstruct data in a storage cluster. Storage nodes and non-volatile solid-state storage can determine when a storage node or non-volatile solid-state storage in a storage cluster becomes unreachable, regardless of whether there are attempts to read data related to that storage node or non-volatile solid-state storage. The storage nodes and non-volatile solid-state storage then work together to recover and reconstruct the data at least partially in a new location. This constitutes proactive reconstruction in that the system reconstructs the data without waiting for the data to be needed for read access initiated by a client system using the storage cluster. These and further details of storage memory and its operation are discussed below.

[0057] Figure 2A is a perspective view of a storage cluster 161, according to one embodiment, having a plurality of storage nodes 150 and internal solid-state memory coupled to each storage node to provide network-attached storage or a storage area network. Network-attached storage, a storage area network, or a storage cluster, or other storage memory may include one or more storage clusters 161, each having one or more storage nodes 150, in a flexible and reconfigurable arrangement of both the physical components and the amount of storage memory provided thereby. The storage cluster 161 is designed to fit into a rack, and one or more racks may be set up and popularized as desired for the storage memory. The storage cluster 161 has a chassis 138 having a plurality of slots 142. It should be understood that the chassis 138 may be referred to as a housing, enclosure, or rack unit. In one embodiment, the chassis 138 has 14 slots 142, but other numbers of slots are also readily conceivable. For example, some embodiments have 4 slots, 8 slots, 16 slots, 32 slots, or other preferred numbers of slots. Each slot 142 can accommodate one storage node 150 in some embodiments. The chassis 138 includes flaps 148 that can be used to mount the chassis 138 in a rack. Fans 144 provide air circulation for cooling the storage nodes 150 and their components, but other cooling components may be used, or an embodiment without cooling components may be devised. A switch fabric 146 connects the storage nodes 150 in the chassis 138 to each other and connects to a network for communication to memory. In one embodiment depicted herein, for illustrative purposes, the left slot 142 of the switch fabric 146 and fan 144 is shown to be occupied by a storage node 150, while the right slot 142 of the switch fabric 146 and fan 144 is empty and available for inserting a storage node 150.This configuration is an example, and one or more storage nodes 150 can occupy slot 142 in various further arrangements. In some embodiments, the arrangement of storage nodes does not need to be contiguous or adjacent. The storage nodes 150 are hot-pluggable, meaning that the storage nodes 150 can be inserted into or removed from slot 142 in the chassis 138 without stopping or powering down the system. When a storage node 150 is inserted into or removed from slot 142, the system recognizes the change and automatically reconfigures to adapt to it. Reconfiguration, in some embodiments, includes restoring redundancy and / or rebalancing data or load.

[0058] Each storage node 150 may have multiple components. In the embodiments shown herein, the storage node 150 includes a CPU 156, i.e., a printed circuit board 159 on which the processor is mounted, a memory 154 coupled to the CPU 156, and a non-volatile solid-state storage 152 coupled to the CPU 156, but in further embodiments, other mountings and / or components may be used. The memory 154 holds instructions executed by the CPU 156 and / or data operated by the CPU 156. As will be further described below, the non-volatile solid-state storage 152 includes flash, or in further embodiments, other types of solid-state memory.

[0059] Referring to Figure 2A, the storage cluster 161 is scalable, meaning that storage capacity with non-uniform storage sizes can be easily added, as described above. In some embodiments, one or more storage nodes 150 can be plugged into or removed from each chassis, and the storage cluster is self-configured. Plug-in storage nodes 150 can have different sizes, whether installed in the chassis at delivery or added later. For example, in one embodiment, storage nodes 150 can have any multiple of 4TB, e.g., 8TB, 12TB, 16TB, 32TB, etc. In further embodiments, storage nodes 150 can have any multiple of other storage amounts or capacities. The storage capacity of each storage node 150 is broadcast and influences the decision of how to stripe the data. For maximum storage efficiency, one embodiment can be self-configured as widely as possible within the stripe, subject to certain requirements for continuous operation with loss of up to one or up to two non-volatile solid-state storage units 152 or storage nodes 150 within the chassis.

[0060] Figure 2B is a block diagram showing a communication interconnect 173 and a power distribution bus 172 connecting multiple storage nodes 150. Referring back to Figure 2A, the communication interconnect 173 may, in some embodiments, be included in or implemented with the switch fabric 146. If multiple storage clusters 161 occupy a rack, in some embodiments, the communication interconnect 173 may be included in or implemented with a top-of-rack switch. As illustrated in Figure 2B, the storage cluster 161 is enclosed within a single chassis 138. External ports 176 are connected to the storage nodes 150 via the communication interconnect 173, and external ports 174 are connected directly to the storage nodes. External power ports 178 are connected to the power distribution bus 172. The storage nodes 150 may include non-volatile solid-state storage 152 of varying amounts and capacities, as described with reference to Figure 2A. In addition, one or more storage nodes 150 may be dedicated compute storage nodes, as illustrated in Figure 2B. The authorization 168 is implemented on the non-volatile solid-state storage 152, for example, as a list or other data structure stored in memory. In some embodiments, the authorization is stored within the non-volatile solid-state storage 152 and supported by software running on the controller of the non-volatile solid-state storage 152 or on another processor. In further embodiments, the authorization 168 is implemented on the storage node 150, for example, as a list or other data structure stored in memory 154 and supported by software running on the CPU 156 of the storage node 150. In some embodiments, the authorization 168 controls how and where the data is stored in the non-volatile solid-state storage 152. This control helps determine what type of loss-of-data correction coding scheme is applied to the data and which storage node 150 has which portion of the data. Each authorization 168 can be assigned to the non-volatile solid-state storage 152.Each authority can, in various embodiments, control a range of inode numbers, segment numbers, or other data identifiers that are assigned to data by the file system, by the storage node 150, or by the non-volatile solid-state storage 152.

[0061] In some embodiments, all data and metadata have redundancy within the system. Furthermore, all data and metadata have an owner, which may be called an authority. If that authority becomes unreachable, for example, due to a storage node failure, there is a succession plan for how to find the data or metadata. In various embodiments, there are redundant copies of authority 168. In some embodiments, authority 168 has relationships with storage nodes 150 and non-volatile solid-state storage 152. Each authority 168 covering a range of data segment numbers or other identifiers of data may be assigned to a specific non-volatile solid-state storage 152. In some embodiments, authority 168 for all such ranges are distributed across the non-volatile solid-state storage 152 of the storage cluster. Each storage node 150 has a network port that provides access to its non-volatile solid-state storage 152. Data can be stored in segments associated with segment numbers, which in some embodiments are an indirect reference to the configuration of a RAID (Redundant Array of Independent Disks) stripe. Thus, the assignment and use of authority 168 establishes an indirect reference to the data. Indirect referencing, according to some embodiments, may be referred to as the ability to indirectly reference data, in this case via authority 168. A segment identifies a set of non-volatile solid-state storage 152 and a local identifier to the set of non-volatile solid-state storage 152 that may contain data. In some embodiments, the local identifier is an offset to a device and may be reused sequentially by multiple segments. In other embodiments, the local identifier is unique to a particular segment and is never reused. The offset within the non-volatile solid-state storage 152 is applied to the location of data for writing to or reading from the non-volatile solid-state storage 152 (in the form of a RAID stripe).The data is striped across multiple units of non-volatile solid-state storage 152, which may include, or may not include, a non-volatile solid-state storage 152 having authorization 168 for a particular data segment.

[0062] For example, if there is a change in the location of a particular segment of data during data movement or data reconstruction, the authorization 168 for that data segment should be referenced in a non-volatile solid-state storage 152 or storage node 150 that has that authorization 168. To locate specific data, embodiments calculate a hash value of the data segment or apply an inode number or data segment number. The output of this operation points to a non-volatile solid-state storage 152 that has authorization 168 for that particular data. In some embodiments, this operation has two stages. The first stage is to map an entity identifier (ID), such as a segment number, inode number, or directory number, to an authorization identifier. This mapping may involve calculations such as hashing or bitmasking. The second stage is to map the authorization identifier to a specific non-volatile solid-state storage 152, which can be done through explicit mapping. This operation is repeatable, and therefore, once the calculation is performed, the result of the calculation will repeatedly and reliably point to a specific non-volatile solid-state storage 152 that has that authorization 168. The operation may take a set of reachable storage nodes as input. If the set of reachable non-volatile solid-state storage units changes, the optimal set changes. In some embodiments, the persisting value is the current allocation (which is always true), and the calculated value is the target allocation that the cluster attempts to reconfigure. This calculation may be used to determine the optimal non-volatile solid-state storage 152 for authorization, given the presence of a set of non-volatile solid-state storage 152 that are reachable and constitute the same cluster. The calculation also determines an ordered set of peer non-volatile solid-state storage 152 that also record the mapping of authorization to non-volatile solid-state storage so that authorization can be determined even if the allocated non-volatile solid-state storage is unreachable. In some embodiments, if a particular authorization 168 is unavailable, a duplicate or substitute authorization 168 may be referenced.

[0063] Referring to Figures 2A and 2B, two of the many tasks of the CPU 156 on storage node 150 are to split the data to be written and to reassemble the data to be read. When the system determines that data is to be written, the authority 168 for that data is located as described above. If the segment ID of the data has already been determined, the write request is forwarded to the non-volatile solid-state storage 152 which is now determined to be the host of the authority 168 determined from the segment. The host CPU 156 of storage node 150, where the non-volatile solid-state storage 152 and the corresponding authority 168 reside, then decomposes or shards the data and transmits the data to the various non-volatile solid-state storage 152. The transmitted data is written as data stripes according to an erasure correction coding scheme. In some embodiments, data is requested to be pulled, and in other embodiments, data is pushed. Conversely, when data is read, the authority 168 for the segment ID containing that data is located as described above. The host CPU 156 of storage node 150, where non-volatile solid-state storage 152 and corresponding authorization 168 reside, requests data from the non-volatile solid-state storage and corresponding storage node pointed to by authorization. In some embodiments, data is read from flash storage as a data stripe. The host CPU 156 of storage node 150 then reassembles the read data, corrects any errors (if any) according to an appropriate loss-of-error coding scheme, and transfers the reassembled data to the network. In further embodiments, some or all of these tasks can be handled in the non-volatile solid-state storage 152. In some embodiments, a segment host requests that data be sent to storage node 150 by requesting a page from storage and then sending the data to the storage node that made the original request.

[0064] In the embodiment, authorization 168 operates to determine how an operation proceeds for a particular logical element. Each logical element can be operated through a specific authorization across multiple storage controllers of the storage system. Authorization 168 can communicate with multiple storage controllers to cause them to collectively perform an operation for that particular logical element.

[0065] In embodiments, a logical element may be, for example, a file, directory, object bucket, individual object, file or object delimiter, other form of key-value pair database, or table. In embodiments, performing an operation may involve, for example, ensuring consistency with other operations on the same logical element, structural integrity, and / or recoverability; reading metadata and data associated with that logical element; determining which data should be permanently written to the storage system in order to persist any changes for the operation; or determining that metadata and data are stored across modular storage devices attached to multiple storage controllers within the storage system.

[0066] In some embodiments, the operation is a token-based transaction for efficient communication within a distributed system. Each transaction may be accompanied by, or associated with, a token that grants permission to execute the transaction. In some embodiments, authority 168 can maintain the system's pre-transaction state until the operation is complete. Token-based communication can be achieved across the system without global locks and allows for the resumption of operations in the event of interruption or other failure.

[0067] In some systems, such as UNIX-style file systems, data is processed by index nodes or inodes that specify data structures representing objects within the file system. Objects can be, for example, files or directories. Metadata can accompany objects as attributes such as permission data and creation timestamps, among other attributes. Segment numbers can be assigned to all or some of such objects within the file system. In other systems, data segments are processed by segment numbers assigned elsewhere. For illustrative purposes, the unit of distribution is an entity, which can be a file, directory, or segment. That is, an entity is a unit of data or metadata stored by a storage system. Entities are grouped into sets called privileges. Each privilege has a privilege owner, which is a storage node with the exclusive right to update the entities within the privilege. In other words, a storage node contains a privilege, and that privilege contains entities.

[0068] A segment is a logical container of data, according to some embodiments. A segment is an address space between the media address space and a physical flash location; i.e., the data segment number resides in this address space. A segment may also contain metadata that allows data redundancy to be restored (rewritten to a different flash location or device) without higher-level software involvement. In one embodiment, the internal format of a segment includes client data and a media mapping for determining the location of that data. Each data segment is protected from, for example, memory and other failures by dividing the segment into multiple data shards and parity shards, where applicable. The data shards and parity shards are distributed, i.e., striped, across non-volatile solid-state storage 152 coupled to the host CPU 156 (see Figures 2E and 2G) according to an erasure correction coding scheme. The use of the term segment, in some embodiments, refers to the container and its location in the address space of the segment. The use of the term stripe, in some embodiments, refers to the same set of shards as a segment, and includes how the shards are distributed with redundancy or parity information, according to some embodiments.

[0069] A series of address space translations are performed throughout the entire storage system. At the top are directory entries (filenames) that link to inodes. Inodes refer to the medium address space where data is logically stored. Medium addresses can be mapped through a series of indirect media to distribute the load of large files or to implement data services such as deduplication or snapshots. Next, segment addresses are translated to physical flash locations. According to some embodiments, the physical flash locations have an address range limited by the amount of flash in the system. Medium addresses and segment addresses are logical containers, and in some embodiments, use identifiers of 128 bits or more to be substantially infinite, with reusability calculated to be longer than the expected lifespan of the system. In some embodiments, addresses from logical containers are allocated hierarchically. First, a range of address space can be allocated to each non-volatile solid-state storage 152 unit. Within this allocated range, the non-volatile solid-state storage 152 can assign addresses without synchronizing with other non-volatile solid-state storage 152s.

[0070] Data and metadata are stored using a set of fundamental storage layouts optimized for various workload patterns and storage devices. These layouts incorporate multiple redundancy schemes, compression formats, and indexing algorithms. Some of these layouts store information about authorizations and authorization masters, while others store file metadata and file data. Redundancy schemes include error correction codes that tolerate corrupted bits within a single storage device (such as a NAND flash chip), erase codes that tolerate failures of multiple storage nodes, and replication schemes that tolerate data center or regional failures. In some embodiments, low-density parity check ("LDPC") codes are used within a single storage unit. In some embodiments, Reed-Solomon coding is used within a storage cluster, and mirroring is used within a storage grid. Metadata may also be stored using ordered log-structured indexes (such as log-structured merge trees), and large data may not be stored in log-structured layouts.

[0071] To maintain consistency across multiple copies of an entity, storage nodes implicitly agree through computation on two things: (1) the authority containing the entity, and (2) the storage node containing the authority. Assigning entities to authorities can be done by pseudo-randomly assigning entities to authorities, by dividing entities into ranges based on externally generated keys, or by assigning a single entity to each authority. Examples of pseudo-random methods are the hash families of linear hashing and Replication Under Scalable Hashing ("RUSH"), including Controlled Replication Under Scalable Hashing ("CRUSH"). In some embodiments, pseudo-random assignment is used solely to assign authorities to nodes because the set of nodes may change. Since the set of authorities cannot be changed, arbitrary subjective functions may be applied in these embodiments. Some placement schemes automatically place authorities on storage nodes, while others rely on explicit mapping of authorities to storage nodes. In some embodiments, pseudo-random methods are used to map each authority to a set of candidate authority owners. The pseudo-random data distribution function associated with CRUSH can assign privileges to storage nodes and create a list of locations to which privileges are assigned. Each storage node has a copy of the pseudo-random data distribution function, which arrives at the same computation for distribution and can later discover or locate privileges. Each of the pseudo-random schemes requires a reachable set of storage nodes as input to conclude the same target node in some embodiments. Once entities are placed within a privilege, they can be stored on physical devices so that expected failures do not lead to unexpected data loss. In some embodiments, the rebalancing algorithm attempts to store copies of all entities within a privilege on the same set of machines in the same layout.

[0072] Examples of anticipated failures include device failure, stolen equipment, data center fire, and local disasters such as nuclear or geological events. Different failures result in different levels of acceptable data loss. In some embodiments, a stolen storage node may have no impact on the security or reliability of the system, but depending on the system configuration, a local event may result in no data loss, loss of updates for seconds or minutes, or even complete data loss.

[0073] In some embodiments, the placement of data for storage redundancy is independent of the placement of authority for data consistency. In some embodiments, storage nodes containing authority do not contain any persistent storage. Instead, the storage nodes are connected to non-volatile solid-state storage units that do not contain authority. The communication interconnection between the storage nodes and the non-volatile solid-state storage units consists of multiple communication technologies and has non-uniform performance and fault tolerance characteristics. In some embodiments, as described above, the non-volatile solid-state storage units are connected to the storage nodes via PCI Express, the storage nodes are connected together within a single chassis using an Ethernet backplane, and the chassis are connected together to form a storage cluster. In some embodiments, the storage cluster is connected to clients using Ethernet or Fibre Channel. When multiple storage clusters are configured in a storage grid, the multiple storage clusters are connected using the Internet or other long-distance networking links such as “metro-scale” links or private links that do not traverse the Internet.

[0074] The authorization holder has the exclusive right to modify entities, migrate entities from one non-volatile solid-state storage unit to another, and add and remove copies of entities. This allows for the maintenance of redundancy of underlying data. If the authorization holder fails, is scheduled for decommissioning, or is overloaded, the authorization is transferred to a new storage node. In the case of transient failure, it is important to ensure that all non-failed machines agree on the new authorization location. The ambiguity arising from transient failure can be resolved through consensus protocols such as Paxos, hot-warm failover schemes, via manual intervention by a remote system administrator, or automatically by a local hardware administrator (e.g., by physically removing the failed machine from the cluster or pressing a button on the failed machine). In some embodiments, a consensus protocol is used and failover is automatic. According to some embodiments, if too many failure or replication events occur in too short a period of time, the system enters a self-save mode and suspends replication and data migration activities until administrator intervention is received.

[0075] When permissions are transferred between storage nodes and the permission owner updates entities within those permissions, the system transfers messages between the storage nodes and non-volatile solid-state storage units. With respect to persistent messages, messages with different purposes are of different types. Depending on the message type, the system maintains different ordering and durability guarantees. When persistent messages are being processed, they are temporarily stored in multiple durable and non-durable storage hardware technologies. In some embodiments, messages are stored in RAM, NVRAM, and NAND flash devices, and various protocols are used to efficiently utilize each storage medium. Latency-sensitive client requests may persist in replicated NVRAM and then later in NAND, while background rebalancing operations persist directly in NAND.

[0076] Persistent messages are stored permanently before transmission. This allows the system to continue serving client requests despite failures and component replacements. While many hardware components include unique identifiers visible to system administrators, manufacturers, hardware supply chains, and ongoing monitoring and quality control infrastructure, applications running on infrastructure addresses virtualize these addresses. These virtualized addresses remain unchanged throughout the lifespan of the storage system, regardless of component failures and replacements. This allows each component of the storage system to be replaced over time without reconfiguring or interrupting client request processing; in other words, the system supports non-disruptive upgrades.

[0077] In some embodiments, virtualized addresses are stored with sufficient redundancy. The continuous monitoring system correlates hardware and software status with hardware identifiers. This enables the detection and prediction of failures caused by defective components and manufacturing details. In some embodiments, the monitoring system also enables proactive transfer of privileges and entities from affected devices before a failure occurs by removing components from the critical path.

[0078] Figure 2C is a multilevel block diagram showing the contents of storage node 150 and the contents of non-volatile solid-state storage 152 of storage node 150. In some embodiments, data is communicated to and from storage node 150 by network interface controller ("NIC") 202. Each storage node 150 has a CPU 156 and one or more non-volatile solid-state storage 152, as described above. Moving down one level in Figure 2C, each non-volatile solid-state storage 152 has relatively fast non-volatile solid-state memory such as non-volatile random access memory ("NVRAM") 204 and flash memory 206. In some embodiments, NVRAM 204 may be a component that does not require program / erase cycles (DRAM, MRAM, PCM) and may be memory that can support writing much more frequently than the memory is read. Moving to another level in Figure 2C, NVRAM 204 is implemented in one embodiment as fast volatile memory such as dynamic random access memory (DRAM) 216 backed by energy storage 218. The energy storage 218 provides sufficient power to continue supplying power to the DRAM 216 for a sufficient period of time in the event of a power failure, allowing the contents to be transferred to the flash memory 206. In some embodiments, the energy storage 218 is a capacitor, supercapacitor, battery, or other device that provides a suitable supply of energy sufficient to enable the transfer of the contents of the DRAM 216 to a stable storage medium in the event of power loss. The flash memory 206 is implemented as a plurality of flash dies 222, which may be referred to as a package of flash dies 222 or an array of flash dies 222. It should be understood that the flash dies 222 can be packaged in any number of ways, including a single die per package, multiple dies per package (i.e., a multi-chip package), a hybrid package, a bare die on a printed circuit board or other substrate, or an encapsulated die.In the illustrated embodiment, the non-volatile solid-state storage 152 includes a controller 212 or another processor and input / output (I / O) ports 210 coupled to the controller 212. The I / O ports 210 are coupled to the CPU 156 and / or network interface controller 202 of the flash storage node 150. Flash input / output (I / O) ports 220 are coupled to the flash die 222, and a direct memory access (DMA) unit 214 is coupled to the controller 212, DRAM 216, and flash die 222. In the illustrated embodiment, the I / O ports 210, controller 212, DMA unit 214, and flash I / O ports 220 are implemented on a programmable logic device ("PLD") 208, for example, an FPGA. In this embodiment, each flash die 222 has pages organized as 16kB (kilobyte) pages 224 and registers 226 that can write data to or read data from the flash die 222. In further embodiments, other types of solid-state memory are used instead of or in addition to the flash memory exemplified within the flash die 222.

[0079] The storage cluster 161 can generally be compared to a storage array in various embodiments as disclosed herein. The storage nodes 150 are part of a collection that makes up the storage cluster 161. Each storage node 150 owns a slice of data and the computing power necessary to deliver the data. Multiple storage nodes 150 work together to store and retrieve data. Storage memory or storage devices, as generally used in storage arrays, are not heavily involved in the processing and manipulation of data. Storage memory or storage devices in a storage array receive commands to read, write, or erase data. Storage memory or storage devices in a storage array are unaware of the larger system in which they are embedded or what the data means. Storage memory or storage devices in a storage array may include various types of storage memory, such as RAM, solid-state drives, and hard disk drives. The non-volatile solid-state storage units 152 described herein are simultaneously active and have multiple interfaces that serve multiple purposes. In some embodiments, some of the functionality of the storage node 150 is shifted to the storage unit 152, transforming the storage unit 152 into a combination of the storage unit 152 and the storage node 150. Placing the computing (on the storage data) in the storage unit 152 places this computing closer to the data itself. Various system embodiments have a hierarchy of storage node layers with different capabilities. In contrast, in a storage array, the controller owns and is aware of everything about all the data that the controller manages within the shelves or storage devices. In a storage cluster 161, as described herein, multiple controllers within multiple non-volatile solid-state storage 152 units and / or storage nodes 150 cooperate in various ways (e.g., for loss correction coding, data sharding, metadata communication and redundancy, storage capacity expansion or reduction, data recovery, etc.).

[0080] Figure 2D shows a storage server environment using embodiments of the storage node 150 and storage 152 units shown in Figures 2A-2C. In this version, each non-volatile solid-state storage 152 unit has a processor such as a controller 212 (see Figure 2C), an FPGA, flash memory 206, and NVRAM 204 (supercapacitor-backed DRAM 216, see Figures 2B and 2C) on a PCIe (Peripheral Component Interconnect Express) board in chassis 138 (see Figure 2A). The non-volatile solid-state storage 152 unit may be implemented as a single board containing the storage, or it may be the largest acceptable failure domain in the chassis. In some embodiments, up to two non-volatile solid-state storage 152 units may fail, and the device continues without data loss.

[0081] In some embodiments, physical storage is divided into named regions based on application usage. NVRAM 204 is a contiguous block of reserved memory within non-volatile solid-state storage 152, DRAM 216, backed by NAND flash. NVRAM 204 is logically divided into multiple memory regions, each written to two spools (e.g., spool_region). The space within the NVRAM 204 spool is independently managed by each authority 168. Each device provides a certain amount of storage space to each authority 168, which further manages the lifetime and allocation within that space. Examples of spools include distributed transactions or concepts. When primary power to the non-volatile solid-state storage 152 unit fails, an onboard supercapacitor provides a short-duration power hold-up. During this hold-up interval, the contents of NVRAM 204 are flashed to flash memory 206. Upon the next power-up, the contents of NVRAM 204 are recovered from flash memory 206.

[0082] With respect to the storage unit controllers, the responsibilities of the logical "controller" are distributed across each blade, including the authority 168. This distribution of logical control is shown in Figure 2D as the host controller 242, the middle-tier controller 244, and the storage unit controller(s) 246. Although the management of the control plane and storage plane is handled independently, some may be physically located in the same place on the same blade. Each authority 168 functions effectively as an independent controller. Each authority 168 provides its own data and metadata structure, its own background workers, and maintains its own lifecycle.

[0083] Figure 2E is a hardware block diagram of blade 252, showing a control plane 254, compute plane 256, and storage plane 258, as well as authorities 168, that interact with the underlying physical resources, using embodiments of the storage node 150 and storage unit 152 of Figures 2A-2C in the storage server environment of Figure 2D. The control plane 254 is divided into multiple authorities 168 that can run on any of the blades 252 using compute resources in the compute plane 256. The storage plane 258 is divided into a set of devices, each providing access to flash 206 and NVRAM 204 resources. In one embodiment, the compute plane 256 can perform the operation of a storage array controller on one or more devices of the storage plane 258 (e.g., a storage array), as described herein.

[0084] In the compute plane 256 and storage plane 258 of Figure 2E, the authorization 168 interacts with underlying physical resources (i.e., devices). From the perspective of authorization 168, its resources are striped across all physical devices. From the perspective of devices, devices provide resources to all authorizations 168, regardless of where the authorization happens to be running. Each authorization 168 has allocated or is allocated one or more partitions 260 of storage memory in the storage unit 152, for example, partitions 260 in flash memory 206 and NVRAM 204. Each authorization 168 uses those allocated partitions 260 belonging to it to write or read user data. Authorities can be associated with different amounts of physical storage in the system. For example, one authorization 168 may have more partitions 260 or larger partitions 260 in one or more storage units 152 than one or more other authorizations 168.

[0085] Figure 2F illustrates the resilience software layer within a blade 252 of a storage cluster in one embodiment. In the resilience structure, the resilience software is symmetric; that is, the compute module 270 of each blade executes three identical layers of the process depicted in Figure 2F. The storage manager 274 executes read and write requests from other blades 252 to data and metadata stored in the local storage unit 152, NVRAM 204, and flash 206. Authorization 168 fulfills client requests by issuing the necessary reads and writes to the blade 252 on the storage unit 152 where the corresponding data or metadata resides. The endpoint 272 parses client connection requests received from the monitoring software of the switch fabric 146, relays the client connection requests to authority 168, which is responsible for fulfilling them, and relays authority 168's response to the client. The symmetrical three-layer structure enables a high degree of simultaneity in the storage system. Resilience scales out efficiently and reliably in these embodiments. In addition, elasticity implements a unique scale-out technique that maximizes simultaneity by evenly balancing work across all resources regardless of client access patterns and eliminating much of the need for inter-blade coordination that typically occurs in traditional distributed locking.

[0086] Referring further to Figure 2F, the authority 168, which runs within the compute module 270 of blade 252, performs the internal operations necessary to fulfill client requests. One feature of resilience is that the authority 168 is stateless; that is, it caches active data and metadata in the DRAM of its own blade 252 for high-speed access, but stores all updates in its NVRAM 204 partitions on the three separate blades 252 until the updates are written to flash 206. In some embodiments, all storage system writes to NVRAM 204 are triple-redundant across the partitions on the three separate blades 252. With triple-mirrored NVRAM 204 and persistent storage protected by parity and Reed-Solomon RAID checksums, the storage system can withstand the simultaneous failure of two blades 252 without losing access to data, metadata, or either.

[0087] Since the privileged 168s are stateless, they can be migrated between blades 252. Each privileged 168 has a unique identifier. The NVRAM 204 and flash 206 partitions are associated with the identifier of the privileged 168, rather than the blade 252 in which they are partially running. Therefore, when a privileged 168 migrates, it continues to manage the same storage partitions from its new location. When a new blade 252 is installed in one embodiment of the storage cluster, the system automatically rebalances the load by partitioning the storage of the new blade 252 for use by the privileged 168s of the system, migrating selected privileged 168s to the new blade 252, starting endpoints 272 on the new blade 252, and including them in the client connectivity distribution algorithm of the switch fabric 146.

[0088] From their new locations, the migrated privilege 168s persist the contents of their NVRAM 204 partitions on flash 206, process read and write requests from other privilege 168s, and fulfill client requests directed to them by endpoint 272. Similarly, if blade 252 fails or is removed, the system redistributes its privilege 168s among the remaining blades 252 of the system. The redistributed privilege 168s continue to perform their original functions from their new locations.

[0089] Figure 2G illustrates the authority 168s and storage resources within a blade 252 of a storage cluster in one embodiment. Each authority 168 is exclusively responsible for the flash 206 and NVRAM 204 partitions on each blade 252. An authority 168 manages the content and integrity of its partitions independently of other authority 168s. An authority 168 compresses incoming data, temporarily stores it in its NVRAM 204 partitions, and then integrates the data into its flash 206 partition storage segments, providing RAID protection and ensuring its persistence. When an authority 168 writes data to flash 206, the storage manager 274 performs the necessary flash conversions to optimize write performance and maximize media lifetime. In the background, the authority 168 performs "garbage collection," i.e., reclaims space occupied by data that is no longer needed when clients overwrite it. It should be understood that because the partitions of the authority 168s are disparate, there is no need for distributed locking for client and write operations or for performing background functions.

[0090] The embodiments described herein can utilize a variety of software, communication, and / or networking protocols. In addition, the hardware and / or software configurations can be adapted to various protocols. For example, embodiments can utilize Active Directory, a database-based system that provides authentication, directory, policy, and other services in a Windows® environment. In these embodiments, Lightweight Directory Access Protocol (LDAP) is an example of an application protocol for querying and modifying items within a directory service provider such as Active Directory. In some embodiments, a network lock manager ("NLM") is used in cooperation with a network file system ("NFS") to provide System V-style advisory file and record locking over the network. The Server Message Block ("SMB") protocol, one version of which is also known as the Common Internet File System ("CIFS"), may be integrated with the storage systems described herein. SMB operates as an application layer network protocol typically used to provide shared access to files, printers, and serial ports, as well as various communications between nodes on a network. SMB also provides an authenticated inter-process communication mechanism. Amazon® S3 (Simple Storage Service) is a web service provided by Amazon Web Services, and the systems described herein may interface with Amazon S3 via web service interfaces (REST (Representational State Transfer), SOAP (Simple Object Access Protocol), and BitTorrent). The RESTful API (Application Programming Interface) breaks down transactions into a series of smaller modules.Each module addresses a specific fundamental part of a transaction. In particular, the controls or permissions provided in these embodiments for object data may include the use of access control lists ("ACLs"). An ACL is a list of permissions attached to an object, specifying which users or system processes are permitted to access an object, and which actions are permitted for a given object. The system may utilize Internet Protocol version 6 ("IPv6") and IPv4 for communication protocols that provide a system for identifying and locating computers on a network and routing traffic over the Internet. Routing packets between networked systems may include equal-cost multi-path routing ("ECMP"), a routing strategy where next-hop packet forwarding to a single destination can occur over multiple "best paths" that are coupled at the top of the routing metric calculation. Because multipath routing is a hop-by-hop decision limited to a single router, it can be used with most routing protocols. The software may support multi-tenancy, an architecture in which a single instance of a software application serves multiple customers. Each customer may be referred to as a tenant. A tenant may be given the ability to customize parts of the application, although in some embodiments, customization of the application code may not be required. Embodiments may maintain audit logs. Audit logs are documents that record events in a computing system. In addition to documenting which resources were accessed, audit log entries typically include destination and source addresses, timestamps, and user login information to comply with various regulations. Embodiments may support various key management policies, such as encryption key rotation.Furthermore, the system can support dynamic root passwords or any variation that dynamically changes the password.

[0091] Figure 3A shows a diagram of a storage system 306 coupled with a cloud service provider 302 for data communication, according to some embodiments of the present disclosure. Although not depicted in much detail, the storage system 306 depicted in Figure 3A may be similar to the storage systems described above with reference to Figures 1A-1D and 2A-2G. In some embodiments, the storage system 306 depicted in Figure 3A may be embodied as a storage system including unbalanced active / active controllers, a storage system including balanced active / active controllers, a storage system including active / active controllers where fewer resources than all of the resources of each controller are utilized so that each controller has reserve resources that can be used to support failover, a storage system including fully active / active controllers, a storage system including a data set isolated controller, a storage system including a dual-tier architecture with a front-end controller and a back-end integrated storage controller, a storage system including a scale-out cluster of dual controller arrays, and combinations of such embodiments.

[0092] In the example depicted in Figure 3A, the storage system 306 is connected to the cloud service provider 302 via a data communication link 304. Such data communication link 304 may be entirely wired, entirely wireless, or any combination of wired and wireless data communication paths. In such an example, digital information may be exchanged between the storage system 306 and the cloud service provider 302 via the data communication link 304 using one or more data communication protocols. For example, digital information may be exchanged between the storage system 306 and the cloud service provider 302 via the data communication link 304 using Handheld Device Transfer Protocol ("HDTP"), Hypertext Transfer Protocol ("HTTP"), Internet Protocol ("IP"), Real-time Transfer Protocol ("RTP"), Transmission Control Protocol ("TCP"), User Datagram Protocol ("UDP"), Wireless Application Protocol ("WAP"), or other protocols.

[0093] The cloud service provider 302 depicted in Figure 3A can be embodied as a system and computing environment that provides vast services to its users, for example, through the sharing of computing resources via a data communication link 304. The cloud service provider 302 can provide on-demand access to a shared pool of configurable computing resources, such as computer networks, servers, storage, applications, and services.

[0094] In the example depicted in Figure 3A, the cloud service provider 302 can be configured to provide various services to the storage system 306 and its users through various service model implementations. For example, the cloud service provider 302 may be configured to provide services through an infrastructure as a service ("IaaS") service model implementation, a platform as a service ("PaaS") service model implementation, a software as a service ("SaaS") service model implementation, an authentication as a service ("AaaS") service model implementation, or a storage as a service model implementation that provides access to the storage infrastructure for use by the storage system 306 and its users.

[0095] In the example depicted in Figure 3A, the cloud service provider 302 may be embodied, for example, as a private cloud, as a public cloud, or as a combination of a private cloud and a public cloud. In one embodiment in which the cloud service provider 302 is embodied as a private cloud, the cloud service provider 302 may be dedicated to providing services to a single organization rather than providing services to multiple organizations. In one embodiment in which the cloud service provider 302 is embodied as a public cloud, the cloud service provider 302 can provide services to multiple organizations. In a further alternative embodiment, the cloud service provider 302 may be embodied as a hybrid of private and public cloud services having a hybrid cloud deployment.

[0096] Although not explicitly depicted in Figure 3A, readers will understand that a vast number of additional hardware and software components may be required to facilitate the delivery of cloud services to the storage system 306 and its users. For example, the storage system 306 may be coupled to (or include) a cloud storage gateway. Such a cloud storage gateway may be embodied, for example, as a hardware-based or software-based device located on-premises with the storage system 306. Such a cloud storage gateway can act as a bridge between local applications running on the storage system 306 and remote cloud-based storage utilized by the storage system 306. Through the use of a cloud storage gateway, an organization may be able to move its primary iSCSI or NAS to the cloud service provider 302, thereby saving space on its on-premises storage systems. Such a cloud storage gateway may be configured to emulate a disk array, block-based device, file server, or other storage system, capable of translating SCSI commands, file server commands, or other appropriate commands into the REST spatial protocol, which facilitates communication with the cloud service provider 302.

[0097] A cloud migration process may be performed to enable storage system 306 and its users to utilize services provided by cloud service provider 302, during which data, applications, or other elements from the organization's local systems (or from another cloud environment) are moved to cloud service provider 302. Middleware, such as cloud migration tools, can be used to bridge the gap between the cloud service provider 302 environment and the organization's environment in order to successfully migrate the data, applications, or other elements to the cloud service provider 302 environment. To further enable storage system 306 and its users to utilize services provided by cloud service provider 302, a cloud orchestrator may also be used to arrange and coordinate automated tasks, pursuing the creation of an integrated process or workflow. Such a cloud orchestrator can perform tasks such as configuring various components, determining whether those components are cloud components or on-premises components, and managing the interconnections between such components.

[0098] In the example depicted in Figure 3A, as briefly described above, the cloud service provider 302 can be configured to provide services to the storage system 306 and its users through the use of a SaaS service model. For example, the cloud service provider 302 may be configured to provide access to a data analysis application to the storage system 306 and its users. Such a data analysis application may be configured to receive, for example, a large amount of telemetry data transferred to the home by the storage system 306. Such telemetry data can describe various operational characteristics of the storage system 306 and can be analyzed for a wide range of purposes, including determining the health of the storage system 306, identifying workloads running on the storage system 306, predicting when the storage system 306 will run out of various resources, and recommending configuration changes, hardware or software upgrades, workflow migrations, or other actions that can improve the operation of the storage system 306.

[0099] The cloud service provider 302 can also be configured to provide access to a virtualized computing environment to the storage system 306 and its users. Examples of such a virtualized environment may include virtual machines created to emulate actual computers, virtualized desktop environments that separate logical desktops from physical machines, and virtualized file systems that enable uniform access to different types of specific file systems.

[0100] The example depicted in Figure 3A illustrates a storage system 306 being coupled for data communication with a cloud service provider 302, but in other embodiments, the storage system 306 may be part of a hybrid cloud deployment in which private cloud elements (e.g., private cloud services, on-premises infrastructure, etc.) and public cloud elements (e.g., private cloud services, infrastructure, etc., which may be provided by one or more cloud service providers) are combined to form a single solution through orchestration across various platforms. Such a hybrid cloud deployment may leverage hybrid cloud management software, such as Microsoft® Azure® Arc, which centralizes the management of the hybrid cloud deployment on any infrastructure and enables the deployment of services anywhere. In such an example, the hybrid cloud management software may be configured to create, update, and delete resources (both physical and virtual) that make up the hybrid cloud deployment, to allocate compute and storage to specific workloads, to monitor workloads and resources for performance, policy compliance, updates and patches, security status, or to perform various other tasks.

[0101] Readers will understand that various offerings can be made by pairing the storage systems described herein with one or more cloud service providers. For example, disaster recovery as a service ("DRaaS") can be provided, in which cloud resources are used to protect applications and data from disruption caused by disasters, including embodiments in which the storage system can function as a primary data store. In such embodiments, a full system backup can be taken to enable business continuity in the event of a system failure. In such embodiments, cloud data backup technology (either by itself or as part of a larger DRaaS solution) can also be integrated into the overall solution, including the storage systems and cloud service providers described herein.

[0102] The storage systems and cloud service providers described herein can be used to provide a variety of security features. For example, a storage system can encrypt dormant data (and send and receive encrypted data to and from the storage system), and can manage encryption keys, keys for locking and unlocking storage devices, etc., using Key Management-as-a-Service ("KMaaS"). Similarly, a cloud data security gateway or similar mechanism can be used to ensure that data stored in the storage system is not improperly stored in the cloud as part of cloud data backup operations. Furthermore, microsegmentation or identity-based segmentation can be used in the data center, including the storage system, or within the cloud service provider to create secure zones that enable the isolation of workloads from one another in data center and cloud deployments.

[0103] For further explanation, Figure 3B shows a diagram of a storage system 306 according to some embodiments of the present disclosure. Although not depicted in much detail, the storage system 306 depicted in Figure 3B may be similar to the storage systems described above with reference to Figures 1A-1D and 2A-2G, because the storage system may include many of the components described above.

[0104] The storage system 306 depicted in Figure 3B may include a vast amount of storage resources 308, which can be embodied in many forms. For example, the storage resources 308 may include flash memory, such as nanoRAM or other forms of non-volatile random-access memory utilizing carbon nanotubes deposited on a substrate, 3D crosspoint non-volatile memory, single-level cell ("SLC") NAND flash, multi-level cell ("MLC") NAND flash, triple-level cell ("TLC") NAND flash, quad-level cell ("QLC") NAND flash, or others. Similarly, the storage resources 308 may include magnetoresistive random-access memory ("MRAM"), which includes spin transfer torque ("STT") MRAM. The exemplary storage resource 308 may alternatively include other forms of storage resources, including non-volatile phase-change memory ("PCM"), quantum memory enabling the storage and retrieval of photonic quantum information, resistive random-access memory ("ReRAM"), storage class memory ("SCM"), or any combination of the resources described herein. The reader will understand that other forms of computer memory and storage devices, including DRAM, SRAM, EEPROM, universal memory, etc., may be utilized by the storage system described above.The storage resource 308 depicted in Figure 3A can be embodied in various form factors, including but not limited to dual in-line memory modules ("DIMM"), non-volatile dual in-line memory modules ("NVDIMM"), M.2, U.2, and others.

[0105] The storage resource 308 depicted in Figure 3B may include various forms of SCM. The SCM can effectively treat high-speed non-volatile memory (e.g., NAND flash) as an extension of DRAM, such that the entire dataset can be treated as an in-memory dataset residing entirely within DRAM. The SCM may include, for example, a non-volatile medium such as NAND flash. Such NAND flash may be accessed using NVMe, which can use the PCIe bus as its transport, providing relatively low access latency compared to older protocols. In fact, network protocols used for SSDs in all-flash arrays include NVMe using Ethernet (ROCE, NVMe TCP), Fibre Channel (NVMe FC), InfiniBand (iWARP), and others that enable the handling of high-speed non-volatile memory as an extension of DRAM. Given the fact that DRAM is often byte-addressable, while high-speed non-volatile memory such as NAND flash is block-addressable, a controller software / hardware stack may be required to translate block data into bytes stored in the medium. Examples of media and software that can be used as SCM include, for example, 3D XPoint, Intel Memory Drive Technology, Samsung's Z-SSD, and others.

[0106] The storage resource 308 depicted in Figure 3B may also include racetrack memory (also known as domain wall memory). Such racetrack memory can be embodied in a solid-state device as a form of non-volatile solid-state memory that depends not only on the charge of electrons but also on the intrinsic strength and orientation of the magnetic field generated by electrons when they spin. By using spin-coherent current to move magnetic domains along nanoscale permalloy wires, the magnetic domains can pass near a magnetic read / write head positioned close to the wire as the current passes through the wire, thereby modifying the magnetic domains and recording a pattern of bits. Many such wires and read / write elements can be packaged together to create a racetrack memory device.

[0107] The exemplary storage system 306 depicted in Figure 3B can implement various storage architectures. For example, a storage system according to some embodiments of the present disclosure may utilize block storage in which data is stored in blocks, each block essentially acting as an individual hard drive. A storage system according to some embodiments of the present disclosure may utilize object storage in which data is managed as objects. Each object may include the data itself, a variable amount of metadata, and a globally unique identifier, and object storage can be implemented at multiple levels (e.g., device level, system level, interface level). A storage system according to some embodiments of the present disclosure may utilize file storage in which data is stored in a hierarchical structure. Such data may be stored in files and folders and presented in the same format to both the system storing and the system retrieving it.

[0108] The exemplary storage system 306 depicted in Figure 3B may be embodied as a storage system that can add additional storage resources through the use of a scale-up model, through the use of a scale-out model, or through any combination thereof. In the scale-up model, additional storage can be added by adding additional storage devices. However, in the scale-out model, additional storage nodes may be added to a cluster of storage nodes, such storage nodes may include additional processing resources, additional networking resources, and so on.

[0109] The exemplary storage system 306 depicted in Figure 3B can utilize the storage resources described above in various different ways. For example, some parts of the storage resources may be used to function as a write cache, storage resources within the storage system may be used as a read cache, or tiering may be achieved within the storage system by arranging data within the storage system according to one or more tiering policies.

[0110] The storage system 306 depicted in Figure 3B also includes communication resources 310 that may be useful for facilitating data communication between components within the storage system 306, as well as data communication between the storage system 306 and computing devices outside the storage system 306, including embodiments in which these resources are isolated by a relatively wide spread. The communication resources 310 may be configured to facilitate data communication between components within the storage system and computing devices outside the storage system by utilizing various different protocols and data communication fabrics. For example, communication resource 310 may include Fibre Channel ("FC") technology such as FC fabric and FC protocol that can transport SCSI commands over an FC network, FC over Ethernet ("FCoE") technology in which FC frames are encapsulated and transmitted over an Ethernet network, InfiniBand ("IB") technology in which a switched fabric topology is used to facilitate transmission between channel adapters, NVM Express ("NVMe") technology and NVMe over fabric ("NVMeoF") technology that can access non-volatile storage media attached via a PCI Express ("PCIe") bus, and others. In fact, the above storage system may directly or indirectly utilize neutrino communication technology and devices in which information (including binary information) is transmitted using a neutrino beam.

[0111] The communication resource 310 may also include a serial attached SCSI ("SAS"), a serial ATA ("SATA") bus interface for connecting the storage resource 308 in the storage system 306 to a host bus adapter in the storage system 306, Internet Small Computer System Interface ("iSCSI") technology for providing block-level access to the storage resource 308 in the storage system 306, and a mechanism for accessing the storage resource 308 in the storage system 306 by utilizing other communication resources that may be useful for facilitating data communication between components in the storage system 306, as well as data communication between the storage system 306 and computing devices outside the storage system 306.

[0112] The storage system 306 depicted in Figure 3B also includes processing resources 312 that may be useful for executing computer program instructions and performing other computational tasks within the storage system 306. The processing resources 312 may include one or more ASICs customized for some specific purpose, as well as one or more CPUs. The processing resources 312 may also include one or more DSPs, one or more FPGAs, one or more systems on a chip ("SoC"), or other forms of processing resources 312. The storage system 306 can utilize the storage resources 312 to perform a variety of tasks, including, but not limited to, supporting the execution of software resources 314, which will be described in more detail below.

[0113] The storage system 306 depicted in Figure 3B also includes a software resource 314 that can perform a vast number of tasks when executed by the processing resource 312 within the storage system 306. The software resource 314 may include, for example, one or more modules of computer program instructions that are useful for performing various data protection techniques when executed by the processing resource 312 within the storage system 306. Such data protection techniques may be performed, for example, by system software running on the computer hardware within the storage system, by a cloud service provider, or in other ways. Such data protection techniques may include data archiving, data backup, data replication, data snapshots, data and database cloning, and other data protection techniques.

[0114] The software resource 314 may also include software useful in implementing software-defined storage ("SDS"). In such an example, the software resource 314 may include one or more modules of computer program instructions that, when executed, are useful in policy-based provisioning and management of data storage independent of the underlying hardware. Such software resource 314 may be useful in implementing storage virtualization to separate the storage hardware from the software that manages the storage hardware.

[0115] The software resource 314 may also include software useful for facilitating and optimizing I / O operations directed to the storage system 306. For example, the software resource 314 may include software modules that implement various data reduction techniques, such as data compression, data deduplication, and others. The software resource 314 may also include software modules that intelligently group I / O operations to facilitate better use of the underlying storage resource 308, software modules that perform data migration operations for migration from within the storage system, and software modules that perform other functions. Such software resource 314 may be embodied as one or more software containers or in many other ways.

[0116] For further explanation, Figure 3C illustrates an example of a cloud-based storage system 318 according to some embodiments of the present disclosure. In the example depicted in Figure 3C, the cloud-based storage system 318 is entirely created within a cloud computing environment 316, such as Amazon Web Services ("AWS"), Microsoft Azure, Google Cloud Platform, IBM Cloud, Oracle Cloud, and others. The cloud-based storage system 318 may be used to provide services similar to those that may be provided by the storage systems described above.

[0117] The cloud-based storage system 318 depicted in Figure 3C includes two cloud computing instances 320 and 322, each used to support the execution of storage controller applications 324 and 326. The cloud computing instances 320 and 322 can be embodied as instances of cloud computing resources (e.g., virtual machines) that may be provided by the cloud computing environment 316 to support the execution of software applications such as storage controller applications 324 and 326. For example, each of the cloud computing instances 320 and 322 may run on an Azure VM, and each Azure VM may include high-speed temporary storage that can be utilized as a cache (e.g., as a read cache). In one embodiment, the cloud computing instances 320 and 322 can be embodied as Amazon Elastic Compute Cloud ("EC2") instances. In such an example, a virtual machine can be created and configured to run storage controller applications 324 and 326 by booting an Amazon Machine Image ("AMI") containing the storage controller applications 324 and 326.

[0118] In the exemplary method depicted in Figure 3C, the storage controller applications 324 and 326 may be embodied as modules of computer program instructions that perform various storage tasks when executed. For example, the storage controller applications 324 and 326 may be embodied as modules of computer program instructions that perform the same tasks as the controllers 110A and 110B in Figure 1A described above, such as writing data to the cloud-based storage system 318, erasing data from the cloud-based storage system 318, retrieving data from the cloud-based storage system 318, monitoring and reporting storage device utilization and performance, performing redundant operations such as RAID or RAID-like data redundancy, data compression, data encryption, and data deduplication when executed. Since there are two cloud computing instances 320 and 322, each containing a storage controller application 324 or 326, the reader will understand that in some embodiments, one cloud computing instance 320 can operate as the primary controller as described above, and the other cloud computing instance 322 can operate as the secondary controller as described above. The reader will understand that the storage controller applications 324 and 326 depicted in Figure 3C may contain identical source code running within different cloud computing instances 320 and 322, such as separate EC2 instances.

[0119] Readers will understand that other embodiments not involving primary and secondary controllers are within the scope of this disclosure. For example, each cloud computing instance 320, 322 may act as a primary controller for several portions of the address space supported by the cloud-based storage system 318, each cloud computing instance 320, 322 may act as a primary controller from which the services of I / O operations directed to the cloud-based storage system 318 are otherwise partitioned, and so on. In fact, in other embodiments where cost savings may take precedence over performance requirements, there may be only a single cloud computing instance containing the storage controller application.

[0120] The cloud-based storage system 318 depicted in Figure 3C includes cloud computing instances 340a, 340b, and 340n, each having local storage 330, 334, and 338. These cloud computing instances 340a, 340b, and 340n can be embodied as instances of cloud computing resources, which may be provided by the cloud computing environment 316 to support the execution of software applications, for example. While the cloud computing instances 340a, 340b, and 340n in Figure 3C have local storage 330, 334, and 338 resources, cloud computing instances 320 and 322, which support the execution of storage controller applications 324 and 326, do not need to have local storage resources. Therefore, the cloud computing instances 340a, 340b, and 340n in Figure 3C may differ from the aforementioned cloud computing instances 320 and 322. Cloud computing instances 340a, 340b, and 340n having local storage 330, 334, and 338 can be embodied, for example, as EC2 M5 instances with one or more SSDs, as EC2 R5 instances with one or more SSDs, as EC2 I3 instances with one or more SSDs, and so on. In some embodiments, the local storage 330, 334, and 338 must be embodied as solid-state storage (e.g., SSDs) rather than storage utilizing hard disk drives.

[0121] In the example depicted in Figure 3C, each of the cloud computing instances 340a, 340b, and 340n, accompanied by local storage 330, 334, and 338, may include software daemons 328, 332, and 336, which, when executed by the cloud computing instances 340a, 340b, and 340n, can present themselves to the storage controller applications 324 and 326 as if the cloud computing instances 340a, 340b, and 340n were physical storage devices (e.g., one or more SSDs). In such an example, the software daemons 328, 332, and 336 may include computer program instructions similar to those that would normally be contained on a storage device, so that the storage controller applications 324 and 326 can send and receive the same commands that the storage controller would send to the storage device. In this way, the storage controller applications 324 and 326 may include code that is identical (or substantially identical) to the code executed by the controller in the storage system described above. In these and similar embodiments, communication between the storage controller applications 324, 326 and the cloud computing instances 340a, 340b, 340n using local storage 330, 334, 338 can utilize iSCSI, NVMe over TCP, messaging, custom protocols, or any other mechanism.

[0122] In the example depicted in Figure 3C, each of the cloud computing instances 340a, 340b, and 340n, accompanied by local storage 330, 334, and 338, may also be coupled to block storage 342, 344, and 346 provided by the cloud computing environment 316, such as Amazon Elastic Block Store ("EBS") volumes. In such an example, the block storage 342, 344, and 346 provided by the cloud computing environment 316 may be utilized in a similar manner to how the NVRAM devices described above are utilized, so that when a software daemon 328, 332, and 336 (or some other module) running within a particular cloud computing instance 340a, 340b, and 340n receives a request to write data, it may initiate writing data to its attached EBS volume and to its local storage 330, 334, and 338 resources. In some alternative embodiments, data may only be written to local storage resources 330, 334, and 338 within a specific cloud, including instances 340a, 340b, and 340n. In alternative embodiments, instead of using block storage 342, 344, and 346 provided by the cloud computing environment 316 as NVRAM, the actual RAM on each of the cloud computing instances 340a, 340b, and 340n having local storage 330, 334, and 338 may be used as NVRAM, thereby reducing the network utilization costs associated with using EBS volumes as NVRAM. In yet another embodiment, one or more high-performance block storage resources, such as Azure Ultra Disks, may be used as NVRAM.

[0123] When a request to write data is received by a specific cloud computing instance 340a, 340b, or 340n having local storage 330, 334, or 338, the software daemons 328, 332, or 336 may be configured not only to write the data to their own local storage 330, 334, or 338 resources and any appropriate block storage 342, 344, or 346 resources, but also to write the data to a cloud-based object storage 348 attached to that specific cloud computing instance 340a, 340b, or 340n. The cloud-based object storage 348 attached to that specific cloud computing instance 340a, 340b, or 340n may be embodied, for example, as Amazon Simple Storage Service ("S3"). In other embodiments, cloud computing instances 320, 322, each containing a storage controller application 324, 326, respectively, can initiate storage of data to the local storage 330, 334, 338 and cloud-based object storage 348 of the cloud computing instances 340a, 340b, 340n. In other embodiments, the persistent storage tier may be implemented in other ways, rather than using both the cloud computing instances 340a, 340b, 340n (also referred to herein as “virtual drives”) with local storage 330, 334, 338 and cloud-based object storage 348 to store data. For example, one or more Azure Ultra disks can be used to persistently store data (e.g., after the data has been written to the NVRAM tier). In one embodiment where one or more Azure Ultra disks may be used to persistently store data, the use of cloud-based object storage 348 may be eliminated so that the data is persistently stored only on the Azure Ultra disks without writing the data to the object storage tier.

[0124] While the local storage resources 330, 334, and 338 and block storage resources 342, 344, and 346 used by cloud computing instances 340a, 340b, and 340n can support block-level access, the cloud-based object storage 348 attached to specific cloud computing instances 340a, 340b, and 340n only supports object-level access. Therefore, the software daemons 328, 332, and 336 may be configured to retrieve blocks of data, package those blocks into objects, and write the objects to the cloud-based object storage 348 attached to specific cloud computing instances 340a, 340b, and 340n.

[0125] In some embodiments, all data stored by the cloud-based storage system 318 may be stored in both 1) the cloud-based object storage 348 and 2) at least one of the local storage resources 330, 334, 338 or the block storage resources 342, 344, 346 used by the cloud computing instances 340a, 340b, 340n. In such embodiments, the local storage resources 330, 334, 338 and the block storage resources 342, 344, 346 used by the cloud computing instances 340a, 340b, 340n can effectively function as a cache generally containing all data also stored in S3, thereby allowing all data retrieval to be serviced by the cloud computing instances 340a, 340b, 340n without requiring them to access the cloud-based object storage 348. However, readers will understand that in other embodiments, all data stored by the cloud-based storage system 318 may be stored in the cloud-based object storage 348, while less than all data stored by the cloud-based storage system 318 may be stored in at least one of the local storage resources 330, 334, 338 or the block storage resources 342, 344, 346 used by the cloud computing instances 340a, 340b, 340n. In such examples, various policies can be used to determine which subsets of data stored by the cloud-based storage system 318 should reside in both 1) the cloud-based object storage 348 and 2) at least one of the local storage resources 330, 334, 338 or the block storage resources 342, 344, 346 used by the cloud computing instances 340a, 340b, 340n.

[0126] One or more modules of computer program instructions running within a cloud-based storage system 318 (for example, a monitoring module running on its own EC2 instance) may be designed to handle the failure of one or more cloud computing instances 340a, 340b, and 340n having local storage 330, 334, and 338. In such an example, the monitoring module can handle the failure of one or more cloud computing instances 340a, 340b, and 340n having local storage 330, 334, and 338 by creating one or more new cloud computing instances with local storage, retrieving data stored in the failed cloud computing instances 340a, 340b, and 340n from the cloud-based object storage 348, and storing the retrieved data from the cloud-based object storage 348 in the local storage of the newly created cloud computing instances. The reader will understand that many variations of this process can be implemented.

[0127] Readers will understand that various performance aspects of the cloud-based storage system 318 can be monitored (for example, by a monitoring module running on an EC2 instance) so that the cloud-based storage system 318 can be scaled up or scaled out as needed. For example, if the cloud computing instances 320 and 322 used to support the running of storage controller applications 324 and 326 are small and do not adequately service the I / O requests issued by users of the cloud-based storage system 318, the monitoring module may create a new, more powerful cloud computing instance (for example, a cloud computing instance of a type that includes more processing power, more memory, etc.) containing the storage controller applications so that the new, more powerful cloud computing instance can begin to act as the primary controller. Similarly, if the monitoring module determines that the cloud computing instances 320 and 322 used to support the running of storage controller applications 324 and 326 are oversized and cost savings can be achieved by switching to a smaller, less powerful cloud computing instance, the monitoring module may create a new, less powerful (and cheaper) cloud computing instance containing the storage controller applications so that the new, less powerful cloud computing instance can begin to act as the primary controller.

[0128] The storage system described above can perform intelligent data backup technology, which allows it to copy data stored in the storage system and store it in a different location, thereby avoiding data loss in the event of equipment failure or other forms of catastrophic disaster. For example, the storage system described above may be configured to inspect each backup in order to avoid restoring the storage system to an undesirable state. Consider an example in which malware infects the storage system. In such an example, the storage system may include a software resource 314 that can scan each backup to identify backups captured before the malware infected the storage system and backups captured after the malware infected the storage system. In such an example, the storage system may restore itself from backups that do not contain malware, or at least not restore the portion of backups that contain malware. In such an example, the storage system may include software resources 314 that can identify write operations served by the storage system, for example, by identifying write operations originating from a network subnet serviced by the storage system and suspected of distributing malware, by identifying write operations originating from a user serviced by the storage system and suspected of distributing malware, by examining the contents of the write operations against malware fingerprints, and by scanning each backup in many other ways to identify the presence of malware (or viruses, or any other unwanted entities).

[0129] The reader will also understand that backups (often in the form of one or more snapshots) can be used to perform rapid recovery of the storage system. Consider an example where a storage system is infected with ransomware that locks the user out of the storage system. In such an example, software resources 314 within the storage system may be configured to detect the presence of ransomware and may be further configured to restore the storage system to a point in time using backups kept before the ransomware infected the storage system. In such an example, the presence of ransomware may be explicitly detected through the use of software tools utilized by the system, through the use of a key (e.g., a USB drive) inserted into the storage system, or in a similar manner. Similarly, the presence of ransomware may be inferred in response to system activity that satisfies a given fingerprint, for example, no reads or writes to the system for a given period of time.

[0130] The reader will understand that the various components described above can be grouped into one or more optimized computing packages as an integrated infrastructure. Such an integrated infrastructure may include a pool of computing, storage, and networking resources that are shared by multiple applications and can be collectively managed using policy-driven processes. Such an integrated infrastructure can be implemented using an integrated infrastructure reference architecture, using standalone devices, using software-driven hyperintegration techniques (e.g., hyperintegrated infrastructure), or in other ways.

[0131] Readers will understand that the storage systems described in this disclosure may be useful for supporting various types of software applications. In fact, a storage system can be “application-aware” in the sense that it can acquire, maintain, or otherwise access information describing connected applications (e.g., applications that utilize the storage system) and optimize the operation of the storage system based on intelligence about the applications and their usage patterns. For example, the storage system may optimize the data layout, optimize the caching behavior, optimize the “QoS” level, or implement any other optimizations designed to improve the storage performance experienced by applications.

[0132] As an example of one type of application that may be supported by the storage system described herein, the storage system 306 may be useful in supporting artificial intelligence ("AI") applications, database applications, XOps projects (e.g., DevOps projects, DataOps projects, MLOps projects, ModelOps projects, PlatformOps projects), electronic design automation tools, event-driven software applications, high-performance computing applications, simulation applications, high-speed data capture and analysis applications, machine learning applications, media production applications, media serving applications, picture archiving and communication system ("PACS") applications, software development applications, virtual reality applications, augmented reality applications, and many other types of applications by providing storage resources to such applications.

[0133] Given that storage systems include computing resources, storage resources, and a wide variety of other resources, they can be very well-suited to support resource-intensive applications such as AI applications. AI applications can be deployed in a variety of fields, including predictive maintenance in manufacturing and related fields, healthcare applications such as patient data and risk analysis, retail and marketing deployments (e.g., search advertising, social media advertising), supply chain solutions, fintech solutions such as business analytics and reporting tools, and operational deployments such as real-time analytics tools, application performance management tools, and IT infrastructure management tools.

[0134] Such AI applications can enable devices to perceive their environments and take actions that maximize their chances of success for a given purpose. Examples of such AI applications may include IBM Watson®, Microsoft Oxford®, Google DeepMind®, Baidu Minwa®, and others.

[0135] The storage systems described above are also well-suited to supporting other resource-intensive types of applications, such as machine learning applications. Machine learning applications can perform various types of data analysis and automate the construction of analytical models. Using algorithms that learn iteratively from data, machine learning applications can enable computers to learn without being explicitly programmed. One particular area of ​​machine learning is called reinforcement learning, which involves taking optimal actions to maximize rewards in specific situations.

[0136] In addition to the resources already described, the storage system described above may also include a graphics processing unit (GPU), sometimes referred to as a visual processing unit (VPU). Such a GPU may be embodied as a dedicated electronic circuit that rapidly manipulates and modifies memory to accelerate the creation of images in a frame buffer intended for output to a display device. Such a GPU may be contained within one of the computing devices that are part of the storage system described above, which is included as one of many individually extensible components of the storage system. Other examples of individually extensible components of such a storage system may include storage components, memory components, computing components (e.g., CPU, FPGA, ASIC), networking components, software components, and others. In addition to the GPU, the storage system described above may also include a neural network processor (NNP) for use in various embodiments of neural network processing. Such an NNP may be used instead of (or in addition to) a GPU and may be independently scalable.

[0137] As described above, the storage systems described herein can be configured to support artificial intelligence applications, machine learning applications, big data analytics applications, and many other types of applications. The rapid growth in these types of applications is driven by three technologies: deep learning (DL), GPU processors, and big data. Deep learning is a computing model that utilizes large-scale parallel neural networks inspired by the human brain. Instead of experts writing software by hand, deep learning models write their own software by learning from many examples. Such GPUs may contain thousands of cores well suited to running algorithms that roughly represent the parallelism of the human brain.

[0138] Advances in deep neural networks, including the development of multi-layer neural networks, have spurred a new wave of algorithms and tools for data scientists to leverage their data using artificial intelligence (AI). Improved algorithms, larger datasets, and a variety of frameworks (including open-source software libraries for machine learning across a range of tasks) are enabling data scientists to tackle new use cases such as autonomous vehicles, natural language processing and understanding, computer vision, machine reasoning, and strong AI. Applications of AI technology are embodied in a wide range of products, including, for example, Amazon Echo's speech recognition technology that allows users to talk to their machines, Google Translate® enabling machine-based language translation, Spotify's Discover Weekly providing recommendations for new songs and artists users might like based on user usage and traffic analysis, Quill's text generation offering that takes structured data and transforms it into narrative stories, and chatbots that provide real-time, context-specific answers to questions in a dialogue format.

[0139] Data is the heart of modern AI and deep learning algorithms. One problem that must be addressed before training can begin is the collection of labeled data, which is crucial for training accurate AI models. Full-scale AI deployments may require the continuous collection, cleaning, transformation, labeling, and storage of massive amounts of data. Adding additional high-quality data points directly leads to more accurate models and better insights. Data samples may be subjected to a series of processing steps, including but not limited to: 1) bringing data from external sources into the training system and storing the data in its raw form; 2) cleaning and transforming the data into a format convenient for training, including linking data samples to appropriate labels; 3) iterating to explore parameters and models, rapidly testing with smaller datasets, and converging on the most promising model to push to the production cluster; 4) running the training phase to select random batches of input data, including both new and older samples, and feeding them to production GPU servers for computation to update model parameters; and 5) evaluating, including using holdback portions of data not used in training to assess model accuracy against holdout data. This lifecycle can be applied not only to neural networks or deep learning, but to any type of parallelized machine learning. For example, a standard machine learning framework may rely on a CPU instead of a GPU, but the data ingestion and training workflows may be the same. Readers will understand that a single shared storage data hub creates a point of coordination throughout the entire lifecycle, without requiring extra data copies between the ingestion, preprocessing, and training phases. Ingested data is rarely used for only one purpose, and shared storage provides the flexibility to train multiple different models or apply traditional analytics to the data.

[0140] The reader will understand that each stage in an AI data pipeline may have varying requirements from the data hub (e.g., a storage system or a collection of storage systems). A scale-out storage system must provide uncompromising performance for all access types and patterns, from small, metadata-heavy files to large files, from random access patterns to sequential access patterns, and from low to high concurrency. The aforementioned storage system can function as an ideal AI data hub because it can serve unstructured workloads. In the first stage, data is ideally ingested and stored on the same data hub used by subsequent stages to avoid excessive data copying. The next two steps can optionally be performed on standard compute servers, including GPUs, and then in the fourth and final stage, the complete training production job runs on a powerful GPU-accelerated server. Often, a production pipeline exists alongside an experimental pipeline working on the same dataset. Furthermore, GPU-accelerated servers can be used independently for different models, or they can be combined to train on one larger model, even if they span multiple systems for distributed training. If the shared storage tier is slow, data must be copied to local storage at each phase, resulting in wasted time staging data to different servers. An ideal data hub for an AI training pipeline provides performance similar to data stored locally on the server nodes, while also possessing the simplicity and performance to allow all pipeline stages to operate simultaneously.

[0141] In some embodiments, for the storage system described above to function as a data hub or as part of an AI deployment, the storage system may be configured to provide DMA between the storage devices included in the storage system and one or more GPUs used in an AI or big data analytics pipeline. One or more GPUs are coupled to the storage system, for example, via an NVMe-over-Fabric ("NVMe-oF"), bypassing bottlenecks such as the host CPU, and allowing the storage system (or one of its components) to directly access GPU memory. In such an example, the storage system can leverage API hooks to the GPU to transfer data directly to the GPU. For example, the GPU may be embodied as an Nvidia® GPU, and the storage system supports GPUDirect Storage ("GDS") software, or has similar proprietary software, enabling the storage system to transfer data to the GPU via RDMA or a similar mechanism.

[0142] While the preceding paragraphs discuss deep learning applications, readers will understand that the storage systems described herein may also be part of a distributed deep learning ("DDL") platform to support the execution of DDL algorithms. The aforementioned storage systems may also be paired with other technologies, such as TensorFlow, an open-source software library for dataflow programming across a range of tasks that can be used in machine learning applications such as neural networks, to facilitate the development of such machine learning models, applications, etc.

[0143] The storage systems described above can also be used in neuromorphic computing environments. Neuromorphic computing is a form of computing that mimics brain cells. To support neuromorphic computing, an architecture of interconnected "neurons" replaces conventional computing models with low-power signals traveling directly between neurons for more efficient computation. Neuromorphic computing can utilize very-large-scale integration (VLSI) systems, including electronic analog circuits to mimic the neurobiological architecture present in the nervous system, as well as analog, digital, and mixed-mode analog / digital VLSIs, and software systems that implement models of the nervous system for perception, motor control, or multisensory integration.

[0144] The reader will understand that the storage systems described above may be configured to support the storage or use of blockchains and derived items (among other types of data), such as, for example, open-source blockchains and related tools that are part of the IBM® Hyperledger project, authorized blockchains where a certain number of trusted parties are permitted to access the blockchain, blockchain products that enable developers to build their own distributed ledger projects, and others. The blockchains and storage systems described herein may be utilized to support both on-chain and off-chain storage of data.

[0145] Off-chain storage of data can be implemented in various ways and can be done when the data itself is not stored within the blockchain. For example, in one embodiment, a hashing function can be used, and the data itself can be fed into the hashing function to generate a hash value. In such an example, the hash of a large data fragment may be embedded within the transaction instead of the data itself. Readers will understand that in other embodiments, alternatives to blockchain can be used to facilitate distributed storage of information. For example, one alternative to blockchain that can be used is blockweave. While traditional blockchains store all transactions to achieve verification, blockweave enables secure decentralization without using the entire chain, thereby enabling low-cost on-chain storage of data. Such blockweaves can utilize consensus mechanisms based on proof of access (PoA) and proof of work (PoW).

[0146] The storage systems described above can be used alone or in combination with other computing devices to support in-memory computing applications. In-memory computing involves storing information in RAM distributed across a cluster of computers. The reader will understand that the storage systems described above, in particular those configurable with customizable amounts of processing resources, storage resources, and memory resources (e.g., a system with blades containing configurable amounts of each type of resource), may be configured to provide infrastructure capable of supporting in-memory computing. Similarly, the storage systems described above may include component parts (e.g., NVDIMM, 3D crosspoint storage providing persistent high-speed random-access memory) that can actually provide an improved in-memory computing environment compared to an in-memory computing environment that relies on RAM distributed across dedicated servers.

[0147] In some embodiments, the storage system described above can be configured to operate as a hybrid in-memory computing environment that includes a universal interface to all storage media (e.g., RAM, flash storage, 3D crosspoint storage). In such embodiments, a user may not have knowledge of the details of where their data is stored, but can still interact with the data using the same complete and unified API. In such embodiments, the storage system can move data to the fastest available tier (in the background), and includes intelligently positioning the data according to various characteristics of the data or some other heuristic. In such examples, the storage system can even use existing products such as Apache Ignite and GridGain to move data between different storage tiers, or the storage system can use custom software to move data between different storage tiers. The storage systems described herein can implement various optimizations to improve the performance of in-memory computing, such as performing calculations as close to the data as possible.

[0148] The reader will further understand that in some embodiments, the storage systems described above may be paired with other resources to support the applications described above. For example, one infrastructure may include primary computing in the form of servers and workstations specializing in using general-purpose computing on graphics processing units, "GPGPU", to accelerate deep learning applications interconnected to computing engines for training parameters for deep neural networks. Each system may have Ethernet external connectivity, InfiniBand external connectivity, some other form of external connectivity, or some combination thereof. In such an example, the GPUs may be grouped together for a single large-scale training or used independently to train multiple models. The infrastructure may also include storage systems like those described above to provide a scale-out all-flash file or object store that can access data via high-performance protocols such as NFS, S3, etc. The infrastructure may also include redundant top-of-rack Ethernet switches connected to the storage and computers via ports in an MLAG port channel for redundancy, for example. The infrastructure may also include additional computers in the form of white-box servers with GPUs, optionally, for data acquisition, preprocessing, and model debugging. Readers will understand that additional infrastructure is also possible.

[0149] Readers will understand that the storage systems described above can be configured to support other AI-related tools, either independently or in conjunction with other computing machines. For example, a storage system could leverage tools such as ONXX or other open neural network exchange formats, which make it easier to transfer models written with different AI frameworks. Similarly, a storage system could be configured to support tools like Amazon's Gluon, which enables developers to prototype, build, and train deep learning models. In fact, the storage systems described above could be part of a larger platform, such as IBM Cloud Private for Data, which includes integrated data science, data engineering, and application building services.

[0150] Readers will further understand that the storage systems described above can also be deployed as edge solutions. Such edge solutions may be suitable for optimizing cloud computing systems by performing data processing at the edge of the network, close to the data source. Edge computing can push applications, data, and computing power (i.e., services) away from centralized points to the logical extrema of the network. Through the use of edge solutions such as the storage systems described above, computing tasks can be performed using the computing resources provided by such storage systems, data can be stored using the storage resources of the storage systems, and cloud-based services can be accessed through the use of various resources (including networking resources) of the storage systems. By performing computing tasks against edge solutions, storing data against edge solutions, and generally utilizing edge solutions, the consumption of expensive cloud-based resources can be avoided, and in fact, performance improvements can be experienced for heavier reliance on cloud-based resources.

[0151] While many tasks can benefit from the use of edge solutions, some specific uses may be particularly well-suited to deployment in such environments. For example, drones, autonomous vehicles, robots, and other devices may require extremely fast processing, and in practice, sending data to a cloud environment and sending it back for data processing support may simply be too slow. As an additional example, some IoT devices, such as connected video cameras, may not be well-suited to using cloud-based resources simply because the sheer volume of data involved makes sending data to the cloud impractical (not just from a privacy, security, or financial standpoint). Thus, many tasks related to data processing, storage, or communication can indeed be better suited by a platform that includes edge solutions such as the storage systems described above.

[0152] The storage systems described above can function as network edge platforms, combining computing resources, storage resources, networking resources, cloud technologies, and network virtualization technologies, either alone or in combination with other computing resources. As part of the network, the edge can have similar characteristics to other network facilities, from customer premises and backhaul aggregation facilities to points of presence (PoPs) and regional data centers. Readers will understand that network workloads such as virtual network functions (VNFs) and others reside on the network edge platform. Network edge platforms enabled by a combination of containers and virtual machines may rely on controllers and schedulers that are no longer geographically located with the data processing resources. Functionality as microservices can be divided into control planes, user and data planes, or even state machines, allowing for the application of independent optimization and scaling techniques. Such user and data planes can be enabled through both increased accelerators present in server platforms, such as FPGAs and smart NICs, and SDN-enabled merchant silicon and programmable ASICs.

[0153] The storage systems described above may also be optimized for use in big data analytics, and a containerized analytics architecture may be leveraged as part of a configurable data analytics pipeline, for example, to make analytical capabilities more configurable. Big data analytics can generally be described as the process of exploring large and diverse datasets to uncover hidden patterns, unknown correlations, market trends, customer preferences, and other useful information that can help organizations make more informed business decisions. As part of that process, semi-structured and unstructured data, such as internet clickstream data, web server logs, social media content, text from customer emails and survey responses, mobile phone call detail records, IoT sensor data, and other data, may be transformed into structured forms.

[0154] The storage system described above may also support (including implementation as a system interface) applications that perform tasks in response to human speech. For example, the storage system may support the execution of intelligent personal assistant applications such as Amazon Alexa®, Apple Siri®, Google Voice®, Samsung Bixby®, Microsoft Cortana®, and others. While the examples described in the preceding paragraph utilize voice as input, the storage system described above may also support chatbots, talkbots, chattabots, or other artificial conversational entities, or other applications configured to converse via auditory or textual means. Similarly, the storage system may actually run such applications to enable users, such as system administrators, to interact with the storage system via voice. Such applications may generally enable voice interaction, music playback, to-do list creation, alarm setting, podcast streaming, audiobook playback, and the provision of other real-time information such as weather, traffic, and news, but in embodiments of this disclosure, such applications may be used as interfaces to various system management operations.

[0155] The aforementioned storage systems can also implement an AI platform for delivery in the vision of autonomous storage. Such an AI platform may be configured to provide global predictive intelligence by collecting and analyzing a large amount of storage system telemetry data points, enabling easy management, analysis, and support. In fact, such a storage system may be capable of predicting both capacity and performance, as well as generating intelligent advice regarding workload deployment, interaction, and optimization. Such an AI platform may be configured to scan all incoming storage system telemetry data against a library of fingerprints of the problem in order to predict and resolve incidents in real time before they affect the customer environment, and may capture hundreds of performance-related variables used to predict performance loads.

[0156] The storage system described above can support the serialization or concurrent execution of artificial intelligence applications, machine learning applications, data analysis applications, data transformations, and other tasks, which can collectively form an AI ladder. Such an AI ladder can be effectively formed by combining these elements to create a complete data science pipeline with dependencies between the elements of the AI ​​ladder. For example, AI may require some form of machine learning to have been performed, machine learning may require some form of analysis to have been performed, and analysis may require some form of data and information construction to have been performed. Thus, each element can be considered a lang in an AI ladder, which can collectively form a complete and sophisticated AI solution.

[0157] The storage systems described above can also be used, either alone or in combination with other computing environments, to deliver AI to any experience in which AI permeates the broad and expansive aspects of business and life. For example, AI may play a crucial role in the delivery of deep learning solutions, deep reinforcement learning solutions, general-purpose artificial intelligence solutions, autonomous vehicles, cognitive computing solutions, commercial UAVs or drones, conversational user interfaces, enterprise classification, ontology management solutions, machine learning solutions, smart dust, smart robots, and smart workplaces.

[0158] The storage systems described above may also be used, either alone or in combination with other computing environments, to provide a wide range of transparent and immersive experiences (including those using digital twins of various “things” such as people, places, processes, and systems) in which the technology can introduce transparency between people, businesses, and things. Such transparent and immersive experiences may be provided as augmented reality technology, connected homes, virtual reality technology, brain-computer interfaces, human augmentation technology, nanotube electronics, volumetric displays, 4D printing technology, or others.

[0159] The storage systems described above can also be used alone or in combination with other computing environments to support a wide variety of digital platforms. Such digital platforms may include, for example, 5G wireless systems and platforms, digital twin platforms, edge computing platforms, IoT platforms, quantum computing platforms, serverless PaaS, software-defined security, and neuromorphic computing platforms.

[0160] The storage system described above may also be part of a multi-cloud environment where multiple cloud computing and storage services are deployed on a single heterogeneous architecture. To facilitate the operation of such a multi-cloud environment, DevOps tools can be deployed to enable orchestration across clouds. Similarly, continuous development and continuous integration tools can be deployed to standardize processes related to continuous integration and delivery, new feature rollouts, and cloud workload provisioning. By standardizing these processes, a multi-cloud strategy can be implemented that enables the use of the best provider for each workload.

[0161] The storage system described above can be used as part of a platform that enables the use of cryptographic anchors, which can be used to authenticate the origin and content of a product and ensure that it matches the blockchain record associated with the product. Similarly, as part of a set of tools for securing data stored on the storage system, the storage system described above can implement various cryptographic techniques and schemes, including lattice cryptography. Lattice cryptography can involve the structure of cryptographic primitives that include a lattice, either in the structure itself or in the security proof. Unlike public-key cryptography systems such as RSA, Diffie-Hellman, or Elliptic-Curve, which are easily attacked by quantum computers, some lattice-based configurations are considered resistant to attacks by both classical and quantum computers.

[0162] A quantum computer is a device that performs quantum computing. Quantum computing is computation using quantum mechanical phenomena such as superposition and entanglement. Quantum computers differ from conventional computers based on transistors. This is because such conventional computers require data to be encoded into binary numbers (bits), each of which is always one of two distinct states (0 or 1). In contrast to conventional computers, quantum computers use qubits, which can be in superposition of states. A quantum computer maintains a sequence of qubits, and a single qubit can represent 1, 0, or any quantum superposition of any two of those qubit states. A pair of qubits can be in any quantum superposition of four states, and three qubits can be in any superposition of eight states. A quantum computer with n qubits can generally be in any superposition of up to 2^n different states simultaneously, whereas a conventional computer can only be in one of these states at any given time. A quantum Turing machine is a theoretical model of such a computer.

[0163] The storage systems described above can be paired with FPGA acceleration servers as part of a larger AI or ML infrastructure. Such FPGA acceleration servers may reside near the storage systems (e.g., within the same data center) or may be incorporated into an appliance that includes one or more storage systems, one or more FPGA acceleration servers, networking infrastructure supporting communication between the one or more storage systems and the one or more FPGA acceleration servers, and other hardware and software components. Alternatively, FPGA acceleration servers may reside in a cloud computing environment that can be used to perform computation-related tasks for AI and ML jobs. Any of the embodiments described above can be used collectively to function as an FPGA-based AI or ML platform. Readers will understand that in some embodiments of the FPGA-based AI or ML platform, the FPGAs contained within the FPGA acceleration server can be reconfigured for different types of ML models (e.g., LSTM, CNN, GRU). The ability to reconfigure the FPGAs contained within the FPGA acceleration server can enable acceleration of ML or AI applications based on the most optimal numerical precision and memory model being used. Readers will understand that by treating a collection of FPGA-accelerated servers as a pool of FPGAs, any CPU in a data center can utilize the pool of FPGAs as a shared hardware microservice, rather than limiting the servers to dedicated accelerators plugged into them.

[0164] The FPGA and GPU acceleration servers described above can implement computing models in which, instead of holding a small amount of data within machine learning and executing long streams of instructions on top of it, as is done in conventional computing models, CPU models and parameters are pinned to high-bandwidth on-chip memory, and a large amount of data is streamed through that high-bandwidth on-chip memory. FPGAs can be even more efficient than GPUs for this type of computing model because they can be programmed with only the instructions necessary to execute it.

[0165] The storage system described above can be configured to provide parallel storage, for example, through the use of a parallel file system such as BeeGFS. Such a parallel file system may include a distributed metadata architecture. For example, the parallel file system may include components that include multiple metadata servers on which metadata is distributed, as well as services for clients and storage servers.

[0166] The system described above can support the execution of a variety of software applications. Such software applications can be deployed in various ways, including container-based deployment models. Containerized applications can be managed using various tools. For example, containerized applications may be managed using Docker Swarm, Kubernetes, and others. Containerized applications can be used to facilitate serverless cloud-native computing deployment and management models for software applications. To support serverless cloud-native computing deployment and management models for software applications, containers can be used as part of an event processing mechanism (e.g., AWS Lambda) so that various events cause the containerized application to spin up and act as an event handler.

[0167] The systems described above can be deployed in various ways, including in ways that support fifth-generation ("5G") networks. 5G networks can support substantially faster data communication than previous generations of mobile communication networks, potentially leading to a decentralization of data and computing resources, as modern large-scale data centers may become less prominent, for example, being replaced by more local micro-data centers closer to mobile network towers. The systems described above may be included in such local micro-data centers, or may be part of or paired with a multi-access edge computing ("MEC") system. Such MEC systems can enable cloud computing capabilities and IT service environments at the edge of cellular networks. By running applications closer to cellular customers and performing associated processing tasks, network congestion can be reduced, and applications can run more smoothly.

[0168] The storage system described above can be configured to implement NVMe zoned namespaces. Through the use of NVMe zoned namespaces, the logical address space of a namespace is divided into zones. Each zone provides a logical block address range that must be written sequentially and explicitly reset before rewriting, thereby enabling the creation of namespaces that expose the natural boundaries of the device and the offloading management of internal mapping tables to the host. To implement NVMe zoned namespaces ("ZNS"), ZNS SSDs or other forms of zoned block devices can be used to expose the namespace logical address space using zones. When zones are aligned with the internal physical characteristics of the device, several inefficiencies in data placement can be eliminated. In such embodiments, each zone can be mapped to a separate application so that functions such as wear leveling and garbage collection can be performed per zone or per application rather than across the entire device. To support ZNS, the storage controllers described herein can be configured to interact with zoned block devices, for example, through the use of the Linux® kernel zoned block device interface or other tools.

[0169] The storage systems described above may also be configured to implement zoning storage in other ways, such as through the use of shingled magnetic recording (SMR) storage devices. In examples where zoning storage is used, a device-managed embodiment may be deployed, where the storage device manages it within its firmware and hides this complexity by presenting an interface similar to any other storage device. Alternatively, zoning storage can also be implemented through a host-managed embodiment that relies on the operating system to know how to handle the drives and sequentially writes only to specific areas of the drives. Zoning storage can also be implemented using a host-aware embodiment in which a combination of drive-managed and host-managed implementations is deployed.

[0170] The storage systems described herein may be used to form a data lake. The data lake can act as the initial location where an organization's data flows, and such data may be in raw format. Metadata tagging may be implemented to facilitate the retrieval of data elements within the data lake, particularly in embodiments where the data lake includes multiple stores of data in formats that are not readily accessible or readable (e.g., unstructured data, semi-structured data, structured data). From the data lake, data can move downstream to a data warehouse, where the data may be more processed, packaged, and stored in a more consumable format. The storage systems described above may also be used to implement such a data warehouse. In addition, a data mart or data hub can enable even more readily consumed data, and the storage systems described above can be used to provide the necessary underlying storage resources for the data mart or data hub. In embodiments, queries to the data lake may require schema-on-read techniques, which are applied to the plan or schema when the data is retrieved from a stored location, rather than when the plan or schema enters the stored location.

[0171] The storage systems described herein may also be configured to implement a recovery point objective ("RPO"), which may be established by the user, by an administrator, as a system default, as part of a storage class or service in which the storage system participates in distribution, or by any other means. The "recovery point objective" is the target maximum time difference between the last update to the source dataset and the last recoverable replicated dataset update that is correctly recoverable from a continuously or frequently updated copy of the source dataset, given a reason to do so. The update is correctly recoverable if all updates processed on the source dataset prior to the last recoverable replicated dataset update are appropriately taken into account.

[0172] In synchronous replication, the RPO is zero, meaning that under normal operation, all completed updates on the source dataset should exist and be correctly recoverable on the copy dataset. In best-effort, near-synchronous replication, the RPO can be as low as a few seconds. In snapshot-based replication, the RPO can be roughly calculated as the interval between snapshots plus the time required to transfer fixes between the previous already transferred snapshot and the most recent snapshot being replicated.

[0173] If updates accumulate faster than they can be replicated, the RPO can be missed. In the case of snapshot-based replication, if more data to be replicated accumulates between two snapshots than can be replicated between the time a snapshot is taken and the cumulative updates of that snapshot are replicated to a copy, the RPO can be deviated. Again, in snapshot-based replication, if the data to be replicated accumulates faster than can be transferred in the time between subsequent snapshots, replication can begin to escalate the gap between the expected target recovery point and the actual recovery point represented by the last correctly replicated update, causing further delays.

[0174] The storage systems described above may be part of a shared-nothing storage cluster. In a shared-nothing storage cluster, each node in the cluster has local storage and communicates with other nodes in the cluster via a network, and the storage used by the cluster is (generally) provided only by the storage connected to each individual node. A set of nodes synchronously replicating a dataset can be an example of a local storage cluster, as each storage system has shared-nothing and communicates with other storage systems via a network, and these storage systems (generally) do not use other storage that shares access through some kind of interconnection. In contrast, some of the storage systems described above are built as shared storage clusters because there are drive shelves shared by paired controllers. However, the other storage systems described above are built as shared-nothing storage clusters because all storage is local to a particular node (e.g., a blade), and all communication is via a network that links the compute nodes together.

[0175] In other embodiments, other forms of shared-nothing storage clusters may include embodiments in which any node in the cluster has a local copy of all the storage it needs, and data is mirrored to other nodes in the cluster via synchronous replication to ensure that no data is lost, or because other nodes are also using that storage. In such an embodiment, if a new cluster node needs any data, that data can be copied to the new node from another node that has a copy of that data.

[0176] In some embodiments, a mirror-copy-based shared storage cluster can store multiple copies of the stored data across all clusters, with each subset of data replicated to a specific set of nodes, and different subsets of data replicated to different sets of nodes. In some modifications, the embodiment may store all of the cluster's stored data on all nodes, while in other modifications, the nodes may be divided such that a first set of nodes all store the same set of data, and a second set of different nodes all store different sets of data.

[0177] Readers will understand that a RAFT-based database (e.g., etcd) can function like a non-shared storage cluster where all RAFT nodes store all data. However, the amount of data stored in a RAFT cluster can be limited so that extra copies do not consume too much storage. A container server cluster can also replicate all data to all cluster nodes, assuming that containers do not tend to be too large and their bulk data (data manipulated by applications running within the containers) is stored elsewhere, such as in an S3 cluster or an external file server. In such an example, container storage may be provided directly by the cluster via its non-shared storage model, and those containers provide images that form the execution environment for parts of an application or service.

[0178] For further explanation, Figure 3D illustrates an exemplary computing device 350 that may be specifically configured to perform one or more of the processes described herein. As shown in Figure 3D, the computing device 350 may include a communication interface 352, a processor 354, a storage device 356, and an input / output ("I / O") module 358, all connected to each other communicably via a communication infrastructure 360. While an exemplary computing device 350 is shown in Figure 3D, the components illustrated in Figure 3D are not intended to be limiting. Additional or alternative components may be used in other embodiments. The components of the computing device 350 shown in Figure 3D will now be described in more detail.

[0179] The communication interface 352 may be configured to communicate with one or more computing devices. Examples of the communication interface 352 include, but are not limited to, wired network interface connections (such as network interface cards), wireless network interface connections (such as wireless network interface card connections), modems, audio / video connections, and any other suitable interfaces.

[0180] The processor 354 generally represents any type or form of processing unit capable of processing data and / or interpreting, executing, and / or directing the execution of one or more instructions, processes, and / or operations as described herein. The processor 354 can perform operations by executing computer executable instructions 362 (e.g., applications, software, code, and / or other executable data instances) stored in the storage device 356.

[0181] The storage device 356 may include one or more data storage media, devices, or configurations, and any type, form, and combination of data storage media and / or devices may be used. For example, the storage device 356 may include, but is not limited to, any combination of the non-volatile media and / or volatile media described herein. Electronic data, including the data described herein, may be stored temporarily and / or permanently in the storage device 356. For example, data representing computer executable instructions 362 configured to instruct the processor 354 to perform any of the operations described herein may be stored in the storage device 356. In some examples, the data may be located in one or more databases residing within the storage device 356.

[0182] The I / O module 358 may include one or more I / O modules configured to receive user input and provide user output. The I / O module 358 may include any hardware, firmware, software, or combination thereof that supports input and output capabilities. For example, the I / O module 358 may include, but is not limited to, hardware and / or software for capturing user input, including a keyboard or keypad, a touchscreen component (e.g., a touchscreen display), a receiver (e.g., an RF or infrared receiver), a motion sensor, and / or one or more input buttons.

[0183] The I / O module 358 may include, but is not limited to, one or more devices for presenting output to the user, including a graphics engine, a display (e.g., a display screen), one or more output drivers (e.g., display drivers), one or more audio speakers, and one or more audio drivers. In certain embodiments, the I / O module 358 is configured to provide graphical data to the display for presentation to the user. The graphical data may represent one or more graphical user interfaces and / or any other graphical content that may be useful for a particular implementation. In some examples, any of the systems, computing devices, and / or other components described herein may be implemented by the computing device 350.

[0184] For further explanation, Figure 3E illustrates an example of a fleet of storage systems 376 for providing storage services (also referred to herein as “data services”). The fleet of storage systems 376 depicted in Figure 3 includes a plurality of storage systems 374a, 374b, 374c, 374d, and 374n, each of which may be similar to the storage systems described herein. The storage systems 374a, 374b, 374c, 374d, and 374n within the fleet of storage systems 376 can be embodied as the same storage system or as different types of storage systems. For example, two of the storage systems 374a and 374n depicted in Figure 3E are depicted as cloud-based storage systems because the resources that collectively form each of the storage systems 374a and 374n are provided by separate cloud service providers 370 and 372. For example, the first cloud service provider 370 may be Amazon AWS®, while the second cloud service provider 372 may be Microsoft Azure®. In other embodiments, one or more public clouds, private clouds, or combinations thereof may be used to provide the underlying resources used to form a particular storage system within a fleet of storage systems 376.

[0185] An example depicted in Figure 3E includes an edge management service 366 for delivering storage services, according to some embodiments of the present disclosure. The delivered storage services (also referred to herein as “data services”) may include, for example, a service that provides a certain amount of storage to a consumer, a service that provides storage to a consumer in accordance with a predetermined quality of service guarantee, a service that provides storage to a consumer in accordance with a predetermined regulatory requirement, and many other services.

[0186] The edge management service 366 depicted in Figure 3E may be embodied, for example, as one or more modules of computer program instructions that run on computer hardware such as one or more computer processors. Alternatively, the edge management service 366 may be embodied as one or more modules of computer program instructions that run in one or more containers or in some other way on a virtualization execution environment such as one or more virtual machines. In other embodiments, the edge management service 366 may be embodied as a combination of the above embodiments, including embodiments in which one or more modules of computer program instructions included in the edge management service 366 are distributed across multiple physical or virtual execution environments.

[0187] The edge management service 366 may act as a gateway for providing storage services to storage consumers, which leverage storage provided by one or more storage systems 374a, 374b, 374c, 374d, 374n. For example, the edge management service 366 may be configured to provide storage services to host devices 378a, 378b, 378c, 378d, 378n running one or more applications that consume storage services. In such an example, the edge management service 366 can act as a gateway between the host devices 378a, 378b, 378c, 378d, 378n and the storage systems 374a, 374b, 374c, 374d, 374n, rather than requiring the host devices 378a, 378b, 378c, 378d, 378n to directly access the storage systems 374a, 374b, 374c, 374d, 374n.

[0188] In Figure 3E, the edge management service 366 exposes the storage service module 364 to the host devices 378a, 378b, 378c, 378d, and 378n in Figure 3E. In other embodiments, the edge management service 366 may expose the storage service module 364 to other consumers of various storage services. Various storage services can be presented to consumers through one or more user interfaces, through one or more APIs, or through some other mechanisms provided by the storage service module 364. Thus, the storage service module 364 depicted in Figure 3E may be embodied as one or more modules of computer program instructions executed on physical hardware, in a virtualized execution environment, or a combination thereof. Executing such modules enables consumers of storage services to be provided with, select, and access various storage services.

[0189] The edge management service 366 in Figure 3E also includes a system management service module 368. The system management service module 368 in Figure 3E includes one or more modules of computer program instructions that, when executed, perform various operations in cooperation with storage systems 374a, 374b, 374c, 374d, and 374n in order to provide storage services to host devices 378a, 378b, 378c, 378d, and 378n. The system management service module 368 can be configured to perform tasks such as provisioning storage resources from storage systems 374a, 374b, 374c, 374d, and 374n via one or more APIs exposed by storage systems 374a, 374b, 374c, 374d, and 374n; migrating datasets or workloads between storage systems 374a, 374b, 374c, 374d, and 374n via one or more APIs exposed by storage systems 374a, 374b, 374c, 374d, and 374n; and setting one or more tunable parameters (i.e., one or more configurable settings) on storage systems 374a, 374b, 374c, 374d, and 374n via one or more APIs exposed by storage systems 374a, 374b, 374c, 374d, and 374n. For example, many of the services described below relate to embodiments in which storage systems 374a, 374b, 374c, 374d, and 374n are configured to operate in some way. In such examples, the system management service module 368 may be responsible for configuring storage systems 374a, 374b, 374c, 374d, and 374n to operate in the manner described below, using APIs (or some other mechanism) provided by the storage systems 374a, 374b, 374c, 374d, and 374n.

[0190] In addition to configuring storage systems 374a, 374b, 374c, 374d, and 374n, the edge management service 366 itself may be configured to perform various tasks required to provide various storage services. Consider an example where, when selected and applied, the storage service includes a service that obfuscates personally identifiable information ("PII") contained in a dataset when the dataset is accessed. In such an example, storage systems 374a, 374b, 374c, 374d, and 374n may be configured to obfuscate the PII when servicing read requests directed to the dataset. Alternatively, storage systems 374a, 374b, 374c, 374d, and 374n can service reads by returning data containing PII, but the edge management service 366 itself can obfuscate the PII as the data passes through the edge management service 366 on its way from storage systems 374a, 374b, 374c, 374d, and 374n to host devices 378a, 378b, 378c, 378d, and 378n.

[0191] The storage systems 374a, 374b, 374c, 374d, and 374n depicted in Figure 3E can be embodied as one or more of the storage systems (including their variations) described above with reference to Figures 1A to 3D. In practice, the storage systems 374a, 374b, 374c, 374d, and 374n can function as a pool of storage resources, and the individual components within that pool may have different performance characteristics, different storage characteristics, etc. For example, one of the storage systems 374a may be a cloud-based storage system, another storage system 374b may be a storage system that provides block storage, another storage system 374c may be a storage system that provides file storage, another storage system 374d may be a relatively high-performance storage system, while another storage system 374n may be a relatively low-performance storage system, and so on. In an alternative embodiment, only a single storage system may exist.

[0192] The storage systems 374a, 374b, 374c, 374d, and 374n depicted in Figure 3E can also be organized into different failure domains such that a failure in one storage system 374a is completely independent of a failure in another storage system 374b. For example, each storage system can receive power from an independent power system, and each storage system can be connected for data communication via an independent data communication network. Furthermore, storage systems in a first failure domain can be accessed via a first gateway, while storage systems in a second failure domain can be accessed via a second gateway. For example, the first gateway may be a first instance of edge management service 366, and the second gateway may be a second instance of edge management service 366, including embodiments in which each instance is separate or each instance is part of a distributed edge management service 366.

[0193] As an illustrative example of available storage services, a user may be presented with storage services associated with different levels of data protection. For example, a user may be presented with a storage service that assures the user that their associated data will be protected so that, when selected and enforced, various recovery point objectives ("RPO") can be guaranteed. A first available storage service may ensure, for example, that some datasets associated with the user are protected so that any data older than 5 seconds can be recovered in the event of a primary datastore failure, while a second available storage service may ensure that datasets associated with the user are protected so that any data older than 5 minutes can be recovered in the event of a primary datastore failure.

[0194] Additional examples of storage services that may be presented to a user, selected by the user, and ultimately applied to a dataset associated with the user may include one or more data compliance services. Such data compliance services may be embodied as services that can be provided to the consumer of data compliance services (i.e., the user) to ensure, for example, that the user's dataset is managed to comply with various regulatory requirements. For example, one or more data compliance services may be provided to the user to ensure that the user's dataset is managed in a manner that complies with the General Data Protection Regulation ("GDPR"), or one or more data compliance services may be provided to the user to ensure that the user's dataset is managed in a manner that complies with the Sarbanes-Oxley Act of 2002 ("SOX"), or one or more data compliance services may be provided to the user to ensure that the user's dataset is managed in a manner that complies with any other regulatory conduct. In addition, one or more data compliance services may be provided to the user to ensure that the user's dataset is managed to comply with certain non-governmental guidance (for example, to comply with best practices for auditing purposes), and one or more data compliance services may be provided to the user to ensure that the user's dataset is managed to comply with the requirements of a particular client or organization, and so on.

[0195] To provide this specific data compliance service, the data compliance service may be presented to the user (e.g., via a GUI) and selected by the user. In response to receiving a selection of a specific data compliance service, one or more storage service policies may be applied to the dataset associated with the user in order to perform the specific data compliance service. For example, a storage service policy may be applied that requires the dataset to be encrypted before it is stored in a storage system, before it is stored in a cloud environment, or before it is stored elsewhere. To enforce this policy, a requirement may be introduced that the dataset must be encrypted before it is transmitted (e.g., before it is sent to another party). In such an example, a storage service policy may also be introduced that requires that any encryption key used to encrypt the dataset is not stored on the same system that stores the dataset itself. The reader will understand that many other forms of data compliance services may be provided and implemented by embodiments of this disclosure.

[0196] The storage systems 374a, 374b, 374c, 374d, and 374n within the fleet of storage system 376 may be collectively managed, for example, by one or more fleet management modules. The fleet management modules may be part of the system management service module 368 depicted in Figure 3E, or they may be separate from it. The fleet management modules can perform tasks such as monitoring the health of each storage system in the fleet, initiating updates or upgrades on one or more storage systems in the fleet, migrating workloads for load balancing or other performance purposes, and many other tasks. Therefore, and for many other reasons, the storage systems 374a, 374b, 374c, 374d, and 374n may be coupled to one another via one or more data communication links to exchange data among them.

[0197] In some embodiments, one or more storage systems or elements of a storage system (e.g., features, services, operations, components, etc.) such as any of the exemplary storage systems or storage system elements described herein may be implemented in one or more container systems. A container system may include any system that supports the execution of one or more containerized applications or services. Such services may be software deployed as infrastructure for building applications, for running runtime environments, and / or as infrastructure for other services. In the following description, the description of containerized applications generally also applies to containerized services.

[0198] A container can combine one or more elements of a containerized software application with a runtime environment for running those elements of the software application, bundled into a single image. For example, each such container of a containerized application may contain the executable code and various dependencies, libraries, and / or other components of the software application, along with configured access to network configurations and additional resources, which are used by the elements of the software application within that particular container to enable the operation of those elements. A containerized application can be represented as a collection of such containers, each representing all elements of the application together with the various runtime environments required for all those elements to run. As a result, a containerized application may be abstracted from the host operating system as a collection of lightweight and portable packages and configurations, and may be uniformly deployed and run consistently in different computing environments using different container-compatible operating systems or different infrastructures. In some embodiments, a containerized application shares a kernel with the host computer system and runs as an isolated environment (an isolated collection of files and directories, processes, system and network resources, and configured access to additional resources and capabilities) isolated by the host system's operating system in conjunction with a container management framework. When executed, a containerized application can provide one or more containerized workloads and / or services.

[0199] A container system may include and / or utilize a cluster of nodes. For example, a container system may be configured to manage the deployment and execution of a containerized application on one or more nodes in a cluster. A containerized application may utilize node resources such as memory, processing, and / or storage resources provided and / or accessed by the node. Storage resources may include any of the exemplary storage resources described herein, and may include on-node resources such as a local tree of files and directories, off-node resources such as an external networked file system, a database, or an object store, or both on-node and off-node resources. Access to additional resources and capabilities that can be configured for a containerized application's container may include dedicated computing power such as GPUs and AI / ML engines, or dedicated hardware such as sensors and cameras.

[0200] In some embodiments, the container system may include a container orchestration system (which may also be referred to as a container orchestrator, container orchestration platform, etc.) that is designed to be reasonably simple and automated to deploy, scale, and manage containerized applications for a wide range of use cases. In some embodiments, the container system may also include a storage management system configured to provision and manage storage resources (e.g., virtual volumes) for private or shared use by cluster nodes and / or containers of the containerized application.

[0201] Figure 3F illustrates an exemplary container system 380. In this example, the container system 380 includes a container storage system 381 which may be configured to perform one or more storage management operations to organize, provision, and manage storage resources for use by one or more containerized applications 382-1 to 382-L of the container system 380. In particular, the container storage system 381 may organize storage resources into one or more storage pools 383 of storage resources for use by the containerized applications 382-1 to 382-L. The container storage system itself may be implemented as a containerized service.

[0202] Container system 380 may include, or be implemented by, one or more container orchestration systems, including, in particular, Kubernetes®, Mesos®, and Docker Swarm®. The container orchestration system may manage container system 380 running on cluster 384 through services implemented by a control node, represented as 385, and may further manage the container storage system, or the relationships between individual containers and their storage, memory, and CPU limits, networking, and their access to additional resources or services.

[0203] The control plane of the container system 380 may implement services including deploying applications via the controller 386, monitoring applications via the controller 386, providing interfaces via the API server 387, and scheduling deployments via the scheduler 388. In this example, the controller 386, scheduler 388, API server 387, and container storage system 381 are implemented on a single node, node 385. In other examples, for resilience, the control plane may be implemented by multiple redundant nodes, so that if a node providing management services to the container system 380 fails, another redundant node may provide management services to the cluster 384.

[0204] The data plane of container system 380 may include a set of nodes that provide container runtimes for running containerized applications. Individual nodes in cluster 384 may run container runtimes such as Docker® and container managers or node agents such as kubelet (not shown) in Kubernetes, which communicate with the control plane via a local network connectivity agent (sometimes called a proxy), such as agent 389. Agent 389 can route network traffic to and from containers, for example, using Internet Protocol (IP) port numbers. For example, a containerized application may request storage classes from the control plane, the request is handled by the container manager, and the container manager communicates the request to the control plane using agent 389.

[0205] Cluster 384 may include a set of nodes running containers for managed containerized applications. Nodes may be virtual or physical machines. Nodes may also be host systems.

[0206] The container storage system 381 can organize storage resources to provide storage to the container system 380. For example, the container storage system 381 can use the storage pool 383 to provide persistent storage to containerized applications 382-1 to 382-L. The container storage system 381 itself may be deployed as a containerized application by a container orchestration system.

[0207] For example, a container storage system 381 application can be deployed within a cluster 384 and perform management functions to provide storage to a containerized application 382. These management functions may include determining one or more storage pools from available storage resources, provisioning virtual volumes on one or more nodes, replicating data, responding to and recovering from host and network failures, or handling storage operations. A storage pool 383 may contain storage resources from one or more local or remote sources, and these storage resources may be different types of storage, including, for example, block storage, file storage, and object storage.

[0208] The container storage system 381 may also be deployed on a set of nodes where persistent storage can be provided by a container orchestration system. In some examples, the container storage system 381 may be deployed on all nodes in cluster 384, for example, using a Kubernetes DaemonSet. In this example, nodes 390-1 to 390-N provide the container runtime that the container storage system 381 runs on. In other examples, some, but not all, nodes in the cluster may run the container storage system 381.

[0209] The container storage system 381 can handle storage on a node and communicate with the control plane of the container system 380 to provide dynamic volumes, including persistent volumes. Persistent volumes may be mounted on the node as virtual volumes, such as virtual volumes 391-1 and 391-P. After virtual volume 391 is mounted, a containerized application may request, use, or be configured to use the storage provided by virtual volume 391. In this example, the container storage system 381 may have a driver installed on the node's kernel that handles storage operations directed to the virtual volumes. In this example, the driver may receive storage operations directed to the virtual volumes and, in response, may perform storage operations on one or more storage resources in the storage pool 383, either under instructions from additional logic within a container that implements the container storage system 381 as a containerized service, or using additional logic.

[0210] The container storage system 381 can determine available storage resources in response to being deployed as a containerized service. For example, storage resources 392-1 to 392-M may include local storage, remote storage (storage on separate nodes in the cluster), or both local and remote storage. Storage resources may also include storage from external sources, such as various combinations of block storage systems, file storage systems, and object storage systems. Storage resources 392-1 to 392-M may include storage resources of any type (and possibly more) and / or configuration (and possibly more) (e.g., any of the exemplary storage resources described above), and the container storage system 381 may be configured to determine available storage resources in any preferred manner, including based on a configuration file. For example, the configuration file may specify account and authentication information for cloud-based object storage 348 or cloud-based storage system 318. The container storage system 381 may also determine the availability of one or more storage devices 356 or one or more storage systems. The total amount of storage from storage devices 356, storage systems 348, cloud-based storage systems 318, edge management services 366, cloud-based object storage 348, or any other storage resources, or any combination or partial combination of any one or more of such storage resources, may be used to provide the storage pool 383. The storage pool 383 is used to provision storage for one or more virtual volumes mounted on one or more nodes 390 in the cluster 384.

[0211] In some implementations, the container storage system 381 can create multiple storage pools. For example, the container storage system 381 can aggregate storage resources of the same type into individual storage pools. In this example, the storage type may be one of the following: storage device 356, storage array 102, cloud-based storage system 318, storage via edge management service 366, or cloud-based object storage 348. Alternatively, it may be storage configured with a certain level or type of redundancy or distribution, such as a specific combination of striping, mirroring, or erase coding.

[0212] The container storage system 381 can run within cluster 384 as a containerized container storage system service, and instances of containers implementing elements of the containerized container storage system service can run on different nodes within cluster 384. In this example, the containerized container storage system service works with the container orchestration system of container system 380 to handle storage operations, mount virtual volumes to provide storage to nodes, aggregate available storage into storage pool 383, provision storage from storage pool 383 to virtual volumes, generate backup data, and replicate data between nodes, clusters, and environments, among other storage system operations. In some examples, the containerized container storage system service can provide storage services across multiple clusters running in separate computing environments. For example, other storage system operations may include the storage system operations described herein. The persistent storage provided by the containerized container storage system service can be used to implement stateful and / or resilient containerized applications.

[0213] The container storage system 381 may be configured to perform any appropriate storage operations of the storage system. For example, the container storage system 381 may be configured to perform one or more of the exemplary storage management operations described herein in order to manage the storage resources used by the container system.

[0214] In some embodiments, one or more storage operations, including one or more of the exemplary storage management operations described herein, may be containerized. For example, one or more storage operations may be implemented as one or more containerized applications configured to run in order to perform the storage operations. Such containerized storage operations may run in any suitable runtime environment to manage any storage system(s), including any of the exemplary storage systems described herein.

[0215] The storage systems described herein can support various forms of data replication. For example, two or more storage systems may synchronously replicate a dataset between them. In synchronous replication, separate copies of a particular dataset may be maintained by multiple storage systems, but all access to the dataset (e.g., reads) should yield consistent results regardless of which storage system the access is directed to. For example, a read directed to any of the storage systems synchronously replicating the dataset must return the same result. Therefore, updates to versions of the dataset do not need to be performed exactly simultaneously, but precautions must be taken to ensure consistent access to the dataset. For example, if an update (e.g., a write) directed to a dataset is received by the first storage system, the update can only be acknowledged as complete if all storage systems synchronously replicating the dataset have applied the update to their copies of the dataset. In such an example, synchronous replication can be performed through the use of I / O transfers (e.g., a write received by the first storage system is transferred to the second storage system), communication between storage systems (e.g., each storage system indicating that the update is complete), or by other means.

[0216] In other embodiments, a dataset may be replicated through the use of checkpoints. In checkpoint-based replication (also referred to as “nearly synchronized replication”), a set of updates to a dataset (e.g., one or more write operations directed to a dataset) may occur between different checkpoints, such that the dataset is updated to a particular checkpoint only when all updates to the dataset prior to that particular checkpoint have been completed. Consider an example where a first storage system stores a live copy of a dataset that is accessed by a user of the dataset. In this example, assume that the dataset is replicated from the first storage system to a second storage system using checkpoint-based replication. For example, the first storage system may send a first checkpoint (time t=0) to the second storage system, then send a first set of updates to the dataset, then send a second checkpoint (time t=1), then send a second set of updates to the dataset, and then send a third checkpoint (time t=2). In such an example, if the second storage system has performed all updates in the first set of updates but has not yet performed all updates in the second set of updates, the copy of the dataset stored on the second storage system may be up to date up to the second checkpoint. Alternatively, if the second storage system has performed all updates in both the first and second sets of updates, the copy of the dataset stored on the second storage system may be up to date up to the third checkpoint. Readers will understand that various types of checkpoints (e.g., metadata-only checkpoints) may be used, and that checkpoints may be distributed based on various factors (e.g., time, number of operations, RPO setting).

[0217] In other embodiments, datasets can be replicated through snapshot-based replication (also referred to as “asynchronous replication”). In snapshot-based replication, snapshots of a dataset can be sent from a replication source, such as a first storage system, to a replication target, such as a second storage system. In such an embodiment, each snapshot may include the entire dataset or only a subset of the dataset, for example, only the portion of the dataset that has changed since the last snapshot was sent from the replication source to the replication target. Readers will understand that snapshots may be sent on demand based on a policy that takes into account various factors (e.g., time, number of operations, RPO setting) or in some other way.

[0218] The storage systems described above can be configured, either individually or in combination, to function as a continuous data protection store. A continuous data protection store is a feature of a storage system that records updates to a dataset, allowing access to a consistent image of the dataset's previous content at a low temporal granularity (often on the order of seconds, or even less) and retrospectively over a reasonable period (often hours or days). These enable access to a very recent consistent point in time for the dataset, and also to access the point in time of the dataset immediately before an event, such as when part of the dataset is corrupted or otherwise lost, while retaining a number of updates close to the maximum number of updates immediately preceding that event. Conceptually, they are like a sequence of snapshots of the dataset that are taken very frequently and retained over long periods, but continuous data protection stores are often implemented quite differently from snapshots. A storage system implementing a continuous data protection store can further provide means to access these points in time, means to access one or more of these points in time as snapshots or clone copies, or means to revert the dataset to one of these recorded points in time.

[0219] Over time, to reduce overhead, some point in time held within a continuous data protection store can be merged with other nearby points in time, essentially deleting some of these point in time from the store. This reduces the capacity required to store updates. It may also be possible to convert these limited number of point in time into longer-duration snapshots. For example, such a store could hold a low-granularity sequence of point in time several hours prior to the present, and merge or delete some point in time to reduce overhead until an additional day. Going even further back, some of these point in time could be converted into snapshots representing a consistent point-in-time image from just a few hours apart.

[0220] While some embodiments are described primarily in the context of storage systems, those skilled in the art will recognize that embodiments of the present disclosure may also take the form of computer program products placed on a computer-readable storage medium for use with any suitable processing system. Such computer-readable storage medium may be any storage medium for machine-readable information, including magnetic media, optical media, solid-state media, or other suitable media. Examples of such media include magnetic disks in hard drives or diskettes, compact disks for optical drives, magnetic tapes, and other media that those skilled in the art can imagine. Those skilled in the art will immediately recognize that any computer system having suitable programming means can perform the steps described herein as embodied in a computer program product. Furthermore, while some embodiments described herein target software installed and run on computer hardware, alternative embodiments implemented as firmware or hardware are also well within the scope of the present disclosure.

[0221] In some examples, non-temporary computer-readable media for storing computer-readable instructions may be provided in accordance with the principles described herein. Instructions can instruct the processor and / or computing device to perform one or more operations, including one or more of the operations described herein, when executed by the processor of the computing device. Such instructions may be stored and / or transmitted using any of the various known computer-readable media.

[0222] Non-temporary computer-readable media as referred to herein may include any non-temporary storage media involved in providing data (e.g., instructions) that can be read and / or executed by a computing device (e.g., by the processor of the computing device). For example, non-temporary computer-readable media may include, but are not limited to, any combination of non-volatile storage media and / or volatile media. Exemplary non-volatile storage media include, but are not limited to, read-only memory, flash memory, solid-state drives, magnetic storage devices (e.g., hard disks, floppy disks, magnetic disks, etc.), ferroelectric random-access memory ("RAM"), and optical discs (e.g., compact discs, digital video discs, Blu-ray discs, etc.). Exemplary volatile storage media include, but are not limited to, RAM (e.g., dynamic RAM).

[0223] One or more embodiments may be described herein with the help of method steps illustrating the performance and relationships of specified functions. The boundaries and sequences of these functional building blocks and method steps are arbitrarily defined herein for the sake of illustrative purposes. Alternative boundaries and sequences may be defined as long as the specified functions and relationships are adequately performed. Therefore, any such alternative boundaries or sequences are within the scope of the claims and spirit. Furthermore, the boundaries of these functional building blocks are arbitrarily defined for the sake of illustrative purposes. Alternative boundaries may be defined as long as certain important functions are adequately performed. Similarly, blocks in the flowchart may also be arbitrarily defined herein to illustrate certain important functions.

[0224] To the extent of use, the block boundaries and sequences of the flow diagram may be defined in other ways and may still perform certain important functions. Therefore, such alternative definitions of both functional building blocks and flow diagram blocks and sequences are within the scope of the claims and spirit. Those skilled in the art will also recognize that the functional building blocks described herein, as well as other illustrative blocks, modules, and components, may be implemented as illustrated, or by individual components, application-specific integrated circuits, processors running appropriate software, or any combination thereof.

[0225] While specific combinations of various functions and features of one or more embodiments are expressly described herein, other combinations of these features and functions are equally possible. This disclosure is not limited by the specific examples disclosed herein and expressly incorporates these other combinations. For further explanation, Figure 4 shows a flowchart illustrating exemplary methods for a dynamic and personalized user experience according to some embodiments of this disclosure. As used herein, user account personality refers to one of several roles that a single user account may perform with respect to a system (such as a storage system). For example, a single user account may be associated with a personality primarily related to the security and protection of the storage system (i.e., a security personality), a personality primarily related to resolving errors in the system (i.e., a troubleshooting personality), a personality primarily related to procuring resources in the storage system (i.e., a procurement personality), a personality primarily related to deploying resources in the storage system (i.e., a deployment personality), and a personality primarily related to government compliance (i.e., a compliance personality).

[0226] Each user account personality may be associated with a set of visual elements that present system characteristics and interaction elements used to perform tasks related to that personality. The visual elements associated with a personality present system characteristics and interaction elements that support activities performed under that personality. Visual elements associated with the same personality may be grouped together within the system interface. For example, one group of visual elements associated with a security personality may include elements such as menus for authorizing or deauthorizing storage client accounts, system characteristics related to data utilization, and system characteristics related to snapshots created from existing data. Another group of visual elements associated with a procurement personality may include elements such as suggestions for devices or services to add to the storage system based on system characteristics. Another group of visual elements associated with a troubleshooting personality may include elements such as system characteristics describing current errors present in the system and actions suggested to address those errors. Another group of visual elements associated with a deployment personality may include elements such as procedural instructions for adding resources or services to the storage system and the status of the deployment process. Finally, another group of visual elements may be associated with compliance personality and may include elements such as rules to be followed and details of each rule.

[0227] An exemplary method depicted in Figure 4 includes receiving a request from user account 406 to access the system interface 404 of system 400. The system interface 404 is a software mechanism for presenting visual elements to the user via the user computer device 408 and user account 406. The system interface 404 is configured and reconfigured by a user interface (UI) engine 402. The UI engine 402 is hardware, software, or a combination of hardware and software that determines the current personality of user account 406 and reconfigures the system interface 404 based on that personality. The UI engine 402 also services requests from the user of user account 406 to the system interface 404. User account 406 is the user's identification information in the storage system 400. User account 406 is under the user's control and managed by a security administrator. User account 406 can interact with the user on the user computer device 408 and translate commands from the user. System 400 may be any of the storage systems described above. The following description relates to the storage system 400, but the described method can be implemented using any computer system having a UI engine 402 and a system interface 404.

[0228] Receiving a request from user account 406 to access the system interface 404 of system 400 may be performed by detecting that the user has logged into system 400 (via user account 406). The request to access the system interface 404 may also be the initialization of a user account session, which refers to the activity between the time when the user of user account 406 logs into system 400 and the time when the user of user account 406 logs out of system 400.

[0229] The exemplary method depicted in Figure 4 also includes identifying a user account personality from a plurality of user account personalities based on personality indicators for user account 406, wherein each personality indicator is associated with at least one of the plurality of user account personalities. As described above, a user account may have a plurality of user account personalities. Such personalities may include a security personality, a troubleshooting personality, and personalities related to adding resources to the system, such as a procurement personality and a deployment personality. Thus, the personality indicators are information that suggests or otherwise indicates that one or more of the user account personalities should be selected for the current user account session.

[0230] One example of a personality indicator is the time when a request to access system interface 404 is received. The UI engine 402 can assume (i.e., configure) that a user who starts a user account session at a particular time is most likely to embody a particular personality. For example, the UI engine 402 can assume that a user account session started between 8:00 AM and 10:00 AM will embody the security personality. As another example, the UI engine 402 can assume that a user account session started on a weekend or holiday will embody the troubleshooting personality.

[0231] Another example of a personality indicator is the location of the user computer device 408. The UI engine 402 can obtain location information from the user computer device 408. The UI engine 402 can assume that a user who starts a user account session from a particular location is most likely to embody a particular personality. For example, the UI engine 402 can assume that a user account session started from the company headquarters will embody the procurement personality.

[0232] Another example of a personality indicator is the characteristics of the user computer device 408 itself. The UI engine 402 may obtain information from the user computer device 408 such as, for example, the IP address, MAC address, type of network connection (e.g., wired, wireless), security level of the network connection (e.g., within a secure local network, via VPN), type of device (e.g., smartphone, laptop, desktop), or other hardware or software on the device. The UI engine 402 can assume that a user who initiates a user account session from a device with certain characteristics is most likely to embody a particular personality. For example, the UI engine 402 can assume that a user account session initiated from an unprotected Wi-Fi network will embody a troubleshooting personality.

[0233] Personality indicators can be based on previous usage patterns. Specifically, personality indicators may be generated based on the usage patterns of a user account 406 interacting with the system interface 404. Machine learning can be used by the UI engine 402 to track a preferred personality at a specific time, from a specific location, and using a device with specific characteristics. Subsequently, when a user starts a user account session, the UI engine 402 can examine the current combination of personality indicators and compare the combination of personality indicators with previous combinations of personality indicators and the most likely desired user account personality. The desired user account personality can be tracked by detecting users who change their personality (e.g., from a presented list of personalities) or by interacting with the system interface 404 in a way that matches a particular personality.

[0234] An example of a personality indicator based on previous usage patterns is the time when a request to access system interface 404 was received. The UI engine 402 can use machine learning to track when a user initiates a user account session and the preferred personality for that time. The UI engine 402 can then compare the time the current user account session is initiated with previously tracked times and associated personalities. For example, a user might initiate a user account session at 1:00 AM on a Sunday. Based on previous occurrences of users initiating user account sessions at similar times, the UI engine 402 can determine that the user is likely to prefer a troubleshooting personality.

[0235] Another example of a personality indicator based on previous usage patterns is the location of the user computer device 408. The UI engine 402 can use machine learning to track the user's location when they initiate a user account session, and the preferred personality for that location. The UI engine 402 can then compare the user's location when the current user account session is initiated with the previously tracked location and associated personality. For example, a user might initiate a user account session from a company satellite office. Based on previous occurrences of user-initiated user account sessions at the same or similar location, the UI engine 402 can determine that the user is likely to prefer a security personality.

[0236] Another example of a personality indicator based on previous usage patterns is the characteristics of the user computer device 408. The UI engine 402 can use machine learning to track the characteristics of the user computer device 408 when the user initiates a user account session, and the preferred personality for those characteristics. The UI engine 402 can then compare the characteristics of the user computer device 408 when the current user account session is initiated with previously tracked characteristics of the user computer device 408 and its associated personality. For example, a user may initiate a user account session from a smartphone using a VPN. Based on the user's previous occurrences of initiating a user account session using a smartphone and VPN, the UI engine 402 can determine that the user is likely to prefer a troubleshooting personality.

[0237] The personality indicator may also include system characteristics. System characteristics are information about the state or context of the storage system 400. System characteristics may include metrics about the storage system (e.g., uptime level, RPO, etc.) as well as information about the storage system itself (e.g., software version, hardware description, etc.). The personality indicator may also include information extracted from systems other than the user computer device 408 and the storage system 400. For example, system characteristics may include comparative information such as metric percentiles within similar deployments (e.g., the bottom 25% of reliability compared to similar companies).

[0238] A request to access the system interface 404 may itself contain one or more personality indicators. The request may include an explicit request from the user to identify one of the personalities. The request may further contain metadata that the UI engine 402 extracts and identifies as personality indicators. Such metadata may include time, location, and system characteristic personality indicators, as described above.

[0239] Identifying user account personalities based on personality indicators 412 can be done by evaluating the retrieved personality indicators and selecting one user account personality for the user account session based on the entirety of the available personality indicators. The selected user account personality then determines how the system interface 404 will be reconfigured. Identifying user account personalities based on personality indicators 412 may also include identifying secondary (and tertiary, etc.) user account personalities for the current user account session. The identified one or more secondary user account personalities may be presented in a less conspicuous manner than the identified primary user account personality.

[0240] The exemplary method depicted in Figure 4 also includes reconfiguring the system interface 404 based on the identified user account personality 414. Reconfiguring the system interface 404 based on the identified user account personality 414 may be performed by adding, removing, and / or rearranging visual elements within the system interface 404. The reconfigured system interface 404 may exclusively present visual elements associated with the identified user account personality. Alternatively, the reconfigured system interface 404 may include groups of visual elements from different user account personalities, with the identified user account personality being presented more prominently than the others. The system interface 404 is “reconfigured” in the sense that the system interface 404 is modified from its default configuration (based on the default personality or an unselected personality). As described above, visual elements may include system characteristics.

[0241] Groups of visual elements associated with the same personality may share similar visual characteristics regardless of the reconstructed and identified user account personality. Such visual characteristics may include, for example, the same color, the same common interface position (e.g., fixed near the upper left corner), or the same brightness or shading. For example, groups of visual elements associated with a security personality may be rendered in different hues of blue, regardless of the personality identified for the current user account session, and groups associated with a troubleshooting personality may be rendered in different hues of red.

[0242] Reconfiguring the system interface 404 based on identified user account personalities 414 may be performed in the middle of a user account session. Specifically, the UI engine 402 can identify different user account personalities in response to user interactions with the system interface 404. Based on initial interactions, the UI engine 402 can reconfigure the system interface 404 to match subsequently identified user account personalities. For example, a user may start a user account session, and the UI engine 402 may initially identify a troubleshooting personality as a possible desired personality. The user may then begin interacting primarily with visual elements associated with a security personality. After a threshold number of interactions, the UI engine 402 can reconfigure the system interface 404 based on the security personality, for example, by making the visual elements associated with the security personality the primary elements of the system interface 404.

[0243] The exemplary method depicted in Figure 4 also includes granting user account 406 access to the reconfigured system interface 404, which includes presenting the reconfigured system interface 404 to the user of user account 406. Granting user account 406 access to the reconfigured system interface 404 416 may be done by modifying the permissions associated with user account 406 and / or the reconfigured system interface 404 that enable the user account to access the reconfigured system interface 404. Presenting the reconfigured system interface 404 to the user of user account 406 may be done by providing user account 406 with a route to the reconfigured system interface 404 which can be used to retrieve the code representing the reconfigured system interface 404.

[0244] Using the process described above, the reconfigured system interface can be used by multiple users in a collaborative setup. Specifically, a first user can initiate a user account session with a first personality that provides a specific perspective on the current state of system 400. A second user can also initiate a user account session with a second personality that provides a different perspective on the same system. For example, the first user can initiate a user account session under a security personality to observe security errors in the system. A second user can initiate a user account session under a troubleshooting personality to attempt to resolve the errors observed by the first user.

[0245] For further explanation, Figure 5 shows a flowchart illustrating an additional exemplary method of a dynamic, personalized user experience according to some embodiments of the present disclosure. The exemplary method depicted in Figure 5 is similar to the exemplary method shown in Figure 4, and the exemplary method shown in Figure 5 also includes receiving a request from user account 406 to access a system interface 404 of system 400 410; identifying a user account personality from a plurality of user account personalities based on personality indicators for user account 406 412, each of which personality indicators is associated with at least one of the plurality of user account personalities; reconfiguring the system interface 404 based on the identified user account personality 414; and granting user account 406 access to the reconfigured system interface 404, including presenting the reconfigured system interface 404 to the user of user account 406 416.

[0246] In the exemplary method depicted in Figure 5, identifying a user account personality from a plurality of user account personalities based on the personality indicator of user account 406 412 includes receiving a selection of a user account personality from the user of user account 406 502 from a list of plurality of user account personalities. Receiving a selection of a user account personality from a list of plurality of user account personalities 502 may be performed by presenting a list of personalities to the user of the user account and receiving a selection of one of the personalities from the user. Such a list may be presented in response to a request to access the system interface 404. Such a list may also be presented within the system interface 404 when the system interface 404 presents a default configuration or a configuration based on previously identified user account personalities.

[0247] For further explanation, Figure 6 shows a flowchart illustrating an additional exemplary method of a dynamic, personalized user experience according to some embodiments of the present disclosure. The exemplary method depicted in Figure 6 is similar to the exemplary method shown in Figure 4, and the exemplary method shown in Figure 6 also includes receiving a request from user account 406 to access a system interface 404 of system 400 410; identifying a user account personality from a plurality of user account personalities based on personality indicators for user account 406 412, each of which personality indicators is associated with at least one of the plurality of user account personalities; reconfiguring the system interface 404 based on the identified user account personality 414; and granting user account 406 access to the reconfigured system interface 404, including presenting the reconfigured system interface 404 to the user of user account 406 416.

[0248] In the exemplary method depicted in Figure 6, identifying a user account personality from multiple user account personalities based on the personality indicators of user account 406 412 includes selecting a user account personality 602 based on a weighted score applied to each personality indicator. Selecting a user account personality 602 based on the weighted score may be performed by determining the weighted score for each personality indicator using previous usage patterns. The UI engine 402 can store the weighted score for each personality indicator along with the user account personality. Specifically, the UI engine 402 can maintain a data structure that maps personality indicators to user account personalities and the weighted scores of those user account personalities. An exemplary entry may include the personality indicators of a user computer device 408 for the current user account session, which is a smartphone. This personality indicator may be mapped to a troubleshooting personality with a 90% probability as a weighted score (for example, based on the previous occurrence rate of the troubleshooting personality, which is the desired personality when using a smartphone). Some personality indicators may be mapped to multiple personalities with different weighted scores (e.g., 50% procurement personality, 50% deployment personality). The UI engine 402 can take each received personality indicator along with its weighted score and calculate the most likely desired personality. If the user works within the presented personalities or instantly switches personalities from the calculated most likely personality, the weighted score data structure may be updated accordingly.

[0249] For further explanation, Figure 7 shows a flowchart illustrating an additional exemplary method of a dynamic, personalized user experience according to some embodiments of the present disclosure. The exemplary method depicted in Figure 7 is similar to the exemplary method shown in Figure 4, as it similarly includes receiving a request from user account 406 to access the system interface 404 of system 400, 410 identifying a user account personality from a plurality of user account personalities based on personality indicators for user account 406, each of which personality indicators is associated with at least one of the plurality of user account personalities, 414 reconfiguring the system interface 404 based on the identified user account personality, and 416 granting user account 406 access to the reconfigured system interface 404, including presenting the reconfigured system interface 404 to the user of user account 406.

[0250] In the exemplary method depicted in Figure 7, reconfiguring the system interface 404 based on identified user account personalities 414 includes arranging visual elements within the system interface 404 702 such that the visual elements associated with the identified user account personality become the primary elements of the system interface 404. Arranging visual elements within the system interface 404 702 such that the visual elements associated with the identified user account personality become the primary elements of the system interface 404 may be done by applying a visual style to the visual elements associated with the identified user account personality to distinguish these visual elements from other groups of visual elements associated with other user account personalities. "Primary" refers to a visual element or group of visual elements that are visually emphasized relative to other visual elements within the system interface 404. Making the visual elements associated with an identified user account personality the primary elements of the system interface 404 can be done in various ways. For example, the visual elements associated with an identified user account personality may be placed in the center of the system interface 404. As another example, visual elements associated with an identified user account personality may be rendered larger than other groups of visual elements associated with other user account personalities. As yet another example, visual elements associated with an identified user account personality may be rendered in a brighter hue than other groups of visual elements associated with other user account personalities. As yet another example, visual elements associated with an identified user account personality may be presented on the top layer of the window layers within the system interface 404.

[0251] Reconfiguring the system interface 404 based on identified user account personalities 414 may include arranging groups of visual elements associated with any identified secondary user account personality such that the visual elements associated with the identified secondary user account personality become secondary elements of the system interface 404. As described above, visual styles may be applied to the visual elements associated with identified secondary user account personalities to distinguish those visual elements from and less conspicuous than the group of visual elements associated with the primary user account personality.

[0252] For further explanation, Figure 8 shows a flowchart illustrating an additional exemplary method of a dynamic, personalized user experience according to some embodiments of the present disclosure. The exemplary method depicted in Figure 8 is similar to the exemplary method shown in Figure 4, as it similarly includes receiving a request from user account 406 to access a system interface 404 of system 400, identifying a user account personality from a plurality of user account personalities based on personality indicators for user account 406, each of which personality indicators is associated with at least one of the plurality of user account personalities, reconfiguring the system interface 404 based on the identified user account personality, and granting user account 406 access to the reconfigured system interface 404, including presenting the reconfigured system interface 404 to the user of user account 406, so the exemplary method depicted in Figure 8 is similar to the exemplary method shown in Figure 4.

[0253] In the exemplary method depicted in Figure 8, reconfiguring the system interface 404 based on the identified user account personality 414 includes populating static objects within the system interface 404 using the visual elements associated with the identified user account personality 802. Populating static objects within the system interface 404 using the visual elements associated with the identified user account personality 802 can be done by matching the visual elements associated with the user account personality with static objects within the system interface 404.

[0254] For each user account personality, the system interface 404 may maintain several consistent static objects. Such static objects may include different windows fixed in specific locations on the screen and may be populated with different visual elements depending on the identified user account personality. For example, the system interface 404 may include a primary window fixed in the center of the system interface 404 and using 60% of the height and 60% of the width of the system interface 404, and secondary windows fixed at each corner and occupying the remaining space. If the security personality is the identified user account personality, such a primary window may populate with a group of visual elements associated with the security personality, such as menus for authorizing or disauthorizing storage client accounts, system characteristics regarding data utilization, and system characteristics regarding snapshots created from existing data. If the troubleshooting personality is the identified user account personality, such a primary window may populate with a group of visual elements associated with the troubleshooting personality, such as system characteristics regarding current errors present in the system and suggested actions to address the errors. If the procurement personality is an identified user account personality, such a primary window may contain a group of visual elements associated with the procurement personality, such as suggestions for devices or services to add to the storage system based on system characteristics. If the deployment personality is an identified user account personality, such a primary window may contain a group of visual elements associated with the deployment personality, such as procedural instructions for adding resources or services to the storage system and the status of the deployment process.

[0255] Reconfiguring the system interface 404 may be based further on an object relation model. An object relation model is a graph that defines objects and the relationships between them. The object relation model of the personality-based system interface 404 described above may include objects that define the different personalities of a user account and the visual elements of each personality. The object relation model may also define the relationships between the visual elements associated with each personality. The object relation model can then be used to populate the system interface 404 with static objects or to generate dynamic objects within the system interface 404.

[0256] For further explanation, Figure 9 shows a flowchart of an exemplary method for profiling user activity to achieve social and governance purposes, according to some embodiments of the present disclosure. The method of Figure 9 may be carried out, for example, by a storage system 900 as described above or another computing system as can be understood. The method of Figure 9 includes generating a plurality of activity groupings, each containing one or more user accounts, corresponding to specific activities within the system, based on data describing activities within the system (e.g., a storage system 900 or another system as can be understood).

[0257] As described herein, each activity grouping corresponds to a specific action or set of actions that may be performed by a user of the system (e.g., storage system 900). Each action may include, for example, a specific function or feature that can be performed by the system. For example, with respect to storage system 900, an action may include creating a snapshot, performing data replication, setting a data replication policy, or restoring a volume from a snapshot or backup. As another example, an action may include various user interactions with the system, such as accepting a recommendation, rejecting a recommendation, or requesting additional details about a recommendation. An action may also include selecting or accessing a specific part of the system, such as selecting a specific user interface element or accessing a specific page or content resource. In some embodiments, an activity grouping may correspond to a set of actions. For example, a specific feature of the system may be associated with multiple actions, and an activity grouping may correspond to the use of a specific feature. Those skilled in the art will understand that these examples of actions are merely illustrative and that other actions performed by users of the system are also contemplated within the scope of this disclosure.

[0258] Each activity grouping includes one or more user accounts. User accounts are included in the activity grouping and are used to perform specific actions on the activity grouping. In other words, users who access the system through a user account are used to perform specific actions. In some embodiments, generating multiple activity groupings 902 is performed with respect to one or more time windows. For example, generating multiple activity groupings 902 selects user accounts that performed the corresponding action within a time window for each activity grouping. In some embodiments, the same time window is applicable to all actions used in generating the activity groupings 902. In other embodiments, different time windows may be used for different actions. Since activity groupings can correspond to non-mutually exclusive actions, the same user account can be included in multiple activity groupings by performing multiple actions corresponding to different activity groupings. For example, a cohort analysis may be performed to select user accounts that performed a particular action (e.g., within a particular time window) in order to generate an activity grouping for a particular action. Thus, generating activity groupings 902 groups user accounts together based on the specific actions performed through those user accounts.

[0259] As described above, generating multiple activity groupings 902 is based on data describing activities within the system. Therefore, in some embodiments, the data describing activities within the system may include logs describing specific activities performed within the system by specific user accounts. For example, audit logs or other logs may be maintained by storing entries indicating a specific action and the user account used to perform that action when each action is performed within the system. To generate an activity grouping for a specific activity, the log entries for that specific activity can be accessed, and the corresponding user accounts indicated in the accessed log entries are added to the activity grouping.

[0260] In some embodiments, user accounts may be added to an activity group based on any performance of the corresponding activity. In other words, a user account may be added to an activity group if it is used to perform a particular action (e.g., within a time window). In some embodiments, a user account may be added to an activity group in response to some metric associated with performing the account (e.g., a certain frequency, a certain frequency for other users who performed the action) exceeding a threshold.

[0261] The method in Figure 9 also includes generating one or more social groupings for each of a plurality of activity groupings based on the user profiles of one or more user accounts within the corresponding activity grouping, with each of the one or more social groupings corresponding to one or more specific user profile attributes. As described herein, a social grouping is a grouping of user accounts (e.g., within a particular activity grouping) that is grouped based on shared characteristics within a user profile relating to the user account. In other words, each social grouping corresponds to one or more user profile attributes, and each social grouping includes user accounts that have those user profile attributes. Thus, user accounts are first grouped based on the actions performed using those user accounts (e.g., activity groupings), and then further grouped based on user profile attributes (e.g., becoming a social group for each of the activity groupings).

[0262] A user profile describes a user associated with a particular user account. Therefore, user profile attributes include various descriptors of the user. Such user profile attributes include, for example, the user's age, gender identification, sexual orientation, place of residence or nationality, the primary language spoken or understood by the user, and one or more other languages ​​spoken or understood by the user. In some embodiments, user profile attributes may include null values ​​or the absence of such values. For example, certain attributes of the user profile (e.g., name, email address) may be required, while other attributes of the user profile may be optionally entered by the user. Therefore, in some embodiments, certain social groupings may correspond to null or absent values ​​for certain attributes, or certain values ​​indicating that the user does not wish to provide a value.

[0263] In some embodiments, social grouping may correspond to a specific range or set of multiple values ​​for a particular user profile attribute. For example, with respect to age, multiple social groupings may each correspond to a range of ages. In some embodiments, social grouping may correspond to a specific value among multiple selectable values. For example, with respect to gender identification, different social groupings may correspond to male, female, and non-binary, respectively. In some embodiments, social grouping may correspond to a specific combination of user profile attributes or ranges of user attributes. For example, social groupings may correspond to different age ranges for different nationalities. In other words, social groupings can be generated at various levels of granularity according to different design or engineering considerations.904

[0264] Generating one or more social groupings for a particular activity grouping may involve accessing user profiles for user accounts within that particular activity grouping and then performing a cohort analysis to group those user profiles into different social groupings. Thus, each activity grouping may contain one or more activity groupings, and each user account may be included in multiple social groupings for multiple activity groupings. Furthermore, social groupings for a particular user attribute or set of user attributes may be generated for multiple activity groupings or represented in multiple activity groupings. In some embodiments, two or more social groupings for a given activity grouping may be merged based on various criteria. For example, social groupings with a number of user accounts below a threshold may be merged into social groupings that cover aggregated or broader user profile attributes.

[0265] The method in Figure 9 also includes identifying, for a particular user account, one or more activity groupings that have social groupings corresponding to the user profile attributes of that particular user account.906 In other words, for a particular user account, an activity grouping that includes a social grouping corresponding to that particular user account is identified. A particular user account corresponds to or is included in a social grouping, and the user profile of a particular user account includes user profile attributes that match the user profile attributes of that social grouping. For example, if a social grouping corresponds to a particular age range, a user account corresponds to or is included in a social grouping, and the user profile of a user account includes ages that fall within that age range. As another example, if a social grouping corresponds to a particular primary language, the user profile corresponds to the social grouping, and the user account profile indicates a primary language that matches the particular primary language of the social grouping. As a further example of how social grouping corresponds to combinations of attributes, assuming that social grouping corresponds to a specific age range and nationality (e.g., place of residence or origin), the user profile corresponds to or includes the social grouping, and the user account profile indicates a specific nationality and also indicates an age that falls within the specific age range of the social grouping.

[0266] In some embodiments, identifying one or more activity groupings 906 may include identifying activity groupings in which a user account is included in a social grouping that matches a particular criterion. For example, identifying one or more activity groupings 906 may include identifying an activity grouping in which a user account is included in a social grouping that has a certain threshold number of user accounts or has a certain representation in activity groupings that exceed a threshold. For example, identifying one or more activity groupings 906 may include identifying an activity grouping in which a user account is included in the top N social groupings based on user account membership or based on exceeding a certain percentage of user account membership for that activity grouping relative to other social groupings. As another example, identifying one or more activity groupings 906 may include identifying an activity grouping that has the highest similarity between a user-specific user account and the social grouping of that activity grouping. Identifying one or more activity groupings 906 may also be based on a variety of other rules or criteria, as can be understood.

[0267] The method in Figure 9 also includes modifying one or more user experience features of the system based on one or more identified activity groupings.908 In some embodiments, one or more user experience features are modified based on specific activities corresponding to one or more identified activity groupings. Thus, one or more user experience features are modified based on specific activities in an activity grouping whose social grouping includes a particular user account. This allows a user experience feature to be modified based on the actions of other users in a social grouping similar to that of a particular user account.

[0268] In some embodiments, one or more user experience features may include a user interface. For example, assuming that certain actions correspond to identified activity groupings of a user account, the user interface may be modified to present user interface elements that enable the execution of those specific actions. Certain menu elements may be modified to include or highlight user interface elements that facilitate the execution of specific actions. Continuing this example, suppose a social grouping of people over 35 years old is included in an activity grouping for creating snapshots within the storage system, and further, suppose a particular user account has an age of over 35. In response to identifying the snapshot activity grouping of a user account, the user interface for a particular user account may be modified to include a shortcut or menu element for creating snapshots, or to create or modify snapshot policies. As another example, a dashboard element describing various metrics related to creating snapshots may be added to the user interface for a user account. As yet another example, suppose a social grouping of primary Spanish speakers is included in an activity grouping for accessing the chat technical support feature, thereby indicating that Spanish speakers frequently use the chat technical support feature. Furthermore, suppose a particular user account profile indicates Spanish as the user's primary language. The user interface for user accounts may be modified to automatically open or minimize the technical support chat window, or to present a button to open the technical support chat window. In other words, the default or baseline user interface may be modified to include or present specific features or content based on the social grouping of user accounts and their corresponding activity groupings.

[0269] One or more user experience features may also include one or more recommendations. For example, suppose a social group of users aged 50 or older is included in an activity group that rejects recommendations. Furthermore, suppose the user profile indicates that the user of the user account is aged 50 or older. Modifying one or more user experience features 908 may include reducing the frequency of recommendations presented to the user account. As another example, suppose a particular user account is included in a social group within an activity group that requests more detailed information about the recommendations presented (e.g., requests an explanation of why a particular recommendation was generated or presented). Modifying one or more user experience features 908 may include automatically providing more detailed information when recommendations are presented to the user account.

[0270] In some embodiments, one or more user experience features are modified based on specific activities corresponding to activity groupings other than the identified 906 activity groupings. Thus, one or more user experience features are modified based on specific activities in activity groupings whose social grouping does not include a particular user account. This allows user experience features to be modified based on the actions of other users in social groups different from a particular user account.

[0271] If a particular feature of the system is not frequently used by members of a specific social group, including a user account, the user interface for that user account may be modified to highlight or present user interface elements that facilitate the use of that feature. Continuing the above example, where a social group of primary Spanish speakers is included in the activity group for accessing the chat technical support feature, but not in the activity group for accessing the help wiki or database, let's further assume that a particular user account profile indicates Spanish as the user's primary language. The user interface may be modified to highlight or present menu elements for accessing the help wiki (e.g., in conjunction with or instead of the chat interface). Thus, features not commonly used by members of a particular social group may be highlighted or promoted to attract more members of that social group to that particular feature.

[0272] In some embodiments, modifying one or more user experience features of the system 908 includes presenting the modified one or more user experience features 910. For example, a user interface may be encoded and presented to a user device associated with a particular user account. As another example, recommendations may be generated and presented to a user device associated with a particular user account.

[0273] In some embodiments, modifying one or more user experience features 908 may be performed based on a user account configuration. For example, in some embodiments, the user account configuration may include an optional parameter indicating whether one or more user experience features should be modified 908. Thus, whether a user is presented with a modified 908 user experience depends on the settings in the user account configuration. This allows the user to receive either a default or a modified user experience based on their specific user account configuration.

[0274] In some embodiments, modifying one or more user experience features (908) may be based on a system configuration associated with multiple user accounts. For example, the system configuration may indicate whether a user should receive a modified experience feature, or whether users in a particular social group should receive a modified user experience feature. As another example, the system configuration may indicate that some users should receive a default user experience and others should receive a modified user experience. For example, a subset of users (e.g., within the same social group, or across all users) may be selected to receive the default user experience, and another subset of users may be selected to receive a modified user experience, such as when conducting A / B testing or other tests on a particular system feature or user experience feature.

[0275] For further explanation, Figure 10 shows a flowchart of an exemplary method for profiling user activity to achieve social and governance purposes, according to some embodiments of the present disclosure. The method in Figure 10 is similar to the method in Figure 9 in that it includes generating a plurality of activity groupings, each containing one or more user accounts and corresponding to a particular activity in the system, based on data describing activity in the system (902); generating one or more social groupings for each of the plurality of activity groupings based on the user profiles of one or more user accounts in the corresponding activity grouping (904), wherein each of the one or more social groupings corresponds to one or more specific user profile attributes, and for a particular user account, identifying one or more activity groupings that have social groupings corresponding to the user profile attributes of the particular user account (906); and modifying one or more user experience features of the system based on the identified one or more activity groupings (908), including presenting one or more modified user experience features (910).

[0276] The method in Figure 10 differs from that in Figure 9 in that it includes performing a representation analysis 1002 based on one or more social groupings for each of several activity groupings. The following description describes performing a representation analysis (1002) in conjunction with modifying one or more user experience features (908), but those skilled in the art will understand that in some embodiments, performing a representation analysis (1002) can be performed independently of any modified user experience features. For example, the representation analysis may be performed after generating activity groupings (902) and generating social groupings of user accounts included in those activity groupings (904).

[0277] Representation analysis describes the extent to which a particular social grouping is represented in relation to the use of the system. Representation analysis can describe the extent to which a social grouping accesses a particular feature or performs a particular action. For example, in some embodiments, representation analysis can describe feature usage by members of a particular social grouping. Continuing this example, representation analysis may describe feature usage (e.g., most commonly or rarely used features, rarely used features) in particular for age grouping, language grouping, nationality grouping, etc. As another example, representation analysis may describe the usage of a particular feature across different age grouping, language grouping, nationality grouping, etc. Representation analysis can describe the extent to which a particular social grouping as a whole accesses or uses the system. Representation analysis may correspond to a specific time window or multiple time windows. For example, representation analysis may describe feature usage over time for a particular social grouping, the use of a particular feature over time across multiple social groupings, or the overall system usage over time for multiple social groupings.

[0278] Representation analysis may be used to benchmark or evaluate the use of a system or specific system features with respect to social groups. For example, representation analysis can identify specific biases or skews in the use of a particular feature by a particular social group. As another example, representation analysis can be used to determine whether a particular feature is used by a diverse audience. Thus, in some embodiments, the system or particular feature may include a quantifiable evaluation of representation or diversity based on a particular social grouping or social grouping membership for a particular activity grouping.

[0279] In some embodiments, representation analysis may be based on specific representational objectives. Representational objectives may be based on specific governance or compliance objectives for diversity, bias, representation, accessibility, etc. Representational objectives may also be based on internal (e.g., organization-specific) diversity, bias, representation, or accessibility objectives. Representational objectives may indicate the degree of representation of membership in a particular social group (e.g., for a particular activity grouping). For example, representational objectives may indicate the target number of members for a particular social group to use a particular feature. The target number of members may include a specific number of members (e.g., a specified number of user accounts), a target increase in members (e.g., by the number or percentage of members), or a target ratio of membership to other user accounts. Representation analysis may also indicate the degree of diversity based on which social groupings are included in a particular activity grouping or in relation to the entire system. For example, for a particular feature, representation analysis may indicate which social groups use that particular feature (e.g., by including it in the corresponding activity grouping) and assess the degree of diversity for that particular feature based on how many different social groupings use that particular feature.

[0280] Performance analysis can be used to meet various social or governance considerations. For example, performance analysis may be used to meet specific internal goals or objectives for diversity or representation in user engagement with the system. As another example, performance analysis may be used to meet specific governance objectives regarding diversity, usefulness, bias, etc. When a system (e.g., storage system 900) is used by a specific customer or tenant, performance analysis can provide social information and governance information useful to that specific customer or tenant. Performance analysis can also provide insights into the makeup of a specific customer organization, which can then be useful in employment considerations. In particular, performance analysis utilizes information correlating specific system actions with specific user accounts and user profiles to provide more detailed insights into user behavior.

[0281] The advantages and features of the present disclosure can be further explained by the following statements.

[0282] 1. A method of receiving a request to access a system interface for a system from a user account, comprising identifying a user account personality from a plurality of user account personalities based on a personality indicator for the user account, wherein each of the personality indicators is associated with at least one of the plurality of user account personalities, reconfiguring the system interface based on the identified user account personality, presenting the reconfigured system interface to a user of the user account, and permitting access to the reconfigured system interface to the user account.

[0283] 2. The method according to statement 1, wherein identifying a user account personality from a plurality of user account personalities based on a personality indicator includes receiving a selection of a user account personality from a list of the plurality of user account personalities from a user of the user account.

[0284] 3. The method according to statement 2 or 1, wherein identifying a user account personality from a plurality of user account personalities based on a personality indicator includes selecting a user account personality based on a weighted score applied to each personality indicator.

[0285] 4. The method according to statement 3, 2 or 1, wherein reconfiguring a system interface based on the identified user account personality includes arranging visual elements within the system interface such that the visual elements associated with the identified user account personality become primary elements of the system interface.

[0286] 5. The method according to statement 4, 3, 2 or 1, wherein reconfiguring a system interface based on the identified user account personality includes populating static objects within the system interface using visual elements associated with the identified user account personality.

[0287] 6. The method according to statement 5, 4, 3, 2 or 1, wherein reconfiguring the system interface is further based on an object relational model.

[0288] 7. The method according to statement 6, 5, 4, 3, 2 or 1, wherein the personality indicator is generated based on machine learning of previous usage patterns of the user account interacting with the system interface.

[0289] 8. The method described in Statements 7, 6, 5, 4, 3, 2, or 1, wherein multiple user account personalities include a user account personality associated with the security of the system.

[0290] 9. The method according to statements 8, 7, 6, 5, 4, 3, 2, or 1, including a user account personality associated with resolving errors in the system.

[0291] 10. The method according to statements 9, 8, 7, 6, 5, 4, 3, 2, or 1, including a user account personality associated with adding resources to the system.

[0292] One or more embodiments may be described herein with the help of method steps illustrating the performance and relationships of specified functions. The boundaries and sequences of these functional building blocks and method steps are arbitrarily defined herein for the sake of illustrative purposes. Alternative boundaries and sequences may be defined as long as the specified functions and relationships are adequately performed. Therefore, any such alternative boundaries or sequences are within the scope of the claims and spirit. Furthermore, the boundaries of these functional building blocks are arbitrarily defined for the sake of illustrative purposes. Alternative boundaries may be defined as long as certain important functions are adequately performed. Similarly, blocks in the flowchart may also be arbitrarily defined herein to illustrate certain important functions.

[0293] To the extent of use, the block boundaries and sequences of the flow diagram may be defined in other ways and may still perform certain important functions. Therefore, such alternative definitions of both functional building blocks and flow diagram blocks and sequences are within the scope of the claims and spirit. Those skilled in the art will also recognize that the functional building blocks described herein, as well as other illustrative blocks, modules, and components, may be implemented as illustrated, or by individual components, application-specific integrated circuits, processors running appropriate software, or any combination thereof.

[0294] While specific combinations of various functions and features of one or more embodiments are expressly described herein, other combinations of these features and functions are equally possible. This disclosure is not limited by the specific examples disclosed herein and expressly incorporates these other combinations.

Claims

1. It is a method, Receiving requests from user accounts to access system interfaces for the system, Identifying a user account personality from a plurality of user account personalities based on a personality indicator for the user account, wherein each of the personality indicators is associated with at least one of the plurality of user account personalities. Reconfiguring the system interface based on the identified user account personality, A method comprising presenting the reconfigured system interface to a user of the user account, and granting the user account access to the reconfigured system interface.

2. The method according to claim 1, wherein identifying a user account personality from a plurality of user account personalities based on the personality indicator includes receiving a selection of the user account personality from the user of the user account.

3. The method according to claim 1, wherein identifying a user account personality from a plurality of user account personalities based on the personality indicators includes selecting a user account personality based on a weighted score applied to each personality indicator.

4. The method according to claim 1, wherein reconfiguring the system interface based on the identified user account personality includes arranging visual elements within the system interface such that the visual elements associated with the identified user account personality become the primary elements of the system interface.

5. The method according to claim 1, wherein reconfiguring the system interface based on the identified user account personality includes populating the system interface with static objects using visual elements associated with the identified user account personality.

6. The method according to claim 1, wherein the system interface is reconfigured to be based on an object relationship model.

7. The method according to claim 1, wherein the personality indicator is generated based on machine learning of the user account's previous usage patterns interacting with the system interface.

8. The method according to claim 1, wherein the plurality of user account personalities include a user account personality associated with the security of the system.

9. The method according to claim 1, wherein the plurality of user account personalities include a user account personality associated with resolving errors in the system.

10. The method according to claim 1, wherein the plurality of user account personalities include a user account personality associated with adding resources to the system.

11. An apparatus comprising a computer processor and computer memory operably coupled to the computer processor, wherein when the computer memory is executed by the computer processor, the apparatus has The process of receiving a request from a user account to access the system interface for the system, A step of identifying a user account personality from a plurality of user account personalities based on a personality indicator for the user account, wherein each of the personality indicators is associated with at least one of the plurality of user account personalities, The process involves reconfiguring the system interface based on the identified user account personality, An apparatus for performing the steps of granting access to the reconfigured system interface to a user account, which includes presenting the reconfigured system interface to the user account.

12. The apparatus according to claim 11, wherein identifying a user account personality from a plurality of user account personalities based on the personality indicator includes receiving a selection of the user account personality from the user of the user account.

13. The apparatus according to claim 11, wherein identifying a user account personality from a plurality of user account personalities based on the personality indicators includes selecting a user account personality based on a weighted score applied to each personality indicator.

14. The apparatus according to claim 11, wherein reconfiguring the system interface based on the identified user account personality includes placing static objects into the system interface using visual elements associated with the identified user account personality.

15. The apparatus according to claim 11, wherein the system interface is reconfigured to be based on an object relationship model.

16. The apparatus according to claim 11, wherein the personality indicator is generated based on machine learning of the user account's previous usage patterns interacting with the system interface.

17. The apparatus according to claim 11, wherein the personality indicator includes system characteristics.

18. The apparatus according to claim 11, wherein the plurality of user account personalities include a user account personality associated with the security of the system.

19. The apparatus according to claim 11, wherein the plurality of user account personalities include a user account personality associated with resolving errors in the system.

20. A computer program product located on a computer-readable medium, wherein when the computer program product is executed, the computer, The process of receiving a request from a user account to access the system interface for the system, A step of identifying a user account personality from a plurality of user account personalities based on a personality indicator for the user account, wherein each of the personality indicators is associated with at least one of the plurality of user account personalities, The process involves reconfiguring the system interface based on the identified user account personality, A computer program product that causes the user account to perform the steps of granting access to the reconfigured system interface, which includes presenting the reconfigured system interface to the user of the user account.