Dynamic tier selection in data storage systems

By introducing multi-plane dies and dynamically adjusting the number of planes in the storage system, compatibility issues in solid-state drive design have been resolved, resulting in more efficient data storage and optimized storage performance.

CN119336248BActive Publication Date: 2026-01-27PURE STORAGE INC
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410984424.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2023-07-21
Filing Date
2024-07-22
Publication Date
2026-01-27
Estimated Expiration
2044-07-22

AI Technical Summary

Technical Problem

Existing solid-state drive designs struggle to fully utilize the unique characteristics of flash memory and other solid-state storage devices, making it difficult to provide enhanced features, and compatibility issues hinder the optimization of storage performance.

Method used

By introducing multi-plane dies into the storage system, dynamically adjusting the number of planes, and combining erase block sets based on the newly determined number of planes, blocks are allocated using the new block size, thereby optimizing data access.

Benefits of technology

It improves the performance and efficiency of the storage system, optimizes the compatibility and reliability of data storage, and enhances the flexibility and adaptability of the storage system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119336248B_ABST
    Figure CN119336248B_ABST
Patent Text Reader

Abstract

This application relates to dynamic plane selection in data storage systems. A storage system is provided. The storage system includes a plurality of non-volatile memory modules and a storage system controller. One or more non-volatile memory modules include a multi-plane die. A processing device of the storage system controller is configured to determine that a number of planes of the multi-plane die that are concurrently used for accessing data should be changed. In response to determining that the number of planes of the multi-plane die that are concurrently used for accessing data should be changed, the processing device is configured to move one or more portions from an existing erase block to a new erase block, the existing erase block having a size that is different than a size of the new erase block.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to dynamic plane selection in data storage systems. Background Technology

[0002] Solid-state storage, such as flash memory, is currently used in solid-state drives (SSDs) to enhance or replace traditional hard disk drives (HDDs), writable CD (optical disc) or writable DVD (digital versatile disc) drives (collectively known as spinning media), and tape drives for storing large amounts of data. Flash memory and other solid-state storage devices have characteristics that differ from spinning media. However, for compatibility reasons, many SSDs are designed to conform to hard disk drive standards, making it difficult to provide enhanced features or take advantage of the unique aspects of flash memory and other solid-state storage devices.

[0003] The examples emerged in this context. Summary of the Invention

[0004] Embodiments of this disclosure provide a storage system comprising: a plurality of non-volatile memory modules, each non-volatile memory module including a multi-plane die; and a storage system controller operatively coupled to the plurality of storage devices, the storage system controller including processing means configured to: determine the number of planes on the multi-plane die used for accessing data simultaneously that should be changed; and, in response to determining the number of planes on the multi-plane die used for accessing data simultaneously that should be changed, allocate blocks using a new block size by combining sets of erase blocks at the same address in individual planes based on the newly determined number of planes.

[0005] Another embodiment of this disclosure provides a method comprising: determining the number of planes used for accessing data simultaneously on a multiplane die, the multiplane die being part of one of a plurality of nonvolatile memory modules in a storage system; and, in response to determining the number of planes used for accessing data simultaneously on the multiplane die, allocating blocks using a new block size by combining a set of erase blocks at the same address in individual planes based on the newly determined number of planes.

[0006] Another embodiment of this disclosure provides a non-transitory computer-readable storage medium for storing instructions that, when executed, cause a processing device of a storage system controller to: determine that the number of planes used for accessing data simultaneously on a multi-plane die should be changed, the multi-plane die being part of one of a plurality of non-volatile memory modules in the storage system; and, in response to determining that the number of planes used for accessing data simultaneously on the multi-plane die should be changed, allocate blocks using a new block size by combining a set of erase blocks at the same address in individual planes based on the newly determined number of planes. Attached Figure Description

[0007] The described embodiments and their advantages can be best understood through the following description taken in conjunction with the accompanying drawings. These drawings are not intended to limit in any way any changes that may be made to the form and details of the described embodiments by those skilled in the art without departing from the spirit and scope of the described embodiments.

[0008] This disclosure is illustrated by way of example but not by way of limitation, and a fuller understanding can be obtained by referring to the following detailed description when considered in conjunction with the accompanying drawings.

[0009] Figure 1A A first example system for data storage according to some implementation schemes is shown.

[0010] Figure 1B A second example system for data storage according to some implementation schemes is shown.

[0011] Figure 1C A third example system for data storage is shown according to some implementation schemes.

[0012] Figure 1D A fourth example system for data storage is shown according to some implementation schemes.

[0013] Figure 2A This is a perspective view of a storage cluster according to some embodiments, the storage cluster having multiple storage nodes and internal storage devices coupled to each storage node to provide network-attached storage.

[0014] Figure 2B This is a block diagram illustrating an interconnect switch coupling multiple storage nodes according to some embodiments.

[0015] Figure 2C This is a multi-level block diagram based on some embodiments, illustrating the contents of a storage node and the contents of one of the non-volatile solid-state storage cells.

[0016] Figure 2D This illustrates a storage server environment according to some embodiments, which uses storage nodes and storage units as shown in some previous diagrams.

[0017] Figure 2E This is a blade hardware block diagram based on some embodiments, illustrating the control plane, compute and storage plane, and authority that interact with the underlying physical resources.

[0018] Figure 2F A resilient software layer in the blades of a storage cluster according to some embodiments is described.

[0019] Figure 2GThe licensing and storage resources in the blades of a storage cluster according to some embodiments are described.

[0020] Figure 3A A diagram illustrating a storage system coupled to communicate data with a cloud service provider, according to some embodiments of the present disclosure.

[0021] Figure 3B A diagram illustrating a storage system according to some embodiments of the present disclosure is provided.

[0022] Figure 3C Examples of cloud-based storage systems according to some embodiments of the present disclosure are illustrated.

[0023] Figure 3D An exemplary computing device 350 is shown, which may be specifically configured to perform one or more of the processes described herein.

[0024] Figure 3E Examples of a series of storage systems 376 for providing storage services are shown.

[0025] Figure 3F An example container system is shown.

[0026] Figure 4 This is a block diagram illustrating an example storage system according to some embodiments of the present disclosure.

[0027] Figure 5 This is a block diagram illustrating an example non-volatile memory module according to some embodiments of the present disclosure.

[0028] Figure 6A This is a block diagram illustrating an example erase block according to some embodiments of the present disclosure.

[0029] Figure 6B This is a block diagram illustrating an example erase block according to some embodiments of the present disclosure.

[0030] Figure 7 This is a flowchart illustrating a method for performing a storage operation according to some embodiments of the present disclosure. Detailed Implementation

[0031] The following embodiments describe a storage cluster for storing user data, such as user data originating from one or more user or client systems or other sources outside the storage cluster. The storage cluster distributes user data among storage nodes housed within a chassis using erasure coding and redundant copies of metadata. Erasure coding is a data protection method in which data is broken down into fragments, expanded and encoded with redundant data blocks, and stored in a set of different locations, such as disks, storage nodes, or geographic locations. Flash memory is a type of solid-state memory that can be integrated with the embodiments, but the embodiments can be extended to other types of solid-state memory or other storage media, including non-solid-state memory. Control of storage locations and workloads is distributed across storage locations in the cluster peer system. Tasks such as regulating communication between storage nodes, detecting when storage nodes are unavailable, and balancing I / O (input and output) on storage nodes are all processed on a distributed basis. In some embodiments, data is arranged or distributed across multiple storage nodes in the form of data fragments or stripes that support data recovery. Data ownership can be reallocated within the cluster, regardless of input and output patterns. The architecture described in more detail below allows for the failure of storage nodes in the cluster, but the system remains operational because data can be reconstructed from other storage nodes, thus remaining available for input and output operations. In various embodiments, storage nodes may be referred to as cluster nodes, blades, or servers.

[0032] A storage cluster is housed within a chassis, i.e., an enclosure that houses one or more storage nodes. The chassis contains mechanisms for powering each storage node, such as a power distribution bus, and communication mechanisms, such as a communication bus that enables communication between storage nodes. According to some embodiments, the storage cluster can operate as a standalone system in one location. In one embodiment, the chassis houses at least two instances of power distribution and internal and external communication buses that can be independently enabled or disabled. The internal communication bus can be an Ethernet bus; however, other technologies, such as Peripheral Component Interconnect (PCI) High Speed, InfiniBand, etc., are equally applicable. The chassis provides ports for the external communication bus for communication between multiple chassis and with client systems, either directly or via a switch. External communication can use technologies such as Ethernet, InfiniBand, Fibre Channel, etc. In some embodiments, the external communication bus uses different communication bus technologies for inter-chassis and client communication. If switches are deployed within or between chassis, the switches can act as converters between various protocols or technologies. When multiple chassis are connected to define a storage cluster, clients can access the storage cluster using proprietary or standard interfaces such as Network File System (NFS), Common Internet File System (CIFS), Small Computer System Interface (SCSI), or Hypertext Transfer Protocol (HTTP). Protocol conversion from the client may occur at the switch, on the external communication bus of the chassis, or within each storage node.

[0033] Each storage node can be one or more storage servers, and each storage server is connected to one or more non-volatile solid-state memory (SSD) cells, which may be referred to as storage cells. One embodiment includes a single storage server and one to eight SSD cells in each storage node, but this example is not intended to be limiting. The storage server may include a processor, dynamic random access memory (DRAM), and interfaces for internal communication buses and power distribution for each power bus. In some embodiments, within the storage node, interfaces and storage cells share a communication bus, such as PCI High Speed. SSD cells can directly access the internal communication bus interface via the storage node's communication bus, or request access to the bus interface from the storage node. SSD cells house an embedded central processing unit (CPU), a solid-state storage controller, and a number of solid-state high-capacity storage devices, for example, between 2 and 32 terabytes (TB) in some embodiments. SSD cells include embedded volatile storage media such as DRAM and energy storage devices. In some embodiments, the energy storage devices are capacitors, supercapacitors, or batteries, capable of transferring a subset of the DRAM content to a stable storage medium in the event of a power outage. In some embodiments, the non-volatile solid-state memory cell is composed of a storage class memory such as phase-change or other resistive random access memory (RRAM) or magnetoresistive random access memory (MRAM), which can replace DRAM and achieve reduced device hold power.

[0034] In some embodiments, a storage node has one or more non-volatile solid-state storage cells, each having non-volatile random access memory (NVRAM) and flash memory. The NVRAM and flash memory can be addressed independently by the storage node, or more specifically, by the processor of the storage node. The storage node writes user data to the NVRAM for storage in the flash memory. The storage node, such as the processor of the storage node, also writes metadata to the NVRAM, which can be used as the processor's work area. Direct memory access (DMA) is used to transfer user data from the NVRAM to the flash memory for storage. DMA is also used to transfer the contents of the NVRAM to the flash memory in the event of a power outage. In some embodiments, a controller in the non-volatile solid-state storage device manages a mapping or translation table and performs various duties to manage the user data in the flash memory.

[0035] Example methods, apparatuses, and products for configurable storage systems according to embodiments of the present disclosure are described with reference to the accompanying drawings and the following disclosure.

[0036] Figure 1AAn example system for data storage according to some embodiments is shown. For illustrative purposes and not for limitation, system 100 (also referred to herein as a “storage system”) includes a number of elements. It will be noted that system 100 may include the same, more or fewer elements configured in the same or different ways in other embodiments.

[0037] System 100 includes several computing devices 164A-B. These computing devices (also referred to herein as “client devices”) can be, for example, servers, workstations, personal computers, laptops, etc., in a data center. The computing devices 164A-B can be coupled to communicate with one or more storage arrays 102A-B via a storage area network ('SAN') 158 or a local area network ('LAN') 160.

[0038] SAN 158 can be implemented using various data communication architectures, devices, and protocols. For example, the architecture of SAN 158 can include Fibre Channel, Ethernet, wireless bandgap, Serial Attached Small Computer System Interface ('SAS'), and so on. Data communication protocols used for SAN 158 can include Advanced Technology Attachment ('ATA'), Fibre Channel protocol, Small Computer System Interface ('SCSI'), Internet Small Computer System Interface ('iSCSI'), HyperSCSI, Structure-Based Non-Volatile Memory High Speed ​​('NVMe'), and so on. It should be noted that SAN 158 is provided for illustrative purposes and not for limitation. Other data communication couplings can be implemented between computing devices 164A-B and storage arrays 102A-B.

[0039] LAN 160 can also be implemented using various architectures, devices, and protocols. For example, LAN 160 architectures can include Ethernet (802.3), wireless (802.11), and so on. Data communication protocols used for LAN 160 can include Transmission Control Protocol ('TCP'), User Datagram Protocol ('UDP'), Internet Protocol ('IP'), Hypertext Transfer Protocol ('HTTP'), Wireless Access Protocol ('WAP'), Handheld Device Transfer Protocol ('HDTP'), Session Initiation Protocol ('SIP'), Real-Time Protocol ('RTP'), and so on.

[0040] Storage arrays 102A-B can provide persistent data storage for computing devices 164A-B. In some embodiments, storage array 102A may be housed in a chassis (not shown), and storage array 102B may be housed in another chassis (not shown). Storage arrays 102A and 102B may include one or more storage array controllers 110A-D (also referred to herein as "controllers"). Storage array controllers 110A-D may be embodied as modules of an automated computing machine, including computer hardware, computer software, or a combination of computer hardware and software. In some embodiments, storage array controllers 110A-D may be configured to perform various storage tasks. Storage tasks may include writing data received from computing devices 164A-B to storage arrays 102A-B, erasing data from storage arrays 102A-B, retrieving data from storage arrays 102A-B, providing data to computing devices 164A-B, monitoring and reporting storage device utilization and performance, performing redundancy operations (such as Independent Drive Redundancy Array ('RAID') or RAID-type data redundancy operations), compressing data, encrypting data, and so on.

[0041] The storage array controllers 110A-D can be implemented in various ways, including as field-programmable gate arrays ('FPGAs'), programmable logic chips ('PLCs'), application-specific integrated circuits ('ASICs'), systems-on-chips ('SOCs'), or any computing device containing discrete components such as processing devices, central processing units, computer memory, or various adapters. The storage array controllers 110A-D may include, for example, data communication adapters configured to support communication via SAN 158 or LAN 160. In some embodiments, the storage array controllers 110A-D may be independently coupled to LAN 160. In some embodiments, the storage array controllers 110A-D may include I / O controllers, etc., coupled to the storage array controllers 110A-D for data communication with persistent storage resources 170A-B (also referred to herein as "storage resources") via a midplane (not shown). Persistent storage resources 170A-B may include any number of storage drives 171A-F (also referred to herein as "storage devices") and any number of non-volatile random access memory ('NVRAM') devices (not shown).

[0042] In some implementations, the NVRAM devices of persistent storage resources 170A-B can be configured to receive data from storage array controllers 110A-D for storage in storage drives 171A-F. In some examples, the data may originate from computing devices 164A-B. In some examples, writing data to the NVRAM devices can be implemented faster than writing data directly to storage drives 171A-F. In some implementations, storage array controllers 110A-D can be configured to use the NVRAM devices as a fast access buffer for data predetermined to be written to storage drives 171A-F. The latency of write requests using NVRAM devices as buffers can be improved compared to a system where storage array controllers 110A-D directly writes data to storage drives 171A-F. In some implementations, the NVRAM devices can be implemented using computer memory in the form of high-bandwidth, low-latency RAM. NVRAM devices are referred to as "non-volatile" because they can receive or contain a single power source that maintains the state of the RAM after the main power supply to the NVRAM device is lost. Such a power source can be a battery, one or more capacitors, etc. In response to a power outage, the NVRAM device can be configured to write the contents of the RAM to a permanent storage device, such as a storage drive 171A-F.

[0043] In some embodiments, storage drives 171A-F can refer to any device configured to permanently record data, where "permanently" or "permanently" means that the device is able to retain the recorded data after a power outage. In some embodiments, storage drives 171A-F can correspond to non-disk storage media. For example, storage drives 171A-F can be one or more solid-state drives ('SSDs'), flash memory-based storage devices, any type of solid-state non-volatile memory, or any other type of non-mechanical storage device. In other embodiments, storage drives 171A-F can include mechanical or spinning hard disks, such as hard disk drives ('HDDs').

[0044] In some embodiments, storage array controllers 110A-D may be configured to offload device management responsibilities from storage drives 171A-F in storage arrays 102A-B. For example, storage array controllers 110A-D may manage control information that describes the status of one or more memory blocks in storage drives 171A-F. For example, the control information may indicate that a particular memory block has failed and should no longer be written to, that a particular memory block contains boot code from storage array controllers 110A-D, the number of program-erase ('P / E') cycles already executed on the particular memory block, the usage period of data stored in the particular memory block, the type of data stored in the particular memory block, and so on. In some embodiments, the control information may be stored as metadata along with the associated memory block. In other embodiments, the control information for storage drives 171A-F may be stored in a specific memory block of storage drives 171A-F selected by storage array controllers 110A-D. The selected memory block may be marked with an identifier indicating that the selected memory block contains control information. Identifiers are available for use by the memory array controllers 110A-D and memory drives 171A-F to quickly identify memory blocks containing control information. For example, the memory controllers 110A-D can issue commands to locate memory blocks containing control information. It can be noted that the control information may be large enough that portions of the control information can be stored in multiple locations, for example, the control information may be stored in multiple locations for redundancy purposes, or the control information may otherwise be distributed across multiple memory blocks in the memory drives 171A-F.

[0045] In some implementations, the memory array controllers 110A-D can offload device management responsibilities from the memory drives 171A-F of the memory array 102A-B by retrieving control information describing the state of one or more memory blocks in the memory drives 171A-F. Retrieving control information from the memory drives 171A-F can be performed, for example, by the memory array controllers 110A-D querying the location of control information for a specific memory drive 171A-F within the memory drives 171A-F. The memory drives 171A-F can be configured to execute instructions that cause the memory drives 171A-F to identify the location of the control information. These instructions can be executed by a controller (not shown) associated with or otherwise located on the memory drives 171A-F, and can cause the memory drives 171A-F to scan a portion of each memory block to identify the memory block storing the control information for the memory drives 171A-F. Storage drives 171A-F can respond by sending a response message containing the location of control information for storage drives 171A-F to storage array controllers 110A-D. Upon receiving the response message, storage array controllers 110A-D can issue a request to read data stored at the address associated with the location of the control information for storage drives 171A-F.

[0046] In other embodiments, the storage array controllers 110A-D may further offload device management responsibilities from the storage drives 171A-F by performing storage drive management operations in response to receiving control information. Storage drive management operations may include, for example, operations typically performed by the storage drives 171A-F (e.g., a controller (not shown) associated with a particular storage drive 171A-F). These operations may include, for example, ensuring that data is not written to a faulty memory block within the storage drive 171A-F, ensuring that data is written to memory blocks within the storage drive 171A-F to achieve sufficient wear leveling, and so on.

[0047] In some implementations, storage arrays 102A-B may implement two or more storage array controllers 110A-D. For example, storage array 102A may include storage array controller 110A and storage array controller 110B. At any given time, a single storage array controller 110A-D of storage system 100 (e.g., storage array controller 110A) may be designated with a primary state (also referred to herein as a "primary controller"), and other storage array controllers 110A-D (e.g., storage array controller 110A) may be designated with a secondary state (also referred to herein as a "secondary controller"). The primary controller may have specific permissions, such as the permission to modify data in persistent storage resources 170A-B (e.g., the permission to write data to persistent storage resources 170A-B). At least some permissions of the primary controller may override the permissions of the secondary controller. For example, when the primary controller has the permission to modify data in persistent storage resources 170A-B, the secondary controller may not have this permission. The state of storage array controllers 110A-D may change. For example, storage array controller 110A can be specified with a secondary state, and storage array controller 110B can be specified with a primary state.

[0048] In some implementations, for example, a primary controller of storage array controller 110A may act as a primary controller for one or more storage arrays 102A-B, and a second controller, for example, storage array controller 110B, may act as a secondary controller for the one or more storage arrays 102A-B. For example, storage array controller 110A may be a primary controller for storage arrays 102A and 102B, and storage array controller 110B may be a secondary controller for storage arrays 102A and 102B. In some implementations, storage array controllers 110C and 110D (also referred to as "storage processing modules") may not have both primary and secondary states. Storage array controllers 110C and 110D, implemented as storage processing modules, may act as communication interfaces between the primary and secondary controllers (e.g., storage array controllers 110A and 110B, respectively) and storage array 102B. For example, storage array controller 110A of storage array 102A may send write requests to storage array 102B via SAN 158. Write requests can be received by both storage array controllers 110C and 110D of storage array 102B. Storage array controllers 110C and 110D facilitate communication, such as sending write requests to the appropriate storage drives 171A-F. It can be noted that in some embodiments, the storage processing module can be used to increase the number of storage drives controlled by the primary and secondary controllers.

[0049] In some embodiments, memory array controllers 110A-D are communicatively coupled to one or more memory drives 171A-F and one or more NVRAM devices (not shown) included as part of memory arrays 102A-B via a midplane (not shown). Memory array controllers 110A-D may be coupled to the midplane via one or more data communication links, and the midplane may be coupled to the memory drives 171A-F and NVRAM devices via one or more data communication links. The data communication links described herein are collectively shown as data communication links 108A-D and may include, for example, a peripheral component interconnect high-speed ('PCIe' bus.

[0050] Figure 1B An example system for data storage according to some implementation schemes is shown. Figure 1B The storage array controller 101 shown can be similar to that described above. Figure 1A The storage array controllers 110A-D are described. In one example, storage array controller 101 may be similar to storage array controller 110A or storage array controller 110B. For illustrative and not limiting purposes, storage array controller 101 includes a number of elements. It will be noted that storage array controller 101 may include the same, more, or fewer elements configured in the same or different manner in other embodiments. It will be noted that the following may include... Figure 1A The components are described to help illustrate the features of the storage array controller 101.

[0051] The memory array controller 101 may include one or more processing devices 104 and random access memory ('RAM') 111. The processing device 104 (or controller 101) represents one or more general-purpose processing devices, such as a microprocessor, a central processing unit, etc. More specifically, the processing device 104 (or controller 101) may be a Complex Instruction Set Computing ('CISC') microprocessor, a Reduced Instruction Set Computing ('RISC') microprocessor, a Very Long Instruction Word ('VLIW') microprocessor, or a processor implementing other instruction sets or a combination of instruction sets. The processing device 104 (or controller 101) may also be one or more special-purpose processing devices, such as an ASIC, an FPGA, a digital signal processor ('DSP'), a network processor, etc.

[0052] Processing device 104 can be connected to RAM 111 via data communication link 106, which can be embodied as a high-speed memory bus, such as a dual data rate 4 ('DDR4') bus. Operating system 112 is stored in RAM 111. In some embodiments, instructions 113 are stored in RAM 111. Instructions 113 may include computer program instructions for performing operations in a direct-mapped flash memory system. In one embodiment, a direct-mapped flash memory system is a system that directly addresses data blocks within a flash drive without requiring address translation performed by the flash drive's memory controller.

[0053] In some embodiments, the storage array controller 101 includes one or more host bus adapters 103A-C coupled to the processing device 104 via data communication links 105A-C. In some embodiments, the host bus adapters 103A-C may be computer hardware that connects a host system (e.g., the storage array controller) to other networks and storage arrays. In some examples, the host bus adapters 103A-C may be Fibre Channel adapters enabling the storage array controller 101 to connect to a SAN, Ethernet adapters enabling the storage array controller 101 to connect to a LAN, and so on. The host bus adapters 103A-C may be coupled to the processing device 104 via, for example, a PCIe bus data communication link 105A-C.

[0054] In some implementations, the storage array controller 101 may include a host bus adapter 114 coupled to an extender 115. The extender 115 can be used to attach a host system to a larger number of storage drives. For example, in an implementation where the host bus adapter 114 is embodied as a SAS controller, the extender 115 may be a SAS extender for attaching the host bus adapter 114 to storage drives.

[0055] In some implementations, the storage array controller 101 may include a switch 116 coupled to the processing device 104 via a data communication link 109. The switch 116 may be a computer hardware device that can create multiple endpoints from a single endpoint, allowing multiple devices to share that single endpoint. For example, the switch 116 may be a PCIe switch coupled to a PCIe bus (e.g., data communication link 109) and presenting multiple PCIe connection points to a midplane.

[0056] In some implementations, the storage array controller 101 includes a data communication link 107 for coupling the storage array controller 101 to other storage array controllers. In some examples, the data communication link 107 may be a Fast Path Interconnect (QPI) interconnect.

[0057] A conventional storage system using a conventional flash drive can implement a process on the flash drive, which is part of the conventional storage system. For example, a higher-level process of the storage system can initiate and control a process on the flash drive. However, the flash drive of a conventional storage system may contain its own storage controller, which also executes the process. Therefore, for a conventional storage system, both higher-level processes (e.g., initiated by the storage system) and lower-level processes (e.g., initiated by the storage system's storage controller) can be executed.

[0058] To address the shortcomings of traditional storage systems, operations can be performed by higher-level processes instead of lower-level processes. For example, a flash storage system can include flash drives that do not contain a storage controller providing the processes. Therefore, the flash storage system's own operating system can initiate and control the processes. This can be achieved through a direct-mapped flash storage system that directly addresses data blocks within the flash drive without requiring address translation performed by the flash drive's storage controller.

[0059] In some implementations, storage drives 171A-F can be one or more partitioned storage devices. In some implementations, the one or more partitioned storage devices can be shingled HDDs. In some implementations, the one or more storage devices can be flash-based SSDs. In a partitioned storage device, the partition namespace on the partitioned storage device can be addressed by block groups, which are grouped and arranged according to their natural size to form several addressable regions. In some implementations utilizing SSDs, the natural size can be based on the SSD's erase block size. In some implementations, the regions of the partitioned storage device can be defined during the initialization of the partitioned storage device. In some implementations, the regions can be dynamically defined when data is written to the partitioned storage device.

[0060] In some implementations, zones can be heterogeneous, with some zones being page groups and others being multiple page groups. In some implementations, some zones may correspond to erase blocks, while others may correspond to multiple erase blocks. In implementations, a zone can be any combination of different numbers of pages in page groups and / or erase blocks, for heterogeneous mixtures of storage device programming modes, manufacturers, product types, and / or product generations, such as those applied to heterogeneous assemblies, upgrades, distributed storage, etc. In some implementations, a zone can be defined as having usage characteristics, such as attributes that support data with a specific lifetime (e.g., very short lifetime or very long lifetime). Partitioned storage devices can use these attributes to determine how to manage zones within their expected lifetime.

[0061] It should be understood that a region is a virtual construct. No particular region may have a fixed location on the storage device. Before allocation, a region may not have any location on the storage device. A region may correspond to a number representing a block of virtual allocatable space, which in various implementations is the size of an erase block or other block size. When the system allocates or opens a region, the region is allocated to flash or other solid-state storage, and when the system writes to the region, pages are written to the mapped flash or other solid-state storage of the partitioned storage device. When the system closes the region, the associated erase block or other block size is completed. At some point in the future, the system may delete a region, which will free up the allocated space for that region. During its lifetime, a region may be moved to different locations on the partitioned storage device, for example, when the partitioned storage device undergoes internal maintenance.

[0062] In some implementations, zones of a partitioned storage device can be in different states. A zone can be in an empty state, where no data has yet been stored in the zone. An empty zone can be explicitly opened or implicitly opened by writing data to the zone. This is the initial state of a zone on a fresh partitioned storage device, but it can also be the result of a zone reset. In some implementations, the empty zone may have a designated location within the flash memory of the partitioned storage device. In implementations, the location of the empty zone can be selected when the zone is first opened or first written to (or subsequently selected if the write buffer is in memory). A zone can be implicitly or explicitly open, where an open zone can be written to store data using write or append commands. In implementations, an open zone can also be written to using copy commands that copy data from different zones. In some implementations, the partitioned storage device may have a limit on the number of open zones at a given time.

[0063] A closed region is a region that has been partially written to but entered the closed state after an explicit close operation was issued. Closed regions can be reserved for future writes, but some runtime overhead can be reduced by keeping the region open. In some implementations, partitioned storage devices may limit the number of closed regions at any given time. A fully open region is a region that is storing data and can no longer be written to. A region may be fully full after a write operation has written all the data to it, or because the region has completed an operation. Before an operation completes, a region may have been or may not have been fully written to. However, after an operation completes, if a region reset operation is not performed first, it may not be possible to further open or write to this region.

[0064] The mapping from a region to an erase block (or to a shingled track in an HDD) can be arbitrary, dynamic, and hidden from view. Opening a region can be an operation that allows a new region to be dynamically mapped to the underlying storage of the partitioned storage device, and then allows data to be written to the region by appending writes until the region reaches its capacity. The region can be completed at any time, after which no further data may be possible to write to it. When the data stored in the region is no longer needed, the region can be reset, which effectively removes the region's contents from the partitioned storage device, making the physical storage held by the region available for subsequent data storage. Once a region has been written to and completed, the partitioned storage device ensures that the data stored in the region is not lost until the region is reset. During the time between writing data to a region and resetting the region, as part of maintenance operations within the partitioned storage device, regions can be moved between shingled tracks or erase blocks, for example, to keep data refreshed by copying data or to handle memory cell aging in an SSD.

[0065] In some implementations utilizing HDDs, a zone reset can allow shingled tracks to be allocated to new open zones that may open at some point in the future. In some implementations utilizing SSDs, a zone reset may cause the associated physical erase blocks of the zone to be erased and subsequently reused for storing data. In some implementations, partitioned storage devices may limit the number of open zones at a given point in time to reduce the amount of overhead dedicated to keeping zones open.

[0066] The operating system of a flash memory system can identify and maintain a list of allocation units across multiple flash drives in the flash memory system. An allocation unit can be an entire erase block or multiple erase blocks. The operating system can also maintain mappings or address ranges that directly map addresses to erase blocks in the flash drives of the flash memory system.

[0067] A direct mapping to erase blocks in the flash drive can be used for both rewriting and erasing data. For example, operations can be performed on one or more allocation units containing first and second data, where the first data is to be retained and the second data is no longer available to the flash storage system. The operating system can initiate this process to write the first data to a new location within other allocation units, erase the second data, and mark the allocation unit as available for subsequent data. Therefore, this process can be performed solely by the higher-level operating system of the flash storage system, without requiring additional lower-level processes performed by the flash drive's controller.

[0068] The advantages of this process, executed solely by the operating system of the flash memory system, include increased reliability of the flash drives because no unnecessary or redundant write operations are performed during the process. A potentially novel aspect is the concept of initiating and controlling the process on the operating system of the flash memory system. Furthermore, this process can be controlled by the operating system across multiple flash drives. This contrasts with the process being executed by the storage controller of the flash drives.

[0069] A storage system can consist of two storage array controllers that share a set of drives for failover purposes, or it can consist of a single storage array controller that provides storage services using multiple drives, or it can consist of a distributed network of storage array controllers, each with a certain number of drives or a certain number of flash storage, wherein the storage array controllers in the network cooperate to provide complete storage services and cooperate in all aspects of the storage services, including storage allocation and garbage collection.

[0070] Figure 1C A third example system 117 for data storage according to some embodiments is shown. For illustrative and not limiting purposes, system 117 (also referred to herein as a “storage system”) includes a number of elements. It will be noted that system 117 may include the same, more or fewer elements configured in the same or different ways in other embodiments.

[0071] In one embodiment, system 117 includes a dual peripheral component interconnect ('PCI') flash memory device 118 having individually addressable fast write memory. System 117 may include a memory device controller 119. In one embodiment, memory device controllers 119A-D may be a CPU, ASIC, FPGA, or any other circuit system that can implement the control structure required by this disclosure. In one embodiment, system 117 includes flash memory devices (e.g., flash memory devices 120a-n) operatively coupled to various channels of memory device controller 119. Flash memory devices 120a-n may be presented to controllers 119A-D as an addressable set of flash pages, erase blocks, and / or control elements sufficient to allow memory device controllers 119A-D to program and retrieve various aspects of the flash. In one embodiment, the storage device controller 119A-D can perform operations on the flash memory devices 120a-n, including storing and retrieving the data content of pages, arranging and erasing any blocks, tracking statistics related to the use and reuse of flash memory pages, erase blocks and cells, tracking and predicting error codes and faults within the flash memory, controlling voltage levels associated with the programming and retrieval of flash cell contents, and so on.

[0072] In one embodiment, system 117 may include RAM 121 for storing individually addressable, fast-write data. In one embodiment, RAM 121 may be one or more separate discrete devices. In another embodiment, RAM 121 may be integrated into memory device controllers 119A-D or multiple memory device controllers. RAM 121 may also be used for other purposes, such as as temporary program memory for processing devices (e.g., CPU) within memory device controller 119.

[0073] In one embodiment, system 117 may include an energy storage device 122, such as a rechargeable battery or capacitor. The energy storage device 122 may store enough energy to power the storage device controller 119, a number of RAMs (e.g., RAM 121), and a number of flash memories (e.g., flash memories 120a-120n), allowing them sufficient time to write the contents of the RAM to the flash memories. In one embodiment, if the storage device controller detects a loss of external power, then the storage device controllers 119A-D may write the contents of the RAM to the flash memories.

[0074] In one embodiment, system 117 includes two data communication links 123a and 123b. In one embodiment, data communication links 123a and 123b may be PCI interfaces. In another embodiment, data communication links 123a and 123b may be based on other communication standards (e.g., HyperTransport, Infinite Bandwidth, etc.). Data communication links 123a and 123b may be based on Non-Volatile Memory High Speed ​​('NVMe') or Architecture-based NVMe ('NVMf') specifications, which allow external connection to storage device controllers 119A-D from other components in storage system 117. It should be noted that, for convenience, data communication links may be interchangeably referred to herein as PCI buses.

[0075] System 117 may also include an external power supply (not shown), which may be provided via one or both data communication links 123a, 123b, or may be provided independently. Alternative embodiments include a separate flash memory (not shown) dedicated to storing the contents of RAM 121. Storage device controllers 119A-D may present logical devices on a PCI bus, which may include addressable fast write logic devices, or different portions of the logical address space of storage device 118, which may present as PCI memory or persistent storage. In one embodiment, operations to store into the device are directed to RAM 121. In the event of a power failure, storage device controllers 119A-D may write the storage contents associated with the addressable fast write logic to flash memory (e.g., flash memory 120a-n) for long-term persistent storage.

[0076] In one embodiment, the logic device may include a representation of some or all of the contents of flash memory devices 120a-n, wherein the representation allows a storage system (e.g., storage system 117) including storage device 118 to directly address flash memory pages and directly reprogram erase blocks of storage system components outside the storage device via the PCI bus. The representation may also allow one or more external components to control and retrieve other aspects of the flash memory, including some or all of the following: tracking statistics related to the use and reuse of flash memory pages, erase blocks, and cells across all flash memory devices; tracking and predicting error codes and faults within and across flash memory devices; controlling voltage levels associated with the programming and retrieval of flash cell contents; and so on.

[0077] In one embodiment, the energy storage device 122 may be sufficient to ensure the completion of ongoing operations on the flash memory devices 120a-120n. The energy storage device 122 can power the memory device controllers 119A-D and the associated flash memory devices (e.g., 120a-n) for these operations, as well as for storing fast write RAM into the flash memory. The energy storage device 122 can be used to store accumulated statistics and other parameters maintained and tracked by the flash memory devices 120a-n and / or the memory device controller 119. Individual capacitors or energy storage devices (e.g., smaller capacitors located near or embedded within the flash memory devices themselves) may be used for some or all of the operations described herein.

[0078] Various methods can be used to track and optimize the lifespan of storage energy components, such as adjusting voltage levels over time, partially discharging the storage energy device 122 to measure the corresponding discharge characteristics, etc. If available energy decreases over time, the effective available capacity of the addressable fast write storage may decrease to ensure that it can be safely written based on the currently available storage energy.

[0079] Figure 1D A third example storage system 124 for data storage is illustrated according to some embodiments. In one embodiment, storage system 124 includes storage controllers 125a, 125b. In one embodiment, storage controllers 125a, 125b are operatively coupled to dual PCI storage devices. Storage controllers 125a, 125b may be operatively coupled (e.g., via storage network 130) to a number of host computers 127a-n.

[0080] In one embodiment, two storage controllers (e.g., 125a and 125b) provide storage services, such as an SCS block storage array, a file server, an object server, a database, or data analytics services. Storage controllers 125a and 125b can provide services to host computers 127a-n outside the storage system 124 via a number of network interfaces (e.g., 126a-d). Storage controllers 125a and 125b can provide integrated services or applications entirely within the storage system 124, forming a converged storage and computing system. Storage controllers 125a and 125b can utilize fast write memory within or across storage devices 119a-d to record ongoing operations, ensuring that these operations are not lost in the event of power failure, storage controller removal, shutdown of the storage controllers or the storage system, or failure of one or more software or hardware components within the storage system 124.

[0081] In one embodiment, storage controllers 125a and 125b function as PCI masters of one or more PCI buses 128a and 128b. In another embodiment, 128a and 128b may be based on other communication standards (e.g., HyperTransport, Infinite Bandwidth, etc.). Other storage system embodiments may operate storage controllers 125a and 125b as multiple masters of both PCI buses 128a and 128b. Alternatively, a PCI / NVMe / NVMe switching infrastructure or structure may connect multiple storage controllers. Some storage system embodiments may allow storage devices to communicate directly with each other, rather than only with storage controllers. In one embodiment, storage device controller 119a may operate under the instruction of storage controller 125a to retrieve data from RAM (e.g., ...). Figure 1C The data in RAM 121 is synthesized and transferred to the flash memory device. For example, after the memory controller determines that the operation has been fully committed across the memory system, or when the fast write memory on the device reaches a certain usage capacity, or after a specific amount of time, a recalculated version of the RAM contents can be transferred to ensure improved data security or to release addressable fast write capacity for reuse. For example, this mechanism can be used to avoid a second transfer from memory controllers 125a, 125b via a bus (e.g., 128a, 128b). In one embodiment, recalculation may include compressing data, adding indexes or other metadata, combining multiple data segments together, performing erase code calculations, etc.

[0082] In one embodiment, under the instruction of storage controllers 125a and 125b, storage device controllers 119a and 119b can be used to determine the contents stored in RAM (e.g., ...). Figure 1CThe data in RAM 121) is processed and transferred to other storage devices without the involvement of storage controllers 125a and 125b. This operation can be used to mirror data stored in one storage controller 125a to another storage controller 125b, or it can be used to offload compression, data aggregation, and / or erase encoding calculations and transfers to storage devices to reduce the load on the storage controllers or the storage controller interfaces 129a and 129b to the PCI buses 128a and 128b.

[0083] Storage device controllers 119A-D may include mechanisms for implementing high availability primitives for use by other parts of the storage system outside the dual PCI storage device 118. For example, reservation or exclusion primitives may be provided so that in a storage system with two storage controllers providing highly available storage services, one storage controller can prevent the other storage controller from accessing or continuing to access the storage device. This could be used, for example, if a controller detects that the other controller is malfunctioning or that the interconnect between the two storage controllers itself may be malfunctioning.

[0084] In one embodiment, a storage system used with dual PCI direct-mapped storage devices having individually addressable fast-write storage includes a system for managing erase blocks or groups of erase blocks as allocation units, for storing data on behalf of the storage service, for storing metadata associated with the storage service (e.g., indexes, logs, etc.), or for properly managing the storage system itself. Flash pages may be only a few kilobytes in size and can be written as data arrives or after the storage system has held the data for a long period (e.g., exceeding a defined time threshold). To commit data faster or reduce the number of writes to the flash memory device, the storage controller may first write data to the individually addressable fast-write storage on one or more storage devices.

[0085] In one embodiment, storage controllers 125a and 125b may initiate the use of erase blocks within and across storage devices (e.g., 118) based on the usage period and expected remaining lifetime of the storage device, or based on other statistical data. Storage controllers 125a and 125b may also initiate garbage collection and data migration between storage devices based on pages that are no longer needed, manage the lifetime of flash pages and erase blocks, and manage overall system performance.

[0086] In one embodiment, storage system 124 may utilize mirroring and / or erasure coding schemes as part of storing data into addressable fast write storage and / or writing data into allocation units associated with erase blocks. Erasure codes may be used across storage devices, within erase blocks or allocation units, or within and across flash memory devices on a single storage device, to provide redundancy against failures of single or multiple storage devices, or to prevent internal damage to flash memory pages due to flash memory operation or flash memory cell degradation. Different levels of mirroring and erasure coding can be used to recover from multiple failures occurring individually or in combination.

[0087] refer to Figure 2A The embodiment described in -G illustrates a storage cluster storing user data, such as user data originating from one or more user or client systems or other sources outside the storage cluster. The storage cluster distributes user data among storage nodes housed within a chassis or across multiple chassis using erasure coding and redundant copies of metadata. Erasure coding is a data protection or reconstruction method where data is stored across a set of different locations, such as disks, storage nodes, or geographic locations. Flash memory is a type of solid-state memory that can be integrated with the embodiment, but the embodiment can be extended to other types of solid-state memory or other storage media, including non-solid-state memory. Control of storage locations and workloads is distributed across storage locations in the cluster peer system. Tasks such as regulating communication between storage nodes, detecting when storage nodes are unavailable, and balancing I / O (input and output) on storage nodes are all processed on a distributed basis. In some embodiments, data is arranged or distributed across multiple storage nodes in the form of data segments or stripes that support data recovery. Ownership of data can be reallocated within the cluster, regardless of input and output patterns. The architecture described in more detail below allows for the failure of storage nodes in the cluster, but the system remains operational because data can be reconstructed from other storage nodes, thus remaining available for input and output operations. In various embodiments, storage nodes may be referred to as cluster nodes, blades, or servers.

[0088] Storage clusters can be housed within a chassis, i.e., an enclosure that houses one or more storage nodes. The chassis contains mechanisms for powering each storage node, such as a power distribution bus, and communication mechanisms, such as a communication bus that enables communication between storage nodes. According to some embodiments, the storage cluster can operate as a standalone system in one location. In one embodiment, the chassis houses at least two instances of the power distribution and communication buses that can be independently enabled or disabled. The internal communication bus can be an Ethernet bus; however, other technologies, such as PCIe, wireless, etc., are also applicable. The chassis provides ports for external communication buses for communication between multiple chassis and with client systems, either directly or via switches. External communication can use technologies such as Ethernet, wireless, Fibre Channel, etc. In some embodiments, the external communication bus uses different communication bus technologies for inter-chassis and client communication. If switches are deployed within or between chassis, the switches can act as converters between various protocols or technologies. When multiple chassis are connected to define a storage cluster, clients can access the storage cluster using proprietary or standard interfaces, such as Network File System ('NFS'), Common Internet File System ('CIFS'), Small Computer System Interface ('SCSI'), or Hypertext Transfer Protocol ('HTTP')). Translation from the client protocol may occur at the switch, at the external communication bus of the chassis, or within each storage node. In some embodiments, multiple chassis may be coupled or connected to each other via an aggregator switch. Part and / or all of the coupled or connected chassis can be represented as a storage cluster. As discussed above, each chassis may have multiple blades, each with a Media Access Control ('MAC') address; however, in some embodiments, the storage cluster is presented to the external network as having a single cluster IP address and a single MAC address.

[0089] Each storage node can be one or more storage servers, and each storage server is connected to one or more non-volatile solid-state memory (SSD) cells, which may be referred to as storage cells or storage devices. One embodiment includes a single storage server and one to eight SSD cells in each storage node, but this example is not intended to be limiting. The storage server may include a processor, DRAM, and interfaces for internal communication buses and power distribution for each power bus. In some embodiments, within the storage node, interfaces and storage cells share a communication bus, such as PCI High Speed. SSD cells can directly access the internal communication bus interface via the storage node's communication bus, or request access to the bus interface from the storage node. The SSD cell houses an embedded CPU, a solid-state storage controller, and a number of solid-state high-capacity storage devices, for example, between 2 and 32 terabytes ('TB') in some embodiments. The SSD cell includes embedded volatile storage media such as DRAM and energy storage devices. In some embodiments, the energy storage devices are capacitors, supercapacitors, or batteries, capable of transferring a subset of the DRAM content to a stable storage medium in the event of a power outage. In some embodiments, the non-volatile solid-state memory cell is composed of a storage class memory such as phase-change or magnetoresistive random access memory ('MRAM'), which can replace DRAM and achieve reduced device hold power.

[0090] One of the many characteristics of storage nodes and non-volatile solid-state storage devices (NSSSDs) is their ability to proactively reconstruct data within a storage cluster. Storage nodes and NSSSDs can determine when a storage node or NSSSD in the storage cluster becomes inaccessible, regardless of whether an attempt has been made to read data involving that storage node or NSSSD. The storage nodes and NSSSDs then cooperate to recover and reconstruct the data in at least a partially new location. This constitutes proactive reconstruction because the system reconstructs data without waiting for a read access initiated from a client system employing the storage cluster to require the data. These and further details of storage memory and its operation are discussed below.

[0091] Figure 2AThis is a perspective view of a storage cluster 161 according to some embodiments, having a plurality of storage nodes 150 and internal solid-state memory coupled to each storage node to provide network-attached storage or a storage area network. Network-attached storage, a storage area network, or a storage cluster or other storage memory may comprise one or more storage clusters 161, each storage cluster having one or more storage nodes 150, arranged in a flexible and reconfigurable manner of physical components and the amount of storage memory thus provided. Storage clusters 161 are designed to be mounted in racks, and one or more racks can be set up and filled for storage memory as needed. Storage cluster 161 has a chassis 138 with a plurality of slots 142. It should be understood that chassis 138 may be referred to as a housing, enclosure, or rack unit. In one embodiment, chassis 138 has fourteen slots 142, but other numbers of slots are readily designed. For example, some embodiments have four slots, eight slots, sixteen slots, thirty-two slots, or other suitable numbers of slots. In some embodiments, each slot 142 may accommodate one storage node 150. The chassis 138 includes fins 148 for rack mounting. A fan 144 provides air circulation to cool the storage nodes 150 and their components, but other cooling components may be used, or embodiments without cooling components may be designed. A switching structure 146 couples the storage nodes 150 within the chassis 138 together and to a network for communication with the memory. In the embodiment depicted herein, for illustrative purposes, slot 142 to the left of the switching structure 146 and fan 144 is shown as occupied by a storage node 150, while slot 142 to the right of the switching structure 146 and fan 144 is empty and available for insertion of a storage node 150. This configuration is an example; one or more storage nodes 150 may occupy slot 142 in various other arrangements. In some embodiments, the storage node arrangement does not need to be sequential or adjacent. The storage nodes 150 are hot-swappable, meaning that a storage node 150 can be inserted into or removed from slot 142 in the chassis 138 without stopping the system or powering it off. When storage node 150 is inserted into or removed from slot 142, the system automatically reconfigures to recognize and adapt to the change. In some embodiments, reconfiguration includes restoring redundancy and / or rebalancing data or load.

[0092] Each storage node 150 may have multiple components. In the embodiment shown here, storage node 150 includes a printed circuit board 159 filled with a CPU 156 (i.e., a processor), a memory 154 coupled to the CPU 156, and a non-volatile solid-state storage device 152 coupled to the CPU 156; however, other mounting components and / or other components may be used in other embodiments. The memory 154 has instructions executed by the CPU 156 and / or data operated by the CPU 156. As explained further below, the non-volatile solid-state storage device 152 includes flash memory, or in other embodiments, other types of solid-state memory.

[0093] refer to Figure 2A Storage cluster 161 is scalable, meaning that, as described above, storage capacity with non-uniform storage sizes can be easily added. In some embodiments, one or more storage nodes 150 can be inserted into or removed from each chassis, and the storage cluster can be self-configurable. Insertable storage nodes 150, whether installed in the chassis at delivery or added later, can have different sizes. For example, in one embodiment, storage nodes 150 can have any multiple of 4 TB, such as 8TB, 12TB, 16TB, 32TB, etc. In other embodiments, storage nodes 150 can have any other multiple of storage amount or capacity. The storage capacity of each storage node 150 is broadcast and influences decisions on how data is striped. To achieve maximum storage efficiency, embodiments can be self-configurable as wide as possible in striping, but must meet predetermined requirements for continuous operation, where at most one or two cells of non-volatile solid-state storage devices 152 or storage nodes 150 are lost within the chassis.

[0094] Figure 2B This is a block diagram illustrating a communication interconnect 173 and a power distribution bus 172 coupling multiple storage nodes 150. (Return to Reference) Figure 2A In some embodiments, the communication interconnect 173 may be included in or implemented using the switching structure 146. In some embodiments, where multiple storage clusters 161 occupy a single rack, the communication interconnect 173 may be included in or implemented using a top-of-rack switch. Figure 2B As shown, storage cluster 161 is enclosed within a single chassis 138. External port 176 is coupled to storage node 150 via communication interconnect 173, while external port 174 is directly coupled to the storage node. External power port 178 is coupled to power distribution bus 172. Storage node 150 may contain varying numbers and capacities of non-volatile solid-state storage devices 152, as shown in reference [reference missing]. Figure 2A As described above. Additionally, one or more storage nodes 150 may be... Figure 2BThe diagram shows only the computed storage nodes. Grant 168 is implemented on non-volatile solid-state storage device 152, for example as a list or other data structure stored in memory. In some embodiments, grants are stored within non-volatile solid-state storage device 152 and supported by software executing on the controller or other processor of non-volatile solid-state storage device 152. In another embodiment, grant 168 is implemented on storage node 150, for example as a list or other data structure stored in memory 154 and supported by software executing on the CPU 156 of storage node 150. In some embodiments, grant 168 controls the manner and location of data stored in non-volatile solid-state storage device 152. This control helps determine which type of erasure coding scheme is applied to the data, and which storage nodes 150 have which portions of the data. Each grant 168 may be assigned to non-volatile solid-state storage device 152. In various embodiments, each grant may control the range of inode numbers, segment numbers, or other data identifiers assigned to the data by the file system, storage node 150, or non-volatile solid-state storage device 152.

[0095] In some embodiments, each piece of data and each piece of metadata is redundant in the system. Additionally, each piece of data and each piece of metadata has an owner, which may be referred to as an authorizer. If the authorizer is inaccessible, for example due to a storage node failure, there is a series of plans on how to locate the data or the metadata. In various embodiments, redundant copies of authorizer 168 exist. In some embodiments, authorizer 168 is associated with storage node 150 and non-volatile solid-state storage device 152. Each authorizer 168, covering a range of data segment numbers or other identifiers of the data, can be assigned to a specific non-volatile solid-state storage device 152. In some embodiments, all these ranges of authorizer 168 are distributed across the non-volatile solid-state storage devices 152 of the storage cluster. Each storage node 150 has a network port for providing access to the non-volatile solid-state storage devices 152 of the storage node 150. Data may be stored in segments associated with segment numbers, which in some embodiments are indirect methods of RAID (Redundant Array of Independent Disks) striping configurations. Therefore, the assignment and use of authorizer 168 establishes indirect access to the data. According to some embodiments, indirect access may refer to the ability to indirectly reference data, in this case, via authorization 168. A segment identifies a group of non-volatile solid-state storage devices 152 and a local identifier within said group of non-volatile solid-state storage devices 152, said local identifier may contain data. In some embodiments, the local identifier is an offset within the device and may be reused sequentially by multiple segments. In other embodiments, the local identifier is unique for a particular segment and is never reused. The offset within the non-volatile solid-state storage device 152 is used to locate data for writing to or reading from the non-volatile solid-state storage device 152 (in the form of RAID striping). Data is striped across multiple cells of the non-volatile solid-state storage device 152, which may include or differ from the non-volatile solid-state storage device 152 with authorization 168 having a specific data segment.

[0096] If the location of a specific data segment changes, for example during data movement or data reconstruction, the authorization 168 of the data segment should be queried at the non-volatile solid-state storage device 152 or storage node 150 with the authorization 168. To locate a specific piece of data, embodiments calculate a hash value of the data segment or apply an inode number or data segment number. The output of this operation points to the non-volatile solid-state storage device 152 with the authorization 168 of the specific data. In some embodiments, this operation is divided into two phases. The first phase maps an entity identifier (ID) (e.g., a segment number, inode number, or directory number) to an authorization identifier. This mapping may involve calculations such as hashing or bitmasking. The second phase is mapping the authorization identifier to a specific non-volatile solid-state storage device 152, which can be done through explicit mapping. This operation is repeatable, such that when the calculation is performed, the result reliably and repeatedly points to the specific non-volatile solid-state storage device 152 with the authorization 168. This operation may include a set of accessible storage nodes as input. If the set of accessible non-volatile solid-state storage cells changes, the optimal set will also change. In some embodiments, the persistent value is the current assignment (always true), and the computed value is the target assignment that the cluster will attempt to reconfigure. This computation can be used to determine the optimal non-volatile solid-state storage device 152 for authorization when a set of accessible non-volatile solid-state storage devices 152 constituting the same cluster exists. The computation also determines an ordered set of peer-to-peer non-volatile solid-state storage devices 152, which also records the mapping of authorizations to non-volatile solid-state storage devices so that authorizations can be determined even if the assigned non-volatile solid-state storage device is inaccessible. In some embodiments, if a particular authorization 168 is unavailable, a replica or alternative authorization 168 can be queried.

[0097] refer to Figure 2A and 2BTwo of the many tasks of the CPU 156 on storage node 150 are splitting and writing data and reassembling and reading data. When the system determines that data needs to be written, the authorization 168 for the data is located as described above. When the segment ID of the data has been determined, the write request is forwarded to the non-volatile solid-state storage device 152 of the host currently identified as the authorization 168 determined according to the segment. The host CPU 156 of storage node 150 (on which non-volatile solid-state storage device 152 and the corresponding authorization 168 reside) then splits or fragments the data and transmits the data to various non-volatile solid-state storage devices 152. The transmitted data is written as data stripes according to the erasure coding scheme. In some embodiments, data is requested to be pulled, while in other embodiments, data is pushed. Conversely, when reading data, as described above, the authorization 168 containing the segment ID of the data is located. The host CPU 156 of storage node 150 on which non-volatile solid-state storage device 152 and the corresponding authorization 168 reside request data from the non-volatile solid-state storage device and the corresponding storage node to which the authorization points. In some embodiments, data is read as a data stripe from a flash storage device. The host CPU 156 of storage node 150 then reassembles the read data, corrects any errors (if any) according to an appropriate erasure encoding scheme, and forwards the reassembled data to the network. In other embodiments, some or all of these tasks may be processed in non-volatile solid-state storage device 152. In some embodiments, a segment host requests data to be sent to storage node 150 by requesting a page from the storage device and then sending the data to the storage node that issued the original request.

[0098] In one embodiment, authorization 168 operates to determine how the operation will be performed on a particular logical element. Each logical element can be operated through a specific authorization on multiple storage controllers of the storage system. Authorization 168 can communicate with multiple storage controllers, causing the multiple storage controllers to jointly perform operations on these particular logical elements.

[0099] In embodiments, a logical element may be, for example, a file, directory, object bucket, individual objects, descriptive portions of a file or object, other forms of key-value database, or table. In embodiments, performing operations may involve, for example, ensuring consistency, structural integrity, and / or recoverability with other operations targeting the same logical element, reading metadata and data associated with the logical element, determining which data should be persistently written to the storage system to preserve any changes to the operation, or determining the location of metadata and data stored on modular storage devices attached to multiple storage controllers in the storage system.

[0100] In some embodiments, operations are token-based transactions to enable efficient communication within a distributed system. Each transaction may be accompanied by or associated with a token that provides permission to execute the transaction. In some embodiments, authorization 168 can maintain the system's pre-transaction state until the operation is complete. Token-based communication can be completed without global locks on the system and can also restart operations in the event of interruption or other failures.

[0101] In some systems, such as UNIX-like file systems, data is handled by index nodes (or inodes), which specify the data structure representing objects in the file system. For example, an object can be a file or a directory. Metadata may accompany an object, such as permissions data and creation timestamps, as well as other attributes. Segment numbers can be assigned to all or part of such objects in the file system. In other systems, data segments are handled by segment numbers assigned elsewhere. For the purposes of discussion, a distribution unit is an entity, which can be a file, directory, or segment. That is, an entity is a unit of data or metadata stored by the storage system. Entities are grouped into sets called licenses. Each license has a license owner, which is a storage node with exclusive rights to update entities within the license. In other words, a storage node contains licenses, and those licenses contain entities.

[0102] According to some embodiments, a segment is a logical container for data. A segment is an address space between the media address space and the physical flash location, where the data segment number resides. Segments may also contain metadata, which allows data redundancy to be recovered without involving higher-level software (by rewriting to different flash locations or devices). In one embodiment, the internal format of a segment contains client data and a media mapping to determine the location of the data. Where applicable, for example, segments are protected from memory and other failures by decomposing each data segment into several data and parity fragments. Depending on the erasure coding scheme, the data and parity fragments are distributed across the non-volatile solid-state storage device 152 coupled to the host CPU 156, i.e., striping (see [link to erasure coding scheme]). Figure 2E and 2G In some embodiments, the term "segment" refers to a container and its location in the segment's address space. According to some embodiments, the term "strip" refers to a group of fragments identical to a segment, including how the fragments are distributed along with redundancy or parity information.

[0103] A series of address space transformations occur throughout the storage system. At the top are directory entries (filenames) linked to inodes. Inodes point to the media address space, where data is logically stored. Media addresses can be mapped through a series of indirect media to distribute the load of large files or to implement data services such as deduplication or snapshots. Segment addresses are then translated into physical flash locations. According to some embodiments, physical flash locations have address ranges defined by the amount of flash memory in the system. Media addresses and segment addresses are logical containers, and in some embodiments, 128-bit or larger identifiers are used, making the practically unlimited, reusable probability calculated to be longer than the system's expected lifetime. In some embodiments, addresses from logical containers are allocated in a hierarchical manner. Initially, each cell of the non-volatile solid-state storage device 152 can be assigned an address space range. Within this assigned range, the non-volatile solid-state storage device 152 can allocate addresses without synchronization with other non-volatile solid-state storage devices 152.

[0104] Data and metadata are stored using a set of underlying storage layouts optimized for different workload patterns and storage devices. These layouts incorporate various redundancy schemes, compression formats, and indexing algorithms. Some layouts store information about licenses and license masters, while others store file metadata and file data. Redundancy schemes include error correction codes tolerating damaged bits within a single storage device (e.g., a NAND flash chip), erase codes tolerating failures of multiple storage nodes, and replication schemes tolerating data center or region failures. In some embodiments, low-density parity-check ('LDPC') codes are used within a single storage cell. In some embodiments, Reed-Solomon coding is used within the storage cluster, and mirroring is used within the storage grid. Metadata can be stored using ordered log-structured indexes (e.g., log-structured merge trees), although large datasets may not be stored in log-structured layouts.

[0105] To maintain consistency across multiple copies of an entity, storage nodes implicitly agree on two things by computation: (1) the authorization containing the entity, and (2) the storage node containing the authorization. Entities can be assigned to authorizations through pseudo-random methods, by splitting entities into ranges based on externally generated keys, or by placing a single entity into each authorization. Examples of pseudo-random schemes are the RUSH hash series under linear hashing and scalable hashing, including CRUSH under scalable hashing. In some embodiments, pseudo-random assignment is used only to assign authorizations to nodes, as the set of nodes can change. The set of authorizations cannot change, so any subjective function can be applied in these embodiments. Some placement schemes automatically place authorizations on storage nodes, while others rely on an explicit mapping of authorizations to storage nodes. In some embodiments, a pseudo-random scheme is used to map each authorization to a set of candidate authorization owners. A pseudo-random data distribution function associated with CRUSH can assign authorizations to storage nodes and create a list of authorization assignment locations. Each storage node has a copy of the pseudo-random data distribution function and is accessible through the same computations used for distribution, followed by lookup or location of the grant. In some embodiments, each pseudo-random scheme requires a set of accessible storage nodes as input to arrive at the same target node. Once an entity is placed in the grant, it can be stored on the physical device to prevent unexpected data loss due to anticipated failures. In some embodiments, the rebalancing algorithm attempts to store copies of all entities within the grant in the same layout and on the same set of machines.

[0106] Examples of anticipated failures include equipment malfunction, machine theft, data center fires, and regional disasters such as nuclear accidents or geological events. Different failures will result in varying degrees of acceptable data loss. In some embodiments, theft of storage nodes will not affect the security and reliability of the system, while regional events may result in no data loss, missed updates for seconds or minutes, or even complete data loss, depending on the system configuration.

[0107] In these embodiments, placing data for storage redundancy is unrelated to placing licenses for data consistency. In some embodiments, storage nodes containing licenses do not contain any persistent storage devices. Instead, storage nodes are connected to non-volatile solid-state storage cells that do not contain licenses. The communication interconnects between storage nodes and non-volatile solid-state storage cells consist of various communication technologies and have non-uniform performance and fault-tolerant characteristics. In some embodiments, as mentioned above, non-volatile solid-state storage cells are connected to storage nodes via PCI high-speed connections, storage nodes are connected together within a single chassis using an Ethernet backplane, and the chassis are connected together to form a storage cluster. In some embodiments, the storage clusters are connected to clients using Ethernet or Fibre Channel. If multiple storage clusters are configured as a storage grid, then the multiple storage clusters are connected using the Internet or other long-distance networking links, such as “metropolitan area scale” links or dedicated links that do not traverse the Internet.

[0108] The authorized owner has exclusive rights to modify entities, migrate entities from one non-volatile solid-state storage unit to another, and add and delete entity copies. This allows for the maintenance of redundancy in the underlying data. When the authorized owner fails, is about to become unusable, or is overloaded, the authorization is transferred to a new storage node. Transient failures make it crucial to ensure that all fault-free machines agree on the new authorized location. Ambiguity arising from transient failures can be automated through consensus protocols (such as Paxos, hot-to-warm failover schemes), manual intervention by a remote system administrator, or by a local hardware administrator (such as physically removing a failed machine from the cluster or pressing a button on a failed machine). In some embodiments, failover is automatic using a consensus protocol. According to some embodiments, if too many failures or replication events occur in a short period of time, the system will enter a self-protection mode and interrupt replication and data movement activities until administrator intervention.

[0109] The system transmits messages between storage nodes and non-volatile solid-state storage units as authorizations are transferred between storage nodes and as authorization owners update entities within their authorizations. Regarding persistent messages, messages with different purposes have different types. Depending on the message type, the system maintains different ordering and durability guarantees. When processing persistent messages, messages are temporarily stored across various persistent and non-persistent storage hardware technologies. In some embodiments, messages are stored on RAM, NVRAM, and NAND flash devices, and various protocols are used to efficiently utilize each storage medium. Latency-sensitive client requests can persist in replicated NVRAM and then in NAND, while background rebalancing operations persist directly to NAND.

[0110] Persistent messages are persistently stored before transmission. This allows the system to continue serving client requests in the event of failures and component replacements. While many hardware components contain unique identifiers visible to system administrators, manufacturers, the hardware supply chain, and the ongoing monitoring and quality control infrastructure, applications running on the infrastructure virtualize addresses. These virtualized addresses remain unchanged throughout the storage system's lifetime, regardless of component failure or replacement. This allows each component of the storage system to be replaced over time without reconfiguration or disruption of client request processing—that is, the system supports non-disruptive upgrades.

[0111] In some embodiments, virtualized addresses are stored with sufficient redundancy. Continuous monitoring systems associate hardware and software status with hardware identifiers. This allows for the detection and prediction of failures due to faulty components and manufacturing details. In some embodiments, by removing components from the critical path, the monitoring system is also able to proactively transfer authorization and entities from the affected device before a failure occurs.

[0112] Figure 2C This is a multi-level block diagram illustrating the contents of storage node 150 and the contents of non-volatile solid-state storage devices 152 of storage node 150. In some embodiments, data is transferred to and from storage node 150 via a network interface controller ('NIC') 202. Each storage node 150 has a CPU 156 and one or more non-volatile solid-state storage devices 152, as discussed above. Figure 2C Moving down one level, each non-volatile solid-state storage device 152 has a relatively fast non-volatile solid-state memory, such as non-volatile random access memory ('NVRAM') 204, and flash memory 206. In some embodiments, NVRAM 204 may be a component that does not require programming / erasing cycles (DRAM, MRAM, PCM) and may be a memory capable of supporting write frequencies significantly higher than the frequency of reading from memory. Figure 2CMoving one level down, NVRAM 204 is implemented in one embodiment as a high-speed volatile memory, such as Dynamic Random Access Memory (DRAM) 216, backed up by an energy reserve 218. The energy reserve 218 provides sufficient power to keep the DRAM 216 powered for a sufficient period of time to transfer content to the flash memory 206 in the event of a power failure. In some embodiments, the energy reserve 218 is a capacitor, supercapacitor, battery, or other device that provides adequate power to transfer the content of the DRAM 216 to a stable storage medium in the event of a power outage. The flash memory 206 is implemented as a plurality of flash dies 222, which may be referred to as a package of flash dies 222 or an array of flash dies 222. It should be understood that the flash dies 222 can be packaged in any number of ways, including each package containing a single die, each package containing multiple dies (i.e., a multi-chip package), hybrid packages, as dies on a printed circuit board or other substrate, as encapsulated dies, etc. In the illustrated embodiment, the non-volatile solid-state storage device 152 has a controller 212 or other processor and an input / output (I / O) port 210 coupled to the controller 212. The I / O port 210 is coupled to a network interface controller 202 of a CPU 156 and / or a flash memory node 150. A flash input / output (I / O) port 220 is coupled to a flash die 222, and a direct memory access unit (DMA) 214 is coupled to the controller 212, DRAM 216, and flash die 222. In the illustrated embodiment, the I / O port 210, controller 212, DMA unit 214, and flash I / O port 220 are implemented on a programmable logic device ('PLD') 208, such as an FPGA. In this embodiment, each flash die 222 has pages organized as sixteen kB (kilobyte) pages 224, and registers 226 that can be used to write data to or read data from the flash die 222. In other embodiments, other types of solid-state memory are used as an alternative to or supplement to the flash memory shown within flash die 222.

[0113] In the various embodiments disclosed herein, storage cluster 161 can be contrasted with a general storage array. Storage nodes 150 are part of a collection that creates storage cluster 161. Each storage node 150 has a data slice and the computation required to provide the data. Multiple storage nodes 150 cooperate to store and retrieve data. Storage memories or storage devices typically used in storage arrays are less involved in data processing and manipulation. Storage memories or storage devices in a storage array receive commands to read, write, or erase data. Storage memories or storage devices in a storage array are unaware of the larger system they are embedded in, nor are they aware of what the data means. Storage memories or storage devices in a storage array can include various types of storage memories, such as RAM, solid-state drives, hard disk drives, etc. The cells of the non-volatile solid-state storage device 152 described herein have multiple interfaces that are active simultaneously and serve multiple purposes. In some embodiments, some functions of storage nodes 150 are moved to storage cells 152, thereby transforming storage cells 152 into a combination of storage cells 152 and storage nodes 150. Placing computation (relative to storing data) in storage cells 152 brings this computation closer to the data itself. Various system embodiments have a hierarchical structure of storage node layers with different functions. In contrast, in a storage array, the controller owns and knows everything about all the data managed by the controller in the rack or storage device. In storage cluster 161, as described herein, multiple controllers in the cells of multiple non-volatile solid-state storage devices 152 and / or storage nodes 150 cooperate in various ways (e.g., for erasure coding, data sharding, metadata communication and redundancy, storage capacity expansion or contraction, data recovery, etc.).

[0114] Figure 2D This illustrates the storage server environment, which uses Figure 2A -C's embodiment of storage node 150 and storage device 152 unit. In this version, each non-volatile solid-state storage device 152 unit is in chassis 138 (see...) Figure 2A The PCIe (Peripheral Component Interconnect High Speed) board in the ) has, for example, a controller 212 (see Figure 2C The processor, FPGA, flash memory 206, and NVRAM 204 (which is DRAM 216 supporting supercapacitors, see...) Figure 2B and 2C The non-volatile solid-state storage device 152 unit can be implemented as a single board containing the storage device and can be the maximum tolerable fault domain within the chassis. In some embodiments, up to two non-volatile solid-state storage device 152 units can fail, but the device can still continue without losing any data.

[0115] In some embodiments, the physical storage device is divided into named regions based on application usage. NVRAM 204 is a contiguous block of reserved memory in DRAM 216 of the non-volatile solid-state storage device 152 and is supported by NAND flash. NVRAM 204 is logically divided into multiple memory regions, two of which are written as spools (e.g., spool_region). The space within the spool of NVRAM 204 is managed independently by each license 168. Each device provides a certain amount of storage space to each license 168. The license 168 further manages the lifetime and allocation within the space. Examples of spools include distributed transactions or concepts. When the main power supply to the non-volatile solid-state storage device 152 cell fails, an onboard supercapacitor provides a short-term power hold. During this hold interval, the contents of NVRAM 204 are flushed to flash memory 206. Upon the next power-on, the contents of NVRAM 204 are restored from flash memory 206.

[0116] Regarding the memory cell controller, the responsibilities of the logical "controller" are distributed across each blade containing license 168. This distribution of logical control is... Figure 2D The diagram shows host controller 242, middleware controller 244, and storage unit controller 246. Management of the control plane and storage plane is handled independently, but the components may be physically co-located on the same blade. Each license 168 effectively acts as an independent controller. Each license 168 provides its own data and metadata structure, its own background worker, and maintains its own lifecycle.

[0117] Figure 2E Is using Figure 2D Storage server environment Figure 2A A hardware block diagram of blade 252 of an embodiment of storage node 150 and storage cell 152 of -C illustrates a control plane 254, compute and storage planes 256 and 258, and licenses 168 that interact with the underlying physical resources. The control plane 254 is partitioned into several licenses 168, which can run on any blade 252 using compute resources in the compute plane 256. The storage plane 258 is partitioned into a set of devices, each providing access to flash 206 and NVRAM 204 resources. In one embodiment, the compute plane 256 can perform the operation of a memory array controller on one or more devices of the storage plane 258 (e.g., a memory array), as described herein.

[0118] exist Figure 2EIn the compute and storage planes 256, 258, license 168 interacts with the underlying physical resources (i.e., devices). From the perspective of license 168, its resources are striped across all physical devices. From the perspective of a device, it provides resources to all licenses 168, regardless of where the license happens to be operating. Each license 168 has been allocated or has been allocated one or more partitions 260 of storage memory in storage unit 152, such as partitions 260 in flash memory 206 and NVRAM 204. Each license 168 uses its allocated partitions 260 to write or read user data. Licenses can be associated with different amounts of physical storage in the system. For example, the number or size of partitions 260 in one or more storage units 152 of a license 168 may be higher than that of one or more other licenses 168.

[0119] Figure 2F A resilient software layer in blade 252 of a storage cluster according to some embodiments is depicted. In the resilient architecture, the resilient software is symmetrical, i.e., the compute module 270 of each blade runs... Figure 2F The process depicted consists of three identical layers. Storage manager 274 executes read and write requests from other blades 252 for data and metadata stored in local storage unit 152, NVRAM 204, and flash 206. Authorizer 168 fulfills client requests by issuing the necessary read and write requests to the blade 252 where the corresponding data or metadata resides in its storage unit 152. Endpoint 272 parses client connection requests received from the supervisory software of switching structure 146, relays the client connection requests to the responsible authorizer 168, and relays the response of authorizer 168 to the client. This symmetric three-layer architecture enables a high degree of parallelism in the storage system. In these embodiments, elasticity scales horizontally efficiently and reliably. Furthermore, elasticity implements a unique scale-out technique that evenly balances work across all resources regardless of client access patterns and maximizes parallelism by eliminating much of the need for inter-blade coordination that typically occurs in traditional distributed locking.

[0120] Still referencing Figure 2FThe licensees 168, running in the compute module 270 of blade 252, perform the internal operations required to fulfill client requests. A characteristic of this resilience is that the licenses 168 are stateless; that is, they cache active data and metadata in the DRAM of their own blade 252 for fast access, but each update is stored on a partition of their NVRAM 204 on three separate blades 252 until the update is written to flash 206. In some embodiments, all storage system writes to NVRAM 204 are triple-copy, written to partitions on the three separate blades 252. Through triple-mirroring of NVRAM 204 and persistent storage protected by parity and Reed-Solomon RAID checksums, the storage system can withstand the simultaneous failure of two blades 252 without loss of data, metadata, or access to either.

[0121] Because licenses 168 are stateless, they can migrate between blades 252. Each license 168 has a unique identifier. In some embodiments, partitions of NVRAM 204 and flash 206 are associated with the identifier of the license 168, but not with the blade 252 on which they run. Therefore, when a license 168 migrates, it continues to manage the same storage partitions from its new location. When a new blade 252 is installed in one embodiment of the storage cluster, the system automatically rebalances the load by partitioning the storage of the new blade 252 for use by the system's licenses 168, migrating selected licenses 168 to the new blade 252, starting endpoint 272 on the new blade 252, and including it in the client connection distribution algorithm of the switching fabric 146.

[0122] From their new locations, the migrated licenses 168 maintain the contents of their NVRAM 204 partitions on flash 206, handle read and write requests from other licenses 168, and implement client requests directed to them by endpoint 272. Similarly, if blade 252 fails or is removed, the system redistributes its licenses 168 among the remaining blades 252 in the system. The redistributed licenses 168 continue to perform their original functions in their new locations.

[0123] Figure 2GThe illustration depicts licenses 168 and storage resources in blades 252 of a storage cluster according to some embodiments. Each license 168 is specifically responsible for partitions of flash 206 and NVRAM 204 on each blade 252. Licenses 168 manage the content and integrity of their partitions independently of other licenses 168. Licenses 168 compress incoming data and temporarily hold it in their NVRAM 204 partitions, then merge, RAID-protect, and retain the data in storage segments within their flash 206 partitions. When licenses 168 write data to flash 206, storage manager 274 performs necessary flash transformations to optimize write performance and maximize media lifetime. In the background, licenses 168 perform "garbage collection" or reclaim the space occupied by data discarded by clients. It should be understood that because the partitions of licenses 168 are disjoint, distributed locking is not required for client and write or background functions.

[0124] The embodiments described herein can utilize various software, communication, and / or networking protocols. Furthermore, the hardware and / or software configuration can be adapted to accommodate various protocols. For example, the embodiments can utilize Active Directory, which is a protocol used in Windows. TM Database-based systems provide authentication, directories, policies, and other services within the environment. In these embodiments, LDAP (Lightweight Directory Access Protocol) is an example application protocol for querying and modifying items in directory service providers such as Active Directory. In some embodiments, a Network Lock Manager ('NLM') is used as a facility to work with the Network File System ('NFS') to provide System V-style advisory document and record locking over the network. The Server Message Block ('SMB') protocol, one version of which is also known as the Universal Internet File System ('CIFS'), can be integrated with the storage systems discussed herein. SMB operates as an application-layer network protocol, typically used to provide shared access to files, printers, and serial ports, as well as various communications between nodes on the network. SMB also provides an authenticated inter-process communication mechanism. Amazon TMS3 (Simple Storage Service) is a web service provided by Amazon Web Services. The system described herein can interface with Amazon S3 through web service interfaces (REST (Representational State Transfer), SOAP (Simple Object Access Protocol), and BitTorrent). The RESTful API (Application Programming Interface) breaks down transactions into a series of small modules. Each module handles a specific underlying part of the transaction. The control or permissions provided by these embodiments, particularly for object data, may include the use of Access Control Lists ('ACLs'). An ACL is a list of permissions attached to an object, specifying which users or system processes are allowed to access the object and what operations are permitted on a given object. The system can utilize Internet Protocol version 6 ('IPv6') and IPv4 as communication protocols, which provide identification and location systems for computers on a network and route traffic over the Internet. Packet routing between network systems may include Equal Cost Multipath Routing ('ECMP'), a routing strategy where the forwarding of next-hop packets to a single destination can occur on multiple "best paths" that are tied for first place in the routing metric calculation. Multipath routing can be used in conjunction with most routing protocols because it is a per-hop decision limited to a single router. The software can support multi-tenancy, an architecture where a single instance of a software application serves multiple clients. Each client can be called a tenant. In some embodiments, a tenant may be given the ability to customize certain parts of the application, but may not customize the application's code. Implementations can maintain audit logs. Audit logs are documents that record events in a computing system. In addition to recording which resources were accessed, audit log entries typically include destination and source addresses, timestamps, and user login information to comply with various regulations. Implementations can support various key management strategies, such as cryptographic key rotation. Additionally, the system can support dynamic root passwords or some variant of dynamically changeable passwords.

[0125] Figure 3A A diagram illustrates a storage system 306 coupled to a cloud service provider 302 for data communication according to some embodiments of the present disclosure. Although depicted in limited detail, Figure 3A The storage system 306 described in the reference above can be similar to the one described above. Figure 1A-1D and Figure 2A-2G The described storage system. In some embodiments, Figure 3AThe storage system 306 described herein can be embodied as a storage system including an unbalanced active / active controller, a storage system including a balanced active / active controller, a storage system including active / active controllers (where less than all resources of each controller are utilized, such that each controller has reserve resources available to support failover), a storage system including a fully active / active controller, a storage system including controllers with separate datasets, a storage system including a two-tier architecture with a front-end controller and a back-end integrated storage controller, a storage system including a scale-out cluster with dual controller arrays, and combinations of these embodiments.

[0126] exist Figure 3A In the depicted example, storage system 306 is coupled to cloud service provider 302 via data communication link 304. Such data communication link 304 can be entirely wired, entirely wireless, or an aggregation of wired and wireless data communication paths. In this example, digital information can be exchanged between storage system 306 and cloud service provider 302 via data communication link 304 using one or more data communication protocols. For example, digital information can be exchanged between storage system 306 and cloud service provider 302 via data communication link 304 using Handheld Device Transfer Protocol ('HDTP'), Hypertext Transfer Protocol ('HTTP'), Internet Protocol ('IP'), Real-Time Transport Protocol ('RTP'), Transmission Control Protocol ('TCP'), User Datagram Protocol ('UDP'), Wireless Application Protocol ('WAP'), or other protocols.

[0127] Figure 3A The cloud service provider 302 described herein can be embodied, for example, as a system and computing environment that provides a wide range of services to users of the cloud service provider 302 via data communication link 304 through shared computing resources. The cloud service provider 302 can provide on-demand access to a shared pool of configurable computing resources, such as computer networks, servers, storage, applications, and services.

[0128] exist Figure 3A In the illustrated example, cloud service provider 302 can be configured to provide various services to storage system 306 and its users through implementations of various service models. For example, cloud service provider 302 can be configured to provide services through implementations of an Infrastructure as a Service ('IaaS') service model, through implementations of a Platform as a Service ('PaaS') service model, through implementations of a Software as a Service ('SaaS') service model, through implementations of an Authentication as a Service ('AaaS') service model, through implementations of a Storage as a Service model (where cloud service provider 302 provides access to its storage infrastructure for use by storage system 306 and its users), and so on.

[0129] exist Figure 3A In the depicted examples, cloud service provider 302 can be embodied, for example, as a private cloud, a public cloud, or a combination of private and public clouds. In embodiments where cloud service provider 302 is embodied as a private cloud, cloud service provider 302 may be dedicated to providing services to a single organization rather than to multiple organizations. In embodiments where cloud service provider 302 is embodied as a public cloud, cloud service provider 302 may provide services to multiple organizations. In alternative embodiments, cloud service provider 302 can be embodied as a hybrid of private and public cloud services with a hybrid cloud deployment.

[0130] although Figure 3A While not explicitly described, readers will understand that significant additional hardware and software components may be required to facilitate the delivery of cloud services to storage system 306 and its users. For example, storage system 306 may be coupled to (or even contain) a cloud storage gateway. Such a cloud storage gateway may be embodied as a hardware-based or software-based device located within storage system 306. This cloud storage gateway can serve as a bridge between local applications running on storage system 306 and the remote cloud-based storage utilized by storage system 306. By using a cloud storage gateway, organizations can move primary iSCSI or NAS to cloud service provider 302, enabling them to save space on their internal storage systems. Such a cloud storage gateway can be configured to emulate disk arrays, block-based devices, file servers, or other storage systems that can translate SCSI commands, file server commands, or other appropriate commands into REST space protocols that facilitate communication with cloud service provider 302.

[0131] To enable storage system 306 and its users to utilize services provided by cloud service provider 302, a cloud migration process can be performed, during which data, applications, or other elements from the organization's on-premises systems (or even from another cloud environment) are moved to cloud service provider 302. To successfully migrate data, applications, or other elements to the environment of cloud service provider 302, middleware such as cloud migration tools can be used to bridge the gap between the cloud service provider 302's environment and the organization's environment. To further enable storage system 306 and its users to utilize services provided by cloud service provider 302, cloud orchestrators can be used to schedule and coordinate automated tasks to create integrated processes or workflows. Such cloud orchestrators can perform tasks such as configuring various components (whether cloud or on-premises) and managing the interconnections between these components.

[0132] exist Figure 3AIn the illustrated example, as briefly described above, cloud service provider 302 can be configured to provide services to storage system 306 and its users using a SaaS service model. For example, cloud service provider 302 can be configured to provide access to data analytics applications to storage system 306 and its users. Such data analytics applications can be configured to receive large amounts of telemetry data returned by storage system 306. This telemetry data can describe various operational characteristics of storage system 306 and can be analyzed for a variety of purposes, including, for example, determining the health status of storage system 306, identifying workloads performed on storage system 306, predicting when storage system 306 will exhaust various resources, recommending configuration changes, hardware or software upgrades, workflow migrations, or other actions that can improve the operation of storage system 306.

[0133] The cloud service provider 302 can also be configured to provide access to the virtualized computing environment to the storage system 306 and its users. Examples of such virtualized environments may include virtual machines created to emulate physical computers, virtualized desktop environments that separate logical desktops from physical machines, virtualized file systems that allow unified access to different types of specific file systems, and so on.

[0134] although Figure 3A The illustrated example shows a storage system 306 coupled to communicate with a cloud service provider 302. However, in other embodiments, the storage system 306 may be part of a hybrid cloud deployment where private cloud components (e.g., private cloud services, on-premises infrastructure, etc.) and public cloud components (e.g., public cloud services, infrastructure, etc., which may be provided by one or more cloud service providers) are combined to form a single solution and orchestrated across various platforms. Such hybrid cloud deployments may utilize hybrid cloud management software, such as Microsoft... TM Azure TM Arc centralizes the management of hybrid cloud deployments on any infrastructure and supports deploying services anywhere. In such an example, hybrid cloud management software can be configured to create, update, and delete resources (both physical and virtual) that form a hybrid cloud deployment, allocate compute and storage to specific workloads, monitor workload and resource performance, policy compliance, updates and patches, security status, or perform various other tasks.

[0135] Readers will understand that various products can be enabled by pairing the storage system described herein with one or more cloud service providers. For example, Disaster Recovery as a Service ('DRaaS') can be provided, where cloud resources are leveraged to protect applications and data from disaster-induced disruptions, included in embodiments where the storage system can act as primary data storage. In such embodiments, a total system backup can be performed to ensure business continuity in the event of system failure. In such embodiments, cloud data backup technologies (alone or as part of a larger DRaaS solution) can also be integrated into a holistic solution that includes the storage system and cloud service provider described herein.

[0136] The storage systems and cloud service providers described in this article can be used to provide a wide range of security features. For example, storage systems can encrypt data at rest (and send and receive data in encrypted form), and can use Key Management as a Service ('KMaaS') to manage encryption keys, keys for locking and unlocking storage devices, and so on. Similarly, cloud data security gateways or similar mechanisms can be used to ensure that data stored within a storage system is not mistakenly stored in the cloud as part of cloud data backup operations. Furthermore, micro-segmentation or identity-based segmentation can be used within data centers or cloud service providers that include storage systems to create security zones in data center and cloud deployments, thereby isolating workloads from each other.

[0137] To further explain, Figure 3B Figures illustrate a storage system 306 according to some embodiments of the present disclosure. Although depicted in limited detail, Figure 3B The storage system 306 described in the reference above can be similar to the one described above. Figure 1A-1D and Figure 2A-2G The storage system described is because a storage system can contain the many components described above.

[0138] Figure 3BThe storage system 306 depicted herein may comprise a large number of storage resources 308, which may take many forms. For example, storage resources 308 may comprise nano-RAM or another form of non-volatile random access memory utilizing carbon nanotubes deposited on a substrate, 3D cross-point non-volatile memory, flash memory comprising single-level cell ('SLC') NAND flash, multi-level cell ('MLC') NAND flash, three-level cell ('TLC') NAND flash, four-level cell ('QLC') NAND flash, or others. Similarly, storage resources 308 may comprise non-volatile magnetoresistive random access memory ('MRAM'), comprising spin-transfer torque ('STT') MRAM. Example storage resources 308 may alternatively comprise non-volatile phase-change memory ('PCM'), quantum memory allowing the storage and retrieval of photonic quantum information, resistive random access memory ('ReRAM'), storage-class memory ('SCM'), or other forms of storage resources, comprising any combination of the resources described herein. Readers will learn that the storage system described above can utilize other forms of computer memory and storage devices, including DRAM, SRAM, EEPROM, general-purpose memory, and many other types of memory. Figure 3A The storage resource 308 described herein can take various physical forms, including but not limited to dual in-line memory modules ('DIMM'), non-volatile dual in-line memory modules ('NVDIMM'), M.2, U.2, etc.

[0139] Figure 3B The storage resource 308 depicted can contain various forms of SCM. An SCM can effectively treat fast, non-volatile memory (e.g., NAND flash) as an extension of DRAM, allowing the entire dataset to be viewed as an in-memory dataset residing entirely in DRAM. An SCM can contain non-volatile media, such as NAND flash. Such NAND flash can be accessed using NVMe, which uses the PCIe bus as its transport mode, offering relatively low access latency compared to older protocols. In practice, network protocols used for SSDs in an all-flash array can include NVMe using Ethernet (ROCE, NVMe TCP), Fibre Channel (NVMe FC), Infinite Bandwidth (iWARP), and others that treat fast, non-volatile memory as an extension of DRAM. Given that DRAM is typically byte-addressable, while fast, non-volatile memory such as NAND flash is block-addressable, a controller software / hardware stack may be required to translate block data into bytes stored in the medium. Examples of media and software that can be used as an SCM include 3D XPoint, Intel Memory Driver Technology, Samsung's Z-SSD, etc.

[0140] Figure 3BThe storage resource 308 depicted may also include raceway memory (also known as domain wall memory). This raceway memory can be embodied as a non-volatile solid-state memory that relies on the inherent strength and direction of the magnetic field generated by electrons spinning in the solid-state device, as well as their charge. By moving magnetic domains along nanoscale permalloy wires using a spin-coherent current, the domains may pass through a magnetic read / write head located near the wire as the current flows through the wire, thereby altering the domains to record bit patterns. To create a raceway memory device, many such wires and read / write elements can be packaged together.

[0141] Figure 3B The example storage system 306 depicted can implement various storage architectures. For example, storage systems according to some embodiments of this disclosure can utilize block storage, where data is stored in blocks, and each block essentially acts as a separate hard disk. Storage systems according to some embodiments of this disclosure can utilize object storage, where data is managed as objects. Each object can contain the data itself, a variable amount of metadata, and a globally unique identifier, where object storage can be implemented at multiple levels (e.g., device level, system level, interface level). Storage systems according to some embodiments of this disclosure utilize file storage, where data is stored in a hierarchy. Such data can be stored in files and folders and presented to its storage and retrieval systems in the same format.

[0142] Figure 3B The example storage system 306 depicted can be considered a storage system in which additional storage resources can be added using a scale-up model, a scale-out model, or a combination of both. In the scale-up model, additional storage is added by adding additional storage devices. However, in the scale-out model, additional storage nodes are added to a cluster of storage nodes, where these nodes may contain additional processing resources, additional network resources, and so on.

[0143] Figure 3B The example storage system 306 described herein can utilize the storage resources described above in a variety of different ways. For example, a portion of the storage resources can be used as a write cache, storage resources within the storage system can be used as a read cache, or tiering within the storage system can be implemented by placing data within the storage system according to one or more tiering strategies.

[0144] Figure 3BThe storage system 306 depicted also includes communication resources 310, which can be used to facilitate data communication between components within the storage system 306 and data communication between the storage system 306 and computing devices outside the storage system 306, including embodiments where these resources are spatially separated by a relatively large area. Communication resources 310 can be configured to facilitate data communication between components within the storage system and with computing devices outside the storage system using various different protocols and data communication structures. For example, communication resources 310 may include Fibre Channel ('FC') technology, such as the FC structure and FC protocol for transmitting SCSI commands over an FC network, FC over Ethernet ('FCoE') technology (FC frames can be encapsulated and transmitted over an Ethernet network using this technology), Infinite Bandwidth ('IB') technology (where a switching topology is used to facilitate transmission between channel adapters), NVM High Speed ​​('NVMe') technology and NVMe over Fabric ('NVMeoF') technology (through which non-volatile storage media attached via a PCI High Speed ​​('PCIe' bus) can be accessed), and so on. In fact, the storage system described above can directly or indirectly utilize neutrino communication technology and devices, through which information (including binary information) is transmitted using neutrino beams.

[0145] Communication resources 310 may also include mechanisms for accessing storage resources 308 within storage system 306 using Serial Attached SCSI ('SAS'), a Serial ATA ('SATA') bus interface for connecting storage resources 308 in storage system 306 to a host bus adapter within storage system 306, Internet Small Computer System Interface ('iSCSI') technology for providing block-level access to storage resources 308 within storage system 306, and other communication resources for facilitating data communication between components within storage system 306 and data communication between storage system 306 and computing devices outside storage system 306.

[0146] Figure 3B The storage system 306 depicted also includes processing resources 312 that can be used to execute computer program instructions and perform other computational tasks within the storage system 306. Processing resources 312 may include one or more ASICs customized for a particular purpose, and one or more CPUs. Processing resources 312 may also include one or more DSPs, one or more FPGAs, one or more system-on-a-chip ('SoC'), or other forms of processing resources 312. Storage system 306 can utilize storage resources 312 to perform various tasks, including, but not limited to, supporting the execution of software resources 314, which will be described in more detail below.

[0147] Figure 3BThe storage system 306 depicted also includes software resources 314, which can perform a wide range of tasks when executed by processing resources 312 within the storage system 306. Software resources 314 may include, for example, one or more computer program instruction modules, which, when executed by processing resources 312 within the storage system 306, can be used to implement various data protection techniques. Such data protection techniques may be implemented, for example, by system software executing on the computer hardware within the storage system, by a cloud service provider, or otherwise. These data protection techniques may include data archiving, data backup, data replication, data snapshots, data and database cloning, and other data protection techniques.

[0148] Software resource 314 may also contain software that can be used to implement software-defined storage ('SDS'). In such examples, software resource 314 may contain one or more computer program instruction modules that, when executed, can be used for policy-based provisioning and data storage management, independent of the underlying hardware. Such software resource 314 can be used to implement storage virtualization to decouple the storage hardware from the software that manages the storage hardware.

[0149] Software resource 314 may also include software that can be used to facilitate and optimize I / O operations involving storage system 306. For example, software resource 314 may include software modules that perform various data reduction techniques, such as data compression and data deduplication. Software resource 314 may include software modules that intelligently merge I / O operations to facilitate better use of the underlying storage resource 308, software modules that perform data migration operations to migrate data from within the storage system, and software modules that perform other functions. Such software resource 314 may be embodied as one or more software containers or in many other ways.

[0150] To further explain, Figure 3C Examples of cloud-based storage systems 318 according to some embodiments of this disclosure are illustrated. Figure 3C In the illustrated example, the cloud-based storage system 318 is formed entirely within a cloud computing environment 316, such as Amazon Web Services ('AWS'). TM Microsoft Azure TM Google Cloud Management Platform TM IBM Cloud TM Oracle Cloud TM And so on. Cloud-based storage system 318 can be used to provide services similar to those provided by the storage systems described above.

[0151] Figure 3CThe cloud-based storage system 318 depicted includes two cloud computing instances 320 and 322, respectively, for supporting the execution of storage controller applications 324 and 326. The cloud computing instances 320 and 322 may, for example, be instances of cloud computing resources (e.g., virtual machines) provided by a cloud computing environment 316 to support the execution of software applications such as storage controller applications 324 and 326. For example, each of the cloud computing instances 320 and 322 may run on an Azure VM, where each Azure VM may contain a high-speed temporary storage device that can be used as a cache (e.g., a read cache). In one embodiment, the cloud computing instances 320 and 322 may be embodied as Amazon Elastic Compute Cloud ('EC2') instances. In such examples, an Amazon Machine Image ('AMI') containing storage controller applications 324 and 326 can be launched to create and configure virtual machines capable of executing storage controller applications 324 and 326.

[0152] exist Figure 3C In the illustrated method, the storage controller applications 324 and 326 can be embodied as computer program instruction modules that, when executed, perform various storage tasks. For example, the storage controller applications 324 and 326 can be embodied as computer program instruction modules that, when executed, perform tasks similar to those described above. Figure 1A The controllers 110A and 110B perform the same tasks, such as writing data to and from the cloud-based storage system 318, erasing data from and from the cloud-based storage system 318, retrieving data from and from the cloud-based storage system 318, monitoring and reporting storage device utilization and performance, performing redundancy operations (e.g., RAID or RAID-like data redundancy operations), compressing data, encrypting data, deleting duplicate data, etc. The reader will understand that because there are two cloud computing instances 320 and 322, each containing storage controller applications 324 and 326, in some embodiments, one cloud computing instance 320 can be used as a primary controller, as described above, while the other cloud computing instance 322 can be used as a secondary controller, as described above. The reader will understand that... Figure 3C The storage controller applications 324 and 326 described herein can contain the same source code that executes within different cloud computing instances 320 and 322 (e.g., different EC2 instances).

[0153] Readers will understand that other embodiments excluding primary and secondary controllers are within the scope of this disclosure. For example, each cloud computing instance 320, 322 may serve as a primary controller for a portion of the address space supported by the cloud-based storage system 318, each cloud computing instance 320, 322 may serve as a primary controller for services involving I / O operations of the cloud-based storage system 318 that are otherwise partitioned, and so on. In fact, in other embodiments where cost savings may take precedence over performance requirements, there may be only a single cloud computing instance containing the storage controller application.

[0154] Figure 3C The cloud-based storage system 318 described herein includes cloud computing instances 340a, 340b, and 340n with local storage devices 330, 334, and 338. The cloud computing instances 340a, 340b, and 340n can, for example, be instances of cloud computing resources provided by a cloud computing environment 316 to support the execution of software applications. Figure 3C Cloud computing instances 340a, 340b, and 340n can be different from cloud computing instances 320 and 322 described above, because Figure 3C Cloud computing instances 340a, 340b, and 340n have local storage devices 330, 334, and 338 resources, while cloud computing instances 320 and 322, which support the execution of storage controller applications 324 and 326, do not require local storage resources. Cloud computing instances 340a, 340b, and 340n with local storage devices 330, 334, and 338 can, for example, be embodied as EC2 M5 instances containing one or more SSDs, EC2 R5 instances containing one or more SSDs, EC2 I3 instances containing one or more SSDs, and so on. In some embodiments, local storage devices 330, 334, and 338 must be embodied as solid-state storage devices (e.g., SSDs) rather than storage devices utilizing hard disk drives.

[0155] exist Figure 3CIn the illustrated example, each of the cloud computing instances 340a, 340b, and 340n, having local storage devices 330, 334, and 338, may contain software daemons 328, 332, and 336 that, when executed by the cloud computing instances 340a, 340b, and 340n, can present themselves to the storage controller applications 324 and 326 as if the cloud computing instances 340a, 340b, and 340n were physical storage devices (e.g., one or more SSDs). In such examples, the software daemons 328, 332, and 336 may contain computer program instructions similar to those normally contained on storage devices, enabling the storage controller applications 324 and 326 to send and receive the same commands sent by the storage controller to the storage devices. In this way, the storage controller applications 324 and 326 may contain code that is the same (or substantially the same) as the code to be executed by the controller in the storage system described above. In these and similar embodiments, communication between storage controller applications 324, 326 and cloud computing instances 340a, 340b, 340n with local storage devices 330, 334, 338 can utilize iSCSI, NVMe over TCP, messaging, custom protocols, or some other mechanism.

[0156] exist Figure 3C In the depicted example, each of the cloud computing instances 340a, 340b, 340n with local storage devices 330, 334, 338 may also be coupled to block storage devices 342, 344, 346, provided by the cloud computing environment 316, such as as an Amazon Elastic Block Storage ('EBS') volume. In such examples, the block storage devices 342, 344, 346 provided by the cloud computing environment 316 can be utilized in a manner similar to that described above for NVRAM devices, because software daemons 328, 332, 336 (or some other module) executing within a particular cloud computing instance 340a, 340b, 340n can initiate writes of data to their attached EBS volumes and to their local storage devices 330, 334, 338 resources upon receiving a write request. In some alternative embodiments, data may only be written to the local storage devices 330, 334, 338 resources within a particular cloud computing instance 340a, 340b, 340n. In an alternative embodiment, instead of using block storage devices 342, 344, 346 provided by the cloud computing environment 316 as NVRAM, the actual RAM on each of the cloud computing instances 340a, 340b, 340n with local storage devices 330, 334, 338 can be used as NVRAM, thereby reducing the network utilization costs associated with using EBS volumes as NVRAM. In yet another embodiment, high-performance block storage resources such as one or more Azure Ultra disks can be used as NVRAM.

[0157] When a specific cloud computing instance 340a, 340b, or 340n with local storage devices 330, 334, or 338 receives a request to write data, software daemons 328, 332, and 336 can be configured not only to write data to their own local storage devices 330, 334, or 338 and any suitable block storage devices 342, 344, or 346, but also to write data to a cloud-based object storage device 348 attached to the specific cloud computing instance 340a, 340b, or 340n. The cloud-based object storage device 348 attached to the specific cloud computing instance 340a, 340b, or 340n can, for example, be embodied as Amazon Simple Storage Service ('S3'). In other embodiments, cloud instances 320 and 322, each including storage controller applications 324 and 326, can initiate storage of data in local storage devices 330, 334, and 338 of cloud instances 340a, 340b, and 340n, and in cloud-based object storage device 348. In other embodiments, instead of using both cloud instances 340a, 340b, and 340n with local storage devices 330, 334, and 338 (also referred to herein as 'virtual drives') and cloud-based object storage device 348 to store data, a persistent storage tier can be implemented in other ways. For example, one or more Azure Ultra disks can be used to persistently store data (e.g., after the data has been written to the NVRAM tier). In embodiments where one or more Azure Ultra disks are used for persistent data storage, the use of cloud-based object storage device 348 can be eliminated, such that data is only persistently stored in the Azure Ultra disks without being written to the object storage tier.

[0158] Although the local storage devices 330, 334, and 338 and block storage devices 342, 344, and 346 utilized by cloud computing instances 340a, 340b, and 340n can support block-level access, the cloud-based object storage device 348 attached to a specific cloud computing instance 340a, 340b, or 340n only supports object-based access. Therefore, software daemons 328, 332, and 336 can be configured to retrieve data blocks, encapsulate those blocks into objects, and write the objects to the cloud-based object storage device 348 attached to the specific cloud computing instance 340a, 340b, or 340n.

[0159] In some embodiments, all data stored by the cloud-based storage system 318 may be stored in either: 1) a cloud-based object storage device 348, and 2) at least one of the local storage devices 330, 334, 338 or block storage devices 342, 344, 346 resources utilized by the cloud computing instances 340a, 340b, 340n. In such embodiments, the local storage devices 330, 334, 338 resources and the block storage devices 342, 344, 346 resources utilized by the cloud computing instances 340a, 340b, 340n can be effectively used as a cache that typically contains all data also stored in S3, so that all data reads can be serviced by the cloud computing instances 340a, 340b, 340n without requiring the cloud computing instances 340a, 340b, 340n to access the cloud-based object storage device 348. However, the reader will understand that in other embodiments, all data stored by the cloud-based storage system 318 may be stored in the cloud-based object storage device 348, but less than all of the data stored by the cloud-based storage system 318 may be stored in at least one of the local storage devices 330, 334, 338 or block storage devices 342, 344, 346 utilized by the cloud computing instances 340a, 340b, 340n. In such examples, various strategies may be employed to determine which subset of the data stored by the cloud-based storage system 318 should reside in either: 1) the cloud-based object storage device 348, and 2) at least one of the local storage devices 330, 334, 338 or block storage devices 342, 344, 346 utilized by the cloud computing instances 340a, 340b, 340n.

[0160] One or more computer program instruction modules executing within the cloud-based storage system 318 (e.g., a monitoring module executing on its own EC2 instance) can be designed to handle the failure of one or more of the cloud computing instances 340a, 340b, 340n with local storage devices 330, 334, 338. In such an example, the monitoring module can handle the failure of one or more of the cloud computing instances 340a, 340b, 340n with local storage devices by creating one or more new cloud computing instances with local storage devices, retrieving data already stored on the failed cloud computing instances 340a, 340b, 340n from the cloud-based object storage device 348, and storing the data retrieved from the cloud-based object storage device 348 in the local storage device of the newly created cloud computing instance. The reader will understand that many variations of this process can be implemented.

[0161] Readers will understand that various performance aspects of the cloud-based storage system 318 can be monitored (e.g., through a monitoring module running in an EC2 instance) to scale the cloud-based storage system 318 up or down as needed. For example, if the cloud computing instances 320 and 322 used to support the execution of storage controller applications 324 and 326 are too small and cannot adequately serve the I / O requests issued by users of the cloud-based storage system 318, the monitoring module can create new, more powerful cloud computing instances (e.g., cloud computing instances with more processing power, more memory, etc.), including storage controller applications that allow the new, more powerful cloud computing instances to begin operating as primary controllers. Similarly, if the monitoring module determines that the cloud computing instances 320 and 322 used to support the execution of storage controller applications 324 and 326 are too large and that cost savings can be achieved by switching to smaller, less powerful cloud computing instances, the monitoring module can create new, less powerful (and less costly) cloud computing instances, including storage controller applications that allow the new, less powerful cloud computing instances to begin operating as primary master controllers.

[0162] The storage system described above can perform intelligent data backup technology, through which data stored in the storage system can be copied and stored in different locations to avoid data loss in the event of device failure or other forms of mutation. For example, the storage system described above can be configured to inspect each backup to avoid restoring the storage system to an undesirable state. Consider an example where a storage system is contaminated with malware. In such an example, the storage system may include software resource 314 that can scan each backup to identify backups captured before and after malware contamination. In such an example, the storage system can recover itself from backups that do not contain malware, or at least not from backups containing malware. In such an example, the storage system may include software resource 314 that can scan each backup to identify the presence of malware (or viruses, or some other unwanted malware), for example, by identifying write operations served by the storage system and originating from a network subnet suspected of delivering malware, by identifying write operations served by the storage system and originating from a user suspected of delivering malware, by identifying write operations served by the storage system and checking the content of the write operations against the malware fingerprint, and in many other ways.

[0163] Readers will further understand that backups (typically in the form of one or more snapshots) can also be used to perform rapid recovery of storage systems. Consider an example where the storage system is contaminated by ransomware that locks users out of the storage system. In such an example, software resource 314 within the storage system can be configured to detect the presence of ransomware and can be further configured to restore the storage system to a point in time prior to when the ransomware contaminated the storage system using a retained backup. In such an example, the presence of ransomware can be explicitly detected by using software tools used by the system, by using a key inserted into the storage system (e.g., a USB drive), or in a similar manner. Similarly, the presence of ransomware can be inferred in response to system activity satisfying a predetermined fingerprint, for example, no reading or writing to the system within a predetermined time period.

[0164] Readers will understand that the various components described above can be grouped into one or more optimized compute packages as converged infrastructure. Such converged infrastructure can include pools of computing, storage, and network connectivity resources that can be shared by multiple applications and managed collectively using policy-driven processes. This converged infrastructure can be implemented using converged infrastructure reference architectures, stand-alone devices, software-driven hyperconverged approaches (e.g., hyperconverged infrastructure), or other methods.

[0165] Readers will understand that the storage system described in this disclosure can be used to support various types of software applications. In fact, the storage system can be "application-aware" because it can acquire, maintain, or otherwise access information describing connected applications (e.g., applications utilizing the storage system) to optimize its operation based on intelligence about the applications and their utilization patterns. For example, the storage system can optimize data layout, cache behavior, QoS tiers, or perform other optimizations aimed at improving the storage performance experienced by the application.

[0166] As an example of an application that the storage system described herein may support, storage system 306 can be used to support these applications by providing storage resources to artificial intelligence ('AI') applications, database applications, XOps projects (such as DevOps projects, DataOps projects, MLOps projects, ModelOps projects, PlatformOps projects), electronic design automation tools, event-driven software applications, high-performance computing applications, simulation applications, high-speed data capture and analysis applications, machine learning applications, media production applications, media service applications, picture archiving and communication system ('PACS') applications, software development applications, virtual reality applications, augmented reality applications, and many other types of applications.

[0167] Given that storage systems encompass computing resources, storage resources, and various other resources, they may be well-suited to support resource-intensive applications, such as AI applications. AI applications can be deployed across a wide range of sectors, including: predictive maintenance in manufacturing and related fields; healthcare applications such as patient data and risk analytics; retail and marketing deployments (e.g., search advertising, social media advertising); supply chain solutions; fintech solutions such as business analytics and reporting tools; operational deployments such as real-time analytics tools, application performance management tools, and IT infrastructure management tools; and many others.

[0168] Such AI applications enable devices to perceive their environment and take actions, maximizing their chances of success at a given goal. Examples of such AI applications include IBM Watson. TM Microsoft Oxford TM Google DeepMind TM Baidu Minwa TM wait.

[0169] The storage system described above may also be well-suited for supporting other types of resource-intensive applications, such as machine learning applications. Machine learning applications can perform various types of data analysis to automate the construction of analytical models. Using algorithms that iteratively learn from data, machine learning applications enable computers to learn without explicit programming. A specific area of ​​machine learning is called reinforcement learning, which involves taking appropriate actions in specific situations to maximize rewards.

[0170] In addition to the resources already described, the storage system described above may also include a graphics processing unit ('GPU'), sometimes also called a visual processing unit ('VPU'). Such a GPU can be embodied as dedicated electronic circuitry that rapidly manipulates and modifies memory to accelerate the creation of images in a frame buffer for output to a display device. Such a GPU can be included in any computing device that is part of the storage system described above, comprising one of many separately expandable components of the storage system. Other examples of separately expandable components of such a storage system may include storage components, memory components, computing components (e.g., CPU, FPGA, ASIC), network connectivity components, software components, etc. Besides the GPU, the storage system described above may also include a neural network processor ('NNP') for various aspects of neural network processing. Such an NNP can be used in place of (or as a complement to) the GPU, or it can be expanded independently.

[0171] As described above, the storage system described in this paper can be configured to support artificial intelligence applications, machine learning applications, big data analytics applications, and many other types of applications. The rapid growth of these applications is driven by three technologies: deep learning (DL), GPU processors, and big data. Deep learning utilizes computational models of massively parallel neural networks inspired by the human brain. Deep learning models learn their own software by studying a large number of examples, rather than having software handcrafted by experts. Such GPUs can contain thousands of cores perfectly suited for running algorithms that loosely represent the parallel nature of the human brain.

[0172] Advances in deep neural networks, including the development of multi-layered neural networks, have spurred a new wave of algorithms and tools for data scientists to mine data using artificial intelligence (AI). With improved algorithms, larger datasets, and a variety of frameworks (including open-source software libraries for machine learning across a range of tasks), data scientists are tackling new use cases such as autonomous vehicles, natural language processing and understanding, computer vision, machine reasoning, and strong AI. AI technologies are already being implemented in numerous products, including, for example, Amazon Echo's speech recognition technology, which allows users to converse with their machines; and Google Translate. TM It allows machine-based language translation; Spotify's Discover Weekly, which provides recommendations for new songs and artists that users might like based on user usage and traffic analysis; Quill's text generation product, which takes structured data and transforms it into narrative stories; Chatbot, which provides real-time, context-specific question answers in a conversational format; and many other products.

[0173] Data is central to modern AI and deep learning algorithms. One crucial issue that must be addressed before training can begin is collecting labeled data, which is essential for training accurate AI models. A comprehensive AI deployment may be required to continuously collect, clean, transform, label, and store large amounts of data. Adding additional high-quality data points directly translates into more accurate models and better insights. Data samples can undergo a series of processing steps, including but not limited to: 1) ingesting data from external sources into the training system and storing the data in its raw form; 2) cleaning and transforming the data into a training-friendly format, including linking data samples to appropriate labels; 3) exploring parameters and models, rapidly testing and iterating with smaller datasets to converge to the most promising model before pushing it into the production cluster; 4) performing a training phase to select random batches of input data, including both new and old samples, and feeding this data into production GPU servers for computation to update model parameters; and 5) evaluation, including a retained portion using data not used in training, to assess the model accuracy on the retained data. This lifecycle can be applied to any type of parallelized machine learning, not just neural networks or deep learning. For example, standard machine learning frameworks may rely on CPUs instead of GPUs, but the data ingestion and training workflows can be the same. Readers will understand that a single shared storage data center creates a coordination point throughout the entire lifecycle, without requiring additional copies of data to be created between the ingestion, preprocessing, and training phases. Ingested data is rarely used for a single purpose, and shared storage provides the flexibility to train multiple different models or apply traditional analytics to the data.

[0174] Readers will learn that each stage in an AI data pipeline may place different requirements on the data center (e.g., a storage system or collection of storage systems). Scale-out storage systems must provide uncompromised performance across a variety of access types and patterns—from small files and metadata-heavy files to large files, from random access to sequential access, and from low to high concurrency. The storage systems described above can serve as ideal AI data centers because they can serve unstructured workloads. In the first stage, ideally, data is ingested and stored in the same data center that will be used in subsequent stages to avoid excessive data duplication. The next two steps can be done on standard compute servers that optionally include GPUs, and then in the fourth and final stage, a full training production job is run on a powerful GPU-accelerated server. Typically, a production pipeline exists in addition to the experimental pipeline that operates on the same dataset. Furthermore, GPU-accelerated servers can be used independently for different models, or connected together for training on a larger model, or even for distributed training across multiple systems. If the shared storage layer is slow, data must be copied to local storage at each stage, wasting time transferring data to different servers. An ideal data center for an AI training pipeline offers performance similar to data stored locally on server nodes, while also being simple and capable of enabling all pipeline stages to operate simultaneously.

[0175] To enable the storage system described above to be used as part of a data center or AI deployment, in some embodiments, the storage system can be configured to provide DMA between storage devices contained within the storage system and one or more GPUs used in AI or big data analytics pipelines. One or more GPUs can be coupled to the storage system, for example, via structural NVMe ('NVMe-oF'), allowing bottlenecks such as the host CPU to be bypassed, and the storage system (or one of its components) to directly access GPU memory. In such an example, the storage system can utilize the GPU's API hooks to directly transfer data to the GPU. For example, the GPU could be an Nvidia GPU. TM The GPU and the storage system can support GPUDirect Storage ('GDS') software, or have similar proprietary software that enables the storage system to transfer data to the GPU via RDMA or a similar mechanism.

[0176] While the preceding paragraphs discussed deep learning applications, readers should understand that the storage system described in this article can also be part of a distributed deep learning ('DDL') platform to support the execution of DDL algorithms. The storage system described above can also be paired with other technologies (such as TensorFlow, an open-source software library for dataflow programming across a range of tasks, applicable to machine learning applications such as neural networks) to facilitate the development of such machine learning models, applications, etc.

[0177] The storage system described above can also be used in neuromorphic computing environments. Neuromorphic computing is a form of computation that simulates brain cells. To support neuromorphic computing, interconnected "neuronal" architectures replace traditional computational models with low-power signals transmitted directly between neurons, enabling more efficient computation. Neuromorphic computing can utilize very large-scale integrated (VLSI) systems containing electronic analog circuitry to simulate the neurobiological architectures present in the nervous system, as well as software systems that utilize analog, digital, and mixed-mode analog / digital VLSI and implement neural system models for perception, motor control, or multisensory integration.

[0178] Readers will learn that the storage system described above can be configured to support the storage or use (and other types of data) of blockchain and derivatives, such as those provided by IBM. TM This paper covers open-source blockchains and related tools as part of the Hyperledger Project, permissioned blockchains that allow a limited number of trusted parties to access the blockchain, and blockchain products that enable developers to build their own distributed ledger projects. The blockchain and storage systems described in this paper can be used to support both on-chain and off-chain storage of data.

[0179] Off-chain storage of data can be implemented in various ways and can occur even when the data itself is not stored on the blockchain. For example, in one embodiment, a hash function can be utilized, and the data itself can be fed into the hash function to generate a hash value. In such an example, the hash value of a large block of data, rather than the data itself, can be embedded in a transaction. The reader will understand that in other embodiments, alternatives to blockchain can be used to facilitate decentralized storage of information. For example, one possible alternative to blockchain is blockweave. While traditional blockchains store each transaction for verification, blockweave allows for secure decentralization without using the entire chain, thus enabling low-cost on-chain storage of data. This blockweave can leverage consensus mechanisms based on Proof-of-Access (PoA) and Proof-of-Work (PoW).

[0180] The storage systems described above can be used alone or in combination with other computing devices to support in-memory computing applications. In-memory computing involves storing information in RAM distributed across a computer cluster. The reader will understand that the storage systems described above, particularly those configurable with customizable amounts of processing, storage, and memory resources (e.g., those where blades contain configurable amounts of each type of resource), can be configured in a way that provides the infrastructure to support in-memory computing. Similarly, compared to in-memory computing environments that rely on RAM distributed across dedicated servers, the storage systems described above can include components that can actually provide an improved in-memory computing environment (e.g., NVDIMMs, 3D cross-point storage devices providing persistent, fast random access memory).

[0181] In some embodiments, the storage system described above can be configured to operate as a hybrid in-memory computing environment that includes a common interface to all storage media, such as RAM, flash memory, and 3D cross-point memory. In such embodiments, users may not be aware of the details of where their data is stored, but they can still use the same complete, unified API to address the data. In such embodiments, the storage system can (in the background) move data to the fastest available tier—including intelligently placing data based on various characteristics of the data or according to some other heuristic. In such examples, the storage system can even leverage existing products such as Apache Ignite and GridGain to move data between storage tiers, or the storage system can leverage custom software to move data between storage tiers. The storage system described herein can implement various optimizations to improve the performance of in-memory computing, for example, by placing computation as close to the data as possible.

[0182] As the reader will further understand, in some embodiments, the storage system described above can be paired with other resources to support the applications described above. For example, an infrastructure may include primary computing in the form of servers and workstations that specifically utilize general-purpose computing on graphics processing units ('GPGPUs') to accelerate deep learning applications interconnected with computing engines to train parameters for deep neural networks. Each system may have Ethernet external connectivity, InfiniBand external connectivity, some other form of external connectivity, or some combination thereof. In such an example, GPUs may be grouped for a single large training session or used independently to train multiple models. The infrastructure may also include storage systems, such as those described above, to provide, for example, horizontally scalable full-flash file or object storage through which data can be accessed via high-performance protocols such as NFS, S3, etc. The infrastructure may also include, for example, redundant top-of-rack Ethernet switches for storage and computing via port connections in MLAG port channels to achieve redundancy. The infrastructure may also include additional computing in the form of white-box servers, optionally using GPUs, for data ingestion, preprocessing, and model debugging. The reader will understand that additional infrastructure is also possible.

[0183] Readers will see that the storage system described above, whether used alone or in conjunction with other computing machines, can be configured to support other AI-related tools. For example, the storage system can use tools such as ONXX or other open neural network exchange formats to make it easier to transfer models written using different AI frameworks. Similarly, the storage system can be configured to support tools like Amazon's Gluon, allowing developers to prototype, build, and train deep learning models. In fact, the storage system described above can be part of a larger platform, such as IBM's... TM Cloud Private for Data includes integrated data science, data engineering, and application building services.

[0184] Readers will further understand that the storage systems described above can also be deployed as edge solutions. Such edge solutions optimize cloud computing systems by performing data processing at the network edge near the data source. Edge computing pushes applications, data, and computing power (i.e., services) from a centralized point to the logical edge of the network. By using edge solutions such as the storage systems described above, computing tasks can be performed using the computing resources provided by such storage systems, data can be stored using the storage resources of the storage systems, and cloud-based services can be accessed using the various resources of the storage systems, including network connectivity resources. By performing computing tasks on edge solutions, storing data on edge solutions, and primarily using edge solutions, the consumption of expensive cloud-based resources can be avoided, and in fact, performance improvements can be achieved relative to a greater reliance on cloud-based resources.

[0185] While many tasks may benefit from utilizing edge solutions, certain specific uses may be particularly well-suited for deployment in such an environment. For example, devices such as drones, self-driving cars, and robots may require extremely fast processing—in fact, so fast that sending data to a cloud environment and receiving it back for processing support might be too slow. As an additional example, some IoT devices, such as connected cameras, may not be well-suited to leveraging cloud-based resources because sending data to the cloud simply because of the sheer volume involved may be impractical (not only from a privacy, security, or financial perspective). Therefore, many tasks that truly involve data processing, storage, or communication may be better suited to platforms incorporating edge solutions that include storage systems such as those described above.

[0186] The storage system described above can be used independently or in conjunction with other computing resources as a network edge platform that combines computing resources, storage resources, network connectivity resources, cloud technologies, and network virtualization technologies. As part of the network, the edge may exhibit characteristics similar to other network infrastructure, from customer premises and backhaul aggregation facilities to access points (PoPs) and regional data centers. Readers will understand that network workloads such as Virtual Network Functions (VNFs) will reside on the network edge platform. Through a combination of containers and virtual machines, the network edge platform may rely on controllers and schedulers that are no longer geographically located alongside data processing resources. As microservices, these functions can be segmented into control planes, user and data planes, and even state machines, allowing for independent optimization and scaling techniques. Such user and data planes can be implemented using added accelerators (all residing in server platforms such as FPGAs and smart NICs) and through commercially available silicon and programmable ASICs that support SDN.

[0187] The storage system described above can also be optimized for big data analytics uses, including as part of a composable data analytics pipeline, where containerized analytics architectures, for example, make analytical capabilities more composable. Big data analytics can generally be described as the process of examining large and diverse datasets to discover hidden patterns, unknown correlations, market trends, customer preferences, and other useful information that can help organizations make more informed business decisions. As part of this process, semi-structured and unstructured data, such as internet clickstream data, web server logs, social media content, text in customer emails and survey responses, mobile phone call details, IoT sensor data, and other data, can be transformed into a structured form.

[0188] The storage system described above can also support (including implementations as system interfaces) applications that respond to human voice commands to perform tasks. For example, the storage system can support intelligent personal assistant applications, such as Amazon's Alexa. TM Apple Siri TM Google Voice TM Samsung Bixby TM Microsoft Cortana TM While the example described in the preceding sentence uses voice as input, the storage system described above can also support chatbots, talkbots, chatterbots, or other human dialogue entities or applications configured to engage in dialogue via auditory or text methods. Similarly, the storage system can actually execute applications that enable users, such as system administrators, to interact with the storage system via voice. Although such applications are typically capable of voice interaction, music playback, creating to-do lists, setting alarms, streaming podcasts, playing audiobooks, and providing weather, traffic, and other real-time information such as news, in embodiments according to this disclosure, such applications can be used as interfaces to various system management operations.

[0189] The storage systems described above can also implement AI platforms to realize the vision of self-driven storage. Such AI platforms can be configured to provide global predictive intelligence by collecting and analyzing vast amounts of storage system telemetry data points, enabling easy management, analysis, and support. In fact, such storage systems may be able to predict capacity and performance and generate intelligent recommendations for workload deployment, interaction, and optimization. These AI platforms can be configured to scan all incoming storage system telemetry data against a problem fingerprint database to predict and resolve events in real time before they impact the customer's environment, and capture hundreds of performance-related variables for predicting performance loads.

[0190] The storage system described above can support the sequential or simultaneous execution of artificial intelligence applications, machine learning applications, data analytics applications, data transformation, and other tasks that can collectively form the AI ​​ladder. By combining these elements to form a complete data science pipeline, such an AI ladder can be effectively constructed, where dependencies exist between the elements. For example, AI may require some form of machine learning to have occurred, machine learning may require some form of analytics to have occurred, analytics may require some type of data and information architecture to have occurred, and so on. Therefore, each element can be considered a step in the AI ​​ladder, collectively forming a complete and complex AI solution.

[0191] The storage system described above can be used alone or in conjunction with other computing environments to deliver an experience where AI is ubiquitous, permeating a wide range of aspects of business and life. For example, AI may play a significant role in providing deep learning solutions, deep reinforcement learning solutions and general artificial intelligence solutions, autonomous vehicles, cognitive computing solutions, commercial UAVs or drones, conversational user interfaces, enterprise classification, ontology management solutions, machine learning solutions, smart dust, smart robots, and smart workplaces.

[0192] The storage system described above can also be used alone or in conjunction with other computing environments to provide a wide range of transparent immersive experiences (including experiences using digital twins of various "things" such as people, places, processes, systems, etc.), where technology can introduce transparency between people, businesses, and things. Such transparent immersive experiences can be provided through augmented reality, connected homes, virtual reality, brain-computer interfaces, human augmentation technologies, nanotube electronics, volumetric displays, 4D printing, or other means.

[0193] The storage system described above can be used alone or in conjunction with other computing environments to support a variety of digital platforms. For example, such digital platforms can include 5G wireless systems and platforms, digital twin platforms, edge computing platforms, IoT platforms, quantum computing platforms, serverless PaaS, software-defined security, neuromorphic computing platforms, and so on.

[0194] The storage system described above can also be part of a multi-cloud environment, where multiple cloud computing and storage services are deployed within a single heterogeneous architecture. To facilitate operation in such a multi-cloud environment, DevOps tools can be deployed to enable cross-cloud orchestration. Similarly, continuous development and continuous integration tools can be deployed to standardize processes regarding continuous integration and delivery, new feature rollout, and provisioning of cloud workloads. By standardizing these processes, a multi-cloud strategy that leverages the best providers for each workload can be implemented.

[0195] The storage system described above can be used as part of a platform to enable the use of cryptographic anchors that can be used to verify the origin and content of a product, ensuring it matches the blockchain record associated with the product. Similarly, as part of a suite of tools to protect data stored on the storage system, the storage system described above can implement various cryptographic techniques and schemes, including lattice cryptography. Lattice cryptography can involve the construction of cryptographic primitives that include lattices, both in the construction itself and in security proofs. Unlike public-key schemes such as RSA, Diffie-Hellman, or elliptic curve cryptography systems, which are vulnerable to quantum computer attacks, some lattice-based constructions appear to be resistant to attacks from both classical and quantum computers.

[0196] A quantum computer is a device that performs quantum computation. Quantum computation is computation using quantum mechanical phenomena such as superposition and entanglement. Quantum computers differ from conventional transistor-based computers because conventional computers require data to be encoded as binary digits (bits), each of which is always in one of two definite states (0 or 1). In contrast to conventional computers, quantum computers use qubits, which can be in superposition states. A quantum computer maintains a sequence of qubits, where a single qubit can represent one, zero, or any quantum superposition of those two states. A pair of qubits can be in any quantum superposition of four states, and three qubits can be in any superposition of eight states. A quantum computer with n qubits can typically be in any superposition of up to 2^n distinct states simultaneously, while a conventional computer can only be in one state at any given time. The quantum Turing machine is a theoretical model of such a computer.

[0197] The storage systems described above can also be paired with FPGA acceleration servers as part of a larger AI or ML infrastructure. Such FPGA acceleration servers can be located near the storage systems described above (e.g., in the same data center) or even incorporated into a device that includes one or more storage systems, one or more FPGA acceleration servers, network connectivity infrastructure supporting communication between the one or more storage systems and the one or more FPGA acceleration servers, and other hardware and software components. Alternatively, FPGA acceleration servers can reside in a cloud computing environment that can be used to perform computationally relevant tasks for AI and ML jobs. Any of the embodiments described above can collectively serve as an FPGA-based AI or ML platform. Readers will understand that in some embodiments of an FPGA-based AI or ML platform, the FPGAs contained within the FPGA acceleration server can be reconfigured for different types of ML models (e.g., LSTM, CNN, GRU). The ability to reconfigure the FPGAs contained within the FPGA acceleration server can accelerate ML or AI applications based on the optimal numerical precision and memory model used. Readers will understand that by viewing the collection of FPGA acceleration servers as an FPGA pool, any CPU in the data center can use the FPGA pool as a shared hardware microservice, rather than limiting the servers to dedicated accelerators plugged into it.

[0198] The FPGA-accelerated servers and GPU-accelerated servers described above can implement a computing model in which machine learning models and parameters are fixed in high-bandwidth on-chip memory, and large amounts of data flow through this high-bandwidth on-chip memory, rather than storing small amounts of data in the CPU and running long streams of instructions as in more traditional computing models. For this computing model, FPGAs may even be more efficient than GPUs because FPGAs can be programmed using only the instructions required to run this computing model.

[0199] The storage system described above can be configured to provide parallel storage, for example, by using a parallel file system such as BeeGFS. This parallel file system can contain a distributed metadata architecture. For example, a parallel file system can contain metadata distributed across multiple metadata servers, as well as components containing services for clients and storage servers.

[0200] The system described above can support the execution of various software applications. These applications can be deployed in multiple ways, including container-based deployment models. Containerized applications can be managed using various tools. For example, they can be managed using Docker Swarm, Kubernetes, and others. Containerized applications can facilitate serverless, cloud-native computing deployment and management models for software applications. To support these models, containers can be used as part of event handling mechanisms (such as AWS Lambdas), allowing various events to trigger containerized applications to run as event handlers.

[0201] The system described above can be deployed in various ways, including to support fifth-generation ('5G') networks. 5G networks can support data communication much faster than previous generations of mobile communication networks, potentially leading to a fragmentation of data and computing resources. Modern large-scale data centers may become less prominent and could be replaced by local micro data centers closer to mobile network towers. The system described above can be contained within such local micro data centers and can be part of or paired with a multi-access edge computing ('MEC') system. Such MEC systems can realize cloud computing capabilities and IT service environments at the edge of cellular networks. By running applications and performing related processing tasks closer to cellular customers, network congestion can be reduced and application performance improved.

[0202] The storage system described above can also be configured to implement NVMe partitioned namespaces. By using NVMe partitioned namespaces, the logical address space of the namespace is divided into multiple zones. Each zone provides a range of logical block addresses that must be written sequentially and explicitly reset before being overwritten, thereby enabling the creation of namespaces that expose the natural boundaries of the device and offloading the management of internal mapping tables to the host. To implement NVMe partitioned namespaces ('ZNS'), ZNS SSDs or other forms of partitioned block devices that expose the logical address space of the namespace using zones can be used. By aligning zones with the internal physical properties of the device, several inefficiencies in data placement can be eliminated. In such embodiments, each zone can be mapped to, for example, a separate application, allowing functions such as wear leveling and garbage collection to be performed on a zone-by-zone or per-application basis rather than across the entire device. To support ZNS, the storage controller described herein can be configured to use, for example, Linux TM The kernel partition block device interface or other tools interact with the partition block device.

[0203] The storage system described above can also be configured to implement partitioned storage in other ways, such as by using shingled magnetic recording (SMR) storage devices. In examples using partitioned storage, a device management embodiment can be deployed where the storage device hides this complexity by managing it within the firmware, thus presenting an interface similar to any other storage device. Alternatively, partitioned storage can be implemented via a host-managed embodiment that relies on the operating system to know how to handle the drive and only sequentially writes to certain areas of the drive. Similarly, a host-aware embodiment can be used to implement partitioned storage, deploying a combination of drive management and host management implementations.

[0204] The storage system described herein can be used to form a data lake. A data lake can serve as the first place to organize data flows, where such data can be in its raw format. Metadata tagging can be implemented to facilitate searching for data elements within the data lake, particularly in embodiments where the data lake contains multiple data storage areas and the formats of this data are not easily accessible or readable (e.g., unstructured data, semi-structured data, structured data). Data can be transferred downstream from the data lake to a data warehouse, where it can be stored in a format that is easier to process, package, and consume. The storage system described above can also be used to implement such a data warehouse. Additionally, data marts or data centers can allow for more easily consumed data, where the storage system described above can also be used to provide the underlying storage resources required for the data mart or data center. In embodiments, querying a data lake may require a read schema approach, where data is applied to a plan or schema when it is pulled out of storage, rather than when it enters storage.

[0205] The storage system described herein can also be configured to implement Recovery Point Objectives ('RPOs'), which can be established by a user, by an administrator, as a system default, as part of a storage class or service provided by the storage system, or otherwise. A "Recovery Point Objective" is a target for the maximum time difference between the last update to the source dataset and the last recoverable replica dataset update, for reasons that ensure the update will be correctly recoverable from consecutive or frequently updated copies of the source dataset. An update is correctly recoverable if it properly accounts for all updates processed on the source dataset prior to the last recoverable replica dataset update.

[0206] In synchronous replication, the Recovery Point Objective (RPO) will be zero, meaning that under normal operation, all completed updates on the source dataset should exist and be correctly recovered on the replica dataset. In best-effort near-synchronous replication, the RPO can be as low as a few seconds. In snapshot-based replication, the RPO can be roughly calculated as the interval between snapshots plus the time spent transferring modifications between previously transferred snapshots and the most recently copied snapshot.

[0207] If the rate of update accumulation exceeds the rate of replication, the Recovery Point Objective (RPO) may be missed. If the amount of data to be replicated (for snapshot-based replication) accumulated between two snapshots exceeds the amount of data that can be replicated between taking a snapshot and copying the accumulated updates from that snapshot to the copy, the RPO may be missed. Again, in snapshot-based replication, if the data to be replicated accumulates at a rate faster than it is delivered over time between subsequent snapshots, replication may begin to lag further, potentially prolonging the miss between the expected recovery point target and the actual recovery point represented by the last correctly replicated update.

[0208] The storage systems described above can also be part of a shared-nothing (SNO) cluster. In a SNO cluster, each node has local storage and communicates with other nodes in the cluster via a network, where the storage used by the cluster is (typically) provided only by the storage connected to each individual node. A collection of nodes that synchronously replicates a dataset might be an example of a SNO cluster, as each storage system has local storage and communicates with other storage systems via a network, where those storage systems (typically) do not use storage from other places they access through some kind of interconnect. In contrast, some of the storage systems described above are built as shared storage clusters because there are drive racks shared by pairs of controllers. However, other storage systems described above are built as SNO clusters because all storage devices are local to specific nodes (e.g., blades), and all communication takes place through a network that links compute nodes together.

[0209] In other embodiments, other forms of shared-nothing storage clusters may include implementations where any node in the cluster has a local copy of all the storage devices it needs, and data is mirrored to other nodes in the cluster via synchronous replication to ensure that data is not lost, or because other nodes are also using the storage devices. In such embodiments, if a new cluster node needs some data, that data can be copied from other nodes that have data copies to the new node.

[0210] In some embodiments, a mirror-based shared storage cluster can store multiple copies of all the data stored in the cluster, wherein each subset of data is replicated to a specific set of nodes, and different subsets of data are replicated to different sets of nodes. In some variations, embodiments can store all the data stored in the cluster across all nodes, while in other variations, nodes can be partitioned such that a first set of nodes will all store the same dataset, and a second, different set of nodes will all store different datasets.

[0211] Readers will understand that RAFT-based databases (e.g., etcd) can operate like a shared-nothing storage cluster, with all RAFT nodes storing all the data. However, the amount of data stored in a RAFT cluster may be limited, so additional replicas won't consume too much storage. Container server clusters may also be able to replicate all data across all cluster nodes, provided the containers are not too large and their bulk data (data manipulated by applications running within them) is stored elsewhere, such as an S3 cluster or an external file server. In such examples, container storage can be provided directly by the cluster through its shared-nothing storage model, with these containers providing the image of the execution environment that forms part of an application or service.

[0212] To further explain, Figure 3D An exemplary computing device 350 is shown, which may be specifically configured to perform one or more processes described herein. Figure 3D As shown, computing device 350 may include communication interface 352, processor 354, storage device 356, and input / output (“I / O”) module 358, which are communicatively connected to each other via communication infrastructure 360. Although Figure 3D An exemplary computing device 350 is shown, but Figure 3D The components shown are not intended to be limiting. Other or alternative components may be used in other embodiments. A more detailed description will now follow. Figure 3D The components of the computing device 350 shown.

[0213] Communication interface 352 can be configured to communicate with one or more computing devices. Examples of communication interface 352 include, but are not limited to, wired network interfaces (e.g., network interface cards), wireless network interfaces (e.g., wireless network interface cards), modems, audio / video connections, and any other suitable interfaces.

[0214] Processor 354 generally refers to any type or form of processing unit capable of processing data and / or interpreting, executing, and / or directing the execution of one or more of the instructions, procedures, and / or operations described herein. Processor 354 may perform operations by executing computer-executable instructions 362 (e.g., applications, software, code, and / or other executable data instances) stored in storage device 356.

[0215] Storage device 356 may include one or more data storage media, devices, or configurations, and may employ data storage media and / or devices of any type, form, and combination. For example, storage device 356 may include, but is not limited to, any combination of non-volatile media and / or volatile media described herein. Electronic data, including the data described herein, may be temporarily and / or persistently stored in storage device 356. For example, data representing computer-executable instructions 362 configured to instruct processor 354 to perform any of the operations described herein may be stored in storage device 356. In some examples, data may be arranged in one or more databases within storage device 356.

[0216] I / O module 358 may include one or more I / O modules configured to receive user input and provide user output. I / O module 358 may include any hardware, firmware, software, or a combination thereof that supports input and output capabilities. For example, I / O module 358 may include hardware and / or software for capturing user input, including but not limited to a keyboard or keypad, a touchscreen component (e.g., a touchscreen display), a receiver (e.g., an RF or infrared receiver), a motion sensor, and / or one or more input buttons.

[0217] I / O module 358 may include one or more means for presenting output to a user, including, but not limited to, a graphics engine, a display (e.g., a screen), one or more output drivers (e.g., display drivers), one or more audio speakers, and one or more audio drivers. In some embodiments, I / O module 358 is configured to provide graphical data to the display for presentation to the user. The graphical data may represent one or more graphical user interfaces and / or any other graphical content that may serve a particular implementation. In some instances, any of the systems, computing devices, and / or other components described herein may be implemented by computing device 350.

[0218] To further explain, Figure 3E An example of a storage system 376 for providing storage services (also referred to herein as 'data services') is shown. Figure 3EThe group of storage systems 376 described herein includes multiple storage systems 374a, 374b, 374c to 374n, each of which may be similar to the storage system described herein. The storage systems 374a, 374b, 374c to 374n in the group of storage systems 376 may be the same storage system or different types of storage systems. For example, Figure 3E The two storage systems 374a and 374n depicted are described as cloud-based storage systems because the resources that together form each of the storage systems 374a and 374n are provided by different cloud service providers 370 and 372. For example, the first cloud service provider 370 could be Amazon AWS. TM The second cloud service provider, 372, is Microsoft Azure. TM However, in other embodiments, one or more public clouds, private clouds, or combinations thereof may be used to provide underlying resources for a particular storage system in a group forming storage system 376.

[0219] Figure 3E The examples depicted include edge management service 366 for providing storage services according to some embodiments of this disclosure. The provided storage services (also referred to herein as 'data services') may include, for example, services that provide a certain amount of storage to consumers, services that provide storage to consumers according to a pre-defined service tier agreement, services that provide storage to consumers according to pre-defined regulatory requirements, and many other services.

[0220] Figure 3E The edge management service 366 described herein can be embodied as, for example, one or more computer program instruction modules executing on computer hardware (e.g., one or more computer processors). Alternatively, the edge management service 366 can be embodied as one or more computer program instruction modules executing on a virtualized execution environment (e.g., one or more virtual machines), in one or more containers, or otherwise. In other embodiments, the edge management service 366 can be embodied as a combination of the embodiments described above, including embodiments in which one or more computer program instruction modules included in the edge management service 366 are distributed across multiple physical or virtual execution environments.

[0221] Edge management service 366 can operate as a gateway to provide storage services to storage consumers, where the storage services utilize storage provided by one or more storage systems 374a, 374b, 374c to 374n. For example, edge management service 366 can be configured to provide storage services to host devices 378a, 378b, 378c, 378d, 378n that are performing one or more applications consuming storage services. In such an example, edge management service 366 can operate as a gateway between host devices 378a, 378b, 378c, 378d, 378n and storage systems 374a, 374b, 374c to 374n, rather than requiring host devices 378a, 378b, 378c, 378d, 378n to directly access storage systems 374a, 374b, 374c to 374n.

[0222] Figure 3E 366's edge management service exposes its storage service module 364 to... Figure 3E The host devices are 378a, 378b, 378c, 378d, and 378n, but in other embodiments, the edge management service 366 can expose the storage service module 364 to other consumers of various storage services. Various storage services can be presented to consumers via one or more user interfaces, via one or more APIs, or through some other mechanism provided by the storage service module 364. Therefore, Figure 3E The storage service module 364 described herein can be embodied as one or more computer program instruction modules executed on physical hardware, a virtualized execution environment, or a combination thereof, wherein executing such a module enables consumers of storage services to obtain, select, and access various storage services.

[0223] Figure 3E The edge management service 366 also includes the system management service module 368. Figure 3EThe system management service module 368 includes one or more computer program instruction modules. When these instructions are executed, they coordinate with the storage systems 374a, 374b, 374c to 374n to perform various operations to provide storage services to the host devices 378a, 378b, 378c, 378d, and 378n. For example, the system management service module 368 can be configured to perform tasks such as allocating storage resources from storage systems 374a, 374b, 374c to 374n via one or more APIs exposed by storage systems 374a, 374b, 374c to 374n, migrating datasets or workloads between storage systems 374a, 374b, 374c to 374n via one or more APIs exposed by storage systems 374a, 374b, 374c to 374n, setting one or more tunable parameters (i.e., one or more configurable settings) on storage systems 374a, 374b, 374c to 374n via one or more APIs exposed by storage systems 374a, 374b, 374c to 374n, and so on. For example, many of the services described below relate to embodiments in which storage systems 374a, 374b, 374c to 374n are configured to operate in a certain way. In such an example, the system management service module 368 may be responsible for configuring the storage systems 374a, 374b, 374c to 374n to operate in the manner described below using the APIs (or some other mechanism) provided by the storage systems 374a, 374b, 374c to 374n.

[0224] In addition to configuring storage systems 374a, 374b, 374c through 374n, the edge management service 366 itself can also be configured to perform various tasks required to provide various storage services. Consider an example where the storage service includes a service that, when selected and applied, obfuscates personally identifiable information ('PII') contained in a dataset when the dataset is accessed. In such an example, storage systems 374a, 374b, 374c through 374n can be configured to obfuscate the PII when serving a read request to the dataset. Alternatively, storage systems 374a, 374b, 374c through 374n can service reads by returning data containing the PII, but the edge management service 366 itself can obfuscate the PII as the data passes through the edge management service 366 en route from storage systems 374a, 374b, 374c through 374n to host devices 378a, 378b, 378c, 378d, 378n.

[0225] Figure 3E The storage systems 374a, 374b, 374c to 374n described in the reference above can be embodied in the above reference. Figure 1A-3DThe description includes one or more storage systems, including variations thereof. In fact, storage systems 374a, 374b, 374c to 374n can be used as a storage resource pool, wherein the components in the pool have different performance characteristics, different storage characteristics, etc. For example, one storage system 374a may be a cloud-based storage system, another storage system 374b may be a storage system providing block storage, another storage system 374c may be a storage system providing file storage, another storage system 374d may be a relatively high-performance storage system, and another storage system 374n may be a relatively low-performance storage system, etc. In alternative embodiments, only one storage system may exist.

[0226] Figure 3E The storage systems 374a, 374b, 374c to 374n depicted can also be organized into different fault domains, such that a failure of one storage system 374a should be completely independent of a failure of another storage system 374b. For example, each storage system can receive power from an independent power system, and each storage device can communicate data through an independent data communication network. Furthermore, storage systems in the first fault domain can be accessed via a first gateway, while storage systems in the second fault domain can be accessed via a second gateway. For example, the first gateway can be a first instance of the edge management service 366, and the second gateway can be a second instance of the edge management service 366, including embodiments where each instance is different or each instance is part of the distributed edge management service 366.

[0227] As an illustrative example of available storage services, storage services associated with different levels of data protection can be presented to the user. For example, a storage service can be provided to the user that, when selected and executed, guarantees that the data associated with the user will be protected, thus ensuring various recovery point objectives ('RPO'). For instance, a first available storage service can ensure that some datasets associated with the user will be protected, enabling recovery of any data exceeding 5 seconds in the event of a failure of the primary data store, while a second available storage service can ensure that the datasets associated with the user will be protected, enabling recovery of any data exceeding 5 minutes in the event of a failure of the primary data store.

[0228] Another example of a storage service that can be presented to, selected by, and ultimately applied to a user's associated dataset can include one or more data compliance services. Such data compliance services can manifest as, for example, providing data compliance services to consumers (i.e., users) to ensure that the user's dataset is managed in a manner that complies with various regulatory requirements. For instance, one or more data compliance services can be provided to ensure that the user's dataset is managed in accordance with the General Data Protection Regulation ('GDPR'), one or more data compliance services can be provided to ensure that the user's dataset is managed in accordance with the Sarbanes-Oxley Act of 2002 ('SOX'), or one or more data compliance services can be provided to ensure that the user's dataset is managed in accordance with another regulatory act. Additionally, one or more data compliance services can be provided to ensure that the user's dataset is managed in accordance with a non-governmental guideline (e.g., best practices for audit purposes), or one or more data compliance services can be provided to ensure that the user's dataset is managed in a manner that meets the requirements of a specific client or organization, etc.

[0229] To provide this specific data compliance service, the service can be presented to the user (e.g., via a GUI) and selected by the user. In response to receiving a selection for a specific data compliance service, one or more storage service policies can be applied to the dataset associated with the user to perform the specific data compliance service. For example, a storage service policy can be applied requiring the dataset to be encrypted before being stored in a storage system, a cloud environment, or elsewhere. To enforce this policy, a requirement can be enforced not only to encrypt the dataset during storage but also to encrypt the dataset before transmission (e.g., sending the dataset to another party). In such an example, a storage service policy can also be implemented requiring that any encryption keys used to encrypt the dataset are not stored on the same system storing the dataset itself. The reader will appreciate that many other forms of data compliance services can be provided and implemented according to embodiments of this disclosure.

[0230] Storage systems 374a, 374b, 374c to 374n within the cluster of storage system 376 can, for example, be jointly managed by one or more cluster management modules. The cluster management module can be... Figure 3EThe system management service module 368, as depicted, may be part of or separate from it. The group management module can perform tasks such as monitoring the health of each storage system in the group, initiating updates or upgrades to one or more storage systems in the group, migrating workloads for load balancing or other performance purposes, and many other tasks. Therefore, for many other reasons, storage systems 374a, 374b, 374c to 374n can be coupled to each other via one or more data communication links to exchange data between storage systems 374a, 374b, 374c to 374n.

[0231] In some embodiments, one or more storage systems or one or more elements of storage systems (e.g., features, services, operations, components, etc. of the storage system), such as any illustrative storage system or storage system element described herein, may be implemented in one or more container systems. A container system may contain any system that supports the execution of one or more containerized applications or services. Such services may be software deployed for building applications, operating runtime environments, and / or serving as infrastructure for other services. In the following discussion, the description of containerized applications generally also applies to containerized services.

[0232] Containers can combine one or more elements of a containerized software application with a runtime environment for operating the software application elements bundled within a single image. For example, each such container of a containerized application may contain the software application's executable code and various dependencies, libraries and / or other components, as well as network configurations, and be configured to access additional resources used by the elements of the software application within that specific container to operate those elements. A containerized application can be represented as a collection of such containers that together represent all the elements of the application combined with the various runtime environments required for all these elements to run. Therefore, a containerized application can be abstracted from the host operating system as a combination of lightweight and portable packages and configurations, where it can be uniformly deployed and consistently executed in different computing environments using different container-compatible operating systems or different infrastructures. In some embodiments, a containerized application shares a kernel with the host computer system and executes as an isolated environment (an isolated collection of files and directories, processes, system and network resources, configured to access additional resources and functionality), isolated by the host system's operating system in conjunction with a container management framework. A containerized application, when executed, can provide one or more containerized workloads and / or services.

[0233] Container systems can contain and / or utilize clusters of nodes. For example, a container system can be configured to manage the deployment and execution of containerized applications on one or more nodes in a cluster. Containerized applications can utilize the resources of the nodes, such as memory, processing, and / or storage resources provided and / or accessed by the nodes. Storage resources can include any illustrative storage resources described herein and can include on-node resources such as local files and directory trees, off-node resources such as externally networked file systems, databases, or object stores, or both on-node and off-node resources. Access to additional resources and functionality that can be configured for containers of containerized applications can include dedicated computing capabilities such as GPUs and AI / ML engines, or dedicated hardware such as sensors and cameras.

[0234] In some embodiments, a container system may include a container orchestration system (also referred to as a container orchestrator, container orchestration platform, etc.) designed to be reasonably simple and, for many use cases, automated for deploying, scaling, and managing containerized applications. In some embodiments, a container system may include a storage management system configured to provide and manage storage resources (e.g., virtual volumes) for private or shared use by cluster nodes and / or containers of containerized applications.

[0235] Figure 3F An example container system 380 is illustrated. In this example, container system 380 includes a container storage system 381, which can be configured to perform one or more storage management operations to organize, provide, and manage storage resources for use by one or more containerized applications 382-1 to 382-L of container system 380. Specifically, container storage system 381 can organize storage resources into one or more storage pools 383 for use by containerized applications 382-1 to 382-L. The container storage system itself can be implemented as a containerized service.

[0236] Container system 380 may contain or be implemented by one or more container orchestration systems, including Kubernetes. TM Mesos TM Docker Swarm TM The container orchestration system can manage container systems 380 running on cluster 384 through services implemented by the control node, described as 385, and can further manage the container storage system or the relationships between individual containers and their storage devices, memory and CPU limits, network connectivity, and their access to additional resources or services.

[0237] The control plane of container system 380 can implement services including: deploying applications via controller 386, monitoring applications via controller 386, providing interfaces via API server 387, and scheduling deployments via scheduler 388. In this example, controller 386, scheduler 388, API server 387, and container storage system 381 are implemented on a single node (node ​​385). In other examples, for flexibility, the control plane can be implemented by multiple redundant nodes, where if the node providing management services for container system 380 fails, another redundant node can provide management services for cluster 384.

[0238] The data plane of the container system 380 can contain a set of nodes that provide container runtimes for executing containerized applications. Individual nodes within the cluster 384 can execute container runtimes, such as Docker. TM And it executes the container manager or node agent, such as a kubelet (not depicted) in Kubernetes that communicates with the control plane via a local network connection (sometimes called a proxy server, such as proxy 389). Proxy 389 can use, for example, Internet Protocol (IP) port numbers to route network traffic to and from containers. For example, a containerized application might request storage classes from the control plane, where the request is handled by the container manager, and the container manager uses proxy 389 to forward the request to the control plane.

[0239] Cluster 384 can contain a set of nodes that run containers for managed containerized applications. Nodes can be virtual machines or physical machines. Nodes can be host systems.

[0240] Container storage system 381 can orchestrate storage resources to provide storage to container system 380. For example, container storage system 381 can use storage pool 383 to provide persistent storage for containerized applications 382-1 to 382-L. Container storage system 381 itself can be deployed as a containerized application by a container orchestration system.

[0241] For example, container storage system 381 can be deployed within cluster 384 and perform management functions to provide storage for containerized application 382. Management functions may include identifying one or more storage pools from available storage resources, allocating virtual volumes on one or more nodes, replicating data, responding to and recovering from host and network failures, or handling storage operations. Storage pool 383 may contain storage resources from one or more local or remote sources, where the storage resources can be different types of storage devices, such as block storage devices, file storage devices, and object storage devices.

[0242] Container storage system 381 can also be deployed on a set of nodes for which a container orchestration system can provide persistent storage. In some examples, container storage system 381 can be deployed on all nodes of cluster 384 using, for example, a Kubernetes DaemonSet. In this example, nodes 390-1 to 390-N provide the container runtime executed by container storage system 381. In other examples, some, but not all, nodes in the cluster can execute container storage system 381.

[0243] Container storage system 381 can handle storage on nodes and communicate with the control plane of container system 380 to provide dynamic volumes, including persistent volumes. Persistent volumes can be mounted on nodes as virtual volumes, such as virtual volumes 391-1 and 391-P. After mounting virtual volume 391, containerized applications can request and use, or otherwise configure, the storage provided by virtual volume 391. In this example, container storage system 381 can mount a driver on the node's kernel, where the driver handles storage operations directed to the virtual volume. In this example, the driver can receive storage operations directed to the virtual volume, and in response, the driver can perform storage operations on one or more storage resources within storage pool 383, possibly under the guidance of or using additional logic within the container that implements container storage system 381 as a containerized service.

[0244] Container storage system 381 can determine available storage resources in response to being deployed as a containerized service. For example, storage resources 392-1 to 392-M can include local storage devices, remote storage devices (storage devices on individual nodes in a cluster), or both local and remote storage devices. Storage resources can also include storage devices from external sources, such as various combinations of block storage systems, file storage systems, and object storage systems. Storage resources 392-1 to 392-M can include any type and / or configured storage resources (e.g., any illustrative storage resources described above), and container storage system 381 can be configured to include configuration files in any suitable manner to determine available storage resources. For example, a configuration file can specify account and authentication information for cloud-based object storage device 348 or cloud-based storage system 318. Container storage system 381 can also determine the availability of one or more storage devices 356 or one or more storage systems. Aggregated storage from one or more of storage devices 356, storage systems, cloud-based storage systems 318, edge management services 366, cloud-based object storage devices 348, or any other storage resources, or any combination or sub-combination of these storage resources, may be used to provide storage pool 383. Storage pool 383 is used to provide storage for one or more virtual volumes mounted on one or more nodes 390 within cluster 384.

[0245] In some implementations, container storage system 381 can create multiple storage pools. For example, container storage system 381 can aggregate storage resources of the same type into a single storage pool. In this example, the storage type can be one of the following: storage device 356, storage array 102, cloud-based storage system 318, storage via edge management service 366, or cloud-based object storage device 348. Alternatively, it can be a tiered or typed redundant or distributed storage device, such as a specific combination of striping, mirroring, or erasure coding.

[0246] Container storage system 381 can execute as a containerized container storage system service within cluster 384, where container instances of the elements implementing the containerized container storage system service can operate on different nodes in cluster 384. In this example, the containerized container storage system service can combine the container orchestration system operations of container system 380 to handle storage operations, mount virtual volumes to provide storage to nodes, aggregate available storage devices into storage pool 383, allocate storage to virtual volumes from storage pool 383, generate backup data, replicate data between nodes, clusters, and environments, and perform other storage system operations. In some examples, the containerized container storage system service can provide storage services across multiple clusters operating in different computing environments. For example, other storage system operations may include the storage system operations described herein. The persistent storage provided by the containerized container storage system service can be used to implement stateful and / or flexible containerized applications.

[0247] The container storage system 381 can be configured to perform any suitable storage operation of the storage system. For example, the container storage system 381 can be configured to perform one or more illustrative storage management operations described herein to manage the storage resources used by the container system.

[0248] In some embodiments, one or more storage operations, including one or more illustrative storage management operations described herein, may be containerized. For example, one or more storage operations may be implemented as one or more containerized applications configured to perform storage operations. Such containerized storage operations can be executed in any suitable runtime environment to manage any storage system, including any illustrative storage system described herein.

[0249] The storage system described in this paper can support various forms of data replication. For example, two or more storage systems can synchronously replicate a dataset to each other. In synchronous replication, different copies of a particular dataset can be maintained by multiple storage systems, but all accesses to the dataset (e.g., reads) should produce consistent results, regardless of which storage system the access is directed to. For example, a read directed to any storage system that is synchronously replicating the dataset should return the same result. Therefore, while updates to dataset versions do not need to occur exactly simultaneously, precautions must be taken to ensure consistent access to the dataset. For example, if the first storage system receives an update (e.g., a write) to the dataset, the update can only be confirmed as complete if all storage systems synchronously replicating the dataset have applied the update to their copies of the dataset. In such an example, synchronous replication can be performed using I / O forwarding (e.g., a write received at the first storage system is forwarded to the second storage system), communication between storage systems (e.g., each storage system indicates that it has completed an update), or other methods.

[0250] In other embodiments, datasets can be replicated using checkpoints. In checkpoint-based replication (also known as 'near-synchronous replication'), sets of updates to the dataset (e.g., one or more write operations pointing to the dataset) may occur between different checkpoints, such that the dataset is only updated to a particular checkpoint if all updates to the dataset prior to that checkpoint have been completed. Consider an example where a first storage system stores a live copy of the dataset that a user is accessing. In this example, assume that the dataset is copied from the first storage system to a second storage system using checkpoint-based replication. For example, the first storage system might send a first checkpoint (at time t = 0) to the second storage system, followed by a first set of updates to the dataset, then a second checkpoint (at time t = 1), then a second set of updates to the dataset, and then a third checkpoint (at time t = 2). In such an example, if the second storage system has performed all updates in the first set of updates but not all updates in the second set, the copy of the dataset stored on the second storage system may be up-to-date until the second checkpoint. Alternatively, if the second storage system has performed all updates in both the first and second update sets, the copy of the dataset stored on the second storage system can be up-to-date until the third checkpoint. The reader will understand that various types of checkpoints can be used (e.g., metadata-only checkpoints), and checkpoints can be expanded based on various factors (e.g., time, number of operations, RPO settings), etc.

[0251] In other embodiments, the dataset can be replicated via snapshot-based replication (also known as 'asynchronous replication'). In snapshot-based replication, snapshots of the dataset can be sent from a replication source (e.g., a first storage system) to a replication target (e.g., a second storage system). In such embodiments, each snapshot can contain the entire dataset or a subset of the dataset, for example, only the portion of the dataset that has changed since the last snapshot was sent from the replication source to the replication target. The reader will understand that snapshots can be sent on demand based on strategies or other methods that take into account various factors (e.g., time, number of operations, RPO settings).

[0252] The storage systems described above can be configured individually or in combination as continuous data protection storage devices. Continuous data protection storage is a feature of storage systems that records updates to a dataset in such a way that a consistent picture of the dataset's previous contents can be accessed at a low temporal granularity (typically within seconds or even less) and backtracked over a reasonable time period (typically hours or days). This allows access to the latest consistent point in time of the dataset, and also allows access to points in time of the dataset that may have just occurred before an event, such as an event that caused partial corruption or otherwise loss of the dataset, while retaining the maximum number of updates close to said event. Conceptually, they are like a series of snapshots of a dataset taken frequently and stored for a long time, although the implementation of continuous data protection storage devices is generally quite different from snapshots. Storage systems implementing continuous data protection storage can also provide methods for accessing these points in time, accessing one or more points in time as snapshots or cloned copies, or restoring the dataset to one of these recorded points in time.

[0253] Over time, to reduce overhead, some points in time stored in continuous data protection storage can be merged with other nearby points in time, essentially deleting some of these points from the storage area. This reduces the capacity required to update the storage area. Alternatively, a limited number of these points in time can be converted into snapshots with a longer duration. For example, such storage might retain a low-granularity sequence of points in time several hours from now, merging or deleting some points to reduce overhead for another day. Going back in time, some of these points in time can be converted into snapshots, representing a consistent image of points in time every few hours.

[0254] Figure 4This is a block diagram illustrating an example storage system 400 according to some embodiments. Storage system 400 includes a computing device 440, a storage system controller 410 (e.g., a storage controller, storage node, node, storage array controller, storage controller node, etc.), storage nodes 420, and a network 405. System controller 410 and storage nodes 420 may include a processing device 411 and one or more non-volatile memory modules 413. Processing device 411 may be a device capable of executing instructions and / or performing various operations, actions, functions, etc. (e.g., one or more CPUs, ASICs, FPGAs, multi-core processors, processing cores, circuitry, etc.), as discussed in more detail below. Non-volatile memory modules 413 may be devices for storing data such that the data in the device is retained even when power is no longer supplied. Examples of non-volatile memory modules 413 may include, but are not limited to, SSDs, NVMe drives, flash drives, etc. Storage system controller 410 also includes a relocation module 415, which will be discussed in more detail below.

[0255] In some embodiments, the non-volatile memory module 413 may be removed from the storage node 420 and / or the storage system controller 410. For example, the non-volatile memory module 413 may be removed and replaced with another non-volatile memory module. The non-volatile memory module 413 may also be removable while the storage node 420 and / or the storage system controller 410 is in operation. For example, the non-volatile memory module 413 may be hot-swappable. The non-volatile memory module 413 may also be heterogeneous (e.g., non-uniform). For example, non-volatile memory modules may have different brands, models, manufacturers, capacities, and memory types (e.g., SLC, MLC, TLC, QLC, PLC, etc.).

[0256] In one embodiment, the non-volatile memory module 413 may include a multi-planar die. The multi-planar die may be a flash die (e.g., a flash memory die, a semiconductor die) containing multiple planes (e.g., multiple layers) of flash memory cells. Multi-planar dies will be discussed in more detail below.

[0257] Network 405 can utilize various technologies, including wireless connectivity, direct local area network (LAN) connectivity, wide area network (WAN) connectivity such as the Internet, routers, storage area networks, Ethernet, etc. Network 405 may include one or more LANs, which may also be wireless. Network 405 may further include Remote Direct Memory Access (RDMA) hardware and / or software, Transmission Control Protocol / Internet Protocol (TCP / IP) hardware and / or software, routers, repeaters, switches, mesh networks, etc. Protocols such as Fibre Channel, Fibre Channel over Ethernet (FCoE), iSCSI, etc., can be used in network 405. Network 405 can connect to a set of communication protocols used for the Internet (e.g., Transmission Control Protocol (TCP)) and Internet Protocol (IP) or TCP / IP. In one embodiment, network 405 represents a storage area network (SAN) that provides access to integrated block-level data storage devices. The SAN can be used to enhance the storage devices accessible to the computing device, making the non-volatile memory module 413 appear as a locally attached storage device to the computing device 440.

[0258] Computing device 440 refers to any number of fixed or mobile computing devices, such as desktop personal computers (PCs), servers, tablets, server clusters, workstations, laptops, handheld computers, personal digital assistants (PDAs), smartphones, etc. Generally, computing device 440 may also include one or more processing devices, which may further include one or more processor cores. Each processor core includes a circuitry for executing instructions according to a predefined general-purpose instruction set. For example, an x86 instruction set architecture may be selected. Alternatively, a different architecture may be selected. Or any other general-purpose instruction set architecture. The processor core can access the cache memory subsystem to obtain data and computer program instructions. The cache subsystem can be coupled to a memory hierarchy that includes random access memory (RAM) and storage devices. The computing device can utilize the memory system controller 410 and / or memory node 420 to read, write, store, and / or access data in the memory system 400.

[0259] In one embodiment, the storage system controller 410 can be connected to other storage system controllers ( Figure 4(Not shown in the diagram) Parallel operation. For example, storage system controller 410 can be combined with other storage system controllers as a distributed system operation, where storage operations can be distributed among storage system controller 410 and other storage system controllers (e.g., load balancing, based on scheduling algorithms / mechanisms, or based on availability distribution, etc.). In one embodiment, storage system controller 410 is designated as the "primary" storage system controller, which performs most or all I / O operations on storage device group 430. However, in the event of a software crash, hardware failure, or other error, the "secondary" storage system controller (… Figure 4 (Not shown in the image) The primary master controller is promoted and takes over all responsibilities for servicing the storage device group 430.

[0260] The storage system controller 410 may include firmware, software, and / or hardware configured to provide access to the non-volatile memory module 413. For example, the storage system controller 410 may include an operating system, a processor (e.g., a processing device, CPU, ASIC, FPGA), memory (e.g., RAM, NVRAM, DRAM, etc.), and / or applications (…). Figure 4 (Not shown in the image).

[0261] In one embodiment, when writing a dataset to multiple planes of a multi-plane die, the storage system controller 410 can divide the dataset into different parts (e.g., different subsets). Each part (e.g., subset) of the dataset can be written to one plane of the multi-plane die. For example, if there are four planes in the multi-plane die, the dataset can be divided into four parts, and each part can be written to a different plane of the multi-plane die. This can allow data to be written to the multi-plane die faster and / or more efficiently.

[0262] In one embodiment, the storage system controller 410 (e.g., relocation module 415) can determine that the number of planes used simultaneously for accessing data should be changed. As described above, each non-volatile memory module 413 (e.g., SSD, flash drive, NVMe drive) may contain one or more dies (e.g., flash dies), and said one or more dies may be multi-plane dies (e.g., dies with multiple levels or planes). The storage system controller 410 can use multiple planes simultaneously to access data (e.g., read and / or write data) (e.g., data can be written to multiple planes simultaneously). The relocation module 415 can determine that more or fewer planes should be used to access data simultaneously. For example, the relocation module 415 can determine that more planes should be used to write data simultaneously (e.g., more planes should be written in parallel).

[0263] In one embodiment, when an additional (e.g., a new) non-volatile memory module 413 is added to the storage system 400, the storage system controller 410 (e.g., relocation module 415) can determine that the number of planes simultaneously used for accessing data should be changed. For example, the new non-volatile memory module 413 can be added to an existing storage node 420 and / or the storage system controller 410. In another example, a new storage system controller 410 and / or a new storage node 420 can be added to the storage system 400.

[0264] In another embodiment, when the non-volatile memory module 413 is removed from the storage system 400, the storage system controller 410 (e.g., relocation module 415) may determine that the number of planes simultaneously used for accessing data should be changed. For example, the non-volatile memory module 413 may be removed from the existing storage node 420 and / or storage system controller 410. In another example, the storage system controller 410 and / or storage node 420 may be removed from the storage system 400.

[0265] In one embodiment, the storage system controller 410 (e.g., relocation module 415) can determine the number of planes simultaneously used for data access that should be changed based on a request indicating that the number of planes should be changed. For example, a user (e.g., a system / network administrator, engineer, etc.) can determine that the number of planes simultaneously used for data access should be changed (e.g., increased or decreased). The user can provide user input indicating that the number of planes should be changed, and the user input can cause a request (e.g., a message indicating the number of planes to be changed or other instructions) to be sent to the relocation module 415.

[0266] In one embodiment, the number of planes used for simultaneous data access can be increased or decreased based on various criteria, factors, parameters, etc. For example, the number of planes used for simultaneous data access can be based on one or more of the following: the erase block size of storage system 400 (e.g., erase block size), the amount of storage space in storage system 400 (e.g., total storage space in storage system 400), the predicted usage period / lifetime of the data, the type of data (e.g., the content of the data), and / or other characteristics of the data stored in storage system 400.

[0267] In one embodiment, the number of planes simultaneously used for data access can be increased or decreased based on power usage constraints, requirements, limitations, conditions, etc. For example, using more planes (simultaneously used for data access) can improve power efficiency, as discussed in more detail below. This allows the storage system controller 410 (e.g., relocation module 415) to use less power based on power usage constraints. For example, if there is a specific amount of power that the storage system 400 should use (e.g., maximum power), the storage system controller 410 can keep the storage system 400 within that specific amount of power by increasing the number of planes simultaneously used for data access. In another example, if electricity prices (e.g., utility rates) are more expensive at different times of the day (e.g., more expensive during the day than at night), the storage system controller 410 can increase the number of planes during the more expensive times and decrease the number of planes during the less expensive times. The storage system controller 410 can automatically change (e.g., adjust) the number of planes simultaneously used for data access, increasing or decreasing said number based on power usage constraints, requirements, limitations, conditions, etc. For example, storage system 410 can continuously monitor power usage, utility rates or other conditions / parameters, and can automatically change the number of planes used to access data simultaneously.

[0268] In one embodiment, a storage system controller 410 (e.g., a relocation module 415) can move, relocate, copy, shift, transfer, or otherwise transfer one or more portions of an existing erase block to a new erase block. For example, the storage system controller 410 may have existing erase blocks allocated in the storage system 400. Data used by the computing device 440 may be stored in existing erase blocks. The relocation module 410 can identify, select, and determine one or more portions of an existing erase block (e.g., information blocks, fragments, sections, segments, etc.) that should be moved to the new erase block. The relocation module 410 can identify, select, and determine one or more portions of an existing erase block based on various criteria, factors, parameters, etc. For example, the relocation module 410 can identify one or more portions of an existing erase block based on metadata. The metadata may contain information about the erase block and / or information about the characteristics of the data in the erase block, which will be discussed in more detail below.

[0269] In one embodiment, one or more portions of an existing erase block may contain real-time data. For example, real-time data may be data that has not yet been marked for deletion / erasure (e.g., data that a user or computing device has not yet indicated should be deleted). In another example, real-time data may be data that is currently in use (e.g., data that is still accessible to the user or computing device 440). An existing erase block may also contain expired data. For example, expired data may be data that has been marked for deletion but has not yet been garbage collected by the storage system 400 (e.g., data that has not been garbage collected by one or more of the non-volatile memory module 413, storage node 420, and storage system controller 410). Expired data may also be referred to as obsolete data, inactive data, stagnant data, obsolete data, etc.

[0270] In one embodiment, after the storage system controller 410 (e.g., relocation module 415) determines that the number of planes simultaneously used for data access should be changed, one or more portions of an existing erase block (which may contain real-time data) can be moved to a new erase block at different times. For example, the relocation module 415 may move the data immediately after determining that the number of planes simultaneously used for data access should be changed. In another example, the relocation module 415 may move the data some time after determining that the number of planes simultaneously used for data access should be changed (e.g., the data may be moved in the background or as part of a background process).

[0271] In one embodiment, the size of the new erase block can differ from the size of one or more existing erase blocks. For example, if the number of planes used for simultaneous data access is increased, the size of the new erase block can be larger than the size of one or more existing erase blocks. In another example, if the number of planes used for simultaneous data access is decreased, the size of the new erase block can be smaller than the size of one or more existing erase blocks. The size of the new erase block can also be based on various other parameters, criteria, conditions, etc. For example, the size of the new erase block can be based on the characteristics of the data to be stored in the new erase block (e.g., expected lifetime, content, etc.).

[0272] In one embodiment, storage system controller 410 (e.g., relocation module 415) may move, relocate, or otherwise replace portions of one or more existing erased blocks with one or more new erased blocks during a garbage collection operation. The garbage collection operation may include reclaiming (e.g., reusing) blocks containing invalid data (or no data) and allocating / redistributing these blocks for writing new data. For example, a block may be garbage collected when it is used in a new erased block (e.g., the block is erased and / or included as one of the blocks in the new erased block). Garbage collection may also include moving real-time data (e.g., blocks containing real-time data) to other locations (e.g., to other erased blocks). Storage system controller 410 and / or storage node 420 may perform garbage collection periodically. For example, storage system controller 410 and / or storage node 420 may perform garbage collection operations at regular intervals / cycles or based on a schedule.

[0273] In one embodiment, storage system controller 410 (e.g., relocation module 415) may move, relocate, or move portions of one or more existing erased blocks to one or more new erased blocks during a defragmentation operation (e.g., a defragmentation process). The defragmentation operation includes moving real-time data (e.g., blocks containing real-time data) to other locations (e.g., to other erased blocks). Storage system controller 410 and / or storage node 420 may perform defragmentation periodically. For example, storage system controller 410 and / or storage node 420 may perform defragmentation operations at regular intervals / cycles or based on a schedule.

[0274] In one embodiment, using multiple planes 511 simultaneously to access data can improve... Figure 4 The performance of the storage system 400 shown is illustrated. For example, accessing data simultaneously from multiple planes can make the storage system access data faster, as discussed in more detail below.

[0275] In one embodiment, accessing data simultaneously from multiple planes (of a multi-plane die) can reduce the amount of energy consumed and / or used by the storage system 400. For example, accessing data simultaneously from multiple planes can reduce the power usage of the storage system compared to accessing data serially from multiple planes of a multi-plane die, as discussed in more detail below.

[0276] In one embodiment, accessing data simultaneously from multiple planes (of a multi-plane die) can reduce the amount of heat generated by the storage system 400. For example, accessing data simultaneously from multiple planes can reduce the amount of heat generated by the storage system compared to accessing data serially from multiple planes of a multi-plane die, as discussed in more detail below.

[0277] As data storage demands increase and / or change, it may be useful for a data storage system (e.g., storage system 400) to operate faster and / or more efficiently. For example, the ability to read / write data faster is often useful and / or desirable. Changing the number of planes used simultaneously for accessing data in a multi-plane die (e.g., increasing the number of planes) can allow storage system 400 to operate faster and / or more efficiently (e.g., writing more data over a period of time). In another example, the energy usage of storage system 400 can be more efficient. However, since changing the number of planes used may also change the size of erase blocks, problems may arise when allocating / deallocating erase blocks. Real-time data in an existing erase block (e.g., one or more portions of an existing erase block containing real-time data) may prevent the existing erase block from being reused in a new erase block. For example, if there is still real-time data in an existing erase block, the existing erase block (e.g., a physical block that is part of an existing erase block) may not be deallocated, reused, etc., in other erase blocks (e.g., in a new erase block), as discussed in more detail below.

[0278] The examples, embodiments, and implementations described herein allow storage system 400 to allocate and / or deallocate erase blocks more efficiently. For example, relocation module 415 can identify erase blocks containing real-time data and move the real-time data. Moving real-time data from existing erase blocks allows blocks in an erase block to be used in a new erase block. Relocation module 415 can use metadata and / or data characteristics in existing erase blocks to identify the real-time data that should be moved. This can reduce and / or prevent real-time data from preventing the reuse of blocks in existing erase blocks.

[0279] Figure 5 This is a block diagram illustrating an example non-volatile memory module 413 according to some embodiments. As discussed above, the non-volatile memory module 413 can be a means of storing data such that the data in the means is retained even when power is no longer supplied to the means. The non-volatile memory module 413 can be stored from the memory node and / or the memory system controller (e.g., Figure 4 Removed from storage node 420 and / or storage system controller 410 shown. Non-volatile memory module 413 may also be heterogeneous (e.g., non-uniform, the brand, model and / or storage capacity of non-volatile memory module 413 may be different).

[0280] The non-volatile memory module 413 includes one or more multi-planar dies 510 (e.g., one or more multi-planar dies 510). A multi-planar die 510 may be a die containing multiple planes 511 (e.g., a flash memory die, a semiconductor die). For example, a multi-planar die 510 may contain planes 511 stacked on top of each other. Each plane 511 may contain a flash memory cell (e.g., a single-layer cell, a multi-layer cell, etc.). Each multi-planar die 510 contains four planes 511 (e.g., four layers of flash memory cells). In one embodiment, a plane 511 may be a cell on the die where operations (e.g., read or write) can be performed. A plane 511 may also be referred to as a layer.

[0281] Although the multi-planar die 510 is shown as having four planes, in other embodiments, the multi-planar die 510 may contain any number of planes (e.g., it may contain eight planes, twenty planes, one hundred planes, etc.). Additionally, although... Figure 5 The non-volatile memory module 413 shown includes a multi-planar die 510 having the same number of planes 511, but in other embodiments, the non-volatile memory module 413 may include multi-planar dies 510 with different numbers of planes 511.

[0282] As described above, multiple planes 511 can be used simultaneously to access data. For example, multiple planes 511 can be used to write data to the multi-plane die 510 simultaneously (e.g., data can be written to two or more planes 511 at the same time). In another example, multiple planes 511 can be used to read data from the multi-plane die 510 simultaneously (e.g., data can be read from two or more planes 511 at the same time).

[0283] In one embodiment, using multiple planes 511 simultaneously to access data can improve the storage system (e.g., Figure 4 The performance of the storage system 400 shown. For example, reading data simultaneously from multiple planes 511 allows the storage system to read data faster than reading data serially from multiple planes 511. When reading data serially from planes 511 (e.g., one plane 511 at a time), the amount of time to read data can be the time to access data from a plane 511 multiplied by the number of planes 511 accessed (e.g., time * number of planes). When reading data simultaneously from multiple planes 511, the amount of time to read data can be the same as the time to access data from a single plane 511 because multiple planes 511 are accessed simultaneously.

[0284] In one embodiment, using multiple planes 511 simultaneously to access data can improve the efficiency of the storage system. For example, serially reading data from plane 511 can use a certain amount of power / energy. Reading data from multiple planes 511 at once can use a smaller / less amount of power / energy compared to serially reading data from plane 511, because the overhead of accessing the multi-plane die can be shared among the planes 511 of the multi-plane die 510.

[0285] In one embodiment, using multiple planes 511 simultaneously to access data can reduce the amount of heat generated by the storage system. As discussed above, using multiple planes 511 to access data simultaneously can reduce the amount of power used by the storage system. Using less power also results in the storage system generating less heat.

[0286] Figure 6A This is a block diagram illustrating example erase blocks 610A, 610B, 610C, 610D, and 620 according to some embodiments. An erase block can be a set of blocks (e.g., one or more physical blocks) grouped together (e.g., logically grouped together). An erase block can be a unit on which a storage system can allocate blocks for writing data. For example, a storage system can allocate one erase block at a time to write new data.

[0287] Erasing blocks can contain real-time data and / or expired data. For example, eraser block 610A contains a portion 611 with real-time data and six portions 615 with expired data. Erasing block 610B contains four portions 611 with real-time data and three portions 615 with expired data. Erasing block 610C contains five portions 611 with real-time data and two portions 615 with expired data. Erasing block 610D contains one portion 611 with real-time data and six portions 615 with expired data. Real-time data can be data that is still in use and / or has not yet been marked (e.g., labeled, tiled, etc.) for deletion. Expired data can be data that has been marked for deletion but has not yet been deleted / erased (e.g., not yet garbage collected).

[0288] As discussed above, the storage system can change the number of planes used simultaneously for data access in a multi-plane die. When the storage system increases the number of planes used simultaneously, it can also increase the size of the erase blocks. For example, erase blocks 610A, 610B, 610C, and 610D can contain seven sections of the same size. If the storage system increases the number of planes used simultaneously, the size of the erase block can increase to the size of erase block 620 (e.g., it can have fourteen sections). For example, a newly allocated erase block (after increasing the number of planes used simultaneously) can be larger than the old / existing erase blocks in the storage system. In one embodiment, after the storage system increases the number of planes used simultaneously for data access, smaller and larger erase blocks may exist in the storage system. For example, when the storage system begins allocating newer / larger erase blocks, there may be smaller erase blocks that have not yet been deallocated because they still contain real-time data.

[0289] It can be beneficial for storage systems (e.g., storage system controllers, relocation modules, etc.) to track, monitor, and manage real-time data in erased blocks to allow for more efficient use of blocks (e.g., data blocks, physical blocks, etc.) within the storage system. In one embodiment, the storage system controller may use metadata 650 to track, monitor, and manage real-time data in erased blocks. For example, metadata 650 may include a list of erased blocks in the storage system, a list of physical blocks contained in an erased block (e.g., a list of physical blocks for each erased block), the size of the erased block, and / or which portions of the erased block contain real-time data and which portions of the erased block contain expired data. Metadata 650 may also include other information, such as the characteristics of the data in the erased block (e.g., the type / content of the data in the erased block, the expected lifetime of the data in the erased block, etc.).

[0290] In one embodiment, a storage system (e.g., a storage system controller, a relocation module, etc.) can relocate or move real-time data to allow the deallocation of existing erased blocks and their use for allocating new erased blocks (e.g., allowing the use of blocks from existing erased blocks in new erased blocks). The storage system controller can use metadata 650 to determine, identify, select, etc., portions of erased blocks that should be moved, relocated, copied, etc., to other (e.g., new) erased blocks. For example, the storage system controller can identify erased blocks with a real-time data volume below a threshold. The storage system controller can move or relocate real-time data on the identified erased blocks to one or more other erased blocks. For example, the storage system controller can move / relocate real-time data from existing erased blocks to new erased blocks (which may be larger than the existing erased blocks). Figure 6AAs shown, erase blocks 610A and 610D each have a portion 611 (e.g., shaded portion 611) containing real-time data. A storage system (e.g., a storage system controller, relocation module, etc.) can identify erase blocks 610A and 610D because the amount of real-time data in erase blocks 610A and 610D is below a threshold. The storage system controller can copy the portion 611 (e.g., shaded portion 611) containing real-time data from erase blocks 610A and 610D to a new erase block 620. Erase block 620 (e.g., the new erase block) also includes a portion 617. Portion 617 may contain real-time data (e.g., real-time data from other existing erase blocks, data recently received from the client / computing device, etc.), and / or may contain padding data (e.g., data used to complete or fill the erase block when insufficient data is received from the user / computing device).

[0291] In one embodiment, moving the live (e.g., shadowed) portion 611 allows the storage system to deallocate erased blocks 610A and 610D and allocate a new erased block 620 using blocks / portions from erased blocks 610A and 610D. This allows the storage system to use blocks in the storage system more efficiently. For example, moving the live (e.g., shadowed) portion 611 allows the storage system to allocate a new erased block 620 without having to wait for the live (e.g., shadowed) portion 611 to be deleted and / or marked for deletion. This allows the storage system to allocate new erased blocks faster without wasting portions of erased blocks 610A and 610D containing invalidated data.

[0292] Figure 6B This is a block diagram illustrating example erase blocks 630 and 640 according to some embodiments. As discussed above, an erase block can be a set of blocks (e.g., one or more physical blocks) grouped together (e.g., logically grouped together). An erase block can contain real-time data and / or failed data. For example, erase block 630 includes a portion 611 containing real-time data and six portions 615 containing failed data.

[0293] As discussed above, a storage system can change the number of planes simultaneously used for data access in a multi-plane die. When the storage system reduces the number of planes used simultaneously, it can also reduce the size of the erase blocks. For example, erase block 630 may contain fourteen sections of the same size. If the storage system reduces the number of planes used simultaneously, the size of the erase block can be reduced to the size of erase block 640 (e.g., it may have seven sections). For example, a newly allocated erase block (after reducing the number of planes used simultaneously) may be smaller than the old / existing erase blocks in the storage system. In one embodiment, after the storage system reduces the number of planes used for data access simultaneously, smaller and larger erase blocks may exist in the storage system. For example, the storage system may begin allocating newer / larger erase blocks, and there may be larger erase blocks that have not yet been deallocated because they still contain real-time data.

[0294] It can be beneficial for storage systems (e.g., storage system controllers, relocation modules, etc.) to track, monitor, and manage real-time data in erased blocks to allow for more efficient use of blocks (e.g., data blocks, physical blocks, etc.) within the storage system. In one embodiment, the storage system controller may use metadata 650 to track, monitor, and manage real-time data in erased blocks. As discussed above, metadata 650 may include a list of erased blocks, a list of physical blocks contained within an erased block, the size of the erased block, which portions of the erased block contain real-time data and which portions of the erased block contain invalid data, and / or other information such as the characteristics of the data in the erased block.

[0295] In one embodiment, the storage system (e.g., storage system controller, relocation module, etc.) can relocate or move live data to allow the deallocation of existing erase blocks and their use for allocating new erase blocks. The storage system controller can use metadata 650 to determine, identify, select, etc., portions of erase blocks that should be moved, relocated, copied, etc., to other (e.g., new) erase blocks. The storage system controller can move or relocate live data on the identified erase blocks to one or more other erase blocks. Figure 6B As shown, erase block 630 has two portions 611 (e.g., shaded portions 611) containing real-time data. A storage system (e.g., a storage system controller, a relocation module, etc.) can identify erase block 630 because the amount of real-time data in erase block 630 is below a threshold. The storage system controller can copy the portion 611 (e.g., shaded portions 611) containing real-time data from erase block 630 to a new erase block 640.

[0296] In one embodiment, moving the active (e.g., shadow) portion 611 allows the storage system to deallocate erased block 630 and use blocks / portions from erased block 630 to allocate new erased block 640 (and other erased blocks). This allows the storage system to use blocks in the storage system more efficiently. For example, moving the live (e.g., shadow) portion 611 allows the storage system to allocate new erased block 640 without having to wait for the live (e.g., shadow) portion 611 to be deleted and / or marked for deletion. This allows the storage system to allocate new erased blocks faster without wasting portions of erased block 630 containing invalid data.

[0297] Figure 7 This is a flowchart illustrating a method 700 for performing storage operations according to some embodiments of the present disclosure. Method 700 can be executed by processing logic, which includes hardware (e.g., circuitry, dedicated logic, programmable logic, microcode, etc.), software (e.g., instructions running on a processing device to execute hardware emulation), firmware, or a combination thereof. In one embodiment, method 700 can be executed by a storage system controller, relocation module, processing device, and / or authorization of the storage system shown in Figures 1-4.

[0298] The method begins at block 705, where method 700 can determine whether the number of planes (in a multi-plane die) used for simultaneous data access has changed. For example, method 700 can determine whether a new non-volatile memory module (e.g., a new SSD, a new flash drive, a new NVMe drive, etc.) has been added to the storage system. In another example, method 700 can determine whether a request to change the number of planes has been received.

[0299] In box 710, method 700 may optionally identify one or more existing erase blocks. For example, method 700 may use metadata to identify erase blocks containing amounts of real-time data below a threshold. In box 715, method 700 may move (e.g., copy, relocate, etc.) one or more portions of an erase block to a new erase block. For example, method 700 may copy a portion of an erase block containing real-time data to a new erase block. The new erase block may have a different size than the existing erase blocks (e.g., it may be larger or smaller than the existing erase blocks).

[0300] In block 720, method 700 may optionally use an erase block of the same size as the new erase block to write data (e.g., additional data) to the storage system. For example, if the erase block size has been increased to a larger size, data can be written to the non-volatile memory module using an erase block of a larger size.

[0301] Although some embodiments have been described primarily in the context of storage systems, those skilled in the art will recognize that embodiments of this disclosure may also take the form of a computer program product disposed on a computer-readable storage medium for use by any suitable processing system. Such a computer-readable storage medium can be any storage medium for machine-readable information, including magnetic media, optical media, solid-state media, or other suitable media. Examples of such media include disks in hard disk drives or floppy disks, optical disks in optical disk drives, magnetic tapes, and other media contemplated by those skilled in the art. Those skilled in the art will readily recognize that any computer system with appropriate programming means is capable of performing the steps embodied in the computer program product described herein. Those skilled in the art will also recognize that while some embodiments described in this specification are oriented toward software installed and executed on computer hardware, alternative embodiments as firmware or hardware implementations are entirely within the scope of this disclosure.

[0302] In some instances, a non-transitory computer-readable medium may be provided for storing computer-readable instructions, based on the principles described herein. When executed by a processor of a computing device, the instructions may direct the processor and / or the computing device to perform one or more operations, including one or more operations described herein. Such instructions may be stored and / or transmitted using any of a variety of known computer-readable media.

[0303] As used herein, a non-transitory computer-readable medium can include any non-transitory storage medium that contributes to providing data (e.g., instructions) that can be read and / or executed by a computing device (e.g., by a processor of the computing device). For example, a non-transitory computer-readable medium can include, but is not limited to, any combination of non-volatile storage media and / or volatile storage media. Exemplary non-volatile storage media include, but are not limited to, read-only memory, flash memory, solid-state drives, magnetic storage devices (e.g., hard disks, floppy disks, magnetic tapes, etc.), ferroelectric random access memory (“RAM”), and optical discs (e.g., compact discs, digital video discs, Blu-ray discs, etc.). Exemplary volatile storage media include, but are not limited to, RAM (e.g., dynamic RAM).

[0304] The advantages and features of this disclosure can be further described by the following statements:

[0305] Statement 1. A storage system comprising:

[0306] Multiple non-volatile memory modules, each non-volatile memory module comprising a multi-planar die;

[0307] A storage system controller operatively coupled to a plurality of storage devices, the storage system controller including a processing unit configured to:

[0308] Determine the number of planes used for accessing data that should be changed simultaneously on the multi-plane die;

[0309] In response to determining that the number of planes used for accessing data should be changed while the multi-plane die is being processed, one or more portions are moved from an existing erase block to a new erase block, the size of which differs from the size of the new erase block.

[0310] Statement 2. The storage system according to Statement 1, wherein, in order to determine the number of planes used for accessing data while the multi-plane die should be changed, the processing device is further configured to:

[0311] It has been determined that an additional non-volatile memory module has been added to the memory system.

[0312] Statement 3. The storage system according to statements 1 to 2, wherein, in order to determine the number of planes used for accessing data simultaneously on the multi-plane die that should be changed, the processing device is further configured to:

[0313] Receive a request to change the number of planes used to access data while simultaneously changing the number of planes on the multi-plane die.

[0314] Statement 4. The storage system according to statements 1 to 3, wherein one or more portions of the existing erase block include real-time data and other portions of the existing erase block include invalid data.

[0315] Statement 5. The storage system according to statements 1 to 4, wherein the processing device is further configured to:

[0316] The metadata identifies one or more portions of the existing erase block, which indicates real-time data within the existing erase block.

[0317] Statement 6. The storage system according to statements 1 to 5, wherein the size of the new erase block is larger than the size of the existing erase block.

[0318] Statement 7. The storage system according to statements 1 to 6, wherein the number of planes used for accessing data is increased while the number of multi-plane dies is increased.

[0319] Statement 8. The storage system according to statements 1 to 7, wherein the processing device is further configured to:

[0320] Data is written to the non-volatile memory module, wherein different portions of the data are written to different planes of the multi-plane die, and the number of different portions is equal to the number of planes.

[0321] Statement 9. The storage system according to statements 1 to 8, wherein during garbage collection operations of the storage system, one or more portions are moved from the existing erase block to the new erase block.

[0322] Statement 10. The storage system according to statements 1 to 9, wherein during the defragmentation operation of the storage system, one or more portions are moved from the existing erase block to the new erase block.

[0323] Statement 11. The storage system according to statements 1 to 10, wherein simultaneously accessing data from the multiple planes of the multiplane die reduces the amount of power used by the storage system compared to serially accessing data from multiple planes of the multiplane die.

[0324] Statement 12. The storage system according to statements 1 to 11, wherein simultaneously accessing data from the multiple planes of the multiplane die reduces the amount of heat generated by the storage system compared to serially accessing data from multiple planes of the multiplane die.

[0325] Statement 13. A method comprising:

[0326] Determine the number of planes used for accessing data simultaneously on a multi-plane die, which is part of one of a plurality of non-volatile memory modules in a storage system;

[0327] In response to determining that the number of planes used for accessing data should be changed while the multi-plane die is being processed, one or more portions are moved from an existing erase block to a new erase block, the size of which differs from the size of the new erase block.

[0328] Statement 14. The method according to Statement 13, wherein determining the number of planes used for accessing data while the multi-plane die should be changed includes:

[0329] It has been determined that an additional non-volatile memory module has been added to the memory system.

[0330] Statement 15. The method according to Statements 13 to 14, wherein one or more portions of the existing erase block include real-time data and other portions of the existing erase block include failure data.

[0331] Statement 16. The method according to statements 13 to 15 further includes:

[0332] The metadata identifies one or more portions of the existing erase block, which indicates real-time data within the existing erase block.

[0333] Statement 17. The method according to statements 13 to 16, wherein during garbage collection operations of the storage system, one or more portions are moved from the existing erase block to the new erase block.

[0334] Statement 18. The method according to statements 13 to 17, wherein during the defragmentation operation of the storage system, one or more portions are moved from the existing erase block to the new erase block.

[0335] Statement 19. The method according to statements 13 to 18, wherein simultaneously accessing data from the multiple planes of the multiplane die reduces the amount of power used by the storage system compared to serially accessing data from multiple planes of the multiplane die.

[0336] Statement 20. A non-transitory computer-readable storage medium for storing instructions, which, when executed, cause a processing apparatus of a storage system controller to perform the following operations:

[0337] Determine the number of planes used for accessing data simultaneously on a multi-plane die, which is part of one of a plurality of non-volatile memory modules in a storage system;

[0338] In response to determining that the number of planes used for accessing data should be changed while the multi-plane die is being processed, one or more portions are moved from an existing erase block to a new erase block, the size of which differs from the size of the new erase block.

[0339] This document describes one or more embodiments by means of method steps illustrating the execution of specified functions and their relationships. For ease of description, the boundaries and order of these functional building blocks and method steps are arbitrarily defined herein. Alternative boundaries and orders may be defined as long as the specified functions and their relationships are properly performed. Therefore, any such alternative boundaries or orders are within the scope and spirit of the claims. Furthermore, the boundaries of these functional building blocks are arbitrarily defined herein for ease of description. Alternative boundaries may be defined as long as certain essential functions are properly performed. Similarly, flowchart blocks may also be arbitrarily defined herein to illustrate certain essential functionalities.

[0340] Within the scope of use, flowchart block boundaries and sequences may be defined in other ways while still performing certain important functionalities. Therefore, this alternative definition of functional building blocks and flowchart blocks and sequences is within the scope and spirit of the claims. Those skilled in the art will also recognize that the functional building blocks and other illustrative blocks, modules, and components described herein can be implemented as illustrated, or by discrete components, application-specific integrated circuits, processors executing appropriate software, etc., or any combination thereof.

[0341] While specific combinations of various functions and features of one or more embodiments are explicitly described herein, other combinations of these features and functions are equally possible. This disclosure is not limited to the specific examples disclosed herein, and such other combinations are expressly incorporated.

Claims

1. A storage system comprising: Multiple non-volatile memory modules, each non-volatile memory module comprising a multi-planar die; A storage system controller operatively coupled to a plurality of storage devices, the storage system controller including a processing unit configured to: Determine the number of planes used for accessing data that should be changed simultaneously on the multi-plane die; In response to determining that the number of planes used for accessing data on the multi-plane die should be changed, blocks are allocated using a new block size by combining the set of erase blocks at the same address in individual planes based on the newly determined number of planes.

2. The storage system of claim 1, wherein, in order to determine the number of planes used for accessing data simultaneously on the multi-plane die that should be changed, the processing device is further configured to: It has been determined that an additional non-volatile memory module has been added to the memory system.

3. The storage system of claim 1, wherein, in order to determine the number of planes used for accessing data simultaneously on the multi-plane die that should be changed, the processing device is further configured to: Receive a request to change the number of planes used to access data while simultaneously changing the number of planes on the multi-plane die.

4. The storage system according to claim 1, further comprising: Move one or more portions from an existing erase block to a new erase block, the existing erase block having a new block size.

5. The storage system of claim 4, wherein the processing device is further configured to: The metadata identifies one or more portions of the existing erase block, which indicates real-time data within the existing erase block.

6. The storage system of claim 4, wherein during a garbage collection operation of the storage system, one or more portions are moved from the existing erase block to the new erase block.

7. The storage system of claim 4, wherein during the defragmentation operation of the storage system, one or more portions are moved from the existing erase block to the new erase block.

8. The storage system of claim 1, wherein the new block size is larger than the previous block size.

9. The storage system of claim 1, wherein the number of planes used for accessing data is increased while the number of multi-plane dies is increased.

10. The storage system of claim 1, wherein the processing device is further configured to: Data is written to the non-volatile memory module, wherein different portions of the data are written to different planes of the multi-plane die, and the number of different portions is equal to the number of planes.

11. The storage system of claim 1, wherein simultaneously accessing data from the multiple planes of the multiplane die reduces the amount of power used by the storage system compared to serially accessing data from multiple planes of the multiplane die.

12. The storage system of claim 1, wherein simultaneously accessing data from the multiple planes of the multiplane die reduces the amount of heat generated by the storage system compared to serially accessing data from multiple planes of the multiplane die.

13. A method comprising: Determine the number of planes used for accessing data simultaneously on a multi-plane die, which is part of one of a plurality of non-volatile memory modules in a storage system; In response to determining that the number of planes used for accessing data on the multi-plane die should be changed, blocks are allocated using a new block size by combining the set of erase blocks at the same address in individual planes based on the newly determined number of planes.

14. The method of claim 13, wherein determining the number of planes used for accessing data while the multi-planar die should be changed comprises: It has been determined that an additional non-volatile memory module has been added to the memory system.

15. The method of claim 13, further comprising: Move one or more portions from an existing erase block to a new erase block, the existing erase block having a new block size.

16. The method of claim 15, further comprising: The metadata identifies one or more portions of the existing erase block, which indicates real-time data within the existing erase block.

17. The method of claim 15, wherein during a garbage collection operation of the storage system, one or more portions are moved from the existing erase block to the new erase block.

18. The method of claim 15, wherein during the defragmentation operation of the storage system, one or more portions are moved from the existing erase block to the new erase block.

19. The method of claim 13, wherein simultaneously accessing data from the multiple planes of the multiplane die reduces the amount of power used by the storage system compared to serially accessing data from multiple planes of the multiplane die.

20. A non-transitory computer-readable storage medium for storing instructions, which, when executed, cause a processing apparatus of a storage system controller to perform the following operations: Determine the number of planes used for accessing data simultaneously on a multi-plane die, which is part of one of a plurality of non-volatile memory modules in a storage system; In response to determining that the number of planes used for accessing data on the multi-plane die should be changed, blocks are allocated using a new block size by combining the set of erase blocks at the same address in individual planes based on the newly determined number of planes.

Citation Information

Patent Citations

  • Data redundancy in a hot pluggable, large symmetric multi-processor system

    US7093158B2

  • Dynamic restriping in nonvolatile memory systems

    US9286002B1