Heterogeneous support for elastic groups
By forming resilient groups in the storage system, limiting the number of blades and organizing write groups, and using NVRAM insertion and garbage collection mechanisms, the problem of increased failure probability during storage system expansion is solved, achieving stable data recovery and system reliability.
Patent Information
- Application Number
- CN202280054641.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2021-07-19
- Filing Date
- 2022-05-25
- Publication Date
- 2026-01-27
- Estimated Expiration
- 2042-05-25
AI Technical Summary
During the expansion of existing storage systems, the probability of failure and the risk of data loss increase exponentially with the addition of blades or storage devices, resulting in insufficient data loss survivability and data recovery capabilities.
By forming elastic groups, limiting the number of blades in each group, and organizing write groups and elastic groups in the storage system, NVRAM insertion is used to avoid write amplification. Combined with garbage collection and authority election mechanisms, system stability and data recovery capabilities are ensured.
It effectively reduces the risk of data loss due to blade failure in multi-chassis storage clusters, maintains the stability of data recovery capabilities and system reliability, and adapts to the expansion needs of storage systems.
Smart Images

Figure CN117795468B_ABST
Abstract
Description
[0001] Cross-reference
[0002] This application claims the rights of U.S. Partial Continuation Patent Application No. 17 / 379,762, filed July 19, 2021, which is hereby incorporated herein by reference. Background Technology
[0003] For example, erase codes in storage systems such as storage arrays and storage clusters are typically configured with N+2 redundancy to withstand the failure of two blades or other storage devices (e.g., storage cells or drives (e.g., solid-state drives, hard disk drives, optical disk drives, etc.)), or other specified redundancy levels. As blades or other devices are added to a storage system, the probability of failure (and consequent data loss) of three or more blades or other storage devices or memory and computing devices increases exponentially. This trend is exacerbated by multi-chassis storage clusters, such as those with 150 blades in 10 chassis. Due to this trend, loss survivability and data recovery do not scale adequately with the expansion of the storage system. Therefore, solutions to overcome the shortcomings described above are needed in the art. Attached Figure Description
[0004] The described embodiments and their advantages are best understood through the following description taken in conjunction with the accompanying drawings. These drawings are in no way intended to limit any changes in form and detail that may be made to the described embodiments by those skilled in the art without departing from the spirit and scope thereof.
[0005] This disclosure is illustrative by way of example rather than limitation, and a more complete understanding can be obtained by referring to the following detailed description when considered in conjunction with the figures described below.
[0006] Figure 1A This describes a first instance system for data storage based on some implementation schemes.
[0007] Figure 1B This describes a second instance system for data storage based on some implementation schemes.
[0008] Figure 1C This describes a third instance system for data storage based on some implementation schemes.
[0009] Figure 1D This describes a fourth instance system for data storage based on some implementation schemes.
[0010] Figure 2A This is a perspective view of a storage cluster according to some embodiments, the storage cluster having multiple storage nodes and internal storage coupled to each storage node to provide network-attached storage.
[0011] Figure 2B This is a block diagram illustrating an interconnect switch that couples multiple storage nodes according to some embodiments.
[0012] Figure 2C This is a multi-level block diagram based on some embodiments, illustrating the contents of a storage node and the contents of one of the non-volatile solid-state storage cells.
[0013] Figure 2D This document illustrates a storage server environment according to some embodiments, which utilizes the storage nodes and storage units shown in Figures 1 to 3.
[0014] Figure 2E These are blade hardware block diagrams based on some embodiments, demonstrating the authority of the control plane, compute and storage plane, and interaction with underlying physical resources.
[0015] Figure 2F Depicts a resilient software layer in the blades of a storage cluster according to some embodiments.
[0016] Figure 2G Describes authority and storage resources in the blades of a storage cluster according to some embodiments.
[0017] Figure 3A The diagram illustrates a storage system according to some embodiments of the present disclosure, the storage system being coupled for data communication with a cloud service provider.
[0018] Figure 3B A diagram illustrating a storage system according to some embodiments of the present disclosure is provided.
[0019] Figure 3C Examples of cloud-based storage systems according to some embodiments of this disclosure are described.
[0020] Figure 3D This describes an exemplary computing device that can be specifically configured to perform one or more processes described herein.
[0021] Figure 3E This describes an instance of a group of storage systems used to provide storage services.
[0022] Figure 4 Describes elastic groups in a storage system that support data recovery in the event of the loss of up to a specified number of blades in the elastic group.
[0023] Figure 5 This is a scenario involving a change in the geometry of the storage system, where blades are added to the storage cluster, resulting in a change from the previous version of the Elastic Group to the current version.
[0024] Figure 6It is a system and action diagram that shows the authority in the distributed computing resources of the storage system and the switches that communicate with the chassis and blades to form flexible groups.
[0025] Figure 7 Describes garbage collection for reclaiming storage memory and relocating data.
[0026] Figure 8 It is a system and action diagram showing the details of garbage collection, including the recovery of coordinated memory, data scanning, and relocation.
[0027] Figure 9 Describes the quorum group used for the startup process of using elastic groups.
[0028] Figure 10 Depict witness groups and authoritative elections using flexible groups.
[0029] Figure 11 It is a system and action diagram that shows the details of the authoritative election and allocation of majority votes by witness groups and involving blades across the storage system.
[0030] Figure 12 This is a flowchart of a method for operating a storage system with elastic groups.
[0031] Figure 13 This is a flowchart of a method for operating a storage system with garbage collection in a resilient group.
[0032] Figure 14 This is a description of exemplary computing devices that can implement the embodiments described herein.
[0033] Figure 15 Depicts a flexible group formed by blades with varying amounts of memory.
[0034] Figure 16 A conservative estimate of the amount of memory space available in the elastic group.
[0035] Figure 17 Describing with Figure 15 and 16 The garbage collection module describes various options related to flexible groups.
[0036] Figure 18 This is a flowchart of a method for operating a storage system to form a resilient group.
[0037] Figure 19 A storage system according to an embodiment is described, the storage system having one or more computing resource elastic groups defined in a computing region and one or more storage resource elastic groups defined in a storage region.
[0038] Figure 20 An embodiment of a storage system is described, which uses resources in an elastic group to form and write data stripes.
[0039] Figure 21 This is a flowchart of a method for a storage system to use resources in an elastic group, according to an embodiment. Detailed Implementation
[0040] The various embodiments of the storage systems described herein form elastic groups, each elastic group having a specified subset of resources, such as the total number of blades in a storage system, like a storage array or storage cluster. In some embodiments (see...), Figures 1A to 3E (and 4 to 18), the entire blade may be a member of a resilient group, and in some embodiments (see 4 to 18), the ...) the blade may be a member of a resilient Figures 1A to 3E(and 19 to 21), a subset of blades and storage devices may be members of an elastic group. A suggested maximum number of blades for an elastic group is 29 blades, or one less than the number of blades required to fill two chassis, but other numbers and arrangements of blades or other storage devices or memory and computing devices may be used. If another blade is added, and any elastic group will have more than the specified maximum number of blades, then the embodiment of the storage cluster modifies the elastic group. Elastic groups are also modified by removing blades or moving blades to different slots, or by adding or removing chassis with one or more blades from the storage system (i.e., changing the cluster geometry that meets specified criteria). In some embodiments, flash writes (for segmentation) and NVRAM (non-volatile random access memory) insertions (i.e., writes to NVRAM) should not cross the boundaries of write groups, and each write group is selected within an elastic group. NVRAM insertions will attempt to select blades within the same chassis to avoid overloading the top-of-rack (ToR) switch and causing write amplification in some embodiments. This organization of write groups and elastic groups results in a stable, non-deteriorating failure probability for three or more blades within a write group or elastic group, even as more blades are added to the storage cluster. When the cluster geometry changes, the authority (in some embodiments of the storage cluster) lists segments that have been divided into two or more new elastic groups, and garbage collection remaps the segments to keep each segment in one of the elastic groups. For the startup process, in some embodiments, a majority group is formed within the elastic group and remains the same as the elastic group in a stable state. In some embodiments, when the elastic group changes, a process is defined for the authority to flush NVRAM and remap segments and transactions and commit records. In some embodiments, a witness group selected as the first elastic group of the cluster is used to define the process for authority election and authority lease renewal. In various embodiments, when blades have different amounts of memory, there are multiple possibilities for forming elastic groups and performing garbage collection. References below. Figures 1A to 3E Various storage systems are described, and embodiments thereof may be found in the references. Figures 4 to 13 And the elastic groups described in 15 to 18 operate as storage systems.
[0041] Figure 1AThis illustration describes an example system for data storage according to some implementation schemes. For illustrative and non-limiting purposes, system 100 (also referred to herein as a “storage system”) includes a number of elements. It can be noted that system 100 may include the same, more, or fewer elements configured in the same or different manner in other implementation schemes. System 100 includes several computing devices 164A to B. The computing devices (also referred to herein as “client devices”) may be embodied as, for example, servers, workstations, personal computers, laptop computers, or the like in a data center. The computing devices 164A to B may be coupled for data communication with one or more storage arrays 102A to B via a storage area network ('SAN') 158 or a local area network ('LAN') 160.
[0042] SAN 158 can be implemented using various data communication architectures, devices, and protocols. For example, the architecture of SAN 158 may include Fibre Channel, Ethernet, unlimited bandwidth, Serial Attached Small Computer System Interface ('SAS'), or similar. Data communication protocols used with SAN 158 may include Advanced Technology Attachment ('ATA'), Fibre Channel protocol, Small Computer System Interface ('SCSI'), Internet Small Computer System Interface ('iSCSI'), HyperSCSI, Architectural Non-Volatile Fast Memory ('NVMe'), or similar. It should be noted that SAN 158 is provided for illustrative purposes and not for limitation. Other data communication couplings may be implemented between computing devices 164A-B and storage arrays 102A-B.
[0043] LAN 160 can also be implemented using various architectures, devices, and protocols. For example, architectures used for LAN 160 may include Ethernet (802.3), wireless (802.11), or similar protocols. Data communication protocols used in LAN 160 may include Transmission Control Protocol ('TCP'), User Datagram Protocol ('UDP'), Internet Protocol ('IP'), Hypertext Transfer Protocol ('HTTP'), Wireless Access Protocol ('WAP'), Handheld Device Transfer Protocol ('HDTP'), Session Initiation Protocol ('SIP'), Real-Time Protocol ('RTP'), or similar protocols.
[0044] Storage arrays 102A to 102B provide persistent data storage for computing devices 164A to 164B. In some embodiments, storage array 102A may be housed in a chassis (not shown), and storage array 102B may be housed in another chassis (not shown). Storage arrays 102A and 102B may include one or more storage array controllers 110A to 110D (also referred to herein as “controllers”). Storage array controllers 110A to 110D may be embodied as modules of an automated computing apparatus comprising computer hardware, computer software, or a combination of computer hardware and software. In some embodiments, storage array controllers 110A to 110D may be configured to perform various storage tasks. Storage tasks may include writing data received from computing devices 164A to B to storage arrays 102A to B, erasing data from storage arrays 102A to B, retrieving data from storage arrays 102A to B and providing the data to computing devices 164A to B, monitoring and reporting disk utilization and performance, performing redundancy operations such as redundant array of independent drives ('RAID') or similar RAID data redundancy operations, compressing data, encrypting data, and so on.
[0045] Storage array controllers 110A to D can be implemented in various ways, including as field-programmable gate arrays ('FPGAs'), programmable logic chips ('PLCs'), application-specific integrated circuits ('ASICs'), systems-on-a-chip ('SOCs'), or any computing device containing discrete components such as processing devices, central processing units, computer memory, or various adapters. Storage array controllers 110A to D may include, for example, data communication adapters configured to support communication via SAN 158 or LAN 160. In some embodiments, storage array controllers 110A to D may be independently coupled to LAN 160. In embodiments, storage array controllers 110A to D may include I / O controllers or the like that couple the storage array controllers 110A to D for data communication to persistent storage resources 170A to B (also referred to herein as "storage resources") via a midplane (not shown). The persistent storage resources 170A to B primarily comprise any number of storage drives 171A to F (also referred to herein as “storage devices”) and any number of non-volatile random access memory ('NVRAM') devices (not shown).
[0046] In some implementations, the NVRAM devices of persistent storage resources 170A to B can be configured to receive data from storage array controllers 110A to D that will be stored in storage drives 171A to F. In some instances, the data may originate from computing devices 164A to B. In some instances, writing data to NVRAM devices can be performed faster than writing data directly to storage drives 171A to F. In implementations, storage array controllers 110A to D can be configured to utilize NVRAM devices as a fast-accessible buffer for data destined to be written to storage drives 171A to F. The latency of write requests using NVRAM devices as buffers can be improved compared to systems in which storage array controllers 110A to D directly write data to storage drives 171A to F. In some implementations, the NVRAM devices can be implemented using computer memory in the form of high-bandwidth, low-latency RAM. NVRAM devices are referred to as "non-volatile" because they can receive or contain a single power source that maintains the state of the RAM after the main power of the NVRAM device is lost. This power source can be a battery, one or more capacitors, or the like. In response to power loss, NVRAM devices can be configured to write the contents of RAM to persistent storage devices, such as storage drives 171A to F.
[0047] In implementations, storage drives 171A to F may refer to any device configured to persistently record data, where "persistently" or "persistently" refers to the device's ability to retain recorded data after power loss. In some implementations, storage drives 171A to F may correspond to non-disk storage media. For example, storage drives 171A to F may be one or more solid-state drives ('SSDs'), flash-based storage devices, any type of solid-state non-volatile memory, or any other type of non-mechanical storage device. In other implementations, storage drives 171A to F may include mechanical or spinning hard disks, such as hard disk drives ('HDDs').
[0048] In some embodiments, memory array controllers 110A to D may be configured to offload device management responsibilities from memory drives 171A to F in memory arrays 102A to B. For example, memory array controllers 110A to D may manage control information that describes the state of one or more memory blocks in memory drives 171A to F. The control information may indicate, for example, that a particular memory block has failed and should no longer be written to, that the particular memory block contains boot code for memory array controllers 110A to D, the number of program erase ('P / E') cycles executed on the particular memory block, the age of data stored in the particular memory block, the type of data stored in the particular memory block, and so on. In some embodiments, the control information may be stored as metadata for the associated memory block. In other embodiments, the control information for memory drives 171A to F may be stored in one or more specific memory blocks of memory drives 171A to F selected by memory array controllers 110A to D. Selected memory blocks may be tagged with identifiers indicating that the selected memory blocks contain control information. The memory blocks containing control information can be quickly identified by the memory array controllers 110A to D in conjunction with memory drives 171A to F using the identifier. For example, the memory controllers 110A to D can issue commands to locate the memory blocks containing control information. It can be noted that the control information can be so large that portions of the control information can be stored in multiple locations, for example, for redundancy purposes, or the control information can be distributed across multiple memory blocks in the memory drives 171A to F in other ways.
[0049] In the implementation, memory array controllers 110A to D can offload device management responsibilities from memory drives 171A to F of memory arrays 102A to B by retrieving control information describing the state of one or more memory blocks in memory drives 171A to F. Retrieving control information from memory drives 171A to F can be performed, for example, by memory array controllers 110A to D querying memory drives 171A to F for the location of control information for a specific memory drive 171A to F. Memory drives 171A to F can be configured to execute instructions that enable memory drives 171A to F to identify the location of the control information. These instructions can be executed by a controller (not shown) associated with or otherwise located on memory drives 171A to F, and can cause memory drives 171A to F to scan a portion of each memory block to identify the memory block storing the control information for memory drives 171A to F. Storage drives 171A to F can respond by sending a response message to storage array controllers 110A to D, the response message containing the location of control information for storage drives 171A to F. In response to receiving the response message, storage array controllers 110A to D can issue a request to read data stored at the address associated with the location of the control information for storage drives 171A to F.
[0050] In other embodiments, the memory array controllers 110A to D can further offload device management responsibilities from memory drives 171A to F by performing memory drive management operations in response to receiving control information. Memory drive management operations may include, for example, operations typically performed by memory drives 171A to F (e.g., a controller (not shown) associated with a particular memory drive 171A to F). These operations may include, for example, ensuring that data is not written to faulty memory blocks within memory drives 171A to F, ensuring that data is written to memory blocks within memory drives 171A to F in a manner that achieves adequate wear leveling, and so on.
[0051] In implementations, storage arrays 102A to B may implement two or more storage array controllers 110A to D. For example, storage array 102A may include storage array controller 110A and storage array controller 110B. In a given example, a single storage array controller 110A to D of storage system 100 (e.g., storage array controller 110A) may be designated as having primary status (also referred to herein as the "primary controller"), and other storage array controllers 110A to D (e.g., storage array controller 110A) may be designated as having secondary status (also referred to herein as the "secondary controller"). The primary controller may have specific rights, such as permission to modify data in persistent storage resources 170A to B (e.g., to write data to persistent storage resources 170A to B). At least some rights of the primary controller may supersede the rights of the secondary controller. For example, when the primary controller has the right to modify data in persistent storage resources 170A to B, the secondary controller may not have permission to modify data in persistent storage resources 170A to B. The status of storage array controllers 110A to D can be changed. For example, storage array controller 110A can be designated as having a secondary status, and storage array controller 110B can be designated as having a primary status.
[0052] In some implementations, for example, a primary controller of storage array controller 110A may serve as the primary controller for one or more storage arrays 102A to B, and a secondary controller of storage array controller 110B may serve as the secondary controller for one or more storage arrays 102A to B. For example, storage array controller 110A may be the primary controller for both storage arrays 102A and 102B, and storage array controller 110B may be the secondary controller for both storage arrays 102A and 102B. In some implementations, storage array controllers 110C and 110D (also referred to as “storage processing modules”) may be neither primary nor secondary. Storage array controllers 110C and 110D, implemented as storage processing modules, may serve as communication interfaces between the primary and secondary controllers (e.g., storage array controllers 110A and 110B, respectively) and storage array 102B. For example, storage array controller 110A of storage array 102A may send write requests to storage array 102B via SAN 158. Write requests can be received by both storage array controllers 110C and 110D of storage array 102B. Storage array controllers 110C and 110D facilitate communication, for example, sending write requests to the appropriate storage drives 171A to F. It can be noted that in some embodiments, the storage processing module can be used to increase the number of storage drives controlled by the primary and secondary controllers.
[0053] In the implementation, memory array controllers 110A to D are communicatively coupled to one or more memory drives 171A to F and one or more NVRAM devices (not shown) included as part of memory arrays 102A to B via a midplane (not shown). Memory array controllers 110A to D may be coupled to the midplane via one or more data communication links, and the midplane may be coupled to the memory drives 171A to F and the NVRAM devices via one or more data communication links. The data communication links described herein are collectively illustrated by data communication links 108A to D and may include, for example, a peripheral component fast interconnect ('PCIe' bus.
[0054] Figure 1B This describes an example system for data storage based on some implementation schemes. Figure 1B The storage array controller 101 described herein may be similar to that described above. Figure 1A The described storage array controllers are 110A to D. In one example, storage array controller 101 may be similar to storage array controller 110A or storage array controller 110B. For illustrative and not limiting purposes, storage array controller 101 includes a number of elements. It should be noted that storage array controller 101 may contain the same, more, or fewer elements configured in the same or different manner in other embodiments. It should be noted that... Figure 1A The components may be included below to help illustrate the features of the storage array controller 101.
[0055] The storage array controller 101 may include one or more processing devices 104 and random access memory ('RAM') 111. The processing device 104 (or controller 101) represents one or more general-purpose processing devices, such as a microprocessor, central processing unit, or the like. More specifically, the processing device 104 (or controller 101) may be a Complex Instruction Set Computing ('CISC') microprocessor, a Reduced Instruction Set Computing ('RISC') microprocessor, a Very Long Instruction Word ('VLIW') microprocessor, or a processor implementing other instruction sets or combinations of instruction sets. The processing device 104 (or controller 101) may also be one or more special-purpose processing devices, such as an ASIC, FPGA, digital signal processor ('DSP'), network processor, or the like.
[0056] Processing device 104 can be connected to RAM 111 via data communication link 106, which may be embodied as a high-speed memory bus, such as a Double Data Rate 4 ('DDR4') bus. Operating system 112 is stored in RAM 111. In some embodiments, instructions 113 are stored in RAM 111. Instructions 113 may include computer program instructions for performing operations in a direct-mapped flash memory system. In one embodiment, a direct-mapped flash memory system is a system that directly addresses data blocks within a flash drive without requiring address translation performed by the flash drive's memory controller.
[0057] In one embodiment, the storage array controller 101 includes one or more host bus adapters 103A to C, which are coupled to the processing device 104 via data communication links 105A to C. In another embodiment, the host bus adapters 103A to C may be computer hardware that connects a host system (e.g., the storage array controller) to other networks and storage arrays. In some instances, the host bus adapters 103A to C may be Fibre Channel adapters enabling the storage array controller 101 to connect to a SAN, Ethernet adapters enabling the storage array controller 101 to connect to a LAN, or the like. The host bus adapters 103A to C may be coupled to the processing device 104 via, for example, a PCIe bus.
[0058] In one implementation, the storage array controller 101 may include a host bus adapter 114 coupled to an extender 115. The extender 115 can be used to attach a host system to a larger number of storage drives. The extender 115 may be, for example, a SAS extender, which enables the host bus adapter 114 to be attached to storage drives in an implementation where the host bus adapter 114 is embodied as a SAS controller.
[0059] In one implementation, the storage array controller 101 may include a switch 116 coupled to the processing device 104 via a data communication link 109. The switch 116 may be a computer hardware device capable of creating multiple endpoints from a single endpoint, thereby enabling multiple devices to share a single endpoint. For example, the switch 116 may be a PCIe switch coupled to a PCIe bus (e.g., data communication link 109) and presenting multiple PCIe connection points to a midplane.
[0060] In one implementation, the storage array controller 101 includes a data communication link 107 for coupling the storage array controller 101 to other storage array controllers. In some instances, the data communication link 107 may be a Fast Path Interconnect (QPI) interconnect.
[0061] Conventional storage systems using conventional flash drives can implement processes across the flash drive itself, which is part of the conventional storage system. For example, higher-level processes of the storage system can span the flash drive boot and control processes. However, the flash drive of a conventional storage system may contain its own storage controller, which also performs the processes. Therefore, for a conventional storage system, both higher-level processes (e.g., booted by the storage system) and lower-level processes (e.g., booted by the storage system's storage controller) can be performed.
[0062] To address various shortcomings of traditional storage systems, operations can be performed by higher-level processes instead of lower-level processes. For example, a flash memory system may include flash drives that do not contain a storage controller providing the processes. Therefore, the operating system of the flash memory system itself can initiate and control the processes. This can be achieved with a directly mapped flash memory system that addresses data blocks within the flash drive directly and without requiring address translation performed by the flash drive's storage controller.
[0063] In some implementations, storage drives 171A through F may be one or more partitioned storage devices. In some implementations, the one or more partitioned storage devices may be shingled HDDs. In some implementations, the one or more storage devices may be flash-based SSDs. In partitioned storage devices, the partition namespace on the partitioned storage device can be addressed by groups of blocks grouped and aligned according to their natural size, thereby forming several addressable regions. In implementations utilizing SSDs, the natural size may be based on the SSD's erase block size. In some implementations, regions of the partitioned storage device can be defined during the initialization of the partitioned storage device. In some implementations, regions can be dynamically defined as data is written to the partitioned storage device.
[0064] In some implementations, regions may be heterogeneous, with some regions being individual page groups and others being multiple page groups. In some implementations, some regions may correspond to erase blocks, and others may correspond to multiple erase blocks. In some implementations, for heterogeneous mixtures of programming models, manufacturers, product types, and / or product generations of storage devices applied to heterogeneous assemblies, upgrades, distributed storage, etc., regions may be any combination of different numbers of pages in page groups and / or erase blocks. In some implementations, regions may be defined as having usage characteristics, such as the ability to support data with a specific kind of lifetime (e.g., very short lifetime or very long lifetime). These characteristics can be used by the partitioned storage device to determine how the regions will be managed within their expected lifetime.
[0065] It should be understood that regions are virtual constructs. Any particular region may not have a fixed location on the storage device. Before allocation, a region may not have any location on the storage device. A region may correspond to a number representing a block of virtual allocatable space, which in various embodiments is the size of an erase block or other block size. When the system allocates or opens a region, the region is allocated to flash memory or other solid-state storage, and when the system writes to a region, pages are written to the mapped flash memory or other solid-state storage of the partitioned storage device. When the system closes a region, the associated erase block or other block size is completed. At some point in the future, the system may delete a region, which will release the allocated space of the region. During its lifetime, a region may be moved to different locations on the partitioned storage device, for example, when the partitioned storage device undergoes internal maintenance.
[0066] In implementations, regions of a partitioned storage device can be in different states. A region may be empty, meaning data is not yet stored in that region. An empty region can be explicitly opened or implicitly opened by writing data to it. This is the initial state of a region on a new partitioned storage device, but it could also be the result of a region reset. In some implementations, an empty region may have a designated location within the flash memory of the partitioned storage device. In some implementations, the location of an empty region can be selected when the region is first opened or first written to (or later if the write is buffered in memory). A region can be implicitly or explicitly open, where an open region can be written to store data using write or append commands. In some implementations, a copy command that copies data from different regions can also be used to write to an open region. In some implementations, the partitioned storage device may have a limit on the number of open regions at a given time.
[0067] A closed region is a region that has been partially written to but entered the closed state after an explicit close operation was issued. Closed regions can be left for future writes, but this reduces some runtime overhead incurred by keeping regions open. In implementations, partitioned storage devices may limit the number of closed regions at a given time. A complete region is a region that is storing data and can no longer be written to. A region may be in a complete state after a write operation has completed the data to the entire region, or as a result of a region completion operation. Before completion, a region may have been fully written to or may not have been fully written to. However, after completion, the region may not be able to be opened for further writing without first performing a region reset operation.
[0068] The mapping from regions to erase blocks (or to shingled tracks in an HDD) can be arbitrary, dynamic, and hidden. Opening a region can allow for the dynamic mapping of a new region to the underlying storage of the partitioned storage device, followed by the writing of data by appending writes to the region until it reaches its capacity. A region can end at any point in time, after which no further data can be written to it. When the data stored in the region is no longer needed, the region can be reset, effectively removing the region's contents from the partitioned storage device, making the physical storage held by the region available for subsequent data storage. Once a region has been written to and completed, the partitioned storage device ensures that the data stored in the region is not lost until the region is reset. During the time between writing data to the region and resetting the region, as part of maintenance operations within the partitioned storage device, the region can be moved between shingled tracks or erase blocks, for example, to keep data refreshed by copying data or to handle memory cell aging in an SSD.
[0069] In HDD-based implementations, resetting a region allows shingled tracks to be allocated to new, open regions that may be opened at a future point in time. In SSD-based implementations, resetting a region can cause the region's associated physical erase blocks to be erased and subsequently reused for data storage. In some implementations, partitioned storage devices may limit the number of open regions at a given time to reduce the amount of space dedicated to keeping regions open.
[0070] The operating system of a flash memory system can identify and maintain a list of allocation units spanning multiple flash drives across the flash memory system. An allocation unit can be an entire erase block or multiple erase blocks. The operating system can maintain a mapping or address range that directly maps addresses to erase blocks in the flash memory system's flash drives.
[0071] Erasable blocks directly mapped to the flash drive can be used to rewrite and erase data. For example, operations can be performed on one or more allocation units containing first and second data, where the first data is to be retained and the second data is no longer used by the flash system. The operating system can initiate a process to write the first data to a new location within other allocation units, erase the second data, and mark the allocation unit as available for subsequent data. Therefore, this process can be performed solely by the higher-level operating system of the flash system, without requiring additional lower-level processes from the flash drive controller.
[0072] The advantage of having the process executed solely by the operating system of the flash memory system includes increased reliability of the flash memory drives, as no unnecessary or redundant write operations are performed during the process. A potentially novel feature of this paper is the concept of initiating and controlling the process on the operating system of the flash memory system. Furthermore, the process can be controlled by the operating system across multiple flash memory drives. This contrasts with processes executed by the storage controller of the flash memory drives.
[0073] The storage system may consist of two storage array controllers sharing a set of drives for failover purposes, or it may consist of a single storage array controller providing storage services using multiple drives, or it may consist of a distributed network of storage array controllers, each storage array controller having a certain number of drives or a certain number of flash memory devices, wherein the storage array controllers in the network cooperate to provide complete storage services and cooperate in all aspects of the storage services, including storage allocation and garbage collection.
[0074] Figure 1C This describes a third instance system 117 for data storage according to some embodiments. For illustrative and not limiting purposes, system 117 (also referred to herein as a “storage system”) includes a number of elements. It may be noted that system 117 may contain the same, more or fewer elements configured in the same or different ways in other embodiments.
[0075] In one embodiment, system 117 includes a dual peripheral component interconnect ('PCI') flash memory device 118 having individually addressable fast write storage. System 117 may include a memory device controller 119. In one embodiment, memory device controllers 119A to D may be a CPU, ASIC, FPGA, or any other circuit system capable of implementing the control architecture required according to this disclosure. In one embodiment, system 117 includes flash memory devices (e.g., flash memory devices 120a to n) operatively coupled to various channels of memory device controller 119. Flash memory devices 120a to n may be presented to controllers 119A to D as an addressable set of flash memory pages, erase blocks, and / or control elements sufficient to allow memory device controllers 119A to D to program and retrieve various aspects of the flash memory. In one embodiment, the storage device controllers 119A to D can perform operations on the flash memory devices 120a to n, including storing and retrieving data content of pages, arranging and erasing any blocks, tracking statistics related to the use and reuse of flash memory pages, erase blocks and cells, tracking and predicting error codes and faults within the flash memory, and controlling voltage levels associated with programming and retrieving the contents of flash memory cells, etc.
[0076] In one embodiment, system 117 may include RAM 121 for separately storing addressable, fast-write data. In one embodiment, RAM 121 may be one or more separate discrete devices. In another embodiment, RAM 121 may be integrated into memory device controllers 119A to D or more memory device controllers. RAM 121 may also be used for other purposes, such as temporary program memory for processing devices (e.g., CPU) within memory device controller 119.
[0077] In one embodiment, system 117 may include an energy storage device 122, such as a rechargeable battery or capacitor. The energy storage device 122 may store sufficient energy to power the storage device controller 119, a certain amount of RAM (e.g., RAM 121), and a certain amount of flash memory (e.g., flash memory 120a to 120n) to allow sufficient time for the contents of the RAM to be written to the flash memory. In one embodiment, if the storage device controller detects a loss of external power, then the storage device controllers 119A to D may write the contents of the RAM to the flash memory.
[0078] In one embodiment, system 117 includes two data communication links 123a and 123b. In one embodiment, data communication links 123a and 123b may be PCI interfaces. In another embodiment, data communication links 123a and 123b may be based on other communication standards (e.g., HyperTransport, Infinite Bandwidth, etc.). Data communication links 123a and 123b may be based on Non-Volatile Fast Memory ('NVMe') or Architectural NVMe ('NVMf') specifications, which allow external connection from other components in storage system 117 to storage device controllers 119A through D. It should be noted that, for convenience, data communication links are interchangeably referred to herein as PCI buses.
[0079] System 117 may also include an external power supply (not shown), which may be provided via one or two data communication links 123a, 123b, or may be provided independently. Alternative embodiments include a separate flash memory (not shown) dedicated to storing the contents of RAM 121. Storage device controllers 119A to D may present logic devices on a PCI bus, which may include addressable fast-write logic devices, or different portions of the logical address space of storage device 118, which may present as PCI memory or persistent storage. In one embodiment, operations stored in the device are directed to RAM 121. In the event of a power failure, storage device controllers 119A to D may write the storage contents associated with the addressable fast-write logic memory to flash memory (e.g., flash memory 120a to n) for long-term persistent storage.
[0080] In one embodiment, the logic device may include some or all of the contents of flash memory devices 120a to n, wherein the presentation allows a memory system (e.g., memory system 117) including memory device 118 to directly address flash memory pages and directly reprogram erase blocks from memory system components outside the memory devices via the PCI bus. The presentation may also allow one or more external components to control and retrieve other aspects of the flash memory, including some or all of the following: tracking statistics related to the use and reuse of flash memory pages, erase blocks, and cells across all flash memory devices; tracking and predicting error codes and faults within and across flash memory devices; controlling voltage levels associated with programming and retrieving the contents of flash memory cells; and so on.
[0081] In one embodiment, energy storage device 122 may be sufficient to ensure the completion of ongoing operations on flash memory devices 120a to 120n. Energy storage device 122 may power storage device controllers 119A to D and associated flash memory devices (e.g., 120a to n) for said operations and for storing fast writes from RAM to flash memory. Energy storage device 122 may be used to store accumulated statistics and other parameters saved and tracked by flash memory devices 120a to n and / or storage device controller 119. Individual capacitors or energy storage devices (e.g., smaller capacitors located near or embedded within the flash memory devices themselves) may be used for some or all of the operations described herein.
[0082] Various methods can be used to track and optimize the lifespan of energy storage components, such as adjusting voltage levels over time, partially discharging energy storage device 122 to measure corresponding discharge characteristics, etc. If available energy decreases over time, the effective available capacity of the addressable fast write memory device can be reduced to ensure that it can be safely written to based on currently available energy.
[0083] Figure 1D This describes a third instance storage system 124 for data storage according to some embodiments. In one embodiment, storage system 124 includes storage controllers 125a, 125b. In one embodiment, storage controllers 125a, 125b are operatively coupled to a dual PCI storage device. Storage controllers 125a, 125b are operatively coupled (e.g., via storage network 130) to a number of host computers 127a to n.
[0084] In one embodiment, two storage controllers (e.g., 125a and 125b) provide storage services, such as an SCS block storage array, file server, object server, database, or data analytics services. Storage controllers 125a and 125b can provide services to host computers 127a to n outside the storage system 124 via a number of network interfaces (e.g., 126a to d). Storage controllers 125a and 125b can provide integrated services or applications entirely within the storage system 124, forming a converged storage and computing system. Storage controllers 125a and 125b can utilize fast write memory within or across storage devices 119a to d to record ongoing operations to ensure that operations are not lost in the event of a power failure, storage controller removal, storage controller or storage system shutdown, or failure of one or more software or hardware components within the storage system 124.
[0085] In one embodiment, storage controllers 125a and 125b operate as PCI masters of one or the other PCI bus 128a and 128b. In another embodiment, 128a and 128b may be based on other communication standards (e.g., HyperTransport, Infinite Bandwidth, etc.). Other storage system embodiments may operate storage controllers 125a and 125b as multiple masters of two PCI buses 128a and 128b. Alternatively, a PCI / NVMe / NVMe switching infrastructure or structure may connect multiple storage controllers. Some storage system embodiments may allow storage devices to communicate directly with each other, rather than only with storage controllers. In one embodiment, storage device controller 119a may operate under the guidance of storage controller 125a to access data stored in RAM (e.g., ...). Figure 1C The data in RAM 121 is synthesized and transferred to the flash memory device. For example, after the memory controller has determined that the operation has been fully committed across the memory system, or when the fast write memory on the device has reached a certain used capacity, or after a certain period of time, a recalculated version of the RAM contents may be transferred to ensure improved data security or to free up addressable fast write capacity for reuse. For example, this mechanism can be used to avoid a second transfer from memory controllers 125a, 125b via a bus (e.g., 128a, 128b). In one embodiment, recalculation may include compressing data, attaching indexes or other metadata, combining multiple data segments together, performing erase code calculations, etc.
[0086] In one embodiment, under the guidance of storage controllers 125a, 125b, storage device controllers 119a, 119b are operable to retrieve data from RAM (e.g., storage devices stored in RAM). Figure 1CThe data in RAM 121) is processed and transferred to other storage devices without involving storage controllers 125a and 125b. This operation can be used to mirror data stored in one storage controller 125a to another storage controller 125b, or it can be used to offload compression, data aggregation, and / or erase encoding calculations and transfers to storage devices to reduce the load on the storage controllers or the storage controller interfaces 129a and 129b to the PCI buses 128a and 128b.
[0087] Storage device controllers 119A to D may include mechanisms for implementing high availability primitives for use by other parts of the storage system outside the dual PCI storage device 118. For example, reserve or exclusion primitives may be provided, allowing one storage controller in a storage system with two storage controllers providing highly available storage services to prevent the other storage controller from accessing or continuing to access the storage device. This can be used, for example, if one controller detects that the other controller is not functioning correctly or if the interconnect between the two storage controllers itself may be malfunctioning.
[0088] In one embodiment, a storage system for use with dual PCI direct-mapped storage devices having individually addressable fast-write storage includes a system for managing erase blocks or erase block groups as allocation units for storing data on behalf of the storage service, or for storing metadata associated with the storage service (e.g., indexes, logs, etc.), or for properly managing the storage system itself. Flash pages, which can be several gigabytes in size, can be written when data arrives or when the storage system intends to retain the data for a long period (e.g., exceeding a defined time threshold). To commit data faster or to reduce the number of writes to the flash memory device, the storage controller may first write the data to an individually addressable fast-write storage device on another storage device.
[0089] In one embodiment, storage controllers 125a and 125b may initiate the use of erase blocks within and across storage devices (e.g., 118) based on the age and expected remaining lifetime of the storage devices or based on other statistical data. Storage controllers 125a and 125b may also initiate garbage collection and data migration between storage devices based on pages that are no longer needed, manage flash memory page and erase block lifetimes, and manage overall system performance.
[0090] In one embodiment, the storage system 124 may utilize mirroring and / or erasure coding schemes as part of storing data into an addressable, fast-write storage device and / or as part of writing data into an allocation unit associated with an erase block. Erasure codes may be used across storage devices, and within erase blocks or allocation units, or within and across flash memory devices on a single storage device, to provide redundancy against single or multiple storage device failures, or to prevent internal damage to flash pages caused by flash memory operations or flash cell degradation. Various levels of mirroring and erasure coding can be used for recovery from multiple types of failures occurring individually or in combination.
[0091] refer to Figure 2A The embodiments described to G illustrate a storage cluster that stores user data (e.g., user data originating from one or more user or client systems or other sources outside the storage cluster). The storage cluster uses erasure coding and redundant copies of metadata across storage nodes housed within a chassis or across multiple chassis to distribute user data. Erasure coding refers to a data protection or reconstruction method where data is stored across a set of different locations (e.g., disks, storage nodes, or geographic locations). Flash memory is a type of solid-state storage that can be integrated with the embodiments, but the embodiments can be extended to other types of solid-state storage or other storage media, including non-solid-state storage. Control over storage location and workload spans the distribution of storage locations in a clustered peer-to-peer system. For example, tasks such as mediating communication between various storage nodes, detecting when storage nodes become unavailable, and balancing I / O (input and output) across various storage nodes are all handled on a distributed basis. In some embodiments, data is arranged or distributed across multiple storage nodes in the form of data segments or stripes that support data recovery. Data ownership can be reassigned within the cluster, independent of input and output patterns. This architecture, described in more detail below, allows the system to remain operational even if a storage node in the cluster fails, because the data can be reconstructed from other storage nodes and thus remains available for input and output operations. In various embodiments, the storage node may be referred to as a cluster node, blade, or server.
[0092] Storage clusters can be housed within a chassis (i.e., a enclosure that houses one or more storage nodes). The chassis contains mechanisms for supplying power to each storage node (e.g., a power distribution bus) and communication mechanisms (e.g., a communication bus) enabling communication between storage nodes. According to some embodiments, the storage cluster can operate as a standalone system in one location. In one embodiment, the chassis contains at least two examples of both power distribution and communication buses that can be independently enabled or disabled. The internal communication bus can be an Ethernet bus; however, other technologies (e.g., PCIe, wireless bandwidth, and others) are also applicable. The chassis provides ports for external communication buses to enable communication between multiple chassis and with client systems, either directly or via a switch. External communication can use technologies such as Ethernet, wireless bandwidth, Fibre Channel, etc. In some embodiments, the external communication bus uses different communication bus technologies for inter-chassis and client communication. If a switch is deployed within or between chassis, the switch can act as a converter between various protocols or technologies. When multiple chassis are connected to define a storage cluster, the storage cluster can be accessed by clients using proprietary or standard interfaces, such as Network File System ('NFS'), Common Internet File System ('CIFS'), Small Computer System Interface ('SCSI'), or Hypertext Transfer Protocol ('HTTP')). The conversion from client protocols can occur at a switch, at the chassis external communication bus, or within each storage node. In some embodiments, multiple chassis can be coupled or connected to each other via an aggregator switch. A portion and / or all of the coupled or connected chassis can be designated as a storage cluster. As discussed above, each chassis may have multiple blades, each blade having a Media Access Control ('MAC') address; however, in some embodiments, the storage cluster presents itself to the external network as having a single cluster IP address and a single MAC address.
[0093] Each storage node may be one or more storage servers, and each storage server is connected to one or more non-volatile solid-state storage cells, which may be referred to as storage cells or storage devices. One embodiment includes a single storage server and between one and eight non-volatile solid-state storage cells in each storage node; however, this example is not intended to be limiting. The storage server may include a processor, DRAM, and interfaces for power distribution to an internal communication bus and each power bus. In some embodiments, within the storage node, the interfaces and storage cells share a communication bus, such as PCI Express. The non-volatile solid-state storage cells can directly access the internal communication bus interface via the storage node's communication bus, or request the storage node to access the bus interface. The non-volatile solid-state storage cell contains an embedded CPU, a solid-state storage controller, and a certain amount of solid-state high-capacity storage, for example, between 2 and 32 terabytes ('TB') in some embodiments. The non-volatile solid-state storage cell includes embedded volatile storage media, such as DRAM, and an energy storage device. In some embodiments, the energy storage device is a capacitor, supercapacitor, or battery, which enables the transfer of a subset of the DRAM content to a stable storage medium in the event of power loss. In some embodiments, the non-volatile solid-state memory cell is composed of a storage-type memory, such as a phase-change or magnetoresistive random access memory ('MRAM') that replaces DRAM and implements a device for reducing power hold-up.
[0094] One of the many characteristics of storage nodes and non-volatile solid-state storage devices (NSSSDs) is the ability to proactively reconstruct data within a storage cluster. Storage nodes and NSSSDs can determine when a storage node or NSSSD in a storage cluster is unreachable, regardless of whether an attempt is made to read data involving that storage node or NSSSD. The storage nodes and NSSSDs then cooperate to recover and reconstruct the data in at least a partial new location. This constitutes proactive reconstruction because the system does not need to wait for a read access initiated from a client system employing the storage cluster to require the data before reconstructing it. These and further details of memory and its operation are discussed below.
[0095] Figure 2AThis is a perspective view of a storage cluster 161 according to some embodiments, the storage cluster 161 having a plurality of storage nodes 150 and internal solid-state memory coupled to each storage node to provide network-attached storage or a storage area network. A network-attached storage device, storage area network, or storage cluster or other memory may comprise one or more storage clusters 161, each storage cluster 161 having one or more storage nodes 150, arranged in a flexible and reconfigurable manner for both the physical components and the amount of memory provided. The storage cluster 161 is designed to be mounted in racks and one or more racks can be set up and filled as needed for the memory. The storage cluster 161 has a chassis 138 with a plurality of slots 142. It should be understood that the chassis 138 may be referred to as a shell, enclosure, or rack unit. In one embodiment, the chassis 138 has 14 slots 142, but other numbers of slots can be easily designed. For example, some embodiments have four slots, eight slots, sixteen slots, thirty-two slots, or other suitable numbers of slots. In some embodiments, each slot 142 may accommodate one storage node 150. The chassis 138 includes a baffle 148 for mounting the chassis 138 on a rack. A fan 144 provides air circulation for cooling the storage nodes 150 and their components, but other cooling components may be used, or embodiments may be designed without cooling components. A switch structure 146 couples the storage nodes 150 within the chassis 138 together and to a network for communication with the memory. In the embodiment depicted herein, slot 142 to the left of the switch structure 146 and fan 144 is shown occupied by a storage node 150, while slot 142 to the right of the switch structure 146 and fan 144 is empty and available for insertion of a storage node 150 for illustrative purposes. This configuration is an example, and one or more storage nodes 150 may occupy slot 142 in various other arrangements. In some embodiments, the storage node arrangement need not be sequential or adjacent. The storage nodes 150 are hot-swappable, meaning that a storage node 150 can be inserted into or removed from slot 142 in the chassis 138 without stopping or shutting down the system. After a storage node 150 is inserted into or removed from slot 142, the system automatically reconfigures to recognize and adapt to the change. In some embodiments, reconfiguration includes restoring redundancy and / or rebalancing data or load.
[0096] Each storage node 150 may have multiple components. In the embodiment shown herein, storage node 150 includes a printed circuit board 159 filled with a CPU 156 (i.e., a processor), a memory 154 coupled to the CPU 156, and a non-volatile solid-state storage device 152 coupled to the CPU 156; however, other mounting and / or components may be used in other embodiments. The memory 154 has instructions executed by the CPU 156 and / or data operated by the CPU 156. As further explained below, the non-volatile solid-state storage device 152 includes flash memory, or in other embodiments, other types of solid-state memory.
[0097] refer to Figure 2A Storage cluster 161 is scalable, meaning that storage capacity with uneven storage sizes can be easily added, as described above. In some embodiments, one or more storage nodes 150 can be inserted into or removed from each chassis, and the storage cluster is self-configurable. Insertable storage nodes 150, whether installed in the chassis at delivery or added later, can have different sizes. For example, in one embodiment, storage nodes 150 can have any multiple of 4TB, such as 8TB, 12TB, 16TB, 32TB, etc. In other embodiments, storage nodes 150 can have any multiple of other storage amounts or capacities. The storage capacity of each storage node 150 is broadcast and influences decisions on how data is striped. For maximum storage efficiency, embodiments can be self-configurable as wide as possible within stripes, complying with predetermined requirements to continue operation in the event of the loss of up to one or two non-volatile solid-state storage device 152 units or storage nodes 150 within the chassis.
[0098] Figure 2B This is a block diagram showing the communication interconnect 173 and power distribution bus 172 that couple multiple storage nodes 150. (Return to Reference) Figure 2A In some embodiments, the communication interconnect 173 may be included in or implemented using the switch structure 146. In some embodiments, where multiple storage clusters 161 occupy a rack, the communication interconnect 173 may be included in or implemented using a top-of-rack switch. Figure 2B The description states that storage cluster 161 is enclosed within a single chassis 138. External port 176 is coupled to storage node 150 via communication interconnect 173, while external port 174 is directly coupled to the storage node. External power port 178 is coupled to power distribution bus 172. Storage node 150 may contain varying amounts and capacities of non-volatile solid-state storage devices 152, as described in reference [reference missing]. Figure 2A Description. Additionally, as... Figure 2BAs explained, one or more storage nodes 150 may be computation-only storage nodes. Authority 168 is implemented on non-volatile solid-state storage device 152, for example, as a list or other data structure stored in memory. In some embodiments, authority is stored within non-volatile solid-state storage device 152 and supported by software executing on the controller or other processor of non-volatile solid-state storage device 152. In other embodiments, authority 168 is implemented on storage node 150, for example, as a list or other data structure stored in memory 154 and supported by software executing on the CPU 156 of storage node 150. In some embodiments, authority 168 controls how and where data is stored in non-volatile solid-state storage device 152. This control helps determine which type of erasure coding scheme is applied to the data and which storage nodes 150 have which portions of the data. Each authority 168 may be assigned to non-volatile solid-state storage device 152. In various embodiments, each authority can control a series of inode numbers, segment numbers, or other data identifiers assigned to data by the file system, storage node 150, or non-volatile solid-state storage device 152.
[0099] In some embodiments, each piece of data and each piece of metadata is redundant in the system. Additionally, each piece of data and each piece of metadata has an owner, which may be referred to as an authority. If the authority becomes unreachable, for example due to a storage node failure, there is a successor plan for how to find the data or the metadata. In various embodiments, redundant copies of authority 168 exist. In some embodiments, authority 168 is associated with storage node 150 and non-volatile solid-state storage device 152. Each authority 168, covering a range of data segment numbers or other identifiers of the data, may be assigned to a specific non-volatile solid-state storage device 152. In some embodiments, all these ranges of authorities 168 are distributed across the non-volatile solid-state storage devices 152 of the storage cluster. Each storage node 150 has a network port providing access to the non-volatile solid-state storage devices 152 of that storage node 150. In some embodiments, data may be stored in segments associated with segment numbers, and the segment numbers are indirections in the configuration of RAID (Redundant Array of Independent Disks) stripes. Therefore, the assignment and use of authority 168 establishes indirections to the data. According to some embodiments, indirection may be referred to as the ability to indirectly (in this case via authority 168) reference data. A segment identifies a set of non-volatile solid-state storage devices 152 and a local identifier within said set of non-volatile solid-state storage devices 152 that may contain data. In some embodiments, the local identifier is an offset within the device and can be reused sequentially by multiple segments. In other embodiments, the local identifier is unique for a particular segment and is never reused. Offsets within the non-volatile solid-state storage devices 152 are used to locate data (in the form of RAID stripes) to be written to or read from the non-volatile solid-state storage devices 152. Data is striped across multiple cells of the non-volatile solid-state storage devices 152, which may contain or differ from non-volatile solid-state storage devices 152 having an authority 168 for a particular data segment.
[0100] If the location of a specific data segment changes, for example during data movement or data reconstruction, then the authority 168 of the data segment should be consulted at the non-volatile solid-state storage device 152 or storage node 150 that has the authority 168. To locate a specific piece of data, embodiments calculate a hash value of the data segment or apply an inode number or data segment number. The output of this operation points to the non-volatile solid-state storage device 152 that has the authority 168 for the specific piece of data. In some embodiments, this operation has two phases. The first phase maps an entity identifier (ID) (e.g., segment number, inode number, or directory number) to an authority identifier. This mapping may include, for example, the calculation of a hash or bitmask. The second phase is mapping the authority identifier to a specific non-volatile solid-state storage device 152, which can be done through explicit mapping. The operation is repeatable such that when the calculation is performed, the result of the calculation reliably and repeatedly points to the specific non-volatile solid-state storage device 152 that has the authority 168. The operation may include a group of reachable storage nodes as input. If the set of reachable non-volatile solid-state storage units changes, then the optimal set also changes. In some embodiments, the stored value is the current assignment (always true), and the calculated value is the target assignment that the cluster will attempt to reconfigure towards. This calculation can be used to determine the optimal non-volatile solid-state storage device 152 for an authority when there is a set of reachable non-volatile solid-state storage devices 152 that constitute the same cluster. The calculation also determines an ordered set of peer non-volatile solid-state storage devices 152, which will also record the authority-to-non-volatile solid-state storage mapping, so that the authority can be determined even if the assigned non-volatile solid-state storage device is unreachable. If a particular authority 168 is unavailable in some embodiments, then a replica or alternative authority 168 can be consulted.
[0101] refer to Figure 2A and 2BTwo of the many tasks performed by the CPU 156 on storage node 150 are breaking down written data and reassembling read data. When the system determines that data is to be written, the authority 168 for the data is located as described above. When the segment ID of the data is determined, the write request is forwarded to the non-volatile solid-state storage device 152 of the host that is currently identified as the authority 168 from which the segment was determined. The non-volatile solid-state storage device 152 and the host CPU 156 of the storage node 150 on which the corresponding authority 168 resides then break down or fragment the data and transmit the data out to the various non-volatile solid-state storage devices 152. The transmitted data is written as data stripes according to the erasure coding scheme. In some embodiments, data is requested to be retrieved, and in other embodiments, data is pushed. Conversely, when reading data, the authority 168 containing the segment ID of the data is located as described above. The host CPU 156 of the storage node 150, on which the non-volatile solid-state storage device 152 and its corresponding authority 168 reside, requests data from the non-volatile solid-state storage device and its corresponding storage node as directed by the authority. In some embodiments, the data is read from flash memory as data stripes. The host CPU 156 of the storage node 150 then reassembles the read data, corrects any errors (if any) according to an appropriate erasure coding scheme, and forwards the reassembled data to the network. In other embodiments, some or all of these tasks may be handled within the non-volatile solid-state storage device 152. In some embodiments, a segment host requests data to be sent to storage node 150 by requesting a page from the storage device and then sending the data to the storage node that made the original request.
[0102] In this embodiment, authority 168 operates to determine how operations will be performed on specific logic elements. Each logic element can be operated by a specific authority across multiple storage controllers in the storage system. Authority 168 can communicate with multiple storage controllers, enabling the multiple storage controllers to collectively perform operations on those specific logic elements.
[0103] In embodiments, a logical element may be, for example, a file, directory, object bucket, individual object, a delimited portion of a file or object, or other forms of key-value database or table. In embodiments, performing operations may involve, for example, ensuring the consistency, structural integrity, and / or recoverability of other operations targeting the same logical element, reading metadata and data associated with the logical element, determining which data should be persistently written to the storage system to preserve any changes to the operation, or determining where the metadata and data are stored across modular storage devices attached to multiple storage controllers in the storage system.
[0104] In some embodiments, operations are token-based transactions for efficient communication within a distributed system. Each transaction may be accompanied by or associated with a token that grants permission to execute the transaction. In some embodiments, authority 168 is able to maintain the system's pre-transaction state until the operation is complete. Token-based communication can be completed without crossing global locks across the system and can also resume operations in the event of interruption or other failures.
[0105] In some systems, such as UNIX-style file systems, data is handled using index nodes (or inodes), which specify the data structure representing objects within the file system. For example, an object can be a file or a directory. Metadata may accompany the object as attributes (such as permission data, creation timestamps, and other attributes). Segment numbers can be assigned to all or part of such objects within the file system. In other systems, data segments are handled using segment numbers assigned elsewhere. For the purposes of discussion, the unit of allocation is an entity, and an entity can be a file, directory, or segment. That is, an entity is a unit of data or metadata stored by the storage system. Entities are grouped into groups called authorities. Each authority has an authority owner, who is a storage node with exclusive rights to update the entities within that authority. In other words, a storage node contains authorities, and those authorities in turn contain entities.
[0106] According to some embodiments, a segment is a logical container for data. A segment is an address space between media address spaces, and the physical flash memory location (i.e., the data segment number) is within this address space. The segment may also contain metadata that enables data redundancy to be recovered (rewritten to a different flash memory location or device) without involving higher-level software. In one embodiment, the internal format of a segment contains client data and a media mapping to determine the location of the data. Where applicable, each data segment is protected from memory and other failures, for example, by dividing the segment into several data and parity fragments. Depending on the erasure coding scheme, the data and parity fragments are coupled across the host CPU 156 (see...). Figure 2E and 2G The non-volatile solid-state storage device 152 is distributed, i.e., striped. In some embodiments, the term "segment" refers to a container and its location in the address space of a segment. According to some embodiments, the term "strip" refers to fragments within the same group of segments, and includes how the fragments are distributed along with redundancy or parity information.
[0107] A series of address space translations occur across the entire storage system. At the top are directory entries (filenames) linked to inodes. Inodes point to the media address space where data is logically stored. Media addresses can be mapped through a series of indirect media to distribute the load of large files or to implement data services such as deduplication or snapshots. Next, segment addresses are translated into physical flash memory locations. According to some embodiments, physical flash memory locations have address ranges delimited by the amount of flash memory in the system. Media addresses and segment addresses are logical containers and, in some embodiments, use 128-bit or larger identifiers so that they are virtually unlimited, where the possibility of reuse is calculated to be longer than the expected lifespan of the system. In some embodiments, addresses from logical containers are allocated hierarchically. Initially, each non-volatile solid-state storage device 152 cell can be assigned a range of address space. Within this assigned range, the non-volatile solid-state storage device 152 can allocate addresses without synchronization with other non-volatile solid-state storage devices 152.
[0108] Data and metadata are stored using a set of underlying storage layouts optimized for different workload patterns and storage devices. These layouts incorporate various redundancy schemes, compression formats, and indexing algorithms. Some of these layouts store information about the authority and the authority master, while others store file metadata and file data. Redundancy schemes include error correction codes that allow for damaged bits within a single storage device (e.g., NAND flash memory chips), erase codes that allow for the failure of multiple storage nodes, and replication schemes that allow for data center or region failures. In some embodiments, low-density parity-check ('LDPC') codes are used within a single storage cell. In some embodiments, Reed-Solomon encoding is used within the storage cluster, and mirroring is used within the storage grid. Ordered log-structured indexes (e.g., log-structured merge trees) can be used to store metadata, and large amounts of data may not be stored in log-structured layouts.
[0109] To maintain consistency across multiple replicas of an entity, storage nodes implicitly agree on two things: (1) the authority containing the entity, and (2) the storage node containing the authority. Assigning an entity to an authority can be done by pseudo-randomly assigning the entity, by partitioning the entity into ranges based on externally generated keys, or by placing a single entity into each authority. Examples of pseudo-random schemes are hashes of the 'RUSH' family under linear hashing and scalable hashing, including 'CRUSH' under scalable hashing. In some embodiments, pseudo-random assignment is used only to assign authorities to nodes because the node group can change. The authority group cannot change, so any subjective functionality can be applied in these embodiments. Some placement schemes automatically place authorities on storage nodes, while others rely on an explicit mapping from authorities to storage nodes. In some embodiments, a pseudo-random scheme is used to map from each authority to a set of candidate authority owners. The pseudo-random data allocation function associated with CRUSH assigns authorities to storage nodes and creates a list of where authorities are assigned. Each storage node has a copy of the pseudo-random data allocation function, and the same computations used for allocating and subsequently finding or locating authorities are derived. In some embodiments, each pseudo-random scheme requires a set of reachable storage nodes as input to infer the same target node. Once an entity has been placed in an authority, it can be stored on physical devices so that anticipated failures do not result in unexpected data loss. In some embodiments, the rebalancing algorithm attempts to store copies of all entities within an authority in the same layout and on the same set of machines.
[0110] Examples of anticipated failures include device failure, machine theft, data center fires, and regional disasters such as nuclear or geological events. Different failures result in different levels of acceptable data loss. In some embodiments, theft of storage nodes does not affect the security or reliability of the system; depending on the system configuration, regional events can result in no data loss, lost updates for seconds or minutes, or even complete data loss.
[0111] In some embodiments, the placement of redundant data is independent of the placement of authorities for data consistency. In some embodiments, the storage node containing the authority does not contain any persistent storage. Instead, the storage node is connected to a non-volatile solid-state storage cell that does not contain an authority. The communication interconnect between the storage node and the non-volatile solid-state storage cell consists of various communication technologies and exhibits non-uniform performance and fault tolerance. In some embodiments, as mentioned above, the non-volatile solid-state storage cell is connected to the storage node via PCI express, the storage nodes are connected together in a single chassis using an Ethernet backplane, and the chassis are connected together to form a storage cluster. In some embodiments, the storage cluster is connected to the client using Ethernet or Fibre Channel. If multiple storage clusters are configured in a storage grid, the Internet or other long-distance network links (e.g., “metro-scale” links or dedicated links that do not traverse the Internet) are used to connect the multiple storage clusters.
[0112] The authoritative owner has exclusive rights to modify entities, migrate entities from one non-volatile solid-state storage unit to another, and add and delete copies of entities. This allows for the maintenance of redundancy in the underlying data. When the authoritative owner fails, is about to retire, or is overloaded, authority is transferred to a new storage node. Transient failures make it crucial to ensure that all fault-free machines agree on the new authoritative position. The ambiguity arising from transient failures can be automated through consensus protocols (e.g., Paxos, hot-cold failover schemes) or through manual intervention by a remote system administrator or local hardware administrator (e.g., by physically removing the failed machine from the cluster or pressing a button on the failed machine). In some embodiments, a consensus protocol is used, and failover is automatic. According to some embodiments, if too many failures or replication events occur within a short period, the system enters a self-preservation mode and stops replication and data movement activities until administrator intervention.
[0113] The system transmits messages between storage nodes and non-volatile solid-state storage units as authority is transferred between storage nodes and authority owners update entities within their authority. For persistent messages, messages with different purposes are classified as different types. Depending on the message type, the system maintains different ordering and persistence guarantees. When persistent messages are processed, they are temporarily stored across multiple persistent and non-persistent storage hardware technologies. In some embodiments, messages are stored in RAM, NVRAM, and NAND flash memory devices, using various protocols to efficiently utilize each storage medium. Latency-sensitive client requests may be stored in replicated NVRAM and subsequently in NAND, while background rebalancing operations are performed directly on NAND.
[0114] Persistent messages are persistently stored before transmission. This allows the system to continue serving client requests even in the event of failures and component replacements. While many hardware components contain unique identifiers visible to system administrators, manufacturers, the hardware supply chain, and persistent monitoring and quality control infrastructure, applications running on top of these infrastructure addresses virtualize those addresses. These virtualized addresses do not change throughout the storage system's lifetime, regardless of component failures or replacements. This allows each component of the storage system to be replaced over time without reconfiguration or interruption of client request processing; that is, the system supports non-disruptive upgrades.
[0115] In some embodiments, virtualized addresses are stored with sufficient redundancy. The continuous monitoring system correlates hardware and software status with hardware identifiers. This allows for the detection and prediction of failures due to faulty components and manufacturing details. In some embodiments, the monitoring system also proactively transfers authority and entities away from the affected device before a failure occurs by removing components from the critical path.
[0116] Figure 2C This is a multi-level block diagram illustrating the contents of storage node 150 and the contents of its non-volatile solid-state storage devices 152. In some embodiments, data is transferred to and from storage node 150 by a network interface controller (NIC) 202. Each storage node 150 has a CPU 156 and one or more non-volatile solid-state storage devices 152, as discussed above. Figure 2C Moving down one level, each non-volatile solid-state storage device 152 has a relatively fast non-volatile solid-state memory, such as non-volatile random access memory ('NVRAM') 204 and flash memory 206. In some embodiments, NVRAM 204 may be a component that does not require programming / erasing cycles (DRAM, MRAM, PCM) and may be a memory that can support being written to more frequently than the memory is read. Figure 2CMoving down another level, NVRAM 204 is implemented in one embodiment as a high-speed volatile memory, such as dynamic random access memory (DRAM) 216 supported by an energy reserve 218. The energy reserve 218 provides sufficient power to keep the DRAM 216 powered for a sufficient period to transfer content to the flash memory 206 in the event of a power failure. In some embodiments, the energy reserve 218 is a capacitor, supercapacitor, battery, or other device that provides a suitable energy supply sufficient to transfer the content of the DRAM 216 to a stable storage medium in the event of a power loss. The flash memory 206 is implemented as a plurality of flash memory dies 222, which may be referred to as a package of flash memory dies 222 or an array of flash memory dies 222. It should be understood that the flash memory dies 222 can be packaged in any number of ways, such as one die per package, multiple dies per package (i.e., multi-chip package), hybrid package, as dies on a printed circuit board or other substrate, as encapsulated dies, etc. In the illustrated embodiment, the non-volatile solid-state storage device 152 has a controller 212 or other processor and an input / output (I / O) port 210 coupled to the controller 212. The I / O port 210 is coupled to the CPU 156 and / or network interface controller 202 of the flash memory node 150. A flash memory input / output (I / O) port 220 is coupled to a flash memory die 222, and a direct memory access unit (DMA) 214 is coupled to the controller 212, DRAM 216, and flash memory die 222. In the illustrated embodiment, the I / O port 210, controller 212, DMA unit 214, and flash memory I / O port 220 are implemented on a programmable logic device ('PLD') 208, such as an FPGA. In this embodiment, each flash memory die 222 has pages organized as 16kB (kilobyte) pages 224 and registers 226 through which data can be written to or read from the flash memory die 222. In other embodiments, other types of solid-state memory are used instead of or in addition to the flash memory described within flash memory die 222.
[0117] In the various embodiments disclosed herein, storage cluster 161 can be compared to a general storage array. Storage nodes 150 are part of the collection that creates storage cluster 161. Each storage node 150 has a data slice and the computation required to provide the data. Multiple storage nodes 150 cooperate to store and retrieve data. Memory or storage devices typically used in storage arrays are less involved in processing and manipulating data. Memory or storage devices in a storage array receive commands to read, write, or erase data. Memory or storage devices in a storage array are unaware of the larger system they are embedded in, or what the data means. Memory or storage devices in a storage array can include various types of memory, such as RAM, solid-state drives, hard disk drives, etc. The non-volatile solid-state storage device 152 unit described herein has multiple interfaces that are active simultaneously and serve multiple purposes. In some embodiments, some functions of storage nodes 150 are offloaded to storage units 152, thereby transforming storage units 152 into a combination of storage units 152 and storage nodes 150. Placing computation (relative to storing data) in storage units 152 brings this computation closer to the data itself. Various system embodiments have a hierarchy of storage node layers with varying capabilities. In contrast, in a storage array, the controller possesses and knows everything about all the data managed by the controller within the rack or storage device. In storage cluster 161, as described herein, multiple controllers in multiple non-volatile solid-state storage device units 152 and / or storage nodes 150 cooperate in various ways (e.g., for erasure coding, data fragmentation, metadata communication and redundancy, storage capacity expansion or contraction, data recovery, etc.).
[0118] Figure 2D Demonstrates the storage server environment, its usage Figure 2A An embodiment of storage node 150 and storage device 152 unit to C. In this version, each non-volatile solid-state storage device 152 unit is in chassis 138 (see...). Figure 2A The PCIe (Peripheral Component Rapid Interconnect) board in the ) has a processor (e.g., controller 212 (see) Figure 2C ), FPGA, Flash 206 and NVRAM 204 (which is DRAM 216 supported by supercapacitors, see ) ... Figure 2B and 2C The non-volatile solid-state storage device 152 cell can be implemented as a single board containing the storage device and can be the maximum permissible fault domain within the chassis. In some embodiments, up to two non-volatile solid-state storage device 152 cells can fail, and the device will continue without data loss.
[0119] In some embodiments, physical storage is divided into named regions based on application usage. NVRAM 204 is a contiguous block of reserved memory in non-volatile solid-state storage device 152 DRAM 216 and is backed by NAND flash memory. NVRAM 204 is logically divided into multiple memory regions (e.g., spool_region) for two writes as spools. The space within the NVRAM 204 spool is managed independently by each authority 168. Each device provides a certain amount of storage space to each authority 168. The authority 168 further manages the lifetime and allocation within the space. Instances of spools include distributed transactions or concepts. When the main power of the non-volatile solid-state storage device 152 cell fails, an onboard supercapacitor provides short-duration power retention. During this retention interval, the contents of NVRAM 204 are flushed to flash memory 206. Upon the next power-on, the contents of NVRAM 204 are restored from flash memory 206.
[0120] As for the memory cell controller, the logic "controller's" responsibilities span across each blade distribution containing the authority 168. This distribution of logic control is... Figure 2D The diagram shows a host controller 242, a middleware controller 244, and a storage unit controller 246. Management of the control plane and storage plane is handled independently, but components can be physically co-located on the same blade. Each authority 168 effectively functions as an independent controller. Each authority 168 provides its own data and metadata structure, its own background worker, and maintains its own lifecycle.
[0121] Figure 2E Is Figure 2D Use in storage server environment Figure 2A The hardware block diagram of blade 252 in an embodiment of storage node 150 and storage unit 152 to C illustrates a control plane 254, compute and storage planes 256 and 258, and authorities 168 that interact with the underlying physical resources. The control plane 254 is partitioned into several authorities 168, which can use the compute resources in the compute plane 256 to run on any blade 252. The storage plane 258 is partitioned into a set of devices, each providing access to flash memory 206 and NVRAM 204 resources. In one embodiment, the compute plane 256 can perform the operation of a memory array controller on one or more devices (e.g., a memory array) of the storage plane 258, as described herein.
[0122] exist Figure 2EIn the compute and storage planes 256 and 258, authority 168 interacts with the underlying physical resources (i.e., devices). From the perspective of authority 168, its resources are striped across all physical devices. From the perspective of the devices, they provide resources to all authorities 168, regardless of where the authority is operating. Each authority 168 has been allocated or has been allocated one or more partitions 260 of memory in storage unit 152, such as partitions 260 in flash memory 206 and NVRAM 204. Each authority 168 uses its allocated partitions 260 for writing or reading user data. Authority may be associated with different amounts of physical storage in the system. For example, an authority 168 may have a larger number of partitions 260 or larger partitions 260 in one or more storage units 152 than one or more other authorities 168.
[0123] Figure 2F A resilient software layer is depicted in blade 252 of a storage cluster according to some embodiments. In the resilient architecture, the resilient software is symmetrical, i.e., the compute module 270 of each blade runs... Figure 2F The process described herein consists of three identical layers. Storage manager 274 executes read and write requests from other blades 252 for data and metadata stored in local storage unit 152 NVRAM 204 and flash memory 206. Authority 168 satisfies client requests by issuing necessary read and write requests to blade 252, corresponding to data or metadata residing in storage unit 152 of the blade 252. Endpoint 272 parses client connection requests received from the monitoring software of switch architecture 146, relays the client connection requests to the authority 168 responsible for implementation, and relays the response of authority 168 to the client. This symmetrical three-layer architecture enables high concurrency of the storage system. In these embodiments, elastic, efficient, and reliable horizontal scaling is achieved. Furthermore, elastic implementation employs a unique horizontal scaling technique that balances workloads across all resources regardless of client access patterns and maximizes concurrency by eliminating many of the inter-blade coordination requirements typically associated with conventional distributed locking.
[0124] Still referencing Figure 2FThe authorities 168, running in the compute module 270 of blade 252, perform the internal operations required to satisfy client requests. A characteristic of this resilience is that the authorities 168 are stateless; that is, they cache active data and metadata in their own blade 252 DRAM for fast access, but each authority stores each update in its NVRAM 204 partitions on three separate blades 252 until the update has been written to flash memory 206. In some embodiments, all storage system writes to the NVRAM 204 are written triplicate to the partitions on the three separate blades 252. With the triple-mirrorized NVRAM 204 and persistent storage protected by parity and Reed-Solomon RAID checksums, the storage system can withstand the simultaneous failure of two blades 252 without loss of data, metadata, or access to either.
[0125] Because authorities 168 are stateless, they can migrate between blades 252. Each authority 168 has a unique identifier. NVRAM 204 and flash 206 partitions are associated with the identifier of the authority 168, not with the blade 252 on which they run. Therefore, when an authority 168 migrates, it continues to manage the same storage partitions from its new location. When a new blade 252 is installed in an embodiment of the storage cluster, the system automatically rebalances the load by: partitioning the storage of the new blade 252 for use by the system's authorities 168, migrating selected authorities 168 to the new blade 252, starting endpoints 272 on the new blade 252, and including them in the client connection allocation algorithm of the switch architecture 146.
[0126] From their new locations, the migrated authorities 168 store the contents of their NVRAM 204 partitions on flash memory 206, handle read and write requests from other authorities 168, and fulfill client requests directed to them by endpoint 272. Similarly, if blade 252 fails or is removed, the system reallocates its authorities 168 among the remaining blades 252 in the system. The reallocated authorities 168 continue to perform their original functions from their new locations.
[0127] Figure 2GThe diagram depicts authorities 168 and storage resources in blades 252 of a storage cluster according to some embodiments. Each authority 168 is specifically responsible for partitioning the flash memory 206 and NVRAM 204 on each blade 252. Authority 168 manages the content and integrity of its partitions independently of other authorities 168. Authority 168 compresses incoming data and temporarily stores it in its NVRAM 204 partition, then merges, RAID-protects, and saves the data in storage segments within its flash memory 206 partition. When authority 168 writes data to flash memory 206, storage manager 274 performs necessary flash conversions to optimize write performance and maximize media lifetime. In the background, authority 168 performs "garbage collection," or reclaims space occupied by data discarded by clients, by rewriting the data. It should be understood that because the partitions of authority 168 are disjoint, distributed locking is not required for client and write or background functions.
[0128] The embodiments described herein can utilize various software, communication, and / or network protocols. Furthermore, hardware and / or software configurations can be adjusted to accommodate various protocols. For example, embodiments may utilize Active Directory, which is used in Windows. TM Database-based systems provide authentication, directories, policies, and other services in the environment. In these embodiments, LDAP (Lightweight Directory Access Protocol) is an example application protocol used to query and modify items in a directory service provider such as Active Directory. In some embodiments, a Network Lock Manager ('NLM') is used to work with a Network File System ('NFS') to provide System V-style document and record locking on the network. The Server Message Block ('SMB') protocol (one version of which is also known as the Common Internet File System ('CIFS')) can be integrated with the storage systems discussed herein. SMB operates as an application-layer network protocol, typically used to provide shared access to files, printers, and serial ports, as well as a wide variety of communications between nodes on the network. SMB also provides an authenticated inter-process communication mechanism. AMAZON TMS3 (Simple Storage Service) is a web service provided by Amazon Web Services, and the system described herein can interface with Amazon S3 via web service interfaces (REST (Representative State Transfer), SOAP (Simple Object Access Protocol), and BitTorrent). The RESTful API (Application Programming Interface) breaks down transactions into a series of small modules. Each module handles a specific underlying part of the transaction. Controls or permissions provided by these embodiments (especially for object data) may include the use of Access Control Lists ('ACLs'). An ACL is a list of permissions attached to an object, specifying which users or system processes are granted access to the object and what operations are allowed on a given object. The system may utilize Internet Protocol version 6 ('IPv6') and IPv4 as communication protocols, which provide identification and location systems for computers on a network and route traffic across the Internet. Packet routing between network systems may include Equal Cost Multipath Routing ('ECMP'), a routing strategy where the forwarding of next-hop packets to a single destination can occur on multiple "best paths" that are tied for first place in the routing metric calculation. Multipath routing can be used in conjunction with most routing protocols because it is limited to hop-by-hop decisions by a single router. The software can support multitenancy, an architecture where a single instance of a software application serves multiple clients. Each client can be referred to as a tenant. In some embodiments, tenants may be given the ability to customize parts of the application, but not the application's code. Implementations can maintain audit logs. Audit logs are documents that record events in a computing system. In addition to recording which resources are accessed, audit log entries typically include destination and source addresses, timestamps, and user login information to comply with various regulations. Implementations can support various key management strategies, such as cryptographic key rotation. Additionally, the system can support dynamic root passwords or some changes to dynamically change passwords.
[0129] Figure 3A Figures illustrate a storage system 306 according to some embodiments of the present disclosure, the storage system 306 being coupled for data communication with a cloud service provider 302. Although described in limited detail, Figure 3A The storage system 306 described herein may be similar to the one referenced above. Figures 1A to 1D and Figures 2A to 2G The described storage system. In some embodiments, Figure 3AThe storage system 306 described herein may be embodied as a storage system including an unbalanced active / active controller, a storage system including a balanced active / active controller, a storage system including an active / active controller (where less than all the resources of each controller are utilized, such that each controller has reserve resources available to support failover), a storage system including a fully active / active controller, a storage system including a data set isolation controller, a storage system including a two-tier architecture with a front-end controller and a back-end integrated storage controller, a storage system including a scale-out cluster of dual-controller arrays, and combinations of these embodiments.
[0130] exist Figure 3A In the example depicted, storage system 306 is coupled to cloud service provider 302 via data communication link 304. Data communication link 304 may be a dedicated data communication link, a data communication path provided by using one or more data communication networks, such as a wide area network ('WAN') or a LAN, or some other mechanism capable of transmitting digital information between storage system 306 and cloud service provider 302. Such data communication link 304 may be entirely wired, entirely wireless, or some aggregation of wired and wireless data communication paths. In such an example, one or more data communication protocols may be used to exchange digital information between storage system 306 and cloud service provider 302 via data communication link 304. For example, handheld device transfer protocol ('HDTP'), hypertext transfer protocol ('HTTP'), Internet protocol ('IP'), real-time transfer protocol ('RTP'), transmission control protocol ('TCP'), user datagram protocol ('UDP'), wireless application protocol ('WAP'), or other protocols may be used to exchange digital information between storage system 306 and cloud service provider 302 via data communication link 304.
[0131] Figure 3A The cloud service provider 302 depicted herein can be embodied as, for example, a system and computing environment that provides a large number of services to users of the cloud service provider 302 by sharing computing resources via data communication link 304. The cloud service provider 302 can provide on-demand access to a shared pool of configurable computing resources, such as computer networks, servers, storage devices, applications, and services. This shared pool of configurable resources can be quickly provisioned and released to users of the cloud service provider 302 with minimal management effort. Generally, users of the cloud service provider 302 are unaware of the exact computing resources used by the cloud service provider 302 to provide services. Although in many cases such a cloud service provider 302 may be accessible via the Internet, those skilled in the art will recognize that any system that abstracts the use of shared resources to provide services to users via any data communication link can be considered a cloud service provider 302.
[0132] exist Figure 3A In the examples depicted, cloud service provider 302 can be configured to provide various services to storage system 306 and its users by implementing various service models. For example, cloud service provider 302 can be configured to provide services by implementing an Infrastructure as a Service ('IaaS') service model, by implementing a Platform as a Service ('PaaS') service model, by implementing a Software as a Service ('SaaS') service model, by implementing an Authentication as a Service ('AaaS') service model, by implementing a Storage as a Service model (whereby cloud service provider 302 provides access to its storage infrastructure for use by storage system 306 and its users), and so on. Readers will understand that cloud service provider 302 can be configured to provide additional services to storage system 306 and its users by implementing additional service models, as the service models described above are included for illustrative purposes only and do not imply any limitation on the services that cloud service provider 302 can provide or the service models that cloud service provider 302 can implement.
[0133] exist Figure 3A In the examples depicted, cloud service provider 302 may be embodied as, for example, a private cloud, a public cloud, or a combination of private and public clouds. In embodiments where cloud service provider 302 is embodied as a private cloud, cloud service provider 302 may be dedicated to providing services to a single organization, rather than providing services to multiple organizations. In embodiments where cloud service provider 302 is embodied as a public cloud, cloud service provider 302 may provide services to multiple organizations. In other alternative embodiments, cloud service provider 302 may be embodied as a hybrid of private and public cloud services with a hybrid cloud deployment.
[0134] Although not in Figure 3A While explicitly described, readers will understand that numerous additional hardware and software components may be necessary to facilitate the delivery of cloud services to storage system 306 and its users. For example, storage system 306 may be coupled to (or even contain) a cloud storage gateway. This cloud storage gateway may be embodied as, for example, a hardware-based or software-based appliance located within storage system 306. This cloud storage gateway can operate as a bridge between local applications running on storage system 306 and remote cloud-based storage utilized by storage system 306. By using a cloud storage gateway, organizations can move their primary iSCSI or NAS to cloud service provider 302, thereby saving space on their internal storage systems. This cloud storage gateway can be configured to emulate disk arrays, block-based devices, file servers, or other storage systems, translating SCSI commands, file server commands, or other appropriate commands into REST space protocols that facilitate communication with cloud service provider 302.
[0135] To enable storage system 306 and its users to utilize services provided by cloud service provider 302, a cloud migration process can occur, during which data, applications, or other elements from an organization's on-premises systems (or even from another cloud environment) are moved to cloud service provider 302. To successfully migrate data, applications, or other elements to the environment of cloud service provider 302, middleware such as cloud migration tools can be used to bridge the gap between the environment of cloud service provider 302 and the organization's environment. Such cloud migration tools can also be configured to address the potentially high network costs and long transfer times associated with migrating large amounts of data to cloud service provider 302, as well as the security issues associated with sensitive data being transmitted to cloud service provider 302 via data communication networks. To further enable storage system 306 and its users to utilize services provided by cloud service provider 302, cloud orchestrators can also be used to deploy and coordinate automated tasks to create unified processes or workflows. Such cloud orchestrators can perform tasks such as configuring various components (whether cloud or on-premises) and managing the interconnections between these components. Cloud orchestrators can simplify inter-component communication and connectivity to ensure proper configuration and maintenance of links.
[0136] exist Figure 3A In the examples depicted herein, and as briefly described above, cloud service provider 302 may be configured to provide services to storage system 306 and its users using a SaaS service model, thereby eliminating the need to install and run applications on local computers, which simplifies application maintenance and support. According to various embodiments of this disclosure, such applications may take many forms. For example, cloud service provider 302 may be configured to provide access to a data analytics application to storage system 306 and its users. This data analytics application may be configured to, for example, receive large amounts of telemetry data callbacked from storage system 306. This telemetry data may describe various operational characteristics of storage system 306 and may be analyzed for various purposes, including, for example, determining the health status of storage system 306, identifying workloads performed on storage system 306, predicting when storage system 306 will exhaust various resources, and recommending configuration changes, hardware or software upgrades, workflow migrations, or other actions that may improve the operation of storage system 306.
[0137] The cloud service provider 302 can also be configured to provide access to the virtualized computing environment to the storage system 306 and its users. This virtualized computing environment can manifest as, for example, virtual machines or other virtualized computer hardware platforms, virtual storage devices, virtualized computer network resources, etc. Examples of such virtualized environments may include virtual machines created to simulate physical computers, virtualized desktop environments that separate logical desktops from physical machines, virtualized file systems that allow unified access to different types of specific file systems, and many others.
[0138] although Figure 3A The example depicted illustrates that storage system 306 is coupled for data communication with cloud service provider 302. However, in other embodiments, storage system 306 may be part of a hybrid cloud deployment, where private cloud elements (e.g., private cloud services, on-premises infrastructure, etc.) and public cloud elements (e.g., public cloud services, infrastructure, etc., provided by one or more cloud service providers) are combined to form a single solution orchestrated across various platforms. Such a hybrid cloud deployment may utilize hybrid cloud management software, for example, from Microsoft... TM Azure TM Arc centralizes the management of hybrid cloud deployments on any infrastructure and enables services to be deployed anywhere. In such an instance, hybrid cloud management software can be configured to create, update, and delete resources (both physical and virtual) that form a hybrid cloud deployment, allocate compute and storage to specific workloads, monitor workload and resource performance, policy compliance, updates and patches, security status, or perform various other tasks.
[0139] Readers will understand that various offerings can be made by pairing the storage system described herein with one or more cloud service providers. For example, Disaster Recovery as a Service ('DRaaS') can be offered, where cloud resources are used to protect applications and data from disruptions caused by disasters, included in embodiments where the storage system can be used as a primary data repository. In such embodiments, comprehensive system backups can be performed, allowing business continuity in the event of system failure. In such embodiments, cloud data backup technologies (either on their own or as part of a larger DRaaS solution) can also be integrated into a holistic solution that includes the storage system described herein and the cloud service provider.
[0140] The storage systems and cloud service providers described in this article can be used to provide a wide range of security features. For example, storage systems can encrypt data at rest (and the data can be sent to or from encrypted storage systems) and can utilize Key Management as a Service ('KMaaS') to manage encryption keys, keys used to lock and unlock storage devices, and so on. Similarly, cloud data security gateways or similar mechanisms can be used to ensure that data stored within a storage system is not improperly stored in the cloud as part of cloud data backup operations. Furthermore, micro-segmentation or identity-based segmentation can be used within data centers containing storage systems or within cloud service providers to create secure zones that isolate workloads from each other in data center and cloud deployments.
[0141] To further explain, Figure 3B Figures illustrating a storage system 306 according to some embodiments of the present disclosure are provided. Although the description is not very detailed, Figure 3B The storage system 306 described herein may be similar to the one referenced above. Figures 1A to 1D and Figures 2A to 2G The storage system described is because a storage system may contain many of the components described above.
[0142] Figure 3B The storage system 306 described herein may include a large number of storage resources 308, which may be embodied in various forms. For example, storage resources 308 may include nano-RAM or another form of non-volatile random access memory utilizing carbon nanotubes deposited on a substrate, 3D cross-point non-volatile memory, flash memory (including single-cell ('SLC') NAND flash memory, multi-cell ('MLC') NAND flash memory, three-cell ('TLC') NAND flash memory, four-cell ('QLC') NAND flash memory), or others. Similarly, storage resources 308 may include non-volatile magnetoresistive random access memory ('MRAM'), including spin-transfer torque ('STT') MRAM. Example storage resources 308 may alternatively include non-volatile phase-change memory ('PCM'), quantum memory that allows storage and retrieval of photonic quantum information, resistive random access memory ('ReRAM'), storage-class memory ('SCM'), or other forms of storage resources, including any combination of the resources described herein. Readers will learn that other forms of computer memory and storage devices, including DRAM, SRAM, EEPROM, general-purpose memory, and many others, can be used with the storage system described above. Figure 3A The storage resource 308 depicted may be embodied in various form factors, including but not limited to dual in-line memory modules ('DIMM'), non-volatile dual in-line memory modules ('NVDIMM'), M.2, U.2 and others.
[0143] Figure 3BThe storage resource 308 depicted may comprise various forms of SCM. An SCM effectively treats fast, non-volatile memory (e.g., NAND flash) as an extension of DRAM, allowing the entire dataset to be viewed as an in-memory dataset residing entirely in DRAM. The SCM may comprise non-volatile media, such as, for example, NAND flash. This NAND flash can be accessed using NVMe, which can utilize a PCIe bus for its transport, thus providing relatively low access latency compared to older protocols. In fact, network protocols used for SSDs in all-flash arrays may include NVMe using Ethernet (ROCE, NVMe TCP), Fibre Channel (NVMe FC), Infinite Bandwidth (iWARP), and others that treat fast, non-volatile memory as an extension of DRAM. Given that DRAM is typically byte-addressable and fast, non-volatile memory (e.g., NAND flash) is block-addressable, a controller software / hardware stack may be required to translate block data into bytes stored in the media. Examples of media and software that can be used as SCMs include, for example, 3D XPoint, Intel Memory Drive Technology, Samsung's Z-SSD, and others.
[0144] Figure 3B The storage resource 308 depicted may also include raceway memory (also known as domain wall memory). This raceway memory can take the form of non-volatile solid-state memory, relying on the inherent strength and orientation of the magnetic field generated by electrons as they spin within the solid-state device, in addition to their charge. By using a spin-coherent current to move magnetic domains along nanoscale permalloy wires, the domains can be altered to a bit-recording pattern by a magnetic read / write head positioned near the wire as the current passes through the wire. To fabricate a raceway memory device, numerous such wires and read / write elements can be packaged together.
[0145] Figure 3B The instance storage system 306 described herein can implement various storage architectures. For example, storage systems according to some embodiments of this disclosure may utilize block storage, where data is stored in blocks, and each block essentially serves as an individual hard disk drive. Storage systems according to some embodiments of this disclosure may utilize object storage, where data is managed as objects. Each object may contain the data itself, variable metadata, and a globally unique identifier, wherein object storage may be implemented at multiple levels (e.g., device level, system level, interface level). Storage systems according to some embodiments of this disclosure utilize file storage, where data is stored in a hierarchical structure. Such data may be stored in files and folders and presented in the same format to both the system storing it and the system retrieving it.
[0146] Figure 3BThe instance storage system 306 described herein can be embodied in a storage system in which additional storage resources can be added using a vertical scaling model, an external scaling model, or some combination thereof. In the vertical scaling model, additional storage is added by adding additional storage devices. However, in the external scaling model, additional storage nodes are added to a cluster of storage nodes, where such storage nodes may contain additional processing resources, additional network resources, and so on.
[0147] Figure 3B The instance storage system 306 described above can utilize the storage resources in various ways. For example, portions of the storage resources can be used as write caches, storage resources within the storage system can be used as read caches, or tiering can be implemented within the storage system by placing data within the storage system according to one or more tiering strategies.
[0148] Figure 3B The storage system 306 depicted also includes communication resources 310, which can be used to facilitate data communication between components within the storage system 306 and data communication between the storage system 306 and computing devices outside the storage system 306, including embodiments where the resources are spatially separated over a relatively large area. Communication resources 310 can be configured to utilize various different protocols and data communication architectures to facilitate data communication between components within the storage system and computing devices outside the storage system. For example, communication resources 310 may include Fibre Channel ('FC') technology, such as the FC architecture and FC protocol for transmitting SCSI commands over an FC network, FC over Ethernet ('FCoE') technology for encapsulating and transmitting FC frames over Ethernet, Infinite Bandwidth ('IB') technology that utilizes a switching topology to facilitate transmission between channel adapters, NVM Express ('NVMe') technology, and NVMe over Structure ('NVMeoF') technology for accessing non-volatile storage media attached via a PCIe bus, and others. In fact, the storage system described above can directly or indirectly utilize neutrino communication technology and devices to transmit information (including binary information) using neutrino beams.
[0149] Communication resources 310 may also include mechanisms for accessing storage resources 308 within storage system 306 using Serial Attached SCSI ('SAS'), a Serial ATA ('SATA') bus interface for connecting storage resources 308 within storage system 306 to a host bus adapter within storage system 306, Internet Minicomputer System Interface ('iSCSI') technology for providing block-level access to storage resources 308 within storage system 306, and other communication resources that can be used to facilitate data communication between components within storage system 306 and data communication between storage system 306 and computing devices outside storage system 306.
[0150] Figure 3B The storage system 306 depicted also includes processing resources 312, which can be used to execute computer program instructions and perform other computational tasks within the storage system 306. Processing resources 312 may include one or more ASICs and one or more CPUs customized for specific purposes. Processing resources 312 may also include one or more DSPs, one or more FPGAs, one or more system-on-a-chip ('SOC') or other forms of processing resources 312. Storage system 306 can utilize storage resources 312 to perform various tasks, including, but not limited to, supporting the execution of software resources 314, which will be described in more detail below.
[0151] Figure 3B The storage system 306 depicted also includes software resource 314, which, when executed by processing resource 312 within the storage system 306, can perform a wide range of tasks. Software resource 314 may include, for example, one or more computer program instruction modules, which, when executed by processing resource 312 within the storage system 306, can be used to implement various data protection technologies. Such data protection technologies may be implemented, for example, through system software running on the computer hardware within the storage system, through a cloud service provider, or otherwise. These data protection technologies may include data archiving, data backup, data replication, data snapshots, data and database cloning, and other data protection technologies.
[0152] Software resource 314 may also include software that can be used to implement software-defined storage ('SDS'). In such an instance, software resource 314 may include one or more computer program instruction modules that, when executed, can be used for policy-based data storage provisioning and management independent of the underlying hardware. This software resource 314 can be used to implement storage virtualization to separate the storage hardware from the software that manages the storage hardware.
[0153] Software resource 314 may also include software that can facilitate and optimize I / O operations booted to storage system 306. For example, software resource 314 may include software modules that perform various data reduction techniques, such as, for instance, data compression, data deduplication, and others. Software resource 314 may include software modules that intelligently group I / O operations to facilitate better use of the underlying storage resource 308, software modules that perform data migration operations to migrate data from within the storage system, and software modules that perform other functions. Such software resource 314 may be embodied as one or more software containers or in many other ways.
[0154] To further explain, Figure 3C Examples of a cloud-based storage system 318 according to some embodiments of this disclosure are described. Figure 3C In the examples described, the cloud-based storage system 318 is entirely within a cloud computing environment 316, such as, for example, Amazon Web Services ('AWS'). TM Microsoft Azure TM Google Cloud Platform TM IBM Cloud TM Oracle Cloud TM And others created therein. The cloud-based storage system 318 can be used to provide services similar to those provided by the storage systems described above.
[0155] Figure 3C The cloud-based storage system 318 depicted includes two cloud computing instances 320 and 322, each instance used to support the execution of storage controller applications 324 and 326. Cloud computing instances 320 and 322 may be instances of cloud computing resources (e.g., virtual machines) available from a cloud computing environment 316 to support the execution of software applications such as storage controller applications 324 and 326. For example, each of cloud computing instances 320 and 322 may execute on an Azure VM, where each Azure VM may contain high-speed temporary storage that can be used as a cache (e.g., as a read cache). In one embodiment, cloud computing instances 320 and 322 may be instances of Amazon Elastic Compute Cloud ('EC2'). In such an instance, an Amazon Machine Image ('AMI') containing storage controller applications 324 and 326 may be launched to create and configure virtual machines that can execute storage controller applications 324 and 326.
[0156] exist Figure 3CIn the example methods described above, the storage controller application programs 324 and 326 can be embodied as computer program instruction modules that, when executed, perform various storage tasks. For example, the storage controller application programs 324 and 326 can be embodied as computer program instruction modules that, when executed, perform tasks similar to those described above. Figure 1A The controllers 110A and 110B perform the same tasks, such as writing data to and from the cloud-based storage system 318, erasing data from and from the cloud-based storage system 318, retrieving data from and from the cloud-based storage system 318, monitoring and reporting disk utilization and performance, performing redundancy operations (e.g., RAID or similar data redundancy operations), compressing data, encrypting data, deduplicating data, etc. The reader will understand that because there are two cloud computing instances 320 and 322, each containing storage controller applications 324 and 326, in some embodiments, one cloud computing instance 320 may operate as the primary controller described above, while the other cloud computing instance 322 may operate as the secondary controller described above. The reader will understand that... Figure 3C The storage controller applications 324 and 326 described herein may contain the same source code that executes within different cloud computing instances 320 and 322 (e.g., different EC2 instances).
[0157] Readers will understand that other embodiments that do not include primary and secondary controllers are also within the scope of this disclosure. For example, each cloud computing instance 320, 322 may operate as a primary controller for a portion of the address space supported by the cloud-based storage system 318, wherein services for I / O operations routed to the cloud-based storage system 318 are divided in some other way, and so on. In fact, in other embodiments where cost savings may take precedence over performance requirements, there may only be a single cloud computing instance containing a storage controller application.
[0158] Figure 3C The cloud-based storage system 318 described herein includes cloud computing instances 340a, 340b, and 340n having local storage devices 330, 334, and 338. The cloud computing instances 340a, 340b, and 340n may be embodied as, for example, instances of cloud computing resources, which may be provided by the cloud computing environment 316 to support the execution of software applications. Figure 3C The cloud computing examples 340a, 340b, and 340n may differ from the cloud computing examples 320 and 322 described above, because Figure 3CCloud computing examples 340a, 340b, and 340n have local storage devices 330, 334, and 338 resources, while cloud computing examples 320 and 322, which support the execution of storage controller applications 324 and 326, do not require local storage device resources. Cloud computing examples 340a, 340b, and 340n with local storage devices 330, 334, and 338 can be embodied, for example, as EC2 M5 examples containing one or more SSDs, EC2 R5 examples containing one or more SSDs, EC2 I3 examples containing one or more SSDs, etc. In some embodiments, local storage devices 330, 334, and 338 must be embodied as solid-state storage devices (e.g., SSDs) rather than storage devices utilizing hard disk drives.
[0159] exist Figure 3C In the examples depicted, each of the cloud computing instances 340a, 340b, and 340n having local storage devices 330, 334, and 338 may include software daemons 328, 332, and 336. When executed by the cloud computing instances 340a, 340b, and 340n, the software daemons 328, 332, and 336 may present themselves to the storage controller applications 324 and 326 as if the cloud computing instances 340a, 340b, and 340n were physical storage devices (e.g., one or more SSDs). In such instances, the software daemons 328, 332, and 336 may contain computer program instructions similar to those typically contained on storage devices, enabling the storage controller applications 324 and 326 to send and receive the same commands that the storage controller would send to the storage devices. In this way, the storage controller applications 324 and 326 may contain code that is the same (or substantially the same) as the code that will be executed by the controller in the storage system described above. In these and similar embodiments, communication between storage controller applications 324, 326 and cloud computing examples 340a, 340b, 340n having local storage devices 330, 334, 338 may utilize iSCSI, NVMe over TCP, messaging, custom protocols, or some other mechanism.
[0160] exist Figure 3CIn the examples depicted, each of the cloud computing instances 340a, 340b, and 340n having local storage devices 330, 334, and 338 can also be coupled to block storage devices 342, 344, and 346 provided by the cloud computing environment 316, such as, for example, Amazon Elastic Block Storage ('EBS') volumes. In such instances, the block storage devices 342, 344, and 346 provided by the cloud computing environment 316 can be utilized in a manner similar to how the NVRAM devices described above are utilized, because the software daemons 328, 332, and 336 (or some other module) executing within a particular cloud computing instance 340a, 340b, or 340n can, upon receiving a request to write data, initiate the writing of data to its attached EBS volume and to its local storage device 330, 334, or 338 resources. In some alternative embodiments, data may be written only to local storage devices 330, 334, and 338 resources within specific cloud computing instances 340a, 340b, and 340n. In alternative embodiments, instead of using block storage devices 342, 344, and 346 provided by the cloud computing environment 316 as NVRAM, the actual RAM on each of the cloud computing instances 340a, 340b, and 340n with local storage devices 330, 334, and 338 is used as NVRAM, thereby reducing the network utilization costs associated with using EBS volumes as NVRAM. In yet another embodiment, high-performance block storage resources, such as one or more Azure Ultra disks, may be used as NVRAM.
[0161] Storage controller applications 324 and 326 can be used to perform various tasks, such as deduplicating, encrypting, or otherwise potentially updating the data to be written, before sending the request to one or more of the cloud computing examples 340a, 340b, and 340n with local storage devices 330, 334, and 338. These tasks include deduplicating the data in the request, compressing the data in the request, determining where to write the data in the request, and so on. In some embodiments, any cloud computing example 320 or 322 can receive a request to read data from a cloud-based storage system 318 and can ultimately send the request to one or more of the cloud computing examples 340a, 340b, and 340n with local storage devices 330, 334, and 338.
[0162] When a request to write data is received by a specific cloud computing instance 340a, 340b, 340n having local storage devices 330, 334, 338, software daemons 328, 332, 336 can be configured to write data not only to their own local storage devices 330, 334, 338 and any suitable block storage devices 342, 344, 346, but also to write data to a cloud-based object storage device 348 attached to the specific cloud computing instance 340a, 340b, 340n. The cloud-based object storage device 348 attached to the specific cloud computing instance 340a, 340b, 340n can be, for example, Amazon Simple Storage Service ('S3'). In other embodiments, cloud computing instances 320 and 322, each including storage controller applications 324 and 326, can initiate the storage of data in local storage devices 330, 334, and 338 of cloud computing instances 340a, 340b, and 340n, and in cloud-based object storage device 348. In other embodiments, instead of using both cloud computing instances 340a, 340b, and 340n with local storage devices 330, 334, and 338 (also referred to herein as "virtual drives") and cloud-based object storage device 348 to store data, persistent storage tiers can be implemented in other ways. For example, one or more Azure Ultra disks can be used to persistently store data (e.g., after the data has been written to the NVRAM tier).
[0163] While the local storage devices 330, 334, and 338 and the block storage devices 342, 344, and 346 utilized by cloud computing examples 340a, 340b, and 340n support block-level access, the cloud-based object storage device 348 attached to a specific cloud computing example 340a, 340b, or 340n only supports object-based access. Therefore, software daemons 328, 332, and 336 can be configured to acquire data blocks, package those data blocks into objects, and write the objects to the cloud-based object storage device 348 attached to the specific cloud computing example 340a, 340b, or 340n.
[0164] Consider an instance where data is written to local storage devices 330, 334, and 338, and block storage devices 342, 344, and 346, which are utilized in 1MB blocks by cloud computing examples 340a, 340b, and 340n. In such an instance, suppose a user of the cloud-based storage system 318 issues a request to write data, which, after being compressed and deduplicated by storage controller applications 324 and 326, results in 5MB of data needing to be written. In such an instance, writing data to the local storage devices 330, 334, and 338, and block storage devices 342, 344, and 346 utilized by cloud computing examples 340a, 340b, and 340n is relatively straightforward, as five 1MB blocks are written to these resources. In such an example, software daemons 328, 332, and 336 can also be configured to create five objects containing different 1MB data blocks. Therefore, in some embodiments, each object written to the cloud-based object storage device 348 may be the same (or nearly the same) in size. The reader will understand that in such an example, metadata associated with the data itself may be included in each object (e.g., the first 1MB of the object is the data, and the remainder is the metadata associated with the data). The reader will understand that the cloud-based object storage device 348 may be incorporated into the cloud-based storage system 318 to increase the persistence of the cloud-based storage system 318.
[0165] In some embodiments, all data stored by the cloud-based storage system 318 may be stored in either: 1) a cloud-based object storage device 348, and 2) at least one of the local storage devices 330, 334, 338 or block storage devices 342, 344, 346 resources utilized by cloud computing examples 340a, 340b, 340n. In such embodiments, the local storage devices 330, 334, 338 and block storage devices 342, 344, 346 resources utilized by cloud computing examples 340a, 340b, 340n can be efficiently operated as a cache that typically contains and stores all data in S3, such that all data reads can be served by cloud computing examples 340a, 340b, 340n without requiring cloud computing examples 340a, 340b, 340n to access the cloud-based object storage device 348. However, the reader will understand that in other embodiments, all data stored by the cloud-based storage system 318 may be stored in the cloud-based object storage device 348, but less than all data stored by the cloud-based storage system 318 may be stored in at least one of the local storage devices 330, 334, 338 resources or block storage devices 342, 344, 346 resources utilized by cloud computing examples 340a, 340b, 340n. In such instances, various strategies may be employed to determine which subset of the data stored by the cloud-based storage system 318 should reside in both: 1) the cloud-based object storage device 348, and 2) at least one of the local storage devices 330, 334, 338 resources or block storage devices 342, 344, 346 resources utilized by cloud computing examples 340a, 340b, 340n.
[0166] One or more computer program instruction modules executing within the cloud-based storage system 318 (e.g., a monitoring module executing on its own EC2 instance) can be designed to handle the failure of one or more of the cloud computing instances 340a, 340b, 340n having local storage devices 330, 334, 338. In such an instance, the monitoring module can handle the failure of one or more of the cloud computing instances 340a, 340b, 340n having local storage devices by creating one or more new cloud computing instances having local storage devices, retrieving data stored on the failed cloud computing instances 340a, 340b, 340n from the cloud-based object storage device 348, and storing the data retrieved from the cloud-based object storage device 348 on the local storage device of the newly created cloud computing instance. The reader will understand that many variations of this process can be implemented.
[0167] Readers will understand that various performance aspects of the cloud-based storage system 318 can be monitored (e.g., through a monitoring module executing in an EC2 instance), allowing the cloud-based storage system 318 to scale vertically or horizontally as needed. For example, if the cloud computing instances 320 and 322 used to support the execution of storage controller applications 324 and 326 are too small and do not adequately serve the I / O requests issued by users of the cloud-based storage system 318, the monitoring module can create a new, more powerful cloud computing instance containing the storage controller applications (e.g., a type of cloud computing instance containing more processing power, more storage, etc.), allowing the new, more powerful cloud computing instance to begin operating as the primary controller. Similarly, if the monitoring module determines that the cloud computing instances 320 and 322 used to support the execution of storage controller applications 324 and 326 are too large and that cost savings can be achieved by switching to smaller, weaker cloud computing instances, the monitoring module can create a new, weaker (and cheaper) cloud computing instance containing the storage controller applications, allowing the new, weaker cloud computing instance to begin operating as the primary controller.
[0168] The storage system described above can implement intelligent data backup technology, through which data stored in the storage system can be copied and stored in different locations to avoid data loss in the event of equipment failure or some other form of disaster. For example, the storage system described above can be configured to inspect each backup to avoid restoring the storage system to an undesirable state. Consider an instance where malware infects the storage system. In such an instance, the storage system may include software resource 314 that can scan each backup to identify backups captured before and after malware infection. In such an instance, the storage system can recover itself from backups that do not contain malware, or at least not recover portions of backups containing malware. In such an instance, the storage system may include software resource 314 that can scan each backup to identify the presence of malware (or virus, or some other unwanted substance), for example, by identifying write operations served by the storage system and originating from a network subnet suspected of delivering malware, by identifying write operations served by the storage system and originating from a user suspected of delivering malware, by identifying write operations served by the storage system and checking the content of the write operations against the fingerprint of malware, and in many other ways.
[0169] Readers will further understand that backups (typically in the form of one or more snapshots) can also be used to perform rapid recovery of storage systems. Consider an instance where the storage system is infected with ransomware that locks users out of the storage system. In such an instance, software resource 314 within the storage system can be configured to detect the presence of ransomware and can be further configured to use retained backups to restore the storage system to a point in time prior to the time the ransomware infected the storage system. In such an instance, the presence of ransomware can be explicitly detected by using software tools exploited by the system, by using a key inserted into the storage system (e.g., a USB drive), or in a similar manner. Similarly, the presence of ransomware can be inferred in response to system activity satisfying a predetermined fingerprint (e.g., for example, no reads or writes to the system within a predetermined time period).
[0170] Readers will understand that the various components described above can be grouped into one or more optimized computing packages as aggregated infrastructure. Such aggregated infrastructure may include pools of computing, storage, and network resources that can be shared by multiple applications and managed collectively using policy-driven processes. This aggregated infrastructure can be implemented using an aggregated infrastructure reference architecture, stand-alone appliances, a software-driven hyper-aggregation approach (e.g., hyper-aggregated infrastructure), or other methods.
[0171] Readers will understand that the storage systems described in this disclosure can be used to support a wide variety of software applications. In fact, a storage system can be “application-aware” in the sense that it can acquire, maintain, or otherwise access information describing the connected applications (e.g., applications utilizing the storage system) to optimize the operation of the storage system based on intelligence about the applications and their utilization patterns. For example, the storage system can optimize data layout, optimize cache behavior, optimize 'QoS' levels, or perform some other optimization designed to improve the storage performance experienced by the applications.
[0172] As an example of a type of application that can be supported by the storage system described herein, storage system 306 can be used to support these applications by providing storage resources to artificial intelligence ('AI') applications, database applications, XOps initiatives (e.g., DevOps, DataOps, MLOps, ModelOps, PlatformOps), electronic design automation tools, event-driven software applications, high-performance computing applications, simulation applications, high-speed data capture and analysis applications, machine learning applications, media production applications, media service applications, picture archiving and communication system ('PACS') applications, software development applications, virtual reality applications, augmented reality applications, and many other types of applications.
[0173] Given that storage systems encompass computing resources, storage resources, and various other resources, they are well-suited to support resource-intensive applications, such as, for example, AI applications. AI applications can be deployed across a wide range of sectors, including: predictive maintenance in manufacturing and related fields, healthcare applications (such as patient data and risk analytics), retail and marketing deployments (such as search advertising and social media advertising), supply chain solutions, fintech solutions (such as business analytics and reporting tools), operational deployments (such as real-time analytics tools), application performance management tools, IT infrastructure management tools, and many others.
[0174] This type of AI application enables devices to perceive their environment and take actions that maximize their chances of success in achieving a given goal. An example of such an AI application could be IBM Watson. TM Microsoft Oxford TM Google DeepMind TM Baidu Minwa TM And others.
[0175] The storage system described above is also well-suited for supporting other types of resource-intensive applications, such as, for example, machine learning applications. Machine learning applications perform various types of data analysis to automate the construction of analytical models. Using algorithms that iteratively learn from data, machine learning applications enable computers to learn without being explicitly programmed. A specific area of machine learning is called reinforcement learning, which involves taking appropriate actions in specific situations to maximize rewards.
[0176] In addition to the resources already described, the storage system described above may also include a graphics processing unit ('GPU'), occasionally referred to as a vision processing unit ('VPU'). This GPU may be embodied as specialized electronic circuitry that rapidly manipulates and modifies memory to accelerate the creation of images in the frame buffer for output to a display device. This GPU may be included within any computing device that is part of the storage system described above, and may include one of many individual scalable components of the storage system, other instances of which may include storage components, memory components, computing components (e.g., CPU, FPGA, ASIC), network components, software components, and others. In addition to the GPU, the storage system described above may also include a neural network processor ('NNP') for various aspects of neural network processing. This NNP may be used in place of (or to complement) the GPU, and may also be independently scalable.
[0177] As described above, the storage system described in this paper can be configured to support artificial intelligence applications, machine learning applications, big data analytics applications, and many other types of applications. The rapid growth of these applications is driven by three technologies: deep learning (DL), GPU processors, and big data. Deep learning utilizes computational models of massively parallel neural networks inspired by the human brain. Deep learning models learn their own software from a vast number of instances, rather than having experts hand-write the software. Such GPUs can contain thousands of cores perfectly suited for running algorithms that loosely represent the parallel nature of the human brain.
[0178] Advances in deep neural networks (including the development of multilayer neural networks) have sparked a new wave of algorithms and tools for data scientists to leverage their data using artificial intelligence (AI). With improved algorithms, larger datasets, and a variety of frameworks (including open-source software libraries for machine learning across a range of tasks), data scientists are tackling new use cases such as autonomous vehicles, natural language processing and understanding, computer vision, machine reasoning, strong AI, and many others. Applications of this technology include: machine and vehicle object detection, recognition, and avoidance; visual recognition, classification, and labeling; algorithmic financial trading strategy performance management; simultaneous localization and mapping; predictive maintenance of high-value equipment; cybersecurity threat mitigation; automation of specialized knowledge; image recognition and classification; question answering; robotics; text analysis (extraction, classification) and text generation and translation; and many others. AI technology has been applied in a wide range of products, including, for example: Amazon Echo's voice recognition technology, which allows users to talk to their machines; Google Translate™, which enables machine-based language translation; Spotify's Weekly Discover, which provides recommendations for new songs and artists that users may like based on user usage and traffic analysis; Quill's text generation product, which takes structured data and transforms it into narrative stories; chatbots, which provide real-time, context-specific answers to questions in a conversational format; and many others.
[0179] Data is central to modern AI and deep learning algorithms. Before training can begin, a critical issue that must be addressed is the collection of labeled data, which is essential for training accurate AI models. A comprehensive AI deployment may be required to continuously collect, clean, transform, label, and store large amounts of data. Adding additional high-quality data points directly translates into more accurate models and better insights. Data samples may undergo a series of processing steps, including but not limited to: 1) ingesting data from external sources into the training system and storing the data in its raw form; 2) cleaning and transforming the data into a training-friendly format, including linking data samples to appropriate labels; 3) exploring parameters and models, quickly testing with smaller datasets, and iterating to converge to the most promising model for deployment to the production cluster; 4) performing a training phase to select random batches of input data, including both new and old samples, and feeding those to production GPU servers for computation to update model parameters; and 5) evaluation, including using a reserved portion of the data not used during training to assess the model accuracy on the reserved data. This lifecycle is applicable to any type of parallel machine learning, not just neural networks or deep learning. For example, standard machine learning frameworks may rely on CPUs instead of GPUs, but the data ingestion and training workflows can remain the same. Readers will understand that a single shared storage data center creates a coordination point throughout the entire lifecycle, eliminating the need for additional copies of data during the ingestion, preprocessing, and training phases. Ingested data is rarely used for a single purpose, and shared storage provides the flexibility to train multiple different models or apply traditional analytics to the data.
[0180] Readers will learn that each stage in an AI data pipeline can place different requirements on the data center (e.g., a storage system or a collection of storage systems). Scale-out storage systems must deliver uncompromising performance across various access types and modes, from small files with large metadata volumes to large files, from random access to sequential access, and from low to high concurrency. The storage system described above serves as an ideal AI data center because it can serve unstructured workloads. In the first stage, data is ideally ingested and stored on the same data center that will be used in subsequent stages to avoid excessive data duplication. The next two steps can be performed on standard compute servers that optionally include GPUs, followed by a fourth and final stage where a full training production job is run on a powerful GPU-accelerated server. Typically, an experimental pipeline operating on the same dataset exists alongside a production pipeline. Furthermore, GPU-accelerated servers can be used independently for different models, or combined to train on a larger model, or even for distributed training across multiple systems. If the shared storage layer is slow, each stage must copy data to local storage, resulting in wasted time paving data across different servers. The ideal data center for an AI training pipeline delivers performance similar to data stored locally on server nodes, while also offering the simplicity and performance to enable all pipeline stages to operate simultaneously.
[0181] To enable the storage system described above to be used as a data center or as part of an AI deployment, in some embodiments, the storage system may be configured to provide DMA between storage devices contained within the storage system and one or more GPUs used in AI or big data analytics pipelines. One or more GPUs may be coupled to the storage system, for example, via structural NVMe ('NVMe-oF'), allowing bottlenecks such as those of the host CPU to be bypassed, and the storage system (or a component contained therein) to directly access GPU memory. In such instances, the storage system may utilize GPU API hooks to transfer data directly to the GPU. For example, the GPU may be an Nvidia GPU. TM The GPU and the storage system can support GPUDirect storage ('GDS') software, or have similar proprietary software that enables the storage system to transfer data to the GPU via RDMA or a similar mechanism.
[0182] Although the preceding paragraphs discussed deep learning applications, the reader will understand that the storage system described in this article can also be part of a distributed deep learning ('DDL') platform to support the execution of DDL algorithms. The storage system described above can also be paired with other technologies such as TensorFlow, an open-source software library for dataflow programming across a range of tasks, which can be used in machine learning applications such as neural networks to facilitate the development of such machine learning models, applications, etc.
[0183] The storage system described above can also be used in neuromorphic computing environments. Neuromorphic computing is a form of computation that mimics brain cells. To support neuromorphic computing, the architecture of interconnected "neurons" replaces traditional computational models with low-power signals that are transmitted directly between neurons for more efficient computation. Neuromorphic computing can utilize very large-scale integrated (VLSI) systems containing electronic analog circuitry to mimic the neurobiological architecture present in the nervous system, as well as analog, digital, and mixed-mode analog / digital VLSI and software systems for implementing neural system models for perception, motor control, or multisensory integration.
[0184] Readers will understand that the storage system described above can be configured to support the storage or use (among other types of data) of blockchain and related projects, such as, for example, open-source blockchains and those used by IBM. TM This document covers various aspects of the Hyperledger Project, including tools, permissioned blockchains that allow access to a limited number of trusted parties, blockchain products that enable developers to build their own distributed ledger projects, and others. The blockchains and storage systems described herein can be used to support both on-chain and off-chain storage of data.
[0185] Off-chain storage of data can be implemented in various ways and can occur even when the data itself is not stored on the blockchain. For example, in one embodiment, a hash function can be utilized, and the data itself can be fed into the hash function to generate a hash value. In such instances, the hash of a large amount of data can be embedded within a transaction, rather than the data itself. The reader will understand that in other embodiments, alternatives to blockchain can be used to facilitate decentralized storage of information. For example, one usable alternative to blockchain is blockweave. While conventional blockchains store each transaction for verification, blockweave allows for secure decentralized storage of data without using the entire chain, thereby achieving low-cost on-chain storage of data. This blockweave can utilize consensus mechanisms based on Proof-of-Access (PoA) and Proof-of-Work (PoW).
[0186] The storage systems described above can be used alone or in combination with other computing devices to support in-memory computing applications. In-memory computing involves storing information in RAM distributed across a computer cluster. The reader will understand that the storage systems described above (especially those configurable with customizable amounts of processing, storage, and memory resources—e.g., those systems containing blades of each type of resource with configurable amounts) can be configured in a way that supports the infrastructure for in-memory computing. Similarly, the storage systems described above may include component portions (e.g., NVDIMMs, 3D cross-point storage devices providing persistent, fast random access memory) that can effectively provide an improved in-memory computing environment compared to in-memory computing environments that rely on RAM distributed across dedicated servers.
[0187] In some embodiments, the storage system described above can be configured to operate as a hybrid in-memory computing environment that includes a common interface to all storage media, such as RAM, flash memory, and 3D cross-point storage devices. In such embodiments, users may not know the details of where their data is stored, but they can still use the same complete, unified API to address the data. In such embodiments, the storage system can (in the background) move data to the fastest available tier, including intelligently placing the data based on various characteristics of the data or some other heuristic. In such instances, the storage system can even leverage existing products such as Apache Ignite and GridGain to move data between various storage tiers, or the storage system can leverage custom software to move data between various storage tiers. The storage system described herein can implement various optimizations to improve the performance of in-memory computing, for example, by bringing the computation as close as possible to where the data occurs.
[0188] As the reader will further understand, in some embodiments, the storage system described above can be paired with other resources to support the applications described above. For example, an infrastructure may include primary computing in the form of servers and workstations dedicated to using general-purpose computing on graphics processing units ('GPGPUs') to accelerate deep learning applications interconnected to a computing engine to train parameters of deep neural networks. Each system may have Ethernet external connectivity, unlimited bandwidth external connectivity, some other form of external connectivity, or some combination thereof. In such instances, GPUs may be grouped for a single large training session or used independently to train multiple models. The infrastructure may also include, for example, the storage system described above to provide, for example, horizontally scalable all-flash file or object storage, through which data can be accessed via high-performance protocols such as NFS, S3, etc. The infrastructure may also include, for example, redundant top-of-rack Ethernet switches connected to the storage devices and computing via ports in MLAG port channels for redundancy. The infrastructure may also include additional computing in the form of white-box servers, optionally with GPUs, for data ingestion, preprocessing, and model debugging. The reader will understand that additional infrastructure is also possible.
[0189] Readers will understand that the storage system described above (whether standalone or in conjunction with other computing devices) can be configured to support other AI-related tools. For example, the storage system can utilize tools such as ONXX or other open neural network exchange formats, making it easier to transfer models written in different AI frameworks. Similarly, the storage system can be configured to support tools such as Amazon's Gluon, which allows developers to prototype, build, and train deep learning models. In fact, the storage system described above can be part of a larger platform, such as IBM's… TM A private data cloud that includes integrated data science, data engineering, and application building services.
[0190] Readers will further understand that the storage system described above can also be deployed as an edge solution. Such an edge solution can be positioned to optimize cloud computing systems by performing data processing at the network edge, close to the data source. Edge computing pushes applications, data, and computing power (i.e., services) away from a central point to the logical edge of the network. By using an edge solution with a storage system such as the one described above, computing resources provided by such a storage system can be used to perform computing tasks, storage resources can be used to store data, and cloud-based services can be accessed by using various resources of the storage system, including network resources. By performing computing tasks on an edge solution, storing data on an edge solution, and generally utilizing edge solutions, the consumption of expensive cloud-based resources can be avoided, and in fact, performance improvements can be experienced relative to a greater reliance on cloud-based resources.
[0191] While many tasks can benefit from the use of edge solutions, certain specific uses may be particularly well-suited for deployment in such environments. For example, devices such as drones, autonomous vehicles, robots, and others may require extremely high processing speeds—in fact, so high that sending data up to the cloud and back to receive processing support might simply be too slow. As an additional example, some IoT devices (such as connected cameras) may not be well-suited to utilizing cloud-based resources because, simply because of the sheer volume of data involved, sending data to the cloud may be impractical (not only from a privacy, security, or financial perspective). Therefore, many tasks truly involving data processing, storage, or communication may be better suited to platforms incorporating edge solutions (such as the storage systems described above).
[0192] The storage system described above can be used alone or in combination with other computing resources as a network edge platform that integrates computing resources, storage resources, network resources, cloud technologies, and network virtualization technologies. As part of the network, the edge can exhibit characteristics similar to other network infrastructures, from customer premises and backhaul aggregation facilities to points of presence (POPs) and regional data centers. As the reader will understand, network workloads (such as Virtual Network Functions (VNFs) and others) will reside on the network edge platform. Implemented through a combination of containers and virtual machines, the network edge platform can rely on controllers and schedulers that are no longer geographically co-located with data processing resources. As microservices, these functions can be partitioned into control planes, user and data planes, or even state machines, allowing for independent optimization and scaling techniques. Such user and data planes can be implemented through added accelerators (both residing in server platforms, such as FPGAs and smart NICs) and commercially available chips and programmable ASICs enabled by SDN.
[0193] The storage system described above can also be optimized for big data analytics, including utilization as part of a composable data analysis pipeline, where containerized analytics architectures, for example, make analytical capabilities more composable. Big data analytics can generally be described as the process of examining large and diverse datasets to uncover hidden patterns, unknown correlations, market trends, customer preferences, and other useful information that can help organizations make more informed business decisions. As part of this process, semi-structured and unstructured data (e.g., for example, internet clickstream data, web server logs, social media content, text from customer emails and survey responses, mobile phone call details, IoT sensor data, and other data) can be transformed into a structured form.
[0194] The storage system described above can also support (including implementations as system interfaces) applications that respond to human voice commands to perform tasks. For example, the storage system can support intelligent personal assistant applications, such as Amazon's Alexa. TM Apple Siri TM Google Voice TM Samsung Bixby TM Microsoft Cortana TM And others. While the example described in the preceding sentence utilizes voice as input, the storage system described above may also support chatbots, conversational bots, chatbots, or human dialogue entities, or other applications configured to engage in dialogue via auditory or text methods. Similarly, the storage system may actually execute such applications to enable users, such as system administrators, to interact with the storage system via voice. Such applications typically enable voice interaction, music playback, creating to-do lists, setting alarms, streaming podcasts, playing audiobooks, and providing weather, traffic, and other real-time information (e.g., news), but in embodiments according to this disclosure, such applications may serve as interfaces for various system management operations.
[0195] The storage system described above can also implement an AI platform to fulfill the vision of autonomous storage devices. This AI platform can be configured to deliver global predictive intelligence by collecting and analyzing vast amounts of storage system telemetry data points for easy management, analysis, and support. In fact, such a storage system can predict both capacity and performance, and generate intelligent recommendations regarding workload deployment, interaction, and optimization. This AI platform can be configured to scan all incoming storage system telemetry data against an issued fingerprint database to predict and resolve events in real time before they impact the customer environment, and capture hundreds of performance-related variables for predicting performance loads.
[0196] The storage system described above can support the serialization or simultaneous execution of artificial intelligence applications, machine learning applications, data analytics applications, data transformation, and other tasks that collectively form an AI ladder. This AI ladder can be effectively formed by combining these elements to create a complete data science pipeline, where dependencies exist between the elements. For example, AI may require some form of machine learning to have occurred, machine learning may require some form of analytics to have occurred, analytics may require some form of data and information architecture to have occurred, and so on. Therefore, each element can be considered a step in the AI ladder, collectively forming a complete and complex AI solution.
[0197] The storage system described above can also be used, either alone or in combination with other computing environments, to deliver ubiquitous AI experiences, where AI permeates a wide range of business and life. For example, AI can play a significant role in delivering deep learning solutions, deep reinforcement learning solutions, artificial general intelligence solutions, autonomous vehicles, cognitive computing solutions, commercial UAVs or drones, conversational user interfaces, enterprise taxonomy, ontology management solutions, machine learning solutions, smart dust, smart robots, smart workplaces, and many others.
[0198] The storage system described above can also be used, alone or in combination with other computing environments, to deliver a wide range of transparent immersive experiences (including those using digital twins of various “things” such as people, places, processes, systems, etc.), where technology can introduce transparency between people, businesses, and things. Such transparent immersive experiences can be delivered as augmented reality, connected homes, virtual reality, brain-computer interfaces, human augmentation technologies, nanotube electronics, volumetric displays, 4D printing, or others.
[0199] The storage system described above can also be used alone or in combination with other computing environments to support various digital platforms. Such digital platforms may include, for example, 5G wireless systems and platforms, digital twin platforms, edge computing platforms, IoT platforms, quantum computing platforms, serverless PaaS, software-defined security, neuromorphic computing platforms, etc.
[0200] The storage system described above can also be part of a multi-cloud environment, where multiple cloud computing and storage services are deployed in a single heterogeneous architecture. To facilitate operation in such a multi-cloud environment, DevOps tools can be deployed to enable orchestration across clouds. Similarly, continuous development and continuous integration tools can be deployed to standardize processes around continuous integration and delivery, new feature rollout, and provisioning of cloud workloads. By standardizing these processes, a multi-cloud strategy can be implemented that leverages the best provider for each workload.
[0201] The storage system described above can be used as part of a platform to enable the use of cryptographic anchors, which can be used to authenticate the origin and content of a product to ensure it matches the blockchain record associated with the product. Similarly, as part of a suite of tools to protect data stored on the storage system, the storage system described above can implement various cryptographic techniques and schemes, including lattice cryptography. Lattice cryptography can involve the construction of cryptographic primitives that incorporate lattices in the construction itself or in proofs of security. Unlike public-key schemes such as RSA, Diffie-Hellman, or elliptic curve cryptosystems, which are vulnerable to quantum computer attacks, some lattice-based constructions appear to be resistant to attacks by both classical and quantum computers.
[0202] A quantum computer is a device that performs quantum computation. Quantum computation is computation that utilizes quantum mechanical phenomena, such as superposition and entanglement. Quantum computers differ from conventional transistor-based computers because these computers encode data as binary digits (bits), each of which is always in one of two finite states (0 or 1). In contrast, quantum computers use qubits, which can be in superpositions of states. A quantum computer maintains a series of qubits, where a single qubit can represent any quantum superposition of 1, 0, or those two qubit states. A pair of qubits can be in any quantum superposition of four states, and three qubits can be in any superposition of eight states. A quantum computer with n qubits can typically be in any superposition of up to 2^n different states simultaneously, while a conventional computer can only be in one of these states at any given time. The quantum Turing machine is a theoretical model of such a computer.
[0203] The storage system described above can also be paired with FPGA-accelerated servers as part of a larger AI or ML infrastructure. Such FPGA-accelerated servers can reside near the storage system described above (e.g., in the same data center), or even be incorporated into an appliance that includes one or more storage systems, one or more FPGA-accelerated servers, network infrastructure supporting communication between the one or more storage systems and the one or more FPGA-accelerated servers, and other hardware and software components. Alternatively, the FPGA-accelerated servers can reside within a cloud computing environment that can be used to perform computationally relevant tasks for AI and ML jobs. Any of the embodiments described above can be used together as an FPGA-based AI or ML platform. The reader will understand that in some embodiments of an FPGA-based AI or ML platform, the FPGA contained within the FPGA-accelerated server can be reconfigured for different types of ML models (e.g., LSTM, CNN, GRU). The ability to reconfigure the FPGA contained within the FPGA-accelerated server can accelerate ML or AI applications based on optimal numerical accuracy and the memory model used. Readers will learn that by viewing a collection of FPGA-accelerated servers as an FPGA pool, any CPU in the data center can use the FPGA pool as a shared hardware microservice, rather than limiting servers to dedicated accelerators plugged into it.
[0204] The FPGA-accelerated and GPU-accelerated servers described above implement computational models where, instead of storing small amounts of data in the CPU and running long streams of instructions on it, as in more traditional computational models, machine learning models and parameters are pinned to high-bandwidth on-chip memory, with large data streams passing through it. For this type of computational model, FPGAs may even be more efficient than GPUs because FPGAs can be programmed with only the instructions required to run such computational models.
[0205] The storage system described above can be configured to provide parallel storage, for example, by using a parallel file system such as BeeGFS. This parallel file system can contain a distributed metadata architecture. For example, a parallel file system can contain multiple metadata servers spanning its distributed metadata, as well as components containing services for clients and storage servers.
[0206] The system described above supports the execution of a wide range of software applications. These applications can be deployed in various ways, including container-based deployment models. Various tools can be used to manage containerized applications. For example, Docker Swarm, Kubernetes, and others can be used to manage containerized applications. Containerized applications can be used to facilitate serverless, cloud-native computing deployment and management models for software applications. To support serverless, cloud-native computing deployment and management models for software applications, containers can be used as part of event handling mechanisms (e.g., AWS Lambdas), causing various events to trigger the launch of containerized applications to operate as event handlers.
[0207] The system described above can be deployed in various ways, including to support fifth-generation ('5G') networks. 5G networks support generally faster data communication than previous generations of mobile communication networks, thus leading to the decentralization of data and computing resources. Modern large-scale data centers can become less prominent and can be replaced by more localized micro data centers, for example, closer to mobile network towers. The system described above can be contained within such localized micro data centers and can be part of or paired with a multi-access edge computing ('MEC') system. Such MEC systems enable cloud computing capabilities and IT service environments at the edge of cellular networks. By running applications and performing related processing tasks closer to cellular customers, network congestion is reduced, and applications can perform better.
[0208] The storage system described above can also be configured to implement NVMe partitioned namespaces. By using NVMe partitioned namespaces, the logical address space of the namespace is divided into multiple regions. Each region provides a logical block address range that must be written sequentially and explicitly reset before being overwritten, thereby creating a namespace that exposes the natural boundaries of the device and offloading the management of the internal mapping table to the host. To implement NVMe partitioned namespaces ('ZNS'), ZNS SSDs or some other form of partitioned block device can be utilized, which expose the logical address space of the namespace using regions. Several inefficiencies in data placement can be eliminated when the regions are aligned with the internal physical properties of the device. In such an embodiment, each region can be mapped to, for example, a separate application, so that functions such as wear leveling and garbage collection can be performed on a region-by-region or application-by-application basis, rather than across the entire device. To support ZNS, the storage controller described herein can be configured to use, for example, Linux TM Kernel partition block device interface or other tools to interact with partition block devices.
[0209] The storage system described above can also be configured to implement partitioned storage in other ways, for example, by using shingled magnetic recording (SMR) storage devices. In instances using partitioned storage, device-managed embodiments can be deployed, where the storage device hides this complexity by managing it in the firmware, thus presenting an interface like any other storage device. Alternatively, partitioned storage can be implemented via host-managed embodiments, which depend on the operating system to know how to handle the drive and only write to certain areas of the drive sequentially. Partitioned storage can similarly be implemented using host-aware embodiments, where a combination of drive management and host management implementations is deployed.
[0210] The storage system described herein can be used to form a data lake. A data lake can operate as the first point of entry for an organization's data flow, where this data can be in its raw format. Metadata tagging can be implemented to facilitate the searching of data elements within the data lake, particularly in embodiments where the data lake contains multiple data stores (e.g., unstructured data, semi-structured data, structured data) in formats that are not easily accessible or readable. From the data lake, data can be downstream into a data warehouse, where data can be stored in formats that allow for deep processing, packaging, and consumption. The storage system described above can also be used to implement such a data warehouse. Additionally, data marts or data centers can allow for more easily consumed data, where the storage system described above can also be used to provide the underlying storage resources required for data marts or data centers. In embodiments, queries to the data lake may require a read-time schema approach, where data is applied to a plan or schema as it is pulled from the storage location, rather than as it enters the storage location.
[0211] The storage system described herein can also be configured to implement Recovery Point Objectives ('RPOs'), which can be established by a user, by an administrator, as a system default, as part of a storage class or service that the storage system is participating in delivering, or in some other way. A "Recovery Point Objective" is a target for the maximum time difference between the last update to the source dataset and the last recoverable replicated dataset update, such that, if justifiable, the update will be correctly recovered from consecutive or frequently updated copies of the source dataset. Updates can be correctly recovered if all updates processed on the source dataset prior to the last recoverable replicated dataset update are properly accounted for.
[0212] In synchronous replication, the Recovery Point Objective (RPO) will be zero, meaning that under normal operation, all completed updates on the source dataset should exist and be correctly recovered on the replicated dataset. In near-synchronous replication, the RPO can be as low as a few seconds. In snapshot-based replication, the RPO can be roughly calculated as the time interval between snapshots plus the time between transmitting previously transmitted snapshots and the most recent snapshot to be replicated.
[0213] If updates accumulate faster than they are replicated, an Recovery Point Objective (RPO) may be missed. For snapshot-based replication, if the amount of data to be replicated accumulates between two snapshots greater than the amount that can be replicated between taking a snapshot and replicating the accumulated updates from those snapshots, an RPO may be missed. Again, in snapshot-based replication, if the rate at which data to be replicated accumulates is faster than the rate at which it can be transferred over the time between subsequent snapshots, replication may begin to lag further, which can prolong the time between the expected recovery point target and the actual recovery point represented by the last correctly replicated update.
[0214] The storage system described above can also be part of a shared-nothing (SNO) cluster. In a SNO cluster, each node in the cluster has local storage and communicates with other nodes in the cluster via a network, where the storage used by the cluster is (generally) provided only by the storage devices connected to each other node. A set of nodes that synchronously replicates a dataset can be an example of a SNO cluster because each storage system has local storage and communicates with other storage systems via a network, where those storage systems (generally) do not use storage devices from other places that they share access to through some kind of interconnect. In contrast, some of the storage systems described above are themselves built as shared storage clusters because there are drive racks shared by paired controllers. However, other storage systems described above are built as shared-nothing clusters because all storage devices are local to specific nodes (e.g., blades), and all communication is through a network that links compute nodes together.
[0215] In other embodiments, other forms of shared-nothing storage clusters may include embodiments where any node in the cluster has a local copy of all the storage devices it needs, and where data is mirrored to other nodes in the cluster via synchronous replication to ensure that data is not lost, or because other nodes are also using the storage devices. In such an embodiment, if a new cluster node needs some data, the data can be copied from other nodes that have copies of the data to the new node.
[0216] In some embodiments, a mirror-based shared storage cluster can store multiple copies of all storage data in the cluster, wherein each subset of data is replicated to a specific set of nodes, and different subsets of data are replicated to different sets of nodes. In some variations, embodiments may store all storage data of the cluster across all nodes, while in other variations, nodes may be partitioned such that a first set of nodes will all store the same dataset, and a second, different set of nodes will all store different datasets.
[0217] Readers will understand that RAFT-based databases (such as etcd) can operate like a shared-nothing storage cluster, where all RAFT nodes store all data. However, the amount of data stored in a RAFT cluster can be limited, so additional replicas won't consume too much storage. Assuming containers don't tend to be too large, and their bulk data (data manipulated by applications running within containers) is stored elsewhere, such as in an S3 cluster or an external file server, then a container server cluster might also be able to replicate all data across all cluster nodes. In such instances, container storage can be provided directly by the cluster through its shared-nothing storage model, where those containers provide an image of the execution environment that forms part of an application or service.
[0218] To further explain, Figure 3D This describes an exemplary computing device 350 that can be specifically configured to perform one or more processes described herein. For example... Figure 3D As shown, computing device 350 may include a communication interface 352, a processor 354, a storage device 356, and an input / output ('I / O') module 358 that are communicatively connected to each other via communication infrastructure 360. Although in Figure 3D The demonstration computing device 350 was shown in the middle, but Figure 3D The components described herein are not intended to be limiting. Additional or alternative components may be used in other embodiments. A more detailed description will now follow. Figure 3D The components of the computing device 350 are shown in the image.
[0219] Communication interface 352 can be configured to communicate with one or more computing devices. Examples of communication interface 352 include, but are not limited to, wired network interfaces (e.g., network interface cards), wireless network interfaces (e.g., wireless network interface cards), modems, audio / video connections, and any other suitable interfaces.
[0220] Processor 354 generally refers to any type or form of processing unit capable of processing data and / or interpreting, executing and / or directing the execution of one or more of the instructions, procedures and / or operations described herein. Processor 354 may perform operations by executing computer-executable instructions 362 (e.g., application programs, software, code and / or other executable data items) stored in storage device 356.
[0221] Storage device 356 may include one or more data storage media, devices, or configurations, and may employ data storage media and / or devices of any type, form, and combination. For example, storage device 356 may include, but is not limited to, any combination of non-volatile media and / or volatile media described herein. Electronic data (including the data described herein) may be temporarily and / or permanently stored in storage device 356. For example, data representing computer-executable instructions 362 configured to instruct processor 354 to perform any of the operations described herein may be stored within storage device 356. In some instances, data may be arranged in one or more databases residing within storage device 356.
[0222] I / O module 358 may include one or more I / O modules configured to receive user input and provide user output. I / O module 358 may include any hardware, firmware, software, or a combination thereof that supports input and output capabilities. For example, I / O module 358 may include hardware and / or software for capturing user input, including but not limited to a keyboard or keypad, a touchscreen component (e.g., a touchscreen display), a receiver (e.g., an RF or infrared receiver), a motion sensor, and / or one or more input buttons.
[0223] I / O module 358 may include one or more devices for presenting output to a user, including, but not limited to, a graphics engine, a display (e.g., a screen), one or more output drivers (e.g., display drivers), one or more audio speakers, and one or more audio drivers. In some embodiments, I / O module 358 is configured to provide graphical data to a display for presentation to a user. The graphical data may represent one or more graphical user interfaces and / or any other graphical content that may serve a particular implementation. In some instances, any of the systems, computing devices, and / or other components described herein may be implemented by computing device 350.
[0224] To further explain, Figure 3EThis section illustrates an example of a storage system group 376 used to provide storage services (also referred to herein as 'data services'). The storage system group 376 depicted in Figure 3 comprises multiple storage systems 374a, 374b, 374c, 374d, and 374n, each of which may be similar to the storage systems described herein. The storage systems 374a, 374b, 374c, 374d, and 374n in the storage system group 376 may represent the same storage system or different types of storage systems. For example, Figure 3E The two storage systems 374a and 374n depicted are described as cloud-based storage systems because the resources that together form each of the storage systems 374a and 374n are provided by different cloud service providers 370 and 372. For example, the first cloud service provider 370 could be Amazon AWS. TM The second cloud service provider, 372, is Microsoft Azure. TM However, in other embodiments, one or more public clouds, private clouds, or combinations thereof may be used to provide underlying resources for forming a particular storage system in storage system group 376.
[0225] According to some embodiments of this disclosure Figure 3E The examples described herein include edge management service 382 for delivering storage services. The delivered storage services (also referred to herein as 'data services') may include, for example, services that provide a certain amount of storage to consumers, services that provide storage to consumers according to a pre-defined service level agreement, services that provide storage to consumers according to pre-defined regulatory requirements, and many others.
[0226] Figure 3E The edge management service 382 described herein may be embodied as, for example, one or more computer program instruction modules executing on computer hardware such as one or more computer processors. Alternatively, the edge management service 382 may be embodied as one or more computer program instruction modules executing on, for example, a virtualized execution environment of one or more virtual machines, in one or more containers, or otherwise. In other embodiments, the edge management service 382 may be embodied as a combination of the embodiments described above, including embodiments in which one or more computer program instruction modules contained in the edge management service 382 are distributed across multiple physical or virtual execution environments.
[0227] Edge management service 382 can operate as a gateway for providing storage services to storage consumers, where the storage services utilize storage devices provided by one or more storage systems 374a, 374b, 374c, 374d, 374n. For example, edge management service 382 can be configured to provide storage services to host devices 378a, 378b, 378c, 378d, 378n that are executing one or more applications consuming the storage services. In such an instance, edge management service 382 can operate as a gateway between host devices 378a, 378b, 378c, 378d, 378n and storage systems 374a, 374b, 374c, 374d, 374n, rather than requiring host devices 378a, 378b, 378c, 378d, 378n to directly access storage systems 374a, 374b, 374c, 374d, 374n.
[0228] Figure 3E Edge management service 382 Figure 3E The host devices 378a, 378b, 378c, 378d, and 378n expose the storage service module 380, but in other embodiments, the edge management service 382 may expose the storage service module 380 to other consumers of various storage services. Various storage services may be presented to consumers via one or more user interfaces, via one or more APIs, or through some other mechanism provided by the storage service module 380. Therefore, Figure 3E The storage service module 380 described herein may be embodied as one or more computer program instruction modules executed on physical hardware, a virtualized execution environment, or a combination thereof, wherein executing such a module enables consumers of storage services to be provided with, select, and access various storage services.
[0229] Figure 3E The edge management service 382 also includes the system management service module 384. Figure 3EThe system management service module 384 includes one or more computer program instruction modules. When executed, the module coordinates with the storage systems 374a, 374b, 374c, 374d, and 374n to perform various operations to provide storage services to the host devices 378a, 378b, 378c, 378d, and 378n. The system management service module 384 can be configured to perform tasks, such as supplying storage resources from storage systems 374a, 374b, 374c, 374d, and 374n via one or more APIs exposed by storage systems 374a, 374b, 374c, 374d, and 374n; migrating datasets or workloads within storage systems 374a, 374b, 374c, 374d, and 374n via one or more APIs exposed by storage systems 374a, 374b, 374c, 374d, and 374n; setting one or more adjustable parameters (i.e., one or more configurable settings) on storage systems 374a, 374b, 374c, 374d, and 374n via one or more APIs exposed by storage systems 374a, 374b, 374c, 374d, and 374n; and so on. For example, many of the services described below relate to embodiments of storage systems 374a, 374b, 374c, 374d, and 374n configured to operate in a certain manner. In such instances, the system management service module 384 may be responsible for configuring storage systems 374a, 374b, 374c, 374d, and 374n to operate in the manner described below using APIs (or some other mechanism) provided by storage systems 374a, 374b, 374c, 374d, and 374n.
[0230] In addition to configuring storage systems 374a, 374b, 374c, 374d, and 374n, the edge management service 382 itself can be configured to perform various tasks required to provide various storage services. Consider an instance where the storage service includes a service that, when selected and applied, causes personally identifiable information ('PII') contained in a dataset to be obfuscated when accessing the dataset. In such an instance, storage systems 374a, 374b, 374c, 374d, and 374n can be configured to obfuscate the PII when the service is directed to a read request of the dataset. Alternatively, storage systems 374a, 374b, 374c, 374d, and 374n can serve reads by returning data containing the PII, but the edge management service 382 itself can obfuscate the PII when the data passes through the edge management service 382 on its way from storage systems 374a, 374b, 374c, 374d, and 374n to host devices 378a, 378b, 378c, 378d, and 378n.
[0231] Figure 3EThe storage systems 374a, 374b, 374c, 374d, and 374n described in the reference above can be seen as examples of these systems. Figures 1A to 3D The description refers to one or more of the storage systems, including variations thereof. In fact, storage systems 374a, 374b, 374c, 374d, and 374n can be used as a storage resource pool, where individual components within the pool have different performance characteristics, different storage characteristics, etc. For example, one of storage systems 374a can be a cloud-based storage system, another storage system 374b can be a block storage system, another storage system 374c can be a file storage system, another storage system 374d can be a relatively high-performance storage system, and another storage system 374n can be a relatively low-performance storage system, and so on. In alternative embodiments, only a single storage system may exist.
[0232] Figure 3E The storage systems 374a, 374b, 374c, 374d, and 374n depicted can also be organized into different fault domains, such that a failure of one storage system 374a should be completely independent of a failure of another storage system 374b. For example, each storage system may receive power from an independent power system, each storage system may be coupled for data communication via an independent data communication network, and so on. Furthermore, storage systems in a first fault domain may be accessed via a first gateway, while storage systems in a second fault domain may be accessed via a second gateway. For example, the first gateway may be a first instance of edge management service 382, and the second gateway may be a second instance of edge management service 382, including embodiments in which each instance is different or each instance is part of distributed edge management service 382.
[0233] As an illustrative example of available storage services, storage services can be presented to users associated with different levels of data protection. For example, a storage service can be presented to a user that, when selected and implemented, guarantees that the data associated with the user will be protected, thereby ensuring various recovery point objectives ('RPO'). A first available storage service can ensure, for example, that some datasets associated with the user will be protected, such that any data older than 5 seconds can be recovered in the event of a failure of the primary data store, while a second available storage service can ensure that the datasets associated with the user will be protected, such that any data older than 5 minutes can be recovered in the event of a failure of the primary data store.
[0234] Additional instances of storage services that can be presented to users, selected by users, and ultimately applied to datasets associated with users may include one or more data compliance services. Such data compliance services may manifest as, for example, providing data compliance services to consumers (i.e., users) to ensure that users' datasets are managed in a manner that complies with various regulatory requirements. For example, one or more data compliance services may be provided to users to ensure that users' datasets are managed in accordance with the General Data Protection Regulation ('GDPR'), one or more data compliance services may be provided to users to ensure that users' datasets are managed in accordance with the Sarbanes-Oxley Act of 2002 ('SOX'), or one or more data compliance services may be provided to users to ensure that users' datasets are managed in accordance with some other regulation. Additionally, one or more data compliance services may be provided to users to ensure that users' datasets are managed in accordance with some non-governmental guidance (e.g., best practices for auditing purposes), one or more data compliance services may be provided to users to ensure that users' datasets are managed in a manner that meets the requirements of a specific client or organization, and so on.
[0235] Consider examples of specific data compliance services designed to ensure that users' datasets are managed in a manner consistent with the requirements set forth in the GDPR. While a complete list of GDPR requirements can be found in the regulation itself, for illustrative purposes, the examples set forth in the GDPR require that a pseudonymization process be applied to stored data in a way that transforms personal data in a manner that, without the use of additional information, the resulting data cannot be attributed to a particular data subject. For example, data encryption techniques can be applied to make the original data incomprehensible and irreversible without access to the correct decryption key. Therefore, the GDPR may require that the decryption key be stored separately from the pseudonymous data. A specific data compliance service may be provided to ensure compliance with the requirements set forth in this paragraph.
[0236] To provide such a specific data compliance service, the data compliance service may be presented to the user (e.g., via a GUI) and selected by the user. In response to receiving a selection for a specific data compliance service, one or more storage service policies may be applied to the dataset associated with the user to implement the specific data compliance service. For example, a storage service policy requiring the dataset to be encrypted before being stored in a storage system, before being stored in a cloud environment, or elsewhere may be applied. To implement such a policy, not only may the requirement to encrypt the dataset at storage be implemented, but also the requirement to encrypt the dataset before transmitting it (e.g., sending the dataset to another party) may be implemented. In such instances, a storage service policy requiring that any encryption keys used to encrypt the dataset are not stored on the same system storing the dataset itself may also be implemented. The reader will understand that many other forms of data compliance services can be provided and implemented according to embodiments of this disclosure.
[0237] Storage systems 374a, 374b, 374c, 374d, and 374n within storage system group 376 can be managed by one or more group management modules. The group management module can be... Figure 3E The system management service module 384 described herein may be part of or separate from it. The group management module can perform tasks such as monitoring the health status of each storage system in the group, initiating updates or upgrades on one or more storage systems in the group, migrating workloads for load balancing or other performance purposes, and many other tasks. Therefore, and for many other reasons, storage systems 374a, 374b, 374c, 374d, and 374n may be coupled to each other via one or more data communication links to exchange data between storage systems 374a, 374b, 374c, 374d, and 374n.
[0238] The storage system described herein supports various forms of data replication. For example, two or more storage systems can synchronously replicate datasets to each other. In synchronous replication, multiple storage systems may maintain different copies of a particular dataset, but all accesses to the dataset (e.g., reads) should produce consistent results, regardless of which storage system the access is routed to. For example, a read routed to any storage system of the synchronously replicated dataset should return the same result. Therefore, while updates to versions of the dataset do not need to occur exactly simultaneously, precautions must be taken to ensure consistent access to the dataset. For example, if an update (e.g., a write) routed to the dataset is received by the first storage system, the update can only be considered complete if all storage systems of the synchronously replicated dataset have applied the update to their copies of the dataset. In such instances, synchronous replication can be implemented using I / O forwarding (e.g., a write received by the first storage system is forwarded to the second storage system), communication between storage systems (e.g., each storage system instructs itself that it has completed an update), or other means.
[0239] In other embodiments, datasets can be replicated using checkpoints. In checkpoint-based replication (also known as "near-synchronous replication"), a set of updates to the dataset (e.g., one or more write operations bootstrapping the dataset) can occur between different checkpoints, such that the dataset is updated to a particular checkpoint only if all updates to the dataset prior to that checkpoint have been completed. Consider an instance where a first storage system stores a live copy of the dataset being accessed by users of the dataset. In such an instance, suppose checkpoint-based replication is used to copy the dataset from the first storage system to a second storage system. For example, the first storage system might send a first checkpoint (at time t = 0) to the second storage system, followed by a first set of updates to the dataset, and then a second checkpoint (at time t = 1). The process involves a second set of updates to the dataset, followed by a third checkpoint (at time t=2). In such an instance, if the second storage system has executed all updates from the first set but not all updates from the second set, then the copy of the dataset stored on the second storage system is up-to-date until the second checkpoint. Alternatively, if the second storage system has executed all updates from both the first and second sets, then the copy of the dataset stored on the second storage system is up-to-date until the third checkpoint. The reader will understand that various types of checkpoints can be used (e.g., metadata-only checkpoints), and checkpoints can be distributed based on various factors (e.g., time, number of operations, RPO settings), etc.
[0240] In other embodiments, the dataset can be replicated via snapshot-based replication (also known as "asynchronous replication"). In snapshot-based replication, snapshots of the dataset can be sent from a replication source, such as a first storage system, to a replication target, such as a second storage system. In such embodiments, each snapshot may contain the entire dataset or a subset of the dataset, for example, only the portion of the dataset that has changed since the last snapshot was sent from the replication source to the replication target. The reader will understand that snapshots can be sent on demand, based on a strategy that takes into account various factors (e.g., time, number of operations, RPO settings), or in some other manner.
[0241] The storage systems described above can be configured individually or in combination to serve as continuous data protection storage. Continuous data protection storage is a feature of storage systems that records updates to a dataset, allowing access to a consistent picture of the dataset's previous contents at low temporal granularity (typically about a few seconds or even less) and back to a reasonable time period (typically hours or days). This allows access to the most recent consistent point in time of the dataset, and also allows access to points in time that may immediately precede an event (e.g., causing partial corruption or additional loss of the dataset), while maintaining the maximum number of updates close to said event. Conceptually, they are like a series of snapshots of a dataset taken very frequently and stored for a long time, but continuous data protection storage is typically implemented very differently from snapshots. Storage systems implementing continuous data protection storage can further provide means of accessing these points in time, accessing one or more of these points in time as snapshots or clones, or restoring the dataset to one of those recorded points in time.
[0242] Over time, to reduce overhead, some points in time stored in continuous data protection storage can be merged with other nearby points in time, essentially deleting some of these points from the storage. This reduces the capacity required for storage updates. A limited number of these points in time can also be converted into snapshots with longer durations. For example, such storage can maintain a low-granularity sequence of points in time several hours from now, where some points in time are merged or deleted to reduce overhead by up to an additional day. Going back further in time, some of these points in time can be converted into snapshots representing a consistent image of points in time, occurring only every few hours.
[0243] Although some embodiments are described primarily in the context of storage systems, those skilled in the art will recognize that embodiments of this disclosure may also take the form of a computer program product disposed on a computer-readable storage medium for use with any suitable processing system. Such a computer-readable storage medium may be any storage medium for machine-readable information, including magnetic media, optical media, solid-state media, or other suitable media. Examples of such media include disks or floppy disks in hard disk drives, optical disks for optical disk drives, magnetic tapes, and other media that will be apparent to those skilled in the art. Those skilled in the art will readily recognize that any computer system with suitable programming elements will be able to perform the steps described herein, as embodied in the computer program product. Those skilled in the art will also recognize that while some embodiments described in this specification are oriented toward software installed and executed on computer hardware, alternative embodiments implemented as firmware or hardware are fully within the scope of this disclosure.
[0244] In some instances, a non-transitory computer-readable medium may be provided for storing computer-readable instructions, based on the principles described herein. When executed by a processor of a computing device, the instructions may direct the processor and / or the computing device to perform one or more operations, including one or more operations described herein. Such instructions may be stored and / or transmitted using any of a variety of known computer-readable media.
[0245] The term "non-transitory computer-readable media" as used herein may include any non-transitory storage medium that contributes to providing data (e.g., instructions) that can be read and / or executed by a computing device (e.g., a processor of the computing device). For example, non-transitory computer-readable media may include, but is not limited to, non-volatile storage media and / or any combination of volatile storage media. Exemplary non-volatile storage media include, but are not limited to, read-only memory, flash memory, solid-state drives, magnetic storage devices (e.g., hard disks, floppy disks, magnetic tapes, etc.), ferroelectric random access memory ("RAM"), and optical discs (e.g., optical discs, digital video discs, Blu-ray discs, etc.). Exemplary volatile storage media include, but are not limited to, RAM (e.g., dynamic RAM).
[0246] Now we turn to solutions to reliability issues affecting multi-chassis storage systems and storage systems with a large number of blades, described as having resilient groups. The blades of the storage system are divided into groups called resilient groups, and the storage system implements writes to select target blades belonging to the same group. Implementations are grouped into four mechanisms presented below in various embodiments.
[0247] 1) The formation of elastic groups and how to improve the reliability of storage systems through storage system expansion.
[0248] 2) Write path, and how to reliably extend data writes through storage system expansion: Flash writes (segmentation) and NVRAM insertions should not cross the boundaries of write groups.
[0249] 3) Majority group, and how to reliably upgrade the startup through storage system expansion: The majority of blades participate in the consensus algorithm. The majority group is used by the startup process.
[0250] 4) Authoritative election, and how to reliably expand the system through storage system expansion: Witness algorithms can be executed using witness groups within the boundaries of elastic groups.
[0251] Figure 4 This describes an elastic group 402 in the storage system that supports data recovery in the event of the loss of up to a specified number of blades 252 in the elastic group 402. In some embodiments, the formation of the elastic group is determined solely by the cluster geometry, which includes the location of each blade in each chassis. An instance cluster geometry has only one current version of the elastic group. The elastic group can change if blades are added, removed, or moved to different slots. In some embodiments, no more than one blade can be added or removed from the cluster geometry at a time. In some embodiments, the elastic group is calculated by the array leader (running on a top-of-rack (ToR) switch) and pushed as part of the cluster geometry update.
[0252] For write paths, in some embodiments, flash segment formation selects target blades from a single elastic group. NVRAM insertion first selects target blades from a single elastic group. Additionally, NVRAM insertion attempts to select blades also in the same chassis to avoid overloading the ToR switch. When the cluster geometry changes, the array leader recalculates the elastic groups. If a change exists, the array leader rolls out the new version of the cluster geometry to each blade in the storage system. Each leader rescans the base storage segments and gathers a list of segments split into two or more new elastic groups. Garbage collection uses this list and remaps segments back to the same elastic group for super redundancy, where all data segments are in their respective elastic groups, and no data segment spans or crosses two or more elastic groups.
[0253] like Figure 4As explained, in some embodiments, one or more write groups 404 are formed within each elastic group 402 for data striping (i.e., writing data into data stripes). For example, each elastic group 402 may contain fewer than all blade / storage nodes of a chassis, all blade / storage nodes of a chassis, or blades and storage nodes of two chassis, or fewer or more of those. It should be understood that the number and / or arrangement of blade / memory-equipped components for the elastic group 402 may be user-definable, adaptable, or determined based on component characteristics, system architecture, or other criteria.
[0254] In one embodiment, write group 404 is a data stripe spanning a group of blades 252 (or other physical devices with memory) on which data striping is performed. In some embodiments, write group is a RAID stripe spanning a group of blades 252. Error-coded data written in write group 404 as data stripes can be recovered through error correction in the event of the loss of up to a specified number of blades 252. For example, using N+2 redundancy and error coding / error correction, data can be recovered in the event of the loss of two blades. Other redundancy levels can be easily defined for the storage system. In some embodiments, depending on the redundancy level, each write group 404 may extend across a number of blades 252, less than or equal to the total number of blades 252 in the elastic group 402 to which write group 404 belongs. For example, an elastic group 402 of ten blades may support write groups 404 of three blades 252 (e.g., dual mirroring), seven blades 252 (e.g., N+2 redundancy, N=5), and / or ten blades 252 (e.g., N+2 redundancy, N=8), and other possibilities. Because write groups 404 are established in this manner, within elastic groups 402, in the event of the loss of up to a specified number of blades 252 equal to the error coding and error correction in each write group 404, any and all data in elastic group 402 can be recovered through error correction. For write groups 404 with N+2 redundancy in elastic group 402, all data in elastic group 402 remains recoverable even if up to two blades 252 fail, become unresponsive, or are otherwise lost.
[0255] Some embodiments of the storage system have ownership and authority over the data 168, see, for example, the references above. Figures 2B to 2GThe storage cluster is described. Authority 168 can be a software or firmware entity or agent that executes, generates, and accesses metadata on one or more processors, having ownership and access rights to a specified range of user data (e.g., a series of segment numbers or one or more inode numbers). Data owned and accessed by Authority 168 is striped in one or more write groups 404. In versions of storage systems with Authority 168 and Elastic Group 402, an Authority on a blade 252 in Elastic Group 402 can access write groups 404 within the same Elastic Group 402, as described by... Figure 4 Some arrows are shown in the diagram. Furthermore, the authority on blade 252 and elastic group 402 can access write group 404 in another elastic group 402, as shown by... Figure 4 The other arrows in the image are shown further.
[0256] Figure 5 This is a scenario involving a storage system geometry change 506, where blades 252 are added to storage cluster 161, resulting in a change from a previous version of Elastic Group 502 to a current version of Elastic Group 504. In this example, the storage system enforces a rule that Elastic Group 402 cannot be as large as two chassis 138 filled with blades 252. In one embodiment, where chassis 138 has a maximum capacity of 15 blades 252, this rule would mean that the maximum number of blades 252 in Elastic Group 402 is 29 blades in both chassis 138 (e.g., one less than 2 × 15). Adding blades 252 to storage cluster 161 makes the total blade count greater than the maximum value of Elastic Group 402, causing the storage system to reassign blades to Elastic Group 402 that does not violate this rule. Therefore, the 29 blades 252 of the previous version of Elastic Group 502 are reassigned to the current version of Elastic Group 504, labeled "Elastic Group 1" and "Elastic Group 2". Following the teachings of this document, it is easy to design other rules and scenarios for the elastic groups 402 and geometry changes 506 of the storage system. For example, a storage cluster with 150 blades can have 10 elastic groups, each with 15 blades. Removing blades 252 from storage cluster 161 can result in more than one elastic group 402 of blades 252 being reassigned to a smaller number or one elastic group 402. Adding or removing chassis with blades 252 to the storage system can result in the blades 252 being reassigned to elastic groups 402. By specifying the maximum number and / or associated arrangement of blades 252 for elastic groups 402, and by assigning blades 252 according to elastic groups 402, the storage system can expand or shrink its storage capacity, and has component loss survivability and data recoverability that scale appropriately without an exponential increase in the probability of data loss.
[0257] Figure 6This is a system and action diagram showing the authority 168 in the distributed computing resource 602 of the storage system and the switch 146 communicating with the chassis 138 and blades 252 to form a resilient group 402. (See above for reference.) Figure 2F and Figure 2G An embodiment of switch 146 is described. (See above for reference.) Figures 2A to 2G Embodiments of distributed computing resource 602 and memory 604 are described in section 3B. Other embodiments of switch 146, distributed computing 602, and memory 604 can be easily designed based on the teachings herein.
[0258] continue Figure 6 The switch 146 has an array leader 606 that communicates with chassis 138 and blades 252, implemented in software, firmware, hardware, or a combination thereof. The array leader 606 evaluates the number and arrangement of blades 252 in one or more chassis 138, determines membership of a resilient group 402, and transmits information about the membership of the resilient group 402 to the chassis 138 and blades 252. In various embodiments, this determination may be rule-based, table-based, derived from data structures, or artificial intelligence, etc. In other embodiments, the blades 252 may form the resilient group 402 through a decision-making process in distributed computing 602.
[0259] Figure 7 Describes garbage collection in the context of Elastic Group 402 reclaiming memory and relocating data. In this scenario, such as... Figure 5 Adding blade 252 to cluster 161 causes the storage system to reassign blade 252 from the previous version of elastic group 502 to the current version of elastic group 504. Data stripe 702 from the previous version of elastic group 502 is split across the two current versions of elastic groups 504. That is, some data in data stripe 702 is physically located in memory on some (e.g., one or more) blades 252 in elastic group 1, and some data in data stripe 702 is physically located in memory on some (e.g., one or more) blades 252 in elastic group 2. If the data in the data stripe is not reassigned, the data will be vulnerable to the loss of two blades in elastic group 1 and one blade in elastic group 2, or vice versa, and the data written to the data stripe in group 404 from each current elastic group 504 (i.e., elastic group 1 and elastic group 2) will be recoverable (see [link to relevant documentation]). Figure 4To remedy this situation and ensure that the data conforms to the current version of FlexGroup 504, the data in data stripe 702 is relocated (i.e., reassigned, redistributed, or transferred) to one, another, or both of the current versions of FlexGroup 504, but not across two FlexGroups 402. That is, in some versions, the data in data stripe 702 is transferred to FlexGroup 1, or the data in data stripe 702 is transferred to FlexGroup 2. In some versions, the data stripe may be split into two data stripes, with the data in one of the two new data stripes transferred to FlexGroup 1, and the data in the other of the two new data stripes transferred to FlexGroup 2. The situation should not be that data stripe 702 or any subdivided new data stripe is split across the current FlexGroup 504. Although for other embodiments, this relocation may be defined as a separate process, Figure 7 In the embodiment shown, garbage collection performs data relocation with the help of Authority168. See below for reference. Figure 8 Describe such a version. It's easy to design. Figure 7 The scene shown in the image changes, with other geometric changes and other arrangements of the flexible group 402.
[0260] Figure 8 This is a system and action diagram illustrating the details of garbage collection 802, including the recovery of coordinated memory and the scanning 810 and relocation 822 of data. Standard garbage collection, a known process for moving data, followed by erasing and reclaiming memory (e.g., in software, firmware, hardware, or a combination thereof), is modified herein to relocate 822 of real-time data 808 within or to another elastic group 402, such that memory with obsolete data 806 can be erased 816 and reclaimed 818. By relocating data relative to elastic group 402, the storage system, when faced with, for example, regarding... Figure 4 The discussed failure and extended scenarios aim to improve reliability and data recoverability. By combining garbage collection and the relocation of data from the current Elastic Group 504, the storage system improves efficiency compared to these being separate processes.
[0261] exist Figure 8In the embodiments and scenarios illustrated, garbage collection 802 collaborates with authority 168 to scan 810, access 812, and recover 814 memory. Authority 168 provides addressing information about where the data is located and information (e.g., metadata) about the data stripe 804, write group 404, and elastic group 402 to which the data is associated or mapped, for garbage collection 802 to use in determining whether to move the data or where to move it. Typically, physical locations in memory are in an erase or write state, and if in a write state, in some embodiments, the data is live data (i.e., readable, valid data) or obsolete data (i.e., invalid data because it has been overwritten or deleted). Memory blocks with a sufficient number (e.g., a threshold) of erase or obsolete data locations or addresses can be erased or reclaimed after any remaining live data is relocated, i.e., removed / copied from the block and obsolete within the block. In collaboration with Authority 168, garbage collection 802 can identify candidate regions for reclamation of 818 in memory through a top-down process, starting with data segments and working down through address indirection levels to the physical addressing of memory. Alternatively, garbage collection 802 can use a bottom-up process, starting with the physical addressing of memory and working up through address indirection to the data segment number, and associated with data stripes 804, write groups 404, and flexible groups 402. Other embodiments of the mechanism for identifying regions for reclamation of 818 can be easily developed following the teachings herein.
[0262] once Figure 8Garbage collection 802 has identified candidate regions in the memory, and determines whether said region contains obsolete data 806, real-time data 808, erased, unused / unwritten addresses, or some combination thereof. Regions containing completely obsolete data 806 (e.g., solid-state memory blocks) can be erased 816 and reclaimed 818 without data movement. For example, regions with a larger amount of obsolete data 806 or unused addresses but containing some real-time data 808 are candidate regions for relocating real-time data 808. To relocate 822 real-time data 808, garbage collection 802 reads real-time data 808, copies real-time data 808, and writes real-time data 808 to the same current elastic group 504 (if real-time data 808 is found there), or writes from a previous version of elastic group 502 to the current version of elastic group 504 (if real-time data 808 is found in a previous version of elastic group 502). After relocation 822, garbage collection 802 discards the data in the previous location 820, then erases 816 and reclaims 818 the memory occupied by obsolete data 806. The relocated real-time data 808 is remapped 824, where garbage collection 802 collaborates with authority 168 to write metadata 826, which reassigns real-time data 808 to write group 404 in elastic group 402, as described above.
[0263] Figure 9 A majority group 904 is depicted for the boot process 902 using elastic group 402. The storage system forms the majority group 904 for the boot process 902 based on elastic group 402, such that even when expanded by storage system expansion, there are sufficient resources within each elastic group 402 for successful and reliable operation. When no data stripes 702, 804 span two or more elastic groups 402, the majority group 904 for the boot process 902 is the same as the elastic group 402. When there are boot information data stripes spanning two or more elastic groups 402 (e.g., due to a recent cluster geometry change 506), the majority group 904 is a superset of said elastic group 402, which is formed by merging blades 252 from two or more elastic groups 252 into one, and may contain all storage nodes 150 or blades 252. The startup process 902 of the majority group 904 is coordinated across all blade / storage nodes in the majority group 904 to ensure that all data stripes with startup information (such as from the commit log and basic storage segments) can be read correctly.
[0264] In some embodiments, in a steady state, the majority group 904 is the same as the elastic group 402. In any given state, the formation of the majority group 904 is determined by the current version of the elastic group 504 and the historical formation of the elastic group 402. To initiate the startup process 902, various services must reach a majority (N-2 blades, where N is the group size, distinct from the N+2 redundancy N) in each majority group 904. When the elastic group 402 changes, the array leader 606 calculates a new majority group 904 from the current elastic group 402 and past majority groups 904 (if any). The array leader 606 saves the new majority group 904 to distributed storage. After the new majority group 904 is saved, the array leader 606 pushes the new cluster geometry to the blades 252. Each authority 168 refreshes NVRAM and remaps segments until all base storage segments and commit logs reach super redundancy (i.e., consistency with the elastic group 402). The array leader 606 continues to poll the NFS state until it reaches super redundancy. Next, array leader 606 causes majority group 904 to exit the distributed storage.
[0265] Figure 10 Witness group 1006 and authority election 1002 using elastic group 402 are depicted. Storage node 150 coordinates and selects witness group 1006 as the first elastic group 402 of the storage cluster. Witness group 1006 holds authority election 1002 and renews authority lease 1004 (defined as the time span during which authority 168 is valid for user data I / O). In some embodiments, elastic group 402 and witness group 1006 are updated when the cluster geometry changes.
[0266] Figure 11 This is a system and action diagram illustrating the details of the authority election 1002 and allocation, organized by Witness Group 1006 and involving majority vote 1102 across blades 252 of the storage system. At startup, members of Witness Group 1006 receive majority votes from all blades 252 in the cluster geometry (e.g., N / 2+1, where N is the number of blades in the storage cluster, distinct from the redundant N+2). Once majority votes 1102 are received, the highest-ranking blade 252 in Elastic Group 402 determines the authority 168 placement among blades 252 (from a static dictionary), as determined by... Figure 11 The arrow in the diagram indicates that an endorsement message with the authority number to be initiated will be sent to blade 252. Upon receiving an authority endorsement message, each blade 252 checks and ensures that it comes from the same geometry version, and initiates authority 168 (e.g., to execute threads in a multi-threaded processor environment, access metadata, and data in memory 604). Witness group 1006 will continue to renew authority lease 1004 until elastic group 402 changes again.
[0267] Figure 12 This is a flowchart of a method for operating a storage system with resilient groups. The method can be implemented using the storage system described herein and its variations, more specifically, using processors within the storage system. In action 1202, a resilient group of blades in the storage system is formed. (See reference...) Figure 4 Discuss elastic groups. In action 1204, write the data stripe into the elastic group. In the determination action 1206, query whether there is a geometry change in the storage system. If the answer is no, then there is no geometry change, and the process returns to action 1204 to continue writing the data stripe. If the answer is yes, then there is a geometry change in the storage system, and the process continues to determination action 1208.
[0268] In action 1208, it is determined whether the geometry change meets the criteria used to correct the elastic group. For example, for elastic groups, these criteria may be based on rules regarding the maximum number and arrangement of blades relative to the chassis. If the answer is no, the geometry change does not meet the criteria, and the process returns to action 1204 to continue writing data stripes. If the answer is yes, the geometry change meets the criteria, and the process continues to action 1210. In action 1210, the blades are reassigned from the elastic group of the previous version to the elastic group of the current version. In various embodiments, the number of elastic groups and the number and arrangement of blades in the elastic groups are case- and rule-specific. The storage system propagates the assignment to the blades and retains version information.
[0269] In action 1212, data is transferred from a previous version of the elastic group to the current version of the elastic group. In various embodiments, the data transfer is coordinated by garbage collection with the help of the authority that owns and manages access to the data in memory (see [link to relevant documentation]). Figure 13 This can be a separate process from garbage collection. In action 1214, it is determined whether all data has been transferred from the previous version of the elastic group to the current version of the elastic group. If the answer is no, then not all data has been transferred to the current version of the elastic group, and the process returns to action 1212 to continue transferring data. If the answer is yes, then all data has been transferred to the current version of the elastic group, and the process continues to action 1216. In action 1216, the previous version of the elastic group is exited. No new data is written to any previous version of the elastic group, and in some embodiments, all references to previous version of the elastic group can be deleted. In some embodiments, once exited, data cannot be read from any previous version of the elastic group because all data has been transferred to the current version of the elastic group.
[0270] Figure 13This is a flowchart of a method for operating a storage system with garbage collection in a resilient group. The method can be implemented using the storage system described herein and its variations, more specifically, using processors within the storage system. In various embodiments, it can be used... Figure 12 The method described above is implemented in this way. In determining action 1302, it is determined whether a new elastic group of the current version has been formed. If the answer is no, then no new elastic group exists, and the process loops back to determining action 1302 (or alternatively branches elsewhere to perform other tasks). If the answer is yes, then a new elastic group of the current version exists, for example, because a change in the geometry of the storage system has occurred or been detected, and the arrangement of the blades and chassis meets the criteria (see [reference]). Figure 12 Action 1208), the process continues to action 1304.
[0271] In Action 1304, garbage collection, aided by authorities in the storage system, coordinates memory recovery and data scanning and relocation through attachment mapping. Authorities execute in the blades and manage metadata and own and access specified ranges of user data in the storage system's memory through addressing schemes, error coding, and error correction. In Action 1306, it is determined whether live data in a previous version of an Elastic Group spans two or more current versions of Elastic Groups. Figure 7 This demonstrates an example of data striping across two elastic groups after a change in the geometry of the storage cluster results in a change to the elastic groups. If the answer is no, then there is no live data spanning two or more current elastic groups in the previous version of the elastic group, and the process continues to determination action 1312. If the answer is yes, then there is live data spanning two or more current elastic groups in the previous elastic group, and the process continues to action 1308.
[0272] In action 1308, real-time data is relocated from a previous version of the elastic group to the current version of the elastic group. Relocation may involve reading real-time data from the previous version of the elastic group, copying the real-time data, writing the real-time data to the current version of the elastic group, and then discarding the data in the previous version of the elastic group, making the memory location reclaimable by garbage collection. In action 1310, garbage collection erases and recovers memory holding stale data. For example, solid-state memory blocks containing only stale data or stale data combined with erased state locations may be erased and then made available for writing as recovered memory.
[0273] In action 1312, it is determined whether the real-time data in the current version of the elastic group occupies the target location for garbage collection recovery of memory. If the answer is no, then no data exists in that case, and the process branches back to action 1304 for garbage collection to continue coordinating memory recovery with data scanning and relocation. If the answer is yes, then the real-time data in the current version of the elastic group occupies the target location for garbage collection recovery of memory, and the process continues to action 1314. In action 1314, the real-time data is relocated elsewhere in the current version of the elastic group. As mentioned above, relocation involves reading, copying, and writing the real-time data to a new location in the current version of the elastic group, and discarding the real-time data at the earlier location. However, this time, the relocation is within the same elastic group because the elastic group where the real-time data is found is the current version of the elastic group. In action 1316, garbage collection erases and recovers the memory holding the outdated data.
[0274] It should be understood that the methods described herein can be implemented using digital processing systems, such as conventional general-purpose computer systems. Alternatively, a dedicated computer designed or programmed to perform only one function can be used. Figure 14 This is an illustration demonstrating exemplary computing devices capable of implementing the embodiments described herein. According to some embodiments, Figure 14 The computing device can be used in embodiments that perform resilient grouping and garbage collection functionality. The computing device includes a central processing unit (CPU) 1401 (coupled to memory 1403 via bus 1405) and a mass storage device 1407. Mass storage device 1407 represents a persistent data storage device, such as a floppy disk drive or a fixed disk drive, and in some embodiments, it may be local or remote. Memory 1403 may include read-only memory, random access memory, etc. In some embodiments, applications residing on the computing device may be stored on or accessed via a computer-readable medium, such as memory 1403 or mass storage device 1407. Applications may also be in the form of modulated electronic signals accessed via a network modem or other network interface of the computing device. It should be understood that in some embodiments, CPU 1401 may be embodied in a general-purpose processor, a dedicated processor, or a dedicated programming logic device.
[0275] Display 1411 communicates with CPU 1401, memory 1403, and mass storage device 1407 via bus 1405. Display 1411 is configured to display any visualization tools or reports associated with the system described herein. Input / output device 1409 is coupled to bus 1405 to pass information from command selection to CPU 1401. It should be understood that data to and from external devices can be passed via input / output device 1409. CPU 1401 can be defined to perform the functionality described herein to implement the reference... Figures 1A to 13 The functionality described. In some embodiments, the code embodying this functionality may be stored in memory 1403 or mass storage device 1407 for execution by a processor such as CPU 1401. The operating system on the computing device may be MSDOS. TM MS-WINDOWS TM OS / 2 TM UNIX TM LINUX TM Or other known operating systems. It should be understood that the embodiments described herein can also be integrated with virtualized computing systems implemented using physical computing resources.
[0276] Figure 15 Resilient groups 1508, 1510, and 1512 are depicted formed by blades 252 with different amounts of memory 604 (1504, 1506). This can occur when different types of blades 252 exist in the storage system, such as when the system is upgraded, a failed blade is physically replaced by a newer blade 252 with more memory 604, a spare blade is brought online, or a new chassis is connected to create or expand a multi-chassis system. If there are sufficient numbers of each type of blade to form multiple resilient groups, then the system's computing resources 1502 can form one resilient group 1508 with one type of blade 252 (e.g., 15 blades each with 8TB of storage capacity) and another resilient group 1510 with another type of blade 252 (e.g., 7 blades each with 52TB of storage capacity). As described above, in the event of failure of two blades 252 in a resilient group, or in the event of failure of some other predetermined number of blades 252 in a resilient group during a change, each resilient group supports data recovery, for example, by erasing the code.
[0277] Alternatively, the system's computing resources 1502 can be formed into a flexible group 1512 with two types of blades. A longer data segment 1518 is formed using all 1504 of the memory 604 of one type of blade (e.g., all 8TB of memory 1504 in each of the first blade subset of the storage system and 8TB of memory from each of the second blade subset of the storage system (e.g., blades 252 each with 52TB of memory). A shorter data segment 1520 is formed using the remaining amount of memory in the second blade subset, for example, 44TB of memory 1516 remaining after subtracting the 8TB of memory 1504 used in the longer data segment 1518 from the total 52TB of memory in the blades. Because the longer data segment 1518 has higher storage efficiency than the shorter data segment 1520 (i.e., higher storage efficiency relative to redundancy overhead for stored user data), this formation of data segments 1518 and 1520 has overall higher storage efficiency than data segments that could be formed in two flexible groups 1508 and 1510. Some embodiments may preferably use one or the other, and some embodiments may be able to migrate data from two flexible groups to one flexible group, or vice versa, or otherwise reorganize the blades and flexible groups as described above or otherwise. It is easy to design. Figure 15 The scenarios described herein vary, including additional amounts of memory, flexible groups, blades, data segment formation and allocation within flexible groups, and single-chassis or multi-chassis systems. It should be understood that the Authority 168 participates in the formation of data segments and the reading and writing of data within those segments, as discussed above.
[0278] Figure 16 A conservative estimate of the amount of memory space available in elastic group 1602 is depicted. The computing resources 1502 of the storage system track memory usage in blades 252 and determine which blade 252 in elastic group 1602 is the fullest, or equivalently, the least empty. Figure 16 A crossed-out line indicates a full memory space, meaning a memory space that has been written to but not erased, and contains either real-time data or invalid, outdated data that has not been erased. Figure 16 Empty spaces, rather than crossed lines, are used to indicate empty memory space, i.e., memory space that has been erased and is available for writing. In some embodiments, the amount of empty available space 1604 in the fullest or least empty blade of elastic group 1602 is multiplied by the number of blades in elastic group 1602 to produce a conservative estimate of the available memory space in elastic group 1602. This represents the minimum possible amount of memory space available for writing, and it should be understood that the actual amount of memory space available in elastic group 1602 may be larger. However, this conservative estimate is well used in various embodiments to make data writing and garbage collection decisions, such as Figure 17 It is displayed in the middle.
[0279] Figure 17 Describing with Figure 15 and 16 The garbage collection module 1702 describes various options related to elastic groups. Garbage collection performs real-time data reading from memory 604 and real-time data writing to memory 604. It should be understood that garbage collection is performed to merge obsolete data blocks for block erasure of solid-state memory and memory recovery to make space available for writing. In one embodiment, memory space tracker 1704 tracks used and available memory space and suggests source selector 1706 for data reading by garbage collection module 1702 and destination selector 1708 for data writing. Authority 168 participates in data reading, data writing, and decisions about which data is real-time relative to which data is obsolete, and identifies candidates for memory erasure. Garbage collection module 1702 can read data into and write data from the same elastic group 1710. Alternatively, garbage collection module 1702 can read data from one elastic group 1710 and write data to a different elastic group 1710. In some embodiments, the elastic group to which the data is written may even be an elastic group with lower storage efficiency than the elastic group from which the data is read.
[0280] In some embodiments, decisions regarding data writes are handled through a weighted random distribution of data writes. In some embodiments, weights are increased, i.e., increased towards elastic groups with conservative estimates of larger available memory space, and weights are decreased, i.e., moved away from but still included towards elastic groups with conservative estimates of smaller available memory space. Using this weighted random distribution of data writes, emptier elastic groups 1710 tend to be written to more frequently, thus filling up faster than non-empty elastic groups 1710, thereby balancing memory usage and available space across the blades 252 of the storage system. In some embodiments, the weighted random distribution of data writes is applied to data ingestion from the storage system.
[0281] The following are options for various embodiments of the garbage collection module 1702 in some examples. Some versions select from two or more of these options, while others pursue a single option. Garbage collection can be performed between the same elastic group. Garbage collection can be performed between different elastic groups, i.e., reading from one elastic group and writing to another. Garbage collection can use a weighted random distribution for data writes, as described above. Garbage collection can be performed in the elastic group with the lowest available space, i.e., reading and writing in the elastic group with the lowest available space. Garbage collection can read from the elastic group with the lowest available space and write to the elastic group with the largest available space, or write using a weighted random distribution weighted toward the elastic group with the largest available space. The estimation of available space can be referenced as above. Figure 16 The described implementation uses blades with the minimum amount of available memory space for writing to generate a conservative estimate of the available memory space in the elastic group, or it can be based on a more detailed tracking of the space available in each blade. These estimates of the available space in the elastic group can be used to determine the elastic group with the minimum available space and the elastic group with the maximum available space. These determinations of available space and elastic groups can then be used to make decisions about the weighted random distribution of garbage collection and / or data writing. Garbage collection can read from elastic groups with higher storage efficiency and write to elastic groups with lower storage efficiency. Following the teachings of this document, other options and combinations of the features described above can be easily designed.
[0282] Figure 18 This is a flowchart of a method for operating a storage system to form a resilient group. The method can be performed using various embodiments of the storage system described herein and its variations, as well as other storage systems with memory and redundancy. In action 1802, the amount of memory on the blades is determined. This is the total amount of erased or unerased memory, not just the amount of memory available for writing. In decision action 1804, it is determined whether the blades have different amounts of memory. If the answer is yes, then the blades have different amounts of memory, and the process continues to action 1806. If the answer is no, then the blades do not have different amounts of memory, i.e., all blades have the same amount of memory, and the process continues to action 1808.
[0283] In action 1806, flexible groups are formed based on the different amounts of memory in the blades of the storage system. For example, one or more flexible groups can be formed from blades with the same amount of memory. One or more flexible groups can be formed from blades with different amounts of memory, wherein data segments of different lengths are formed within the flexible groups to maximize the use of different amounts of memory and storage efficiency.
[0284] In action 1808, elastic groups are formed based on the same amount of memory in the blades of the storage system. For example, each elastic group may have a threshold number or more blades. Elastic groups do not overlap; that is, a blade can belong to only one elastic group. Each elastic group thus formed supports data recovery in the event of failure of two blades or, in other embodiments, a specific number (e.g., one or three) of blades. Extensions to the above method for performing garbage collection within elastic groups can be easily designed, as referenced above. Figure 17 describe.
[0285] Figure 19 A storage system according to an embodiment is described, the storage system having one or more compute resource elastic groups defined in a compute region and one or more storage resource elastic groups defined in a storage region. This storage system embodiment and its variations change the mechanism of elastic groups from having blade 252 groups (see...). Figure 2E to 2G (Sections 4 through 7, 11, 12, 15, and 16) are extended to a finer granularity in terms of members and memberships within elastic groups, and compute and storage resources are decoupled based on system and data survivability under failure conditions. This flexibility in defining elastic groups and the memberships of various system resources within them further enhances the scalability of the storage system. Compute resources can be scaled independently of storage resources, and vice versa, where system reliability and survivability under failure conditions can be addressed by reconfiguring these finer-grained elastic groups. For example, storage resources can be expanded by adding or upgrading modular storage devices 1922 and a storage resource elastic group 1912 redefined independently of compute resource elastic group 1908. Similarly, compute resources can be expanded by upgrading or adding blades 1904 (while keeping modular storage devices 1922 as is) and a compute resource elastic group 1908 redefined independently of storage resource elastic group 1912. In other embodiments, expanding both compute and storage resources can trigger a reconfiguration of both compute resource elasticity group 1908 and storage resource elasticity group 1912 to maintain system reliability and survivability at desired levels. Some embodiments perform group-level space and load balancing and may have different elasticity groups associated with group-level space and load. Some embodiments decouple compute and storage from removable drives and define elasticity groups associated with compute availability and storage / data persistence. In some embodiments, the elasticity group determines whether the cluster is available when some blades or devices fail and determines whether this affects storage / data persistence as removable drives recover data.
[0286] continue Figure 19The embodiments described herein illustrate a storage system 1900 having multiple blades 1904, which may span a single chassis or multiple chassis and be heterogeneous or homogeneous in terms of capacity and type of flash memory, etc. For example, each blade 1904 may be for hybrid computing and storage, storage only, or computing only, wherein the combination of blades 1904 in a particular storage system 1900 embodiment collectively has the computing and storage resources required for distributed data storage, and supports heterogeneity or homogeneity of each of the various resources. In this embodiment, each blade may have zero or more modular storage devices 1922, one blade is illustrated as having two modular storage devices 1922, while other blades may have four modular storage devices 1922, and so on. In some embodiments, the modular storage devices 1922 are pluggable and may be hot-swappable, or in some embodiments may be fixed within the blade 1904. In various embodiments, the modular storage device 1922 may be homogeneous or heterogeneous across a given blade 1904, and the blade 1904 across the storage system 1900 may be homogeneous or heterogeneous. The solid-state memory 1930 may be homogeneous or heterogeneous in terms of memory type and / or memory amount (i.e., storage capacity) among the given modular storage devices 1922, between the modular storage devices 1922 on the blade 1904, and / or across the blade 1904 in the storage system 1900.
[0287] Storage system 1900 defines a compute region 1906 spanning blade 1904, within which one or more compute resource elastic groups 1908 are defined, and these are both configurable and reconfigurable. The compute region generally contains resources for compute within storage system 1900. This includes compute to form writes (see...). Figure 20 The computing resources can also be used for operating systems, applications, etc. Each computing resource elastic group 1908 includes various computing resources from multiple blades 1904, such as CPUs, processor local memory 1916, and communication modules 1918 for connecting and communicating with modular storage devices 1922 via network 1920. In other embodiments, such resources can be subdivided into other elastic groups, which can be easily designed as taught herein. For example, CPU 1914, processor local memory 1916, communication module 1918, authority 168 (see [link to relevant documentation]). Figure 2B , 2E Resources of types 2G, 4G, and / or others can be included as resource groups within elastic groups, or subdivided into different regions and different elastic groups. A compute resource elastic group 1908 does not necessarily span all blades 1904 in the storage system 1900, but one or more compute resource elastic groups can be defined across blade 1904 groups. The storage system 1900 supports various configurations of elastic groups.
[0288] Storage system 1900 defines a storage region 1910 spanning modular storage devices 1922, within which one or more flexible groups of storage resources 1912 are defined, and these are both configurable and reconfigurable. Storage region 1910 generally contains resources for storing data and metadata within the storage system. This includes resources for writing (see...). Figure 20 Storage devices that read data stripes (e.g., RAID stripes), such as solid-state memory 1930 in various embodiments. Each storage resource elastic group 1912 includes various storage resources from multiple modular storage devices 1922. In some embodiments, storage region 1910 is subdivided into multiple regions with various granularities, and storage resource elastic groups 1912 may be formed in or within storage region 1910. Figure 19 In the embodiments depicted herein, there exists an NVRAM region 1932 comprising NVRAM 1928 in modular storage device 1922, a memory region 1934 comprising solid-state memory 1930 of modular storage device 1922, a temporary storage region 1936 comprising a subset of solid-state memory 1930 of modular storage device 1922, and a long-term storage region 1938 comprising another subset of solid-state memory 1930 of modular storage device 1922. For other embodiments, various combinations or other granularities of regions and flexible groups within or in regions of storage region 1910 can be readily designed, in accordance with the teachings herein.
[0289] For example, in other embodiments, NVRAM 1928 may be located on blade 1904 instead of NVRAM 1928 in modular storage device 1922, or as a supplement thereto. Controller 1926 in modular storage device 1922 may be located in a controller region within storage region 1910. Communication module 1924 of modular storage device 1922 may be located in a communication resource elastic group or communication resource region within storage region 1910. Memory in storage buffer region 1936 may be, for example, single-level cell (SLC) flash memory, used for fast buffering during data striping, followed by moving the data of the data stripe to long-term memory in long-term storage region 1938, which may be, for example, multi-level cell memory, such as quad-level cell (QLC) flash memory (which has longer write times and higher data bit density). This subdivision of storage region 1910 and the establishment of elastic groups within this subdivision can be used for background or post-processing of stored data, including deduplication, compression, and encryption. Each elastic region has a defined redundancy level for fault tolerance. The granularity and membership flexibility of various resources within the various defined elastic groups in different regions provide greater system flexibility, particularly in upgrades and expansions, to achieve the desired level of system and data reliability and resource loss survivability, which will be discussed below. Figure 20 Further description.
[0290] Figure 20 An embodiment of a storage system is described, which uses resources in elastic groups to form and write data stripes. In this example, the storage system defines multiple compute resource elastic groups 1908 in compute region 1906 and multiple storage resource elastic groups 1912 in storage region 1910. Compute resources in each compute resource elastic group 1908 perform action 2002 to form a data stripe. Additionally, storage resources in each storage resource elastic group 1912 perform action 2004 to write the data stripe. Each data stripe is transferred from the formed compute region 1906 to the written storage region 1910. This broad description of storage system operation is readily implementable in various storage systems, where additional modifications to the storage system define compute resource elastic groups 1908 and storage resource elastic groups 1912, and support various configurations of the elastic groups as described herein.
[0291] In various operational scenarios, storage systems experience failures, and this paper demonstrates the benefits of using various elastic groups in different implementations. When a storage system performs an action where the survivable number of members of an elastic group is lost (e.g., 2006), in... Figure 20The description in the text summarizes the failure of one or more components (i.e., system resources). Each elastic group has a defined redundancy level 2010 such that a specified number of members of the elastic group can fail, be removed, or be corrupted, and the storage system operation can still perform 2008 and continue to operate because the data is recoverable and the elastic group is reconfigurable. For example, an elastic group defined as having one redundant member can be designed to be N+1 redundant and have N+1 members, but can continue to operate and recover data, having N members after losing one member. Similarly, an elastic group defined as having two redundant members can be designed to be N+2 redundant and have N+2 members, but can continue to operate and recover data, having N members after losing two members due to a failure. And, an elastic group defined as having three redundant members can be designed to be N+3 redundant and have N+3 members, but can continue to operate and recover data, having N members after losing three members of the elastic group. In some embodiments, different elastic groups may have different defined redundancy levels. Some embodiments support N+4, N+5, or even higher redundancy, for example, for segment coding. Some embodiments define resilient groups for temporary storage and different resilient groups for long-term storage, and in various embodiments, these resilient groups may have the same or different redundancy levels. It is important to recognize that a failure of a particular component can affect one or more resilient groups, and this example of system survivability is illustrative and not intended to be limiting.
[0292] In an illustrative example, modular storage device 1922 fails, while other modular storage devices 1922 and blade 1904 remain operational. In this example, the failed modular storage device 1922 is considered a loss of storage resource elasticity group 1912, but data can still be recovered from the remaining resources of storage resource elasticity group 1912 because there is sufficient membership or resource redundancy. Other storage resource elasticity groups 1912 with other modular storage devices 1922 as members are unaffected by the failure. It should be understood that computing resource elasticity group 1908 will be unaffected in this scenario.
[0293] In the event of a failure of the entire blade 1904, and depending on the nature of the failure, this may or may not affect the modular storage device 1922 on the blade 1904 in various scenarios. This, in turn, may affect the computing resource elastic group 1908 that has the failed blade, or more specifically, the computing resources of the blade, as a member of the computing resource elastic group 1908, and may or may not affect the storage resource elastic group 1912 that has the modular storage device 1922 mounted on the blade 1904 as a member. In this scenario, other computing resource elastic groups 1908 that have other blades as members will be unaffected.
[0294] It is easy to explore various other failures and failure granularities of individual components within storage devices or blades, as well as other scenarios involving various granularities and compositions of resources, resilient groups, and regions. Through the description in this paper, it is easy to understand how the flexibility of resilient group definitions can, independently or synergistically, benefit the scalability of computing and storage resources to maintain the desired level of system and data reliability and recoverability in the event of failure. For example, storage system 1900 can detect failures of a single processor, a single NVRAM 1928, a single communication path, a single section of a certain type of memory, or even failures of a survivable number of components in each of multiple resilient groups, and then recover the data, followed by reconfiguring one or more resilient groups with appropriate resource content and granularity. The storage system with the reconfigured resilient groups can then continue operating with the desired level of system and data reliability and recoverability.
[0295] Figure 21 This is a flowchart illustrating a method for a storage system to utilize resources within a resilient group according to an embodiment. The method may be embodied in a tangible computer-readable storage medium. In various embodiments, the method may be implemented by the processing device of the storage system. In action 2102, the storage system establishes a resilient group comprising one or more compute resource resilient groups and one or more storage resource resilient groups. Each resilient group has a defined resource redundancy level for the storage system. Examples of various types of resilient groups, various resource redundancy levels of resilient groups, and heterogeneous and homogeneous resource support for resilient groups have been described above, and their variations are readily designable.
[0296] In action 2104, the storage system supports the ability to configure elastic groups with multiples of each type. For example, the storage system can be configured with a single elastic group of multiple types, a multiple of one type of elastic group and a single elastic group of other types, multiples of multiple types of elastic groups, etc. It should be understood that at any given moment of operation, the storage system has a specific configuration of elastic groups, but can be reconfigured to other configurations of elastic groups, for example, in response to failures, changes in membership of blade or modular storage devices, upgrades, expansions, or extensions of the storage system, etc.
[0297] In action 2106, the storage system performs distributed data and metadata storage across modular storage devices based on elastic groups. Depending on the specific configuration, the system may use the compute resources of the elastic group to form data stripes, or do so in parallel across more than one compute resource elastic group. The system may use the storage resources of the elastic group to write data stripes, or do so in parallel across more than one storage resource elastic group. In action 2108, it is determined whether one or more members of the elastic group are faulty. If the answer is no, then there is no fault, and the process branches back to action 2106 to continue performing distributed data and metadata storage. If the answer is yes, then one or more members of the elastic group are faulty, and the process continues to action 2110.
[0298] In action 2110, the storage system recovers data. Because the resilient group experiencing the failure of one or more members is configured with a defined resource redundancy level, the data can be recovered using the remaining resources of the resilient group (and other system resources suitable for a particular embodiment). For example, a resilient group with an N+R configuration can tolerate up to R failures and is able to recover data. In action 2112, the system determines whether the failed one or more members have been recovered. For example, a member of the resilient group may have been offline or become uncommunicative and then come back online or restarted, restoring communication. As another example, a blade containing multiple drives may be physically removed and then placed back into the chassis, resulting in the storage system becoming operationally recoverable. If the failed member has been recovered, the process branches back to action 2106 to continue with distributed data and metadata storage. If the answer is no, then the failed one or more members have not been recovered, and the process branches to action 2114.
[0299] In action 2114, the storage system reconfigures one or more elastic groups. The specific elastic groups to be reconfigured and the new configuration will be situation- and system-specific, and this reconfiguration capability will be supported by the system before, during, and after the reconfiguration. After the elastic groups are reconfigured, the process continues to action 2106 to continue executing distributed data and metadata based on the reconfigured elastic groups.
[0300] It should be understood that an elastic group with an N+R configuration that loses R+1 or more members may not be able to recover data. If the elastic group loses R members or fewer, the system can continue to operate and the data can be rebuilt (e.g., to another drive) to regain >N members in order to withstand failures that would otherwise reduce the write group to below N. Reconfiguring the elastic group supports the desired N+R or other configurations and maintains system survivability and data recovery capabilities.
[0301] The advantages and features of this disclosure can be further described by the following statements:
[0302] 1. A method comprising:
[0303] Multiple elastic groups are established from a storage system with multiple blades, each elastic group having a defined resource redundancy level of the storage system, wherein the multiple elastic groups include at least one computing resource elastic group having computing resources with the multiple blades and at least one storage resource elastic group having storage resources with multiple modular storage devices.
[0304] The storage system's ability to support multiple configurations, each configuration having multiple of each of the multiple elastic groups; and
[0305] Distributed data and metadata storage is performed by the multiple blades across the multiple modular storage devices according to the established multiple elastic groups.
[0306] 2. According to the method of statement 1, wherein performing the distributed data and metadata storage includes:
[0307] The computing resources in the computing resource elastic group are used to form at least one data stripe; and
[0308] The storage resources in the storage resource elastic group are used to write the at least one data stripe.
[0309] 3. The method according to statement 1 further includes:
[0310] For each of the plurality of elastic groups, the defined redundancy level is selected from the group including N+1 redundancy, N+2 redundancy, and N+3 redundancy.
[0311] 4. The method according to statement 1 further includes:
[0312] Select different defined redundancy levels from among the multiple elastic groups.
[0313] 5. The method according to statement 1 further includes:
[0314] Supports heterogeneous computing resources in at least one computing resource elastic group.
[0315] 6. The method according to statement 1 further includes:
[0316] Supports heterogeneous storage resources in at least one storage resource elastic group.
[0317] 7. The method according to statement 1 further includes:
[0318] It supports different defined redundancy levels in the first elastic group of storage resources in the temporary storage area and the second elastic group of storage resources in the long-term storage area.
[0319] 8. A tangible, non-transitory computer-readable medium having instructions thereon, the instructions causing the processor to perform, when executed by a processor, a method comprising:
[0320] Multiple elastic groups are established from a storage system with multiple blades, each elastic group having a defined resource redundancy level of the storage system, wherein the multiple elastic groups include at least one computing resource elastic group having computing resources with the multiple blades and at least one storage resource elastic group having storage resources with multiple modular storage devices.
[0321] The storage system's ability to support multiple configurations, each configuration having multiple of each of the multiple elastic groups; and
[0322] Distributed data and metadata storage is performed by the multiple blades across the multiple modular storage devices according to the established multiple elastic groups.
[0323] 9. The computer-readable medium according to statement 8, wherein the method further comprises:
[0324] For each of the plurality of elastic groups, the defined redundancy level is selected from the group including N+1 redundancy, N+2 redundancy, and N+3 redundancy.
[0325] 10. The computer-readable medium according to statement 8, wherein the method further comprises:
[0326] Select different defined redundancy levels from among the multiple elastic groups.
[0327] 11. The computer-readable medium according to statement 8, wherein the method further comprises:
[0328] Different redundancy levels are defined in the first elastic group of storage resources in the temporary storage area and the second elastic group of storage resources in the long-term storage area.
[0329] 12. A storage system comprising:
[0330] Multiple modular storage devices;
[0331] Multiple blades, which can be configured to establish multiple elastic groups, each elastic group having a defined resource redundancy level;
[0332] The plurality of elastic groups includes at least one computing resource elastic group and at least one storage resource elastic group;
[0333] The plurality of blades are used to support configurations, each configuration having a plurality of each of the plurality of resilient groups; and
[0334] The multiple blades are configured to perform distributed data and metadata storage across the multiple modular storage devices according to the multiple resilient groups.
[0335] 13. The storage system according to statement 12, wherein the plurality of blades and the plurality of modular storage devices will:
[0336] Data stripes are formed using resources within a compute resource elastic group; and
[0337] The data stripe is written using resources from the storage resource elastic group.
[0338] 14. The storage system according to statement 12, wherein in each of the plurality of elastic groups, the defined redundancy level is related to the fault survivability of the elastic group and can be selected from a group including N+1 redundancy, N+2 redundancy and N+3 redundancy.
[0339] 15. The storage system according to statement 12, wherein the plurality of blades supports configurations with different defined redundancy levels in the plurality of resilient groups.
[0340] 16. The storage system according to statement 12, wherein the plurality of blades support heterogeneous computing resources in the at least one computing resource elastic group.
[0341] 17. The storage system according to statement 12, wherein the plurality of blades support heterogeneous storage resources in the at least one storage resource elastic group.
[0342] 18. The storage system according to statement 12, wherein the plurality of blades supports different defined redundancy levels in a first elastic group of storage resources in a temporary storage area and a second elastic group of storage resources in a long-term storage area.
[0343] 19. The storage system of statement 12, wherein the plurality of blades supports different defined redundancy levels in a first elastic group comprising authorities among the plurality of blades and a second elastic group comprising other authorities among the plurality of blades.
[0344] 20. The storage system according to statement 12, wherein the plurality of resilient groups further includes at least one resilient group of non-volatile random access memory (NVRAM).
[0345] This document discloses detailed illustrative embodiments. However, the specific functional details disclosed herein are merely representative for the purpose of describing the embodiments. Furthermore, embodiments may be embodied in many alternative forms and should not be construed as being limited to the embodiments set forth herein.
[0346] It should be understood that although the terms first, second, etc., may be used herein to describe various steps or calculations, these steps or calculations should not be limited by these terms. These terms are used only to distinguish one step or calculation from another. For example, without departing from the scope of this disclosure, a first calculation may be referred to as a second calculation, and similarly, a second step may be referred to as a first step. As used herein, the terms “and / or” and the “ / ” symbol encompass any and all combinations of one or more of the associated listed items.
[0347] As used herein, the singular forms “a,” “an,” and “the” are also intended to include the plural forms, unless the context clearly indicates otherwise. It will be further understood that, when used herein, the terms “comprises,” “comprising,” and / or “includes,” specify the presence of the stated features, integers, steps, operations, elements, and / or components, but do not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof. Therefore, the terminology used herein is for the purpose of describing particular embodiments only and is not intended to be limiting.
[0348] It should also be noted that in some alternative implementations, the functions / actions mentioned may not appear in the order shown in the diagrams. For example, in fact, depending on the functions / actions involved, two diagrams shown consecutively may be executed substantially simultaneously or sometimes in reverse order.
[0349] Considering the above embodiments, it should be understood that the embodiments may employ various computer-implemented operations involving data stored in a computer system. These operations are operations requiring physical manipulation of physical quantities. Typically, but not necessarily, these quantities take the form of electrical or magnetic signals that can be stored, transmitted, combined, compared, and otherwise manipulated. Furthermore, the manipulations performed are generally referred to using terms such as generating, identifying, determining, or comparing. Any operations described herein that form part of the embodiments are useful machine operations. The embodiments also relate to a device or apparatus for performing these operations. The apparatus may be specifically constructed for the desired purpose, or the apparatus may be a general-purpose computer selectively activated or configured by a computer program stored in a computer. Specifically, various general-purpose machines may be used with computer programs written in accordance with the teachings herein, or it may be more convenient to construct more specialized apparatus to perform the desired operations.
[0350] Modules, applications, layers, agents, or other method-operable entities can be implemented as hardware, firmware, or processor-executed software, or a combination thereof. It should be understood that in the software-based embodiments disclosed herein, the software may be embodied in a physical machine, such as a controller. For example, the controller may include a first module and a second module. The controller may be configured to perform various actions, such as methods, applications, layers, or agents.
[0351] The embodiments may also be embodied as computer-readable code on a tangible, non-transitory computer-readable medium. A computer-readable medium is any data storage device capable of storing data that can subsequently be read by a computer system. Examples of computer-readable media include hard disk drives, network-attached storage devices (NAS), read-only memory, random access memory, CD-ROMs, CD-Rs, CD-RWs, magnetic tapes, and other optical and non-optical data storage devices. Computer-readable media may also be distributed across a network-coupled computer system, enabling the distributed storage and execution of computer-readable code. The embodiments described herein can be practiced with a variety of computer system configurations, including handheld devices, tablet computers, microprocessor systems, microprocessor-based or programmable consumer electronics, minicomputers, mainframe computers, and the like. The embodiments can also be practiced in distributed computing environments, where tasks are performed by remote processing devices via wired or wireless network links.
[0352] Although the methods are described in a specific order, it should be understood that other operations may be performed between the described operations, the described operations may be adjusted so that they occur at slightly different times, or the described operations may be distributed across a system that allows processing operations to occur at various time intervals associated with the processing.
[0353] In various embodiments, one or more portions of the methods and mechanisms described herein may form part of a cloud computing environment. In such embodiments, resources may be provided as a service via the Internet according to one or more various models. These models may include Infrastructure as a Service (IaaS), Platform as a Service (PaaS), and Software as a Service (SaaS). In IaaS, computing infrastructure is delivered as a service. In this case, computing devices are typically owned and operated by the service provider. In the PaaS model, software tools and underlying devices used by developers to develop software solutions are provided as a service and hosted by the service provider. SaaS typically includes software licensed by the service provider as an on-demand service. The service provider may host the software or deploy it to customers for a given period of time. Multiple combinations of the above models are possible and considered.
[0354] Various units, circuits, or other components may be described or required to be “configured to” or “configurable to” perform one or more tasks. In this context, the phrase “configured to” or “configurable to” is used to imply a structure by indicating that the unit / circuit / component contains a structure (e.g., a circuit system) that performs one or more tasks during operation. Thus, even when the specified unit / circuit / component is currently inoperable (e.g., not turned on), it may be said that the unit / circuit / component is configured to perform a task or is configurable to perform a task. Units / circuit / components used with the terms “configured to” or “configurable to” include hardware, such as circuits, memory storing program instructions executable to perform operations, etc. The reference to a unit / circuit / component being “configured to” or “configurable to” performing one or more tasks is not expressly intended to invoke paragraph 6 of 35 U.S.C. 112 for the purpose of referring to such unit / circuit / component as “configured to” or “configurable to” performing one or more tasks. Additionally, "configurable to" or "configurable to" may include a general-purpose structure (e.g., a general-purpose circuit system) manipulated by software and / or firmware (e.g., an FPGA or a general-purpose processor executing the software) to operate in a manner capable of performing the tasks discussed. "Configurable to" may also include adapting manufacturing processes (e.g., semiconductor manufacturing facilities) to manufacture devices (e.g., integrated circuits) suitable for implementing or performing one or more tasks. "Configurable to" is expressly intended not to apply to blank media, unprogrammed processors or unprogrammed general-purpose computers, or unprogrammed programmable logic devices, programmable gate arrays, or other unprogrammed devices, unless accompanied by programmable media that endows the unprogrammed device with the ability to be configured to perform the disclosed functions.
[0355] For purposes of explanation, specific embodiments have been described above. However, the foregoing illustrative discussion is not intended to be exhaustive or to limit the invention to the precise forms disclosed. Many modifications and variations are possible in light of the foregoing teachings. Embodiments were chosen and described to best explain the principles of the invention and its practical application, thereby enabling others skilled in the art to best utilize embodiments and various modifications suitable for the particular purpose contemplated. Therefore, these embodiments are to be considered illustrative and non-limiting, and the invention is not limited to the details given herein, but can be modified within the scope of the appended claims and their equivalents.
Claims
1. A method comprising: Multiple elastic groups are established from a storage system with multiple storage nodes. Each elastic group has a defined resource redundancy level of the storage system. The multiple elastic groups include multiple types of elastic groups, including at least one computing resource elastic group with computing resources of the multiple storage nodes and at least one storage resource elastic group with storage resources of multiple modular storage devices. The storage system supports the ability to configure multiple types of elastic groups, wherein each of the multiple configurations of the multiple types of elastic groups includes multiple storage resource elastic groups and multiple compute resource elastic groups; and Distributed data storage is performed by the multiple storage nodes across the multiple modular storage devices according to the multiple resilient groups established.
2. The method according to claim 1, wherein performing the distributed data storage comprises: Use the computing resources in the computing resource elastic group to form at least one data stripe; and The storage resources in the storage resource elastic group are used to write the at least one data stripe.
3. The method according to claim 1, further comprising: For each of the plurality of elastic groups, the defined redundancy level is selected from the group including N+1 redundancy, N+2 redundancy, and N+3 redundancy.
4. The method of claim 1, further comprising: Select different defined redundancy levels from among the multiple elastic groups.
5. The method of claim 1, further comprising: Supports heterogeneous computing resources in at least one computing resource elastic group.
6. The method of claim 1, further comprising: Supports heterogeneous storage resources in at least one storage resource elastic group.
7. The method of claim 1, further comprising: Supports different defined redundancy levels within the first elastic group and the second elastic group.
8. A tangible, non-transitory computer-readable medium having instructions thereon, the instructions causing the processor to perform, when executed by a processor, a method comprising: Multiple elastic groups are established from a storage system with multiple storage nodes. Each elastic group has a defined resource redundancy level of the storage system. The multiple elastic groups include multiple types of elastic groups, including at least one computing resource elastic group with computing resources of the multiple storage nodes and at least one storage resource elastic group with storage resources of multiple modular storage devices. The storage system supports the ability to configure multiple types of elastic groups, wherein each of the multiple configurations of the multiple types of elastic groups includes multiple storage resource elastic groups and multiple compute resource elastic groups; and Distributed data storage is performed by the multiple storage nodes across the multiple modular storage devices according to the multiple resilient groups established.
9. The computer-readable medium of claim 8, wherein the method further comprises: For each of the plurality of elastic groups, the defined redundancy level is selected from the group including N+1 redundancy, N+2 redundancy, and N+3 redundancy.
10. The computer-readable medium of claim 8, wherein the method further comprises: Select different defined redundancy levels from among the multiple elastic groups.
11. The computer-readable medium of claim 8, wherein the method further comprises: Different redundancy levels are defined in the first elastic group and the second elastic group.
12. A storage system comprising: Multiple modular storage devices; Multiple storage nodes, which can be configured to form multiple elastic groups, each elastic group having a defined resource redundancy level; The multiple elastic groups include various types of elastic groups, and the multiple types of elastic groups include at least one computing resource elastic group and at least one storage resource elastic group; The plurality of storage nodes are used to support multiple configurations for the plurality of elastic groups of the plurality of types, wherein each of the plurality of configurations for the plurality of elastic groups of the plurality of types includes a plurality of storage resource elastic groups and a plurality of compute resource elastic groups; and The plurality of storage nodes are configured to perform distributed data storage across the plurality of modular storage devices according to the plurality of elastic groups.
13. The storage system of claim 12, wherein the plurality of storage nodes and the plurality of modular storage devices will: Data stripes are formed using resources from the compute resource elasticity group, and written using resources from the storage resource elasticity group.
14. The storage system of claim 12, wherein in each of the plurality of elastic groups, the defined redundancy level is related to the fault survivability of the elastic group and can be selected from a group including N+1 redundancy, N+2 redundancy and N+3 redundancy.
15. The storage system of claim 12, wherein the plurality of storage nodes support configurations with different defined redundancy levels in the plurality of elastic groups.
16. The storage system of claim 12, wherein the plurality of storage nodes support heterogeneous computing resources in the at least one computing resource elastic group.
17. The storage system of claim 12, wherein the plurality of storage nodes support heterogeneous storage resources in the at least one storage resource elastic group.
18. The storage system of claim 12, wherein the plurality of storage nodes support different defined redundancy levels in a first elastic group and a second elastic group.
19. The storage system of claim 12, wherein the plurality of storage nodes support different defined redundancy levels in a first elastic group including permissions of the plurality of storage nodes and a second elastic group including other permissions of the plurality of storage nodes.
20. The storage system of claim 12, wherein the plurality of resilient groups further comprises at least one resilient group of non-volatile random access memory (NVRAM).
Citation Information
Patent Citations
Resiliency groups
US20190042407A1