Optimizing data mapping by utilizing multiple indirect unit sizes
By using direct data mapping technology, the operating system of the storage system directly manages the flash memory, solving the problem of low efficiency in traditional storage systems and achieving more efficient and reliable data storage.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- PURE STORAGE INC
- Filing Date
- 2024-09-30
- Publication Date
- 2026-05-01
AI Technical Summary
Existing storage systems suffer from inefficiency and unreliability in data storage and management, especially when using flash storage devices, where traditional storage controllers lead to unnecessary operations and wasted resources.
By using direct mapping data mapping technology, the operating system of the storage system can directly manage the flash memory, avoiding the address translation process of the traditional storage controller and optimizing the mapping and management of data in the flash memory.
It improves the performance and reliability of flash memory, reduces unnecessary write operations, extends the lifespan of the memory, and improves the efficiency and reliability of data storage.
Smart Images

Figure CN121970034A_ABST
Abstract
Description
Background Technology
[0001] Storage systems (such as enterprise storage systems) may include centralized or distributed repositories of data used to provide common data management, data protection, and data sharing functions, for example, through connections to computer systems. Attached Figure Description
[0002] Figure 1A This describes the first instance system used for data storage.
[0003] Figure 1B This describes a second instance system used for data storage.
[0004] Figure 1C This describes the third instance system used for data storage.
[0005] Figure 1D This describes the fourth instance system used for data storage.
[0006] Figure 2A It is a perspective view of a storage cluster with multiple storage nodes and internal storage devices coupled to each storage node to provide network-attached storage.
[0007] Figure 2B This is a block diagram illustrating an interconnect switch that couples multiple storage nodes according to some embodiments.
[0008] Figure 2C It is a multi-level block diagram that shows the contents of storage nodes and the contents of one of the non-volatile solid-state storage units.
[0009] Figure 2D This document illustrates a storage server environment using storage nodes and storage units from some previous diagrams, according to some embodiments.
[0010] Figure 2E It is a block diagram showing the control plane, computing and storage plane, and licensed blade hardware that interact with the underlying physical resources.
[0011] Figure 2F Describe the resilient software layer within the blades of a storage cluster.
[0012] Figure 2G Describe the licenses and storage resources in the blades of the storage cluster.
[0013] Figure 3A A diagram is provided showing a storage system coupled for data communication with a cloud service provider, according to some embodiments of the present disclosure.
[0014] Figure 3B A diagram illustrating the storage system.
[0015] Figure 3C Describe an example of a cloud-based storage system.
[0016] Figure 3D This describes an exemplary computing device that can be specifically configured to perform one or more of the processes described herein.
[0017] Figure 3E This describes an instance of a set of storage systems used to provide storage services (also referred to as 'data services' in this document).
[0018] Figure 3F Illustrate an example of a container system.
[0019] Figure 3G Examples of storage nodes for large-scale storage platforms according to embodiments of this disclosure are described.
[0020] Figure 4 This is an illustration of an example of a storage system controller that generates commands for a storage system containing data of size indirect unit (IU) according to embodiments of the present disclosure.
[0021] Figure 5 This is an illustration of an example of a storage device controller for a storage system that uses multiple IU sizes to map data stored in a flash memory, according to embodiments of the present disclosure.
[0022] Figure 6 This is an illustration of an example of a storage device controller for a storage system that uses multiple IU sizes to map subsequent data according to embodiments of the present disclosure.
[0023] Figure 7 This is an illustration of an example of a storage device controller for a storage system that uses multiple IU sizes to map subsequent data according to embodiments of the present disclosure.
[0024] Figure 8 This is an example method of storing data in a storage device using multiple IU sizes according to embodiments of the present disclosure. Detailed Implementation
[0025] by Figure 1A Beginning with reference to the accompanying drawings, an example method, apparatus, and product for optimizing data mapping by utilizing multiple indirect unit sizes according to embodiments of the present disclosure are described. Figure 1A This section describes an example system for data storage according to some embodiments. For illustrative and not limiting purposes, system 100 (also referred to herein as a "storage system") comprises numerous elements. It can be noted that in other embodiments, system 100 may comprise the same, more, or fewer elements configured in the same or different ways.
[0026] System 100 includes several computing devices 164A to B. These computing devices (also referred to herein as "client devices") may be, for example, servers, workstations, personal computers, laptops, or the like in a data center. The computing devices 164A to B may be coupled for data communication with one or more storage arrays 102A to B via a storage area network ('SAN') 158 or a local area network ('LAN') 160.
[0027] SAN 158 can be implemented using various data communication architectures, devices, and protocols. For example, architectures used for SAN 158 may include Fibre Channel, Ethernet, Unlimited Bandwidth, Serial Attached Small Computer System Interface ('SAS'), or similar. Data communication protocols used with SAN 158 may include Advanced Technology Attachment ('ATA'), Fibre Channel protocol, Small Computer System Interface ('SCSI'), Internet Small Computer System Interface ('iSCSI'), Ultra SCSI, Architecture-based Fast Non-Volatile Memory ('NVMe'), or similar. It should be noted that SAN 158 is provided for illustrative purposes and not for limitation. Other data communication couplings may be implemented between computer units 164A-B and storage arrays 102A-B.
[0028] LAN 160 can also be implemented using various architectures, devices, and protocols. For example, the architecture used for LAN 160 can include Ethernet (802.3), wireless (802.11), or similar. Data communication protocols used in LAN 160 can include Transmission Control Protocol ('TCP'), User Datagram Protocol ('UDP'), Internet Protocol ('IP'), Hypertext Transfer Protocol ('HTTP'), Wireless Access Protocol ('WAP'), Handheld Device Transfer Protocol ('HDTP'), Session Initiation Protocol ('SIP'), Real-Time Protocol ('RTP'), or similar protocols. LAN 160 can also connect to the Internet.
[0029] Storage arrays 102A to 102B provide persistent data storage for computing devices 164A to 164B. In some embodiments, storage array 102A may be housed in a chassis (not shown), and storage array 102B may be housed in another chassis (not shown). Storage arrays 102A and 102B may include one or more storage array controllers 110A to 110D (also referred to herein as "controllers"). Storage array controllers 110A to 110D may be modules embodied as automated computing machines comprising computer hardware, computer software, or a combination of computer hardware and software. In some embodiments, storage array controllers 110A to 110D may be configured to perform various storage tasks. Storage tasks may include writing data received from computing devices 164A to B to storage arrays 102A to B, erasing data from storage arrays 102A to B, retrieving data from storage arrays 102A to B and providing data to computing devices 164A to B, monitoring and reporting storage device utilization and performance, performing redundancy operations (such as independent drive redundant array ('RAID') or RAID-like data redundancy operations), compressing data, encrypting data, and so on.
[0030] Storage array controllers 110A to D can be implemented in various ways, including as field-programmable gate arrays ('FPGAs'), programmable logic chips ('PLCs'), application-specific integrated circuits ('ASICs'), monolithic systems ('SOCs'), or any computing device containing discrete components such as processing devices, central processing units, computer memory, or various adapters. Storage array controllers 110A to D may include, for example, data communication adapters configured to support communication via SAN 158 or LAN 160. In some embodiments, storage array controllers 110A to D may be independently coupled to LAN 160. In some embodiments, storage array controllers 110A to D may include I / O controllers or the like that couple the storage array controllers 110A to D for data communication to persistent storage resources 170A to B (also referred to herein as "storage resources") via an intermediate plane (not shown). Persistent storage resources 170A to B may include any number of storage drives 171A to F (also referred to herein as “storage devices”) and any number of non-volatile random access memory ('NVRAM') devices (not shown).
[0031] In some embodiments, one or more of the storage drives 171A to F may be managed flash memory devices. The managed flash memory device (which may also be referred to as a directly managed flash memory device, directly managed storage device, managed storage device, etc.) can provide external devices (e.g., processing devices of storage array controllers (e.g., storage array controllers 110A to D)) with functions, operations, commands, APIs, or other suitable mechanisms to control, manage, and / or interact with the flash memory of the managed flash memory device. This allows the storage device controller to perform fewer operations (e.g., disposal queuing, burst transfer, internal error correction, encryption, line / page voltage level adjustment for flash memory, etc.). Because the storage device can be directly managed, this allows the storage system to optimize, manage, and / or improve various aspects, characteristics, etc., of the flash memory to improve its performance, reliability, and / or lifespan, as discussed in more detail below.
[0032] In some implementations, the NVRAM devices of persistent storage resources 170A to B can be configured to receive data from storage array controllers 110A to D that will be stored in storage drives 171A to F. In some instances, the data may originate from computing devices 164A to B. In some instances, writing data to the NVRAM devices can be performed faster than writing data directly to storage drives 171A to F. In some implementations, storage array controllers 110A to D can be configured to utilize the NVRAM devices as a fast-accessible buffer for data intended to be written to storage drives 171A to F. The latency of write requests using NVRAM devices as buffers can be improved compared to systems where storage array controllers 110A to D write data directly to storage drives 171A to F. In some implementations, the NVRAM devices can be implemented using computer memory in the form of high-bandwidth, low-latency RAM. NVRAM devices are referred to as "non-volatile" because they can receive or contain the only power source that maintains the state of the RAM after the NVRAM device loses mains power. This power source can be a battery, one or more capacitors, or the like. In response to a loss of power, the NVRAM device can be configured to write the contents of the RAM to a persistent storage device, such as storage drives 171A to F.
[0033] In some embodiments, storage drives 171A to F may refer to any device configured to persistently record data, where "persistently" or "persistently" means the ability of the device to retain the recorded data after a power loss. In some embodiments, storage drives 171A to F may correspond to non-disk storage media. For example, storage drives 171A to F may be one or more solid-state drives ('SSDs'), flash memory-based storage devices, any type of solid-state non-volatile memory, or any other type of non-mechanical storage device. In other embodiments, storage drives 171A to F may include mechanical or spinning hard disks, such as hard disk drives ('HDDs').
[0034] In some embodiments, storage array controllers 110A to D may be configured to offload management responsibilities from storage drives 171A to F in storage arrays 102A to B. For example, storage array controllers 110A to D may manage control information that describes the state of one or more memory blocks in storage drives 171A to F. The control information may indicate, for example, that a particular memory block has failed and should no longer be written to, that a particular memory block contains boot code from storage array controllers 110A to D, the number of program-erase ('P / E') cycles performed on a particular memory block, the age of data stored in a particular memory block, the type of data stored in a particular memory block, and so on. In some embodiments, the control information may be stored as metadata along with the associated memory block. In other embodiments, the control information for storage drives 171A to F may be stored in one or more specific memory blocks of storage drives 171A to F selected by storage array controllers 110A to D. The selected memory blocks may be marked with identifiers indicating that the selected memory blocks contain control information. Identifiers can be used by memory array controllers 110A to D, together with memory drives 171A to F, to quickly identify memory blocks containing control information. For example, memory controllers 110A to D can issue commands to locate memory blocks containing control information. It can be noted that the control information may be so large that portions of the control information can be stored in multiple locations, that the control information can be stored in multiple locations for example for redundancy purposes, or that the control information can be otherwise distributed across multiple memory blocks in memory drives 171A to F.
[0035] In some implementations, memory array controllers 110A to D can offload management responsibilities from memory drives 171A to F of memory arrays 102A to B by retrieving control information describing the state of one or more memory blocks in memory drives 171A to F. Retrieving control information from memory drives 171A to F can be performed, for example, by memory array controllers 110A to D querying memory drives 171A to F for the location of control information for a specific memory drive 171A to F. Memory drives 171A to F can be configured to execute instructions that enable memory drives 171A to F to identify the location of the control information. These instructions can be executed by a controller (not shown) associated with or otherwise located on memory drives 171A to F and can cause memory drives 171A to F to scan a portion of each memory block to identify the memory block storing the control information for memory drives 171A to F. Storage drives 171A to F can respond by sending a response message containing the location of control information for storage drives 171A to F to storage array controllers 110A to D. Upon receiving the response message, storage array controllers 110A to D can issue a request to read data stored at the address associated with the location of the control information for storage drives 171A to F.
[0036] In other embodiments, storage array controllers 110A to D can further offload device management responsibilities from storage drives 171A to F by performing storage drive management operations in response to receiving control information. Storage drive management operations may include, for example, operations typically performed by storage drives 171A to F (e.g., controllers associated with a particular storage drive 171A to F (not shown)). Storage drive management operations may include, for example, ensuring that data is not written to faulty memory blocks within storage drives 171A to F, ensuring that data is written to memory blocks within storage drives 171A to F in a manner that achieves sufficient wear leveling, and so on.
[0037] In some implementations, storage arrays 102A to B may implement two or more storage array controllers 110A to D. For example, storage array 102A may include storage array controller 110A and storage array controller 110B. At any given time, a single storage array controller 110A to D of storage system 100 (e.g., storage array controller 110A) may be designated as having a primary state (also referred to herein as the "primary controller"), and other storage array controllers 110A to D (e.g., storage array controller 110B) may be designated as having a secondary state (also referred to herein as the "secondary controller"). The primary controller may have specific rights, such as permission to modify data in persistent storage resources 170A to B (e.g., to write data to persistent storage resources 170A to B). At least some rights of the primary controller may supersede the rights of the secondary controllers. For example, when the primary controller has permission to modify data in persistent storage resources 170A to B, the secondary controller may not have said rights. The state of storage array controllers 110A to D may be changeable. For example, storage array controller 110A may be designated to have a secondary state, and storage array controller 110B may be designated to have a primary state.
[0038] In some implementations, a primary controller (e.g., storage array controller 110A) may serve as the primary controller for one or more storage arrays 102A to B, and a secondary controller (e.g., storage array controller 110B) may serve as the secondary controller for one or more storage arrays 102A to B. For example, storage array controller 110A may be the primary controller for both storage arrays 102A and 102B, and storage array controller 110B may be the secondary controller for both storage arrays 102A and 102B. In some implementations, storage array controllers 110C and 110D (also referred to as "storage processing modules") may not have a primary or secondary state. Storage array controllers 110C and 110D, implemented as storage processing modules, may serve as communication interfaces between the primary and secondary controllers (e.g., storage array controllers 110A and 110B, respectively) and storage array 102B. For example, storage array controller 110A of storage array 102A may send write requests to storage array 102B via SAN 158. Write requests can be received by both storage array controllers 110C and 110D of storage array 102B. Storage array controllers 110C and 110D facilitate communication, for example, sending write requests to the appropriate storage drives 171A to F. It can be noted that in some embodiments, the storage processing module can be used to increase the number of storage drives controlled by the primary and secondary controllers.
[0039] In some implementations, memory array controllers 110A to D are communicatively coupled to one or more memory drives 171A to F and to one or more NVRAM devices (not shown) included as part of memory arrays 102A to B via an intermediate plane (not shown). Memory array controllers 110A to D may be coupled to the intermediate plane via one or more data communication links, and the intermediate plane may be coupled to the memory drives 171A to F and the NVRAM devices via one or more data communication links. For example, the data communication links described herein are collectively illustrated by data communication links 108A to D and may include a Fast Peripheral Component Interconnect ('PCIe') bus.
[0040] Figure 1B This describes an example system for data storage based on some implementation schemes. Figure 1B The storage array controller 101 described herein may be similar to that described above. Figure 1A The described storage array controllers are 110A to D. In one example, storage array controller 101 may be similar to storage array controller 110A or storage array controller 110B. For illustrative and not limiting purposes, storage array controller 101 includes numerous elements. It can be noted that in other embodiments, storage array controller 101 may include the same, more, or fewer elements configured in the same or different ways. It can be noted that the following may include... Figure 1A The components are described to help illustrate the features of the storage array controller 101.
[0041] The storage array controller 101 may include one or more processing devices 104 and random access memory ('RAM') 111. The processing device 104 (controller 101) represents one or more general-purpose processing devices, such as a microprocessor, central processing unit, or the like. More specifically, the processing device 104 (or controller 101) may be a Complex Instruction Set Computing ('CISC') microprocessor, a Reduced Instruction Set Computing ('RISC') microprocessor, a Very Long Instruction Word ('VLIW') microprocessor, or a processor implementing other instruction sets, or several processors implementing combinations of instruction sets. The processing device 104 (controller 101) may also be one or more special-purpose processing devices, such as an ASIC, FPGA, digital signal processor ('DSP'), network processor, or the like.
[0042] Processing device 104 can be connected to RAM 111 via data communication link 106, which can be embodied as a high-speed memory bus, such as a Double Data Rate 4 ('DDR4') bus. Operating system 112 is stored in RAM 111. In some embodiments, instructions 113 are stored in RAM 111. Instructions 113 may include computer program instructions for performing operations in a direct-mapped flash memory storage system. In one embodiment, a direct-mapped flash memory storage system is a system that directly addresses data blocks within a flash drive without requiring address translation performed by the flash drive's memory controller.
[0043] In some embodiments, the storage array controller 101 includes one or more host bus adapters 103A to C coupled to the processing device 104 via data communication links 105A to C. In some embodiments, the host bus adapters 103A to C may be computer hardware that connects a host system (e.g., the storage array controller) to other networks and storage arrays. In some instances, the host bus adapters 103A to C may be Fibre Channel adapters enabling the storage array controller 101 to connect to a SAN, Ethernet adapters enabling the storage array controller 101 to connect to a LAN, or the like. The host bus adapters 103A to C may be coupled to the processing device 104 via data communication links 105A to C (e.g., a PCIe bus).
[0044] In some implementations, the storage array controller 101 may include a host bus adapter 114 coupled to an extender 115. The extender 115 can be used to attach a host system to a larger number of storage drives. In implementations where the host bus adapter 114 is embodied as a SAS controller, the extender 115 may be, for example, a SAS extender for enabling the host bus adapter 114 to be attached to storage drives.
[0045] In some implementations, the storage array controller 101 may include a switch 116 coupled to the processing device 104 via a data communication link 109. The switch 116 may be a computer hardware device capable of creating multiple endpoints from a single endpoint, thereby enabling multiple devices to share a single endpoint. The switch 116 may be, for example, a PCIe switch coupled to a PCIe bus (e.g., data communication link 109) and presenting multiple PCIe connection points to an intermediate plane.
[0046] In some implementations, the storage array controller 101 includes a data communication link 107 for coupling the storage array controller 101 to other storage array controllers. In some instances, the data communication link 107 may be a Fast Path Interconnect (QPI) interconnect.
[0047] Conventional storage systems using conventional flash drives can implement processes across the flash drive, which is part of the conventional storage system. For example, higher-level processes of the storage system can be initiated and controlled across the flash drive. However, the flash drive of a conventional storage system may contain its own storage controller that also performs the processes. Therefore, for a conventional storage system, both higher-level processes (e.g., initiated by the storage system) and lower-level processes (e.g., initiated by the storage system's storage controller) can be performed.
[0048] To address the various shortcomings of traditional storage systems, operations can be performed by higher-level processes rather than lower-level ones. For example, a flash storage system can contain flash drives that do not include a storage controller providing the processes. Therefore, the operating system of the flash storage system itself can initiate and control the processes. This can be accomplished by a direct-mapped flash storage system that directly addresses data blocks within the flash drive without requiring address translation performed by the flash drive's storage controller.
[0049] In some embodiments, storage drives 171A to F may be one or more partitioned storage devices. In some embodiments, one or more partitioned storage devices may be shingled HDDs. In some embodiments, one or more storage devices may be flash-based SSDs. In a partitioned storage device, the partition namespace on the partitioned storage device may be addressed by block groups, which are grouped and aligned according to their natural size to form several addressable regions. In some embodiments utilizing SSDs, the natural size may be based on the SSD's erase block size. In some embodiments, the regions of the partitioned storage device may be defined during the initialization of the partitioned storage device. In some embodiments, the regions may be dynamically defined as data is written to the partitioned storage device.
[0050] In some implementations, zones may be heterogeneous, with some zones each being a page group and others being multiple page groups. In some implementations, some zones may correspond to a single erase block, and others may correspond to multiple erase blocks. In implementations, for heterogeneous mixtures of programming models, manufacturers, product types, and / or product generations suitable for storage devices such as heterogeneous assembly, updates, distributed storage, etc., zones may be any combination of different numbers of pages in units of page groups and / or erase blocks. In some implementations, a zone may be defined as having usage characteristics, such as the ability to support data with a specific kind of lifetime (e.g., very short lifetime or very long lifetime). These characteristics may be used by the partitioned storage device to determine how the zone will be managed during its expected lifetime.
[0051] It should be understood that a region is a virtual construct. Any particular region may not have a fixed location on the storage device. Before allocation, a region may not have any location on the storage device. A region may correspond to a number, representing a block of virtual allocatable space, which may be the size of an erase block or, in various embodiments, another block size. When the system allocates or opens a region, the region is allocated to flash memory or other solid-state storage, and as the system writes to the region, pages are written to that mapped flash memory or other solid-state storage of the partitioned storage device. When the system closes the region, the associated erase block or other block of size is completed. At some point in the future, the system may delete the region, which will release the allocated space of the region. During its lifetime, for example, when the partitioned storage device undergoes internal maintenance, a region may be moved to a different location on the partitioned storage device.
[0052] In some implementations, zones of a partitioned storage device can be in different states. A zone can be in an empty state, where data has not yet been stored in that zone. An empty zone can be explicitly or implicitly opened by writing data to it. This is the initial state of a zone on a newly partitioned storage device, but can also be the result of a zone reset. In some implementations, empty zones may have a designated location within the flash memory of the partitioned storage device. In some implementations, the location of the zone can be selected when the empty zone is first opened or first written to (or later if the write is buffered in memory). A zone can be explicitly or implicitly in an open state, where a zone in an open state can be written to store data using write or append commands. In some implementations, a zone in an open state can also be written to using copy commands that copy data from different zones. In some implementations, the partitioned storage device may have a limit on the number of open zones at a given time.
[0053] A closed region is a region that has been partially written to but entered the closed state after an explicit close operation was issued. A closed region can be left available for future writes, but this reduces some of the runtime overhead incurred by keeping the region open. In some implementations, the partitioned storage device may limit the number of closed regions at a given time. A full region is a region that is storing data and can no longer be written to. A region may be full after a write operation has filled the entire region or due to a region completion operation. Before completion, the region may have been or may not have been fully written to. However, after completion, without first performing a region reset operation, the region may not be open for further writes.
[0054] The mapping from a region to an erase block (or to a shingled track in an HDD) can be arbitrary, dynamic, and hidden. Opening a region can be an operation that allows a new region to be dynamically mapped to the underlying storage of the partitioned storage device, followed by writing data to the region via append writes until the region reaches its capacity. A region can be completed at any time, after which additional data may not be able to be written to it. When the data stored at the region is no longer needed, the region can be reset, effectively removing the region's contents from the partitioned storage device, thus making the physical storage held by that region available for subsequent data storage. Once a region has been written to and completed, the partitioned storage device ensures that the data stored at the region is not lost until the region is reset. During the time between writing data to a region and resetting the region, the region can move between shingled tracks or erase blocks as part of maintenance operations within the partitioned storage device, such as copying data to keep it refreshed or disposing of aging memory cells in an SSD.
[0055] In some implementations utilizing HDDs, resetting a region allows the allocation of shingled tracks to new open regions that can be opened at a future time. In some implementations utilizing SSDs, resetting a region can cause the region's associated physical erase block to be erased and subsequently reused for data storage. In some implementations, partitioned storage devices may limit the number of open regions at a given point in time to reduce the number of open regions dedicated to keeping regions open.
[0056] The operating system of a flash memory storage system can identify and maintain a list of allocation units across multiple flash drives in the flash memory storage system. An allocation unit can be all erase blocks or multiple erase blocks. The operating system can maintain a mapping or address range that directly maps addresses to erase blocks in the flash drives of the flash memory storage system.
[0057] Erasable blocks directly mapped to the flash drive can be used to rewrite and erase data. For example, an operation can be performed on one or more allocation units containing first and second data, where the first data will be retained and the second data is no longer used by the flash storage system. The operating system can initiate a process to write the first data to a new location within another allocation unit and erase the second data, marking the allocation unit as available for subsequent data. Therefore, the process can be executed only by the higher-level operating system of the flash storage system; additional lower-level processes do not need to be executed by the flash drive controller.
[0058] The advantage of having the process executed solely by the operating system of the flash memory storage system includes improved reliability of the flash drives, as no unnecessary or redundant write operations are performed during the process. A potentially novel aspect here is the concept of initiating and controlling the process at the operating system level of the flash memory storage system. Furthermore, the process can be controlled by the operating system across multiple flash drives. This contrasts with processes executed by the storage controllers of the flash drives.
[0059] The storage system may consist of two storage array controllers that share a set of drives for failover purposes, or it may consist of a single storage array controller that provides storage services using multiple drives, or it may consist of a distributed network of storage array controllers, each having a certain number of drives or a certain amount of flash storage, wherein the storage array controllers in the network cooperate to provide complete storage services and cooperate in all aspects of storage services, including storage allocation and collection of discarded items.
[0060] Figure 1C This section describes a third instance system 117 for data storage according to some embodiments. For illustrative and not limiting purposes, system 117 (also referred to herein as a “storage system”) comprises numerous elements. It can be noted that in other embodiments, system 117 may comprise the same, more, or fewer elements configured in the same or different ways.
[0061] In one embodiment, system 117 includes a dual peripheral component interconnect ('PCI') flash memory device 118 having individually addressable fast write memory. System 117 may include a memory device controller 119. In one embodiment, memory device controllers 119A to D may be a CPU, ASIC, FPGA, or any other circuit system capable of implementing the necessary control structures according to this disclosure. In one embodiment, system 117 includes flash memory devices (e.g., flash memory devices 120a to n) operatively coupled to various channels of memory device controller 119. Flash memory devices 120a to n may be presented to controllers 119A to D as an addressable set of flash pages, erase blocks, and / or control elements sufficient to allow memory device controllers 119A to D to program and retrieve various aspects of the flash memory. In one embodiment, storage device controllers 119A to D can perform operations on flash memory devices 120a to n, including storing and retrieving data content of pages, arranging and erasing any blocks, tracking statistics related to the use and reuse of flash memory pages, erase blocks and cells, tracking and predicting error codes and faults in the flash memory, controlling and programming and retrieving voltage levels associated with the contents of flash memory cells, etc.
[0062] In one embodiment, system 117 may include RAM 121 to store individually addressable, fast-write data. In one embodiment, RAM 121 may be one or more separate discrete devices. In another embodiment, RAM 121 may be integrated into storage device controllers 119A to D or more storage device controllers. RAM 121 may also be used for other purposes, such as as temporary program memory for processing devices (e.g., CPU) in storage device controller 119.
[0063] In one embodiment, system 117 may include an energy storage device 122, such as a rechargeable battery or capacitor. The energy storage device 122 may store enough energy to power storage device controller 119, a certain amount of RAM (e.g., RAM 121), and a certain amount of flash memory (e.g., flash memory 120a to 120n) for a sufficient time to write the contents of RAM to the flash memory. In one embodiment, if the storage device controller detects a loss of external power, then storage device controllers 119A to D may write the contents of RAM to the flash memory.
[0064] In one embodiment, system 117 includes two data communication links 123a and 123b. In one embodiment, data communication links 123a and 123b may be PCI interfaces. In another embodiment, data communication links 123a and 123b may be based on other communication standards (e.g., HyperTransport, Infinite Bandwidth, etc.). Data communication links 123a and 123b may be based on Fast Non-Volatile Memory ('NVMe') or Architecture-based NVMe ('NVMf') specifications, which allow external connection from other components in storage system 117 to storage device controllers 119A through D. It should be noted that, for convenience, data communication links may be interchangeably referred to herein as PCI buses.
[0065] System 117 may also include an external power supply (not shown), which may be provided via one or two data communication links 123a, 123b, or may be provided separately. Alternative embodiments include a separate flash memory (not shown) dedicated to storing the contents of RAM 121. Storage device controllers 119A to D may present a different portion of the logical address space of a logic device (which may include addressable fast write logic) or storage device 118 (which may be presented as PCI memory or as persistent storage) via a PCI bus. In one embodiment, operations to store in the device are directed to RAM 121. In the event of a power failure, storage device controllers 119A to D may write the stored content associated with the addressable fast write logic to flash memory (e.g., flash memory 120a to n) for long-term persistent storage.
[0066] In one embodiment, the logic device may include some or all of the contents of flash memory devices 120a to n, wherein the presentation allows a storage system (e.g., storage system 117) including storage device 118 to directly address flash memory pages and directly reprogram erase blocks from storage system components external to the storage device via the PCI bus. The presentation may also allow one or more external components to control and retrieve other aspects of the flash memory, including some or all of the following: tracking statistics related to the use and reuse of flash memory pages, erase blocks, and cells across all flash memory devices; tracking and predicting error codes and faults within and across flash memory devices; controlling voltage levels associated with programming and retrieving the contents of flash memory cells; etc.
[0067] In one embodiment, energy storage device 122 may be sufficient to ensure the completion of ongoing operations on flash memory devices 120a to 120n. Energy storage device 122 may power storage device controllers 119A to D and associated flash memory devices (e.g., 120a to n) for those operations, as well as for storing fast write RAM into flash memory. Energy storage device 122 may be used to store accumulated statistics and other parameters maintained and tracked by flash memory devices 120a to n and / or storage device controllers 119. Individual capacitors or energy storage devices (e.g., smaller capacitors located near or embedded within the flash memory devices themselves) may be used for some or all of the operations described herein.
[0068] Various methods can be used to track and optimize the lifespan of energy storage components, such as adjusting voltage levels over time and partially discharging the energy storage device 122 to measure the corresponding discharge characteristics. If available energy decreases over time, the effective available capacity of the addressable fast write storage device can be reduced to ensure that it can be safely written to based on the currently available stored energy.
[0069] Figure 1D This describes a third instance storage system 124 for data storage according to some embodiments. In one embodiment, storage system 124 includes storage controllers 125a, 125b. In one embodiment, storage controllers 125a, 125b are operatively coupled to dual PCI storage devices. Storage controllers 125a, 125b are operatively coupled (e.g., via storage network 130) to a number of host computers 127a to n.
[0070] In one embodiment, two storage controllers (e.g., 125a and 125b) provide storage services, such as SCS block storage arrays, file servers, object servers, databases, or data analytics services. Storage controllers 125a and 125b can provide services to host computers 127a to n outside the storage system 124 via a number of network interfaces (e.g., 126a to d). Storage controllers 125a and 125b can provide integrated services or applications entirely within the storage system 124, thus forming a converged storage and computing system. Storage controllers 125a and 125b can utilize fast write memory within or across storage devices 119a to d to record ongoing operations to ensure no loss of operation in the event of power failure, storage controller removal, storage controller or storage system shutdown, or failure of one or more software or hardware components within the storage system 124.
[0071] In one embodiment, storage controllers 125a and 125b operate as a PCI master of one or more PCI buses 128a and 128b. In another embodiment, 128a and 128b may be based on other communication standards (e.g., HyperTransport, Infinite Bandwidth, etc.). Other storage system embodiments may operate storage controllers 125a and 125b as multiple masters of both PCI buses 128a and 128b. Alternatively, a PCI / NVMe / NVMe switching infrastructure or architecture may connect multiple storage controllers. Some storage system embodiments may allow storage devices to communicate directly with each other, rather than only with storage controllers. In one embodiment, storage device controller 119a may operate under the guidance of storage controller 125a to access data stored in RAM (e.g., ...). Figure 1C The data in RAM 121 is recombined and transferred to the flash memory device. For example, the recombined version of the RAM content may be transferred after the memory controller has determined that the operation has been fully committed across the memory system, or when the fast write memory on the device has reached a certain usage capacity, or after a certain amount of time, to ensure improved data security or free up addressable fast write capacity for reuse. This mechanism can be used, for example, to avoid secondary transfers from memory controllers 125a, 125b via buses (e.g., 128a, 128b). In one embodiment, recombining may include compressing data, adding indexes or other metadata, combining multiple data segments together, performing erasure coding calculations, etc.
[0072] In one embodiment, under the guidance of storage controllers 125a, 125b, storage device controllers 119a, 119b are operable to retrieve data from RAM (e.g., stored in RAM). Figure 1CThe data in RAM 121) is processed and transferred to other storage devices without involving storage controllers 125a and 125b. This operation can be used to mirror data stored in one storage controller 125a to another storage controller 125b, or it can be used to offload compression, data aggregation, and / or erasure coding calculations and transfers to storage devices to reduce the load on the storage controllers or the storage controller interfaces 129a and 129b to the PCI buses 128a and 128b.
[0073] Storage device controllers 119A to D may include mechanisms for implementing high availability primitives for use by other components of the storage system outside the dual PCI storage device 118. For example, reserve or exclusion primitives may be provided, allowing one storage controller in a storage system with two storage controllers providing highly available storage services to prevent the other storage controller from accessing or continuing to access the storage device. This could be used, for example, in situations where one controller detects that the other controller is not functioning correctly or that the interconnect between the two storage controllers itself may not be functioning correctly.
[0074] In one embodiment, a storage system used with dual PCI direct-mapped storage devices having separate addressable fast-write storage includes a system for managing erase blocks or erase block groups as allocation units for storing data on behalf of the storage service, storing metadata associated with the storage service (e.g., indexes, logs, etc.), or for the proper management of the storage system itself. Flash pages, which can be several kilobytes in size, can be written when data arrives or when the storage system will retain the data for a long time interval (e.g., exceeding a defined time threshold). To commit data faster or to reduce the number of writes to the flash memory device, the storage controller may first write data to separate addressable fast-write storage devices on one or more storage devices.
[0075] In one embodiment, storage controllers 125a and 125b may initiate the use of erase blocks within and across storage devices (e.g., 118) based on the age and expected remaining lifespan of the storage device or based on other statistics. Storage controllers 125a and 125b may also initiate obsolete item collection and data migration between storage devices based on no longer needed pages and for managing the lifespan of flash memory pages and erase blocks, as well as for managing overall system performance.
[0076] In one embodiment, storage system 124 may utilize mirroring and / or erasure coding schemes as part of storing data into an addressable, fast-write storage device and / or as part of writing data into an allocation unit associated with an erase block. Erasure codes may be used across storage devices and within erase blocks or allocation units, or within and across flash memory devices on a single storage device, to provide redundancy against failures of single or multiple storage devices or to protect them from internal damage to flash memory pages caused by flash memory operation or flash memory cell degradation. Mirroring and erasure coding at various levels can be used for recovery from multiple types of failures occurring individually or in combination.
[0077] refer to Figure 2A The embodiments described in G illustrate a storage cluster for storing user data, such as user data originating from one or more user or client systems or from other sources outside the storage cluster. The storage cluster uses erasure coding of metadata and redundant copies to distribute user data across storage nodes housed within a single chassis or across multiple chassis. Erasure coding refers to a data protection or reconstruction method where data is stored across a set of different locations (e.g., disks, storage nodes, or geographical locations). Flash memory is one type of solid-state memory that can be integrated with the embodiments, but the embodiments can be extended to other types of solid-state memory or other storage media, including non-solid-state memory. Control of storage location and workload distribution across the cluster peer system. For example, tasks such as mediating communication between storage nodes, detecting when a storage node becomes unavailable, and balancing I / O (input and output) across storage nodes are all handled on a distributed basis. In some embodiments, data is laid out or distributed across multiple storage nodes in data segments or stripes that support data recovery. Data ownership can be reassigned within the cluster, independent of input and output formats. This architecture, described in more detail below, allows the system to remain operational even if a storage node in the cluster fails, because data can be reconstructed from other storage nodes and thus remains available for input and output operations. In various embodiments, the storage node may be referred to as a cluster node, blade, or server.
[0078] Storage clusters can be housed within a chassis (i.e., a enclosure housing one or more storage nodes). Mechanisms for supplying power to each storage node (e.g., a power distribution bus) and communication mechanisms (e.g., a communication bus enabling communication between storage nodes) are contained within the chassis. According to some embodiments, the storage cluster can operate as a standalone system in one location. In one embodiment, the chassis houses at least two examples of both the power distribution and communication buses, which can be individually enabled or disabled. The internal communication bus can be an Ethernet bus; however, other technologies such as PCIe, wireless bandwidth, and others are equally suitable. The chassis provides ports for external communication buses for communication between multiple chassis and with client systems, either directly or via switches. External communication can use technologies such as Ethernet, wireless bandwidth, Fibre Channel, etc. In some embodiments, the external communication buses use different communication bus technologies for inter-chassis and client communication. If switches are deployed within or between chassis, the switches can act as a translator between multiple protocols or technologies. When multiple chassis are connected to define a storage cluster, the storage cluster can be accessed by clients using proprietary or standard interfaces such as Network File System ('NFS'), Common Internet File System ('CIFS'), Small Computer System Interface ('SCSI'), or Hypertext Transfer Protocol ('HTTP')). Protocol translation from the client can occur at a switch, on the chassis external communication bus, or within each storage node. In some embodiments, multiple chassis can be coupled or connected to each other via an aggregator switch. A portion and / or all of the coupled or connected chassis can be designated as a storage cluster. As discussed above, each chassis may have multiple blades, each blade having a Media Access Control ('MAC') address; however, in some embodiments, the storage cluster presents itself to the external network as having a single cluster IP address and a single MAC address.
[0079] Each storage node may be one or more storage servers, and each storage server is connected to one or more non-volatile solid-state memory cells, which may be referred to as storage cells or storage devices. One embodiment includes a single storage server in each storage node and among 1 to 8 non-volatile solid-state memory cells; however, this single example is not intended to be limiting. The storage server may include a processor, DRAM, and an interface for internal communication buses and power distribution for each of the power buses. In some embodiments, within a storage node, interfaces and storage cells share a communication bus, such as a PCI Express. Non-volatile solid-state memory cells can directly access the internal communication bus interface via the storage node's communication bus, or request access to the bus interface from the storage node. The non-volatile solid-state memory cell contains an embedded CPU, a solid-state storage controller, and a number of solid-state high-capacity storage devices, for example, between 2 and 32 terabytes ('TB') in some embodiments. Embedded volatile storage media (e.g., DRAM) and energy storage devices are included in the non-volatile solid-state memory cell. In some embodiments, the energy storage device is a capacitor, supercapacitor, or battery that enables the transfer of a subset of DRAM content to a stable storage medium in the event of a power loss. In some embodiments, the non-volatile solid-state memory cell is configured to have a storage-type memory, such as a phase-change or magnetoresistive random access memory ('MRAM') that replaces DRAM and enables power-reduced retention devices.
[0080] One of the many characteristics of storage nodes and non-volatile solid-state storage devices (NSSSDs) is the ability to proactively reconstruct data within a storage cluster. Storage nodes and NSSSDs can determine when a storage node or NSSSD in the cluster is unreachable, independent of whether an attempt is made to read data involving that storage node or NSSSD. The storage nodes and NSSSDs then cooperate to recover and reconstruct the data in at least a partial new location. This constitutes proactive reconstruction because the system does not need to wait until a read access initiated from a client system employing the storage cluster requires the data before reconstructing it. These and other details regarding storage and its operation are discussed below.
[0081] Figure 2AThis is a perspective view of a storage cluster 161 according to some embodiments, having a plurality of storage nodes 150 and internal solid-state memory coupled to each storage node to provide a network-attached storage device or a storage area network. A network-attached storage device, storage area network, or storage cluster or other storage memory may comprise one or more storage clusters 161, each storage cluster having one or more storage nodes 150, arranging both the physical components and the amount of storage memory provided in a flexible and reconfigurable manner. Storage clusters 161 are designed to be mounted in racks, and one or more racks can be configured and filled as needed for storage memory. Storage cluster 161 has a chassis 138 with a plurality of slots 142. It should be understood that chassis 138 may be referred to as a housing, enclosure, or rack unit. In one embodiment, chassis 138 has fourteen slots 142, but other numbers of slots can be easily designed. For example, some embodiments have 4 slots, 8 slots, 16 slots, 32 slots, or other suitable numbers of slots. In some embodiments, each slot 142 may accommodate one storage node 150. Chassis 138 includes fins 148 for mounting chassis 138 onto a rack. Fan 144 provides air circulation for cooling storage nodes 150 and their components, but other cooling components may be used, or embodiments may be designed without cooling components. Switch architecture 146 couples the storage nodes 150 together within chassis 138 and to a network for communication with the memory. In the embodiment depicted herein, for illustrative purposes, slot 142 to the left of switch architecture 146 and fan 144 is shown as occupied by storage nodes 150, while slot 142 to the right of switch architecture 146 and fan 144 is empty and available for insertion of storage nodes 150. This configuration is one example, and in various other arrangements, one or more storage nodes 150 may occupy slot 142. In some embodiments, the storage nodes need not be arranged sequentially or adjacently. Storage node 150 is hot-swappable, meaning that storage node 150 can be inserted into or removed from slot 142 in chassis 138 without stopping or powering down the system. After insertion or removal of storage node 150 from slot 142, the system automatically reconfigures to recognize and adapt to the change. In some embodiments, reconfiguration includes restoring redundancy and / or rebalancing data or load.
[0082] Each storage node 150 may have multiple components. In the embodiment shown herein, storage node 150 includes a printed circuit board 159 filled with a CPU 156 (i.e., a processor), a memory 154 coupled to the CPU 156, and a non-volatile solid-state storage device 152 coupled to the CPU 156; however, in other embodiments, other mounting components and / or components may be used. The memory 154 has instructions executed by the CPU 156 and / or data operated on by the CPU 156. As explained further below, the non-volatile solid-state storage device 152 includes flash memory, or in other embodiments, other types of solid-state memory. In some embodiments, the non-volatile solid-state storage device 152 may include one or more managed flash memory devices, as previously described.
[0083] refer to Figure 2A The storage cluster 161 is scalable, meaning that storage capacity with non-uniform storage sizes can be easily added, as described above. One or more storage nodes 150 can be inserted into or removed from each chassis, and in some embodiments, the storage cluster is self-configurable. Insertable storage nodes 150 can have different sizes, whether installed in the chassis at delivery or added later. For example, in one embodiment, storage nodes 150 can have any multiple of 4TB, such as 8TB, 12TB, 16TB, 32TB, etc. In other embodiments, storage nodes 150 can have any other multiple of storage amount or capacity. The storage capacity of each storage node 150 is broadcast and influences decisions on how data is striped. For maximum storage efficiency, embodiments can self-configure as wide as possible in stripes, subject to predetermined continuous operation requirements, wherein up to one or more non-volatile solid-state storage device 152 cells or storage nodes 150 are lost within the chassis.
[0084] Figure 2B This is a block diagram illustrating the communication interconnect 173 and power distribution bus 172 that couple multiple storage nodes 150. (Return to Reference) Figure 2A In some embodiments, the communication interconnect 173 may be included in or implemented with the switch architecture 146. In cases where multiple storage clusters 161 occupy a rack, in some embodiments, the communication interconnect 173 may be included in or implemented with the top of a rack switch. Figure 2B The description states that storage cluster 161 is enclosed within a single chassis 138. External port 176 is coupled to storage node 150 via communication interconnect 173, while external port 174 is directly coupled to the storage node. External power port 178 is coupled to power distribution bus 172. Storage node 150 may contain varying amounts and capacities of non-volatile solid-state storage devices 152, as described in reference [reference missing]. Figure 2ADescription. Additionally, one or more storage nodes 150 may be compute-only storage nodes, such as... Figure 2B The description continues. Authorization 168 is implemented on non-volatile solid-state storage device 152, for example, as a list or other data structure stored in memory. In some embodiments, authorization is stored within non-volatile solid-state storage device 152 and supported by software executing on the controller or other processor of non-volatile solid-state storage device 152. In another embodiment, authorization 168 is implemented on storage node 150, for example, as a list or other data structure stored in memory 154 and supported by software executing on the CPU 156 of storage node 150. In some embodiments, authorization 168 controls how data is stored in non-volatile solid-state storage device 152 and the location where data is stored in non-volatile solid-state storage device 152. This control helps determine which type of erasure coding scheme is applied to the data and which storage nodes 150 have which portions of the data. Each authorization 168 may be assigned to non-volatile solid-state storage device 152. In various embodiments, each authorization can control the range of inode numbers, segment numbers, or other data identifiers assigned to data by the file system, storage node 150, or non-volatile solid-state storage device 152.
[0085] In some embodiments, each piece of data and each piece of metadata is redundant in the system. Additionally, each piece of data and each piece of metadata has an owner, which may be referred to as an authorizer. If that authorizer is unreachable, for example, due to a storage node failure, then there is a successor plan for how to locate that data or that metadata. In various embodiments, redundant copies of authorizer 168 exist. In some embodiments, authorizer 168 is associated with storage node 150 and non-volatile solid-state storage device 152. Each authorizer 168 covering a range of data segment numbers or other identifiers may be assigned to a specific non-volatile solid-state storage device 152. In some embodiments, authorizers 168 for all such ranges are distributed across the non-volatile solid-state storage devices 152 of the storage cluster. Each storage node 150 has a network port providing access to the non-volatile solid-state storage device 152 of that storage node 150. Data may be segmented, and in some embodiments, the segment is associated with a segment number, and that segment number is an indirection of the RAID (Redundant Array of Independent Disks) stripe configuration. The assignment and use of authorizer 168 thus establishes indirection to the data. According to some embodiments, indirection may be referred to as the ability to indirectly (in this case, via authorization 168) reference data. A segment identifies a set of non-volatile solid-state storage devices 152 and a local identifier within that set of non-volatile solid-state storage devices 152, which may contain data. In some embodiments, the local identifier is an offset within the device and can be reused sequentially across multiple segments. In other embodiments, the local identifier is unique to a particular segment and is never reused. Offsets in the non-volatile solid-state storage devices 152 are applied to locate data for writing to or reading from the non-volatile solid-state storage devices 152 (in the form of RAID striping). Data is striped across multiple cells of the non-volatile solid-state storage devices 152, which may include non-volatile solid-state storage devices 152 with authorization 168 for a particular data segment or different from the non-volatile solid-state storage devices 152.
[0086] If the location of a specific data segment changes, for example, during data movement or data reconstruction, the authorization 168 for that data segment should be consulted at the non-volatile solid-state storage device 152 or storage node 150 that has that authorization 168. To locate a specific piece of data, embodiments calculate a hash value of the data segment or apply an inode number or data segment number. The output of this operation points to the non-volatile solid-state storage device 152 that has the authorization 168 for that specific data. In some embodiments, this operation has two phases. The first phase maps an entity identifier (ID) (e.g., a segment number, inode number, or directory number) to an authorization identifier. This mapping may include, for example, the calculation of a hash or bitmask. The second phase is mapping the authorization identifier to a specific non-volatile solid-state storage device 152, which can be done through explicit mapping. The operation is repeatable, such that when the calculation is performed, the result of the calculation reliably and repeatedly points to the specific non-volatile solid-state storage device 152 that has that authorization 168. The operation may include a set of reachable storage nodes as input. If the set of reachable non-volatile solid-state storage cells changes, then the optimal set changes. In some embodiments, the held value is the current assignment (which is always true) and the calculated value is the target assignment that the cluster will attempt to reconfigure. This calculation can be used to determine the optimal non-volatile solid-state storage device 152 to be authorized when there is a set of non-volatile solid-state storage devices 152 that are reachable and constitute the same cluster. The calculation also determines an ordered set of peer non-volatile solid-state storage devices 152, which also records the authorizations for the mapping of non-volatile solid-state storage devices, so that authorizations can be determined even when the assigned non-volatile solid-state storage device is unreachable. In some embodiments, if a particular authorization 168 is unavailable, then duplicate or alternative authorizations 168 can be consulted.
[0087] refer to Figure 2A and 2BTwo of the many tasks performed by the CPU 156 on storage node 150 are decomposing written data and reassembling read data. When the system determines that data will be written, the authorization 168 for that data is positioned as described above. When the segment ID of the data has been determined, the write request is forwarded to the non-volatile solid-state storage device 152 of the host currently identified as the authorization 168 from which the segment was determined. The non-volatile solid-state storage device 152 and the host CPU 156 of the storage node 150 on which the corresponding authorization 168 resides then decompose or slice the data and transmit the data to various non-volatile solid-state storage devices 152. The transmitted data is written as data stripes according to an erasure coding scheme. In some embodiments, data is requested to be pulled, and in other embodiments, data is pushed. Conversely, when data is read, the authorization 168 for the segment ID containing the data is positioned as described above. The host CPU 156 of the storage node 150, on which the non-volatile solid-state storage device 152 and the corresponding license 168 reside, requests data from the non-volatile solid-state storage device and the corresponding storage node pointed to by the license. In some embodiments, the data is read as a data stripe from the flash memory device. The host CPU 156 of the storage node 150 then reassembles the read data, correcting any errors (if any) according to an appropriate erasure coding scheme and forwarding the reassembled data to the network. In other embodiments, some or all of these tasks may be handled within the non-volatile solid-state storage device 152. In some embodiments, a segment host requests data to be sent to the storage node 150 by requesting a page from the storage device and then sending the data to the storage node that issued the original request.
[0088] In this embodiment, authorization 168 operates to determine how the operation will be performed on a specific logical element. Each of the logical elements can be operated on via specific authorization from multiple storage controllers across the storage system. Authorization 168 can communicate with multiple storage controllers, causing the multiple storage controllers to jointly perform operations on those specific logical elements.
[0089] In embodiments, a logical element may be, for example, a file, directory, object bucket, individual object, a description portion of a file or object, other forms of key-value database or table. In embodiments, performing operations may involve, for example, ensuring the consistency, structural integrity, and / or recoverability of other operations against the same logical element, reading metadata and data associated with that logical element, determining what data should be persistently written to the storage system to preserve any changes to the operations, or determining where the metadata and data will be stored across modular storage devices attached to multiple storage controllers in the storage system.
[0090] In some embodiments, operations are token-based transactions efficiently transmitted within a distributed system. Each transaction may be accompanied by or associated with a token, which grants permission to execute the transaction. In some embodiments, authorization 168 can maintain the system's pre-transaction state until the operation is complete. Token-based communication can be completed without global locking across the system and can also restart operations in the event of corruption or other failures.
[0091] In some systems, such as UNIX-style file systems, data is handled using index nodes (or inodes), which are data structures that represent objects within the file system. For example, an object can be a file or a directory. Metadata may accompany an object as attributes, such as permission data, creation timestamps, and other properties. Segment numbers can be assigned to all or part of this object within the file system. In other systems, data segments are handled using segment numbers assigned elsewhere. For illustrative purposes, a unit of distribution is an entity, and an entity can be a file, directory, or segment. That is, an entity is a unit of data or metadata stored by the storage system. Entities are grouped into sets called licenses. Each license has a license owner, who is a storage node with exclusive rights to update entities within the license. In other words, a storage node contains licenses, and a license contains entities.
[0092] According to some embodiments, a segment is a logical container for data. A segment is an address space between a media address space and a physical flash memory location; that is, the data segment number resides in this address space. A segment may also contain metadata that enables data redundancy recovery (rewriting to different flash memory locations or devices) without involving higher-level software. In one embodiment, the internal format of a segment contains client data and a media mapping to determine the location of that data. Where applicable, the segments are protected, for example, from the effects of memory and other failures, by breaking each data segment down into several data and parity fragments. The data and parity fragments are distributed (i.e., striped) across the non-volatile solid-state storage device 152 coupled to the host CPU 156 according to an erasure coding scheme (see [reference]). Figure 2E and 2G In some embodiments, the term "segment" refers to a container and its location in the segment's address space. According to some embodiments, the term "strip" refers to the same set of fragments as a segment, including how fragmentation, redundancy, or parity information is distributed.
[0093] A series of address space translations occur across the entire storage system. At the top are directory entries (filenames) linked to inodes. Inodes point to the media address space where data is logically stored. Media addresses can be mapped through a series of indirect media mappings to extend the load of large files or implement data services such as deduplication or snapshots. Next, segment addresses are translated to physical flash memory locations. According to some embodiments, physical flash memory locations have address ranges defined by the amount of flash memory in the system. Media addresses and segment addresses are logical containers and, in some embodiments, use 128-bit or larger identifiers for practical infinity, where the likelihood of reuse is calculated to be longer than the expected lifespan of the system. In some embodiments, addresses from logical containers are allocated hierarchically. Initially, each non-volatile solid-state storage device 152 cell can be assigned an address space range. Within this assigned range, the non-volatile solid-state storage device 152 can allocate addresses without synchronization with other non-volatile solid-state storage devices 152.
[0094] Data and metadata are stored using a set of underlying storage layouts optimized for different workload types and storage devices. These layouts incorporate various redundancy schemes, compression formats, and indexing algorithms. Some layouts store information about licenses and licensees, while others store file metadata and file data. Redundancy schemes include error correction codes that tolerate corrupted bits within a single storage device (e.g., NAND flash memory chips), erasure codes that tolerate failures of multiple storage nodes, and replication schemes that tolerate data center or region failures. In some embodiments, low-density parity-check ('LDPC') codes are used within a single storage cell. In some embodiments, Reed-Solomon coding is used within the storage cluster, and mirroring is used within the storage grid. Metadata can be stored using an ordered log structure index (e.g., a log structure merge tree), and large amounts of data may not be stored in the log structure layout.
[0095] To maintain consistency across multiple replicas of an entity, storage nodes implicitly agree on two things through computation: (1) the authorization containing the entity, and (2) the storage node containing the authorization. Entity-to-authority assignment can be accomplished by pseudo-randomly assigning entities to authorizations, by splitting entities into ranges based on externally generated keys, or by placing individual entities into each authorization. Instances of pseudo-random schemes are hashes of the 'RUSH' family under linear hashing and scalable hashing, including controlled replication under scalable hashing ('CRUSH'). In some embodiments, pseudo-random assignment is used only to assign authorizations to nodes because the set of nodes can change. The set of authorizations cannot change, so any subjective function can be applied in these embodiments. Some placement schemes automatically place authorizations on storage nodes, while others rely on an explicit mapping of authorizations to storage nodes. In some embodiments, a pseudo-random scheme is used to map from each authorization to a set of candidate authorization owners. A pseudo-random data distribution function associated with CRUSH can assign authorizations to storage nodes and create a list of where authorizations are assigned. Each storage node has a copy of the pseudo-random data distribution function and yields the same computation for distribution and subsequent lookup or location of grants. In some embodiments, each of the pseudo-random schemes requires a set of reachable storage nodes as input to arrive at the same target node. Once an entity has been placed in a grant, it can be stored on physical devices such that anticipated failures will not result in unexpected data loss. In some embodiments, the rebalancing algorithm attempts to store copies of all entities within a grant in the same layout and on the same set of machines.
[0096] Examples of anticipated failures include equipment malfunctions, machine theft, data center fires, and regional disasters such as nuclear or geological events. Different failures result in varying degrees of acceptable data loss. In some embodiments, theft of storage nodes does not affect the security or reliability of the system; depending on the system configuration, regional events may result in no data loss, loss of updates within seconds or minutes, or even complete data loss.
[0097] In these embodiments, the placement of redundant data is independent of the placement of authorizations for data consistency. In some embodiments, the authorized storage nodes do not contain any persistent storage devices. Instead, the storage nodes are connected to non-volatile solid-state storage units without authorizations. The communication interconnects between the storage nodes and the non-volatile solid-state storage units consist of various communication technologies and have non-uniform performance and fault tolerance characteristics. In some embodiments, as mentioned above, the non-volatile solid-state storage units are quickly connected to the storage nodes via PCI, the storage nodes are connected together in a single chassis using an Ethernet backplane, and the chassis are connected together to form a storage cluster. In some embodiments, the storage clusters are connected to clients using Ethernet or Fibre Channel. If multiple storage clusters are configured as a storage grid, then the multiple storage clusters are connected using the Internet or other long-distance networking links (e.g., "metro-scale" links or private links that do not cross the Internet).
[0098] The authorized owner has exclusive rights to modify entities, migrate entities from one non-volatile solid-state storage unit to another, and add and remove copies of entities. This allows for the maintenance of redundancy in the underlying data. When the authorized owner fails, is deactivated, or becomes overloaded, the authorization is transferred to a new storage node. Transient failures make it critical to ensure that all fault-free machines agree on the new authorized location. Ambiguity arising from transient failures can be resolved automatically via consensus protocols (such as Paxos) or hot-swap schemes through manual intervention by a remote system administrator or a local hardware administrator (e.g., by physically removing the failed machine from the cluster or pressing a button on the failed machine). In some embodiments, a consensus protocol is used, and failover is automatic. According to some embodiments, if too many failures or replication events occur within a short period, the system enters a self-protection mode and stops replication and data movement activities until administrator intervention.
[0099] When authorization is transferred between storage nodes and the authorization owner to update an entity's authorization, the system transmits messages between the storage nodes and non-volatile solid-state storage units. Regarding persistent messages, messages with different purposes have different types. Depending on the message type, the system maintains different ordering and durability guarantees. When processing persistent messages, messages are temporarily stored using multiple persistent and non-persistent storage hardware technologies. In some embodiments, messages are stored in RAM, NVRAM, and NAND flash memory devices, using various protocols to efficiently utilize each storage medium. Latency-sensitive client requests may be held in replicated NVRAM and then later in NAND, while background rebalancing operations are held directly to NAND.
[0100] Persistent messages are persistently stored before being transmitted. This allows the system to continue serving client requests regardless of failures and component replacements. Although many hardware components contain unique identifiers visible to system administrators, manufacturers, the hardware supply chain, and ongoing monitoring and quality control infrastructure, applications running on these infrastructure addresses virtualize those addresses. These virtualized addresses do not change throughout the storage system's lifetime, regardless of component failures or replacements. This allows for the replacement of every component of the storage system over time without reconfiguration or disruption of client request processing; that is, the system supports non-destructive upgrades.
[0101] In some embodiments, virtualized addresses are stored with sufficient redundancy. The continuous monitoring system associates hardware and software status with hardware identifiers. This allows for the detection and prediction of failures due to faulty components and manufacturing details. In some embodiments, the monitoring system is also capable of proactively relocating authorizations and entities away from affected devices before a failure occurs by removing components from the critical path.
[0102] Figure 2C This is a multi-level block diagram illustrating the contents of storage node 150 and the contents of its non-volatile solid-state storage devices 152. In some embodiments, data is transferred to and from storage node 150 by a network interface controller ('NIC') 202. Each storage node 150 has a CPU 156 and one or more non-volatile solid-state storage devices 152, as discussed above. Figure 2C Moving down one level, each non-volatile solid-state storage device 152 has relatively fast non-volatile solid-state memory, such as non-volatile random access memory ('NVRAM') 204 and flash memory 206. In some embodiments, NVRAM 204 may be a component that does not require programming / erasing cycles (DRAM, MRAM, PCM) and may be a memory capable of supporting writes much more frequently than reads. Figure 2CMoving down another level, in one embodiment, NVRAM 204 is implemented as a high-speed volatile memory, such as dynamic random access memory (DRAM) 216, backed up by energy reserve 218. Energy reserve 218 provides sufficient power to keep DRAM 216 powered on long enough to transfer content to flash memory 206 in the event of a power failure. In some embodiments, energy reserve 218 is a capacitor, supercapacitor, battery, or other device that provides a suitable energy supply sufficient to transfer the content of DRAM 216 to a stable storage medium in the event of a power failure. Flash memory 206 is implemented as a plurality of flash dies 222, which may be referred to as a flash die 222 package or an array of flash die 222. It should be understood that flash dies 222 can be packaged in any number of ways, including single die per package, multiple dies per package (i.e., multi-chip package), hybrid packages, as exposed dies on a printed circuit board or other substrate, as encapsulated dies, etc. In the illustrated embodiment, the non-volatile solid-state storage device 152 has a controller 212 or other processor and an input / output (I / O) port 210 coupled to the controller 212. The I / O port 210 is coupled to the CPU 156 and / or network interface controller 202 of the flash memory storage node 150. A flash input / output (I / O) port 220 is coupled to a flash memory die 222, and a direct memory access unit (DMA) 214 is coupled to the controller 212, DRAM 216, and flash memory die 222. In the illustrated embodiment, the I / O port 210, controller 212, DMA unit 214, and flash I / O port 220 are implemented on a programmable logic device ('PLD') 208 (e.g., an FPGA). In this embodiment, each flash memory die 222 has pages organized as 16 kB (kilobyte) pages 224 and registers 226 through which data can be written to or read from the flash memory die 222. In another embodiment, other types of solid-state memory are used to replace or supplement the flash memory described within flash memory die 222.
[0103] In the various embodiments disclosed herein, storage cluster 161 is generally contrasted with a storage array. Storage nodes 150 are part of the collection that creates storage cluster 161. Each storage node 150 has a data slice and the computation required to provide said data. Multiple storage nodes 150 cooperate to store and retrieve data. As is typically used in a storage array, storage memories or storage devices are less involved in processing and manipulating data. Storage memories or storage devices in a storage array receive commands to read, write, or erase data. Storage memories or storage devices in a storage array are unaware of the larger system in which they are embedded, or what the data means. Storage memories or storage devices in a storage array can include various types of storage memories, such as RAM, solid-state drives, hard disk drives, etc. The non-volatile solid-state storage device 152 unit described herein has multiple interfaces that are simultaneously active and serve multiple purposes. In some embodiments, a function of storage node 150 is shifted into storage cell 152, thereby transforming storage cell 152 into a combination of storage cell 152 and storage node 150. Placing computation (relative to storing data) in storage unit 152 brings that computation closer to the data itself. Various system embodiments have hierarchical layers of storage nodes with varying capabilities. In contrast, in a storage array, the controller possesses and understands everything about all the data managed by the controller in a shelf or storage device. In storage cluster 161, as described herein, multiple controllers in multiple non-volatile solid-state storage device units 152 and / or storage nodes 150 cooperate in various ways (e.g., for erasure coding, data sharding, metadata communication and redundancy, storage capacity expansion or contraction, data recovery, etc.).
[0104] Figure 2D Demonstration and use Figure 2A The storage server environment of the embodiment of storage node 150 and storage device 152 unit to C. In this version, each non-volatile solid-state storage device 152 unit has a chassis 138 (see Figure 2A For example, controller 212 on a PCIe (Fast Peripheral Component Interconnect) board (see...) Figure 2C The processor, FPGA, flash memory 206, and NVRAM 204 (which is supercapacitor-supported DRAM 216, see [link]) Figure 2B and 2C The non-volatile solid-state storage device 152 unit can be implemented as a single board containing storage devices and can be the maximum tolerable fault domain within the chassis. In some embodiments, up to two non-volatile solid-state storage device 152 units may fail and the device will continue to function without data loss.
[0105] In some embodiments, the physical storage device is divided into named regions based on application usage. NVRAM 204 is a contiguous block of memory reserved in DRAM 216 of the non-volatile solid-state storage device 152 and is powered by NAND flash memory. NVRAM 204 is logically divided into multiple memory regions, written as two spools (e.g., spool_region). The space within the NVRAM 204 spool is managed independently by each license 168. Each device provides a certain amount of storage space to each license 168. That license 168 further manages the lifetime and allocation within that space. Instances of spools encompass distributed transactions or concepts. When a mains power failure occurs to the non-volatile solid-state storage device 152 cell, an onboard supercapacitor provides power retention for a short duration. During this retention interval, the contents of NVRAM 204 are flushed to flash memory 206. Upon the next power-on, the contents of NVRAM 204 are restored from flash memory 206.
[0106] Regarding the memory cell controller, the responsibilities of the logical "controller" are distributed across each of the 168 authorized blades. This logical control is distributed across... Figure 2D The diagram shows a host controller 242, an intermediate layer controller 244, and a storage unit controller 246. The control plane and storage plane are managed independently, but the components can be physically co-located on the same blade. Each license 168 effectively functions as an independent controller. Each license 168 provides its own data and metadata structure, its own background workers, and maintains its own lifecycle.
[0107] Figure 2E It is displayed in Figure 2D Use in storage server environment Figure 2A This is a hardware block diagram of the control plane 254, compute and storage planes 256, 258, and blade 252 of the embodiment of storage node 150 and storage unit 152 interacting with the underlying physical resources. The control plane 254 is partitioned into several licenses 168 that can run on any of the blades 252 using the compute resources in the compute plane 256. The storage plane 258 is partitioned into a group of devices, each of which provides access to flash memory 206 and NVRAM 204 resources. In one embodiment, the compute plane 256 can perform the operation of a storage array controller on one or more devices of the storage plane 258 (e.g., a storage array), as described herein.
[0108] exist Figure 2EIn the compute and storage planes 256 and 258, license 168 interacts with the underlying physical resources (i.e., devices). From the viewpoint of license 168, its resources are striped across all physical devices. From the viewpoint of the device, it provides resources to all licenses 168, regardless of where the license happens to be operating. Each license 168 has been allocated or has been allocated one or more partitions 260 of storage memory in storage unit 152, such as partitions 260 in flash memory 206 and NVRAM 204. Each license 168 uses its allocated partitions 260 for writing or reading user data. Licenses can be associated with different amounts of physical storage in the system. For example, a license 168 may have a larger number or larger size of partitions 260 in one or more storage units 152 than one or more other licenses 168.
[0109] Figure 2F A resilient software layer is depicted in blade 252 of a storage cluster according to some embodiments. In the resilient architecture, the resilient software is symmetrical, i.e., the compute module 270 of each blade runs... Figure 2F The process described herein consists of three identical layers. Storage manager 274 executes read and write requests from other blades 252 for data and metadata stored in local storage unit 152 NVRAM 204 and flash memory 206. Authorizer 168 fulfills client requests by issuing the necessary read and write requests to the blade 252 where the corresponding data or metadata resides in its storage unit 152. Endpoint 272 parses client connection requests received from the switch architecture 146 monitoring software, relays the client connection requests to the responsible authorizer 168, and relays the response from authorizer 168 to the client. This symmetrical three-tier architecture enables high concurrency in the storage system. In these embodiments, elastic, efficient, and reliable horizontal scaling is achieved. Furthermore, elastic implementation of a unique horizontal scaling technique evenly balances work across all resources, regardless of client access type, and maximizes concurrency by eliminating most of the inter-blade coordination required for conventional distributed locking.
[0110] Still referencing Figure 2FThe licensee 168, running within the compute module 270 of blade 252, performs the internal operations required to fulfill the client request. A characteristic of this resilience is that the licensee 168 is stateless; that is, it caches valid data and metadata in the DRAM of its own blade 252 for fast access, but stores each update in partitions of its NVRAM 204 on the three separate blades 252 until the update has been written to flash memory 206. In some embodiments, all storage system writes to the NVRAM 204 are triplicated to the partitions on the three separate blades 252. Utilizing triple-mirrorized NVRAM 204 and persistent storage protected by parity and Reed-Solomon RAID checksums, the storage system can withstand concurrent failures of two blades 252 without loss of data, metadata, or access to either.
[0111] Because license 168 is stateless, it can be migrated between blades 252. Each license 168 has a unique identifier. In some embodiments, partitions of NVRAM 204 and flash memory 206 are associated with the identifier of license 168, not with the blade 252 in which it operates. Therefore, when license 168 migrates, it continues to manage the same storage partitions from its new location. When a new blade 252 is installed in an embodiment of the storage cluster, the system automatically rebalances the load by partitioning the storage device of the new blade 252 for use by the system's licenses 168, migrating the selected licenses 168 to the new blade 252, enabling endpoint 272 on the new blade 252, and including it in the client connection distribution algorithm of the switch architecture 146.
[0112] The migrated license 168 retains the contents of its NVRAM 204 partition on flash memory 206 from its new location, processes read and write requests from other licenses 168, and completes client requests directed to it by endpoint 272. Similarly, if blade 252 fails or is removed, the system redistributes its license 168 among the remaining blades 252 in the system. The redistributed license 168 continues to perform its original functions from its new location.
[0113] Figure 2GThis describes licenses 168 and storage resources in blades 252 of a storage cluster according to some embodiments. Each license 168 is specifically responsible for partitioning the flash memory 206 and NVRAM 204 on each blade 252. License 168 manages the contents and integrity of its partitions independently of other licenses 168. License 168 compresses incoming data and temporarily stores it in its NVRAM 204 partition, and then merges, RAID-protects, and stores the data in segments of the storage device in its flash memory 206 partition. While license 168 writes data to flash memory 206, storage manager 274 performs necessary flash translation to optimize write performance and maximize media lifetime. In the background, license 168 "performs obsolete item collection" or reclaims space occupied by data that has become obsolete by clients overwriting data. It should be understood that because the partitions of license 168 are disjoint, distributed locking is not required for client and write or background functions.
[0114] The embodiments described herein can utilize various software, communication, and / or networking protocols. Furthermore, the hardware and / or software configuration can be adapted to accommodate various protocols. For example, embodiments may utilize Active Directory, which is available in Windows. TM Database-based systems provide authentication, directories, policies, and other services in the environment. In these embodiments, LDAP (Lightweight Directory Access Protocol) is an instance application protocol used to query and modify items in a directory service provider (such as Active Directory). In some embodiments, a Network Lock Manager ('NLM') is used as a facility that works with the Network File System ('NFS') to provide System V-style document and record locking over the network. The Server Message Block ('SMB') protocol (one version of which is also known as the Universal Internet File System ('CIFS')) can be integrated with the storage systems discussed herein. SMB operation is an application-layer network protocol commonly used to provide shared access to files, printers, and serial ports, as well as miscellaneous communication between nodes on a network. SMB also provides an authenticated inter-process communication mechanism. AMAZON TMS3 (Simple Storage Service) is a web service provided by Amazon Web Services, and the systems described herein can interface with Amazon S3 via web service interfaces (REST (Representative State Transfer), SOAP (Simple Object Access Protocol), and BitTorrent). The RESTful API (Application Programming Interface) breaks down transactions into a series of small modules. Each module addresses a specific underlying part of the transaction. Controls or permissions provided by these embodiments, particularly for object data, may include the use of Access Control Lists ('ACLs'). An ACL is a list of permissions attached to an object, specifying which users or system processes are authorized to access the object and what operations are allowed on a given object. The system may utilize Internet Protocol version 6 ('IPv6') and IPv4 for communication protocols that provide identification and location systems for computers on the network and for routing services across the Internet. Packet routing between networked systems may include Equal Cost Multipath ('ECMP'), a routing strategy where next-hop packet forwarding to a single destination can occur via multiple "best paths" that are at the top of the routing metric calculation. Multipath routing can be used in conjunction with most routing protocols because it is limited to per-hop decisions by a single router. The software can support multi-tenancy, an architecture where a single instance of a software application serves multiple clients. Each client can be called a tenant. In some embodiments, tenants may be given the ability to customize parts of the application, but may not be able to customize the application's code. Implementations can maintain audit logs. Audit logs are documents that record events in a computing system. In addition to recording what resources were accessed, audit log entries typically include destination and source addresses, timestamps, and user login information to comply with various regulations. Implementations can support various key management strategies, such as cryptographic key rotation. Additionally, the system can support dynamic root passwords or dynamic changes to a password.
[0115] Figure 3A A diagram is provided showing a storage system 306 coupled for data communication with a cloud service provider 302, according to some embodiments of this disclosure. Although depicted in limited detail, Figure 3A The storage system 306 described herein may be similar to the one referenced above. Figures 1A to 1D and Figures 2A to 2G The described storage system. In some embodiments, Figure 3AThe storage system 306 described herein may be embodied as a storage system including an unbalanced active / active controller, a storage system including a balanced active / active controller, a storage system including an active / active controller (where not all resources of each controller are utilized, such that each controller has reserved resources available to support failover), a storage system including a fully active / active controller, a storage system including a dataset isolation controller, a storage system including a two-tier architecture with a front-end controller and a back-end integrated storage controller, a storage system including a scale-out cluster with a dual-controller array, and combinations of such embodiments.
[0116] exist Figure 3A In the example depicted, storage system 306 is coupled to cloud service provider 302 via data communication link 304. This data communication link 304 may be a fully wired, fully wireless, or some aggregation of wired and wireless data communication paths. In this example, digital information may be exchanged between storage system 306 and cloud service provider 302 via data communication link 304 using one or more data communication protocols. For example, digital information may be exchanged between storage system 306 and cloud service provider 302 via data communication link 304 using Handheld Device Transfer Protocol ('HDTP'), Hypertext Transfer Protocol ('HTTP'), Internet Protocol ('IP'), Real-Time Transport Protocol ('RTP'), Transmission Control Protocol ('TCP'), User Datagram Protocol ('UDP'), Wireless Application Protocol ('WAP'), or other protocols.
[0117] For example, Figure 3A The cloud service provider 302 described herein can be embodied as a system and computing environment that provides a large number of services to users of the cloud service provider 302 by sharing computing resources via data communication link 304. The cloud service provider 302 can provide on-demand access to a pool of shared configurable computing resources (such as computer networks, servers, storage devices, applications and services, etc.).
[0118] exist Figure 3A In the example described, cloud service provider 302 can be configured to provide various services to storage system 306 and its users by implementing various service models. For example, cloud service provider 302 can be configured to provide services by implementing an Infrastructure as a Service ('IaaS') service model, by implementing a Platform as a Service ('PaaS') service model, by implementing a Software as a Service ('SaaS') service model, by implementing an Authentication as a Service ('AaaS') service model, by implementing a Storage as a Service model (whereby cloud service provider 302 provides access to its storage infrastructure for use by storage system 306 and its users), and so on.
[0119] exist Figure 3A In the examples depicted, cloud service provider 302 may be embodied as, for example, a private cloud, a public cloud, or a combination of private and public clouds. In embodiments where cloud service provider 302 is embodied as a private cloud, cloud service provider 302 may be dedicated to providing services to a single organization rather than to multiple organizations. In embodiments where cloud service provider 302 is embodied as a public cloud, cloud service provider 302 may provide services to multiple organizations. In alternative embodiments, cloud service provider 302 may be embodied as a hybrid of private and public cloud services with a hybrid cloud deployment.
[0120] although Figure 3A While not explicitly described, readers will understand that significant additional hardware and software components may be required to facilitate the delivery of cloud services to storage system 306 and its users. For example, storage system 306 may be coupled to (or even contain) a cloud storage gateway. This cloud storage gateway may be manifested as a hardware-based or software-based device deployed on-premises with storage system 306. This cloud storage gateway can operate as a bridge between local applications running on storage system 306 and remote cloud-based storage devices utilized by storage system 306. By using a cloud storage gateway, organizations can move their primary iSCSI or NAS to cloud service provider 302, thereby saving space on their on-premises storage systems. This cloud storage gateway may be configured to emulate disk arrays, block-based devices, file servers, or other storage systems that can translate SCSI commands, file server commands, or other appropriate commands into REST space protocols that facilitate communication with cloud service provider 302.
[0121] To enable storage system 306 and its users to use services provided by cloud service provider 302, a cloud migration process can be performed. During this process, data, applications, or other elements from the organization's on-premises systems (or even from another cloud environment) are moved to cloud service provider 302. To successfully migrate data, applications, or other elements to the environment of cloud service provider 302, middleware such as cloud migration tools can be used to bridge the gap between the environment of cloud service provider 302 and the organization's environment. To further enable storage system 306 and its users to use services provided by cloud service provider 302, a cloud orchestrator can be used to deploy and coordinate automated tasks to create merged processes or workflows. This cloud orchestrator can perform tasks such as configuring various components, whether cloud or on-premises, and managing the interconnections between such components.
[0122] exist Figure 3AIn the examples depicted, and as briefly described above, cloud service provider 302 may be configured to provide services to storage system 306 and its users using a SaaS service model. For instance, cloud service provider 302 may be configured to provide storage system 306 and its users with access to data analytics applications. Such data analytics applications may be configured, for example, to receive large amounts of telemetry data transmitted back from storage system 306. This telemetry data can describe various operational characteristics of storage system 306 and can be analyzed for a wide range of purposes, including, for example, determining the health of storage system 306, identifying workloads performed on storage system 306, predicting when storage system 306 will run out of resources, recommending configuration changes, hardware or software upgrades, workflow migrations, or other actions that can improve the operation of storage system 306.
[0123] The cloud service provider 302 can also be configured to provide access to the virtualized computing environment to the storage system 306 and its users. Instances of such virtualized environments may include virtual machines created to emulate physical computers, virtualized desktop environments that separate logical desktops from physical machines, virtualized file systems that allow unified access to different types of specific file systems, and many others.
[0124] although Figure 3A The example depicted illustrates that storage system 306 is coupled for data communication with cloud service provider 302. However, in other embodiments, storage system 306 may be part of a hybrid cloud deployment, where private cloud elements (e.g., private cloud services, on-premises infrastructure, etc.) and public cloud elements (e.g., public cloud services, infrastructure, etc., that may be provided by one or more cloud service providers) are combined to form a single solution orchestrated across various platforms. This hybrid cloud deployment may utilize hybrid cloud management software, such as (for example) from Microsoft. TM Azure TM Arc centralizes the management of hybrid cloud deployments across any infrastructure and enables service deployments anywhere. In this example, the hybrid cloud management software can be configured to create, update, and delete resources (both physical and virtual) that form a hybrid cloud deployment, allocate compute and storage to specific workloads, monitor workload and resource performance, policy compliance, updates and patches, security status, or perform various other tasks.
[0125] Readers will understand that a variety of products can be achieved by pairing the storage systems described herein with one or more cloud service providers. For example, in embodiments where cloud resources are used to protect applications and data from damage caused by disasters, and where the storage system included can be used as the primary data store, Disaster Recovery as a Service ('DRaaS') can be provided. In such embodiments, comprehensive system backups can be performed, allowing business continuity to be maintained in the event of system failure. In such embodiments, cloud data backup technologies (either independently or as part of a larger DRaaS solution) can also be integrated into the overall solution that includes the storage systems and cloud service providers described herein.
[0126] The storage systems and cloud service providers described in this article can be used to provide a wide range of security features. For example, storage systems can encrypt data at rest (and the data can be sent to and from the encrypted storage system) and can use Key Management as a Service ('KMaaS') to manage encryption keys, keys for locking and unlocking storage devices, and so on. Similarly, cloud data security gateways or similar mechanisms can be used to ensure that data stored within a storage system is not ultimately stored incorrectly in the cloud, as part of cloud data backup operations. Furthermore, micro-segmentation or identity-based segmentation can be used within the data center containing the storage system or within the cloud service provider to create security zones in data center and cloud deployments, allowing workloads to be isolated from each other.
[0127] To further explain, Figure 3B Figures are shown of a storage system 306 according to some embodiments of the present disclosure. Although depicted in little detail, Figure 3B The storage system 306 described herein may be similar to the one referenced above. Figures 1A to 1D and Figures 2A to 2G The storage system described is because a storage system can contain many of the components described above.
[0128] Figure 3BThe storage system 306 depicted may include a large number of storage resources 308, which may be embodied in many forms. For example, storage resources 308 may include nano-RAM or another form of non-volatile random access memory utilizing carbon nanotubes deposited on a substrate, 3D cross-point non-volatile memory, flash memory (including single-level cell ('SLC') NAND flash memory with one bit of data per cell, multi-level cell ('MLC') NAND flash memory with two bits of data per cell, three-level cell ('TLC') NAND flash memory with three bits of data per cell, four-level cell ('QLC') NAND flash memory with four bits of data per cell, five-level cell ('PLC') NAND flash memory with five bits of data per cell, or other programming modes of flash memory with different numbers of bits of data per cell). Similarly, storage resources 308 may include non-volatile magnetoresistive random access memory ('MRAM'), including spin-transfer torque ('STT') MRAM. Example storage resource 308 may alternatively include non-volatile phase-change memory ('PCM'), quantum memory that allows the storage and retrieval of photonic quantum information, resistive random access memory ('ReRAM'), storage-class memory ('SCM'), or other forms of storage resources, including any combination of the resources described herein. The reader will understand that other forms of computer memory and storage devices may be utilized by the storage systems described above, including DRAM, SRAM, EEPROM, general-purpose memory, and many others. Figure 3A The storage resource 308 depicted can be embodied in various form factors, including but not limited to dual in-line memory modules ('DIMM'), non-volatile dual in-line memory modules ('NVDIMM'), M.2, U.2, and others. In some embodiments, the storage resource 308 may include one or more managed flash memory devices, as previously described.
[0129] Figure 3BThe storage resource 308 described herein may comprise various forms of SCM. An SCM effectively treats fast non-volatile memory (e.g., NAND flash) as an extension of DRAM, allowing the entire dataset to be viewed as a dataset entirely residing in DRAM. The SCM may comprise non-volatile media, such as (for example) NAND flash. This NAND flash can be accessed using NVMe, which can use the PCIe bus for its transmission, thus providing relatively low access latency compared to older protocols. In practice, network protocols used for SSDs in all-flash arrays may include NVMe using Ethernet (ROCE, NVMe TCP), Fibre Channel (NVMe FC), Infinite Bandwidth (iWARP), and others that make it possible to treat fast non-volatile memory as an extension of DRAM. Given that DRAM is typically byte-addressable and fast non-volatile memory (e.g., NAND flash) is block-addressable, a controller software / hardware stack may be required to translate block data into bytes stored in the media. Examples of media and software that can be used as SCMs include, for example, 3D XPoint, Intel memory drive technology, Samsung's Z-SSD, and others.
[0130] Figure 3B The storage resource 308 depicted may also include raceway memory (also known as domain-wall memory). This raceway memory can take the form of non-volatile solid-state memory, which, in addition to the charge of electrons, depends on the intrinsic strength and orientation of the magnetic field created by electrons in their spin within the solid-state device. By moving magnetic domains along nanopermalloy wires using a spin-coherent current, the domains can be transferred through the wires by a magnetic read / write head positioned near the wires, thus altering the domains to record bit patterns. To create a raceway memory device, many such wires and read / write elements can be packaged together.
[0131] Figure 3B The instance storage system 306 described herein can implement various storage architectures. For example, storage systems according to some embodiments of this disclosure may utilize block storage devices, where data is stored in blocks, and each block is essentially used as an individual hard drive. Storage systems according to some embodiments of this disclosure may utilize object storage devices, where data is managed as objects. Each object may include the data itself, variable metadata, and a globally unique identifier, wherein object storage may be implemented at multiple levels (e.g., device level, system level, interface level). Storage systems according to some embodiments of this disclosure utilize file storage, where data is stored in a hierarchical structure. This data may be stored in files and folders and presented to both the system storing it and the system retrieving it in the same format.
[0132] Figure 3B The instance storage system 306 described herein can be embodied as a storage system in which additional storage resources can be added using a vertical scaling model, additional storage resources can be added using a horizontal scaling model, or a combination thereof. In the vertical scaling model, additional storage is added by adding additional storage devices. However, in the horizontal scaling model, additional storage nodes can be added to a cluster of storage nodes, where such storage nodes may contain additional processing resources, additional networking resources, and so on.
[0133] Figure 3B The instance storage system 306 described above can utilize the storage resources described above in various ways. For example, a portion of the storage resources can be used as a write cache, storage resources within the storage system can be used as a read cache, or tiering can be implemented within the storage system by placing data within the storage system according to one or more tiering strategies.
[0134] Figure 3B The storage system 306 depicted also includes communication resources 310, which can be used to facilitate data communication between components within the storage system 306 and data communication between the storage system 306 and computing devices outside the storage system 306, including embodiments where those resources are separated by relatively large areas. Communication resources 310 can be configured to utilize various different protocols and data communication architectures to facilitate data communication between components within the storage system and computing devices outside the storage system. For example, communication resources 310 may include: Fibre Channel ('FC') technology, such as an FC architecture and FC protocol that can transmit SCSI commands via an FC network; Ethernet-based FC ('FCoE') technology, through which FC frames are encapsulated and transmitted via an Ethernet network; Infinite Bandwidth ('IB') technology, in which a switching architecture topology is used to facilitate transmission between channel adapters; NVM Fast ('NVMe') technology and architecture-based NVMe ('NVMeoF') technology, through which non-volatile storage media attached via a PCI Fast ('PCIe') bus can be accessed; and others. In fact, the storage system described above can directly or indirectly use neutrino communication technology and devices to transmit information (including binary information) using neutrino beams.
[0135] Communication resources 310 may also include mechanisms for accessing storage resources 308 within storage system 306 using Serial Attached SCSI ('SAS'), a Serial ATA ('SATA') bus interface for connecting storage resources 308 within storage system 306 to a host bus adapter within storage system 306, Internet Minicomputer System Interface ('iSCSI') technology for providing block-level access to storage resources 308 within storage system 306, and other communication resources that can be used to facilitate data communication between components within storage system 306 and data communication between storage system 306 and computing devices outside storage system 306.
[0136] Figure 3B The storage system 306 depicted also includes processing resources 312 that can be used to execute computer program instructions and perform other computational tasks within the storage system 306. Processing resources 312 may include one or more ASICs and one or more CPUs customized for a particular purpose. Processing resources 312 may also include one or more DSPs, one or more FPGAs, one or more system-on-a-chip ('SoC') or other forms of processing resources 312. The storage system 306 can utilize processing resources 312 to perform various tasks, including, but not limited to, supporting the execution of software resources 314, which will be described in more detail below.
[0137] Figure 3B The storage system 306 depicted also includes software resources 314, which, when executed by processing resources 312 within the storage system 306, can perform a wide range of tasks. Software resources 314 may include, for example, one or more computer program instruction modules, which, when executed by processing resources 312 within the storage system 306, are used to implement various data protection technologies. Such data protection technologies may be implemented, for example, by system software running on the computer hardware within the storage system, by a cloud service provider, or otherwise. These data protection technologies may include data archiving, data backup, data replication, data snapshots, data and database cloning, and other data protection technologies.
[0138] Software resource 314 may also include software for implementing software-defined storage ('SDS'). In this example, software resource 314 may include one or more computer program instruction modules that, when executed, are used for policy-based deployment and management of data storage, independent of the underlying hardware. Such software resource 314 can be used to implement storage virtualization to decouple storage hardware from the software that manages the storage hardware.
[0139] Software resource 314 may also include software for facilitating and optimizing I / O operations routed to storage system 306. For example, software resource 314 may include software modules that perform various data reduction techniques (e.g., data compression, data deduplication, and others). Software resource 314 may include software modules that intelligently group I / O operations to facilitate better use of the underlying storage resource 308, software modules that perform data migration operations to migrate data from within the storage system, and software modules that perform other functions. Such software resource 314 may be embodied as one or more software containers or in many other ways.
[0140] To further explain, Figure 3C Examples of a cloud-based storage system 318 according to some embodiments of this disclosure are presented. Figure 3C In the example described, the cloud-based storage system 318 is entirely built within a cloud computing environment 316, such as (for example) Amazon Web Services ('AWS'). TM Microsoft Azure TM Google Cloud Platform TM IBM Cloud TM Oracle Cloud TM And others. Cloud-based storage system 318 can be used to provide services similar to those provided by the storage system described above.
[0141] Figure 3C The cloud-based storage system 318 described herein includes two cloud computing examples 320 and 322, each used to support the execution of storage controller applications 324 and 326. For example, cloud computing examples 320 and 322 may be embodied as examples of cloud computing resources (e.g., virtual machines) provided by a cloud computing environment 316 to support the execution of software applications (e.g., storage controller applications 324 and 326). For instance, each of cloud computing examples 320 and 322 may execute on an Azure VM, where each Azure VM may contain a high-speed temporary storage device that can be used as a cache (e.g., as a read cache). In one embodiment, cloud computing examples 320 and 322 may be embodied as Amazon Elastic Compute Cloud ('EC2') examples. In this instance, an Amazon Machine Image ('AMI') containing storage controller applications 324 and 326 may be initiated to create and configure virtual machines capable of executing storage controller applications 324 and 326.
[0142] exist Figure 3CIn the example methods described above, the storage controller applications 324 and 326 can be embodied as computer program instruction modules, which, when executed, perform various storage tasks. For example, the storage controller applications 324 and 326 can be embodied as computer program instruction modules, which, when executed, perform tasks as described above. Figure 1A The controllers 110A and 110B perform the same tasks, such as writing data to and from the cloud-based storage system 318, erasing data from and from the cloud-based storage system 318, retrieving data from and from the cloud-based storage system 318, monitoring and reporting storage device utilization and performance, performing redundancy operations (e.g., RAID or RAID-like data redundancy operations), compressing data, encrypting data, deduplicating data, etc. The reader will understand that because there are two cloud computing examples 320 and 322, each containing storage controller applications 324 and 326, in some embodiments, one cloud computing example 320 may operate as the primary controller as described above, while the other cloud computing example 322 may operate as the secondary controller as described above. The reader will understand that... Figure 3C The storage controller applications 324 and 326 described herein may contain the same source code that executes within different cloud computing examples 320 and 322 (e.g., different EC2 examples).
[0143] Readers will understand that other embodiments not including primary and secondary controllers are within the scope of this disclosure. For example, each cloud computing example 320, 322 may operate as a primary controller for a portion of the address space supported by the cloud-based storage system 318, each cloud computing example 320, 322 may operate as a primary controller in which services directing I / O operations to the cloud-based storage system 318 are otherwise partitioned, and so on. In fact, in other embodiments where cost savings may take precedence over performance requirements, there may be only a single cloud computing example containing a storage controller application.
[0144] Figure 3C The cloud-based storage system 318 described herein includes cloud computing examples 340a, 340b, and 340n having local storage devices 330, 334, and 338. For example, cloud computing examples 340a, 340b, and 340n may be embodied as examples of cloud computing resources that can be provided by cloud computing environment 316 to support the execution of software applications. Figure 3C The cloud computing examples 340a, 340b, and 340n may differ from the cloud computing examples 320 and 322 described above because... Figure 3CCloud computing examples 340a, 340b, and 340n have local storage devices 330, 334, and 338 resources, while cloud computing examples 320 and 322, which support the execution of storage controller applications 324 and 326, do not require local storage resources. For example, cloud computing examples 340a, 340b, and 340n with local storage devices 330, 334, and 338 can be embodied as EC2 M5 examples containing one or more SSDs, EC2 R5 examples containing one or more SSDs, EC2 I3 examples containing one or more SSDs, and so on. In some embodiments, local storage devices 330, 334, and 338 must be embodied as solid-state storage devices (e.g., SSDs) rather than storage devices using hard disk drives.
[0145] exist Figure 3C In the examples depicted, each of the cloud computing examples 340a, 340b, and 340n, having local storage devices 330, 334, and 338, may include software daemons 328, 332, and 336, which, when executed by the cloud computing examples 340a, 340b, and 340n, may present themselves to the storage controller applications 324 and 326 as if the cloud computing examples 340a, 340b, and 340n were physical storage devices (e.g., one or more SSDs). In this example, the software daemons 328, 332, and 336 may contain computer program instructions similar to those typically contained on storage devices, enabling the storage controller applications 324 and 326 to send and receive the same commands that the storage controller would send to the storage devices. In this way, the storage controller applications 324 and 326 may contain code that is the same (or substantially the same) as the code that will be executed by the controller in the storage system described above. In these and similar embodiments, communication between storage controller applications 324, 326 and cloud computing examples 340a, 340b, 340n having local storage devices 330, 334, 338 may utilize iSCSI, TCP-based NVMe, messaging, custom protocols, or some other mechanism.
[0146] exist Figure 3CIn the examples depicted, each of the cloud computing examples 340a, 340b, and 340n, having local storage devices 330, 334, and 338, can also be coupled to block storage devices 342, 344, and 346 provided by the cloud computing environment 316, for example (as an example) as an Amazon Elastic Block Storage ('EBS') volume. In this example, the block storage devices 342, 344, and 346 provided by the cloud computing environment 316 can be utilized in a manner similar to how NVRAM devices described above are utilized, because the software daemons 328, 332, and 336 (or some other module) executing within a particular cloud computing example 340a, 340b, or 340n can initiate writing data to its attached EBS volume and to its local storage device 330, 334, or 338 resources upon receiving a request to write data. In some alternative embodiments, data may only be written to local storage devices 330, 334, and 338 resources within specific cloud computing examples 340a, 340b, and 340n. In alternative embodiments, instead of using block storage devices 342, 344, and 346 provided by the cloud computing environment 316 as NVRAM, the actual RAM on each of the cloud computing examples 340a, 340b, and 340n with local storage devices 330, 334, and 338 may be used as NVRAM, thereby reducing the network utilization costs associated with using EBS volumes as NVRAM. In yet another embodiment, high-performance block storage resources, such as one or more Azure HyperDisks, may be used as NVRAM.
[0147] When a request to write data is received by a specific cloud computing example 340a, 340b, 340n having local storage devices 330, 334, 338, software daemons 328, 332, 336 can be configured to write data not only to their own local storage devices 330, 334, 338 and any appropriate block storage devices 342, 344, 346, but also to write data to a cloud-based object storage device 348 attached to the specific cloud computing example 340a, 340b, 340n. For example, the cloud-based object storage device 348 attached to the specific cloud computing example 340a, 340b, 340n can be embodied as Amazon Simple Storage Service ('S3'). In other embodiments, cloud computing examples 320, 322, each including storage controller applications 324, 326, can initiate the storage of data in local storage devices 330, 334, 338 and cloud-based object storage device 348 of cloud computing examples 340a, 340b, 340n. In other embodiments, instead of using cloud computing examples 340a, 340b, 340n and cloud-based object storage device 348 with local storage devices 330, 334, 338 (also referred to herein as 'virtual drives') to store data, a persistent storage tier can be implemented in other ways. For example, one or more Azure HyperDisks can be used to persistently store data (e.g., after the data has been written to the NVRAM tier). In embodiments where one or more Azure HyperDisks are used for persistent data storage, the use of cloud-based object storage device 348 can be eliminated, so that data is only persistently stored in Azure HyperDisks without having to write data to the object storage tier.
[0148] While the local storage devices 330, 334, and 338 and the block storage devices 342, 344, and 346 utilized by cloud computing examples 340a, 340b, and 340n support block-level access, the cloud-based object storage device 348 attached to a specific cloud computing example 340a, 340b, or 340n only supports object-based access. Software daemons 328, 332, and 336 can therefore be configured to take data blocks, encapsulate those blocks into objects, and write the objects to the cloud-based object storage device 348 attached to the specific cloud computing example 340a, 340b, or 340n.
[0149] In some embodiments, all data stored by the cloud-based storage system 318 may be stored in either: 1) a cloud-based object storage device 348; and 2) at least one of the local storage devices 330, 334, 338 or block storage devices 342, 344, 346 resources utilized by cloud computing examples 340a, 340b, 340n. In such embodiments, the local storage devices 330, 334, 338 resources and the block storage devices 342, 344, 346 resources utilized by cloud computing examples 340a, 340b, 340n can be effectively operated as a cache that typically contains all data still stored in S3, such that all data reads can be serviced by cloud computing examples 340a, 340b, 340n without requiring cloud computing examples 340a, 340b, 340n to access the cloud-based object storage device 348. However, the reader will understand that in other embodiments, all data stored by the cloud-based storage system 318 may be stored in the cloud-based object storage device 348, but not all data stored by the cloud-based storage system 318 may be stored in at least one of the local storage devices 330, 334, 338 or block storage devices 342, 344, 346 resources utilized by the cloud computing examples 340a, 340b, 340n. In this example, various strategies may be used to determine which subset of the data stored by the cloud-based storage system 318 should reside in either: 1) the cloud-based object storage device 348; and 2) at least one of the local storage devices 330, 334, 338 or block storage devices 342, 344, 346 resources utilized by the cloud computing examples 340a, 340b, 340n.
[0150] One or more computer program instruction modules executing within the cloud-based storage system 318 (e.g., a monitoring module executing on its own EC2 instance) can be designed to handle the failure of one or more of the cloud computing instances 340a, 340b, 340n having local storage devices 330, 334, 338. In this example, the monitoring module can handle the failure of one or more of the cloud computing instances 340a, 340b, 340n having local storage devices 330, 334, 338 by creating one or more new cloud computing instances having local storage devices, retrieving data stored on the failed cloud computing instances 340a, 340b, 340n from the cloud-based object storage device 348, and storing the data retrieved from the cloud-based object storage device 348 in the local storage device on the newly created cloud computing instance. The reader will understand that many variations of this process can be implemented.
[0151] Readers will understand that various performance aspects of the cloud-based storage system 318 can be monitored (e.g., by the monitoring module executing in the EC2 example), enabling the cloud-based storage system 318 to scale vertically or horizontally as needed. For example, if the cloud computing instances 320 and 322 used to support the execution of storage controller applications 324 and 326 are too small to serve the I / O requests issued by users of the cloud-based storage system 318, the monitoring module can create a new, more powerful cloud computing instance (e.g., a type of cloud computing instance with more processing power, more storage, etc.), which includes the storage controller application, allowing the new, more powerful cloud computing instance to begin operating as the primary controller. Similarly, if the monitoring module determines that the cloud computing instances 320 and 322 used to support the execution of storage controller applications 324 and 326 are too large and cost savings can be achieved by switching to a smaller, less powerful cloud computing instance, the monitoring module can create a new, less powerful (and cheaper) cloud computing instance, which includes the storage controller application, allowing the new, less powerful cloud computing instance to begin operating as the primary controller.
[0152] The storage system described above can implement intelligent data backup technology, which replicates data stored in the storage system and stores the data in disparate locations to prevent data loss in the event of equipment failure or other forms of disaster. For example, the storage system described above can be configured to inspect each backup to avoid restoring the storage system to an unintended state. Consider an example where malware has infected the storage system. In this example, the storage system may include software resource 314 that can scan each backup to identify backups captured before and after malware infection. In this example, the storage system can restore itself from backups that do not contain malware—or at least not from portions of backups containing malware. In this example, the storage system may include software resource 314 that can scan each backup to identify the presence of malware (or virus or something else) for example by identifying write operations served by the storage system and originating from a network subnet suspected of having been delivered malware, by identifying write operations served by the storage system and originating from a user suspected of having been delivered malware, by identifying write operations served by the storage system and checking the content of the write operations against the fingerprint of malware, and in many other ways.
[0153] Readers will further understand that backups (typically in the form of one or more snapshots) can also be used to perform rapid recovery of the storage system. Consider an instance where the storage system is infected with ransomware that locks users out of the storage system. In this instance, software resource 314 within the storage system can be configured to detect the presence of ransomware and can be further configured to use a retained backup to restore the storage system to a point in time prior to the time the ransomware infected the storage system. In this instance, the presence of ransomware can be explicitly detected using software tools exploited by the system, by using a key inserted into the storage system (e.g., a USB drive), or in a similar manner. Similarly, the presence of ransomware can be inferred in response to system activity satisfying a predetermined fingerprint (e.g., no reads or writes to the system within a predetermined period of time).
[0154] Readers will understand that the components described above can be grouped into one or more optimized compute packages as converged infrastructure. This converged infrastructure can comprise a pool of computing, storage, and networking resources that can be shared by multiple applications and managed collectively using policy-driven processes. Such converged infrastructure can be implemented using a converged infrastructure reference architecture, as standalone devices, as a software-driven hyperconverged approach (e.g., hyperconverged infrastructure), or otherwise.
[0155] Readers will understand that the storage systems described in this disclosure can be used to support a wide variety of software applications. In fact, a storage system can be 'application-aware' in the sense that it can acquire, maintain, or otherwise access information describing a connected application (e.g., an application utilizing the storage system) to intelligently optimize its operation based on information about the application and its usage patterns. For example, the storage system may optimize data layout, optimize cache behavior, optimize 'QoS' tiers, or perform other optimizations designed to improve storage performance experienced by the application.
[0156] As an example of a type of application that can be supported by the storage system described herein, storage system 306 can be used to support such applications by providing storage resources to applications such as: artificial intelligence ('AI') applications, database applications, XOps projects (e.g., DevOps projects, DataOps projects, MLOps projects, ModelOps projects, PlatformOps projects), electronic design automation tools, event-driven software applications, high-performance computing applications, simulation applications, high-speed data capture and analysis applications, machine learning applications, media production applications, media service applications, picture archiving and communication system ('PACS') applications, software development applications, virtual reality applications, augmented reality applications, and many other types of applications.
[0157] Given that storage systems encompass computing resources, storage resources, and a wide variety of other resources, they may be well-suited to support resource-intensive applications, such as (for example) AI applications. AI applications can be deployed across a wide range of sectors, including: predictive maintenance in manufacturing and related fields; healthcare applications such as patient data and risk analytics; retail and marketing deployments (e.g., search notifications, social media notifications); supply chain solutions; fintech solutions such as business analytics and reporting tools; operational deployments such as real-time analytics tools, application performance management tools, IT infrastructure management tools; and many others.
[0158] Such AI applications enable devices to perceive their environment and take actions that maximize their chances of success in achieving a goal. Examples of such AI applications include IBM Watson. TM Microsoft Oxford TM Google DeepMind TM Baidu Minwa TM and others.
[0159] The storage system described above may also be well-suited for supporting other types of resource-intensive applications, such as (for example) machine learning applications. Machine learning applications perform various types of data analysis to automate the construction of analytical models. Using algorithms that iteratively learn from data, machine learning applications enable computers to learn without being explicitly programmed. A specific area of machine learning is called reinforcement learning, which involves taking appropriate actions in specific situations to maximize rewards.
[0160] In addition to the resources already described, the storage system described above may also include a graphics processing unit ('GPU'), sometimes referred to as a vision processing unit ('VPU'). Such a GPU may be embodied as dedicated electronic circuitry that rapidly manipulates and modifies memory to accelerate the creation of images in the frame buffer intended for output to a display device. Such a GPU may be included within any computing device that is part of the storage system described above, and may include one of many individual scalable components of the storage system, wherein other instances of these individual scalable components may include storage components, memory components, computing components (e.g., CPU, FPGA, ASIC), networking components, software components, and others. Besides GPUs, the storage system described above may also include neural network processors ('NNPs') for various aspects of neural network processing. Such NNPs may be used in place of GPUs (or other than GPUs), and may also be independently scalable.
[0161] As described above, the storage system described in this paper can be configured to support artificial intelligence applications, machine learning applications, big data analytics applications, and many other types of applications. The rapid growth of these applications is driven by three technologies: deep learning (DL), GPU processors, and big data. Deep learning is a computational model that uses massively parallel neural networks inspired by the human brain. Instead of experts manually crafting software, deep learning models learn their own software by learning from a large number of instances. Such GPUs can contain thousands of cores, making them well-suited for running algorithms that loosely represent the parallel nature of the human brain.
[0162] Advances in deep neural networks, including the development of multi-layered neural networks, have spurred a wave of new algorithms and tools for data scientists to mine their data using artificial intelligence (AI). Using improved algorithms, larger datasets, and various frameworks (including open-source software libraries for machine learning across a range of tasks), data scientists are addressing new use cases such as autonomous vehicles, natural language processing and understanding, computer vision, machine reasoning, strong AI, and many others. AI technologies are already being implemented in a wide variety of products, including, for example, Amazon Echo's speech recognition technology, which allows users to converse with their machines; and Google Translate. TM This includes technologies such as machine-based language translation; Spotify's Discover Weekly, which recommends new songs and artists that users might like based on user activity and traffic analysis; Quill's text generation product, which takes structured data and transforms it into narrative stories; chatbots, which provide real-time, context-specific answers to questions in a conversational format; and many others.
[0163] Data is at the heart of modern AI and deep learning algorithms. Before training can begin, a crucial issue to address is collecting labeled data essential for training accurate AI models. This may require a full-scale AI deployment to continuously collect, clean, transform, label, and store massive amounts of data. Adding additional high-quality data points directly translates into more accurate models and better insights. Data samples can undergo a series of processing steps, including but not limited to: 1) ingesting data from external sources into the training system and storing the data in its raw form; 2) cleaning and transforming the data in a training-friendly format, including linking data samples to appropriate labels; 3) exploring parameters and models, quickly testing with smaller datasets, and iterating to converge the most promising model to advance it to the production cluster; 4) performing a training phase to select several batches of random input data, including both new and older samples, and feeding those to the production GPU server for computation to update model parameters; and 5) evaluation, including using the retained portion of data not used for training, to assess the accuracy of the model while preserving the data. This lifecycle can be applied to any type of parallel machine learning, not just neural networks or deep learning. For example, standard machine learning frameworks may rely on CPUs (rather than GPUs), but the data ingestion and training workflows can remain the same. Readers will understand that a single shared-store data hub creates a coordination point throughout the entire lifecycle, eliminating the need for additional data copies during the ingestion, preprocessing, and training phases. Ingested data is rarely used for a single purpose, and shared storage provides the flexibility to train multiple different models or apply traditional analytics to the data.
[0164] Readers will learn that each stage in an AI data pipeline may have different requirements from the data central point (e.g., a storage system or collection of storage systems). Scale-out storage systems must provide unimpeded performance for all types and formats of access (from small metadata-heavy systems to large files, from random to sequential access, and from low to high concurrency). The storage system described above serves as an ideal AI data central point because it can serve unstructured workloads. In the first stage, data is ideally ingested and stored at the same central point that will be used in subsequent stages to avoid additional data duplication. The next two steps can be performed on standard compute servers that optionally include GPUs, and then, in the fourth and final stages, the complete training production job runs on powerful GPU-accelerated servers. Typically, a production pipeline runs alongside an experimental pipeline that operates on the same dataset. Furthermore, GPU-accelerated servers can be used independently for different models or combined to train on a larger model, or even for distributed training across multiple systems. If the shared storage tier is slow, data must be copied to local storage for each stage, resulting in wasted time pausing data across different servers. An ideal data centralization point for an AI training pipeline provides performance similar to data stored locally on server nodes, while also offering the simplicity and performance to enable concurrent operation across all pipeline stages.
[0165] To enable the storage system described above to serve as a data central point or as part of an AI deployment, in some embodiments, the storage system may be configured to provide DMA between storage devices contained within the storage system and one or more GPUs used in an AI or big data analytics pipeline. One or more GPUs may be coupled to the storage system, for example, via architecture-based NVMe ('NVMe-oF'), allowing bottlenecks such as those of the host CPU to be bypassed and the storage system (or one of its contained components) to directly access GPU memory. In this example, the storage system may utilize API hooks to the GPU to transfer data directly to the GPU. For example, the GPU may be embodied as an Nvidia GPU. TM The GPU and the storage system may support GPUDirect Storage ('GDS') software or have similar proprietary software that enables the storage system to transfer data to the GPU via RDMA or a similar mechanism.
[0166] While the preceding paragraphs discuss deep learning applications, the reader will understand that the storage system described in this article can also be part of a distributed deep learning ('DDL') platform to support the execution of DDL algorithms. The storage system described above can also be paired with other technologies, such as TensorFlow, open-source software libraries for dataflow programming across a range of tasks that can be used in machine learning applications (such as neural networks), to facilitate the development of such machine learning models, applications, and so on.
[0167] The storage system described above can also be used in neuromorphic computing environments. Neuromorphic computing is a form of computation that simulates brain cells. To support neuromorphic computing, an architecture of interconnected "neurons" replaces traditional computational models with low-power signals transmitted directly between neurons, enabling more efficient computation. Neuromorphic computing can utilize very large-scale integration (VLSI) systems containing electronic analog circuitry to simulate the neurobiological architecture present in neural systems, as well as analog, digital, and mixed-mode analog / digital VLSI, and software systems that implement models of neural systems for perception, motor control, or multisensory integration.
[0168] Readers will understand that the storage system described above can be configured to support the storage or use of blockchain and its derivatives (as well as other types of data), for example (for instance) as IBM TM This document covers open-source blockchains and related tools from the Hyperledger Project, permissioned blockchains where a limited number of trusted parties are allowed access to the blockchain, blockchain products that enable developers to build their own distributed ledger projects, and others. The blockchains and storage systems described herein can be used to support both on-chain and off-chain storage of data.
[0169] Off-chain storage of data can be implemented in various ways and can occur even when the data itself is not stored within the blockchain. For example, in one embodiment, a hash function can be used, and the data itself can be fed into the hash function to generate a hash value. In this example, the hash of a large number of data entries can be embedded within a transaction, rather than the data itself. The reader will understand that in other embodiments, alternatives to blockchain can be used to facilitate decentralized storage of information. For example, one usable alternative to blockchain is blockweave. While a regular blockchain stores each transaction for confirmation, blockweave allows for secure, decentralized on-chain storage of data without using the entire chain. Such blockweave can utilize consensus mechanisms based on Proof-of-Access (PoA) and Proof-of-Work (PoW).
[0170] The storage systems described above can be used alone or in combination with other computing devices to support in-memory computing applications. In-memory computing involves storing information in RAM distributed across a computer cluster. The reader will understand that the storage systems described above, especially those configurable with customizable amounts of processing, storage, and memory resources (e.g., those where blades contain configurable amounts of each type of resource), can provide an infrastructure capable of supporting in-memory computing. Similarly, the storage systems described above may include components (e.g., NVDIMMs providing persistent high-speed random access memory, 3D cross-point storage devices) that can effectively provide an improved in-memory computing environment compared to in-memory computing environments that rely on RAM distributed across dedicated servers.
[0171] In some embodiments, the storage system described above can be configured to operate as a hybrid memory computing environment containing a common interface to all storage media, such as RAM, flash memory, and 3D cross-point memory. In such embodiments, users may not know the details of where their data is stored, but they can still use the same complete and unified API to address the data. In such embodiments, the storage system can (in the background) move data to the fastest available tier—including various characteristics dependent on the data or relying on some other heuristic to intelligently place the data. In this example, the storage system can even use existing products (such as Apache Ignite and GridGain) to move data between various storage tiers, or the storage system can use custom software to move data between various storage tiers. The storage system described herein can implement various optimizations to improve the performance of computation in memory, such as (for example) bringing computation as close to the data as possible.
[0172] As the reader will further understand, in some embodiments, the storage system described above can be paired with other resources to support the applications described above. For example, an infrastructure may include a main computer in the form of servers and workstations, which are dedicated to using general-purpose computing on graphics processing units ('GPGPUs') to accelerate deep learning applications interconnected to a computing engine to train parameters of deep neural networks. Each system may have Ethernet external connectivity, unlimited bandwidth external connectivity, some other form of external connectivity, or some combination thereof. In this example, GPUs may be grouped for a single large training run or used independently to train multiple models. The infrastructure may also include storage systems (such as those described above) to provide, for example, horizontally scalable all-flash file or object storage areas through which data can be accessed via high-performance protocols (such as NFS, S3, etc.). The infrastructure may also include redundant top-of-rack Ethernet switches, for example, connected to storage devices and computers via ports in MLAG port channels, for redundancy. The infrastructure may also include additional computing, optionally in the form of white-box servers with GPUs, for data ingestion, preprocessing, and model debugging. The reader will understand that additional infrastructure is also possible.
[0173] Readers will see that the storage systems described above, either alone or in conjunction with other computing machines, can be configured to support other AI-related tools. For example, the storage system can use tools such as ONXX or other open neural network exchange formats that make it easier to transfer models written in different AI frameworks. Similarly, the storage system can be configured to support tools such as Amazon's Gluon, which allows developers to prototype, build, and train deep learning models. In fact, the storage systems described above can be part of larger platforms, such as IBM... TM Private cloud data includes integrated data science, data engineering, and application building services.
[0174] Readers will further understand that the storage system described above can also be deployed as an edge solution. This edge solution can be positioned appropriately to optimize cloud computing systems by performing data processing at the edge of the network, near the source of the data. Edge computing pushes applications, data, and computing power (i.e., services) from a central point to the logical edge of the network. By using edge solutions, such as the storage system described above, computing tasks can be performed using the computing resources provided by such storage systems, data can be stored using the storage resources of the storage system, and cloud-based services can be accessed using the various resources of the storage system, including networking resources. By performing computing tasks on edge solutions, storing data on edge solutions, and generally using edge solutions, expensive cloud-based resources can be avoided, and in fact, performance improvements can be experienced compared to a heavier reliance on cloud-based resources.
[0175] While many tasks can benefit from leveraging edge solutions, certain use cases are particularly well-suited for deployment in this environment. For example, devices such as drones, self-driving cars, robots, and others may require extremely fast processing—in fact, so fast that sending data up to the cloud and receiving it back for processing support can be incredibly slow. As an additional example, some IoT devices (such as connected cameras) may not be well-suited for cloud-based resources because sending data to the cloud simply due to the sheer volume of data involved may be impractical (not just from a privacy, security, or financial perspective). Therefore, many tasks that truly involve data processing, storage, or communication may be better suited to platforms that incorporate edge solutions, such as the storage systems described above.
[0176] The storage system described above can be used alone or in combination with other computing resources as a network edge platform that integrates computing resources, storage resources, networking resources, cloud technologies, and network virtualization technologies. As part of the network, the edge can have characteristics similar to other network infrastructure from customer premises and backhaul aggregation facilities to access points (PoPs) and regional data centers. Readers will understand that network workloads (such as Virtual Network Functions (VNFs) and others) will reside on the network edge platform. Implemented through a combination of containers and virtual machines, the network edge platform can rely on controllers and schedulers that are no longer geographically co-located with data processing resources. Functionality can be broken down as microservices into the control plane, user and data planes, or even state machines, allowing for independent application optimization and scaling techniques. Such user and data planes can be implemented with added accelerators residing in server platforms (such as FPGAs and smart NICs) and implemented using SDN-enabled commercial silicon and programmable ASICs.
[0177] The storage system described above can also be optimized for big data analytics, including components used as composable data analysis pipelines, where containerized analytics architectures, for example, make analytical capabilities more composable. Big data analytics can be broadly described as the process of examining large and diverse datasets to discover hidden patterns, unknown correlations, market trends, customer preferences, and other useful information that can help organizations make more informed business decisions. As part of that process, semi-structured and unstructured data (e.g., internet clickstream data, web server logs, social media content, text from customer emails and survey responses, mobile phone call details, IoT sensor data, and other data) can be transformed into structured forms.
[0178] The storage system described above can also support (including implementations as a system interface) applications that respond to human voice commands to perform tasks. For example, the storage system can support intelligent personal assistant applications, such as (for example) Amazon's Alexa. TM Apple Siri TM Google Voice TM Samsung Bixby TM Microsoft Cortana TM And others. While the examples described in the preceding sentences use voice as input, the storage system described above may also support chatbots, conversational bots, chattterbots, or human dialogue entities, or other applications configured to converse via auditory or text methods. Similarly, the storage system may actually execute such applications to enable users (e.g., system administrators) to interact with the storage system via voice. Such applications are typically capable of voice interaction, music playback, creating to-do lists, setting alarms, streaming podcasts, playing audiobooks, and providing weather, traffic, and other real-time information, such as news; however, in embodiments according to this disclosure, such applications may serve as interfaces for various system management operations.
[0179] The storage systems described above can also implement AI platforms to realize the vision of autonomous storage. Such AI platforms can be configured to provide global predictive intelligence by collecting and analyzing vast amounts of storage system telemetry data points, enabling easy management, analysis, and support. In practice, such storage systems may be able to predict both capacity and performance, and generate intelligent recommendations regarding workload deployment, interaction, and optimization. These AI platforms can be configured to scan all incoming storage system telemetry data against a problem fingerprint database to predict and resolve incidents in real time before they impact the customer's environment, and capture hundreds of performance-related variables used to predict performance loads.
[0180] The storage system described above can support the sequential or simultaneous execution of artificial intelligence applications, machine learning applications, data analytics applications, data transformation, and other tasks that can collectively form the AI ladder. This AI ladder can be effectively formed by combining such elements to create a complete data science pipeline, where the elements of the AI ladder are interdependent. For example, AI may require some form of machine learning, machine learning may require some form of analytics, analytics may require some form of data and information architecture architecture, and so on. Therefore, each element can be considered a step in the AI ladder, which together form a complete and complex AI solution.
[0181] The storage systems described above can also be used, alone or in combination with other computing environments, to deliver an experience where AI permeates a wide range of business and life. For example, AI may play a significant role in providing deep learning solutions, deep reinforcement learning solutions, general artificial intelligence solutions, autonomous vehicles, cognitive computing solutions, commercial UAVs or drones, conversational user interfaces, enterprise classification, ontology management solutions, machine learning solutions, smart dust, smart robots, smart workplaces, and many others.
[0182] The storage system described above can also be used, alone or in combination with other computing environments, to provide a wide range of transparent immersive experiences, including experiences using digital twins of various “things” (such as people, places, processes, systems, etc.), where technology can introduce transparency between people, businesses, and things. This transparent immersive experience can be provided through augmented reality, connected homes, virtual reality, brain-computer interfaces, human augmentation technologies, nanotube electronics, volumetric displays, 4D printing, or others.
[0183] The storage system described above can also be used alone or in combination with other computing environments to support a wide variety of digital platforms. Such digital platforms may include, for example, 5G wireless systems and platforms, digital twin platforms, edge computing platforms, IoT platforms, quantum computing platforms, serverless PaaS, software-defined security, neuromorphic computing platforms, and so on.
[0184] The storage system described above can also be part of a multi-cloud environment, where multiple cloud computing and storage services are deployed in a single heterogeneous architecture. To facilitate operation in this multi-cloud environment, DevOps tools can be deployed to enable cross-cloud orchestration. Similarly, continuous development and continuous integration tools can be deployed to standardize processes surrounding continuous integration and delivery, new feature rollouts, and cloud workload deployment. By standardizing these processes, a multi-cloud strategy can be implemented, enabling the best provider to be utilized for each workload.
[0185] The storage system described above can be used as part of a platform to enable the use of cryptographic anchors, which can be used to authenticate the origin and content of products to ensure they match the blockchain records associated with the products. Similarly, as part of a suite of tools to protect data stored on the storage system, the storage system described above can implement various cryptographic techniques and schemes, including grid cryptography. Grid cryptography involves the construction of cryptographic primitives involving grids, both in the construction itself and in security proofs. Unlike public-key schemes such as RSA, Diffie-Hellman, or elliptic curve cryptosystems, which are vulnerable to quantum computer attacks, some grid-based constructions exhibit resistance to attacks from both classical and quantum computers.
[0186] A quantum computer is a device that performs quantum computation. Quantum computing uses quantum mechanical phenomena, such as superposition and entanglement, to perform calculations. Quantum computers differ from conventional transistor-based computers because these computers require data to be encoded as binary digits (bits), where each binary digit is always in one of two definite states (0 or 1). Unlike conventional computers, quantum computers use qubits that can be in superpositions of states. A quantum computer maintains a sequence of qubits, where a single qubit can represent one, zero, or any quantum superposition of those two qubit states. A pair of qubits can be in any quantum superposition of four states, and three qubits can be in any superposition of eight states. A quantum computer with n qubits can typically be in any superposition of up to 2^n different states simultaneously, while a conventional computer can only be in one of these states at any given time. The quantum Turing machine is the theoretical model for this type of computer.
[0187] The storage systems described above can also be paired with FPGA acceleration servers as part of a larger AI or ML infrastructure. Such FPGA acceleration servers can reside near the storage systems described above (e.g., in the same data center) or even be incorporated into a device that includes one or more storage systems, one or more FPGA acceleration servers, networking infrastructure supporting communication between the one or more storage systems and the one or more FPGA acceleration servers, and other hardware and software components. Alternatively, FPGA acceleration servers can reside within a cloud computing environment that can be used to perform computationally relevant tasks for AI and ML jobs. Any of the embodiments described above can be used collectively as an FPGA-based AI or ML platform. Readers will understand that in some embodiments of an FPGA-based AI or ML platform, the FPGA contained within the FPGA acceleration server can be reconfigured for different types of ML models (e.g., LSTM, CNN, GRU). The ability to reconfigure the FPGA contained within the FPGA acceleration server can accelerate ML or AI applications based on the best numerical precision and memory model used. Readers will understand that by viewing the collection of FPGA acceleration servers as an FPGA pool, any CPU in the data center can use the FPGA pool as a shared hardware microservice, rather than limiting servers to dedicated accelerators plugged into it.
[0188] The FPGA-accelerated servers and GPU-accelerated servers described above can implement computational models where, unlike more traditional computational models where a small amount of data is held in the CPU and a long stream of instructions is executed via it, machine learning models and parameters are fixed in a high-bandwidth single-chip memory, through which a large amount of data flows. For this computational model, FPGAs can even be more efficient than GPUs because FPGAs can be programmed with only the instructions required to run this computational model.
[0189] The storage system described above can be configured to provide parallel storage, for example, by using a parallel file system such as BeeGFS. Such a parallel file system can contain a distributed metadata architecture. For example, a parallel file system can contain multiple metadata servers across its distributed metadata, as well as components containing services for clients and storage servers.
[0190] The system described above can support the execution of a large number of software applications. These applications can be deployed in various ways, including container-based deployment models. Various tools can be used to manage containerized applications. For example, Docker Swarm, Kubernetes, and others can be used to manage containerized applications. Containerized applications can be used to facilitate serverless, cloud-native computing deployment and management models for software applications. To support serverless, cloud-native computing deployment and management models for software applications, containers can be used as part of event handling mechanisms (such as AWS Lambdas), causing various events to trigger the launch of containerized applications to operate as event handlers.
[0191] The system described above can be deployed in various ways, including to support fifth-generation ('5G') networks. 5G networks support data communication that is significantly faster than previous generations of mobile communication networks, thus leading to the decomposition of data and computing resources. This is because modern large-scale data centers may become less prominent and may be replaced, for example, by more localized micro data centers located closer to mobile network towers. The system described above can be included in such localized micro data centers and can be part of or paired with a multi-access edge computing ('MEC') system. Such MEC systems enable cloud computing capabilities and IT service environments at the edge of cellular networks. By running applications and performing related processing tasks closer to cellular customers, network congestion is reduced, and applications can perform better.
[0192] The storage system described above can also be configured to implement NVMe partitioned namespaces. By using NVMe partitioned namespaces, the logical address space of the namespace is divided into zones. Each zone provides a logical block address range that must be written sequentially and explicitly reset before being overwritten, thereby enabling the creation of natural boundaries to expose the device and offloading the management of internal mapping tables to the host's namespace. To implement NVMe partitioned namespaces ('ZNS'), ZNS SSDs or some other form of partitioned block device can be utilized, which use zones to expose the namespace's logical address space. By aligning zones to the internal physical properties of the device, several inefficiencies in data placement can be eliminated. In such embodiments, for example, each zone can be mapped to a separate application, enabling functions such as wear leveling and discarded item collection to be performed on a zone-by-zone or application-by-application basis (rather than across the entire device). To support ZNS, the storage controller described herein can be configured to use, for example, Linux TM Kernel partition block device interface or other tools to interact with partition block devices.
[0193] The storage system described above can also be configured to implement partitioned storage in other ways, for example (by using shingled magnetic recording (SMR) storage devices. In instances where partitioned storage is used, device-managed embodiments can be deployed, where the storage device hides this complexity by managing it in the firmware, thus presenting an interface like any other storage device. Alternatively, partitioned storage can be implemented via host-managed embodiments, which depend on the operating system knowing how to handle the drive and only sequentially writing to certain areas of the drive. Partitioned storage can similarly be implemented using host-aware embodiments, where a combination of drive management and host management implementations is deployed.
[0194] The storage system described herein can be used to form a data lake. A data lake can operate as the first place where an organization's data flows, where such data can be in its raw format. Metadata tagging can be implemented to facilitate the search for data elements in the data lake, especially in embodiments where the data lake contains multiple data stores in formats that may not be easily accessible or readable (e.g., unstructured data, semi-structured data, structured data). Data can flow from the data lake down to a data warehouse, where it can be further processed, packaged, and stored in consumable formats. The storage system described above can also be used to implement this data warehouse. Additionally, data marts or data centralizations can allow for even more consumable data, where the storage system described above can also be used to provide the underlying storage resources required for the data marts or data centralizations. In embodiments, querying a data lake may require a schema-on-read approach, where data is applied to a schema when it is pulled from storage, rather than when it enters storage.
[0195] The storage system described herein can also be configured to implement Recovery Point Objectives ('RPOs'), which can be established by a user, by an administrator, as a system default, as part of a storage class or service that the storage system participates in delivering, or in some other way. A "Recovery Point Objective" is a target for the maximum time difference between the last update of the source dataset and the last recoverable replicated dataset update, given a rationale that the last recoverable replicated dataset update can be correctly recovered from consecutive or frequently updated copies of the source dataset. Updates can be correctly recovered if all updates processed on the source dataset prior to the last recoverable replicated dataset update are properly considered.
[0196] In synchronous replication, the Recovery Point Objective (RPO) will be zero, meaning that under normal operation, all completed updates on the source dataset should exist and be correctly recovered on the replica dataset. In near-synchronous replication, the RPO can be as low as a few seconds. In snapshot-based replication, the RPO can be roughly calculated as the interval between snapshots plus the time used to transfer modifications between previously transferred snapshots and the most recent snapshot to be replicated.
[0197] If updates accumulate faster than they are replicated, an Recovery Point Objective (RPO) may be missed. Similarly, if the data to be replicated accumulates between two snapshots (for snapshot-based replication) more than is replicable between taking a snapshot and replicating the accumulated updates from that snapshot, an RPO may be missed. Again, in snapshot-based replication, if the data to be replicated accumulates at a rate faster than it is transmitted between subsequent snapshots, replication may begin to lag further, potentially prolonging the miss between the expected recovery point target and the actual recovery point represented by the last correctly replicated update.
[0198] The storage systems described above can also be part of a shared-nothing (SNO) cluster. In a SNO cluster, each node in the cluster has local storage and communicates with other nodes in the cluster via a network, where the storage used by the cluster is (typically) provided only by the storage connected to each other node. A set of nodes that synchronously replicates a dataset can be an example of a SNO cluster because each storage system has local storage and communicates with other storage systems via a network, where those storage systems (typically) do not use storage from other places that they share access to via some kind of interconnect. In contrast, some of the storage systems described above are built as shared storage clusters themselves because there are drive racks shared by paired controllers. However, other storage systems described above are built as SNO clusters because all storage is local for a particular node (e.g., a blade), and all communication is via a network that links compute nodes together.
[0199] In other embodiments, other forms of shared-nothing storage clusters may include embodiments in which any node in the cluster has a local copy of all the storage it needs, and where data is mirrored to other nodes in the cluster via synchronous replication to ensure that data is not lost or because other nodes are also using that storage. In this embodiment, if a new cluster node needs some data, that data can be copied from other nodes that have copies of the data to the new node.
[0200] In some embodiments, a mirror-based shared storage cluster can store multiple copies of all cluster-stored data, wherein each subset of the data is replicated to a specific set of nodes, and different subsets of the data are replicated to several different sets of nodes. In some variations, embodiments may store all data in the cluster across all nodes, while in other variations, nodes may be partitioned such that a first set of nodes will all store the same set of data, while a second, different set of nodes will all store different sets of data.
[0201] Readers will understand that RAFT-based databases (such as etcd) can operate like a shared-nothing storage cluster where all data is stored on all RAFT nodes. However, the amount of data stored in a RAFT cluster may be limited, preventing additional replicas from consuming excessive storage. Container server clusters may also be able to replicate all data across all cluster nodes, provided the containers are not too large and a significant portion of their data (manipulated by applications running within them) is stored elsewhere, such as on an S3 cluster or an external file server. In this example, container storage can be provided directly by the cluster through its shared-nothing storage model, where those containers provide images of the execution environment that form part of an application or service.
[0202] To further explain, Figure 3DThis describes an exemplary computing device 350 that can be specifically configured to perform one or more of the processes described herein. For example... Figure 3D As shown, computing device 350 may include a communication interface 352, a processor 354, a storage device 356, and an input / output (“I / O”) module 358 that are communicatively connected to each other via communication infrastructure 360. Although in Figure 3D The demonstration computing device 350 was shown in China, but Figure 3D The components described herein are not intended to be limiting. Additional or alternative components may be used in other embodiments. A more detailed description will now follow. Figure 3D The components of the computing device 350 shown in the image.
[0203] Communication interface 352 can be configured to communicate with one or more computing devices. Examples of communication interface 352 include (but are not limited to) wired network interfaces (e.g., network interface cards), wireless network interfaces (e.g., wireless network interface cards), modems, audio / video connections, and any other suitable interfaces.
[0204] Processor 354 generally refers to any type or form of processing unit capable of processing data and / or interpreting, executing one or more of the instructions, procedures and / or operations described herein and / or directing their execution. Processor 354 may perform operations by executing computer-executable instructions 362 (e.g., applications, software, code and / or other examples of executable data) stored in storage device 356.
[0205] Storage device 356 may include one or more data storage media, devices, or configurations and may employ any type and form of data storage media and / or devices and combinations thereof. For example, storage device 356 may include, but is not limited to, any combination of non-volatile media and / or volatile media described herein. Electronic data, including the data described herein, may be temporarily and / or permanently stored in storage device 356. For example, data representing computer-executable instructions 362 configured to bootstrap processor 354 to perform any of the operations described herein may be stored within storage device 356. In some instances, data may be arranged in one or more databases residing within storage device 356.
[0206] I / O module 358 may include one or more I / O modules configured to receive user input and provide user output. I / O module 358 may include any hardware, firmware, software, or a combination thereof that supports input and output capabilities. For example, I / O module 358 may include hardware and / or software for capturing user input, including but not limited to a keyboard or keypad, a touchscreen component (e.g., a touchscreen display), a receiver (e.g., an RF or infrared receiver), a motion sensor, and / or one or more input buttons.
[0207] I / O module 358 may include one or more means for presenting output to a user, including, but not limited to, a graphics engine, a display (e.g., a display screen), one or more output drivers (e.g., display drivers), one or more audio speakers, and one or more audio drivers. In some embodiments, I / O module 358 is configured to provide graphical data to a display for presentation to a user. The graphical data may represent one or more graphical user interfaces and / or any other graphical content that may serve a particular implementation. In some instances, any of the systems, computing devices, and / or other components described herein may be implemented by computing device 350.
[0208] To further explain, Figure 3E This describes an instance of a storage system 376 used to provide storage services (also referred to herein as 'data services'). Figure 3E The storage system 376 described herein includes a plurality of storage systems 374a, 374b, 374c to 374n, each of which may be similar to the storage system described herein. The storage systems 374a, 374b, 374c to 374n in the storage system 376 may be the same storage system or different types of storage systems. For example, Figure 3E The two storage systems 374a and 374n described are depicted as cloud-based storage systems because the resources that together form each of the storage systems 374a and 374n are provided by different cloud service providers 370 and 372. For example, the first cloud service provider 370 could be Amazon AWS. TM The second cloud service provider, 372, is Microsoft Azure. TM However, in other embodiments, one or more public clouds, private clouds, or combinations thereof may be used to provide underlying resources for forming a specific storage system in the storage system 376.
[0209] According to some embodiments of this disclosure Figure 3E The examples described herein include edge management services 366 used to deliver storage services. The delivered storage services (also referred to herein as 'data services') may include, for example, services that provide a specific amount of storage to a customer, services that provide storage to a customer in accordance with a pre-defined service level agreement, services that provide storage to a customer in accordance with pre-defined regulatory requirements, and many others.
[0210] Figure 3EThe edge management service 366 described herein may be embodied as, for example, one or more computer program instruction modules executing on computer hardware (e.g., one or more computer processors). Alternatively, the edge management service 366 may be embodied as one or more computer program instruction modules executing on a virtualized execution environment (e.g., one or more virtual machines), in one or more containers, or otherwise. In other embodiments, the edge management service 366 may be embodied as a combination of the embodiments described above, including embodiments in which one or more computer program instruction modules contained in the edge management service 366 are distributed across multiple physical or virtual execution environments.
[0211] Edge management service 366 can operate as a gateway to provide storage services to storage clients, where the storage services utilize storage provided by one or more storage systems 374a, 374b, 374c to 374n. For example, edge management service 366 can be configured to provide storage services to host devices 378a, 378b, 378c, 378d, 378n that are performing one or more applications that consume storage services. In this example, edge management service 366 operates as a gateway between host devices 378a, 378b, 378c, 378d, 378n and storage systems 374a, 374b, 374c to 374n, without requiring host devices 378a, 378b, 378c, 378d, 378n to have direct access to storage systems 374a, 374b, 374c to 374n.
[0212] Figure 3E The edge management service 366 exposes the storage service module 364 to Figure 3E The host devices are 378a, 378b, 378c, 378d, and 378n, but in other embodiments, the edge management service 366 can expose the storage service module 364 to other clients of various storage services. These various storage services can be presented to clients via one or more user interfaces, via one or more APIs, or through some other mechanism provided by the storage service module 364. Therefore, Figure 3E The storage service module 364 described herein may be embodied as one or more computer program instruction modules executed on physical hardware, on a virtualized execution environment, or a combination thereof, wherein execution of such modules enables clients of the storage service to be provided with various storage services and to select and access various storage services.
[0213] Figure 3E The edge management service 366 also includes the system management service module 368. Figure 3EThe system management service module 368 contains one or more computer program instruction modules, which, when executed, coordinate with storage systems 374a, 374b, 374c to 374n to perform various operations to provide storage services to host devices 378a, 378b, 378c, 378d, and 378n. The system management service module 368 can be configured to perform tasks, such as deploying storage resources from storage systems 374a, 374b, 374c to 374n via one or more APIs exposed by storage systems 374a, 374b, 374c to 374n; migrating datasets or workloads within storage systems 374a, 374b, 374c to 374n via one or more APIs exposed by storage systems 374a, 374b, 374c to 374n; setting one or more tunable parameters (i.e., one or more configurable settings) on storage systems 374a, 374b, 374c to 374n via one or more APIs exposed by storage systems 374a, 374b, 374c to 374n; and so on. For example, many of the services described below relate to embodiments in which storage systems 374a, 374b, 374c to 374n are configured to operate in a certain manner. In such instances, the system management service module 368 may be responsible for configuring the storage systems 374a, 374b, 374c to 374n to operate in the manner described below using the APIs (or some other mechanism) provided by the storage systems 374a, 374b, 374c to 374n.
[0214] In addition to configuring storage systems 374a, 374b, 374c to 374n, the edge management service 366 itself can be configured to perform various tasks required to provide various storage services. Consider an instance where the storage service includes a service that, when selected and applied, obfuscates personally identifiable information ('PII') contained in the dataset when the dataset is accessed. In this instance, storage systems 374a, 374b, 374c to 374n can be configured to obfuscate the PII when serving read requests directed to the dataset. Alternatively, storage systems 374a, 374b, 374c to 374n can serve reads by returning data containing the PII, but the edge management service 366 itself can obfuscate the PII as the data is transmitted through the edge management service 366 from storage systems 374a, 374b, 374c to 374n to host devices 378a, 378b, 378c, 378d, 378n.
[0215] Figure 3E The storage systems 374a, 374b, 374c to 374n described in the reference above can be embodied in the above reference. Figures 1A to 3DThe description refers to one or more of the storage systems, including variations thereof. In practice, storage systems 374a, 374b, 374c through 374n can be used as a storage resource pool, where individual components within that pool have different performance characteristics, different storage characteristics, and so on. For example, one of storage systems 374a could be a cloud-based storage system, another storage system 374b could be a block storage system, another storage system 374c could be a file storage system, another storage system 374d could be a relatively high-performance storage system, and another storage system 374n could be a relatively low-performance storage system, and so on. In alternative embodiments, only a single storage system may exist.
[0216] Figure 3E The storage systems 374a, 374b, 374c to 374n described herein can also be organized into different fault domains, such that a failure of one storage system 374a should be completely independent of a failure of another storage system 374b. For example, each of the storage systems can receive power from an independent power supply system, each of the storage systems can be coupled for data communication via an independent data communication network, and so on. Furthermore, storage systems in a first fault domain can be accessed via a first gateway, and storage systems in a second fault domain can be accessed via a second gateway. For example, the first gateway can be a first example of edge management service 366, and the second gateway can be a second example of edge management service 366, including embodiments where each example is distinct or each example is part of a distributed edge management service 366.
[0217] As an illustrative example of available storage services, storage services can be presented to users associated with different levels of data protection. For instance, a storage service can be presented to a user, and when selected and enforced, guarantees that the data associated with that user will be protected to ensure various Recovery Point Objectives ('RPOs'). A first available storage service can ensure, for example, that a dataset associated with the user will be protected so that any data older than 5 seconds can be recovered in the event of a failure of the primary data store, while a second available storage service can ensure that the dataset associated with the user will be protected so that any data older than 5 minutes can be recovered in the event of a failure of the primary data store.
[0218] Additional instances of storage services that can be presented to users, selected by users, and ultimately applied to the user's associated dataset may include one or more data compliance services. Such data compliance services may manifest as, for example, providing data compliance services to clients (i.e., users) to ensure that the user's dataset is managed in a way that complies with various regulatory requirements. For example, one or more data compliance services may be provided to users to ensure that the user's dataset is managed in a way that complies with the General Data Protection Regulation ('GDPR'), one or more data compliance services may be provided to users to ensure that the user's dataset is managed in a way that complies with the Sarbanes-Oxley Act of 2002 ('SOX'), or one or more data compliance services may be provided to users to ensure that the user's dataset is managed in a way that complies with another regulatory act. Additionally, one or more data compliance services may be provided to users to ensure that the user's dataset is managed in a way that complies with a non-governmental guideline (e.g., best practices for audit purposes), one or more data compliance services may be provided to users to ensure that the user's dataset is managed in a way that complies with specific client or organizational requirements, and so on.
[0219] To provide this specific data compliance service, the data compliance service may be presented to the user (e.g., via a GUI) and selected by the user. In response to receiving a selection for the specific data compliance service, one or more storage service policies may be applied to the dataset associated with the user to implement the specific data compliance service. For example, a storage service policy may be applied requiring the dataset to be encrypted before being stored in a storage system, before being stored in a cloud environment, or before being stored elsewhere. To enforce this policy, a provision may be enforced requiring not only encryption of the dataset when it is stored, but also encryption of the dataset before it is transmitted (e.g., sent to another party). In this example, a storage service policy may also be proposed requiring that any encryption keys used to encrypt the dataset are not stored on the same system that stores the dataset itself. The reader will understand that many other forms of data compliance services can be provided and implemented according to embodiments of this disclosure.
[0220] The storage systems 374a, 374b, 374c to 374n in the queue storage system 376 can be managed by one or more queue management modules. The queue management module can be... Figure 3EThe system management service module 368 described herein may be part of or separate from it. The team management module can perform tasks such as monitoring the health of each storage system in the team, initiating updates or upgrades on one or more storage systems in the team, migrating workloads for load balancing or other performance purposes, and many other tasks. Therefore, and for many other reasons, storage systems 374a, 374b, 374c to 374n may be coupled to each other via one or more data communication links to exchange data between storage systems 374a, 374b, 374c to 374n.
[0221] In some embodiments, one or more storage systems or one or more elements of storage systems (e.g., features, services, operations, components, etc. of the storage system), such as any illustrative storage system or storage system element described herein, may be implemented in one or more container systems. A container system may contain any system that supports the execution of one or more containerized applications or services. Such services may be software deployed as infrastructure for building applications, for operating runtime environments, and / or as infrastructure for other services. In the following discussion, the description of containerized applications generally also applies to containerized services.
[0222] Containers can combine one or more elements of a containerized software application with a runtime environment to operate those application elements bundled into a single image. For example, each container of a containerized application may contain the executable code of the software application and various dependencies, libraries, and / or other components, as well as network configurations and configured access to additional resources used by the elements of the software application within that particular container to enable the operation of those elements. A containerized application can be represented as a collection of such containers, which together represent all the elements of the application and the various runtime environments required for all those elements to run. Therefore, a containerized application can be abstracted from the host operating system as a lightweight and portable collection of packages and configurations, wherein the containerized application can be uniformly deployed and consistently executed in different computing environments using different container-compatible operating systems or different infrastructures. In some embodiments, the containerized application shares a kernel with the host computer system and executes as an isolated environment (an isolated set of files and directories, processes, system and network resources, and configured access to additional resources and capabilities) isolated by the host system's operating system in conjunction with a container management framework. When executed, the containerized application can provide one or more containerized workloads and / or services.
[0223] Container systems may contain and / or utilize clusters of nodes. For example, a container system may be configured to manage the deployment and execution of containerized applications on one or more nodes in a cluster. Containerized applications may utilize the resources of the nodes, such as memory, processing, and / or storage resources provided and / or accessed by the nodes. Storage resources may include any illustrative storage resources described herein and may include on-node resources (e.g., local files and directory trees), off-node resources (e.g., external networked file systems, databases, or object storage), or both on-node and off-node resources. Access to additional resources and capabilities that can be configured for containers of containerized applications may include specialized computing capabilities, such as GPUs and AI / ML engines, or specialized hardware, such as sensors and cameras.
[0224] In some embodiments, the container system may include a container orchestration system (which may also be referred to as a container orchestrator, container orchestration platform, etc.) designed to be reasonably simple and automated for many use cases to deploy, scale, and manage containerized applications. In some embodiments, the container system may include a storage management system configured to provide and manage storage resources (e.g., virtual volumes) for private or shared use by cluster nodes and / or containers of containerized applications.
[0225] Figure 3F This example describes container system 380. In this example, container system 380 includes container storage system 381, which can be configured to perform one or more storage management operations to organize, provide, and manage storage resources for use by one or more containerized applications 382-1 to 382-L of container system 380. Specifically, container storage system 381 can organize storage resources into one or more storage pools 383 for use by containerized applications 382-1 to 382-L. The container storage system itself can be implemented as a containerized service.
[0226] Container system 380 may include or be implemented by one or more container orchestration systems, including Kubernetes™, Mesos™, Docker Swarm™, etc. The container orchestration system manages container system 380 running on cluster 384 through services implemented by the control node (described as 385), and can further manage the relationship between container storage systems or individual containers and their storage devices, memory and CPU limits, networking, and their access to additional resources or services.
[0227] The control plane of container system 380 can implement services including: deploying applications via controller 386, monitoring applications via controller 386, providing interfaces via API server 387, and scheduling deployments via scheduler 388. In this example, controller 386, scheduler 388, API server 387, and container storage system 381 are implemented on a single node (node 385). In other instances, for resilience, the control plane can be implemented by multiple redundant nodes, where if the node providing management services for container system 380 fails, another redundant node can provide management services for cluster 384.
[0228] The data plane of container system 380 may contain a set of nodes that provide a container runtime for executing containerized applications. Individual nodes within cluster 384 may execute a container runtime (such as Docker™) and a container manager or node agent (such as a kubelet in Kubernetes, not depicted) that communicates with the control plane via a proxy (sometimes called a proxy server) (e.g., agent 389) connected to a local network. Agent 389 may use, for example, Internet Protocol (IP) port numbers to route network traffic to and from containers. For instance, a containerized application may request storage classes from the control plane, where the request is processed by the container manager, and the container manager uses agent 389 to forward the request to the control plane.
[0229] Cluster 384 can contain a group of nodes running containers for managed containerized applications. Nodes can be virtual machines or physical machines. Nodes can be host systems.
[0230] Container storage system 381 can orchestrate storage resources to provide storage to container system 380. For example, container storage system 381 can use storage pool 383 to provide persistent storage to containerized applications 382-1 to 382-L. Container storage system 381 itself can be deployed as a containerized application by a container orchestration system.
[0231] For example, the container storage system 381 application can be deployed within cluster 384 and perform management functions to provide storage to the containerized application 382. These management functions may include identifying one or more storage pools from available storage resources, providing virtual volumes on one or more nodes, replicating data, responding to and recovering from host and network failures, or handling storage operations. Storage pool 383 may contain storage resources from one or more local or remote sources, where the storage resources can be different types of storage, including (as instances) block storage, file storage, and object storage.
[0232] Container storage system 381 can also be deployed on a set of nodes for which a container orchestration system can provide persistent storage. In some instances, container storage system 381 can be deployed on all nodes of cluster 384 using, for example, a Kubernetes DaemonSet. In this instance, nodes 390-1 to 390-N provide the container runtime executed by container storage system 381. In other instances, some, but not all, nodes in the cluster can execute container storage system 381.
[0233] Container storage system 381 can manage storage on nodes and communicate with the control plane of container system 380 to provide dynamic volumes, including persistent volumes. Persistent volumes can be mounted on nodes as virtual volumes, such as virtual volumes 391-1 and 391-P. After mounting virtual volume 391, containerized applications can request and use, or otherwise configure, the storage provided by virtual volume 391. In this example, container storage system 381 can mount a driver on the node's kernel, where the driver handles storage operations directed to the virtual volume. In this example, the driver can receive storage operations directed to the virtual volume, and in response, the driver can perform storage operations on one or more storage resources within storage pool 383, either guided by additional logic within the container that implements container storage system 381 as a containerized service or using said additional logic.
[0234] Container storage system 381 can determine available storage resources in response to being deployed as a containerized service. For example, storage resources 392-1 to 392-M can include local storage, remote storage (storage on individual nodes in a cluster), or both local and remote storage. Storage resources can also include storage from external sources, such as various combinations of block storage systems, file storage systems, and object storage systems. Storage resources 392-1 to 392-M can include any type and / or configuration of storage resources (such as any of the illustrative storage resources described above), and container storage system 381 can be configured to determine available storage resources in any suitable manner (including based on configuration files). For example, a configuration file can specify account and authentication information for cloud-based object storage device 348 or cloud-based storage system 318. Container storage system 381 can also determine the availability of one or more storage devices 356 or one or more storage systems. The total storage capacity from one or more of the following—storage device 356, storage system, cloud-based storage system 318, edge management service 366, cloud-based object storage device 348, or any other storage resource, or any combination or sub-combination of such storage resources—can be used to provide storage pool 383. Storage pool 383 is used to provide storage for one or more virtual volumes mounted on one or more nodes 390 within cluster 384.
[0235] In some implementations, container storage system 381 can create multiple storage pools. For example, container storage system 381 can aggregate storage resources of the same type into individual storage pools. In this example, the storage type can be one of the following: storage device 356, storage array 102, cloud-based storage system 318, storage via edge management service 366, or cloud-based object storage device 348. Alternatively, it can be storage configured with a certain level or type of redundancy or distribution, such as a specific combination of striping, mirroring, or erasure coding.
[0236] Container storage system 381 can be executed within cluster 384 as a containerized container storage system service, where examples of containers that implement the containerized container storage system service can operate on different nodes within cluster 384. In this example, the containerized container storage system service can combine the container orchestration system operations of container system 380 to handle storage operations, mount virtual volumes to provide storage to nodes, aggregate available storage into storage pool 383, provide storage for virtual volumes from storage pool 383, generate backup data, replicate data between nodes, clusters, and environments, and perform other storage system operations. In some instances, the containerized container storage system service can provide storage services across multiple clusters operating in different computing environments. For example, other storage system operations may include the storage system operations described herein. Persistent storage provided by the containerized container storage system service can be used to implement stateful and / or resilient containerized applications.
[0237] The container storage system 381 can be configured to perform any suitable storage operation of the storage system. For example, the container storage system 381 can be configured to perform one or more of the illustrative storage management operations described herein to manage storage resources used by the container system.
[0238] In some embodiments, one or more storage operations (including one or more of the illustrative storage management operations described herein) may be containerized. For example, one or more storage operations may be implemented as one or more containerized applications configured to perform storage operations. Such containerized storage operations can be executed in any suitable runtime environment to manage any storage system, including any illustrative storage system described herein.
[0239] Figure 3G This section describes an example of a storage node in a storage system architecture 395 for a large-scale storage platform according to embodiments of the present disclosure. The storage system architecture 395 includes nodes 390, as previously described... Figure 3F As described, and large-scale storage platform 396. It should be noted that in some embodiments, in addition to or in place of the components shown in storage system architecture 395, storage system architecture 395 may include, as previously described... Figures 1A to 3F Other storage system components described herein. For illustrative purposes, some components of storage system architecture 395 are not shown.
[0240] Node 390 includes a processing unit 397 that executes a scale-out platform component 398 and a flash device management component 399, which will be described in further detail below. Node 390 may also include storage resources 392, as previously described in... Figure 3F Description at the location.
[0241] The large-scale storage platform 396 may be a cloud-scale or hyperscale storage platform operatively coupled to node 390 via one or more network connections (not shown). The large-scale storage platform 396 can provide enterprise-level computing and / or storage resources across multiple data centers to consumers. In an embodiment, node 390 may store data at storage resource 392 on behalf of the large-scale storage platform 396.
[0242] Node 390 can be optimized for use by a large-scale storage platform 396 by moving the processing of large data blocks from individual storage resources 392 to processing devices 397 (e.g., flash device management components 399). This can be used to form a simple storage node service that decouples the software connecting node 390 to the software in the large-scale storage platform 396, which has similarities but many specific differences among other large-scale storage platforms, from the software that optimizes the use of directly managed and abstract-based management of flash storage devices (e.g., storage resources 392).
[0243] In one embodiment, a node 390 for a large-scale storage platform 396 may be a server with a large number of slots for inserting coupled storage devices (e.g., storage resource 392). In another embodiment, node 390 may be configured with a processing unit 397 executing a flexible server operating system with ample available RAM and network connectivity for connecting node 390 to the remainder of the large-scale storage platform 396 within and across its data center and geographic area. In some embodiments, node 390 may operate as needed by the large-scale storage platform 396. While node 390 may typically store individual fragments of widely distributed and erasure-coded data stripes, and additional computational needs related to serving databases, logging, and various services that can be used to run operations that facilitate large-scale storage services.
[0244] In some embodiments, node 390 may have a large number of connected flash storage devices, including storage devices that are directly managed and those managed based on abstractions, as well as processing devices and memory (not shown) on which software (e.g., scale-out platform component 398 and / or flash device management component 399) runs. In embodiments, scale-out platform component 398 may interact with large-scale storage platform 396 and other large-scale storage platform management infrastructure. Flash device management component 399 may manage storage resources 392 that present storage resources to scale-out platform component 398 in some form.
[0245] In this embodiment, a flexible bulk model can be implemented within node 390, and the performance of the storage system used in the large-scale storage platform 396 can be improved by utilizing flash memory devices with direct management or abstraction-based management. In conventional storage systems, storage devices typically have fixed-size DRAM for the storage table. This makes the storage device inflexible in terms of balancing block size and DRAM, as it is easier to manufacture a series of DRAM configurations for the server portion of the storage node than to manufacture a series of DRAM configurations for pluggable storage devices to be inserted into the hot-swappable drive bays of a server. Furthermore, DRAM occupies valuable space on the circuit board of the pluggable storage device, and it is preferable to place the DRAM in the motherboard of the housing rather than relocating the portion of the storage node / storage device combination to a location where the relative inaccessibility of the DRAM is not an issue.
[0246] In an example embodiment, the flash device management component 399 may interface with and manage a managed flash storage device (e.g., storage resource 392). The flash device management component 399 may provide services to optimally utilize storage resource 392, managing its capacity, performance, durability, lifespan, and / or partial hardware failures. The flash device management component 399 may be able to provide bulk storage services using the memory of processing device 397 and node 390. For example, the flash device management component 399 may present a set of block volumes to other software layers of node 390. In another example, a large-scale storage platform 396 may implement software within and across servers and storage nodes (e.g., node 390) of the large-scale storage platform 396's data center. Storage nodes may interact with the distributed storage platform to store and retrieve data blocks on behalf of the large-scale storage platform 396. Node 390 may then utilize the set of block volumes provided by the flash device management component 399 to write data segments into bulk blocks according to a bulk block model.
[0247] In some embodiments, the flash device management component 399 of node 390 may provide a set of differently optimized storage, such as some volumes supporting a certain amount of block-optimized storage and other volumes presenting a certain amount of small random-write-optimized storage for use with tables or databases required by the scale-out platform component 398. In some embodiments, the flash device management component 399 may provide a certain amount of SLC storage desired for storing data with heavy overwrite rates. In some embodiments, the flash device management component 399 may provide a certain amount of archive-level storage that uses less DRAM, assuming that the data stored in the archive-level storage is not frequently read and does not require low latency.
[0248] In some embodiments, node 390 may need to store a log, which may be a sequentially written file in which data segments of variable size are written. This can ultimately resemble a fragmented, large-block allocation model similar to a large-block allocation model, as long as it can induce the file system writing the log to allocate log file blocks that are appropriately large and properly aligned by logical block addresses.
[0249] In one embodiment, one mode in which the large-scale storage platform 396 can operate its storage nodes (e.g., node 390) is to run services integrated with the remainder of the large-scale storage platform on a simple file system (e.g., the XFS file system). This simple file system may be configured to preferentially allocate aligned extents for regular files (e.g., logs) and may be further configured to support fixed alignments and fixed-size extents on separate volumes for certain types of files (e.g., files that XFS calls live files), wherein some of these fixed alignments / fixed-size extents can be written in parallel for higher performance. The storage node then receives fragments or fragment strips, as well as chunks of log data and / or random write blocks for various types of databases, and converts the data into writes, for example, to one or more XFS file systems. Potentially, the writes of one or more fragments are written as parallel writes to aligned blocks of the XFS live file system, and then stored by XFS in a volume optimized for the writes. The scale-out platform component 398 may further utilize other XFS properties to write log chunks to preferred aligned extents. Fragmentation refers to larger and very wide stripes, with significant parity fragmentation both within and across data centers, as previously described. The role of the flash device management component 399 on node 390 is to receive various writes, store them as optimally as possible, and support the reading of any previously written data.
[0250] The XFS file system also supports issuing TRIM / UNMAPs for deleted blocks. In some embodiments, this can be utilized by the flash device management component 399 to eliminate the potential need to read previously written data from large blocks.
[0251] In some embodiments, node 390 may also require a certain amount of storage for booting the operating system. This may be provided by a separate storage resource (e.g., an SSD) or on-board flash storage. In some embodiments, this may be provided by storage resource 392 when the BIOS of the storage node is able to read the boot block without using flash device management component 399, since this layer may be unavailable before the operating system has booted and execution of flash device management component 399 has begun. To support this, the storage device controller on one or more of the flash storage devices of storage resource 392 may be configured to provide a namespace that supports random access to a relatively small boot file system. If only read support is required, this can be provided by a simple translation table, which is updated in the flash storage device while node 390 is running with flash device management component 399 intact, but the translation table is then used to support reads for booting before flash device management component 399 is booted later in the boot sequence.
[0252] It should be noted that although the storage system architecture 395 is shown as having a single node 390 and a single large-scale storage platform 396, embodiments of this disclosure may include any number of nodes and / or a large-scale storage platform.
[0253] The storage systems described in this paper can support various forms of data replication. For example, two or more storage systems can synchronously replicate datasets to each other. In synchronous replication, dissimilar copies of a particular dataset can be maintained by multiple storage systems, but all accesses to the dataset (e.g., reads) should produce consistent results regardless of which storage system the access is directed to. For example, a read directed to any of the storage systems synchronously replicating the dataset should return the same result. Therefore, while updates to the dataset version do not need to occur exactly simultaneously, precautions must be taken to ensure consistent access to the dataset. For example, if an update (e.g., a write) is received by the first storage system directed to the dataset, the update can only be confirmed as complete if all storage systems synchronously replicating the dataset have applied the update to their copies of the dataset. In this example, synchronous replication can be implemented using I / O forwarding (e.g., a write received at the first storage system is forwarded to the second storage system), communication between storage systems (e.g., each storage system instructs itself that it has completed the update), or other means.
[0254] In other embodiments, datasets can be replicated using checkpoints. In checkpoint-based replication (also known as 'near-synchronous replication'), a set of updates to the dataset (e.g., one or more write operations directed to the dataset) can occur between different checkpoints, such that the dataset is only updated to a particular checkpoint if all updates to the dataset prior to that checkpoint have been completed. Consider an instance where a first storage system stores a live copy of the dataset that a user is accessing. In this instance, assume that the dataset is copied from the first storage system to a second storage system using checkpoint-based replication. For example, the first storage system might send a first checkpoint (at time t = 0) to the second storage system, followed by a first set of updates to the dataset, then a second checkpoint (at time t = 1), then a second set of updates to the dataset, and then a third checkpoint (at time t = 2). In this instance, if the second storage system has executed all updates in the first set of updates but has not yet executed all updates in the second set of updates, then the copy of the dataset stored on the second storage system may be up-to-date up to the second checkpoint. Alternatively, if the second storage system has performed all updates in the first set of updates and the second set of updates, then the copy of the dataset stored on the second storage system can be up-to-date up to the third checkpoint. The reader will recognize that various types of checkpoints can be used (e.g., metadata-only checkpoints), and checkpoints can be distributed based on various factors (e.g., time, number of operations, RPO settings), etc.
[0255] In other embodiments, the dataset can be replicated via snapshot-based replication (also known as 'asynchronous replication'). In snapshot-based replication, snapshots of the dataset can be sent from a replication source (e.g., a first storage system) to a replication target (e.g., a second storage system). In this embodiment, each snapshot may contain the entire dataset or a subset of the dataset, such as (for example) only the portion of the dataset that has changed since the last snapshot was sent from the replication source to the replication target. The reader will understand that snapshots can be sent on demand based on a strategy that takes into account various factors (e.g., time, number of operations, RPO settings) or in some other way.
[0256] The storage systems described above can be configured, either individually or in combination, as continuous data protection storage. Continuous data protection storage is a feature of storage systems that records updates to a dataset in such a way that a consistent image of the dataset's previous contents can be accessed at a low temporal granularity (typically approximating seconds or even less) and extended backward over a reasonable period (typically hours or days). This allows access to the most recent consistent point in time of the dataset, and also allows access to points in time where an event may have just occurred (e.g., causing partial corruption or otherwise loss of the dataset), while maintaining a maximum number of updates close to that event. Conceptually, they are like a sequence of snapshots of a dataset taken very frequently and maintained for a long period, but continuous data protection storage is typically implemented quite differently from snapshots. Storage systems implementing continuous data protection storage can further provide means of accessing these points in time, accessing one or more of these points in time as snapshots or clones, or restoring the dataset back to one of these recorded points in time.
[0257] Over time, to reduce overhead, some points in time maintained in continuous data protection storage can be merged with other nearby points in time, essentially deleting some of these points from storage. This reduces the capacity required for storage updates. A limited number of these points in time can also be converted into snapshots of longer durations. For example, this storage can maintain a low-granularity sequence of points in time that go back several hours from the present, where some points in time are merged or deleted to reduce overhead by up to one extra day. Going back even further in the past, some of these points in time can be converted into snapshots representing a consistent image of points in time, just every few hours.
[0258] Although some embodiments are described primarily in the context of storage systems, readers in the art will recognize that embodiments of this disclosure may also take the form of a computer program product disposed on a computer-readable storage medium for use with any suitable processing system. Such computer-readable storage media can be any storage medium for machine-readable information, including magnetic media, optical media, solid-state media, or other suitable media. Examples of such media include disks in hard drives or floppy disks, optical disks on optical drives, magnetic tapes, and others as will be apparent to those skilled in the art. Those skilled in the art will readily recognize that any computer system with suitable programming elements will be able to perform the steps described herein as embodied in a computer program product. Those skilled in the art will also recognize that while some embodiments described in this specification are oriented towards software installed and executed on computer hardware, alternative embodiments implemented as firmware or hardware are also within the scope of this disclosure.
[0259] In some instances, a non-transitory computer-readable medium may be provided for storing computer-readable instructions, based on the principles described herein. When executed by a processor of a computing device, the instructions may direct the processor and / or the computing device to perform one or more operations, including one or more of the operations described herein. Such instructions may be stored and / or transmitted using any of a variety of known computer-readable media.
[0260] The term "non-transitory computer-readable media" as used herein may include any non-transitory storage medium that participates in providing data (e.g., instructions) that can be read and / or executed by a computing device (e.g., by a processor of the computing device). For example, non-transitory computer-readable media may include, but is not limited to, any combination of non-volatile storage media and / or volatile memory media. Exemplary non-volatile storage media include, but are not limited to, read-only memory, flash memory, solid-state drives, magnetic storage devices (e.g., hard disks, floppy disks, magnetic tapes, etc.), ferroelectric random access memory ("RAM"), and optical discs (e.g., optical discs, digital video discs, Blu-ray discs, etc.). Exemplary volatile storage media include, but are not limited to, RAM (e.g., dynamic RAM).
[0261] The advantages and features of this disclosure can be further described by the following statements:
[0262] 1. A method comprising: receiving, by a storage device controller of a storage device, one or more requests to store data in a flash memory portion of the storage device from a storage system controller; determining, based on information associated with the one or more requests, an indirect cell size to be used for mapping the data in a flash translation layer (FTL); storing the data in the flash memory portion of the storage device; and mapping the data in the FTL using the indirect cell size.
[0263] 2. The method according to statement 1, wherein the information associated with the one or more requests includes an indication of the indirect cell size received from the storage system controller.
[0264] 3. The method according to any one of statements 1 to 2, wherein the information associated with the one or more requests includes a logical address for storing the data, wherein the logical address falls within a logical address range allocated to use the indirection unit size.
[0265] 4. The method according to any one of statements 1 to 3, wherein the indirect unit size is selected based on the type of the data.
[0266] 5. The method according to any one of statements 1 to 4, wherein the storage device controller is configured to map multiple data stored in the flash memory portion using two or more indirect cell sizes.
[0267] 6. The method according to any one of statements 1 to 5, further comprising reallocating one or more blocks of the flash memory portion in the FTL that are mapped using the indirect cell size to use different indirect cell sizes in the FTL.
[0268] 7. The method according to any one of statements 1 to 6, wherein the one or more blocks of the flash memory portion are reallocated while avoiding moving data stored in the one or more blocks of the flash memory portion.
[0269] 8. The method according to any one of statements 1 to 7, wherein the storage device is a managed flash memory storage device.
[0270] 9. A storage device comprising: a flash memory portion; and a storage device controller including a flash translation layer (FTL) operatively coupled to the flash memory portion and configured to perform any of statements 1 to 8.
[0271] 10. A non-transitory computer-readable storage medium for storing instructions, which, when executed, cause a processing device of a storage device controller to perform any of statements 1 to 8.
[0272] Embodiments of this disclosure may include a storage system having a storage device configured to map data stored in the flash memory of the storage device using multiple indirection unit sizes. The storage device controller includes a flash translation layer (FTL) that acts as a layer mapping logical block addresses from the host side or file system to physical addresses in the flash memory of the storage device. The FTL may include indirect mappings or indirect layers that map logical block addresses of data to any underlying physical location in the flash memory. The indirect mappings may have indirection unit sizes corresponding to data objects of a specific size that can be accessed by specific references (e.g., the name, identifier, or pointer of the data object).
[0273] In conventional storage devices, the FTL (Fixed Indirect Unit) translates logical addresses to physical addresses using a single fixed indirect unit (IU) size. For example, an FTL might use a fixed IU size of 4 kilobytes (KiB). As the minimum write size for storing data at the storage device increases, the FTL can continue using the original IU size, leading to the need for new write pages that only overwrite portions of the write unit based on mapping. This increases the complexity and indirect mapping size of the FTL. Conversely, the storage device controller can perform a read-modify-write operation to update the original indirect unit size to a new, larger unit size, which can degrade the performance of the storage device.
[0274] Another issue is that large storage devices with high storage capacity require either very large indirect mappings with small IU sizes or inefficient indirect mappings with large IU sizes, which are coarse-grained mappings that can lead to wasted storage capacity and / or performance penalties. While conventional storage devices can update the entire FTL to use a larger IU size, this would be a slow and potentially inefficient process because subsequent smaller writes that would be handled more efficiently with a smaller IU size will not be handled efficiently with a larger IU size.
[0275] This disclosure provides an improved storage device and storage system by providing a storage device configured to maintain multiple IU sizes within a single storage device. The storage system may include one or more storage system controllers operatively coupled to multiple storage devices configured to maintain multiple IU sizes in a corresponding FTL (Flash Length Transmission) of the storage device. When the storage system controller transmits a command to store data in the storage device, the storage system controller may include an IU size to be used for mapping the data in the FTL. For example, the storage system controller may transmit a command to the storage device to store data in the flash memory of the storage device, and indicate that a 128 KiB IU size will be used to map the data in the FTL.
[0276] In some embodiments, the indication may be implicit. The storage device controller of the storage device may allocate a first logical address range using a first IU size and a second logical address range using a second IU size. For example, the storage device controller may allocate a first logical address range to use a first IU size of 4 KiB and allocate a second logical address range to use a second IU size of 128 KiB. In this embodiment, when the storage device controller receives a write request for stored data, the storage device controller may select the IU size based on a logical address provided as part of the write request for stored data. For example, if the write request for stored data contains logical addresses falling within the first logical address range, then the storage device controller may use a first IU size of 4 KiB to map the data in the FTL. In embodiments, these ranges may (in combination or individually) be larger than the physical size of the storage device (e.g., storage capacity) to avoid statically partitioning the physical space of the storage device.
[0277] Upon receiving a command, the storage device controller can store the data in the flash memory of the storage device. The storage device controller can use the IU size provided by the storage system controller to map physical blocks of flash memory to logical addresses in the FTL.
[0278] Embodiments of this disclosure provide an improved storage system by providing storage devices configured to use multiple IU sizes. By allowing the use of multiple IU sizes, the storage system controller can dynamically select the IU size for storing data based on various parameters. For example, the storage system controller can select a smaller IU size that has improved performance / latency for smaller data writes but higher resource usage for larger data writes, or select a larger IU size that has lower performance / latency for smaller data writes but lower resource usage for larger data writes. The storage system controller can dynamically select the IU size based on the size of the data being written and / or the IU sizes of aligned or logically adjacent data.
[0279] Figure 4 This is a description of an example of a storage system controller of a storage system 400 that generates commands for storing data comprising an indirect unit (IU) according to embodiments of the present disclosure. The storage system 400 may correspond to and include previously described in... Figure 1A to 3G One or more components of the described storage system. For clarity, some components of storage system 400 are not shown.
[0280] Storage system 400 includes a storage system controller 402 operatively coupled to storage device 404 via one or more network connections (not shown). Although a single storage system controller 402 and a single storage device (e.g., storage device 404) are shown, embodiments of storage system 400 may include any number of storage system controllers and / or storage devices. Storage device 404 may include a storage device controller 406 and flash memory 410. Storage device controller 406 includes a flash translation layer 408 that maps logical addresses to physical addresses in flash memory 410. Flash translation layer 408 may be configured to support multiple IU sizes, as previously described. In some embodiments, storage device 404 may be a managed flash memory device, as previously described.
[0281] Storage system controller 402 may receive various types of data to be stored in one or more storage devices of storage system 400. For example, storage system controller 402 may receive data and user-stored metadata from users of storage system 400, such as the logical address of data to be stored in one or more storage devices of storage system 400 (hereinafter also referred to as "user data"). Storage system controller 402 may also receive storage system metadata associated with data to be stored in one or more storage devices of storage system 400.
[0282] When the storage system controller 402 receives data and / or metadata to be stored at the storage device 404, the storage system controller 402 may generate a corresponding command or other type of instruction for the data and / or metadata. Upon receiving a command, the command may cause the storage device controller 406 to store the data associated with the command in the flash memory 410 and map the data stored in the flash memory 410 in the flash translation layer. The command may further include the IU size to be used for mapping the data in the flash translation layer 408.
[0283] refer to Figure 4 The storage system controller 402 has generated two commands (e.g., command 412a and command 412b). Command 412a contains a piece of data (e.g., data A) to be stored in the flash memory 410 and includes an indicated IU size (e.g., IU size 1). Command 412b contains another piece of data (e.g., data B) to be stored in the flash memory 410 and includes another indicated IU size different from IU size 1 (e.g., IU size 2). It should be noted that commands 412a to b are shown for illustrative purposes only and are not physical components of the storage system controller 402. When generating commands 412a to b, the storage system controller 402 may transmit commands 412a to b to the storage device 404, where commands 412a to b are received by the storage device controller 406.
[0284] In embodiments, the storage system controller 402 may select a specific IU size for the data based on one or more parameters associated with the data and / or storage system 400. In some embodiments, the storage system controller 402 may select the IU size based on the type of data to be stored in the flash memory. For example, the storage system controller 402 may select a first IU size for user data and a second IU size for metadata. In embodiments, the storage system controller 402 may select a smaller IU size for mapping metadata and a larger IU size for mapping user data. For example, the storage system controller 402 may select a 4 KiB IU size for mapping metadata and a 128 KiB IU size for mapping user metadata. It should be noted that although... Figure 4 The use of two different IU sizes (e.g., IU size 1 and IU size 2) is described, but embodiments of this disclosure may utilize any number of different IU sizes.
[0285] In some embodiments, the storage system controller 402 may have a defined IU size ratio or percentage that can be used to select the IU size to be used for mapping data. For example, storage device 404 may be specified to allocate 80% of its capacity to 128 KiB mapping and 20% of its capacity to 4 KiB mapping. In embodiments, the IU size ratio may be adjusted based on various storage system parameters. For example, if storage system 400 requires higher performance for smaller-sized writes, the amount of capacity mapped using smaller IU sizes (which provides increased performance for smaller-sized writes at the cost of resource usage) may be increased, and the amount of capacity mapped using larger IU sizes may be decreased. Conversely, if storage system 400 requires reduced resource usage, the amount of capacity mapped using larger IU sizes may be increased, and the amount of capacity mapped using smaller IU sizes may be decreased. In embodiments, adjustments to the amounts of different IU sizes may be performed dynamically and may be initiated by the storage system controller 402 or the controller of storage device 604.
[0286] Figure 5 This is an illustration of an example of a storage device controller in a storage system 500 that maps data stored in a flash memory using multiple IU sizes according to embodiments of the present disclosure. The storage system 500 may correspond to and include... Figure 4 One or more components of storage system 400. For clarity, some components of storage system 500 are not shown. Figure 5 In the middle, storage device 404 has received commands (e.g., from a storage system controller (not shown) having multiple IU sizes. Figure 4 Commands 412a to b), as previously described. The commands...
[0287] Upon receiving a command, the storage device controller 406 of storage device 404 may store the data contained in the command in one or more physical blocks of flash memory 410 of storage device 404. It should be noted that data A 502 and data B 504 are shown for illustrative purposes only and are not physical components of storage device 404 or flash memory 410.
[0288] refer to Figure 4Command 412a indicates that data A will have a first IU size, denoted as IU size 1, and command 412b indicates that data B will have a second IU size, denoted as IU size 2. When data A 502 and data B 504 are mapped in the flash translation layer 408, the storage device controller 406 can use the corresponding IU size indicated in the command to map the logical addresses of data A 502 and data B 504 to the physical blocks of the flash memory 410 storing data A 502 and data B 504. This can result in the flash translation layer 408 having one or more data items mapped using IU size 1 506 and one or more other data items mapped using IU size 2 508. Figure 5 In this example, the flash translation layer 408 has one piece of data (data A502) mapped using IU size 1 506 and another piece of data (data B 504) mapped using IU size 2 508. It should be noted that IU size 1 506 and IU size 2 508 are shown for illustrative purposes only and are not physical components of the storage device controller 406 or the flash translation layer 408.
[0289] Figure 6 This is an illustration of an example of a storage device controller in a storage system 600 that maps subsequent data using multiple IU sizes according to embodiments of the present disclosure. The storage system 600 may correspond to and include... Figure 5 Storage system 500 and Figure 4 One or more components of storage system 400. For clarity, some components of storage system 600 are not shown. Figure 6 In the above, storage device 404 has stored and mapped data A 502 and data B 504, as previously described.
[0290] refer to Figure 6 Storage device 404 may receive subsequent commands from one or more storage system controllers (not shown) to store additional data strips of different IU sizes that will be used to map additional data strips. One or more subsequent commands are used to store data C 602 and data D 604 in the flash memory 410 of storage device 404. The subsequent commands may instruct data C 602 and data D 604 to be stored using IU size 2 508. When data C 602 and data D 604 are stored in flash memory 410, storage device controller 406 maps data C 602 and data D 604 in flash translation layer 408 using IU size 2 508, resulting in data A being mapped in flash translation layer 408 using IU size 1 506 and data B to D being mapped in flash translation layer 408 using IU size 2 508.
[0291] Figure 7This is an illustration of an example of a storage device controller in a storage system 700 that maps subsequent data using multiple IU sizes according to embodiments of the present disclosure. The storage system 700 may correspond to and include... Figure 6 Storage system 600, Figure 5 Storage system 500 and Figure 4 One or more components of storage system 400. For clarity, some components of storage system 700 are not shown. Figure 6 In the storage device 404, data A 502, data B 504, data C 602 and data D 604 have been stored and mapped.
[0292] In some embodiments, the storage device controller 406 may be configured to dynamically modify the IU size used in the flash translation layer 408 while avoiding moving the underlying data stored in the flash memory 410. To this end, the storage device controller 406 may create one or more entries in the FTL with modified IU sizes corresponding to existing physical locations of the stored underlying data. Older entries using the original IU size may then be invalidated or made to reference newly created entries with modified IU sizes. In embodiments, the storage device controller 406 may modify the IU size based on receiving one or more commands from a storage system controller (not shown) instructing the storage device controller 406 to modify one or more IU sizes used for mapping data stored in the flash memory 410. For example, the storage system controller may transmit a command instructing the IU size used for mapping data bars to be changed from IU size 1 to IU size 2.
[0293] In some embodiments, the storage device controller 406 may modify one or more IU sizes based on the amount of resources available for use by the flash translation layer 408. For example, if the amount of memory available for storing the flash translation layer 408 is within a threshold of the maximum memory capacity, the storage device controller 406 may modify a smaller IU size to a larger IU size on one or more data entries to reduce the amount of memory required by the flash translation layer 408. This may result in the physical memory of the storage device 404 being reallocated from a smaller IU size to a larger IU size.
[0294] refer to Figure 7 The storage device controller 406 has modified the IU size used to map data B 504, data C 602, and data D 604 from IU size 2 508 to IU size 1 506. This results in data A through D each being mapped using IU size 1 506 in the flash translation layer 408, whereas there are currently no data strips mapped using IU size 2 508. Although the IU size has been modified, data A 502, data B 504, data C 602, and data D 604 can remain in the same physical block of the flash memory 410 where the data strips were originally stored.
[0295] Figure 8 This is an example method 800 for storing data in a storage device using multiple IU sizes according to embodiments of the present disclosure. Generally, method 800 can be executed by processing logic, which may include hardware (e.g., processing device, circuit system, dedicated logic, programmable logic, microcode, device hardware, integrated circuit, etc.), software (e.g., instructions that run or execute on the processing device), or a combination thereof. In some embodiments, Figure 4 The processing logic executable method 800 of the storage device controller 406.
[0296] Method 800 may begin at block 802, wherein the processing logic receives from the storage system controller one or more requests to store data in the flash memory portion of the storage device.
[0297] At block 804, the processing logic determines the size of the indirect cell to be used for mapping data in the flash translation layer based on information associated with one or more requests.
[0298] At block 806, the processing logic stores the data in the flash memory portion.
[0299] At block 808, the processing logic uses the indirect cell size to map data in the flash translation layer.
[0300] This document describes one or more embodiments with the aid of method steps illustrating the execution of specified functions and their relationships. For ease of description, the boundaries and sequences of these functional building blocks and method steps are arbitrarily defined herein. Alternative boundaries and sequences may be defined provided that the specified functions and relationships are properly performed. Any such alternative boundaries or sequences are therefore within the scope and spirit of the claims. Furthermore, for ease of description, the boundaries of these functional building blocks are arbitrarily defined. Alternative boundaries may be defined provided that a particular important function is properly performed. Similarly, flowchart blocks are also arbitrarily defined herein to illustrate specific important functionalities.
[0301] To the extent used, flowchart block boundaries and sequences may be defined in other ways while still performing specific important functions. Therefore, such alternative definitions of both functional building blocks and flowchart blocks and sequences are within the scope and spirit of the claims. Those skilled in the art will also recognize that the functional building blocks and other illustrative blocks, modules, and components described herein may be implemented as described, or by discrete components, application-specific integrated circuits, processors executing appropriate software, or any combination thereof.
[0302] While specific combinations of various functions and features of one or more embodiments are explicitly described herein, other combinations of these features and functions are also possible. This disclosure is not limited to the specific examples disclosed herein and is expressly incorporated into these other combinations.
Claims
1. A storage device comprising: Flash memory section; and A storage device controller, including a flash translation layer (FTL) operably coupled to the flash memory portion, the storage device controller being configured to: Receive one or more requests from the storage system controller to store data in the flash memory portion; The size of the indirect unit to be used for mapping the data in the FTL is determined based on information associated with the one or more requests; The data is stored in the flash memory portion; and The data is mapped in the FTL using the indirect unit size.
2. The storage device of claim 1, wherein the information associated with the one or more requests includes an indication of the indirect cell size received from the storage system controller.
3. The storage device of claim 1, wherein the information associated with the one or more requests includes a logical address for storing the data, wherein the logical address falls within a logical address range allocated to use the indirection unit size.
4. The storage device of claim 1, wherein the indirect unit size is selected based on the type of the data.
5. The storage device of claim 1, wherein the storage device controller is configured to use two or more indirect cell sizes to map multiple data stored in the flash memory portion.
6. The storage device of claim 1, wherein the storage device controller is further configured to: One or more blocks of the flash memory portion that are mapped using the indirect cell size in the FTL are reallocated to use different indirect cell sizes in the FTL.
7. The storage device of claim 6, wherein the one or more blocks of the flash memory portion are reallocated while avoiding moving data stored in the one or more blocks of the flash memory portion.
8. The storage device of claim 1, wherein the storage device is a managed flash memory storage device.
9. A method comprising: The storage device controller receives one or more requests from the storage system controller to store data in the flash memory portion of the storage device. The size of the indirect cell to be used for mapping the data in the Flash Translation Layer (FTL) is determined based on information associated with the one or more requests; The data is stored in the flash memory portion of the storage device; and The data is mapped using the indirect cell size in the FTL.
10. The method of claim 9, wherein the information associated with the one or more requests includes an indication of the indirect cell size received from the storage system controller.
11. The method of claim 9, wherein the information associated with the one or more requests includes a logical address for storing the data, wherein the logical address falls within a logical address range allocated to use the indirection unit size.
12. The method of claim 9, wherein the indirect unit size is selected based on the type of the data.
13. The method of claim 9, wherein the storage device controller is configured to use two or more indirect cell sizes to map multiple data stored in the flash memory portion.
14. The method of claim 9, further comprising: One or more blocks of the flash memory portion that are mapped using the indirect cell size in the FTL are reallocated to use different indirect cell sizes in the FTL.
15. The method of claim 14, wherein the one or more blocks of the flash memory portion are reallocated while avoiding moving data stored in the one or more blocks of the flash memory portion.
16. The method of claim 9, wherein the storage device is a managed flash memory storage device.
17. A non-transitory computer-readable storage medium for storing instructions, which, when executed, cause a processing apparatus of a storage device controller to: Receive one or more requests from the storage system controller to store data in the flash memory portion of the storage device; The size of the indirect cell to be used for mapping the data in the Flash Translation Layer (FTL) is determined based on information associated with the one or more requests; The data is stored in the flash memory portion of the storage device; and The data is mapped using the indirect cell size in the FTL.
18. The non-transitory computer-readable storage medium of claim 17, wherein the information associated with the one or more requests includes an indication of the indirect cell size received from the storage system controller.
19. The non-transitory computer-readable storage medium of claim 17, wherein the information associated with the one or more requests includes a logical address for storing the data, wherein the logical address falls within a logical address range allocated to use the indirection unit size.
20. The non-transitory computer-readable storage medium of claim 17, wherein the indirect unit size is selected based on the type of the data.