Data storage system with managed flash memory
Through the direct-mapped flash storage system, the operating system directly manages the flash drive, solving the complexity and latency issues of flash drives in traditional storage systems and achieving more efficient data writing and system reliability.
Patent Information
- Application Number
- CN202480014312.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2024-02-27
- Filing Date
- 2024-02-28
- Publication Date
- 2025-10-03
AI Technical Summary
Existing solid-state drives have difficulty effectively utilizing the unique characteristics of flash memory, making it difficult to provide enhanced storage features. In addition, flash drives in traditional storage systems require a storage controller to perform address translation, which increases complexity and latency.
A direct-mapped flash storage system is used to directly manage the flash drive through the operating system, avoiding the storage controller from performing address translation and directly addressing data blocks, thereby improving the reliability and efficiency of the flash drive.
It simplifies the operation process of the flash storage system, improves data writing speed and system reliability, reduces unnecessary writing operations, and improves the overall performance of the storage system.
Smart Images

Figure CN120752620A_ABST
Abstract
Description
Background Art
[0001] A storage system can include different types of storage devices or memories. Solid-state memory, such as flash memory, is currently used in solid-state drives (SSDs) to enhance or replace conventional hard disk drives (HDDs), writable CD (compact disk) or writable DVD (digital versatile disk) drives (collectively referred to as rotating media), and tape drives for storing large amounts of data. Flash memory and other solid-state memories have different characteristics than rotating media. However, for compatibility reasons, many solid-state drives are designed to conform to hard disk drive standards, which makes it difficult to provide enhanced features or take advantage of the unique aspects of flash memory and other solid-state memories.
[0002] The embodiments are presented in this context. BRIEF DESCRIPTION OF THE DRAWINGS
[0003] Figure 1A A first example system for data storage is described.
[0004] Figure 1B A second example system for data storage is described.
[0005] Figure 1C A third example system for data storage is described.
[0006] Figure 1D A fourth example system for data storage is described.
[0007] Figure 2A is a perspective diagram of a storage cluster having multiple storage nodes and internal storage devices coupled to each storage node to provide network attached storage.
[0008] Figure 2B is a block diagram illustrating an interconnect switch coupling a plurality of storage nodes according to some embodiments.
[0009] Figure 2C is a multi-level block diagram showing the contents of a storage node and the contents of one of the non-volatile solid-state storage units.
[0010] Figure 2D A storage server environment is shown using embodiments of the storage nodes and storage units of some previous figures according to some embodiments.
[0011] Figure 2E It is a hardware block diagram showing the control plane, computing and storage planes, and authorized blades that interact with the underlying physical resources.
[0012] Figure 2F Describes the elastic software layer in the blades of a storage cluster.
[0013] Figure 2GDescribes the authorization and storage resources in the blades of the storage cluster.
[0014] Figure 3A A diagram presenting a storage system coupled for data communication with a cloud service provider according to some embodiments of the present disclosure.
[0015] Figure 3B A diagram illustrating a storage system.
[0016] Figure 3C State examples of cloud-based storage systems.
[0017] Figure 3D An exemplary computing device is illustrated that may be specifically configured to perform one or more of the processes described herein.
[0018] Figure 3E An example of a fleet of storage systems for providing storage services (also referred to herein as "data services") is illustrated.
[0019] Figure 3F Describes an example of a container system.
[0020] Figure 4 is a block diagram illustrating an example storage system according to some embodiments of the present disclosure.
[0021] Figure 5 is a block diagram illustrating an example primary storage node and an example secondary storage node 420 according to some embodiments of the present disclosure.
[0022] Figure 6 is a block diagram illustrating an example storage system according to some embodiments of the present disclosure.
[0023] Figure 7 is a block diagram illustrating an example storage system according to some embodiments of the present disclosure.
[0024] Figure 8 is a flowchart illustrating a method for performing a storage operation according to some embodiments of the present disclosure.
[0025] Figure 9 is a block diagram illustrating an example storage system according to some embodiments of the present disclosure.
[0026] Figure 10 is a block diagram illustrating an example storage system according to some embodiments of the present disclosure.
[0027] Figure 11 is a block diagram illustrating an example storage system according to some embodiments of the present disclosure.
[0028] Figure 12 is a flowchart illustrating a method for accessing data according to some embodiments of the present disclosure.
[0029] Figure 13 is a flowchart illustrating a method for configuring a flash memory according to some embodiments of the present disclosure.
[0030] Figure 14 is a block diagram illustrating an instance store according to one or more embodiments of the present disclosure.
[0031] Figure 15 is a block diagram illustrating an instance storage node according to one or more embodiments of the present disclosure.
[0032] Figure 16 is a flow chart illustrating a method for accessing data storage according to some embodiments of the present disclosure. DETAILED DESCRIPTION
[0033] by Figure 1A Initially, example methods, apparatuses, and products for data storage systems according to embodiments of the present disclosure are described with reference to the accompanying drawings. Figure 1A An example system for data storage according to some embodiments is described. For purposes of illustration and not limitation, system 100 (also referred to herein as a "storage system") includes numerous elements. It should be noted that in other embodiments, system 100 may include the same, more, or fewer elements configured in the same or different manners.
[0034] The system 100 includes a plurality of computing devices 164A-B. The computing devices (also referred to herein as "client devices") may be embodied as, for example, servers in a data center, workstations, personal computers, laptops, or the like. The computing devices 164A-B may be coupled for data communication with one or more storage arrays 102A-B via a storage area network ('SAN') 158 or a local area network ('LAN') 160.
[0035] SAN 158 can be implemented using a variety of data communication architectures, devices, and protocols. For example, the architecture used for SAN 158 can include Fibre Channel, Ethernet, InfiniBand, Serial Attached Small Computer System Interface ('SAS'), or the like. Data communication protocols used with SAN 158 can include Advanced Technology Attachment ('ATA'), Fibre Channel protocol, Small Computer System Interface ('SCSI'), Internet Small Computer System Interface ('iSCSI'), Hyper-SCSI, Non-Volatile Memory Express ('NVMe') architecture, or the like. It should be noted that SAN 158 is provided for illustration and not limitation. Other data communication couplings can be implemented between computer devices 164A-B and storage arrays 102A-B.
[0036] LAN 160 may also be implemented using a variety of architectures, devices, and protocols. For example, the architecture used for LAN 160 may include Ethernet (802.3), wireless (802.11), or the like. Data communication protocols used in LAN 160 may include Transmission Control Protocol ('TCP'), User Datagram Protocol ('UDP'), Internet Protocol ('IP'), Hypertext Transfer Protocol ('HTTP'), Wireless Access Protocol ('WAP'), Handheld Device Transport Protocol ('HDTP'), Session Initiation Protocol ('SIP'), Real-Time Protocol ('RTP'), or the like. LAN 160 may also be connected to the Internet 162.
[0037] Storage arrays 102A-B may provide persistent data storage for computing devices 164A-B. In some embodiments, storage array 102A may be housed in a chassis (not shown), and storage array 102B may be housed in another chassis (not shown). Storage arrays 102A and 102B may include one or more storage array controllers 110A-D (also referred to herein as "controllers"). Storage array controllers 110A-D may be embodied as modules of automated computing machinery including computer hardware, computer software, or a combination of computer hardware and software. In some embodiments, storage array controllers 110A-D may be configured to perform various storage tasks. Storage tasks may include writing data received from computing devices 164A-B to storage arrays 102A-B, erasing data from storage arrays 102A-B, retrieving data from storage arrays 102A-B and providing data to computing devices 164A-B, monitoring and reporting storage device utilization and performance, performing redundancy operations (such as redundant array of independent drives ('RAID') or RAID-like data redundancy operations), compressing data, encrypting data, and the like.
[0038] The storage array controllers 110A-D may be implemented in various ways, including as a field programmable gate array ('FPGA'), a programmable logic chip ('PLC'), an application specific integrated circuit ('ASIC'), a system on a chip ('SOC'), or any computing device including discrete components such as a processing device, a central processing unit, computer memory, or various adapters. The storage array controllers 110A-D may include, for example, data communications adapters configured to support communications via the SAN 158 or the LAN 160. In some embodiments, the storage array controllers 110A-D may be independently coupled to the LAN 160. In some embodiments, the storage array controllers 110A-D may include an I / O controller or the like that couples the storage array controllers 110A-D for data communications to persistent storage resources 170A-B (also referred to herein as "storage resources") via a midplane (not shown). The persistent storage resources 170A-B may include any number of storage drives 171A-F (also referred to herein as "storage devices") and any number of non-volatile random access memory ('NVRAM') devices (not shown).
[0039] In some embodiments, the NVRAM devices of persistent storage resources 170A-B may be configured to receive data from storage array controllers 110A-D to be stored in storage drives 171A-F. In some examples, the data may originate from computing devices 164A-B. In some examples, writing data to the NVRAM devices may be faster than writing the data directly to storage drives 171A-F. In some embodiments, storage array controllers 110A-D may be configured to utilize the NVRAM devices as a quickly accessible buffer for data intended for writing to storage drives 171A-F. The latency of write requests using the NVRAM devices as buffers may be improved compared to systems in which storage array controllers 110A-D write data directly to storage drives 171A-F. In some embodiments, the NVRAM devices may be implemented using computer memory in the form of high-bandwidth, low-latency RAM. NVRAM devices are referred to as "non-volatile" because they may receive or contain a sole source of power to maintain the RAM's state after the NVRAM device loses main power. This power source may be a battery, one or more capacitors, or the like. In response to a loss of power, the NVRAM device may be configured to write the contents of the RAM to a persistent storage device, such as storage drives 171A-F.
[0040] In some implementations, storage drives 171A-F may refer to any device configured to persistently record data, where "persistently" or "persistent" refers to the ability of the device to maintain the recorded data after a loss of power. In some implementations, storage drives 171A-F may correspond to non-disk storage media. For example, storage drives 171A-F may be one or more solid-state drives ('SSDs'), flash memory-based storage devices, any type of solid-state non-volatile memory, or any other type of non-mechanical storage device. In other implementations, storage drives 171A-F may include a mechanical or rotating hard disk, such as a hard disk drive ('HDD').
[0041] In some embodiments, the storage array controllers 110A-D may be configured to offload device management responsibilities from the storage drives 171A-F in the storage arrays 102A-B. For example, the storage array controllers 110A-D may manage control information that may describe the status of one or more memory blocks in the storage drives 171A-F. The control information may indicate, for example, that a particular memory block has failed and should no longer be written to, that a particular memory block contains boot code for the storage array controllers 110A-D, the number of program-erase ('P / E') cycles that have been performed on a particular memory block, the age of the data stored in the particular memory block, the type of data stored in the particular memory block, and so on. In some embodiments, the control information may be stored as metadata along with the associated memory blocks. In other embodiments, the control information for the storage drives 171A-F may be stored in one or more specific memory blocks of the storage drives 171A-F selected by the storage array controllers 110A-D. The selected memory blocks may be tagged with an identifier indicating that the selected memory blocks contain the control information. The identifier can be used by the storage array controllers 110A-D in conjunction with the storage drives 171A-F to quickly identify a memory block containing control information. For example, the storage controllers 110A-D can issue a command to locate a memory block containing control information. It can be noted that the control information may be so large that portions of the control information may be stored in multiple locations, that the control information may be stored in multiple locations, for example, for redundancy purposes, or that the control information may be otherwise distributed across multiple memory blocks in the storage drives 171A-F.
[0042] In some implementations, the storage array controllers 110A-D can offload device management responsibilities from the storage drives 171A-F of the storage arrays 102A-B by retrieving control information from the storage drives 171A-F that describes the state of one or more memory blocks in the storage drives 171A-F. Retrieving the control information from the storage drives 171A-F can be accomplished, for example, by the storage array controllers 110A-D querying the storage drives 171A-F for the location of the control information for a particular storage drive 171A-F. The storage drives 171A-F can be configured to execute instructions that enable the storage drives 171A-F to identify the location of the control information. The instructions can be executed by a controller (not shown) associated with or otherwise located on the storage drives 171A-F and can cause the storage drives 171A-F to scan a portion of each memory block to identify the memory blocks that store the control information for the storage drives 171A-F. The storage drives 171A-F may respond by sending a response message containing the location of the control information for the storage drives 171A-F to the storage array controllers 110A-D. In response to receiving the response message, the storage array controllers 110A-D may issue a request to read data stored at the address associated with the location of the control information for the storage drives 171A-F.
[0043] In other embodiments, the storage array controllers 110A-D can further offload device management responsibilities from the storage drives 171A-F by performing storage drive management operations in response to receiving the control information. The storage drive management operations can include, for example, operations typically performed by the storage drives 171A-F, e.g., a controller (not shown) associated with the particular storage drives 171A-F. The storage drive management operations can include, for example, ensuring that data is not written to failed memory blocks within the storage drives 171A-F, ensuring that data is written to memory blocks within the storage drives 171A-F in a manner such that adequate wear leveling is achieved, and the like.
[0044] In some embodiments, storage arrays 102A-B may implement two or more storage array controllers 110A-D. For example, storage array 102A may include storage array controller 110A and storage array controller 110B. In a given example, a single storage array controller 110A-D of storage system 100 (e.g., storage array controller 110A) may be designated as having primary status (also referred to herein as the "primary controller"), and other storage array controllers 110A-D (e.g., storage array controller 110B) may be designated as having secondary status (also referred to herein as the "secondary controller"). The primary controller may have certain rights, such as permission to change data in persistent storage resources 170A-B (e.g., write data to persistent storage resources 170A-B). At least some of the rights of the primary controller may supersede the rights of the secondary controllers. For example, when the primary controller has permission to change data in persistent storage resources 170A-B, the secondary controller may not have such rights. The status of storage array controllers 110A-D may change. For example, storage array controller 110A may be designated as having a secondary status, and storage array controller 110B may be designated as having a primary status.
[0045] In some embodiments, a primary controller (e.g., storage array controller 110A) may serve as the primary controller for one or more storage arrays 102A-B, and a secondary controller (e.g., storage array controller 110B) may serve as the secondary controller for one or more storage arrays 102A-B. For example, storage array controller 110A may be the primary controller for both storage array 102A and storage array 102B, and storage array controller 110B may be the secondary controller for both storage arrays 102A and 102B. In some embodiments, storage array controllers 110C and 110D (also referred to as "storage processing modules") may not have primary or secondary status. Storage array controllers 110C and 110D, implemented as storage processing modules, may serve as the communication interface between the primary and secondary controllers (e.g., storage array controllers 110A and 110B, respectively) and storage array 102B. For example, storage array controller 110A of storage array 102A may send write requests to storage array 102B via SAN 158. The write request may be received by both storage array controllers 110C and 110D of storage array 102B. Storage array controllers 110C and 110D facilitate communication, for example, sending the write request to the appropriate storage drives 171A through F. It may be noted that in some embodiments, storage processing modules may be used to increase the number of storage drives controlled by the primary and secondary controllers.
[0046] In some implementations, the storage array controllers 110A-D are communicatively coupled to one or more storage drives 171A-F and to one or more NVRAM devices (not shown) included as part of the storage arrays 102A-B via a mid-plane (not shown). The storage array controllers 110A-D may be coupled to the mid-plane via one or more data communication links, and the mid-plane may be coupled to the storage drives 171A-F and the NVRAM devices via one or more data communication links. For example, the data communication links described herein are collectively illustrated by data communication links 108A-D and may include a Peripheral Component Interconnect Express ('PCIe') bus.
[0047] Figure 1B An example system for data storage is described according to some embodiments. Figure 1B The storage array controller 101 described in Figure 1A 1. Storage array controllers 110A to D are described in detail below. In one example, storage array controller 101 may be similar to storage array controller 110A or storage array controller 110B. For purposes of illustration and not limitation, storage array controller 101 includes numerous components. It is noted that in other embodiments, storage array controller 101 may include the same, more, or fewer components configured in the same or different manners. It is noted that the following may include Figure 1A 1 to help illustrate the features of storage array controller 101.
[0048] Storage array controller 101 may include one or more processing devices 104 and random access memory (RAM) 111. Processing device 104 (controller 101) represents one or more general-purpose processing devices, such as microprocessors, central processing units, or the like. More specifically, processing device 104 (or controller 101) may be a complex instruction set computing (CISC) microprocessor, a reduced instruction set computing (RISC) microprocessor, a very long instruction word (VLIW) microprocessor, or a processor that implements another instruction set or a combination of instruction sets. Processing device 104 (controller 101) may also be one or more special-purpose processing devices, such as an ASIC, an FPGA, a digital signal processor (DSP), a network processor, or the like.
[0049] Processing device 104 may be connected to RAM 111 via a data communication link 106, which may be embodied as a high-speed memory bus, such as a double data rate 4 ('DDR4') bus. Stored in RAM 111 is an operating system 112. In some embodiments, instructions 113 are stored in RAM 111. Instructions 113 may include computer program instructions for performing operations in a direct-mapped flash storage system. In one embodiment, a direct-mapped flash storage system is a system that directly addresses data blocks within a flash drive without requiring address translation performed by the flash drive's storage controller.
[0050] In some embodiments, the storage array controller 101 includes one or more host bus adapters 103A-C coupled to the processing device 104 via data communication links 105A-C. In some embodiments, the host bus adapters 103A-C may be computer hardware that connects a host system (e.g., a storage array controller) to other networks and storage arrays. In some examples, the host bus adapters 103A-C may be Fibre Channel adapters that enable the storage array controller 101 to connect to a SAN, Ethernet adapters that enable the storage array controller 101 to connect to a LAN, or the like. The host bus adapters 103A-C may be coupled to the processing device 104 via data communication links 105A-C, such as a PCIe bus, for example.
[0051] In some embodiments, storage array controller 101 may include a host bus adapter 114 coupled to an expander 115. Expander 115 may be used to attach a host system to a greater number of storage drives. In embodiments where host bus adapter 114 is embodied as a SAS controller, expander 115 may be, for example, a SAS expander for enabling host bus adapter 114 to attach to storage drives.
[0052] In some embodiments, the storage array controller 101 may include a switch 116 coupled to the processing device 104 via a data communication link 109. The switch 116 may be a computer hardware device that can create multiple endpoints from a single endpoint, thereby enabling multiple devices to share a single endpoint. The switch 116 may be, for example, a PCIe switch that couples to a PCIe bus (e.g., the data communication link 109) and presents multiple PCIe connection points to the midplane.
[0053] In some implementations, storage array controller 101 includes a data communication link 107 for coupling storage array controller 101 to other storage array controllers. In some examples, data communication link 107 may be a Quick Path Interconnect (QPI) interconnect.
[0054] A conventional storage system using conventional flash drives can implement processes across the flash drives that are part of the conventional storage system. For example, a higher-level process of the storage system can initiate and control processes across the flash drives. However, the flash drives of the conventional storage system may include their own storage controllers that also execute the processes. Thus, for a conventional storage system, both higher-level processes (e.g., initiated by the storage system) and lower-level processes (e.g., initiated by the storage system's storage controller) can be executed.
[0055] To address various shortcomings of conventional storage systems, operations can be performed by higher-level processes rather than lower-level processes. For example, a flash storage system may include a flash drive that does not include a storage controller that provides the processes. Therefore, the flash storage system's own operating system can initiate and control the processes. This can be accomplished with a direct-mapped flash storage system that directly addresses data blocks within the flash drive, eliminating the need for address translation performed by the flash drive's storage controller.
[0056] In some embodiments, storage drives 171A-F may be one or more partitioned storage devices. In some embodiments, one or more partitioned storage devices may be shingled HDDs. In some embodiments, one or more storage devices may be flash-based SSDs. In a partitioned storage device, the partition namespace on the partitioned storage device may be addressed by groups of blocks that are grouped and aligned according to a natural size, thereby forming a number of addressable regions. In some embodiments utilizing an SSD, the natural size may be based on the erase block size of the SSD. In some embodiments, the regions of the partitioned storage device may be defined during initialization of the partitioned storage device. In some embodiments, the regions may be dynamically defined as data is written to the partitioned storage device.
[0057] In some embodiments, zones may be heterogeneous, with some zones each being a page group and other zones being multiple page groups. In some embodiments, some zones may correspond to an erase block and other zones may correspond to multiple erase blocks. In embodiments, zones may be any combination of different numbers of pages in page groups and / or erase blocks for a heterogeneous mix of programming modes, manufacturers, product types, and / or product generations of storage devices, such as for heterogeneous assembly, upgrades, distributed storage, and the like. In some embodiments, zones may be defined as having usage characteristics, such as properties that support data with a particular type of endurance (e.g., very short life or very long life). These properties may be used by the partitioned storage device to determine how the zone will be managed over its expected lifetime.
[0058] It should be understood that zones are virtual constructs. Any particular zone may not have a fixed location at the storage device. Prior to allocation, a zone may not have any location at the storage device. A zone may correspond to a number that represents a block of virtually allocatable space, which is the size of an erase block or other block size in various embodiments. When the system allocates or opens a zone, the zone is allocated to flash memory or other solid-state storage memory, and as the system writes to the zone, pages are written to that mapped flash memory or other solid-state storage memory of the partitioned storage device. When the system closes a zone, the associated erase block or other sized block is completed. At some point in the future, the system may delete the zone, which will free up the allocated space of the zone. During its lifetime, for example, when the partitioned storage device undergoes internal maintenance, a zone may be moved to a different location on the partitioned storage device.
[0059] In some embodiments, the zones of a partitioned storage device may be in different states. A zone may be in an empty state, where data has not yet been stored in the zone. An empty zone may be explicitly or implicitly opened by writing data to the zone. This is the initial state of a zone on a newly partitioned storage device, but it may also be the result of a zone reset. In some embodiments, an empty zone may have a specified location within the flash memory of the partitioned storage device. In an embodiment, the location of the zone may be selected when the empty zone is first opened or written to (or later if the write is buffered in memory). A zone may be explicitly or implicitly in an open state, where a zone in an open state may be written to using a write or append command to store data. In an embodiment, a zone in an open state may also be written to using a copy command that copies data from different zones. In some embodiments, a partitioned storage device may have a limit on the number of open zones at a particular time.
[0060] A zone in a closed state is a zone that has been partially written to but has been placed in a closed state after an explicit close operation was issued. A zone in a closed state may be left available for future writes, but may reduce some of the runtime overhead consumed by keeping the zone in an open state. In some embodiments, a partitioned storage device may have a limit on the number of closed zones at a particular time. A zone in a full state is a zone that is storing data and cannot be written to anymore. A zone may be in a full state after a write has written data to the entire zone or due to a zone complete operation. Prior to a complete operation, a zone may or may not have been completely written to. However, after a complete operation, it may not be possible to open the zone for further writes without first performing a zone reset operation.
[0061] The mapping from zones to erase blocks (or to shingled tracks in an HDD) can be arbitrary, dynamic, and hidden from view. The process of opening a zone can be an operation that allows a new zone to be dynamically mapped to the underlying storage of the partitioned storage device and then allows data to be written to the zone by additional writes to the zone until the zone reaches capacity. A zone can be completed at any time, after which additional data may not be written to the zone. When the data stored at a zone is no longer needed, the zone can be reset, which effectively deletes the contents of the zone from the partitioned storage device, making the physical storage held by that zone available for subsequent data storage. Once a zone has been written and completed, the partitioned storage device ensures that the data stored at the zone will not be lost until the zone is reset. In the time between writing data to the zone and resetting the zone, the zone can be moved between shingled tracks or erase blocks as part of maintenance operations within the partitioned storage device, such as by copying data to keep the data refreshed or to handle memory cell aging in an SSD.
[0062] In some embodiments utilizing HDDs, resetting a zone may allow shingled tracks to be assigned to a new open zone that may be opened at some future time. In some embodiments utilizing SSDs, resetting a zone may cause the zone's associated physical erase blocks to be erased and subsequently reused for data storage. In some embodiments, a zoned storage device may have a limit on the number of open zones at a point in time to reduce the amount of overhead dedicated to keeping zones open.
[0063] The operating system of the flash storage system can identify and maintain a list of allocation units across multiple flash drives of the flash storage system. An allocation unit can be a whole erase block or multiple erase blocks. The operating system can maintain a mapping or address range that directly maps addresses to erase blocks of the flash drives of the flash storage system.
[0064] The erase blocks mapped directly to the flash drive can be used to rewrite and erase data. For example, an operation can be performed on one or more allocation units containing first and second data, where the first data is to be retained and the second data is no longer used by the flash storage system. The operating system can initiate a process to write the first data to a new location within another allocation unit, erase the second data, and mark the allocation unit as available for subsequent data. Thus, the process can be executed only by the higher-level operating system of the flash storage system, and additional lower-level processes do not need to be executed by the flash drive's controller.
[0065] Advantages of having the process executed solely by the flash storage system's operating system include improved reliability of the flash storage system's flash drives, as no unnecessary or redundant write operations are performed during the process. A potential novelty here is the concept of initiating and controlling the process within the flash storage system's operating system. Furthermore, the process can be controlled by the operating system across multiple flash drives. This is in contrast to processes being executed by the flash drive's storage controller.
[0066] A storage system may consist of two storage array controllers that share a set of drives for failover purposes, or it may consist of a single storage array controller that provides storage services that utilize multiple drives, or it may consist of a distributed network of storage array controllers, each with a certain number of drives or a certain amount of flash storage, where the storage array controllers in the network cooperate to provide a complete storage service and cooperate in all aspects of the storage service including storage allocation and waste collection.
[0067] Figure 1C A third example system 117 for data storage according to some embodiments is described. For purposes of illustration and not limitation, system 117 (also referred to herein as a "storage system") includes numerous elements. It should be noted that in other embodiments, system 117 may include the same, more, or fewer elements configured in the same or different manners.
[0068] In one embodiment, system 117 includes dual peripheral component interconnect ('PCI') flash memory devices 118 with individually addressable fast write storage. System 117 may include a storage device controller 119. In one embodiment, storage device controllers 119A-D may be a CPU, ASIC, FPGA, or any other circuitry that can implement the necessary control structures according to the present disclosure. In one embodiment, system 117 includes flash memory devices (e.g., including flash memory devices 120a-n) operably coupled to various channels of storage device controller 119. Flash memory devices 120a-n may be presented to controllers 119A-D as an addressable collection of flash memory pages, erase blocks, and / or control elements sufficient to allow storage device controllers 119A-D to program and retrieve various aspects of the flash memory. In one embodiment, the storage device controllers 119A to D can perform operations on the flash memory devices 120a to n, including storing and retrieving the data contents of a page, arranging and erasing any block, tracking statistics related to the use and reuse of flash memory pages, erase blocks and cells, tracking and predicting error codes and failures within the flash memory, controlling voltage levels associated with programming and retrieving the contents of flash memory cells, etc.
[0069] In one embodiment, system 117 may include RAM 121 to store individually addressable, fast-write data. In one embodiment, RAM 121 may be one or more separate discrete devices. In another embodiment, RAM 121 may be integrated into storage controller 119A-D or multiple storage controllers. RAM 121 may also be used for other purposes, such as temporary program storage for a processing device (e.g., a CPU) in storage controller 119.
[0070] In one embodiment, system 117 may include an energy storage device 122, such as a rechargeable battery or capacitor. Energy storage device 122 may store enough energy to power storage controller 119, a certain amount of RAM (e.g., RAM 121), and a certain amount of flash memory (e.g., flash memories 120a-120n) for a sufficient time to write the contents of the RAM to the flash memory. In one embodiment, if the storage controller detects a loss of external power, storage controllers 119A-D may write the contents of the RAM to the flash memory.
[0071] In one embodiment, the system 117 includes two data communication links 123a, 123b. In one embodiment, the data communication links 123a, 123b may be PCI interfaces. In another embodiment, the data communication links 123a, 123b may be based on other communication standards (e.g., HyperTransport, InfiniBand, etc.). The data communication links 123a, 123b may be based on the Non-Volatile Memory Express ('NVMe') or NVMe-based Architecture ('NVMf') specifications, which allow external connections to the storage device controllers 119A-D from other components in the storage system 117. It should be noted that for convenience, the data communication links may be interchangeably referred to herein as PCI buses.
[0072] System 117 may also include an external power supply (not shown), which may be provided via one or both of data communication links 123a, 123b, or may be provided separately. An alternative embodiment includes a separate flash memory (not shown) dedicated to storing the contents of RAM 121. Storage controllers 119A-D may present a logical device (which may include an addressable fast write logical device) or a distinct portion of the logical address space of storage 118 (which may be presented as PCI memory or as a persistent storage device) via the PCI bus. In one embodiment, store operations to the device are directed to RAM 121. In the event of a power failure, storage controllers 119A-D may write the stored contents associated with the addressable fast write logical storage to flash memory (e.g., flash memory 120a-n) for long-term persistent storage.
[0073] In one embodiment, the logic device may include some representation of some or all of the contents of flash memory devices 120a-n, wherein the representation allows a storage system (e.g., storage system 117) including storage device 118 to directly address flash memory pages and directly reprogram erase blocks from storage system components external to the storage device over a PCI bus. The representation may also allow one or more of the external components to control and retrieve other aspects of the flash memory, including some or all of the following: tracking statistics related to the use and reuse of flash memory pages, erase blocks, and cells across all flash memory devices; tracking and predicting error codes and failures within and across flash memory devices; controlling voltage levels associated with programming and retrieving the contents of flash cells; etc.
[0074] In one embodiment, the energy storage device 122 may be sufficient to ensure that ongoing operations on the flash memory devices 120a to 120n are completed. The energy storage device 122 may power the storage device controllers 119A to D and the associated flash memory devices (e.g., 120a to n) for those operations, as well as for storing fast write RAM to the flash memory. The energy storage device 122 may be used to store accumulated statistics and other parameters maintained and tracked by the flash memory devices 120a to n and / or the storage device controller 119. A separate capacitor or energy storage device (e.g., a smaller capacitor near or embedded within the flash memory device itself) may be used for some or all of the operations described herein.
[0075] Various schemes may be used to track and optimize the lifespan of the energy storage component, such as adjusting voltage levels over time, partially discharging the energy storage device 122 to measure corresponding discharge characteristics, etc. If the available energy decreases over time, the effective available capacity of the addressable fast write storage device may be reduced to ensure it can be safely written to based on the currently available stored energy.
[0076] Figure 1D A third example storage system 124 for data storage according to some embodiments is illustrated. In one embodiment, the storage system 124 includes storage controllers 125a, 125b. In one embodiment, the storage controllers 125a, 125b are operatively coupled to dual PCI storage devices. The storage controllers 125a, 125b are operatively coupled (e.g., via a storage network 130) to a number of host computers 127a through n.
[0077] In one embodiment, two storage controllers (e.g., 125a and 125b) provide storage services, such as SCS block storage arrays, file servers, object servers, databases, or data analysis services. Storage controllers 125a, 125b can provide services to host computers 127a to n external to storage system 124 through a number of network interfaces (e.g., 126a to d). Figure 1B 101, storage controllers 125a and 125b may each include one or more processing devices (e.g., one or more processors, CPUs, processing cores, etc.) and memory (e.g., RAM, NVRAM, cache, etc.). Storage controllers 125a, 125b may provide integrated services or applications entirely within storage system 124, thereby forming a converged storage and computing system. Storage controllers 125a, 125b may utilize fast write memory within or across storage devices 119a-d to journal ongoing operations to ensure that operations are not lost in the event of a power failure, storage controller removal, storage controller or storage system shutdown, or a failure of one or more software or hardware components within storage system 124.
[0078] In one embodiment, the storage controllers 125a, 125b operate as PCI masters for one or the other PCI buses 128a, 128b. In another embodiment, 128a, 128b may be based on other communication standards (e.g., HyperTransport, InfiniBand, etc.). Other storage system embodiments may operate the storage controllers 125a, 125b as multi-masters for both PCI buses 128a, 128b. Alternatively, a PCI / NVMe / NVMf switching infrastructure or architecture may connect multiple storage controllers. Some storage system embodiments may allow storage devices to communicate directly with each other, rather than just with the storage controllers. In one embodiment, the storage device controller 119a may operate under guidance from the storage controller 125a to retrieve data from a computer that has been stored in RAM (e.g., Figure 1C The recalculated version of the RAM contents may be synthesized with data from the RAM 121 of the storage controller and transferred to the flash memory device. For example, the recalculated version of the RAM contents may be transferred after the storage controller has determined that the operation has been fully committed across the storage system or when the fast write memory on the device has reached a certain usage capacity or after a certain amount of time to ensure improved security of the data or to free up addressable fast write capacity for reuse. This mechanism may be used, for example, to avoid a second transfer from the storage controller 125a, 125b via a bus (e.g., 128a, 128b). In one embodiment, the recalculation may include compressing the data, appending indexes or other metadata, combining multiple data segments together, performing erasure code calculations, etc.
[0079] In one embodiment, under direction from the storage controllers 125a, 125b, the storage controllers 119a, 119b are operable to retrieve data from a computer stored in RAM (e.g., Figure 1C The memory controller 125a and 125b can be used to calculate data from the data in the RAM 121 and transfer the data to other memory devices without involving the memory controllers 125a and 125b. This operation can be used to mirror data stored in one memory controller 125a to another memory controller 125b, or it can be used to offload compression, data aggregation and / or erasure coding calculations and transfers to the memory devices to reduce the load on the memory controllers or the memory controller interfaces 129a and 129b to the PCI buses 128a and 128b.
[0080] The storage controllers 119A-D may include mechanisms for implementing high availability primitives for use by other components of the storage system external to the dual PCI storage device 118. For example, a reservation or exclusion primitive may be provided so that, in a storage system having two storage controllers providing highly available storage services, one storage controller may prevent the other storage controller from accessing or continuing to access the storage. This may be used, for example, in situations where one controller detects that the other controller is not functioning properly or where the interconnect between the two storage controllers itself may not be functioning properly.
[0081] In one embodiment, a storage system for use with dual PCI direct-mapped storage devices having separately addressable fast-write storage includes a system for managing erase blocks or groups of erase blocks as allocation units for storing data on behalf of a storage service or for storing metadata associated with the storage service (e.g., indexes, logs, etc.), or for proper management of the storage system itself. Flash pages, which may be several kilobytes in size, may be written as data arrives or when the storage system will retain the data for a long time interval (e.g., exceeding a defined time threshold). To commit data more quickly, or to reduce the number of writes to the flash memory devices, the storage controller may first write the data to separately addressable fast-write storage devices on one or more storage devices.
[0082] In one embodiment, the storage controllers 125a, 125b may initiate the use of erase blocks within and across storage devices, such as 118, based on the age and expected remaining useful life of the storage devices, or based on other statistics. The storage controllers 125a, 125b may initiate garbage collection and data migration between storage devices based on pages that are no longer needed and to manage flash page and erase block lifespans and to manage overall system performance.
[0083] In one embodiment, the storage system 124 may utilize mirroring and / or erasure coding schemes as part of storing data in addressable fast-write storage devices and / or as part of writing data to allocation units associated with erase blocks. Erasure coding may be used across storage devices and within erase blocks or allocation units, or within and across flash memory devices on a single storage device, to provide redundancy against single or multiple storage device failures or to protect against corruption within flash memory pages caused by flash memory operations or by flash memory cell degradation. Mirroring and erasure coding at various levels may be used to recover from multiple types of failures occurring alone or in combination.
[0084] refer to Figure 2A The embodiments depicted through G illustrate a storage cluster that stores user data, such as user data originating from one or more users or client systems or other sources external to the storage cluster. The storage cluster distributes user data across storage nodes housed within a chassis or across multiple chassis using erasure coding and redundant copies of metadata. Erasure coding refers to a method of data protection or reconstruction in which data is stored across a set of different locations, such as disks, storage nodes, or geographic locations. Flash memory is one type of solid-state memory that may be integrated with embodiments, but embodiments may be extended to other types of solid-state memory or other storage media, including non-solid-state memory. Control of storage locations and workloads is distributed across storage locations in a clustered peer system. Tasks such as mediating communications between storage nodes, detecting when a storage node becomes unavailable, and balancing I / O (input and output) across storage nodes are all handled on a distributed basis. In some embodiments, data is laid out or distributed across multiple storage nodes in data segments or stripes that support data recovery. Ownership of data may be reassigned within the cluster, independent of the type of input and output. This architecture, described in more detail below, allows a storage node in a cluster to fail while the system remains operational because data can be reconstructed from other storage nodes and thus remain available for input and output operations. In various embodiments, a storage node may be referred to as a cluster node, blade, or server.
[0085] The storage cluster may be housed within a chassis (i.e., a housing that houses one or more storage nodes). The mechanism for providing power to each storage node (e.g., a power distribution bus) and the communication mechanism (e.g., a communication bus that enables communication between storage nodes) are contained within the chassis. According to some embodiments, the storage cluster may operate as a standalone system in one location. In one embodiment, the chassis houses at least two instances of both power distribution and communication buses that can be enabled or disabled independently. The internal communication bus may be an Ethernet bus, however, other technologies such as PCIe, InfiniBand, and others are also suitable. The chassis provides ports for an external communication bus for communication between multiple chassis and with client systems, either directly or through a switch. External communication may use technologies such as Ethernet, InfiniBand, Fibre Channel, etc. In some embodiments, the external communication bus uses different communication bus technologies for inter-chassis and client communication. If a switch is deployed within or between chassis, the switch may serve as a translator between multiple protocols or technologies. When multiple chassis are connected to define a storage cluster, the storage cluster can be accessed by clients using a proprietary interface or a standard interface, such as Network File System ('NFS'), Common Internet File System ('CIFS'), Small Computer System Interface ('SCSI'), or Hypertext Transfer Protocol ('HTTP'). Protocol translation from the client can occur at a switch, a chassis external communication bus, or within each storage node. In some embodiments, multiple chassis can be coupled or connected to each other through an aggregator switch. Some and / or all of the coupled or connected chassis can be designated as a storage cluster. As discussed above, each chassis can have multiple blades, each blade having a Media Access Control ('MAC') address, but in some embodiments, the storage cluster appears to the external network as having a single cluster IP address and a single MAC address.
[0086] Each storage node may be one or more storage servers, and each storage server is connected to one or more non-volatile solid-state memory units, which may be referred to as storage units or storage devices. One embodiment includes a single storage server in each storage node and between one and eight non-volatile solid-state memory units, however, this example is not intended to be limiting. The storage server may include a processor, DRAM, and interfaces for an internal communication bus and power distribution for each of the power buses. In some embodiments, within the storage node, the interface and storage units share a communication bus, such as PCI Express. The non-volatile solid-state memory units can directly access the internal communication bus interface via the storage node communication bus or request access to the bus interface from the storage node. The non-volatile solid-state memory unit contains an embedded CPU, a solid-state storage controller, and a certain amount of solid-state mass storage, for example, between 2 and 32 terabytes (TB) in some embodiments. Embedded volatile storage media (such as DRAM) and energy storage devices are included in the non-volatile solid-state memory unit. In some embodiments, the energy storage device is a capacitor, supercapacitor, or battery that enables a subset of the DRAM contents to be transferred to a stable storage medium in the event of a power loss. In some embodiments, the non-volatile solid-state memory cell is constructed with storage class memory such as phase change or magnetoresistive random access memory ('MRAM') that replaces DRAM and enables reduced power retention devices.
[0087] One of the many features of storage nodes and non-volatile solid-state storage devices is the ability to proactively rebuild data in a storage cluster. The storage nodes and non-volatile solid-state storage devices can determine when a storage node or non-volatile solid-state storage device in a storage cluster is unreachable, independent of whether an attempt is made to read data involving that storage node or non-volatile solid-state storage device. The storage nodes and non-volatile solid-state storage devices then cooperate to recover and rebuild the data in at least a portion of the new location. This constitutes proactive reconstruction because the system does not need to wait until the data is needed for a read access initiated from a client system employing the storage cluster before rebuilding the data. These and other details of the storage memory and its operation are discussed below.
[0088] Figure 2AFIG2 is a perspective view of a storage cluster 161 according to some embodiments, having multiple storage nodes 150 and internal solid-state storage coupled to each storage node to provide a network-attached storage device or storage area network (SAN). A NAS, SAN, storage cluster, or other storage storage may include one or more storage clusters 161, each with one or more storage nodes 150, providing a flexible and reconfigurable arrangement of both the physical components and the amount of storage provided. Storage clusters 161 are designed to fit within racks, and one or more racks can be configured and populated with storage as needed. Storage cluster 161 includes a chassis 138 with multiple slots 142. It should be understood that chassis 138 may be referred to as an enclosure, housing, or rack unit. In one embodiment, chassis 138 has fourteen slots 142, but other numbers of slots can be readily designed. For example, some embodiments have four slots, eight slots, sixteen slots, thirty-two slots, or another suitable number of slots. In some embodiments, each slot 142 can accommodate one storage node 150. The chassis 138 includes fins 148 that can be used to mount the chassis 138 on a rack. Fans 144 provide air circulation for cooling the storage nodes 150 and their components, but other cooling components can be used, or embodiments without cooling components can be designed. The switch fabric 146 couples the storage nodes 150 within the chassis 138 together and to a network for communicating with the storage. In the embodiment depicted herein, for illustrative purposes, the slots 142 to the left of the switch fabric 146 and fans 144 are shown as being occupied by storage nodes 150, while the slots 142 to the right of the switch fabric 146 and fans 144 are empty and can be used to insert storage nodes 150. This configuration is an example, and in various other arrangements, one or more storage nodes 150 can occupy slots 142. In some embodiments, the storage node arrangement need not be sequential or adjacent. The storage nodes 150 are hot-swappable, meaning that they can be inserted into or removed from the slots 142 in the chassis 138 without stopping or powering down the system. After a storage node 150 is inserted or removed from a slot 142, the system automatically reconfigures itself to recognize and adapt to the change. In some embodiments, reconfiguration includes restoring storage redundancy and / or rebalancing data or load.
[0089] Each storage node 150 may have multiple components. In the embodiment shown here, the storage node 150 includes a printed circuit board 159 populated by a CPU 156 (i.e., a processor), a memory 154 coupled to the CPU 156, and a non-volatile solid-state storage device 152 coupled to the CPU 156, but in alternative embodiments, other mountings and / or components may be used. The memory 154 contains instructions executed by the CPU 156 and / or data operated on by the CPU 156. As further explained below, the non-volatile solid-state storage device 152 includes flash memory, or in alternative embodiments, includes other types of solid-state memory.
[0090] refer to Figure 2A , the storage cluster 161 is scalable, meaning that storage capacity with non-uniform storage sizes can be easily added, as described above. One or more storage nodes 150 can be inserted into or removed from each chassis, and in some embodiments, the storage cluster self-configures. The plug-in storage nodes 150 can have different sizes whether they are installed in the chassis at delivery or added later. For example, in one embodiment, the storage node 150 can have any multiple of 4TB, such as 8TB, 12TB, 16TB, 32TB, etc. In other embodiments, the storage node 150 can have any multiple of other storage amounts or capacities. The storage capacity of each storage node 150 is broadcast and affects the decision of how to stripe data. For maximum storage efficiency, embodiments can self-configure in stripes as wide as possible, subject to predetermined continuous operation requirements, with the loss of up to one or up to two non-volatile solid-state storage device 152 units or storage nodes 150 within the chassis.
[0091] Figure 2B 1 is a block diagram showing the communication interconnect 173 and the power distribution bus 172 coupling the plurality of storage nodes 150. Figure 2A In some embodiments, the communication interconnect 173 may be included in or implemented with the switch fabric 146. In the event that multiple storage clusters 161 occupy a rack, in some embodiments, the communication interconnect 173 may be included in or implemented with the top of the rack switch. Figure 2B , the storage cluster 161 is enclosed within a single chassis 138. External ports 176 are coupled to storage nodes 150 via communication interconnects 173, while external ports 174 are directly coupled to storage nodes. External power ports 178 are coupled to power distribution bus 172. Storage nodes 150 may include varying amounts and capacities of non-volatile solid-state storage devices 152, as described in reference to FIG. Figure 2A In addition, one or more storage nodes 150 may be compute-only storage nodes, such as Figure 2B. The authorization 168 is implemented on the non-volatile solid-state storage device 152, for example as a list or other data structure stored in memory. In some embodiments, the authorization is stored within the non-volatile solid-state storage device 152 and is supported by software executed on a controller or other processor of the non-volatile solid-state storage device 152. In another embodiment, the authorization 168 is implemented on the storage node 150, for example as a list or other data structure stored in memory 154 and is supported by software executed on the CPU 156 of the storage node 150. In some embodiments, the authorization 168 controls how data is stored in the non-volatile solid-state storage device 152 and where the data is stored in the non-volatile solid-state storage device 152. This control helps determine which type of erasure coding scheme is applied to the data and which storage nodes 150 have which parts of the data. Each authorization 168 can be assigned to a non-volatile solid-state storage device 152. In various embodiments, each authorization may control a range of inode numbers, segment numbers, or other data identifiers assigned to data by the file system, by the storage node 150 , or by the non-volatile solid-state storage device 152 .
[0092] In some embodiments, each piece of data and each piece of metadata has redundancy in the system. In addition, each piece of data and each piece of metadata has an owner, which can be called a grant. If that grant is unreachable, for example due to a storage node failure, there is a succession plan for how to find that data or that metadata. In various embodiments, there are redundant copies of grants 168. In some embodiments, grants 168 are associated with storage nodes 150 and non-volatile solid-state storage devices 152. Each grant 168 covering a range of data segment numbers or other identifiers of data can be assigned to a specific non-volatile solid-state storage device 152. In some embodiments, grants 168 for all such ranges are distributed across the non-volatile solid-state storage devices 152 of the storage cluster. Each storage node 150 has a network port that provides access to the non-volatile solid-state storage device 152 of that storage node 150. Data can be stored in segments, and in some embodiments, the segments are associated with segment numbers and that segment number is the indirection of the configuration of a RAID (Redundant Array of Independent Disks) stripe. The assignment and use of grants 168 thus establish indirection for the data. According to some embodiments, indirectness may be referred to as the ability to reference data indirectly (in this case, via authorization 168). A segment identifies a set of non-volatile solid-state storage devices 152 and a local identifier to the set of non-volatile solid-state storage devices 152 that may contain data. In some embodiments, the local identifier is an offset into the device and may be reused sequentially by multiple segments. In other embodiments, the local identifier is unique for a particular segment and is never reused. The offset in the non-volatile solid-state storage device 152 is applied to locate data for writing to the non-volatile solid-state storage device 152 or reading from the non-volatile solid-state storage device 152 (in the form of a RAID stripe). Data is striped across multiple units of the non-volatile solid-state storage device 152, which may include a non-volatile solid-state storage device 152 having an authorization 168 for a particular data segment or a non-volatile solid-state storage device 152 that is different from the non-volatile solid-state storage device 152.
[0093] If the location where a particular data segment is located changes, for example, during data movement or data reconstruction, the authorization 168 for that data segment should be consulted at the non-volatile solid-state storage device 152 or storage node 150 that has that authorization 168. In order to locate a particular piece of data, an embodiment calculates a hash value of the data segment or applies an index node number or data segment number. The output of this operation points to the non-volatile solid-state storage device 152 that has the authorization 168 for that particular piece of data. In some embodiments, there are two stages for this operation. The first stage maps an entity identifier (ID) (such as a segment number, index node number, or directory number) to an authorization identifier. This mapping may include, for example, the calculation of a hash or bit mask. The second stage is to map the authorization identifier to a specific non-volatile solid-state storage device 152, which can be accomplished by explicit mapping. The operation is repeatable so that when the calculation is performed, the result of the calculation can repeatedly and reliably point to the specific non-volatile solid-state storage device 152 that has that authorization 168. The operation may include a set of reachable storage nodes as input. If the set of reachable non-volatile solid-state storage units changes, then the best set changes. In some embodiments, the value maintained is the current assignment (which is always true) and the calculated value is the target assignment that the cluster will attempt to reconfigure. This calculation can be used to determine the best non-volatile solid-state storage device 152 for authorization when there is a group of non-volatile solid-state storage devices 152 that are reachable and constitute the same cluster. The calculation also determines an ordered set of peer non-volatile solid-state storage devices 152, which will also record the authorizations mapped to the non-volatile solid-state storage devices so that authorizations can be determined even when the assigned non-volatile solid-state storage device is unreachable. In some embodiments, if a particular authorization 168 is not available, a duplicate or alternative authorization 168 can be consulted.
[0094] refer to Figure 2A and 2B, two of the many tasks of the CPU 156 on the storage node 150 are to decompose write data and reassemble read data. When the system has determined that data is to be written, the authorization 168 for that data is located as described above. When the segment ID of the data has been determined, the write request is forwarded to the non-volatile solid-state storage device 152 of the host that is currently determined to be the authorization 168 determined from the segment. Then, the host CPU 156 of the storage node 150 on which the non-volatile solid-state storage device 152 and the corresponding authorization 168 reside decomposes or splits the data and transmits the data out to the various non-volatile solid-state storage devices 152. The transmitted data is written as data stripes according to the erasure coding scheme. In some embodiments, the request pulls the data, and in other embodiments, the data is pushed. Conversely, when data is read, the authorization 168 for the segment ID containing the data is located as described above. The host CPU 156 of the storage node 150 on which the non-volatile solid-state storage device 152 and the corresponding grant 168 reside requests data from the non-volatile solid-state storage device and the corresponding storage node pointed to by the grant. In some embodiments, the data is read from the flash storage device as a data stripe. The host CPU 156 of the storage node 150 then reassembles the read data, correcting any errors (if any) according to an appropriate erasure coding scheme and forwarding the reassembled data to the network. In other embodiments, some or all of these tasks can be handled in the non-volatile solid-state storage device 152. In some embodiments, the segment host requests data to be sent to the storage node 150 by requesting a page from the storage device and then sending the data to the storage node that issued the original request.
[0095] In an embodiment, authorization 168 operates to determine how operations are to be performed for specific logical elements. Operations can be performed on each of the logical elements using specific authorizations across multiple storage controllers of the storage system. Authorization 168 can communicate with multiple storage controllers so that the multiple storage controllers can collectively perform operations on those specific logical elements.
[0096] In embodiments, a logical element may be, for example, a file, a directory, an object bucket, an individual object, a descriptive portion of a file or object, or other forms of key-value databases or tables. In embodiments, performing an operation may involve, for example, ensuring the consistency, structural integrity, and / or recoverability of other operations on the same logical element, reading metadata and data associated with that logical element, determining what data should be persistently written to the storage system to preserve any changes to the operation, or metadata and data may be determined where to store modular storage devices across multiple storage controllers attached to the storage system.
[0097] In some embodiments, operations are token-based transactions that are efficiently communicated within a distributed system. Each transaction may be accompanied by or associated with a token that provides permission to execute the transaction. In some embodiments, authorization 168 can maintain the pre-transaction state of the system until the operation is completed. Token-based communication can be accomplished without global locks across the system and can also restart operations in the event of corruption or other failures.
[0098] In some systems, such as UNIX-style file systems, data is handled by index nodes (inodes), which specify data structures representing objects in the file system. For example, an object may be a file or a directory. Metadata may accompany the object as attributes, such as permission data and a creation timestamp, among other attributes. Segment numbers may be assigned to all or part of this object in the file system. In other systems, data segments are handled by segment numbers assigned elsewhere. For the purposes of this discussion, the unit of distribution is an entity, and an entity may be a file, a directory, or a segment. That is, an entity is a unit of data or metadata stored by the storage system. Entities are grouped into sets called grants. Each grant has a grant owner, which is a storage node that has exclusive rights to update the entities in the grant. In other words, a storage node contains a grant, and a grant contains an entity.
[0099] According to some embodiments, a segment is a logical container for data. A segment is an address space between the media address space and the physical flash memory locations, i.e., the data segment numbers are in this address space. A segment may also contain metadata that enables data redundancy to be restored (rewritten to a different flash memory location or device) without involving higher level software. In one embodiment, the internal format of a segment contains the client data and the media map to determine the location of that data. Where applicable, each data segment is protected, for example, from the effects of memory and other failures, by breaking it into a number of data and parity slices. The data and parity slices are distributed (i.e., striped) across the non-volatile solid-state storage device 152 coupled to the host CPU 156 according to an erasure coding scheme (see Figure 2E and 2G In some embodiments, the term segment is used to refer to a container and its location in the address space of the segment. In some embodiments, the term stripe is used to refer to the same set of shards as a segment and includes how the shards and redundancy or parity information are distributed.
[0100] A series of address space transformations occur across the entire storage system. At the top are directory entries (file names) that link to index nodes. Index nodes point to the media address space where data is logically stored. Media addresses can be mapped through a series of indirect media to expand the load of large files or implement data services such as deduplication or snapshots. Next, the segment address is translated into a physical flash memory location. According to some embodiments, the physical flash memory location has an address range limited by the amount of flash memory in the system. Media addresses and segment addresses are logical containers, and in some embodiments use 128-bit or larger identifiers so that they are effectively infinite, where the possibility of reuse is calculated to be longer than the expected life of the system. In some embodiments, addresses from logical containers are allocated in a hierarchical manner. Initially, each non-volatile solid-state storage device 152 unit can be assigned an address space range. Within this assigned range, the non-volatile solid-state storage device 152 can assign addresses without synchronizing with other non-volatile solid-state storage devices 152.
[0101] Data and metadata are stored by a set of underlying storage layouts optimized for different workload types and storage devices. These layouts incorporate a variety of redundancy schemes, compression formats, and indexing algorithms. Some of these layouts store information about authorizations and authorization masters, while other layouts store file metadata and file data. Redundancy schemes include error correction codes that tolerate corrupted bits within a single storage device (e.g., a NAND flash chip), erasure codes that tolerate failures of multiple storage nodes, and replication schemes that tolerate data center or regional failures. In some embodiments, low-density parity check ('LDPC') codes are used within a single storage cell. In some embodiments, Reed-Solomon encoding is used within a storage cluster, and mirroring is used within a storage grid. Metadata can be stored using an ordered log-structured index (e.g., a log-structured merge tree), and large data may not be stored in a log-structured layout.
[0102] To maintain consistency across multiple copies of an entity, storage nodes implicitly agree on two things through computation: (1) the grants that contain the entity, and (2) the storage nodes that contain the grants. The assignment of entities to grants can be accomplished by pseudo-randomly assigning entities to grants, by splitting entities into ranges based on an externally generated key, or by placing a single entity into each grant. Examples of pseudo-random schemes are linear hashing and the hashing of the Replication Under Scalable Hashing ('RUSH') family, including Controlled Replication Under Scalable Hashing ('CRUSH'). In some embodiments, pseudo-random assignment is only used to assign grants to nodes because the set of nodes can change. The set of grants cannot change, so any subjective function can be applied in these embodiments. Some placement schemes automatically place grants on storage nodes, while other placement schemes rely on an explicit mapping of grants to storage nodes. In some embodiments, a pseudo-random scheme is used to map from each grant to a set of candidate grant owners. A pseudo-random data distribution function related to CRUSH can assign grants to storage nodes and create a list of where to assign grants. Each storage node has a copy of the pseudo-random data distribution function and can derive the same calculations for distribution and later lookup or location of the grant. In some embodiments, each of the pseudo-random schemes requires a set of reachable storage nodes as input in order to derive the same target node. Once an entity has been placed in a grant, it can be stored on a physical device so that expected failures will not result in unexpected data loss. In some embodiments, the rebalancing algorithm attempts to store replicas of all entities within a grant in the same layout and on the same set of machines.
[0103] Examples of expected failures include device failure, machine theft, data center fire, and regional disasters such as nuclear or geological events. Different failures result in varying degrees of acceptable data loss. In some embodiments, the theft of a storage node affects neither the security nor the reliability of the system, while regional events may result in no data loss, a few seconds or minutes of lost updates, or even complete data loss, depending on the system configuration.
[0104] In an embodiment, the placement of data for storage redundancy is independent of the placement of authorization for data consistency. In some embodiments, the storage node containing the authorization does not contain any persistent storage device. Instead, the storage node is connected to a non-volatile solid-state storage unit that does not contain the authorization. The communication interconnection between the storage node and the non-volatile solid-state storage unit consists of a variety of communication technologies and has non-uniform performance and fault tolerance characteristics. In some embodiments, as mentioned above, the non-volatile solid-state storage unit is connected to the storage node via PCI Express, the storage nodes are connected together in a single chassis using an Ethernet backplane, and the chassis are connected together to form a storage cluster. In some embodiments, the storage cluster is connected to the client using Ethernet or Fibre Channel. If multiple storage clusters are configured into a storage grid, then the multiple storage clusters are connected using the Internet or other long-distance networking links (such as "metro scale" links or private links that do not traverse the Internet).
[0105] The authorized owner has exclusive rights to modify the entity, migrate the entity from one non-volatile solid-state storage unit to another non-volatile solid-state storage unit, and add and remove copies of the entity. This allows the redundancy of the underlying data to be maintained. When the authorized owner fails, is to be decommissioned, or is overloaded, the authorization is transferred to a new storage node. Transient failures make it very important to ensure that all non-faulty machines agree on the new authorized location. Ambiguities caused by transient failures can be automatically implemented through consensus protocols (such as Paxos) and hot-warm failover schemes via manual intervention by a remote system administrator or by a local hardware administrator (for example, by physically removing the failed machine from the cluster, or pressing a button on the failed machine). In some embodiments, a consensus protocol is used and failover is automatic. According to some embodiments, if too many failures or replication events occur in too short a time period, the system enters a self-protection mode and stops replication and data movement activities until the administrator intervenes.
[0106] When authorization is transferred between a storage node and an authorization owner to update an entity's authorization, the system transmits messages between the storage node and the non-volatile solid-state storage unit. Regarding persistent messages, messages with different purposes have different types. Depending on the type of message, the system maintains different sorting and durability guarantees. When processing persistent messages, messages are temporarily stored using a variety of persistent and non-persistent storage hardware technologies. In some embodiments, messages are stored in RAM, NVRAM, and on NAND flash devices, using various protocols to efficiently utilize each storage medium. Latency-sensitive client requests can be persisted in replicated NVRAM and then, later, persisted in NAND, while background rebalancing operations are persisted directly to NAND.
[0107] Persistent messages are persistently stored before being transmitted. This allows the system to continue servicing client requests despite failures and component replacements. While many hardware components contain unique identifiers visible to system administrators, manufacturers, the hardware supply chain, and ongoing monitoring and quality control infrastructure, applications running on infrastructure addresses virtualize these addresses. These virtualized addresses do not change over the life of the storage system, regardless of component failures and replacements. This allows each component of the storage system to be replaced over time without requiring reconfiguration or interrupting client request processing; in other words, the system supports non-disruptive upgrades.
[0108] In some embodiments, virtualized addresses are stored with sufficient redundancy. A continuous monitoring system correlates hardware and software status with hardware identifiers. This allows for the detection and prediction of failures due to faulty components and manufacturing details. In some embodiments, the monitoring system can also proactively transfer authorization and entities away from affected devices before a failure occurs by removing components from critical paths.
[0109] Figure 2C is a multi-level block diagram showing the contents of a storage node 150 and the contents of the non-volatile solid-state storage devices 152 of the storage node 150. In some embodiments, data is passed to and from the storage node 150 by a network interface controller ('NIC') 202. Each storage node 150 has a CPU 156 and one or more non-volatile solid-state storage devices 152, as discussed above. Figure 2C Moving down one level in the non-volatile solid-state storage device 152, each non-volatile solid-state storage device 152 has relatively fast non-volatile solid-state memory, such as non-volatile random access memory ('NVRAM') 204 and flash memory 206. In some embodiments, NVRAM 204 may be a component that does not require program / erase cycles (DRAM, MRAM, PCM) and may be a memory that can support writes much more frequently than reads from the memory. Figure 2CMoving down another level in the NVRAM 204, in one embodiment, NVRAM 204 is implemented as high-speed volatile memory, such as dynamic random access memory (DRAM) 216, backed up by an energy reserve 218. The energy reserve 218 provides sufficient power to keep the DRAM 216 powered long enough to transfer its contents to the flash memory 206 in the event of a power failure. In some embodiments, the energy reserve 218 is a capacitor, supercapacitor, battery, or other device that provides a suitable energy supply sufficient to enable the contents of the DRAM 216 to be transferred to a stable storage medium in the event of a power loss. The flash memory 206 is implemented as a plurality of flash die 222, which may be referred to as a package of flash die 222 or an array of flash die 222. It should be understood that the flash die 222 can be packaged in any number of ways, including a single die per package, multiple die per package (i.e., a multi-chip package), in a hybrid package, as a bare die on a printed circuit board or other substrate, as an encapsulated die, and the like. In the embodiment shown, the non-volatile solid-state storage device 152 has a controller 212 or other processor and an input / output (I / O) port 210 coupled to the controller 212. The I / O port 210 is coupled to the CPU 156 and / or the network interface controller 202 of the flash storage node 150. A flash input / output (I / O) port 220 is coupled to a flash die 222, and a direct memory access unit (DMA) 214 is coupled to the controller 212, DRAM 216, and the flash die 222. In the embodiment shown, the I / O port 210, the controller 212, the DMA unit 214, and the flash I / O port 220 are implemented on a programmable logic device (PLD) 208, such as an FPGA. In this embodiment, each flash die 222 has pages organized into 16kB (kilobyte) pages 224 and registers 226 through which data can be written to or read from the flash die 222. In further embodiments, other types of solid-state memory are used in place of or in addition to the flash memory illustrated within flash die 222 .
[0110] In various embodiments disclosed herein, storage cluster 161 can be generally contrasted with a storage array. Storage nodes 150 are part of the collection that creates storage cluster 161. Each storage node 150 owns a slice of data and the computations required to provide that data. Multiple storage nodes 150 collaborate to store and retrieve data. As typically used in storage arrays, storage memory or storage devices are less involved in processing and manipulating data. The storage memory or storage devices in a storage array receive commands to read, write, or erase data. The storage memory or storage devices in a storage array are unaware of the larger system in which they are embedded or what the data means. The storage memory or storage devices in a storage array can include various types of storage memory, such as RAM, solid-state drives, hard disk drives, etc. The non-volatile solid-state storage device 152 unit described herein has multiple interfaces that are simultaneously active and serve multiple purposes. In some embodiments, certain functionality of a storage node 150 is shifted to a storage unit 152, transforming the storage unit 152 into a combination of a storage unit 152 and a storage node 150. Placing computation (as opposed to storing data) into the storage unit 152 places that computation closer to the data itself. Various system embodiments have a hierarchy of storage node tiers with varying capabilities. In contrast, in a storage array, the controller owns and knows everything about all the data that the controller manages in a shelf or storage device. In a storage cluster 161, as described herein, multiple non-volatile solid-state storage device 152 units and / or multiple controllers in storage nodes 150 cooperate in various ways (e.g., for erasure coding, data sharding, metadata communication and redundancy, storage capacity expansion or contraction, data recovery, etc.).
[0111] Figure 2D Demonstration use Figure 2A 150 and storage device 152 units. In this version, each non-volatile solid-state storage device 152 unit has a storage server environment in the chassis 138 (see Figure 2A ) on a PCIe (Peripheral Component Interconnect Express) board, for example, a controller 212 (see Figure 2C ) of the processor, FPGA, flash memory 206 and NVRAM 204 (which is a supercapacitor-backed DRAM 216, see Figure 2B and 2C ). The non-volatile solid-state storage device 152 unit can be implemented as a single board containing the storage device and can be the largest tolerable failure domain within the chassis. In some embodiments, up to two non-volatile solid-state storage device 152 units can fail and the device will continue without data loss.
[0112] In some embodiments, the physical storage device is divided into named regions based on application usage. NVRAM 204 is a contiguous block of memory retained in the non-volatile solid-state storage device 152 DRAM 216 and is backed by NAND flash memory. NVRAM 204 is logically divided into multiple memory regions, two of which are designated as spools (e.g., spool_region). The space within an NVRAM 204 spool is independently managed by each authorization 168. Each device provides a certain amount of storage space to each authorization 168. That authorization 168 further manages the lifecycle and allocation within that space. Instances of spools include distributed transactions or concepts. When the main power to the non-volatile solid-state storage device 152 unit fails, onboard supercapacitors provide a short-duration power holdover. During this holdover interval, the contents of NVRAM 204 are flushed to the flash memory 206. Upon the next power-on, the contents of NVRAM 204 are restored from the flash memory 206.
[0113] With respect to the storage unit controller, the responsibility of the logical "controller" is distributed across each of the blades containing the authorization 168. This logical control is distributed across Figure 2D Shown are host controller 242, mid-tier controller 244, and storage unit controller 246. The control plane and storage plane management are treated independently, but the components can be physically co-located on the same blade. Each authorization 168 effectively acts as an independent controller. Each authorization 168 provides its own data and metadata structures, its own background workers, and maintains its own life cycle.
[0114] Figure 2E It is displayed in Figure 2D Used in a storage server environment Figure 2A FIG2 is a hardware block diagram of a blade 252, showing an embodiment of a storage node 150 and storage unit 152 interacting with underlying physical resources, a control plane 254, compute and storage planes 256, 258, and authorizations 168. The control plane 254 is partitioned into a number of authorizations 168 that can run on any of the blades 252 using the compute resources in the compute plane 256. The storage plane 258 is partitioned into a set of devices, each of which provides access to flash memory 206 and NVRAM 204 resources. In one embodiment, the compute plane 256 can perform the operations of a storage array controller for one or more devices of the storage plane 258 (e.g., a storage array), as described herein.
[0115] exist Figure 2EIn the compute and storage planes 256 and 258 of the system, authorizations 168 interact with the underlying physical resources (i.e., devices). From the perspective of an authorization 168, its resources are striped across all physical devices. From the perspective of a device, it provides resources to all authorizations 168, regardless of where the authorization happens to be running. Each authorization 168 is allocated or has been allocated one or more partitions 260 of storage memory in the storage unit 152, such as partitions 260 in the flash memory 206 and NVRAM 204. Each authorization 168 uses those allocated partitions 260 to write or read user data. Authorizations can be associated with different amounts of physical storage in the system. For example, one authorization 168 may have a greater number of partitions 260 or larger-sized partitions 260 in one or more storage units 152 than one or more other authorizations 168.
[0116] Figure 2F Depicts the elasticity software layer in the blade 252 of a storage cluster according to some embodiments. In the elastic architecture, the elasticity software is symmetrical, i.e., the compute module 270 of each blade runs Figure 2F The three identical layers of the process depicted in FIG. Storage manager 274 executes read and write requests from other blades 252 for data and metadata stored in local storage unit 152 NVRAM 204 and flash memory 206. Authorization 168 fulfills client requests by issuing the necessary reads and writes to the blade 252 whose storage unit 152 the corresponding data or metadata resides. Endpoint 272 parses client connection requests received from the switch fabric 146 monitoring software, relays the client connection requests to authorization 168 for implementation, and relays the authorization 168 response to the client. The symmetrical three-tiered structure enables a high degree of concurrency in the storage system. In these embodiments, elasticity scales out efficiently and reliably. In addition, elasticity implements a unique scale-out technology that evenly balances work across all resources, regardless of client access type, and maximizes concurrency by eliminating much of the need for inter-blade coordination that typically occurs with conventional distributed locking.
[0117] Still refer to Figure 2F, the authority 168 running in the compute module 270 of blade 252 performs the internal operations required to complete client requests. One feature of resiliency is that the authority 168 is stateless, that is, it caches valid data and metadata in the DRAM of its own blade 252 for fast access, but the authority stores each update in its partition of NVRAM 204 on three separate blades 252 until the update has been written to flash memory 206. In some embodiments, all storage system writes to NVRAM 204 are triplicated to the partitions on the three separate blades 252. With three mirrored NVRAM 204 and persistent storage protected by parity and Reed-Solomon RAID checksums, the storage system can withstand the concurrent failure of two blades 252 without losing data, metadata, or access to either.
[0118] Because authorizations 168 are stateless, they can be migrated between blades 252. Each authorization 168 has a unique identifier. In some embodiments, the partitions of NVRAM 204 and flash memory 206 are associated with the identifier of the authorization 168, rather than the blade 252 on which it is running. Therefore, when an authorization 168 migrates, the authorization 168 continues to manage the same storage partitions from its new location. When a new blade 252 is installed in an embodiment of a storage cluster, the system automatically rebalances the load by partitioning the storage of the new blade 252 for use by the system's authorizations 168, migrating selected authorizations 168 to the new blade 252, enabling endpoints 272 on the new blade 252, and including it in the client connection distribution algorithm of the switch fabric 146.
[0119] The migrated authority 168 maintains the contents of its NVRAM 204 partition from its new location on flash memory 206, processes read and write requests from other authorities 168, and completes client requests directed to it by endpoint 272. Similarly, if a blade 252 fails or is removed, the system redistributes its authorities 168 among the remaining blades 252 in the system. The redistributed authorities 168 continue to perform their original functions from their new location.
[0120] Figure 2GDepicts the authorities 168 and storage resources in blades 252 of a storage cluster, according to some embodiments. Each authority 168 is specifically responsible for a partition of flash memory 206 and NVRAM 204 on each blade 252. An authority 168 manages the content and integrity of its partition independently of other authorities 168. An authority 168 compresses incoming data and temporarily stores it in its partition of NVRAM 204. It then consolidates, RAID-protects, and maintains the data in a segment of storage within its partition of flash memory 206. As authorities 168 write data to flash memory 206, a storage manager 274 performs the necessary flash translation to optimize write performance and maximize media life. In the background, authorities 168 "ghost collection," or reclaiming space occupied by data obsoleted by clients by overwriting it. It should be appreciated that because the partitions of authorities 168 are disjoint, no distributed locking is required to execute clients and writes or perform background functions.
[0121] The embodiments described herein may utilize various software, communication and / or networking protocols. In addition, the configuration of the hardware and / or software may be adjusted to accommodate various protocols. For example, embodiments may utilize Active Directory, which is a Windows TM A database-based system that provides authentication, directory, policy and other services in an environment. In these embodiments, LDAP (Lightweight Directory Access Protocol) is an example application protocol for querying and modifying items in a directory service provider (such as Active Directory). In some embodiments, a Network Lock Manager ('NLM') is used as a facility that works in conjunction with the Network File System ('NFS') to provide System V style advisory file and record locking over a network. The Server Message Block ('SMB') protocol (a version of which is also called the Common Internet File System ('CIFS')) can be integrated with the storage systems discussed herein. SMB operations are application layer network protocols that are commonly used to provide shared access to files, printers, and serial ports, as well as miscellaneous communications between nodes on a network. SMB also provides an authenticated inter-process communication mechanism. AMAZON TMS3 (Simple Storage Service) is a web service provided by Amazon Web Services, and the systems described herein can interface with Amazon S3 via web service interfaces (REST (Representational State Transfer), SOAP (Simple Object Access Protocol), and BitTorrent). A RESTful API (Application Programming Interface) breaks down transactions into a series of small modules. Each module addresses a specific underlying portion of a transaction. The controls or permissions provided by these embodiments, particularly for object data, can include the use of access control lists ('ACLs'). An ACL is a permission list attached to an object that specifies which users or system processes are authorized to access the object and what operations are permitted on a given object. The system can utilize Internet Protocol version 6 ('IPv6') as well as IPv4, a communications protocol for computers on a network to identify and locate systems and route traffic across the Internet. Packet routing between networked systems can include equal-cost multipath routing ('ECMP'), a routing strategy in which next-hop packet forwarding to a single destination occurs via multiple "best paths" that are topped in a routing metric calculation. Multipath routing can be used with most routing protocols because it is limited to per-hop decisions within a single router. The software may support multi-tenancy, which is an architecture in which a single instance of a software application serves multiple customers. Each customer may be referred to as a tenant. In some embodiments, tenants may be given the ability to customize portions of the application, but may not be able to customize the application's code. An embodiment may maintain an audit log. An audit log is a document that records events in a computing system. In addition to recording what resources were accessed, audit log entries typically include destination and source addresses, timestamps, and user login information to comply with various regulations. An embodiment may support various key management policies, such as encryption key rotation. Additionally, the system may support a dynamic root password or some variation of a dynamically changing password.
[0122] Figure 3A A diagram illustrating a storage system 306 coupled for data communication with a cloud service provider 302 according to some embodiments of the present disclosure is set forth. Although depicted in less detail, Figure 3A The storage system 306 depicted in FIG. 3 may be similar to the storage system 306 described above with reference to FIG. Figures 1A to 1D and Figures 2A to 2G In some embodiments, Figure 3AThe storage system 306 depicted in the may be embodied as a storage system including unbalanced active / active controllers, a storage system including balanced active / active controllers, a storage system including active / active controllers (in which less than all resources of each controller are utilized so that each controller has reserved resources available to support failover), a storage system including fully active / active controllers, a storage system including data set isolated controllers, a storage system including a two-tier architecture with a front-end controller and a back-end integrated storage controller, a storage system including a scale-out cluster of dual-controller arrays, and combinations of such embodiments.
[0123] exist Figure 3A In the example depicted in FIG, a storage system 306 is coupled to a cloud service provider 302 via a data communication link 304. This data communication link 304 can be entirely wired, entirely wireless, or some aggregation of wired and wireless data communication paths. In this example, digital information can be exchanged between the storage system 306 and the cloud service provider 302 via the data communication link 304 using one or more data communication protocols. For example, digital information can be exchanged between the storage system 306 and the cloud service provider 302 via the data communication link 304 using Handheld Device Transport Protocol ('HDTP'), Hypertext Transport Protocol ('HTTP'), Internet Protocol ('IP'), Real-time Transport Protocol ('RTP'), Transmission Control Protocol ('TCP'), User Datagram Protocol ('UDP'), Wireless Application Protocol ('WAP'), or other protocols.
[0124] For example, Figure 3A The cloud service provider 302 depicted in FIG may embody a system and computing environment that provides a wide range of services to users of the cloud service provider 302 by sharing computing resources via a data communication link 304. The cloud service provider 302 may provide on-demand access to a pool of shared configurable computing resources, such as computer networks, servers, storage devices, applications and services, and the like.
[0125] exist Figure 3A , cloud service provider 302 can be configured to provide various services to storage system 306 and users of storage system 306 by implementing various service models. For example, cloud service provider 302 can be configured to provide services by implementing an Infrastructure as a Service ('IaaS') service model, by implementing a Platform as a Service ('PaaS') service model, by implementing a Software as a Service ('SaaS') service model, by implementing an Authentication as a Service ('AaaS') service model, by implementing a Storage as a Service model (in which cloud service provider 302 provides access to its storage infrastructure for use by storage system 306 and users of storage system 306), and so on.
[0126] exist Figure 3A In the example depicted in FIG, cloud service provider 302 may be embodied as, for example, a private cloud, a public cloud, or a combination of private and public clouds. In embodiments where cloud service provider 302 is embodied as a private cloud, cloud service provider 302 may be dedicated to providing services to a single organization, rather than to multiple organizations. In embodiments where cloud service provider 302 is embodied as a public cloud, cloud service provider 302 may provide services to multiple organizations. In yet another alternative embodiment, cloud service provider 302 may be embodied as a mix of private and public cloud services using a hybrid cloud deployment.
[0127] although Figure 3A Although not explicitly depicted in the present disclosure, the reader will appreciate that a number of additional hardware and software components may be required to facilitate the delivery of cloud services to storage system 306 and its users. For example, storage system 306 may be coupled to (or even include) a cloud storage gateway. For example, this cloud storage gateway may be embodied as a hardware-based or software-based appliance located on-premises with storage system 306. This cloud storage gateway may operate as a bridge between local applications executing on storage system 306 and remote cloud-based storage devices utilized by storage system 306. By using a cloud storage gateway, an organization can move its primary iSCSI or NAS to cloud service provider 302, thereby enabling the organization to save space on its on-premises storage systems. This cloud storage gateway may be configured to emulate a disk array, block-based device, file server, or other storage system that can translate SCSI commands, file server commands, or other appropriate commands into a REST space protocol that facilitates communication with cloud service provider 302.
[0128] To enable storage system 306 and its users to utilize the services provided by cloud service provider 302, a cloud migration process may be performed, during which data, applications, or other elements from an organization's on-premises systems (or even from another cloud environment) are moved to cloud service provider 302. To successfully migrate the data, applications, or other elements to the cloud service provider's 302 environment, middleware, such as a cloud migration tool, may be used to bridge the gap between the cloud service provider's 302 environment and the organization's environment. To further enable storage system 306 and its users to utilize the services provided by cloud service provider 302, a cloud orchestrator may also be used to arrange and coordinate automated tasks to create a merged process or workflow. This cloud orchestrator may perform tasks such as configuring various components, whether those components are cloud components or locally deployed components, and managing the interconnections between such components.
[0129] exist Figure 3AIn the example depicted in FIG, and as briefly described above, cloud service provider 302 can be configured to provide services to storage system 306 and users of storage system 306 using a SaaS service model. For example, cloud service provider 302 can be configured to provide storage system 306 and users of storage system 306 with access to data analytics applications. Such data analytics applications can be configured, for example, to receive large amounts of telemetry data that is phoned home by storage system 306. Such telemetry data can describe various operational characteristics of storage system 306 and can be analyzed for a number of purposes, including, for example, determining the health of storage system 306, identifying workloads executing on storage system 306, predicting when storage system 306 will run out of various resources, recommending configuration changes, hardware or software upgrades, workflow migrations, or other actions that can improve the operation of storage system 306.
[0130] The cloud service provider 302 may also be configured to provide access to a virtualized computing environment to the storage system 306 and users of the storage system 306. Examples of such virtualized environments may include virtual machines created to emulate actual computers, virtualized desktop environments that separate logical desktops from physical machines, virtualized file systems that allow unified access to different types of specific file systems, and many others.
[0131] although Figure 3A The example depicted in FIG3 illustrates storage system 306 coupled for data communication with cloud service provider 302, but in other embodiments, storage system 306 may be part of a hybrid cloud deployment in which private cloud elements (e.g., private cloud services, on-premises infrastructure, etc.) and public cloud elements (e.g., public cloud services, infrastructure, etc. that may be provided by one or more cloud service providers) are combined to form a single solution, orchestrated among the various platforms. This hybrid cloud deployment may utilize hybrid cloud management software such as, for example, Microsoft Azure. TM Azure TM Arc centralizes the management of hybrid cloud deployments to any infrastructure and enables the deployment of services anywhere. In this instance, hybrid cloud management software can be configured to create, update, and delete resources (both physical and virtual) that form a hybrid cloud deployment, allocate compute and storage to specific workloads, monitor workloads and resources for performance, policy compliance, updates and patches, security status, or perform a variety of other tasks.
[0132] The reader will appreciate that a variety of offerings can be implemented by pairing the storage systems described herein with one or more cloud service providers. For example, where cloud resources are used to protect applications and data from damage caused by disasters, Disaster Recovery as a Service (DRaaS) can be provided, including in embodiments where the storage system can be used as primary data storage. In such embodiments, full system backups can be performed, which allows for business continuity in the event of a system failure. In such embodiments, cloud data backup technology (either independently or as part of a larger DRaaS solution) can also be integrated into an overall solution including the storage system described herein and a cloud service provider.
[0133] The storage systems and cloud service providers described herein can be used to provide a number of security features. For example, the storage system can encrypt data at rest (and data can be sent to and from the encrypted storage system) and can use Key Management as a Service ('KMaaS') to manage encryption keys, keys for locking and unlocking storage devices, and so on. Similarly, a cloud data security gateway or similar mechanism can be used to ensure that data stored within the storage system does not end up being improperly stored in the cloud as part of cloud data backup operations. In addition, micro-segmentation or identity-based segmentation can be used in the data center containing the storage system or within the cloud service provider to create security zones in data center and cloud deployments, which enables workloads to be isolated from each other.
[0134] To explain further, Figure 3B A diagram illustrating a storage system 306 according to some embodiments of the present disclosure is set forth. Although depicted in less detail, Figure 3B The storage system 306 depicted in FIG. 3 may be similar to the storage system 306 described above with reference to FIG. Figures 1A to 1D and Figures 2A to 2G The storage system described above is because the storage system may include many of the components described above.
[0135] Figure 3BThe storage system 306 depicted in FIG. 3 may include a number of storage resources 308, which may be embodied in many forms. For example, the storage resources 308 may include nanoRAM or another form of nonvolatile random access memory utilizing carbon nanotubes deposited on a substrate, 3D crosspoint nonvolatile memory, flash memory (including single-level cell ('SLC') NAND flash, multi-level cell ('MLC') NAND flash, triple-level cell ('TLC') NAND flash, quad-level cell ('QLC') NAND flash), or other. Similarly, the storage resources 308 may include nonvolatile magnetoresistive random access memory ('MRAM'), including spin transfer torque ('STT') MRAM. The example storage resources 308 may alternatively include nonvolatile phase change memory ('PCM'), quantum memory that allows for the storage and retrieval of photon quantum information, resistive random access memory ('ReRAM'), storage class memory ('SCM'), or other forms of storage resources, including any combination of the resources described herein. The reader will appreciate that other forms of computer memory and storage devices may be utilized by the memory system described above, including DRAM, SRAM, EEPROM, general purpose memory, and many others. Figure 3A The storage resources 308 depicted in may be embodied in various form factors, including, but not limited to, dual in-line memory modules ('DIMMs'), non-volatile dual in-line memory modules ('NVDIMMs'), M.2, U.2, and others.
[0136] Figure 3B The storage resources 308 depicted in FIG may include various forms of SCM. The SCM can effectively treat fast non-volatile memory (e.g., NAND flash) as an extension of DRAM, allowing an entire data set to be treated as an in-memory data set residing entirely in DRAM. The SCM may include non-volatile media such as, for example, NAND flash. This NAND flash can be accessed using NVMe, which can use the PCIe bus as its transport, providing relatively low access latency compared to older protocols. In practice, network protocols used for SSDs in all-flash arrays may include NVMe over Ethernet (ROCE, NVMe TCP), Fibre Channel (NVMe FC), InfiniBand (iWARP), and others that make it possible to treat fast non-volatile memory as an extension of DRAM. Given that DRAM is typically byte-addressable and fast non-volatile memory (e.g., NAND flash) is block-addressable, a controller software / hardware stack may be required to convert block data into bytes stored in the media. Examples of media and software that may be used as SCM may include, for example, 3D XPoint, Intel memory drive technology, Samsung's Z-SSD, and others.
[0137] Figure 3B The storage resources 308 depicted in FIG may also include racetrack memory (also known as domain-wall memory). This racetrack memory can be embodied as a form of non-volatile solid-state memory that relies, in addition to the electron's charge, on the intrinsic strength and orientation of the magnetic field created by the electrons in the solid-state device as they spin. By using a spin-coherent current to move magnetic domains along nanopermalloy wires, the domains can be transferred by a magnetic read / write head positioned near the wire as the current passes through the wire, altering the magnetic domains to record bit patterns. To create a racetrack memory device, many such wires and read / write elements can be packaged together.
[0138] Figure 3B The example storage system 306 depicted in FIG. 306 can implement various storage architectures. For example, a storage system according to some embodiments of the present disclosure may utilize block storage, where data is stored in blocks, and each block essentially acts as an individual hard drive. A storage system according to some embodiments of the present disclosure may utilize object storage, where data is managed as objects. Each object may include the data itself, a variable amount of metadata, and a globally unique identifier, where object storage can be implemented at multiple levels (e.g., device level, system level, interface level). A storage system according to some embodiments of the present disclosure utilizes file storage, where data is stored in a hierarchical structure. This data can be stored in files and folders and presented to both the system that stored it and the system that retrieved it in the same format.
[0139] Figure 3B The example storage system 306 depicted in FIG3 may embody a storage system in which additional storage resources may be added using a scale-up model, additional storage resources may be added using a scale-out model, or some combination thereof. In a scale-up model, additional storage may be added by adding additional storage devices. However, in a scale-out model, additional storage nodes may be added to the storage node cluster, where such storage nodes may include additional processing resources, additional networking resources, and the like.
[0140] Figure 3B The example storage system 306 depicted in FIG can utilize the storage resources described above in a variety of different ways. For example, a portion of the storage resources can be used as a write cache, storage resources within the storage system can be used as a read cache, or tiering can be implemented within the storage system by placing data within the storage system according to one or more tiering policies.
[0141] Figure 3BThe storage system 306 depicted in FIG also includes communication resources 310 that can be used to facilitate data communication between components within the storage system 306 and between the storage system 306 and computing devices external to the storage system 306, including embodiments in which those resources are separated by relatively wide areas. The communication resources 310 can be configured to utilize a variety of different protocols and data communication architectures to facilitate data communication between components within the storage system and computing devices external to the storage system. For example, the communication resources 310 can include: Fibre Channel ('FC') technology, such as the FC architecture and FC protocol that can transmit SCSI commands over an FC network; FC over Ethernet ('FCoE') technology, by which FC frames are encapsulated and transmitted over an Ethernet network; InfiniBand ('IB') technology, in which a switched fabric topology is used to facilitate transmissions between channel adapters; NVM Express ('NVMe') technology and NVMe over Fabric ('NVMeoF') technology, by which non-volatile storage media attached via a PCI Express ('PCIe') bus can be accessed; and others. In practice, the storage systems described above may directly or indirectly employ neutrino communication techniques and devices by which information (including binary information) is transmitted using neutrino beams.
[0142] The communication resources 310 may also include mechanisms for accessing storage resources 308 within the storage system 306 utilizing Serial Attached SCSI ('SAS'), a Serial ATA ('SATA') bus interface for connecting storage resources 308 within the storage system 306 to a host bus adapter within the storage system 306, Internet Small Computer System Interface ('iSCSI') technology for providing block-level access to storage resources 308 within the storage system 306, and other communication resources that may be used to facilitate data communication between components within the storage system 306 and data communication between the storage system 306 and computing devices external to the storage system 306.
[0143] Figure 3B The storage system 306 depicted in FIG3 also includes processing resources 312 that can be used to execute computer program instructions and perform other computing tasks within the storage system 306. The processing resources 312 can include one or more ASICs and one or more CPUs customized for a particular purpose. The processing resources 312 can also include one or more DSPs, one or more FPGAs, one or more systems on a chip ('SoCs'), or other forms of processing resources 312. The storage system 306 can utilize the processing resources 312 to perform various tasks, including, but not limited to, supporting the execution of software resources 314, which will be described in more detail below.
[0144] Figure 3BThe storage system 306 depicted in FIG3 also includes software resources 314 that, when executed by processing resources 312 within the storage system 306, can perform a number of tasks. Software resources 314 may include, for example, one or more computer program instruction modules that, when executed by processing resources 312 within the storage system 306, are used to implement various data protection techniques. Such data protection techniques may be implemented, for example, by system software executing on computer hardware within the storage system, by a cloud service provider, or otherwise. Such data protection techniques may include data archiving, data backup, data replication, data snapshots, data and database cloning, and other data protection techniques.
[0145] Software resources 314 may also include software for implementing software-defined storage ('SDS'). In this example, software resources 314 may include one or more computer program instruction modules that, when executed, provide for policy-based provisioning and management of data storage, independent of the underlying hardware. Such software resources 314 may be used to implement storage virtualization to separate storage hardware from the software that manages the storage hardware.
[0146] Software resources 314 may also include software for facilitating and optimizing I / O operations directed to storage system 306. For example, software resources 314 may include software modules that perform various data reduction techniques, such as, for example, data compression, data deduplication, and others. Software resources 314 may include software modules that intelligently group I / O operations together to facilitate better use of underlying storage resources 308, software modules that perform data migration operations to migrate data from within the storage system, and software modules that perform other functions. Such software resources 314 may be embodied as one or more software containers or in many other ways.
[0147] To explain further, Figure 3C An example of a cloud-based storage system 318 according to some embodiments of the present disclosure is presented. Figure 3C In the example depicted in FIG, cloud-based storage system 318 is created entirely within a cloud computing environment 316, such as, for example, Amazon Web Services ('AWS'). TM 、Microsoft Azure TM 、Google Cloud Platform TM 、IBM Cloud TM , Oracle Cloud TM The cloud-based storage system 318 may be used to provide services similar to those that may be provided by the storage systems described above.
[0148] Figure 3CThe cloud-based storage system 318 depicted in FIG includes two cloud computing instances 320 and 322, each for supporting the execution of a storage controller application 324 and 326. For example, the cloud computing instances 320 and 322 may be embodied as instances of cloud computing resources (e.g., virtual machines) that may be provided by the cloud computing environment 316 to support the execution of software applications (e.g., storage controller applications 324 and 326). For example, each of the cloud computing instances 320 and 322 may be executed on an Azure VM, where each Azure VM may include high-speed temporary storage that may be used as a cache (e.g., as a read cache). In one embodiment, the cloud computing instances 320 and 322 may be embodied as Amazon Elastic Compute Cloud ('EC2') instances. In this example, an Amazon Machine Image ('AMI') including the storage controller applications 324 and 326 may be launched to create and configure a virtual machine that can execute the storage controller applications 324 and 326.
[0149] exist Figure 3C In the example method described in the embodiment, the storage controller applications 324, 326 may be embodied as computer program instruction modules that, when executed, perform various storage tasks. For example, the storage controller applications 324, 326 may be embodied as computer program instruction modules that, when executed, perform the same tasks as described above. Figure 1A The controllers 110A, 110B in the cloud-based storage system 318 perform the same tasks, such as writing data to the cloud-based storage system 318, erasing data from the cloud-based storage system 318, retrieving data from the cloud-based storage system 318, monitoring and reporting storage device utilization and performance, performing redundancy operations (such as RAID or RAID-like data redundancy operations), compressing data, encrypting data, deduplicating data, etc. The reader will understand that because there are two cloud computing instances 320, 322 each including a storage controller application 324, 326, in some embodiments, one cloud computing instance 320 can operate as a primary controller as described above, while the other cloud computing instance 322 can operate as a secondary controller as described above. The reader will understand that Figure 3C The storage controller applications 324, 326 depicted in FIG may include the same source code that executes within different cloud computing instances 320, 322 (eg, distinct EC2 instances).
[0150] The reader will appreciate that other embodiments that do not include primary and secondary controllers are within the scope of the present disclosure. For example, each cloud computing instance 320, 322 may operate as a primary controller for a certain portion of the address space supported by the cloud-based storage system 318, each cloud computing instance 320, 322 may operate as a primary controller in which the servicing of I / O operations directed to the cloud-based storage system 318 is divided in some other manner, etc. Indeed, in other embodiments in which cost savings may take precedence over performance requirements, there may be only a single cloud computing instance containing the storage controller application.
[0151] Figure 3C The cloud-based storage system 318 depicted in FIG includes cloud computing instances 340a, 340b, 340n having local storage devices 330, 334, 338. For example, cloud computing instances 340a, 340b, 340n may embody examples of cloud computing resources that may be provided by the cloud computing environment 316 to support the execution of software applications. Figure 3C The cloud computing examples 340a, 340b, 340n may differ from the cloud computing examples 320, 322 described above because Figure 3C The cloud computing instances 340a, 340b, 340n of the storage controller applications 324, 326 have local storage 330, 334, 338 resources, while the cloud computing instances 320, 322 supporting the execution of the storage controller applications 324, 326 do not need to have local storage resources. For example, the cloud computing instances 340a, 340b, 340n with local storage 330, 334, 338 can be embodied as EC2 M5 instances including one or more SSDs, EC2 R5 instances including one or more SSDs, EC2 I3 instances including one or more SSDs, and so on. In some embodiments, the local storage 330, 334, 338 must be embodied as solid-state storage devices (e.g., SSDs) rather than storage devices using hard disk drives.
[0152] exist Figure 3CIn the example depicted in FIG, each of the cloud computing instances 340a, 340b, 340n having local storage devices 330, 334, 338 can include a software daemon 328, 332, 336 that, when executed by the cloud computing instances 340a, 340b, 340n, can present itself to the storage controller application 324, 326 as if the cloud computing instances 340a, 340b, 340n were physical storage devices (e.g., one or more SSDs). In this example, the software daemons 328, 332, 336 can include computer program instructions similar to those typically included on storage devices, such that the storage controller applications 324, 326 can send and receive the same commands that a storage controller would send to the storage devices. In this way, the storage controller applications 324, 326 can include code that is the same (or substantially the same) as that to be executed by the controller in the storage system described above. In these and similar embodiments, communication between the storage controller applications 324, 326 and the cloud computing instances 340a, 340b, 340n having local storage devices 330, 334, 338 may utilize iSCSI, NVMe over TCP, messaging, a custom protocol, or some other mechanism.
[0153] exist Figure 3CIn the example depicted in FIG, each of the cloud computing instances 340a, 340b, 340n having local storage 330, 334, 338 may also be coupled to block storage 342, 344, 346 provided by the cloud computing environment 316, such as, for example, an Amazon Elastic Block Store ('EBS') volume. In this example, the block storage 342, 344, 346 provided by the cloud computing environment 316 may be utilized in a manner similar to how the NVRAM devices described above are utilized, in that a software daemon 328, 332, 336 (or some other module) executing within a particular cloud computing instance 340a, 340b, 340n may initiate writing data to its attached EBS volume as well as writing data to its local storage 330, 334, 338 resource upon receiving a request to write data. In some alternative embodiments, data can only be written to local storage 330, 334, 338 resources within a particular cloud computing instance 340a, 340b, 340n. In an alternative embodiment, rather than using block storage 342, 344, 346 provided by the cloud computing environment 316 as NVRAM, actual RAM on each of the cloud computing instances 340a, 340b, 340n having local storage 330, 334, 338 can be used as NVRAM, thereby reducing the network utilization costs associated with using EBS volumes as NVRAM. In yet another embodiment, high-performance block storage resources such as one or more Azure Ultra Disks can be used as NVRAM.
[0154] When a request to write data is received by a particular cloud computing instance 340a, 340b, 340n having local storage 330, 334, 338, the software daemon 328, 332, 336 can be configured to write the data not only to its own local storage 330, 334, 338 resources and any appropriate block storage 342, 344, 346 resources, but the software daemon 328, 332, 336 can also be configured to write the data to a cloud-based object storage 348 attached to the particular cloud computing instance 340a, 340b, 340n. For example, the cloud-based object storage 348 attached to the particular cloud computing instance 340a, 340b, 340n can be embodied as Amazon Simple Storage Service ('S3'). In other embodiments, each cloud computing instance 320, 322 including a storage controller application 324, 326 can initiate the storage of data in the local storage 330, 334, 338 and cloud-based object storage 348 of the cloud computing instances 340a, 340b, 340n. In other embodiments, rather than using cloud computing instances 340a, 340b, 340n with local storage 330, 334, 338 (also referred to herein as 'virtual drives') and cloud-based object storage 348 to store data, a persistent storage layer can be implemented in other ways. For example, one or more Azure Ultra Disks can be used to persistently store data (e.g., after the data has been written to the NVRAM layer). In embodiments where one or more Azure Ultra Disks can be used to persistently store data, the use of cloud-based object storage 348 can be eliminated so that data is only persistently stored in the Azure Ultra Disks, without having to write the data to the object storage layer.
[0155] While the local storage 330, 334, 338 resources and block storage 342, 344, 346 resources utilized by the cloud computing instances 340a, 340b, 340n can support block-level access, the cloud-based object storage 348 attached to a particular cloud computing instance 340a, 340b, 340n only supports object-based access. The software daemons 328, 332, 336 can therefore be configured to take blocks of data, package those blocks into objects, and write the objects to the cloud-based object storage 348 attached to the particular cloud computing instance 340a, 340b, 340n.
[0156] In some embodiments, all data stored by the cloud-based storage system 318 can be stored in both: 1) a cloud-based object store 348 and 2) at least one of the local storage 330, 334, 338 resources or the block storage 342, 344, 346 resources utilized by the cloud computing instances 340a, 340b, 340n. In such embodiments, the local storage 330, 334, 338 resources and the block storage 342, 344, 346 resources utilized by the cloud computing instances 340a, 340b, 340n can effectively operate as a cache that generally includes all data also stored in S3, such that all reads of data can be serviced by the cloud computing instances 340a, 340b, 340n without requiring the cloud computing instances 340a, 340b, 340n to access the cloud-based object store 348. However, the reader will appreciate that in other embodiments, all data stored by the cloud-based storage system 318 may be stored in the cloud-based object store 348, but not all data stored by the cloud-based storage system 318 may be stored in at least one of the local storage 330, 334, 338 resources or the block storage 342, 344, 346 resources utilized by the cloud computing instances 340a, 340b, 340n. In this example, various policies may be utilized to determine which subset of the data stored by the cloud-based storage system 318 should reside in: 1) the cloud-based object store 348; and 2) at least one of the local storage 330, 334, 338 resources or the block storage 342, 344, 346 resources utilized by the cloud computing instances 340a, 340b, 340n.
[0157] One or more computer program instruction modules executing within the cloud-based storage system 318 (e.g., a monitoring module executing on its own EC2 instance) can be designed to handle the failure of one or more of the cloud computing instances 340a, 340b, 340n having local storage 330, 334, 338. In this example, the monitoring module can handle the failure of one or more of the cloud computing instances 340a, 340b, 340n having local storage 330, 334, 338 by creating one or more new cloud computing instances with local storage, retrieving data stored on the failed cloud computing instances 340a, 340b, 340n from the cloud-based object storage 348, and storing the data retrieved from the cloud-based object storage 348 in the local storage on the newly created cloud computing instances. The reader will appreciate that many variations of this process can be implemented.
[0158] The reader will appreciate that various performance aspects of the cloud-based storage system 318 can be monitored (e.g., by a monitoring module executing in an EC2 instance) so that the cloud-based storage system 318 can be scaled up or scaled out as needed. For example, if the cloud computing instances 320, 322 used to support the execution of the storage controller applications 324, 326 are undersized and insufficient to service I / O requests issued by users of the cloud-based storage system 318, the monitoring module can create a new, more powerful cloud computing instance (e.g., a type of cloud computing instance that includes more processing power, more memory, etc.) that includes the storage controller application so that the new, more powerful cloud computing instance can begin operating as the primary controller. Similarly, if the monitoring module determines that the cloud computing instances 320, 322 used to support the execution of the storage controller applications 324, 326 are oversized and that cost savings can be achieved by switching to a smaller, less powerful cloud computing instance, the monitoring module can create a new, less powerful (and less expensive) cloud computing instance that includes the storage controller application so that the new, less powerful cloud computing instance can begin operating as the primary controller.
[0159] The storage system described above can implement intelligent data backup techniques, which replicate data stored in the storage system and store it in a different location to prevent data loss in the event of an equipment failure or some other form of disaster. For example, the storage system described above can be configured to check each backup to avoid restoring the storage system to an unintended state. Consider an example in which malware infects the storage system. In this example, the storage system can include software resources 314 that can scan each backup to identify backups captured before the malware infected the storage system and those captured after the malware infected the storage system. In this example, the storage system can restore itself from a backup that does not contain the malware—or at least not restore the portion of the backup that does contain the malware. In this example, the storage system may include software resources 314 that can scan each backup to identify the presence of malware (or a virus or something else unexpected), for example, by identifying write operations served by the storage system and originating from a network subnet suspected of having delivered malware, by identifying write operations served by the storage system and originating from users suspected of having delivered malware, by identifying write operations served by the storage system and checking the content of the write operations against a fingerprint of malware, and in many other ways.
[0160] The reader will further appreciate that backups (typically in the form of one or more snapshots) can also be used to perform rapid recovery of the storage system. Consider an example in which a storage system is infected with ransomware that locks users out of the storage system. In this example, software resources 314 within the storage system can be configured to detect the presence of ransomware and can further be configured to use the retained backup to recover the storage system to a point in time before the point in time when the ransomware infected the storage system. In this example, the presence of ransomware can be explicitly detected using a software tool utilized by the system, by using a key inserted into the storage system (e.g., a USB drive), or in a similar manner. Similarly, the presence of ransomware can be inferred in response to system activity meeting a predetermined fingerprint (e.g., no reads or writes to the system within a predetermined time period, for example).
[0161] The reader will appreciate that the various components described above can be grouped into one or more optimized computing packages as a converged infrastructure. This converged infrastructure can include a pool of computer, storage, and networking resources that can be shared by multiple applications and collectively managed using policy-driven processes. Such a converged infrastructure can be implemented using a converged infrastructure reference architecture, as a standalone appliance, using a software-driven hyperconvergence approach (e.g., hyperconverged infrastructure), or in other ways.
[0162] The reader will appreciate that the storage systems described in this disclosure can be used to support various types of software applications. In fact, the storage system can be 'application-aware' in the sense that the storage system can obtain, maintain, or otherwise access information describing connected applications (e.g., applications utilizing the storage system) to optimize the operation of the storage system based on intelligence about the applications and their utilization patterns. For example, the storage system can optimize data layout, optimize cache behavior, optimize 'QoS' tiers, or perform some other optimization designed to improve the storage performance experienced by the application.
[0163] As an example of a type of application that may be supported by the storage system described herein, storage system 306 may be used to support such applications by providing storage resources to: artificial intelligence ('AI') applications, database applications, XOps projects (e.g., DevOps projects, DataOps projects, MLOps projects, ModelOps projects, PlatformOps projects), electronic design automation tools, event-driven software applications, high-performance computing applications, simulation applications, high-speed data capture and analysis applications, machine learning applications, media production applications, media serving applications, picture archiving and communication system ('PACS') applications, software development applications, virtual reality applications, augmented reality applications, and many other types of applications.
[0164] Given that storage systems contain compute resources, storage resources, and a wide variety of other resources, they can be well-suited to supporting resource-intensive applications, such as AI applications. AI applications can be deployed in a variety of areas, including predictive maintenance in manufacturing and related fields; healthcare applications, such as patient data and risk analysis; retail and marketing deployments (e.g., search notifications, social media notifications); supply chain solutions; fintech solutions, such as business analytics and reporting tools; operational deployments, such as real-time analytics tools, application performance management tools, IT infrastructure management tools; and many others.
[0165] Such AI applications may enable a device to perceive its environment and take actions that maximize its chances of success in achieving a goal. Examples of such AI applications may include IBM Watson TM Microsoft Oxford TM , Google DeepMind TM , Baidu Minwa TM and others.
[0166] The storage system described above may also be well-suited to supporting other types of resource-intensive applications, such as, for example, machine learning applications. Machine learning applications can perform various types of data analysis to automate analytical model building. Using algorithms that iteratively learn from data, machine learning applications can enable computers to learn without being explicitly programmed. A specific area of machine learning is called reinforcement learning, which involves taking appropriate actions in a specific situation to maximize rewards.
[0167] In addition to the resources already described, the storage system described above may also include a graphics processing unit ('GPU'), sometimes referred to as a visual processing unit ('VPU'). Such a GPU may be embodied as a specialized electronic circuit that rapidly manipulates and modifies memory to accelerate the creation of images in a frame buffer intended for output to a display device. Such a GPU may be included within any computing device that is part of the storage system described above, including as one of many individually scalable components of the storage system, where other examples of individually scalable components of such a storage system may include storage components, memory components, computing components (e.g., CPUs, FPGAs, ASICs), networking components, software components, and others. In addition to a GPU, the storage system described above may also include a neural network processor ('NNP') for use in various aspects of neural network processing. Such an NNP may be used in place of (or in addition to) a GPU and may also be independently scalable.
[0168] As described above, the storage system described herein can be configured to support artificial intelligence applications, machine learning applications, big data analytics applications, and many other types of applications. The rapid growth of these applications is driven by three technologies: deep learning (DL), GPU processors, and big data. Deep learning is a computational model that uses massively parallel neural networks inspired by the human brain. Instead of experts hand-crafting software, deep learning models write their own software by learning from a large number of examples. Such GPUs can contain thousands of cores, making them well-suited to running algorithms that loosely represent the parallel nature of the human brain.
[0169] Advances in deep neural networks, including the development of multi-layer neural networks, have spurred a wave of new algorithms and tools for data scientists to mine their data with artificial intelligence (AI). Using improved algorithms, larger datasets, and a variety of frameworks (including open source software libraries for machine learning across a range of tasks), data scientists are addressing new use cases such as autonomous vehicles, natural language processing and understanding, computer vision, machine reasoning, strong AI, and many others. Applications of AI technology have been implemented in a wide variety of products, including, for example, Amazon Echo's voice recognition technology, which allows users to talk to their machines; Google Translate, which allows users to speak to their machines; and other AI-powered products. TM , which allows machine-based language translation; Spotify’s Discover Weekly, which provides recommendations for new songs and artists that users might like based on their usage and traffic analytics; Quill’s text generation product, which takes structured data and turns it into narrative stories; Chatbot, which provides real-time, context-specific answers to questions in a conversational format; and many others.
[0170] Data is at the heart of modern AI and deep learning algorithms. Before training can begin, one challenge that must be addressed is collecting labeled data, which is crucial for training accurate AI models. Full-scale AI deployments may require the continuous collection, cleaning, transformation, labeling, and storage of large amounts of data. Adding additional high-quality data points directly translates into more accurate models and better insights. Data samples may undergo a series of processing steps, including (but not limited to): 1) ingesting data from external sources into the training system and storing the data in its raw form; 2) cleaning and transforming the data into a format convenient for training, including linking data samples to appropriate labels; 3) exploring parameters and models, quickly testing with smaller datasets, and iterating to converge on the most promising model to advance to the production cluster; 4) performing a training phase to select batches of random input data, including both new and older samples, and feeding those batches into the production GPU servers for computation to update model parameters; and 5) evaluating, including using a holdout portion of the data not used for training to assess model accuracy on the held-out data. This lifecycle is applicable to any type of parallelized machine learning, not just neural networks or deep learning. For example, a standard machine learning framework may rely on CPUs instead of GPUs, but the data ingestion and training workflows can remain the same. You'll appreciate that a single shared storage data hub creates a coordination point throughout the entire lifecycle, eliminating the need for additional data copies during the ingestion, preprocessing, and training phases. Ingested data is rarely used for just one purpose, and shared storage provides the flexibility to train multiple different models or apply traditional analytics to the data.
[0171] The reader will appreciate that each stage in an AI data pipeline may have different requirements from a data center (e.g., a storage system or collection of storage systems). A scale-out storage system must provide uncompromised performance for all types and patterns of access, from small, metadata-heavy to large files, from random to sequential access patterns, and from low to high concurrency. The storage system described above serves as an ideal AI data center because it can serve unstructured workloads. In the first stage, data is ideally ingested and stored on the same data center that will be used by subsequent stages, avoiding additional data copying. The next two steps can be completed on standard compute servers, optionally including GPUs, and then, in the fourth and final stage, the complete training production job is run on powerful GPU-accelerated servers. Typically, there is a production pipeline alongside the experimental pipeline operating on the same dataset. Furthermore, GPU-accelerated servers can be used independently for different models or combined to train on a larger model, or even distributed across multiple systems for training. If the shared storage layer is slow, data must be copied to local storage for each stage, resulting in wasted time staging data on different servers. The ideal data hub for an AI training pipeline provides performance similar to data stored locally on the server nodes, while also providing the simplicity and performance to enable all pipeline stages to operate concurrently.
[0172] In order for the storage system described above to be used as a data center or as part of an AI deployment, in some embodiments, the storage system may be configured to provide DMA between storage devices included in the storage system and one or more GPUs used in an AI or big data analytics pipeline. One or more GPUs may be coupled to the storage system, for example, via NVMe over Fabric ('NVMe-oF'), so that bottlenecks such as the host CPU can be bypassed and the storage system (or one of the components contained therein) can directly access the GPU memory. In this example, the storage system may utilize an API hook to the GPU to transfer data directly to the GPU. For example, the GPU may be embodied as an Nvidia TM The GPU and the storage system may support GPUDirect Storage ('GDS') software or have similar proprietary software that enables the storage system to transfer data to the GPU via RDMA or similar mechanisms.
[0173] While the preceding paragraphs discuss deep learning applications, readers will appreciate that the storage system described herein can also be part of a distributed deep learning ('DDL') platform to support the execution of DDL algorithms. The storage system described above can also be paired with other technologies (e.g., TensorFlow, an open source software library for data flow programming across a range of tasks that can be used in machine learning applications (e.g., neural networks) to facilitate the development of such machine learning models, applications, and the like.
[0174] The storage system described above can also be used in neuromorphic computing environments. Neuromorphic computing is a form of computing that mimics brain cells. To support neuromorphic computing, an architecture of interconnected "neurons" replaces traditional computing models with low-power signals transmitted directly between neurons, enabling more efficient computing. Neuromorphic computing can use very large-scale integration (VLSI) systems containing electronic analog circuits that simulate the neurobiological architecture found in the nervous system, as well as analog, digital, and mixed-mode analog / digital VLSI, and software systems that implement models of the nervous system for perception, motor control, or multi-sensory integration.
[0175] The reader will appreciate that the storage system described above may be configured to support the storage or use of (and other types of data) blockchains and derivatives, such as, for example, IBM TM The open source blockchain and related tools that are part of the Hyperledger project, permissioned blockchains where a limited number of trusted parties are allowed to access the blockchain, blockchain products that enable developers to build their own distributed ledger projects, and others. The blockchain and storage system described in this article can be used to support both on-chain and off-chain storage of data.
[0176] Off-chain storage of data can be implemented in various ways and can occur when the data itself is not stored within the blockchain. For example, in one embodiment, a hash function can be utilized, and the data itself can be fed into the hash function to produce a hash value. In this example, the hash of a large number of pieces of data can be embedded within the transaction, rather than the data itself. The reader will appreciate that in other embodiments, alternatives to blockchain can be used to facilitate decentralized storage of information. For example, one alternative to blockchain that can be used is a blockweave. While conventional blockchains store each transaction for confirmation, blockweaves allow for secure decentralization without using the entire chain, thereby achieving low-cost on-chain storage of data. Such blockweaves can utilize a consensus mechanism based on Proof of Access (PoA) and Proof of Work (PoW).
[0177] The storage systems described above may be used alone or in combination with other computing devices to support in-memory computing applications. In-memory computing involves storing information in RAM that is distributed across a cluster of computers. The reader will appreciate that the storage systems described above, particularly those that are configurable with customizable amounts of processing resources, storage resources, and memory resources (e.g., those systems where blades contain configurable amounts of each type of resource), may be configured in a manner that can provide an infrastructure capable of supporting in-memory computing. Likewise, the storage systems described above may include component parts (e.g., NVDIMMs, 3D crosspoint storage devices that provide persistent fast random access memory) that may, in effect, provide an improved in-memory computing environment compared to in-memory computing environments that rely on RAM distributed across dedicated servers.
[0178] In some embodiments, the storage system described above can be configured to operate as a hybrid in-memory computing environment that includes a universal interface to all storage media (e.g., RAM, flash storage, 3D cross-point storage). In such embodiments, users may not know the details about where their data is stored, but they can still use the same complete, unified API to address the data. In such embodiments, the storage system can (in the background) move the data to the fastest tier available - including intelligently placing the data based on various characteristics of the data or relying on some other heuristics. In this example, the storage system can even use existing products (such as Apache Ignite and GridGain) to move data between the various storage tiers, or the storage system can use custom software to move data between the various storage tiers. The storage system described herein can implement various optimizations to improve the performance of in-memory computing, such as (for example) performing computing as close to the data as possible.
[0179] The reader will further appreciate that, in some embodiments, the storage systems described above can be paired with other resources to support the applications described above. For example, an infrastructure may include primary compute in the form of servers and workstations dedicated to using general-purpose compute on graphics processing units ('GPGPUs') to accelerate deep learning applications interconnected to a compute engine for training parameters of deep neural networks. Each system may have Ethernet external connectivity, InfiniBand external connectivity, some other form of external connectivity, or some combination thereof. In this example, GPUs can be grouped for a single large training run or used independently to train multiple models. The infrastructure may also include storage systems (such as those described above) to provide, for example, scale-out all-flash file or object storage, through which data can be accessed via high-performance protocols (such as NFS, S3, etc.). The infrastructure may also include redundant top-of-rack Ethernet switches connected to the storage and compute, for example, via ports in an MLAG port channel, for redundancy. The infrastructure may also include additional compute in the form of white-box servers, optionally with GPUs, for data ingestion, preprocessing, and model debugging. The reader will appreciate that additional infrastructure is also possible.
[0180] The reader will appreciate that the storage system described above, alone or in coordination with other computing machinery, can be configured to support other AI-related tools. For example, the storage system can use tools such as ONXX or other open neural network exchange formats that make it easier to transfer models written in different AI frameworks. Similarly, the storage system can be configured to support tools such as Amazon's Gluon, which allows developers to prototype, build, and train deep learning models. In fact, the storage system described above can be part of a larger platform, such as IBM TM Private cloud data, which includes integrated data science, data engineering, and application building services.
[0181] The reader will further appreciate that the storage system described above can also be deployed as an edge solution. This edge solution can be in place to optimize cloud computing systems by performing data processing at the edge of the network, close to the source of the data. Edge computing can push applications, data, and computing power (i.e., services) from centralized points to the logical extremes of the network. By using an edge solution, such as the storage system described above, computing tasks can be performed using the computing resources provided by such a storage system, data can be stored using the storage resources of the storage system, and cloud-based services can be accessed by using various resources (including networking resources) of the storage system. By performing computing tasks on the edge solution, storing data on the edge solution, and generally using the edge solution, it is possible to avoid consuming expensive cloud-based resources and, in fact, experience performance improvements relative to a heavier reliance on cloud-based resources.
[0182] While many tasks can benefit from utilizing edge solutions, some specific uses may be particularly well-suited for deployment in this environment. For example, devices such as drones, self-driving cars, robots, and others may require extremely fast processing—so fast, in fact, that sending data up to a cloud environment and receiving data processing support back may simply be too slow. As an additional example, some IoT devices (such as connected cameras) may not be well-suited to utilizing cloud-based resources because sending data to the cloud may be impractical (and not just from a privacy, security, or financial perspective) simply due to the sheer volume of data involved. Thus, many tasks that truly involve data processing, storage, or communication may be better suited for a platform that includes edge solutions, such as the storage systems described above.
[0183] The storage system described above can be used, alone or in combination with other computing resources, as a network edge platform that combines computing resources, storage resources, networking resources, cloud technologies, network virtualization technologies, and more. As part of the network, the edge can have characteristics similar to other network infrastructure, from customer premises and backhaul aggregation facilities to points of presence (PoPs) and regional data centers. The reader will appreciate that network workloads, such as virtual network functions (VNFs) and others, will reside on the network edge platform. Implemented through a combination of containers and virtual machines, the network edge platform can rely on controllers and schedulers that are no longer geographically co-located with data processing resources. Functionality can be split into microservices across control planes, user and data planes, or even state machines, allowing independent optimization and scaling techniques to be applied. Such user and data planes can be implemented through the addition of accelerators residing in server platforms (such as FPGAs and smart NICs) and implemented with SDN-enabled commodity silicon and programmable ASICs.
[0184] The storage system described above can also be optimized for big data analytics, including being used as part of a composable data analytics pipeline, where a containerized analytics architecture, for example, makes analytics capabilities more composable. Big data analytics can be broadly described as the process of examining large and diverse data sets to discover hidden patterns, unknown correlations, market trends, customer preferences, and other useful information that can help organizations make smarter business decisions. As part of that process, semi-structured and unstructured data (such as, for example, Internet clickstream data, web server logs, social media content, text from customer emails and survey responses, mobile phone call detail records, IoT sensor data, and other data) can be converted into a structured form.
[0185] The storage system described above may also support (including being implemented as a system interface) applications that perform tasks in response to human speech. For example, the storage system may support the execution of intelligent personal assistant applications such as (for example) Amazon's Alexa TM , Apple Siri TM , Google Voice TM , Samsung Bixby TM , Microsoft Cortana TM and others. While the examples described in the previous sentences use voice as input, the storage system described above may also support chatbots, conversational robots, chatterbots, or artificial conversational entities or other applications configured to conduct conversations via auditory or textual methods. Likewise, the storage system may actually execute such an application to enable a user (e.g., a system administrator) to interact with the storage system via voice. Such applications typically enable voice interaction, music playback, making to-do lists, setting alarms, streaming podcasts, playing audiobooks, and providing weather, traffic, and other real-time information, such as news, but in embodiments according to the present disclosure, such applications may be used as interfaces for various system management operations.
[0186] The storage systems described above can also implement AI platforms to realize the vision of self-driving storage. Such AI platforms can be configured to provide global predictive intelligence by collecting and analyzing numerous storage system telemetry data points, enabling easy management, analysis, and support. In fact, such storage systems may be able to predict both capacity and performance, as well as generate intelligent recommendations for workload deployment, interaction, and optimization. Such AI platforms can be configured to scan all incoming storage system telemetry data against a library of problem fingerprints to predict and resolve incidents in real time before they impact customer environments, capturing hundreds of performance-related variables for predicting performance loads.
[0187] The storage system described above can support the serial or simultaneous execution of artificial intelligence applications, machine learning applications, data analytics applications, data transformations, and other tasks that collectively form the AI ladder. This AI ladder can be effectively formed by combining these elements to form a complete data science pipeline, with dependencies between the elements of the AI ladder. For example, AI may require some form of machine learning, which may require some form of analytics, which may require some form of data and information systematization, and so on. Thus, each element can be considered a rung on the AI ladder, which together form a complete and complex AI solution.
[0188] The storage systems described above can also be used, alone or in combination with other computing environments, to deliver an AI-everywhere experience, where AI permeates a wide range of aspects of business and life. For example, AI may play a key role in delivering deep learning solutions, deep reinforcement learning solutions, artificial general intelligence solutions, autonomous vehicles, cognitive computing solutions, commercial UAVs or drones, conversational user interfaces, enterprise taxonomy, ontology management solutions, machine learning solutions, smart dust, smart robots, smart workplaces, and many others.
[0189] The storage system described above can also be used, alone or in combination with other computing environments, to provide a wide range of transparent immersive experiences (including experiences using digital twins of various "things" (e.g., people, places, processes, systems, etc.), where technology can introduce transparency between people, businesses, and things. This transparent immersive experience can be provided as augmented reality, connected homes, virtual reality, brain-computer interfaces, human augmentation, nanotube electronics, volumetric displays, 4D printing, or other technologies.
[0190] The storage systems described above can also be used, alone or in combination with other computing environments, to support a wide variety of digital platforms. These platforms may include, for example, 5G wireless systems and platforms, digital twin platforms, edge computing platforms, IoT platforms, quantum computing platforms, serverless PaaS, software-defined security, neuromorphic computing platforms, and more.
[0191] The storage system described above can also be part of a multi-cloud environment, where multiple cloud computing and storage services are deployed within a single heterogeneous architecture. To facilitate the operation of this multi-cloud environment, DevOps tools can be deployed to enable cross-cloud orchestration. Similarly, continuous development and continuous integration tools can be deployed to standardize processes around continuous integration and delivery, new feature rollouts, and provisioning cloud workloads. By standardizing these processes, a multi-cloud strategy can be implemented that leverages the best provider for each workload.
[0192] The storage system described above can be used as part of a platform to enable the use of cryptographic anchors, which can be used to authenticate the origin and content of a product to ensure that it matches the blockchain record associated with the product. Similarly, as part of a toolkit to protect data stored on the storage system, the storage system described above can implement various cryptographic techniques and schemes, including lattice cryptography. Grid cryptography can involve the construction of cryptographic primitives that involve a lattice, either in the construction itself or in security proofs. Unlike public key schemes such as RSA, Diffie-Hellman, or elliptic curve cryptography, which are vulnerable to quantum computer attacks, some lattice-based constructions appear to be resistant to attacks by both classical and quantum computers.
[0193] A quantum computer is a device that performs quantum computations. Quantum computing uses quantum mechanical phenomena, such as superposition and entanglement, to perform calculations. Quantum computers differ from traditional transistor-based computers because such computers require data to be encoded as binary digits (bits), each of which is always in one of two definite states (0 or 1). Unlike traditional computers, quantum computers use qubits that can be in a superposition of states. A quantum computer maintains a sequence of qubits, where a single qubit can represent one, zero, or any quantum superposition of the states of those two qubits. A pair of qubits can be in any quantum superposition of four states, and three qubits can be in any superposition of eight states. A quantum computer with n qubits can typically be in any superposition of up to 2^n different states simultaneously, while a traditional computer can only be in one of these states at any one time. A quantum Turing machine is a theoretical model of this computer.
[0194] The storage systems described above can also be paired with FPGA acceleration servers as part of a larger AI or ML infrastructure. Such FPGA acceleration servers can reside near the storage systems described above (e.g., in the same data center) or even be incorporated into an appliance that includes one or more storage systems, one or more FPGA acceleration servers, a networking infrastructure that supports communication between the one or more storage systems and the one or more FPGA acceleration servers, and other hardware and software components. Alternatively, the FPGA acceleration servers can reside within a cloud computing environment that can be used to perform the computation-related tasks of AI and ML jobs. Any of the embodiments described above can be used together as an FPGA-based AI or ML platform. The reader will appreciate that in some embodiments of an FPGA-based AI or ML platform, the FPGA contained within the FPGA acceleration server can be reconfigured for different types of ML models (e.g., LSTM, CNN, GRU). The ability to reconfigure the FPGA contained within the FPGA acceleration server can accelerate ML or AI applications based on the optimal numerical precision and memory model used. The reader will appreciate that by treating a collection of FPGA acceleration servers as an FPGA pool, any CPU in the data center can use the FPGA pool as a shared hardware microservice, rather than limiting the server to the dedicated accelerators plugged into it.
[0195] The FPGA-accelerated servers and GPU-accelerated servers described above can implement a computing model in which, rather than maintaining a small amount of data in a CPU and running a long instruction stream through it, as occurs in more traditional computing models, machine learning models and parameters are fixed to high-bandwidth single-chip memory through which a large amount of data flows. For this computing model, FPGAs can be even more efficient than GPUs because they can be programmed with only the instructions required to run such a computing model.
[0196] The storage system described above can be configured to provide parallel storage, for example, by using a parallel file system such as BeeGFS. Such a parallel file system can include a distributed metadata architecture. For example, a parallel file system can include multiple metadata servers across which metadata is distributed, as well as components that include services for clients and storage servers.
[0197] The system described above can support the execution of a wide range of software applications. Such software applications can be deployed in a variety of ways, including container-based deployment models. Containerized applications can be managed using various tools. For example, Docker Swarm, Kubernetes, and others can be used to manage containerized applications. Containerized applications can be used to facilitate serverless, cloud-native computing deployment and management models for software applications. To support serverless, cloud-native computing deployment and management models for software applications, containers can be used as part of an event handling mechanism (e.g., AWS Lambdas), such that various events cause the containerized application to be launched to operate as an event handler.
[0198] The systems described above can be deployed in various ways, including to support fifth-generation ('5G') networks. 5G networks can support substantially faster data communications than previous generations of mobile communication networks and, therefore, can lead to disaggregation of data and computing resources, as modern large-scale data centers can become less prominent and, for example, replaced by more local micro-data centers located near mobile network towers. The systems described above can be included in such local micro-data centers and can be part of or paired with multi-access edge computing ('MEC') systems. Such MEC systems can enable cloud computing capabilities and IT service environments at the edge of cellular networks. By running applications and performing related processing tasks closer to cellular clients, network congestion can be reduced and applications can perform better.
[0199] The storage system described above may also be configured to implement NVMe zoned namespaces. By using NVMe zoned namespaces, the logical address space of the namespace is divided into zones. Each zone provides a range of logical block addresses that must be written sequentially and explicitly reset before being overwritten, thereby enabling the creation of natural boundaries that expose the device and offloading the management of internal mapping tables to the host's namespace. To implement NVMe zoned namespaces ('ZNS'), a ZNS SSD or some other form of zoned block device may be utilized that uses zones to expose the namespace logical address space. In aligning the zones to the internal physical properties of the device, several inefficiencies in data placement may be eliminated. In such embodiments, for example, each zone may be mapped to a separate application so that functions such as wear leveling and garbage collection may be performed on a per-zone or per-application basis (rather than across the entire device). To support ZNS, the storage controller described herein may be configured to implement a ZNS SSD using, for example, a Linux TM Kernel partition block device interface or other tools to interact with partition block devices.
[0200] The storage system described above can also be configured to implement zoned storage in other ways, such as, for example, by using shingled magnetic recording (SMR) storage devices. In instances where zoned storage is used, a device-managed embodiment can be deployed, where the storage device hides this complexity by managing it in firmware, presenting an interface like any other storage device. Alternatively, zoned storage can be implemented via a host-managed embodiment, which relies on the operating system knowing how to handle the drive and only writing sequentially to certain areas of the drive. Zoned storage can similarly be implemented using a host-aware embodiment, where a combination of drive-managed and host-managed implementations is deployed.
[0201] The storage system described herein can be used to form a data lake. A data lake can operate as the first place an organization's data flows to, where such data may be in its original format. Metadata tagging can be implemented to facilitate searching of data elements in the data lake, particularly in embodiments where the data lake contains multiple data stores in formats that may not be easily accessed or read (e.g., unstructured data, semi-structured data, structured data). Data can be streamed from the data lake to a data warehouse, where it can be stored in a more processed, packaged, and consumable format. The storage system described above can also be used to implement such a data warehouse. Additionally, a data mart or data hub can allow for even more easily consumable data, where the storage system described above can also be used to provide the underlying storage resources required for the data mart or data hub. In embodiments, querying the data lake may require a schema-on-read approach, where a schema or pattern is applied to the data as it is pulled from the storage location, rather than as it enters the storage location.
[0202] The storage systems described herein may also be configured to implement a recovery point objective ('RPO'), which may be established by a user, by an administrator, as a system default, as part of a storage class or service that the storage system participates in delivering, or in some other manner. A "recovery point objective" is a target for the maximum time difference between the last update to a source dataset and the last recoverable update to a replicated dataset that, given some reason, can be correctly recovered from a continuously or frequently updated copy of the source dataset. Updates can be correctly recovered if all updates processed on the source dataset prior to the last recoverable update to the replicated dataset are properly accounted for.
[0203] In synchronous replication, the RPO will be zero, meaning that under normal operation, all updates completed on the source dataset should be present and correctly recoverable on the replica dataset. In as-close-to-synchronous replication as possible, the RPO can be as low as a few seconds. In snapshot-based replication, the RPO can be roughly calculated as the interval between snapshots plus the time to transfer modifications between the previously transferred snapshot and the most recent snapshot to be replicated.
[0204] If updates accumulate faster than they can be replicated, then the RPO can be missed. If more data to be replicated accumulates between two snapshots (for snapshot-based replication) than can be replicated between taking a snapshot and replicating the cumulative updates for that snapshot to the replica, then the RPO can be missed. Again in snapshot-based replication, if data to be replicated accumulates at a faster rate than it can be transferred in the time between subsequent snapshots, then replication can begin to fall further behind, which can extend the miss between the expected recovery point objective and the actual recovery point represented by the last correctly replicated update.
[0205] The storage systems described above may also be part of a shared-nothing storage cluster. In a shared-nothing storage cluster, each node of the cluster has local storage and communicates with the other nodes in the cluster over a network, where the storage used by the cluster is (typically) provided only by the storage connected to each individual node. A collection of nodes that synchronously replicate data sets may be an instance of a shared-nothing storage cluster because each storage system has local storage and communicates with the other storage systems over a network, where those storage systems (typically) do not use storage from elsewhere that they have shared access to over some kind of interconnect. In contrast, some of the storage systems described above are inherently built as shared storage clusters because there are drive racks that are shared by paired controllers. However, other storage systems described above are built as shared-nothing storage clusters because all storage is local to a particular node (e.g., a blade) and all communication is through the network that links the compute nodes together.
[0206] In other embodiments, other forms of shared-nothing storage clusters may include embodiments in which any node in the cluster has a local copy of all storage it needs, and in which data is mirrored to other nodes in the cluster via synchronous replication to ensure that data is not lost or lost because other nodes are also using that storage. In this embodiment, if a new cluster node needs some data, that data can be copied to the new node from other nodes that have copies of the data.
[0207] In some embodiments, a shared storage cluster based on mirror replication may store multiple copies of all cluster-stored data, where each subset of the data is replicated to a specific set of nodes, and different subsets of the data are replicated to several different sets of nodes. In some variations, embodiments may store all cluster-stored data on all nodes, while in other variations, the nodes may be partitioned such that a first set of nodes all store the same set of data, while a second, different set of nodes all store different sets of data.
[0208] The reader will appreciate that a RAFT-based database (e.g., etcd) can operate as a shared-nothing storage cluster, where all RAFT nodes store all data. However, the amount of data stored in a RAFT cluster may be limited so that additional replicas do not consume excessive storage. A container server cluster may also be able to replicate all data to all cluster nodes, provided that the containers are not too large and the bulk of their data (the data manipulated by the applications running in the containers) is stored elsewhere, such as in an S3 cluster or an external file server. In this instance, container storage can be provided directly by the cluster through its shared-nothing storage model, where those containers provide images of the execution environments that form part of the applications or services.
[0209] To explain further, Figure 3DAn exemplary computing device 350 is illustrated that may be specifically configured to perform one or more of the processes described herein. Figure 3D As shown in FIG, computing device 350 may include a communication interface 352, a processor 354, a storage device 356, and an input / output ("I / O") module 358 communicatively connected to each other via a communication infrastructure 360. Figure 3D An exemplary computing device 350 is shown in FIG. Figure 3D The components described in the drawings are not intended to be limiting. In other embodiments, additional or alternative components may be used. Figure 3D Components of computing device 350 are shown in FIG.
[0210] Communication interface 352 can be configured to communicate with one or more computing devices. Examples of communication interface 352 include, without limitation, a wired network interface (such as a network interface card), a wireless network interface (such as a wireless network interface card), a modem, an audio / video connection, and any other suitable interface.
[0211] The processor 354 generally represents any type or form of processing unit capable of processing data and / or interpreting, executing, and / or directing the execution of one or more of the instructions, processes, and / or operations described herein. The processor 354 may perform operations by executing computer-executable instructions 362 (e.g., applications, software, code, and / or other executable data instances) stored in the storage device 356.
[0212] Storage device 356 may include one or more data storage media, devices, or configurations and may employ any type, form, and combination of data storage media and / or devices. For example, storage device 356 may include, but is not limited to, any combination of non-volatile media and / or volatile media described herein. Electronic data, including the data described herein, may be temporarily and / or permanently stored in storage device 356. For example, data representing computer-executable instructions 362 configured to direct processor 354 to perform any of the operations described herein may be stored within storage device 356. In some examples, the data may be arranged in one or more databases residing within storage device 356.
[0213] The I / O module 358 may include one or more I / O modules configured to receive user input and provide user output. The I / O module 358 may include any hardware, firmware, software, or combination thereof that supports input and output capabilities. For example, the I / O module 358 may include hardware and / or software for capturing user input, including but not limited to a keyboard or keypad, a touch screen component (e.g., a touch screen display), a receiver (e.g., an RF or infrared receiver), a motion sensor, and / or one or more input buttons.
[0214] The I / O module 358 may include one or more devices for presenting output to a user, including, but not limited to, a graphics engine, a display (e.g., a display screen), one or more output drivers (e.g., a display driver), one or more audio speakers, and one or more audio drivers. In certain embodiments, the I / O module 358 is configured to provide graphical data to the display for presentation to the user. The graphical data may represent one or more graphical user interfaces and / or any other graphical content that may serve a particular implementation. In some examples, any of the systems, computing devices, and / or other components described herein may be implemented by the computing device 350.
[0215] To explain further, Figure 3E An example of a fleet of storage systems 376 for providing storage services (also referred to herein as 'data services') is illustrated. Figure 3E The fleet of storage systems 376 depicted in FIG includes a plurality of storage systems 374a, 374b, 374c through 374n that may each be similar to the storage systems described herein. The storage systems 374a, 374b, 374c through 374n in the fleet of storage systems 376 may be embodied as the same storage system or as different types of storage systems. For example, Figure 3E The two storage systems 374a, 374n depicted in FIG are depicted as cloud-based storage systems because the resources that collectively form each of the storage systems 374a, 374n are provided by different cloud service providers 370, 372. For example, the first cloud service provider 370 may be Amazon AWS. TM , and the second cloud service provider 372 is Microsoft Azure TM , but in other embodiments, one or more public clouds, private clouds, or a combination thereof may be used to provide underlying resources that are used to form a particular storage system in the fleet of storage systems 376.
[0216] According to some embodiments of the present disclosure, Figure 3E The example depicted in includes an edge management service 366 for delivering storage services. The storage services (also referred to herein as 'data services') delivered can include, for example, services that provide a specific amount of storage to a customer, services that provide storage to a customer according to a predetermined service level agreement, services that provide storage to a customer according to predetermined regulatory requirements, and many others.
[0217] Figure 3EThe edge management service 366 depicted in FIG3 may be embodied, for example, as one or more modules of computer program instructions that execute on computer hardware (e.g., one or more computer processors). Alternatively, the edge management service 366 may be embodied as one or more modules of computer program instructions that execute on a virtualized execution environment (e.g., one or more virtual machines), in one or more containers, or in some other manner. In other embodiments, the edge management service 366 may be embodied as a combination of the embodiments described above, including embodiments in which the one or more modules of computer program instructions contained in the edge management service 366 are distributed across multiple physical or virtual execution environments.
[0218] The edge management service 366 can operate as a gateway for providing storage services to storage customers, where the storage services utilize storage provided by one or more storage systems 374a, 374b, 374c, through 374n. For example, the edge management service 366 can be configured to provide storage services to host devices 378a, 378b, 378c, 378d, and 378n that execute one or more applications that consume the storage services. In this example, the edge management service 366 can operate as a gateway between the host devices 378a, 378b, 378c, 378d, and 378n and the storage systems 374a, 374b, 374c, through 374n, without requiring the host devices 378a, 378b, 378c, 378d, and 378n to directly access the storage systems 374a, 374b, 374c, through 374n.
[0219] Figure 3E The edge management service 366 exposes the storage service module 364 to Figure 3E In other embodiments, the edge management service 366 may expose the storage services module 364 to other clients of the respective storage services. The respective storage services may be presented to the client via one or more user interfaces, via one or more APIs, or through some other mechanism provided by the storage services module 364. Figure 3E The storage service module 364 depicted in the figure may be embodied as one or more modules of computer program instructions executed on physical hardware, on a virtualized execution environment, or a combination thereof, wherein the execution of such modules enables customers of the storage service to be provided with, select, and access various storage services.
[0220] Figure 3E The edge management service 366 also includes a system management service module 368. Figure 3EThe system management service module 368 includes one or more computer program instruction modules that, when executed, coordinate with the storage systems 374a, 374b, 374c to 374n to perform various operations to provide storage services to the host devices 378a, 378b, 378c, 378d, and 378n. The system management service module 368 may be configured, for example, to perform tasks such as allocating storage resources from the storage systems 374a, 374b, 374c, through 374n via one or more APIs exposed by the storage systems 374a, 374b, 374c, through 374n, migrating data sets or workloads among the storage systems 374a, 374b, 374c, through 374n via one or more APIs exposed by the storage systems 374a, 374b, 374c, through 374n, setting one or more tunable parameters (i.e., one or more configurable settings) on the storage systems 374a, 374b, 374c, through 374n via one or more APIs exposed by the storage systems 374a, 374b, 374c, through 374n, etc. For example, many of the services described below are related to embodiments in which the storage systems 374a, 374b, 374c, through 374n are configured to operate in a certain manner. In such instances, the system management services module 368 may be responsible for configuring the storage systems 374a, 374b, 374c through 374n to operate in the manner described below using an API (or some other mechanism) provided by the storage systems 374a, 374b, 374c through 374n.
[0221] In addition to configuring the storage systems 374a, 374b, 374c, through 374n, the edge management service 366 itself can be configured to perform various tasks required to provide various storage services. Consider an example in which the storage service includes a service that, when selected and applied, causes personally identifiable information (PII) contained in a data set to be obfuscated when the data set is accessed. In this example, the storage systems 374a, 374b, 374c, through 374n can be configured to obfuscate PII when servicing read requests directed to the data set. Alternatively, the storage systems 374a, 374b, 374c, through 374n can service the read by returning data containing PII, but the edge management service 366 itself can obfuscate the PII as the data passes through the edge management service 366 on its way from the storage systems 374a, 374b, 374c, through 374n to the host devices 378a, 378b, 378c, 378d, and 378n.
[0222] Figure 3E The storage systems 374a, 374b, 374c to 374n depicted in FIG. 3 may be embodied as described above with reference to FIG. Figures 1A to 3DOne or more of the storage systems described herein include variations thereof. In practice, storage systems 374a, 374b, 374c, and 374n may function as a storage resource pool, wherein individual components within that pool may have different performance characteristics, different storage characteristics, and so forth. For example, one of storage systems 374a may be a cloud-based storage system, another storage system 374b may be a storage system that provides block storage, another storage system 374c may be a storage system that provides file storage, another storage system 374d may be a relatively high-performance storage system, and another storage system 374n may be a relatively low-performance storage system, and so forth. In alternative embodiments, only a single storage system may be present.
[0223] Figure 3E The storage systems 374a, 374b, 374c through 374n depicted in FIG3 can also be organized into different failure domains such that a failure of one storage system 374a should be completely independent of a failure of another storage system 374b. For example, each of the storage systems can receive power from an independent power system, each of the storage systems can be coupled for data communication via an independent data communication network, and so on. Furthermore, the storage systems in a first failure domain can be accessed via a first gateway, while the storage systems in a second failure domain can be accessed via a second gateway. For example, the first gateway can be a first instance of the edge management service 366, and the second gateway can be a second instance of the edge management service 366, including embodiments in which each instance is distinct or each instance is part of a distributed edge management service 366.
[0224] As an illustrative example of available storage services, storage services may be presented to a user associated with different levels of data protection. For example, storage services may be presented to a user that, when selected and enforced, assure the user that data associated with that user will be protected such that various recovery point objectives ('RPOs') can be guaranteed. A first available storage service may ensure, for example, that a certain data set associated with a user will be protected such that any data older than 5 seconds can be recovered in the event of a primary data store failure, while a second available storage service may ensure that a data set associated with the user will be protected such that any data older than 5 minutes can be recovered in the event of a primary data store failure.
[0225] Additional examples of storage services that can be presented to a user, selected by the user, and ultimately applied to a dataset associated with the user may include one or more data compliance services. Such data compliance services may be embodied as services that can, for example, provide data compliance services to customers (i.e., users) to ensure that the user's dataset is managed in a manner that complies with various regulatory requirements. For example, one or more data compliance services may be provided to a user to ensure that the user's dataset is managed in a manner that complies with the General Data Protection Regulation ('GDPR'), one or more data compliance services may be provided to a user to ensure that the user's dataset is managed in a manner that complies with the Sarbanes-Oxley Act of 2002 ('SOX'), or one or more data compliance services may be provided to a user to ensure that the user's dataset is managed in a manner that complies with some other regulatory act. In addition, one or more data compliance services may be provided to a user to ensure that the user's dataset is managed in a manner that complies with certain non-governmental guidance (e.g., best practices for auditing purposes), one or more data compliance services may be provided to a user to ensure that the user's dataset is managed in a manner that complies with specific client or organizational requirements, and so on.
[0226] To provide this particular data compliance service, the data compliance service may be presented to a user (e.g., via a GUI) and selected by the user. In response to receiving a selection of a particular data compliance service, one or more storage service policies may be applied to a data set associated with the user to implement the particular data compliance service. For example, a storage service policy may be applied that requires that a data set be encrypted before being stored in a storage system, before being stored in a cloud environment, or before being stored elsewhere. To enforce this policy, provisions may be enforced that not only require that the data set be encrypted when the data set is stored, but also that the data set be encrypted before being transmitted (e.g., sent to another party). In this instance, a storage service policy may also be set that requires that any encryption key used to encrypt the data set be not stored on the same system that stores the data set itself. The reader will understand that many other forms of data compliance services may be provided and implemented in accordance with embodiments of the present disclosure.
[0227] The storage systems 374a, 374b, 374c to 374n in the fleet storage system 376 may be jointly managed by one or more fleet management modules. The fleet management module may be Figure 3EThe fleet management module may be part of or separate from the system management services module 368 depicted in FIG. The fleet management module may perform tasks such as monitoring the health of each storage system in the fleet, initiating updates or upgrades on one or more storage systems in the fleet, migrating workloads for load balancing or other performance purposes, and many other tasks. For this reason, and for many other reasons, the storage systems 374 a, 374 b, 374 c through 374 n may be coupled to each other via one or more data communication links to exchange data between the storage systems 374 a, 374 b, 374 c through 374 n.
[0228] In some embodiments, one or more storage systems or one or more elements of a storage system (e.g., features, services, operations, components, etc. of a storage system), such as any of the illustrative storage systems or storage system elements described herein, can be implemented in one or more container systems. A container system can include any system that supports the execution of one or more containerized applications or services. Such a service can be software deployed as infrastructure for building applications, for operating a runtime environment, and / or as infrastructure for other services. In the following discussion, descriptions of containerized applications generally apply equally to containerized services.
[0229] Containers can combine one or more elements of a containerized software application with a runtime environment for operating those elements of the software application bundled into a single image. For example, each container of a containerized application can contain the software application's executable code and various dependencies, libraries, and / or other components, along with network configuration and configured access to additional resources, which are used by the elements of the software application within a particular container to enable the operation of those elements. A containerized application can be represented as a collection of such containers, which together represent all elements of the application combined with the various runtime environments required for all of those elements to operate. Thus, a containerized application can be abstracted from the host operating system as a combined set of lightweight and portable packages and configurations, where the containerized application can be uniformly deployed and executed consistently in different computing environments using different container-compatible operating systems or different infrastructures. In some embodiments, the containerized application shares a kernel with the host computer system and executes as an isolated environment (an isolated set of files and directories, processes, system and network resources, and configured access to additional resources and capabilities) that is isolated by the host system's operating system in conjunction with a container management framework. When executed, the containerized application can provide one or more containerized workloads and / or services.
[0230] A container system may include and / or utilize a cluster of nodes. For example, a container system may be configured to manage the deployment and execution of containerized applications on one or more nodes in a cluster. Containerized applications may utilize resources of a node (e.g., memory), processing and / or storage resources provided and / or accessed by the node. Storage resources may include any of the illustrative storage resources described herein and may include on-node resources (e.g., a local tree of files and directories), off-node resources (e.g., an external networked file system, database, or object store), or both on-node and off-node resources. Access to additional resources and capabilities of containers that may be configured for containerized applications may include specialized computing capabilities (e.g., GPUs and AI / ML engines) or specialized hardware (e.g., sensors and cameras).
[0231] In some embodiments, the container system may include a container orchestration system (which may also be referred to as a container orchestrator, container orchestration platform, etc.) designed to make it relatively simple and automated for many use cases to deploy, scale, and manage containerized applications. In some embodiments, the container system may include a storage management system configured to provision and manage storage resources (e.g., virtual volumes) for private or shared use by cluster nodes and / or containers of the containerized application.
[0232] Figure 3F An example container system 380 is illustrated. In this example, the container system 380 includes a container storage system 381 that can be configured to perform one or more storage management operations to organize, allocate, and manage storage resources for use by one or more containerized applications 382-1 through 382-L of the container system 380. In particular, the container storage system 381 can organize storage resources into one or more storage pools 383 of storage resources for use by the containerized applications 382-1 through 382-L. The container storage system itself can be implemented as a containerized service.
[0233] Container system 380 may include or be implemented by one or more container orchestration systems, including Kubernetes TM 、Mesos TM 、Docker Swarm TM The container orchestration system may manage the container system 380 running on a cluster 384 through services implemented by a control node depicted as 385, and may further manage the relationship between the container storage system or individual containers and their storage, memory and CPU limits, networking, and their access to additional resources or services.
[0234] The control plane of container system 380 can implement services including deploying applications via controller 386, monitoring applications via controller 386, providing interfaces via API server 387, and scheduling deployments via scheduler 388. In this example, controller 386, scheduler 388, API server 387, and container storage system 381 are implemented on a single node, node 385. In other examples, for resiliency, the control plane can be implemented by multiple redundant nodes, where if a node providing management services for container system 380 fails, another redundant node can provide management services for cluster 384.
[0235] The data plane of the container system 380 may include a set of nodes that provide a container runtime for executing containerized applications. Individual nodes within the cluster 384 may execute a container runtime, such as Docker. TM , and executes a container manager or node agent, such as the kubelet in Kubernetes (not depicted), which communicates with the control plane via a local network-connected agent (sometimes referred to as a proxy), such as proxy 389. Proxy 389 can route network traffic to and from containers using, for example, Internet Protocol (IP) port numbers. For example, a containerized application can request a storage class from the control plane, with the request being handled by the container manager, which then passes the request to the control plane using proxy 389.
[0236] Cluster 384 may include a set of nodes that run containers of managed containerized applications. Nodes may be virtual or physical machines. Nodes may be host systems.
[0237] Container storage system 381 can orchestrate storage resources to provide storage to container system 380. For example, container storage system 381 can provide persistent storage to containerized applications 382-1 through 382-L using storage pool 383. Container storage system 381 itself can be deployed as a containerized application by a container orchestration system.
[0238] For example, a container storage system 381 application may be deployed within a cluster 384 and perform management functions for providing storage to containerized applications 382. Management functions may include determining one or more storage pools from available storage resources, assigning virtual volumes on one or more nodes, replicating data, responding to and recovering from host and network failures, or handling storage operations. Storage pool 383 may include storage resources from one or more local or remote sources, where the storage resources may be different types of storage, including, for example, block storage, file storage, and object storage.
[0239] Container storage system 381 can also be deployed on a set of nodes for which a container orchestration system can provide persistent storage. In some examples, container storage system 381 can be deployed on all nodes in cluster 384 using, for example, a Kubernetes DaemonSet. In this example, nodes 390-1 through 390-N provide the container runtime in which container storage system 381 executes. In other examples, some, but not all, nodes in the cluster can execute container storage system 381.
[0240] Container storage system 381 can handle storage on the node and communicate with the control plane of container system 380 to provide dynamic volumes, including persistent volumes. Persistent volumes can be mounted on the node as virtual volumes, such as virtual volumes 391-1 and 391-P. After virtual volume 391 is mounted, containerized applications can request and use, or be otherwise configured to use, the storage provided by virtual volume 391. In this example, container storage system 381 can install a driver on the node's kernel, where the driver handles storage operations directed to the virtual volume. In this example, the driver can receive storage operations directed to the virtual volume and, in response, perform storage operations on one or more storage resources within storage pool 383, possibly under the direction of or using additional logic within the container that implements container storage system 381 as a containerized service.
[0241] Container storage system 381 can determine available storage resources in response to being deployed as a containerized service. For example, storage resources 392-1 through 392-M can include local storage, remote storage (storage on separate nodes in the cluster), or both local and remote storage. Storage resources can also include storage from external sources, such as various combinations of block storage systems, file storage systems, and object storage systems. Storage resources 392-1 through 392-M can include storage resources of any type and / or configuration (e.g., any of the illustrative storage resources described above), and container storage system 381 can be configured to determine available storage resources in any suitable manner, including based on a configuration file. For example, the configuration file can specify account and authentication information for cloud-based object storage device 348 or cloud-based storage system 318. Container storage system 381 can also determine the availability of one or more storage devices 356 or one or more storage systems. Aggregate storage from one or more of storage 356, storage systems, cloud-based storage system 318, edge management service 366, cloud-based object storage 348, or any other storage resources, or any combination or subcombination of such storage resources, may be used to provide storage pool 383. Storage pool 383 is used to assign storage to one or more virtual volumes mounted on one or more of nodes 390 within cluster 384.
[0242] In some embodiments, container storage system 381 can create multiple storage pools. For example, container storage system 381 can aggregate storage resources of the same type into separate storage pools. In this example, the storage type can be one of the following: storage device 356, storage array 102, cloud-based storage system 318, storage via edge management service 366, or cloud-based object storage device 348. Alternatively, it can be storage configured with a specific level or type of redundancy or distribution, such as a specific combination of striping, mirroring, or erasure coding.
[0243] Container storage system 381 can be executed within cluster 384 as a containerized container storage system service, where instances of containers implementing elements of the containerized container storage system service can operate on different nodes within cluster 384. In this example, the containerized container storage system service can integrate with the container orchestration system operations of container system 380 to handle storage operations, mount virtual volumes to provide storage to nodes, aggregate available storage into storage pool 383, allocate storage to virtual volumes from storage pool 383, generate backup data, replicate data between nodes, clusters, and environments, and other storage system operations. In some examples, the containerized container storage system service can provide storage services across multiple clusters operating in disparate computing environments. For example, other storage system operations may include the storage system operations described herein. The persistent storage provided by the containerized container storage system service can be used to implement stateful and / or resilient containerized applications.
[0244] The container storage system 381 can be configured to perform any suitable storage operation of a storage system. For example, the container storage system 381 can be configured to perform one or more of the illustrative storage management operations described herein to manage storage resources used by the container system.
[0245] In some embodiments, one or more storage operations, including one or more of the illustrative storage management operations described herein, may be containerized. For example, one or more storage operations may be implemented as one or more containerized applications configured to be executed to perform the storage operations. Such containerized storage operations may be executed in any suitable runtime environment to manage any storage system, including any of the illustrative storage systems described herein.
[0246] The storage systems described herein can support various forms of data replication. For example, two or more storage systems can synchronously replicate a dataset between each other. In synchronous replication, distinct copies of a particular dataset may be maintained by multiple storage systems, but all accesses (e.g., reads) to the dataset should produce consistent results regardless of which storage system the access is directed to. For example, reads directed to any storage system that is synchronously replicating the dataset should return the same result. Thus, while updates to dataset versions need not occur completely simultaneously, precautions must be taken to ensure consistent access to the dataset. For example, if an update (e.g., a write) directed to a dataset is received by a first storage system, the update can only be considered complete if all storage systems that are synchronously replicating the dataset have applied the update to their copies of the dataset. In this example, synchronous replication can be implemented using I / O forwarding (e.g., a write received at a first storage system is forwarded to a second storage system), communication between storage systems (e.g., each storage system indicates that it has completed the update), or other methods.
[0247] In other embodiments, a data set may be replicated using checkpoints. In checkpoint-based replication (also known as 'almost synchronous replication'), a set of updates to a data set (e.g., one or more write operations directed to the data set) may occur between different checkpoints, such that the data set is updated to a particular checkpoint only when all updates to the data set prior to the particular checkpoint have completed. Consider an example in which a first storage system stores a real-time copy of a data set that is being accessed by a data set user. In this example, assume that the data set is replicated from the first storage system to a second storage system using checkpoint-based replication. For example, the first storage system may send a first checkpoint (at time t=0) to the second storage system, followed by a first set of updates to the data set, followed by a second checkpoint (at time t=1), followed by a second set of updates to the data set, followed by a third checkpoint (at time t=2). In this example, if the second storage system has executed all updates in the first set of updates but has not yet executed all updates in the second set of updates, then the copy of the data set stored on the second storage system may be up to date up to the second checkpoint. Alternatively, if the second storage system has performed all updates in the first set of updates and the second set of updates, then the copy of the data set stored on the second storage system may be up to date up to the third checkpoint. The reader will recognize that various types of checkpoints may be used (e.g., metadata-only checkpoints), checkpoints may be spread based on various factors (e.g., time, number of operations, RPO settings), etc.
[0248] In other embodiments, a dataset may be replicated via snapshot-based replication (also known as 'asynchronous replication'). In snapshot-based replication, a snapshot of a dataset may be sent from a replication source (e.g., a first storage system) to a replication target (e.g., a second storage system). In this embodiment, each snapshot may include the entire dataset or a subset of the dataset, such as, for example, only the portion of the dataset that has changed since the last snapshot was sent from the replication source to the replication target. The reader will appreciate that snapshots may be sent on demand based on a policy that takes into account various factors (e.g., time, number of operations, RPO settings), or in some other manner.
[0249] The storage systems described above can be configured, alone or in combination, to function as continuous data protection storage. Continuous data protection storage is a feature of a storage system that records updates to a dataset in such a way that a consistent image of the previous contents of the dataset can be accessed at a low temporal granularity (typically on the order of seconds or even less), and extended back for a reasonable period of time (typically hours or days). This allows access to the most recent consistent point in time for a dataset, and also allows access to a point in time for a dataset where an event may have just occurred (e.g., causing a portion of the dataset to be corrupted or otherwise lost), while maintaining a maximum number of updates close to that event. Conceptually, they are like a series of snapshots of a dataset that are taken very frequently and maintained for a long period of time, but continuous data protection storage is typically implemented quite differently from snapshots. Storage systems that implement continuous data protection storage may further provide a means of accessing these points in time, accessing one or more of these points in time as snapshots or clones, or restoring the dataset back to one of these recorded points in time.
[0250] Over time, to reduce overhead, some of the time points maintained in the continuous data protection storage can be merged with other nearby time points, essentially removing some of these time points from the storage. This can reduce the capacity required to store updates. A limited number of these time points can also be converted into snapshots of longer duration. For example, the storage can maintain a low-granularity sequence of time points going back a few hours, with some time points merged or removed to reduce overhead by up to an additional day. Going even further back in time, some of these time points can be converted into snapshots representing consistent time images from only every few hours.
[0251] Although some embodiments are primarily described in the context of storage systems, readers skilled in the art will recognize that embodiments of the present disclosure may also take the form of a computer program product disposed on a computer-readable storage medium for use with any suitable processing system. Such computer-readable storage media may be any storage medium for machine-readable information, including magnetic media, optical media, solid-state media, or other suitable media. Examples of such media include magnetic disks in hard drives or floppy disks, optical disks in optical drives, magnetic tape, and others as will occur to those skilled in the art. Those skilled in the art will immediately recognize that any computer system with appropriate programming means will be capable of performing the steps described herein as embodied in a computer program product. Those skilled in the art will also recognize that although some embodiments described in this specification are directed to software installed and executed on computer hardware, alternative embodiments implemented as firmware or hardware are also within the scope of the present disclosure.
[0252] In some examples, a non-transitory computer-readable medium storing computer-readable instructions may be provided according to the principles described herein. When executed by a processor of a computing device, the instructions may direct the processor and / or the computing device to perform one or more operations, including one or more of the operations described herein. Any of a variety of known computer-readable media may be used to store and / or transmit such instructions.
[0253] As used herein, non-transitory computer-readable media may include any non-transitory storage medium that participates in providing data (e.g., instructions) that can be read and / or executed by a computing device (e.g., by a processor of the computing device). For example, non-transitory computer-readable media may include, but are not limited to, any combination of non-volatile storage media and / or volatile memory media. Exemplary non-volatile storage media include, but are not limited to, read-only memory, flash memory, solid-state drives, magnetic storage devices (e.g., hard disks, floppy disks, magnetic tape, etc.), ferroelectric random access memory ("RAM"), and optical disks (e.g., compact disks, digital video disks, Blu-ray disks, etc.). Exemplary volatile storage media include, but are not limited to, RAM (e.g., dynamic RAM).
[0254] Figure 4 4 is a block diagram illustrating an example storage system 400 according to some embodiments of the present disclosure. As discussed above, the storage system 400 may be used to store data and / or provide access to data to clients / computing devices. Figure 4, storage system 400 includes storage clusters 407A, 407B, and 407C. A storage cluster may be a group of storage nodes (e.g., one or more chassis, one or more racks, etc.). Storage clusters 407A, 407B, and 407C may be communicatively coupled to each other via a network 705 (e.g., one or more of a LAN, a WAN, a wireless network, a wired network, the Internet, etc.). Figure 4 Storage clusters 407A, 407B, and 407C are described as communicating via network 705. However, in other embodiments, some of storage clusters 407A, 407B, and 407C may be directly coupled to each other via a communication bridge, an interconnect, a bus, wires, etc. For example, storage cluster 407A may be directly coupled to storage cluster 407B. In another example, storage cluster 407A may be directly coupled to storage cluster 407B and coupled to storage cluster 407B via network 705.
[0255] Storage clusters 407A to 407C include different types of storage nodes and / or may include any number of storage nodes. For example, storage node 407A may include a group of primary storage nodes 410 (e.g., one or more primary storage nodes 410). Storage node 407B may include a group of secondary storage nodes (e.g., one or more secondary storage nodes 420). Storage node 407C may include both a group of primary storage nodes 410 and a group of secondary storage nodes 420. Primary storage node 410 may be referred to as a head node, a head storage node, a control node, a control storage node, a primary node, a primary storage node, a master node, etc. Secondary storage node 420 may be referred to as a slave node, a slave storage node, a secondary node, an extension node, an extension storage node, etc.
[0256] In one embodiment, the primary storage node 410 includes one or more processing devices 411 (e.g., one or more CPUs, ASICs, FPGAs, multi-core processors, processing cores, etc.). The processing device 411 can be used by the primary storage node 410 to perform various operations, actions, functions, etc., as discussed in more detail below. The processing device 411 can be referred to as a primary processing device. The primary storage node 410 can also include one or more non-volatile memory modules 413. The non-volatile memory module 413 can be a device that stores data so that the data in the device remains unchanged even when power is no longer supplied to the device. Examples of the non-volatile memory module 413 can include, but are not limited to, an SSD, an NVME drive, a flash drive, etc.
[0257] In one embodiment, the secondary storage node 420 includes a processing device 421 (e.g., one or more CPUs, ASICs, FPGAs, multi-core processors, processing cores, etc.). The processing device 421 can be used by the secondary storage node 420 to perform various operations, actions, functions, etc., as discussed in more detail below. The processing device 421 can be referred to as a secondary processing device. The secondary storage node 420 can also include one or more non-volatile memory modules 423 (e.g., SSDs, NVME drives, flash drives, etc.).
[0258] In some embodiments, the nonvolatile memory modules 413 and / or 423 may be removable from the primary storage node 410 and / or the secondary storage node 420. For example, the nonvolatile memory modules 413 and / or 423 may be removed, relocated, and / or replaced with other nonvolatile memory modules. In another example, the nonvolatile memory modules 413 and / or 423 may be movable from one storage node to another (e.g., from the primary storage node 410 to the secondary storage node 420, from the secondary storage node 420 to the primary storage node 410, from the first primary storage node 410 to the second primary storage node 410, from the first secondary storage node 420 to the second secondary storage node 420, etc.). The nonvolatile memory modules 413 and / or 423 may also be removable while the primary storage node 410 and / or the secondary storage node 420 are operating. For example, the nonvolatile memory modules 413 and / or 423 may be hot-swappable. The non-volatile memory modules 413 and / or 423 may also be heterogeneous (e.g., non-uniform). For example, the non-volatile memory modules 413 and / or 423 may have different brands, models, manufacturers, capacities, and memory types (e.g., SLC, MLC, TLC, QLC, PLC, etc.).
[0259] In one embodiment, when the non-volatile memory modules 413 and / or 423 are removed, relocated, and / or replaced, the data stored on the non-volatile memory modules 413 and / or 423 can be reconsolidated, reused, and / or reaccessed. For example, if the non-volatile memory module 413 is moved from a first primary storage node 410 to another primary storage node 410, the data on the moved non-volatile memory module 413 can be reconsolidated into the storage system 400.
[0260] In one embodiment, non-volatile memory module 423 (e.g., some or all of non-volatile memory module 423) may have lower performance when compared to non-volatile memory module 413. For example, non-volatile memory module 423 may have a higher / longer access time than non-volatile memory module 413. In another example, non-volatile memory module 423 may have less bandwidth / throughput (e.g., read throughput and / or write throughput) when compared to non-volatile memory module 423.
[0261] In some embodiments, the primary storage node 410 may operate as a controller / manager of the storage system. For example, the primary storage node 410 may determine which storage operations should be performed and may instruct one or more secondary storage nodes to perform the storage operations (e.g., a rebuild / reconstruction operation for rebuilding data, a compression operation for compressing data, a garbage collection operation for freeing data blocks, etc.). In another example, the primary storage node 410 may delegate / offload storage operations to the secondary storage node 420. For example, the authorization 517 may manage / control a portion of the data and / or a subset of the non-volatile memory module 423 of the secondary storage node 420. The authorization 517 may offload / delegate storage operations related to or associated with the portion of the data and / or the subset of the non-volatile memory module 423 to the secondary storage node 420.
[0262] In one embodiment, when writing memory to the non-volatile memory modules 413 and / or 423, one or more of the primary storage node 410 and the secondary storage node 420 may use different write paths. A write path may refer to components, memories, circuits that data may pass through when storing / accessing data. For example, the primary storage node 410 may receive data from a client / computing device and may first store the data in NVRAM and then write the data to the non-volatile memory module 413. One example write path may be writing data directly from NVRAM to MLC flash memory (e.g., TLC memory, QLC memory, five-level cell (PLC) memory). Another example write path may be writing data from NVRAM to SLC flash memory and then to MLC flash memory. Another example write path may include writing data directly to MLC flash memory (e.g., bypassing NVRAM).
[0263] In one embodiment, the primary storage node 410 may determine that the programming mode of one or more portions of the nonvolatile memory modules 413 and / or 423 should be changed. For example, a portion of the nonvolatile memory module 413 may use the MLC programming mode (e.g., the nonvolatile memory module 413 may include MLC cells operating in the MLC mode). The primary storage node 410 may determine that the portion of the nonvolatile memory module 413 uses a different programming mode (e.g., the MLC cells should operate in the SLC mode). The primary storage node 410 may change the programming mode of different portions of the nonvolatile memory modules 413 and / or 423 based on various parameters. For example, the primary storage node 410 may change different portions of the nonvolatile memory modules 413 and / or 423 to the SLC mode and back to the QLC mode based on program / erase cycles of the cells in the nonvolatile memory modules 413 and / or 423.
[0264] In one embodiment, processing device 421 may have less computing power (e.g., less processing power) than processing device 411. For example, processing device 421 may have a lower clock frequency than processing device 411. In another example, processing device 421 may have fewer processing cores than processing device 411. In yet another example, processing device 421 may have less memory (e.g., less cache, fewer memory registers, etc.) than processing device 411. In another example, processing device 421 may be able to perform fewer operations / computations during a period of time when compared to processing device 411.
[0265] In some embodiments, processing device 421 may be different from processing device 411. For example, processing device 421 may be a processor of a different brand, model, etc. than processing device 411. In other embodiments, processing device 421 may be the same processor (e.g., the same brand, model, etc.) as processing device 411. However, processing device 421 may operate at a lower clock speed and / or may use less energy / power when compared to processing device 411.
[0266] In some embodiments, the primary storage node 410 and / or the secondary storage node 420 may include other components, devices, circuits, etc. to perform various other actions, operations, functions, tasks, etc. For example, the primary storage node 410 and / or the secondary storage node 420 may include a graphics processing unit (GPU), add-on boards (e.g., circuit boards, printed circuit boards, etc., which may be used to add additional non-volatile memory modules to the storage node), peripheral components, etc. These other components may be hot-swappable, modular, configurable, etc., and may be used to add capabilities to the storage node.
[0267] In one embodiment, the primary storage node 410 may include one or more authorizations ( Figure 4 Not described in the examples). Authorizations may control how and where data is stored in the non-volatile memory modules 413 and 423 of the storage system 400. For example, authorizations may determine how segments, erase blocks, allocation units, etc. are created and used. In another example, authorizations may determine how stripes (e.g., RAID stripes), parity data, etc. are created. Each primary storage node 410 may include any number of authorizations. For example, a primary storage node 410 may include one authorization, ten authorizations, or any appropriate number of authorizations. Each authorization may be responsible for managing a group of blocks / segments (e.g., a series of blocks / segments), one or more non-volatile memory modules 413 and / or 423. Authorizations may perform the above in conjunction with Figures 2A to 2G The function, operation, action, task, etc. discussed.
[0268] In one embodiment, if a second primary storage node 410 fails, the first primary storage node 410 may be able to continue or take over the operations, functions, actions, tasks, etc. of the second primary storage node 410. For example, if the second primary storage node 410 crashes, resets / reboots, or is otherwise inoperable, the first primary storage node 410 may be able to continue or take over the operations, functions, actions, tasks, etc. of the second primary storage node 410. Additionally, the operations, functions, actions, tasks, etc. of the second primary storage node 410 may be distributed (e.g., assigned to, spread to) multiple other primary storage nodes 410. For example, two, ten, or some other suitable number of primary storage nodes 410 may each take over a portion of the operations, functions, actions, tasks, etc. of the second primary storage node 410.
[0269] In one embodiment, authorizations may be transferred, moved, reallocated, reassigned, etc. from one primary storage node 410 to another primary storage node 410. For example, authorized data structures, metadata, procedures, services, etc. may be replicated from one primary storage node to another primary storage node. In another example, authorized data structures, metadata, procedures, services, etc. may be replicated from one primary storage node 410 to a secondary storage node 420 (e.g., when the secondary storage node 420 is converted to a primary storage node, as discussed above).
[0270] In one embodiment, the secondary storage node 420 may initially lack authorization. For example, when the secondary storage node 420 initiates operation (e.g., initially boots up or starts), there may be no authorization on the secondary storage node 420 (e.g., the secondary storage node 420 may not have any authorization). The secondary storage node 420 may include one or more authorizations after it transitions to the primary storage node 410.
[0271] In one embodiment, a secondary storage node 420 can be converted to a primary storage node (e.g., can operate as a primary storage node). For example, if a primary storage node 410 fails, one or more secondary storage nodes 420 can be used to perform the operations, functions, actions, tasks, etc. of the failed primary storage node 410. The conversion from a secondary storage node to a primary storage node can be temporary. For example, a secondary primary node 420 can operate as a primary storage node for a period of time until the failed primary storage node 410 is restored, or until a new primary storage node 410 is added to the storage system.
[0272] In one embodiment, one or more device hosts ( Figure 4 The secondary storage node 420 may include one or more device hosts (not shown) that can perform one or more storage operations. For example, the secondary storage node 420 may include one or more device hosts. The device host may be hardware, software (e.g., a service, process, thread, daemon, driver, etc.), firmware, or a combination thereof that can manage and control one or more non-volatile memory modules 423, etc. The device host will be described in more detail below.
[0273] In various embodiments, the storage system 400 allows for dynamic configuration of primary storage nodes 410 and secondary storage nodes 420. For example, primary storage nodes 410 and / or secondary processing storage 420 can be added and / or removed from the storage cluster. In addition, the non-volatile memory modules can also be removable and changeable, thereby further improving the configurability of the storage system. Because the primary storage node 410 can control / manage the operations of the storage system 400 (e.g., storage operations), more secondary storage nodes 420 can be used in the storage system. The secondary storage nodes 420 can include non-volatile memory modules with lower processing power and / or lower performance. This can allow the storage system 400 to operate more efficiently (e.g., using less electricity and with reduced costs) when providing access to data. The secondary storage nodes 420 can also allow the storage capacity of the storage system to be increased more efficiently (e.g., using less electricity and with reduced costs).
[0274] Figure 5 is a block diagram illustrating an example primary storage node 410 and an example secondary storage node 420 according to some embodiments of the present disclosure. As discussed above, the primary storage node 410 and / or the secondary storage node 420 can be included in one or more storage clusters (e.g., the primary storage node 410 and the secondary storage node 420 can be located in the same storage cluster or in different storage clusters). The primary storage node 410 and the secondary storage node 420 can be communicatively coupled to each other (e.g., via a network coupling or directly coupled via an interconnect / wire).
[0275] As discussed above, the primary storage node 410 includes one or more processing devices 411 (e.g., one or more CPUs, ASICs, FPGAs, multi-core processors, processing cores, etc.) that can be used to perform various operations, actions, functions, etc. The primary storage node 410 may also include one or more non-volatile memory modules 413 (e.g., SSDs, NVME drives, flash drives, etc.). The primary storage node 410 may further include a device host 515 and an authorization 517.
[0276] In one embodiment, the device host 515 may be hardware, software (e.g., a service, process, thread, daemon, driver, etc.), firmware, or a combination thereof that can manage and control one or more non-volatile memory modules 413. The device host 515 may be executed by the processing device 411 to manage the one or more non-volatile memory modules 413. The device host 515 may communicate with the non-volatile memory modules 413 to read data from the non-volatile memory modules and / or write data to the non-volatile memory modules. In another example, the device host 515 may perform storage operations (e.g., compressing data, reconstructing / rebuilding data, etc.) on the data stored in the non-volatile memory modules 413. The primary storage node 410 may include any number of device hosts 515. For example, the primary storage node 410 may include one device host 515 for each non-volatile memory module 413, or may include fewer device hosts 515 than non-volatile memory modules 413.
[0277] In one embodiment, the authorization 517 can manage and / or control data stored on the non-volatile memory modules 413 and / or 423 of the primary storage nodes 410 and / or 420. Each authorization 168 can be assigned to one or more non-volatile memory modules 413 and / or 423. Each authorization can control / manage a range of inode numbers, segment numbers, block numbers, or other data identifiers assigned to the data. The authorization can determine where the data is written (e.g., which blocks of the non-volatile memory modules are used, which storage nodes or which non-volatile memory modules have which portions of the data), how the data is written (e.g., how to create RAID stripes, which erasure coding scheme is used, whether the data should be compressed, etc.), and whether other operations should be performed on the data and / or non-volatile memory modules 413 and / or 423 (e.g., whether the data should be moved, compressed, etc., whether garbage collection should be performed, etc.).
[0278] The secondary storage node 420 includes one or more processing devices 421 (e.g., one or more CPUs, ASICs, FPGAs, multi-core processors, processing cores, etc.) that can be used to perform various operations, actions, functions, etc. As discussed above, the processing device 421 may have lower processing / computing capabilities than the processing device 411. The secondary storage node 420 may also include one or more non-volatile memory modules 423 (e.g., SSDs, NVME drives, flash drives, etc.). The primary storage node 420 may further include a device host 525 for managing / controlling the non-volatile memory modules 423. The device host 525 may be hardware, software (e.g., services, processes, threads, daemons, drivers, etc.), firmware, or a combination thereof that can manage and control one or more non-volatile memory modules 423, etc. For example, the device host 525 may perform operations similar to those of the device host 515.
[0279] Figure 6 7 is a block diagram illustrating an example storage system 600 according to some embodiments of the present disclosure. The storage system 600 includes a computing device 630 (e.g., a client device, a host computer, etc.), a primary storage node 410, and a secondary storage node 705. The computing device 630 can be communicatively coupled to the primary storage node 410 via a network 705. As discussed above, the primary storage node 410 and / or the secondary storage node 420 can be included in one or more storage clusters. The primary storage node 410 and the secondary storage node 420 can be communicatively coupled to each other (e.g., via a network 707 or directly via interconnects / wires).
[0280] As discussed above, the primary storage node 410 includes one or more processing devices (e.g., one or more CPUs, ASICs, FPGAs, multi-core processors, processing cores, etc.) that can be used to perform various operations, actions, functions, etc. The primary storage node 410 may also include one or more non-volatile memory modules (e.g., SSDs, NVME drives, flash drives, etc.), one or more device hosts, and / or one or more authorizations. As discussed above, the device host can be hardware, software (e.g., services, processes, threads, daemons, drivers, etc.), firmware, or a combination thereof that can manage and control one or more non-volatile memory modules, etc. As also discussed above, the authorization 517 can manage and / or control data (e.g., a series or group of blocks / segments) stored on the non-volatile memory modules of the primary storage nodes 410 and / or 420.
[0281] like Figure 6, computing device 630 and primary storage node 410 may communicate via network 705. Thus, in one embodiment, primary storage node 410 may be the primary / primary point of contact for communications (e.g., messages, requests, etc.) originating from outside primary storage node 410 and / or secondary storage node 420. For example, primary storage node 410 may be the primary storage node for receiving communications from computing device 630. In another example, a request to write, read, modify data (stored on primary storage node 410 and / or secondary storage node 420), etc., may first be received by one of primary storage nodes 410. The request may then be forwarded to the appropriate storage node (e.g., primary storage node 410 and / or secondary storage node 420 that has the requested data).
[0282] The intermediate secondary storage nodes 420 may be communicatively coupled to the primary storage node 410 via a network 707. Thus, the intermediate secondary storage nodes 420 may be accessible via the network 707. The primary storage node 410 (e.g., one or more authorities) may receive requests to access data (e.g., to read, write, modify data, etc.) and may send instructions to another primary storage node 410 and / or one or more of the intermediate secondary storage nodes 420 via the network 707 to perform various storage operations. When accessing data on the intermediate secondary storage nodes 420, the primary storage node 410 may be aware of and / or may account for some delay / latency because the request to access data (or perform other storage operations) is first received by the primary storage node 410 and then forwarded to the intermediate secondary storage node 420. For example, one or more authorities may have data, information, metadata, tables, etc. that may indicate the delay / latency of accessing non-volatile memory modules in different storage nodes.
[0283] The left secondary storage node 420 may be directly coupled to the left primary storage node 410. For example, the left secondary storage node 420 may be directly coupled to the left primary storage node 410 via an interconnect, a bus, a wire, a bridge, etc. Therefore, the left secondary storage node 420 may not be accessible via the network 707. In order for messages and / or data to reach the left secondary storage node 420, the messages and / or data should be forwarded by the left primary storage node 410 to the left secondary storage node 420. When accessing data on the left secondary storage node 420, the primary storage node 410 may be aware of and / or may account for some delay / latency because a request to access data (or perform other storage operations) should first be forwarded to the left primary storage node 410 and then to the left secondary storage node 420. For example, one or more grants may have data, information, metadata, tables, etc. that may indicate the delay / latency of accessing non-volatile memory modules in different storage nodes.
[0284] Figure 77 is a block diagram illustrating an example storage system 700 according to some embodiments of the present disclosure. The storage system 700 includes a primary storage node 410 and a secondary storage node 705. As discussed above, the primary storage node 410 and the secondary storage node 420 can be communicatively coupled to each other (e.g., via a network 707 or directly via interconnects / wires).
[0285] As discussed above, the primary storage node 410 includes one or more processing devices (e.g., one or more CPUs, ASICs, FPGAs, multi-core processors, processing cores, etc.) that can be used to perform various operations, actions, functions, etc. The primary storage node 410 may also include one or more non-volatile memory modules (e.g., SSDs, NVME drives, flash drives, etc.), one or more device hosts, and / or one or more authorizations. As discussed above, the device host can be hardware, software (e.g., services, processes, threads, daemons, drivers, etc.), firmware, or a combination thereof that can manage and control one or more non-volatile memory modules, etc. As also discussed above, the authorization 517 can manage and / or control data (e.g., a series or group of blocks / segments) stored on the non-volatile memory modules of the primary storage nodes 410 and / or 420.
[0286] In one embodiment, the top primary storage node 410 may fail (e.g., may be reset, crash, or otherwise become inoperable). Therefore, operations, actions, functions, tasks, etc. performed by the top primary storage node 410 may be transferred to the top three secondary storage nodes 420 (illustrated by the dotted boxes). For example, the authorization of the primary storage node 410 may be distributed to the top three secondary storage nodes 420. The top three secondary storage nodes 420 may operate as primary storage nodes (e.g., may be converted to primary storage nodes). The top three secondary storage nodes 420 may operate as primary storage nodes until instructed otherwise. For example, the bottom primary storage node 410 may determine that the top primary storage node 410 is now available and may instruct the top three secondary storage nodes 420 to revert back to operating as secondary storage nodes. The top three secondary storage nodes 420 may also operate as primary storage nodes for a period of time (e.g., minutes, hours, days, weeks, etc.).
[0287] Figure 8 is a flow chart illustrating a method 800 for performing a storage operation according to some embodiments of the present disclosure. The method 800 may be performed by processing logic, which may include hardware (e.g., circuitry, dedicated logic, programmable logic, microcode, etc.), software (e.g., instructions running on a processing device to perform hardware simulation), firmware, or a combination thereof. In one embodiment, the method 800 may be performed by, for example, Figures 4 to 7 The main storage nodes, processing devices and / or authorized execution described in.
[0288] The method begins at block 805, where the method 800 may optionally provide one or more primary storage nodes. At block 810, the method 800 may optionally provide one or more secondary storage nodes. At block 815, the method 800 may include receiving one or more requests to access data in the storage system. For example, the requests may be received from a client / computing device external to the storage system. At block 820, the method 800 may include transmitting a set of instructions to the one or more secondary storage nodes. For example, the method 800 may determine which secondary storage nodes should perform the set of storage operations and transmit the set of instructions to the identified secondary storage nodes.
[0289] Scalable storage system
[0290] As discussed above, a storage system (which may also be referred to as a data storage system) may be a combination of hardware, software, and firmware that implements the capability to store, retrieve, and / or access data on behalf of other devices, services, and / or applications. For example, a storage system may provide and / or implement data storage and retrieval capabilities / functionality for server systems / computers, applications, databases, and the like. A storage system may have, implement, or provide a set of features and performance management capabilities, and various mechanisms that may improve reliability by enabling continued operation and reducing the likelihood of data loss in the event of hardware, software, and / or other failures.
[0291] A storage system may include various storage system elements. Storage system elements may be physical and / or software components of the storage system. For example, storage system elements may include circuit boards, power supplies, fans, interconnects, etc. In another example, storage system elements may include software modules, operating systems, APIs, etc. A storage system may also include persistent storage devices or memories that can retain their contents in the event of a software crash, reboot, power outage, storage system element failure, and / or other errors or failures.
[0292] A storage system may include a controller (which may be referred to as a storage system controller) that implements, coordinates, executes, manages, and / or advertises the functions, capabilities, and / or operations of the storage system. For example, the storage system controller may manage the allocation and placement of data within the storage system, provide error correction and deduplication capabilities, etc. Various examples of storage system controllers are illustrated in the figures and discussed herein.
[0293] The storage system also includes storage devices for storing data in the storage system. Storage devices may also be referred to as non-volatile memory devices, non-volatile memory modules, drives, SSDs, flash drives, NVME drives, etc. The storage devices are accessible to the rest of the storage system and present that persistent storage to the storage system regardless of the hardware and software / firmware components used.
[0294] In one embodiment, as discussed above, a storage device may be managed by devices and / or software external to the storage device. The storage device may be managed by one or more storage system controllers of the storage system. For example, block / page allocation, garbage collection, etc. may be controlled and / or managed by the storage system controller. In another example, managing or optimizing other aspects / functionality of a storage device managed by devices and / or software external to the storage device may be referred to as managed storage.
[0295] Each storage device may also include a controller (e.g., one or more processing devices located within the storage device) that can control or manage the storage device itself. For example, the storage device may include a CPU, an ASIC, or an FPGA that provides or implements capabilities, functions, etc. for internally managing storage media (e.g., persistent storage) and for interacting with the storage device (e.g., for communicating with the storage device, transmitting instructions / messages to the storage device, etc.).
[0296] As discussed above, in Figure 1D In the present invention, a storage system may include multiple storage system controllers that can operate in conjunction with each other to manage the storage system. For example, the storage system may include two storage system controllers (or some other suitable number of storage system controllers) that are communicatively coupled to storage devices in the storage system. For example, the storage device may include two interfaces (e.g., to ports, connections, etc.), and each interface may be coupled to one or two storage system controllers. In another example, the storage device and the two storage system controllers may be coupled to a bus / network and may communicate with each other via the bus / network. Having multiple (e.g., two or more) storage system controllers in a storage system can provide various advantages. For example, during an outage of one storage system controller (e.g., one storage system controller fails or is being upgraded), the other storage system controller (or other storage system controllers) can continue to provide access to the storage devices in the storage system.
[0297] Paired storage system controllers can operate in various ways to manage storage devices in a storage system. For example, paired storage system controllers can divide storage devices in a storage system between the storage system controllers. In another example, one storage system controller can manage and / or perform most (or all) services, or can perform all I / O-related services (e.g., servicing read / write requests) until the other storage system controller has reason to take over (e.g., due to a failure of the first storage system controller). In yet other examples, paired storage system controllers can share access to storage devices, coordinate their operations, actions, etc. in some manner, and avoid inconsistencies in data stored on the storage devices.
[0298] Figure 9is a block diagram illustrating an example storage system 900 according to one or more embodiments of the present disclosure. Storage system 900 may be a scale-out storage system (e.g., a scalable storage system). Storage system 900 includes a combined controller 910 and a network 920. Network 920 may include buses, wires, fiber optic cables, other networks, and / or other connections. Network 920 may carry communications (e.g., data, messages, packets, frames, etc.) between combined controllers 910.
[0299] As discussed above, the storage system 900 includes a combined controller 910 (which may be referred to as an integrated controller, a unified controller, etc.). The combined controller 910 may be a storage node, a blade, a removable circuit board, a line card, etc., in which the functionality of a storage system controller and a storage device are combined into a unified / single element or component. For example, the combined controller 910 may perform / implement both storage system controller capabilities and storage device functionality, as well as network connectivity for interaction with other combined controllers in the storage system.
[0300] In one embodiment, the combination controller 910 may be able to access the storage devices 915 and / or flash memory 917 of other combination controllers 910 in the storage system. For example, a first combination controller 910 may be able to read and access the storage devices 915 and / or flash memory 917 of a second combination controller via a network 920, and so on. The first combination controller 910 may directly access the storage devices 915 and / or flash memory 917 of the second combination controller 910, or may communicate with the processing device 911 of the second combination controller 910 to access the storage devices 915 and / or flash memory 917 of the second combination controller 910.
[0301] In one embodiment, the storage system 900 may include storage system elements that do not have integrated storage device controller functionality / capabilities but do include storage system controller functionality / capabilities. For example, the storage system 900 may include additional cards, blades, storage nodes, etc. that do not have storage devices. This may allow the storage system 900 to expand or extend processing power / capabilities, as discussed in more detail below.
[0302] In another embodiment, the storage system 900 may include storage system elements that include storage device controller functionality / capabilities but do not include storage system controller functionality / capabilities. For example, the storage system 900 may include additional cards, blades, storage nodes, etc. that include storage devices but do not include a processing device for performing storage system controller functions. This may allow the storage system 900 to expand or scale storage capacity, as discussed in more detail below.
[0303] In one embodiment, storage system 900 may be an expandable storage system or a scale-out storage system. In a scale-out storage system, performance and / or storage can be expanded by adding storage system controllers. Storage system controllers may cooperate or coordinate with each other to divide tasks, functions, operations, etc., to perform various storage system services and / or storage system capabilities. For example, a storage system may contain hundreds, thousands, or any suitable number of storage system controllers coordinating with each other.
[0304] In a scale-out storage system (e.g., a scalable storage system), communicatively coupling all storage devices to all storage system controllers in the storage system can be more difficult (when compared to a storage system that uses paired storage system controllers). The combined controller 910 can allow the storage system to become and / or operate as a scale-out storage system. The combined controllers 910 can include built-in storage, such as storage device 915, and can communicate with each other over a network to manage storage system capabilities and handle failures, upgrades, and any other conditions.
[0305] In one embodiment, the combined controller 910 may include a processing device 911 and a memory 912 (e.g., RAM, NVRAM, cache, etc.) for performing storage system controller functions. The combined controller 910 may also include a processing device 913 and a memory 914 (e.g., RAM, NVRAM, cache, etc.) for performing storage device controller functions. In another embodiment, the combined controller may include a processing device that performs both storage system controller functions and storage device controller functions. The combined controller 910 may also include an energy storage device 916 (e.g., a supercapacitor, a battery, or some other suitable power source), the processing device 913, the memory 914, the processing device 911, and the memory 912.
[0306] like Figure 9 As shown in FIG. 1 , the combined controller 910 includes one or more storage devices 915. The storage device 915 may be a logical grouping of an energy storage device 916, a processing device 913, a memory 914 (e.g., RAM), and a flash memory 917 (e.g., a flash chip, a flash die, etc.) of the combined controller 910. In other embodiments, as shown in FIG. Figure 10 As described in the specification, the storage device may refer to a component or device including a processing device, a memory, and a flash memory. The storage device 915 may also be a hot-swappable or hot-pluggable device.
[0307] In one embodiment, the storage device 915 may be accessible to multiple combination controllers 910. For example, a first combination controller 910 may be able to access (e.g., read data from, write data to) the storage device 915 of a second combination controller 910. The first combination controller 910 may access the storage device 915 of the second combination controller 910 directly, or may access the storage device 915 of the second combination controller 910 via the storage system controller (e.g., processing device 911) of the second combination controller 910. In one embodiment, the first combination controller 910 may be able to address a read or write request to a specific flash memory cell, page, block, etc. of the storage device 915 of the second combination controller 910.
[0308] In one embodiment, one or more of the storage devices 915 may be managed flash storage devices. Managed flash storage devices (which may also be referred to as directly managed flash storage devices, directly managed storage devices, managed storage devices, etc.) may provide functions, operations, commands, APIs, or some other appropriate mechanism to an external device (e.g., the processing device 911 (of the combined controller 910)) to control, manage, and / or interact with the flash memory of the managed flash storage device. This may allow the storage device controller to perform fewer operations (e.g., handling queues, failovers, internal error correction, encryption, voltage level adjustment of rows / pages of flash memory, etc.). Because the storage devices 915 can be directly managed by the combined controller 910 (e.g., by the processing device 911), this allows the storage system to optimize, manage, and / or improve various aspects, characteristics, etc. of the flash memory 917, improving the performance, reliability, and / or lifespan of the flash memory, as discussed in more detail below.
[0309] In one embodiment, the storage device 915 may provide or expose functions, APIs, commands, etc. that allow external devices (e.g., the combo controller 910) to better optimize and manage flash memory. For example, the storage device 915 may use high-bit-per-cell flash memory (e.g., TLC flash memory / memory, QLC flash memory / memory, five-level cell (PLC) flash memory / memory, etc.) for storing data (e.g., as data storage or mass storage). If the characteristics of the flash memory are not properly accounted for and managed by the storage system 900 and / or the combo controller 910, the performance of such flash memory may degrade rapidly and / or have reduced performance.
[0310] Storage device 915 may also include memory for internal use by storage device 915 (separate from the flash memory of storage device 915). Figure 1C, the storage device 915 may include RAM 121 (e.g., DRAM). The memory of the storage device 915 may also be used directly as temporary storage or transactional memory for the processing device 911 (which may perform the function of a storage system controller). For example, the energy storage 916 may be used to make the memory of the storage device 915 persistent (or non-volatile). In the event of a power failure, the processing device 913 (which may perform the function of a storage device controller) may write or store the contents of the memory 914 to the memory of the storage device 915.
[0311] Storage device 915 may also use memory for various other purposes and / or functions. For example, memory may be used to store tables (e.g., mapping tables) and / or other metadata associated with data stored in the flash memory of storage device 915. Storage device 915 may emulate sector-based read / write operations on a disk by mapping numbered sectors (e.g., 512-byte or 4096-byte blocks) to blocks / pages in the flash memory. Data stored in the flash memory may be organized based on certain characteristics of the flash memory that facilitate writing / rewriting data to the flash memory and maintaining the lifespan / health of the flash memory over time. Characteristics of the flash memory may also be affected by internal garbage collection and wear leveling operations. In one embodiment, storage device 915 may be designed, configured, etc. to operate within storage system 900 and managed by processing device 911 (e.g., by a storage system controller). This may allow storage device 915 to use less memory (e.g., RAM, DRAM, etc.) for tables and / or other metadata, as discussed in more detail below.
[0312] In one embodiment, storage device 915 may include some other form of persistent memory, such as 3D X-Point flash memory or low-bit-per-cell flash memory (e.g., SLC or MLC), which can be written much faster than high-bit-per-cell flash memory (e.g., TLC, QLC, PLC, etc.). High-bit-per-cell flash memory can be programmed as low-bit-per-cell flash memory. For example, QLC flash memory of storage device 915 can be programmed in SLC mode (e.g., can be programmed / used as SLC memory). By using QLC flash memory as SLC flash memory, storage device 915 can reduce or even eliminate the use of memory (e.g., DRAM) used to store tables and / or other metadata. For example, using a portion (e.g., one percent) of the QLC flash memory as SLC memory can allow more tables and / or metadata to be stored in the SLC memory than in the memory of storage device 915. In some embodiments, storage device 915 may have a combination of memory (e.g., DRAM) and QLC flash memory used as SLC flash memory (e.g., QLC flash memory programmed in SLC mode).
[0313] In one embodiment, storage device 915 may not include memory (e.g., DRAM) and may not include other forms of persistent memory (e.g., QLC flash memory programmed in SLC mode) for storing tables / metadata. Storage system 900 (e.g., processing device 911) may manage reads and / or writes to storage device 915 such that tables / metadata do not need to be stored in storage device 915. Instead, storage system 900 may optimize, track, and / or manage the use of flash memory on storage device 915.
[0314] In one embodiment, when managing the flash memory of the storage device 915 (e.g., when managing a managed flash memory device), the combination controller 910 (e.g., the processing device 911) may optimize various parameters and / or characteristics of the flash memory. The storage device 915 (e.g., the managed flash memory device) may provide information, interfaces, functions, etc. to the combination controller 910 to allow the combination controller 910 to optimize various parameters / characteristics of the flash memory.
[0315] In one embodiment, the first parameter may be the lifetime / use period of the flash memory (e.g., how long the flash memory can last before exhausting program / erase cycles, how many times a block has been programmed / erased, etc.). As discussed above, conventional storage devices (e.g., SSDs) emulate sector-based reads / writes by mapping numbered sectors to blocks / pages in the flash memory. This can lead to fragmentation because erase blocks typically use garbage collection to free up entire erase blocks for reprogramming. Fragmentation can increase the rate at which garbage collection is performed to free up erase blocks, which increases the frequency with which erase blocks are erased and reprogrammed. By allowing the use of I / O models other than sectors (e.g., using regions or exposing the flash memory to the combo controller 910 as programmable erase blocks), fragmentation within the storage device 915 can be reduced and / or eliminated. Additionally, some other I / O models can allow the combo controller to manage fragmentation within the storage device 915.
[0316] In another embodiment, directly managing and / or using memory (e.g., DRAM) or other fast persistent memory (e.g., SLC memory) may be another parameter / characteristic. For example, the storage system 900 may use temporary data when providing storage services to clients of the storage system 900. If a failure occurs, this temporary data may need to be restored. The combined controller 910 may mark the temporary data as temporary, and the storage device 915 may store the temporary data in memory (which may be backed by the energy storage 916), which frees up more flash memory for use by the storage system and reduces wear and fragmentation on the flash memory. The combined controller 910 may also directly handle the storage of transient data and control when the transient data is transferred to the flash memory of the storage device 915. In addition, the storage system 900 may also use other tables / metadata that are frequently updated but do not use much storage space. It may be useful to allow the combined controller 910 to determine whether this data is written in a mode with fewer bits per cell (e.g., SLC mode) or a mode with more bits per cell (e.g., QLC mode).
[0317] In other embodiments, the amount of memory used by the storage device 915 (e.g., RAM, DRAM, etc.) can be another parameter. By reducing data fragmentation and / or by using the combination controller 915 to manage the flash memory of the storage device 915, the amount of memory used and / or required by the storage device 915 can be reduced. For example, a traditional storage device (e.g., an SSD) may require RAM to store a mapping table needed to locate sectors that are scattered / distributed across erase blocks. If the combination controller 910 manages where data is stored in the flash blocks, then the storage device 915 does not need to store the mapping table because the combination controller 910 can store or know the mapping table.
[0318] In one embodiment, the underreported capacity of storage device 915 may be another parameter / characteristic that storage system 900 may manage and / or optimize. As discussed above, storage system 900 (e.g., combined controller 910) may reduce the fragmentation of data within storage device 915. This may also reduce junk collection overhead, improve junk collection efficiency, and increase the lifespan of the flash memory. Traditional storage devices (e.g., SSDs) may report less storage capacity than the available capacity of the flash memory, leaving portions of the flash memory unusable due to erase blocks that cannot or have not been junk collected. By directly managing the flash memory of storage device 915, storage system 900 may allow for fewer underreports (e.g., more available storage capacity) of storage device 915.
[0319] In one embodiment, storage device 915 may allow a storage system (e.g., combination controller 910) to control, improve, and / or optimize the performance of storage device 915. For example, garbage collection of blocks in storage device 915 may be directly managed by combination controller 910 (e.g., by processing device 911) and / or may be performed by combination controller 910. In another example, combination controller 910 may manage which flash memory dies or planes are written to. In another example, combination controller 910 may determine which internal buses in storage device 915 are used sequentially or in parallel to write to which portions of the flash memory. Additionally, flash memory (e.g., flash chips) typically support methods for interrupting and restarting very high-latency operations (e.g., large writes and erases of erased blocks). This may allow other operations (e.g., reads) to be performed during the interruption of a high-latency operation. The storage system may use various rules, parameters, criteria, conditions, etc. to determine when an instruction should be interrupted. The group controller 910 may also provide hints / information to the storage device 915 indicating that if a current operation is taking too long, then another storage device 915 may be used to perform the read operation (e.g., data may be reconstructed by accessing an erroneously encoded portion from the other storage device 915). The storage device 915 may provide various other methods, techniques, mechanisms, etc. for the storage system 900 to improve the performance of the storage system 900 (e.g., increase data throughput).
[0320] In traditional storage devices (e.g., SSDs), a storage controller may perform various functions such as mapping and remapping, garbage collection, wear leveling, etc. This can hide the management of flash memory from other systems that may use traditional storage devices (e.g., laptops, desktop computers, etc.). In one embodiment, the storage system 900 (e.g., the combined controller 910) can offload many or most functions performed by the storage controller to the combined controller 910 (e.g., to the processing device 911). If most operations / functions are offloaded from the storage controller and the storage controller performs operation scheduling, NVME queue scheduling, internal queue scheduling, and bus management, the storage controller can be a simpler device (e.g., a less expensive device).
[0321] In one embodiment, storage device 915 may allow the storage system (e.g., group controller 910) to manage, improve, and / or optimize the reliability of storage device 915. This may be achieved by better integrating the behavior / operation of storage device 915 with the operation of group controller (e.g., processing device 911). For example, if a block / page is bad (e.g., due to aging, read disturb, manufacturing defects, etc.), storage device 915 may be unable to recover the data in the bad block / page. However, group controller 910 may be able to recover the data from the bad block / page using RAID parity, erasure lines, or other error correction mechanisms. Additionally, group controller 910 may take into account the aging and read disturb of the flash memory when performing garbage collection and / or wear leveling operations. By moving / offloading garbage collection and / or wear leveling operations to group controller 910, storage system 900 may be able to improve the reliability and / or lifespan of the flash memory in storage device 915.
[0322] As discussed above, the storage system 900 can provide, implement, and / or coordinate various storage system services for clients / hosts. For example, the combined controller 910 can operate in conjunction to implement a SCSI or NVME-based storage array, an NFS or SMB-based file server, an object server, a database service, a service for running scalable storage-related applications, or multiple or a combination of such services. In another example, the combined controller 910 can provide data management services such as erasure coding, RAID, etc. In another example, the combined controller 910 can provide data services such as snapshots, replication, cloning, etc. In another example, the combined controller 910 can provide management services such as creating, modifying, or removing logical elements (such as volumes, snapshots, replication relationships, file systems, object storage, application or host system relationships, etc.). In addition, the combined controller 910 can also perform various functions, services, and operations to manage and / or optimize the use of the flash memory 917, as discussed above (for example, optimizing or directly managing erase block erase / reprogram cycles, fragmentation, operating in different programming modes, etc.).
[0323] In one embodiment, the functions / services provided, executed, or implemented by the composite controller 910 may be divided into multiple layers. The functions of each layer may be distributed across multiple composite controllers 910. One layer of functions / services is a higher layer that may provide logical / virtual capabilities that allow volumes, file systems, databases, object stores, key-value database services, etc.
[0324] etc., for storing data, snapshots, cloning, and copying the data. This layer may be referred to as authorization. The second layer of services / functions may help ensure that data is recoverable in the event of a failure, such as by using RAID or erasure codes. The second layer may also provide waste collection. The third layer of services / functions may allow the storage system 900 to optimize the use of storage devices 915 (e.g., managed flash storage devices). For example, the third layer of functions / services may know the queuing model of a particular storage device, or whether or how much of a device can be written as SLC or QLC, or how to track the loss of a particular erase block. In some cases, the third layer may be responsible for some waste collection, or may be responsible for some first attempts at error recovery, as well as other potential services. In one embodiment, the second and third layers may be combined into a single layer, which is referred to as the device host or device host layer.
[0325] Figure 10 1 is a block diagram illustrating an example storage system 1000 according to one or more embodiments of the present disclosure. Storage system 1000 may be a scale-out storage system (e.g., an expandable storage system). Storage system 1000 includes storage nodes 1010, a network 1020, and a switch 1030. Network 1020 may include a bus, wires, fiber optic cables, other networks, and / or other connections. Network 1020 may carry communications (e.g., data, messages, packets, frames, etc.) between storage nodes 1010 via interfaces 1018 (e.g., storage node interfaces) of storage nodes 1010. Interface 1018 may be a communication interface (e.g., a PCIe interface, an NVME interface, an NVME fiber interface, a bus interface, a network interface, a fiber interface, etc.).
[0326] like Figure 10 As shown in FIG. 1 , storage system 1000 includes storage nodes 1010. Storage nodes 1010 may be storage nodes, blades, removable circuit boards, line cards, and the like that perform the functions of a storage system controller. For example, processing device 1011 may perform the functions of a storage system controller. Each storage node 1010 may include multiple storage devices 1015. Storage devices 1015 are coupled to storage nodes 1010 via interfaces 1019 (e.g., communication interfaces, PCIe interfaces, NVMe interfaces, NVMe fiber interfaces, bus interfaces, network interfaces, etc.). In one embodiment, storage system 1000 may be a scalable storage system or a scale-out storage system.
[0327] In one embodiment, the switch 1030 can be coupled to other networks, other computing devices (e.g., server computers, client devices, etc.), and / or other storage systems. For example, the switch 1030 can allow client devices to store data in the storage system 1000 and access the data. In another example, the switch 1030 can be coupled to another network that is coupled to a second storage system ( Figure 10 This allows the two storage systems to communicate with each other.
[0328] Each storage device 1015 includes energy storage 1016, processing device 1013, and memory 1014. Processing device 1013 can perform storage device controller functions / operations. Storage device 1015 can also be hot-swappable or hot-pluggable devices. For example, storage device 1015 can be moved from one interface 1019 of storage node 1010 to another interface 1019. In another example, storage device 1015 can be moved from one storage node 1010 to another storage node 1010. Energy storage 1016, processing device 1013, and memory 1014 can perform the same functions, operations, etc., and can be used with the same storage devices as described above. Figure 9 For example, the memory 1014 may be used to temporarily store data / writes, store mapping tables, etc.
[0329] Storage nodes 1010 may be able to access storage devices 1015 and / or flash memory 1017 of other storage nodes 1010 in the storage system. In one embodiment, storage system 1000 may include storage system elements (e.g., storage nodes) that do not include storage devices 1015. In another embodiment, storage system 1000 may include storage system elements (e.g., storage nodes) that include storage devices 1015 but do not include storage system controller functionality / capability.
[0330] In a scale-out storage system, performance and / or storage can be expanded by adding storage system controllers. Storage system controllers can cooperate or coordinate with each other to divide tasks, functions, operations, etc., to perform various storage system services and / or storage system capabilities. For example, a storage system may have hundreds, thousands, or any suitable number of storage system controllers coordinating with each other.
[0331] In one embodiment, one or more of the storage devices 1015 may be managed flash storage devices. A managed flash storage device may provide functions, operations, commands, APIs, or some other suitable mechanism for an external device (e.g., the processing device 1011 (of the storage node 1010)) to control, manage, and / or interact with the flash memory of the managed flash storage device. In one embodiment, the storage device 1015 may provide or expose functions, APIs, commands, etc. that allow an external device (e.g., the storage node 1010) to better optimize and manage flash memory.
[0332] In one embodiment, storage device 1015 may be designed, configured, etc., to operate within storage system 1000 and be managed by processing device 1011 (e.g., by a storage system controller). This may allow storage device 1015 to use less memory (e.g., RAM, DRAM, etc.) for tables and / or other metadata.
[0333] In one embodiment, storage device 1015 may include some other form of persistent memory, such as 3D X-Point flash memory or flash memory with fewer bits per cell (e.g., SLC or MLC), which can be written much faster than flash memory with more bits per cell (e.g., TLC, QLC, PLC, etc.). In some embodiments, storage device 1015 may have a combination of memory (e.g., DRAM) and QLC flash memory used as SLC flash memory (e.g., QLC flash memory programmed in SLC mode). In one embodiment, storage device 1015 may not include memory (e.g., DRAM) and may not include other forms of persistent memory for storing tables / metadata (e.g., QLC flash memory programmed in SLC mode).
[0334] In one embodiment, the storage device 1015 may be accessible to multiple storage nodes 1010. For example, a first storage node 1010 may be able to access (e.g., read data from, write data to) the storage device 1015 of a second storage node 1010. The first storage node 1010 may directly access the storage device 1015 of the second storage node 1010, or may access the storage device 1015 of the second storage node 1010 via a storage system controller (e.g., processing device 1011) of the second storage node 1010. In one embodiment, the first storage node 1010 may be able to address a read or write request to a specific flash memory cell, page, block, etc., of the storage device 1015 of the second storage node 1010.
[0335] In one embodiment, when managing the flash memory of the storage device 1015 (e.g., when managing a managed flash memory storage device), the storage node 1010 (e.g., the processing device 1011) may optimize various parameters and / or characteristics of the flash memory. The storage device 1015 (e.g., the managed flash memory storage device) may provide information, interfaces, functions, etc. to the storage node 1010 to allow the storage node 1010 to optimize those various parameters / characteristics of the flash memory, similar to the above description in conjunction with Figure 9 Parameters / features discussed.
[0336] In one embodiment, storage device 1015 may allow a storage system (e.g., storage node 1010) to control, improve, and / or optimize the performance of storage device 1015. For example, obsolete item collection for blocks in storage device 1015 may be directly managed by storage node 1010 (e.g., by processing device 1011) and / or may be performed by storage node 1010. In another example, a storage node may manage which flash memory dies or planes are written to. In another example, storage node 1010 may determine which internal buses within storage device 1015 are used to write to which portions of the flash memory, either sequentially or in parallel. Additionally, flash memory (e.g., flash chips) typically support methods for interrupting and restarting very high-latency operations (e.g., large writes and erases of erased blocks). This may allow other operations, such as reads, to be performed during the interruption of a high-latency operation. The storage system may use various rules, parameters, criteria, conditions, etc. to determine when an instruction should be interrupted. The storage node 1010 may also provide hints / information to the storage device 1015 indicating that if a current operation is taking too long, then another storage device 1015 may be used to perform the read operation (e.g., data may be reconstructed by accessing an erroneously encoded portion from the other storage device 1015). The storage device 1015 may provide various other methods, techniques, mechanisms, etc. to the storage system 1000 to improve the performance of the storage system 1000 (e.g., increase data throughput).
[0337] In traditional storage devices (e.g., SSDs), a storage controller may perform various functions such as mapping and remapping, garbage collection, wear leveling, etc. This can hide the management of flash memory from other systems that may use traditional storage devices (e.g., laptops, desktop computers, etc.). In one embodiment, storage system 1000 (e.g., storage node 1010) can move many or most functions performed by a storage controller to storage node 1010 (e.g., to processing device 1011). If most operations / functions are offloaded from the storage controller and the storage controller performs operation scheduling, NVME queue scheduling, internal queue scheduling, and bus management, the storage controller can be a simpler device (e.g., a less expensive device).
[0338] In one embodiment, storage device 1015 may allow a storage system (e.g., storage node 1010) to manage, improve, and / or optimize the reliability of storage device 1015. This may be achieved by better integrating the behavior / operation of storage device 1015 with the operation of storage node (e.g., processing device 1011). For example, if a block / page is bad (e.g., due to aging, read disturb, manufacturing defects, etc.), storage device 1015 may be unable to recover the data in the bad block / page. However, storage node 1010 may be able to recover data from the bad block / page using RAID parity, erasure lines, or other error correction mechanisms. Additionally, storage node 1010 may take into account the aging and read disturb of the flash memory when performing obsolete item collection and / or wear leveling operations. By moving / offloading obsolete item collection and / or wear leveling operations to storage node 1010, storage system 1000 may be able to improve the reliability and / or lifespan of the flash memory in storage device 1015.
[0339] In one embodiment, the storage system 1000 may be able to operate more efficiently and / or less expensively. For example, the interface 1019 allows the number of storage devices 1015 per storage node 1010 to be configured based on the needs of the clients / users of the storage system. The storage system 1000 may also provide greater flexibility because storage capacity can be replaced or added relatively quickly / easily (e.g., by unplugging an existing storage device 1015 and plugging a new storage device 1015 into the interface 1019). This also allows for faster / easier replacement of storage nodes 1010 while still retaining data in the storage devices 1015. For example, an older storage node 1010 may be replaced with a new storage node 1010 having a faster processing device and / or more memory. The storage devices 1015 from the old storage node 1010 may be plugged into or connected to the new storage node 1010. This allows the storage system 1000 to provide access to data in the storage devices 1010 without requiring a rebuild or copy of the data to the new storage devices.
[0340] As discussed above, to ensure that the storage system operates and that all data remains accessible to clients in the event of a temporary or permanent storage node failure, the storage system 1000 may use RAID, erasure coding, or other error correction / recovery techniques to distribute data across the storage nodes 1010. In one embodiment, because multiple storage devices 910 may be coupled to each storage node 1010, the storage system 1000 may include more than one storage device 910. Figure 9The storage system 900 illustrated in FIG. 1 may have fewer storage system controllers. Because there are fewer storage system controllers, RAID configurations, erasure coding, or other error correction / recovery techniques may take into account a smaller number of storage system controllers. For example, stripes of different widths may be used, different error correction codes may be used, local redundancy and zigzag codes may be used, and the like.
[0341] In one embodiment, the layering of services / functions of the storage system 1000 may be similar to the above-mentioned combination of Figure 9 For example, there may be a first layer (e.g., authorization) that runs across the storage nodes 1010, and a second layer (e.g., device host) that runs within the storage node 1010 to manage the flash memory 1010 in the storage device 1015.
[0342] Figure 11 11 is a block diagram illustrating an example storage system 1100 according to one or more embodiments of the present disclosure. As discussed, the storage system may include storage nodes, combo controllers, and the like. The storage nodes and / or combo controllers may be located in a housing or chassis. For example, housing 1110 may include multiple combo controllers 910, and housings 1120 and 1150 may include multiple storage nodes 1010. Storage system 1100 includes housings 1110 through 1160. The housings may be referred to as expansions, expansion housings, expansion chassis, and the like.
[0343] In one embodiment, the housing 1150 may be an expansion housing / chassis that may include additional storage nodes 1010 that may be added to the storage system 1100. Each of the storage nodes 1010 includes a storage device (e.g., a hot-swappable managed flash storage device), a processing device (e.g., a CPU), and memory (e.g., RAM, NVRAM, cache, etc.). As discussed above, the storage node 1010 may include a device interface (e.g., a communication interface) for connecting to the storage device.
[0344] In one embodiment, the storage system 1100 can be expanded by adding one or more housings 1140. The housing 1140 includes additional device storage nodes 1141 that use processing devices with lower speed / power and / or less memory (e.g., less RAM). The device storage nodes 1141 can use these lower-power or slower processing devices to operate as device hosts and / or provide storage device controller functionality. For example, the device storage node 1141 may not provide storage system controller functionality (e.g., may not have the ability to run authorization). The housing 1140 and / or device storage nodes 1141 can allow the storage system 1100 to expand the system's storage capacity more quickly, more efficiently, and / or more cheaply.
[0345] In one embodiment, housing 1140 may include a communication interface (e.g., one or more interfaces) that allows the housing to communicate with other housings, combo controllers, and / or storage nodes of storage system 1100. The communication interface may also be coupled to a storage device (e.g., a managed flash storage device) of device storage node 411 via an interface of the storage device (e.g., via a PCIe interface, an NVME interface, a bus interface, etc.).
[0346] In one embodiment, housing 1140 may have reduced network capabilities when compared to housings 1110 and 1120. For example, housing 1140 may have fewer network interfaces and / or may have less network bandwidth. Because device storage nodes 1141 lack storage system controller functionality, they may not be able to communicate with client devices or host devices via a front-end network. Instead, drive storage nodes 1141 may communicate or interact with other housings via a back-end network (e.g., a network internal to storage system 1100) (e.g., network 705).
[0347] As also discussed above, a storage system may include paired storage system controllers that operate in various ways to manage storage devices. For example, paired storage system controllers may divide the storage devices in the storage system between the storage system controllers, or a first storage system controller may perform most or all functions / services while a second storage system controller operates a backup. In other examples, paired storage system controllers may share access to the storage devices, coordinate their operations, actions, etc. in some manner to avoid inconsistencies in data stored on the storage devices. High availability system 1131 may be an example of paired storage system controllers. For example, high availability system 1131 may be Figure 1D An example of storage system 124 is illustrated in FIG. A high availability system 1131 (eg, a paired storage system controller) can be located in housing 1130 .
[0348] In one embodiment, the communication or network interface of the high availability system 1131 can be coupled to the network 705 (e.g., a cross-storage system controller network). This can add storage capacity from the high availability system 1131 to the storage system 1100 while providing support for handling errors / failures in the storage system controller of the high availability system 1131 in a scale-out storage system (e.g., by providing failover capability), thereby adding all capacity from the chassis to the scale-out storage system, including full support for storage system controller failure handling by failing over to other storage system controllers to access storage devices in the chassis.
[0349] In one embodiment, the high availability system 1131 may not provide storage system controller functionality, but may provide storage device controller functionality (e.g., may operate as a device host). Management of the storage devices in the high availability system 1131 may be divided between two controllers in the high availability system 1131. Alternatively, a first controller may manage the storage devices, and if the first controller fails, the second controller may take over management of the storage devices.
[0350] In one embodiment, the housing 1160 may be a housing (e.g., a chassis) that includes a device interface (e.g., an NVMe interface, a PCI interface, etc.) for connecting to the storage device 1161. The housing may also include a switch, a bus, and / or a network interface, such as a Fibre Channel interface, a network interface (e.g., a Remote Direct Memory Access (RDMA) over Converged Ethernet (RDM) interface), that can be coupled to the device interface to allow other housings, storage nodes, and / or a composite controller to access the storage device 1161. The housing 1160 may be connected to other housings, storage nodes, and / or a composite controller using fabric-based NVMe, Fibre Channel, or RDM.
[0351] Housing 1160 may not include storage device controller functionality. For example, housing 1160 may not include a processing device that provides storage device controller functionality. Housing 1160 may also not include storage system controller functionality. And may not include storage system controller functionality. For example, housing 1160 may not include a processing device that provides storage system controller functionality. In one embodiment, a processing device from another housing (e.g., a storage system controller) may handle the management of flash memory in storage device 1161. For example, storage node 1010 in housing 1120 may handle the management of flash memory in storage device 1161. Housing 1160 may also be referred to as a dumb expansion chassis / housing. Similar to a pair of controller systems, a failure of storage node 1010 may not affect access to storage device 1161 of housing 1160 because another storage node 1010 may take over management of storage device 1161 previously managed by the failed / shutdown / replaced storage node.
[0352] Figure 12 is a flow chart illustrating a method for accessing data according to some embodiments of the present disclosure. The method 1200 may be performed by processing logic including hardware (e.g., circuitry, dedicated logic, programmable logic, microcode, etc.), software (e.g., instructions running on a processing device to perform hardware simulation), firmware, or a combination thereof. In one embodiment, the method 1200 may be performed by a processor such as Figures 9 to 11 The storage system, storage node, combined controller, etc. described in are executed.
[0353] refer to Figure 12, method 1200 illustrates example functionality used by various embodiments. Although specific functional blocks ("blocks") are disclosed in method 1200, such blocks are examples. That is, embodiments are well suited to performing various other blocks or variations of blocks set forth in method 1200. It should be understood that the blocks in method 1200 may be performed in an order different from that presented, that not all blocks in method 1200 may be performed, and that other blocks (which may not be included in the method 1200) may be performed in the same order as described. Figure 12 (in Chinese) can be Figure 12 Execute between the boxes described in .
[0354] Method 1200 begins at block 1205, where a request to access data stored in a storage system is received. The storage system includes a plurality of storage nodes. The plurality of storage nodes are configured to provide storage services for the storage system, as discussed above. Each storage node includes a processing device, a memory, a network interface, and a plurality of device interfaces that support hot-swappable managed flash storage devices. The managed flash storage devices are managed by the respective storage node in cooperation with other storage nodes of the storage system to provide storage services. The respective managed flash storage devices include a storage device controller that supports a set of commands for access by the respective storage node and optimizes the use of the flash memory for use in the storage system, as discussed above.
[0355] In one embodiment, the request to access data may be a request to read data stored in a storage system. In another embodiment, the request to access data may be a request to write data to a storage system.
[0356] At block 1210, a first storage node is identified based on the request. For example, if the request is to access data, the storage system may identify the first storage node because the first storage node may include a storage device that stores the data. In another example, if the request is to write data to the storage system, the storage system may identify the first storage node because the first storage node may have a storage device with sufficient space to store the data.
[0357] At block 1215, the first storage node is instructed to provide access to the data. For example, if the request is to access the data, the storage system may identify the first storage node because the first storage node may include a storage device that stores the data. In another example, if the request is to write the data to the storage system, the storage system may identify the first storage node because the first storage node may have a storage device with sufficient space to store the data.
[0358] Figure 13is a flow chart illustrating a method 1300 for configuring a flash memory according to some embodiments of the present disclosure. The method 1300 may be performed by processing logic including hardware (e.g., circuitry, dedicated logic, programmable logic, microcode, etc.), software (e.g., instructions running on a processing device to perform hardware simulation), firmware, or a combination thereof. In one embodiment, the method 1300 may be performed by a processor such as Figures 9 to 11 The storage system, storage node, combined controller, etc. described in are executed.
[0359] refer to Figure 13 , method 1300 illustrates example functionality used by various embodiments. Although specific functional blocks ("blocks") are disclosed in method 1300, such blocks are examples. That is, embodiments are well suited to performing various other blocks or variations of blocks set forth in method 1300. It should be understood that the blocks in method 1300 may be performed in an order different from that presented, that not all blocks in method 1300 may be performed, and that other blocks (which may not be included in the method 1300) may be performed in an order different from that presented. Figure 13 (in Chinese) can be Figure 13 Execute between the boxes described in .
[0360] Method 1300 begins at block 1305, where the storage system can detect that a first managed flash storage device has been added to the storage system. A managed flash storage device (which can also be referred to as a directly managed flash storage device, a directly managed storage device, a managed storage device, etc.) can provide functionality, operations, commands, APIs, or some other suitable mechanism for an external device (e.g., a storage node) to control, manage, and / or interact with the flash memory of the managed flash storage device, as discussed above. The storage system can detect that the managed flash storage device has been connected to an interface (e.g., a device interface, a bus interface), etc., of a storage node.
[0361] At block 1310, the storage system may identify, determine, select, or the like one or more parameters for managing the flash memory of the first managed flash storage device. For example, the storage system may identify reliability, memory capacity, or the like as parameters for configuring and / or using the flash memory of the first managed flash storage device. The parameters may be identified based on user input, a configuration file / settings of the storage system, or the like.
[0362] At block 1315 , the flash memory of the first managed flash storage device is configured based on one or more parameters. For example, the programming mode, block size, and frequency of garbage collection of the flash memory may be configured based on the one or more parameters.
[0363] Figure 1414 is a block diagram illustrating an example storage device 1400 according to one or more embodiments of the present disclosure. Storage device 1400 includes a processing device 1410, a memory 1420, an energy storage device 1430, a flash memory 1440, and a flash management module 1450. Processing device 1410 may be referred to as a storage device controller. In one embodiment, storage device 1400 may be a hot-swappable managed flash storage device. Storage device 1400 may be included as part of a storage system. For example, storage device 1400 may be included in a housing, a storage node, a combined controller, etc.
[0364] In one embodiment, the storage device 1400 may be an abstraction-based storage device. The abstraction-based storage device may provide a model 1453 that may be used by an external device (e.g., a storage node, a storage system controller, etc.) to manage how the flash memory 1440 operates without requiring direct interaction with the flash memory 1440. For example, the model may provide programming modes, error correction parameters, obsolete item collection parameters, block size, the amount of memory 1420 available for temporary storage of data, etc. The processing device 1410 may manage the flash memory 1440 based on the model 1453 and based on requests received from external devices (e.g., storage nodes) to access / use the flash memory 1440.
[0365] In one embodiment, because storage device 1400 can allow external devices to request one or more data stores, storage device 1400 can be referred to as an abstraction-based storage device. A data store can be an abstract unit of storage, as discussed in more detail below. For example, an external device can request a data store of a specific size for writing data to storage device 1400. The size of the data store can be different from the internal block / page size used by storage device 1400. In one embodiment, the various models 1453 used by storage device 1400 can be based on the size of the requested data store.
[0366] In one embodiment, storage device 1400 can use a partition namespace to provide an abstraction for flash memory 1440. The partition namespace can present large blocks of flash memory 1440 that are mapped in some way to erase blocks of storage device 1400. This allows external devices or storage systems to efficiently allocate, write, and deallocate erase blocks, or some storage device internal structures associated with erase blocks, while allowing processing device 1410 to manage the life cycle of flash memory 1440. Additionally, mecha...
Claims
1. A storage system comprising: A plurality of storage nodes configured to provide storage services for the storage system, wherein each storage node comprises: a processing device, a memory, a network interface, and a plurality of device interfaces supporting hot-swappable managed flash storage devices, wherein the managed flash storage devices are managed by the storage node in cooperation with other storage nodes of the storage system to provide the storage service, and wherein the respective managed flash storage devices include: Flash memory and a storage device controller that supports a set of commands for access by the corresponding storage nodes and optimizes use of the flash memory for use in the storage system, wherein at least a subset of the managed flash storage devices further includes memory and energy storage for staging writes and for storing metadata for the storage system.
2. The storage system of claim 1 , wherein upon a power failure, the energy storage is utilized by the storage device controllers of the subset of the managed flash storage devices to transfer temporary writes and metadata to the flash memories of the managed flash storage devices. 3 . The storage system according to claim 1 , wherein the storage service comprises at least one of a file service, a block storage service, an object storage service, and a key-value database service. 4 . The storage system of claim 1 , wherein the network interface is configured to interface with one or more client devices.
5. The storage system according to claim 1, wherein the service of the storage system further comprises: Data management services, including erasure coding across managed flash storage devices and across storage nodes; and Wear leveling is performed on a plurality of the managed flash storage devices of the storage system.
6. The storage system according to claim 1, wherein the service of the storage system further comprises: Data services, including snapshots, cloning, and replication.
7. The storage system of claim 1, wherein a plurality of the managed flash storage devices comprise directly managed flash storage devices, and wherein the plurality of storage nodes address read and write requests to specific flash memory cells.
8. The storage system of claim 1, wherein the plurality of managed storage devices comprises abstraction-based managed flash storage devices, and wherein the storage nodes control the amount of flash memory used in a specific manner.
9. The system of claim 1 , wherein the storage system comprises a first housing comprising: A first set of storage nodes and at least one networking switch interconnect the first set of storage nodes of the first enclosure and in data communication with one or more client devices.
10. The system of claim 9, wherein the at least one networking switch is further connected to at least one additional storage system.
11. The system of claim 9, further comprising an additional enclosure comprising an additional interface for hot-pluggable managed flash storage devices addressable from one or more storage nodes of the first enclosure.
12. The system of claim 11 , wherein the additional housing comprises an extended storage node housing comprising an additional storage node, the additional storage node comprising an additional interface for hot-swappable managed flash storage devices, and wherein the additional storage node comprises an additional CPU and memory for the storage system.
13. A system according to claim 11, wherein the additional shell includes a high availability (HA) shell, which includes at least two storage system controllers connected to a device interconnect bus, wherein the device interconnect bus provides access from the at least two storage system controllers to a hot-swappable managed flash storage device of the additional shell, and wherein the managed flash memory of the managed flash storage device is managed by a combination of the HA shell and the storage nodes of other shells of the storage system.
14. A system according to claim 11, wherein the additional shell includes an extended drive extension shell, which includes at least two interfaces connected to at least one shell of the storage system including a storage node, wherein the two interfaces are further interfaced with an internal bus of the drive extension shell including a connection to an interface for a hot-swappable managed flash storage device, and wherein the flash memory of the hot-swappable managed flash storage device inserted into the drive extension shell is managed by the storage nodes of other shells of the storage system.
15. A method comprising: receiving, by a storage system comprising a plurality of storage nodes, a request to access data of the storage system, wherein the plurality of storage nodes are configured to provide storage services for the storage system, and wherein a respective storage node comprises a processing device, a memory, a network interface, and a plurality of device interfaces supporting a hot-swappable managed flash storage device, wherein the managed flash storage device is managed by the respective storage node in cooperation with other storage nodes of the storage system to provide the storage services, and wherein the respective managed flash storage device comprises a storage device controller that supports a set of commands for access by the respective storage node and for optimizing use of the flash storage for use in the storage system; identifying a first storage node based on the request, wherein the first storage node includes a first managed flash storage device for storing the data; and The first storage node is instructed to provide access to the data.
16. The method of claim 15, wherein providing access to the data comprises one or more of: Reading the data from the first managed flash storage device; and The data is written to the first managed flash storage device.
17. The method of claim 15, further comprising: detecting that a first managed flash storage device has been added to the storage system; identifying one or more parameters for managing the flash memory of the first managed flash storage device; and The flash memory of the first managed flash storage device is configured based on the one or more parameters.
18. The method of claim 15, wherein the one or more parameters include one or more of the following: a lifespan of a portion of the flash memory; an amount of memory available in the first managed flash storage device; performance of the first managed flash storage device; and Reliability of a portion of the flash memory.
19. The method of claim 15, wherein a plurality of the managed flash storage devices comprise directly managed flash storage devices, and wherein the plurality of storage nodes address read and write requests to specific flash memory cells.
20. A non-transitory computer-readable storage medium storing instructions that, when executed, cause a processing device to: receiving, by a storage system comprising a plurality of storage nodes, a request to access data of the storage system, wherein the plurality of storage nodes are configured to provide storage services for the storage system, and wherein a respective storage node comprises a processing device, a memory, a network interface, and a plurality of device interfaces supporting a hot-swappable managed flash storage device, wherein the managed flash storage device is managed by the respective storage node in cooperation with other storage nodes of the storage system to provide the storage services, and wherein the respective managed flash storage device comprises a storage device controller that supports a set of commands for access by the respective storage node and for optimizing use of the flash storage for use in the storage system; identifying a first storage node based on the request, wherein the first storage node includes a first managed flash storage device for storing the data; and The first storage node is instructed to provide access to the data.
Citation Information
Cited By
Data lake storage optimization method and platform for enterprise business integration
CN120973835A