Storage system and data storage method
By using flash drives of different granularities to build logical blocks in RAID groups, the capacity inconsistency problem caused by different ZNS SSD types is solved, achieving logical block capacity consistency and simplifying metadata management, avoiding storage space waste, and improving the reliability and space utilization of the storage system.
Patent Information
- Application Number
- CN202410649779.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-05-23
- Publication Date
- 2025-12-02
AI Technical Summary
When building a RAID group using different types of ZNS SSDs, inconsistent logical block capacities increase the complexity of metadata management and may lead to wasted storage space.
By using two flash drives with different granularities to build logical blocks, the logical block capacity is made consistent, simplifying metadata management. Furthermore, by combining flash drives to build storage pools, the consistency of logical block capacity is ensured.
It achieves consistency in logical block capacity, simplifies metadata management complexity, avoids storage space waste, and improves the reliability and space utilization of the storage system.
Smart Images

Figure CN121050641A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of storage, and more particularly to a storage system and a data storage method. Background Technology
[0002] Solid-state drives (SSDs) typically use expensive dynamic random access memory (DRAM) to cache the flash translation layer (FTL) mapping table and require a significant amount of storage space for garbage collection (GC). To reduce these overheads, SSDs with a zoned namespace (ZNS) interface were developed. See also... Figure 1 In ZNS interface SSDs (ZNS SSDs for short), the storage space is divided into multiple zones, and each zone includes one or more blocks. Therefore, its mapping granularity is much larger than that of a page.
[0003] Storage system (reference) Figure 2 In Redundant Array of Independent Disks (RAID) using ZNS SSDs, storage space is typically provided to the RAID group at the zone level. Since the granularity of a zone is larger than that of the original page level, if the RAID group contains multiple types of ZNS SSDs and the zone sizes contained in different types of zone SSDs are inconsistent, it is difficult to ensure that the capacity of different logical blocks within the RAID group is consistent.
[0004] When the capacity of logical blocks within a RAID group is inconsistent, it is necessary to record and manage the capacity of multiple logical blocks simultaneously, which leads to complex metadata management. Summary of the Invention
[0005] This application provides a storage system and a data storage method to reduce the complexity of metadata management.
[0006] Firstly, this application provides a storage system including a RAID group, which comprises multiple logical blocks. The storage system includes flash drives 1 and 2. When flash drives 1 and 2 provide storage space to the logical blocks at different granularities, logical block 0 can be constructed jointly by flash drives 1 and 2, ensuring that the capacity of logical block 0 is the same as that of the other logical blocks. This same capacity is not absolute; the difference only needs to be within an acceptable range. Thus, for this RAID group, only one logical block capacity needs to be recorded, simplifying metadata management complexity. Furthermore, when the logical block capacities are consistent, the number of data stripe units in the data stripes based on the logical blocks is the same. Therefore, only one data stripe type needs to be recorded, further simplifying metadata management complexity. Simultaneously, since the number of data stripe units in the data stripes is the same, storage space waste can be avoided.
[0007] In one possible implementation, the granularity at which flash drive 1 provides storage space to logical block 0 is a page, wherein flash drive 1 includes multiple erase blocks, and one erase block includes multiple pages.
[0008] In one possible implementation, the granularity of the storage space provided by the flash drive 2 to logical block 0 is partitioning, wherein the storage space of the flash drive 2 is divided into several partitions, and one partition includes one or more erase blocks; or the granularity of the storage space provided by the flash drive 2 to logical block 0 is PLOG, in which case the storage space of the flash drive 2 is divided into several PLOGs, and one PLOG includes one or more erase blocks.
[0009] In one possible implementation, the number of flash drives 2 in the storage system is at least two, or more than two, and at least two of the flash drives 2 include partitions of different sizes. In this case, flash drives 1 and flash drives 2 jointly construct logical block 0, making logical block 0 have the same capacity as other logical blocks, thereby simplifying the complexity of metadata management.
[0010] In one possible implementation, the storage space of logical block 1 in the multiple logical blocks comes from flash drive 3, and the granularity of flash drive 3 providing storage space to logical block 1 is partitioning. The storage space of logical block 2 in the multiple logical blocks comes from flash drive 4, and the granularity of flash drive 4 providing storage space to logical block 3 is page.
[0011] In one possible implementation, flash drive 1 and flash drive 2 are physically independent flash drives; or flash drive 1 and flash drive 2 are different parts of a hybrid flash drive; wherein flash drive 1 is configured with a portion of the flash memory chips included in the hybrid flash drive, and flash drive 2 is configured with another portion of the flash memory chips included in the hybrid flash drive. Both of these implementations are acceptable, and no specific limitation is made in this embodiment.
[0012] In one possible implementation, the storage space of each logical block in the multiple logical blocks comes from different flash drives. This ensures that if one or two flash drives used to build the RAID group fail, the data can be reconstructed using the data stored on the other flash drives, thereby improving the reliability of the storage system.
[0013] In one possible implementation, the RAID group is located in the disk enclosure, which also includes a control unit. The control unit constructs the RAID group and, after constructing the RAID group, establishes a mapping relationship between the logical block address of logical block 0 and flash disks 1 and 2, so that it can be used for subsequent writing and reading of data from the flash disks.
[0014] In one possible implementation, the control unit is also used to receive write requests and, based on the mapping relationship established after the RAID group is constructed, determine the flash drive 1 and flash drive 2 corresponding to the logical block address carrying data in the write request, and then store the data in the pages of flash drive 1 including the erase block and the partitions included in flash drive 2.
[0015] In one possible implementation, the control unit is further configured to send a write command 1 to flash drive 1 before storing data to flash drive 1 and flash drive 2. The write command 1 carries a logical block address 1 and the length of data 1, so that flash drive 1 determines the erase block that can be written to based on the length of data 1, writes data 1 to the pages included in the determined erase block, and establishes a mapping relationship between the pages and logical block address 1. The control unit is also configured to send a write command 2 to flash drive 2. The write command 2 carries a logical block address 2 and the length of data 2, so that flash drive 2 determines the partition that can be written to based on the length of data 2, and establishes a mapping relationship between the partition identifier, the offset within the partition, and logical block address 2 after writing data 2 to the partition.
[0016] In one possible implementation, the RAID group is located in the disk enclosure, and the storage system also includes a controller that is communicatively connected to the disk enclosure. The controller is used to build the RAID group and, after building the RAID group, establish a mapping relationship between logical block 1 and flash disk 1 and flash disk 2.
[0017] Secondly, this application also provides a data storage method applied to a storage system, the storage system including a Redundant Array of Independent Disks (RAID) group, the RAID group including multiple logical blocks, the multiple logical blocks having the same capacity, the method including: obtaining a write request, the write request carrying data; storing the data respectively into flash drive 1 and flash drive 2 corresponding to logical block 1 in the multiple logical blocks, flash drive 1 providing storage space for logical block 1 at granularity 1, and flash drive 2 providing storage space for logical block 1 at granularity 2, the granularity 1 and granularity 2 being different.
[0018] In one possible implementation, flash drive 1 includes multiple erase blocks, and each erase block includes multiple pages; the storage space of flash drive 2 is divided into several zones, and each zone includes one or more erase blocks; granularity 1 is a page, and granularity 2 is a zone.
[0019] In one possible implementation, the request also carries a logical block address pointing to logical block 0, and storing data into flash drives 1 and 2 corresponding to logical block 0 in multiple logical blocks includes: determining flash drives 1 and 2 corresponding to the logical block address according to the mapping relationship between the logical block address and flash drives 1 and 2; and storing data into partition 1 of flash drive 2 and pages included in erase block 1 of flash drive 1.
[0020] In one possible implementation, the method further includes: before storing the data into partition 1 of flash drive 2 and pages included in erase block 1 of flash drive 1, respectively, sending a write instruction 1 to flash drive 1, the write instruction 1 carrying a logical block address 1 and the length of data 1; sending a write instruction 2 to flash drive 2, the write instruction 2 carrying a logical block address 2 and the length of data 2; wherein the data includes data 1 and data 2, and the logical block address includes logical block address 1 and logical block address 2.
[0021] Thirdly, this application also provides a computer-readable storage medium including instructions that, when executed on a computer, cause the computer to perform the data storage method as described in the second aspect above.
[0022] Fourthly, this application also provides a computer program product that, when run on a computer, causes the computer to perform the data storage method as described in the second aspect above. Attached Figure Description
[0023] Figure 1 This is a schematic diagram illustrating the host-side presentation of ZNS SSDs in existing technologies.
[0024] Figure 2 A diagram illustrating the construction of a RAID group using different types of ZNS SSDs;
[0025] Figure 3 This is a schematic diagram illustrating the host-side presentation of an HDD that supports the Block interface protocol in the existing technology.
[0026] Figure 4 This is a schematic diagram illustrating the host-side presentation of an SSD that supports the Block interface protocol in existing technologies.
[0027] Figure 5 In order to be in Figure 2 A schematic diagram illustrating data striping on a constructed RAID group;
[0028] Figure 6 In order to be in Figure 2 This diagram illustrates the waste of storage space that can occur when data is striped from a RAID group.
[0029] Figure 7 A schematic diagram of the structure of a storage system provided in this application;
[0030] Figure 8 A schematic diagram of the hardware structure of an SSD provided in this application;
[0031] Figure 9 This is a schematic diagram of the internal storage structure of a flash memory chip included in an SSD.
[0032] Figure 10 A schematic diagram of constructing a RAID group is provided for this application;
[0033] Figure 11 A schematic diagram of another storage system provided in this application;
[0034] Figure 12 A schematic diagram of another storage system provided in this application;
[0035] Figure 13 A schematic diagram of another storage system provided in this application;
[0036] Figure 14 A schematic diagram of another storage system provided in this application;
[0037] Figure 15 A schematic diagram of another storage system provided in this application;
[0038] Figure 16 A schematic diagram of another storage system provided in this application;
[0039] Figure 17 A schematic diagram of another storage system provided in this application;
[0040] Figure 18 A schematic diagram of another storage system provided in this application. Detailed Implementation
[0041] To make the objectives, technical solutions, and advantages of this application clearer, the specific embodiments of this application will be described in further detail below with reference to the accompanying drawings.
[0042] To better understand the storage system and the RAID group constructed by the storage system provided in the embodiments of this application, these interface protocols will be further described below.
[0043] Figure 3 This is an abstract connection diagram between a hard disk drive (HDD) supporting the Block interface protocol and the host device. An HDD supporting the Block interface protocol is displayed on the host side as a one-dimensional array of the entire LBA (Logical Block Address) within a namespace. This allows the host to read, write, and overwrite data in any order without considering the underlying physical implementation, thus simplifying storage management for the host software. For HDDs supporting the Block interface protocol, there is a static mapping between physical block addresses and logical block addresses.
[0044] For SSDs, read and write operations are typically performed at the page level, while erasure is performed at the block level. The granularity of read / write operations and erasure operations within the disk is not consistent. When data needs to be overwritten, a blank page needs to be found; that is, the updated data is written to a blank page located elsewhere, while the data in the original page is marked as garbage. In other words, SSDs employ a remote data update strategy, which means that the mapping between logical block addresses seen by the host and physical block addresses within the disk is not static but dynamically changing. Please see [link to relevant documentation]. Figure 4 To ensure compatibility with the Block interface protocol, an FTL layer is added within the SSD. To support this FTL layer, the SSD requires expensive DRAM to cache the FTL mapping table. The larger the SSD capacity, the larger the mapping table, and the larger the required DRAM capacity. Furthermore, SSDs typically also include over-provisioning (OP) space for garbage collection. Over-provisioning space refers to the capacity that users cannot operate on, and its size is the actual SSD capacity minus the user-available capacity. Therefore, these operations not only cause significant performance fluctuations and write amplification in the SSD but also require large-capacity DRAM cache and OP, significantly increasing hardware costs.
[0045] To reduce the high overhead associated with Block interface protocol compatibility, ZNS SSDs divide the LBA (Local Address Mapping) seen by the host into multiple partitions. Within each partition, writes are generally sequential; random writes and in-place updates are not allowed. If data updates are needed, the entire partition must be reset before writing can begin again from the beginning. In this scenario, the address mapping table maintained by this storage device is at the partition level, while storage devices supporting the Block interface protocol maintain it at the page level. The partition granularity is much larger than the page granularity, resulting in less metadata and reduced DRAM usage. Please also refer to [further details omitted]. Figure 2 ,exist Figure 2 In ZNS SSDs, the storage space is divided into multiple physical partitions, while the LBA (Leveled Base) seen by the host is divided into multiple logical partitions. ZNS SSDs align the granular boundaries of physical and logical partitions, meaning the mapping between physical and logical partitions is static and one-to-one. This allows ZNS SSDs to transfer all data management responsibility to the host, with garbage collection entirely handled by the host. Specifically, when reclaiming a logical partition, the host moves the valid data within that logical partition to another empty logical partition. Because the granularity of the logical and physical partitions is consistent, the SSD only needs to directly erase the physical partition to be reclaimed, without needing to perform garbage collection within the SSD. Therefore, there is no need to reserve capacity space, and the computing performance requirements of the processor within the SSD are also reduced.
[0046] While ZNS SSDs overcome some of the drawbacks of Block SSDs—such as requiring only a coarse-grained address mapping table for each partition, resulting in minimal DRAM requirements—and the fact that partition resets invalidate all blocks within a partition, eliminating the need for garbage collection and thus reducing performance fluctuations and excess provisioning space, they still have shortcomings in other applications. For instance, certain issues may arise when building RAID arrays using different types of ZNS SSDs.
[0047] When different types of ZNS SSDs are used to build a RAID group, meaning the partition space sizes of different ZNS SSDs differ, the logical block capacities within the resulting RAID group may vary. For example, some logical blocks in a RAID group may have large capacities, while others may have small capacities. It's important to note that as SSDs evolve towards larger drives (SSDs with storage capacities greater than 30TB are typically referred to as large SSDs), SSD capacity increases generally occur through two methods: either increasing the number of pages within a single block, or keeping the number of pages within a single block the same but increasing the capacity of the pages themselves. In this evolutionary process, two different types of SSDs may exist: ordinary drives and large drives. The storage capacity of blocks in ordinary drives and large drives will differ, further leading to different zone sizes based on those blocks, resulting in the aforementioned different types of ZNS SSDs.
[0048] exist Figure 2 The example demonstrates the situation where logical block capacities are inconsistent. When logical block capacities are inconsistent, a RAID group needs to record the capacities of multiple logical blocks. Furthermore, in Figure 2 Based on this, please refer to Figure 5 When a RAID group is divided into multiple data stripes, there are two types of data stripes. The first type uses four data stripes (D1, D2, D3, and D4) to calculate the parity shards P and Q (4+2 mode). The second type uses two data stripes (D10 and D11) to calculate the parity shards P and Q (2+2 mode). Compared to a single RAID mode, when a RAID group includes more than one type of data stripe, it is necessary to record and manage the lengths of multiple data stripes. Furthermore, different data stripe lengths result in different storage locations for the parity shards, further increasing the complexity of metadata management. The more ZNS SSDs and data stripe types within the RAID group, the greater the management difficulty.
[0049] On the other hand, when there are many types of ZNS SSDs, it is easy to have unusable storage space, resulting in wasted storage space. For example, please see... Figure 6 If a ZNS SSD with a partition granularity much larger than other member disks exists in the flash drive used to build the RAID group, it cannot support the minimum 2+2 mode. That is, if there are no other ZNS SSDs with the same partition granularity to form a data stripe with it, then the logical block built by the partition using this ZNS SSD may have wasted space.
[0050] To resolve the above technical issues, please refer to [link / reference]. Figure 7 This application provides a storage system 700, which includes a RAID group 701. The RAID group 701 includes multiple logical blocks. The storage space of logical block 7011 among the multiple logical blocks comes from at least flash drives 702 and 703. Flash drive 702 provides storage space to logical block 7011 at a first granularity, and flash drive 703 provides storage space to logical block 7011 at a second granularity. In this application embodiment, the first logical block is constructed using flash drives with two different granularities, so that the capacity of the constructed first logical block is the same as the capacity of the other logical blocks among the multiple logical blocks. Thus, only the capacity size of one type of logical block needs to be recorded in such a RAID group, thereby reducing the complexity of metadata management.
[0051] First, let's introduce the hardware structure of a flash drive. Please refer to [link / reference]. Figure 8 The SSD includes a main controller 801, a flash array 802 composed of multiple flash memory chips, and a cache 803. The main controller 801 is connected to the host, the flash array 802, and the cache 803. The main controller 801 is responsible for complex tasks such as managing data storage and maintaining SSD performance and lifespan. For example, the main controller 801 receives access commands from the host, parses them, converts them into commands that allow direct access to the flash array 802, sends them to the flash array 802, retrieves the access results, and returns the results to the host. The main controller 801 is typically presented as an application-specific integrated circuit (ASIC), but can also be implemented based on a field-programmable gate array (FPGA) or a central processing unit (CPU). In practical applications, considering factors such as cost, performance, and power consumption, the controller is usually implemented as an ASIC chip. The cache 803 is optional in the SSD and is usually implemented using dynamic random access memory (DRAM) to store various data generated during operation, which helps improve the controller's response speed to inference device commands. The main controller 801 also includes an interface and several channel controllers. The interface is used for communication with the host. With several channel controllers, the main controller 801 can operate multiple flash memory chips in parallel, thereby increasing the bandwidth of the underlying layer.
[0052] The flash memory array 802 is used to store various types of data. Specifically, the flash memory array 802 may include one or more flash memory dies, each of which is typically presented as a chip. The specific type of flash memory is NAND flash memory. It should be noted that although flash memory is used as an example here, other types of non-volatile memory can also be used, such as phase-change memory (PCM) and resistive random access memory (RRAM), without affecting the technical solution of this application. The distribution of the storage space of a single flash memory die can be found in [reference needed]. Figure 9 A flash memory chip is internally divided into two regions (Plane). A plane contains multiple blocks, and a block consists of several pages. For example, every 128 pages make up a block, and every 2048 blocks make up a plane. The address of a page is called the Physical Block Address (PBA). Depending on the manufacturer and manufacturing process, typical page values are 4 kilobytes (KB), 8 KB, etc., while typical block values are 8 megabytes (MB), 16 MB, etc. With the increasing demand for high-capacity storage, the storage capacity of storage devices continues to grow, and the value of a block can reach hundreds of megabytes.
[0053] The following is combined with Figure 9 This illustrates the granularity at which flash drives 702 and 703 provide storage space to logical block 7011. Flash drive 702 includes multiple erase blocks to... Figure 9 Taking the block shown as an example, a block is 8MB in size, contains 1000 pages, and each page is 8KB in size. The storage space of the flash drive 703 is divided into several partitions, each partition containing one or more erase blocks. Figure 9Taking the block shown as an example, a block is 8MB in size, and a partition contains 10 blocks, with a total size of 80MB. Since the first granularity and the second granularity are different, the first granularity can be larger or smaller than the second granularity. When the first granularity is smaller than the second granularity, the first granularity can be a page, and the second granularity is a partition. In this embodiment, the flash drive that provides storage space to logical block 7011 at the page granularity is called a Block SSD, and the flash drive that provides storage space to logical block 7011 at the partition granularity is called a ZNS SSD. In some possible implementations, the storage space of flash drive 703 is divided into several PLOGs (Persistence Layer LOGs), and a PLOG includes one or more erase blocks. In this case, the second granularity is the PLOG, and the flash drive that provides storage space to logical block 7011 at the Plog granularity is called a Plog SSD.
[0054] In this embodiment, flash drives 702 and 703 can be physically independent flash drives, or they can be two parts located in a flash memory device. This flash memory device can be a hybrid flash drive, comprising multiple flash memory chips. A first portion of the flash memory chips is configured as flash drive 702, and a second portion of the flash memory chips is configured as flash drive 703. In specific implementation, the hybrid flash drive includes two interfaces, interface 1 and interface 2, which are physically independent. It also includes two main controllers, main controller 1 and main controller 2. Interface 1 is connected to main controller 1, and interface 2 is connected to main controller 2. Interface 1 is used for reading and writing data to flash drive 702, and interface 2 is used for reading and writing data to flash drive 703. In this implementation, when flash drive 702 is a Block SSD, the first interface can be any of the following: Serial Advanced Technology Attachment (SATA), Serial Attached SCSI (SAS), or Peripheral Component Interconnect Express (PCIe). When flash drive 703 is a ZNS SSD, the second interface can be a Non-Volatile Memory Express (NVMe) interface, or any other interface that supports the ZNS instruction set, which is a set of storage device-specific instructions used to manage the region namespace in the ZNS SSD. When flash drive 703 is a PLOG SSD, the second interface can also be a PLOG interface. In some possible implementations, the hybrid flash drive includes a universal interface that can support both flash drive 702 and flash drive 703 for reading and writing data; this universal interface can be an interface that supports the ZNS instruction set. In the following description, we will use the example of flash drives 702 and 703 being physically independent flash drives.
[0055] The following description uses the 702 flash drive as an example (Block SSD) and the 703 flash drive as an example (ZNS SSD).
[0056] After introducing the concepts of first-level and second-level granularity, we will further explain the concept of logical blocks included in a RAID group by examining the RAID group construction process. Assuming we need to create a RAID type of RAID 6 with an EC redundancy ratio of 4+2, and the size of the logical blocks to be built is 5MB, then based on the above information, we know that we need to build six logical blocks, each with a capacity of 5MB. Here, we will represent these six logical blocks as logical blocks 0 through 5, with logical block 0 being logical block 7011. Please refer to [link / reference]. Figure 10 In addition to flash drive 703ZNS SSD1, storage system 700 also includes other ZNS SSDs, specifically ZNS SSD1-ZNSSSD6. Similarly, in addition to flash drive 702Block SSD1, it also includes other Block SSDs, specifically BlockSSD1-Block SSD3. The granularity of the storage space provided by ZNS SSD1-ZNS SSD6 to logical blocks 0-5 differs. This difference can be that the granularity of the storage space provided by ZNS SSD1-ZNS SSD6 to the logical blocks is completely different; or the granularity of the storage space provided by ZNS SSD1-ZNS SSD6 to the logical blocks can be partially the same and partially different. For example, ZNS SSD1-ZNSSSD3 all provide 80MB of storage space to the logical blocks, while ZNS SSD3-ZNS SSD6 all provide 120MB of storage space to the logical blocks, but this granularity differs from that of ZNS SSD1-ZNSSSD6. The granularity of storage space provided by SSD3 to logical blocks is 80MB, while the granularity of storage space provided by Block SSD1 to Block SSD3 to logical blocks can be the same or different. When the granularity of storage space provided by Block SSD1 to Block SSD3 to logical blocks is different, it can be partially the same and partially different.
[0057] Before building the RAID group, the storage space of each ZNS SSD is partitioned. Each ZNS SSD's storage space is divided into several partitions, which are called physical partitions. All physical partitions within a ZNS SSD are of the same size. For example, the physical partition size of ZNS SSD1 is 4MB, ZNS SSD2 is 2MB, ZNS SSD3 is 1.5MB, ZNS SSD4 is 1MB, ZNS SSD5 is 1MB, and ZNS SSD6 is 1MB. Block SSD1-Block SSD3 each contain multiple erase blocks, called physical erase blocks. Each physical erase block contains multiple pages.
[0058] After determining the physical partitions of ZNS SSD1-ZNS SSD6, these physical partitions are then mapped to logical partitions, and the physical erase blocks of Block SSD1-Block SSD3 are mapped to logical erase blocks. The logical partitions of ZNS SSD1-ZNS SSD6 and the logical erase blocks of Block SSD1-Block SSD3 then constitute a storage pool, which is used to provide storage space upwards. In the specific implementation, logical blocks can be constructed based on the storage pool. Therefore, in this embodiment, the logical blocks are constructed from the mapped logical partitions and / or logical erase blocks. Taking logical block 0 as an example, firstly, logical partition 01 is retrieved from the storage pool. Logical partition 01 is 4MB in size, which cannot provide 5MB of storage space. If another logical partition of the same size is retrieved, the total size of the two logical partitions would be 8MB, exceeding 5MB. In this case, 128 pages from logical erase block 01 can be retrieved from the storage pool. Each page is 8KB in size. Thus, the 128 pages from logical erase block 01 can form a 5MB logical block 0 with logical partition 01. Here, logical partition 01 corresponds to physical partition 01 in ZNS SSD1, and logical erase block 01 corresponds to the Block. In SSD1, there is a physical erase block 01. For logical block 1, logical partitions 11 and 12 can be taken from the storage pool. Logical partitions 11 and 12 are both 2MB in size, totaling 4MB. If another logical partition of the same size is taken, the total would be 6MB, exceeding 5MB. In this case, 128 pages from logical erase block 11 can be taken from the storage pool. Each page is 8KB. Thus, logical erase block 11, logical partitions 11 and 12 can construct a 5MB logical block 1. Logical partitions 11 and 12 correspond to physical partitions 11 and 12 in ZNS SSD2, respectively. Logical erase block 11 corresponds to a block. One physical erase block in SSD2; for logical block 2, logical partitions 21, 22, and 23, each 1.5MB in size, can be taken from the storage pool, totaling 4.5MB. Additionally, 64 pages (8KB each) need to be taken from logical erase block 21. Thus, logical erase block 21, along with logical partitions 21, 22, and 23, can form a 5MB logical block. The remaining logical blocks can be constructed in this way, ensuring that all logical blocks have the same size. "Significant size" means the logical blocks are roughly the same size, with acceptable margins of error, not necessarily absolutely identical.
[0059] After construction is complete, each logical block in the logical block group is allocated a set of logical block addresses. This set of logical block addresses constitutes the logical block addresses of the completed logical block group. As can be seen from the logical block construction process described above, the mapping relationship between logical blocks and physical partitions and / or physical erase blocks is not a simple one. Taking logical block 0 as an example, the storage space of logical block 0 consists of physical partitions on ZNS SSD1 and physical erase blocks on Block SSD1. That is, logical block 0 corresponds to physical partitions 01 and 01 on ZNS SSD1 and physical erase block 01 on ZNS Block1. Therefore, it is necessary to record the correspondence between the logical block addresses of logical block 0 and ZNS SSD1 and ZNS Block1 for later data reading. Similarly, the mapping relationship between the logical block addresses of other logical blocks and other ZNS SSDs and / or other Block SSDs also needs to be recorded for later data reading. This mapping relationship can be seen in Table 1 below. It should be noted here that... Figure 10 For illustrative purposes, there is a one-to-one correspondence between logical partitions and physical partitions, and between logical erase blocks and physical erase blocks. However, when actually writing data to a ZNS SSD, it is not necessary to write data to the physical erase block corresponding to the logical partition according to the mapping relationship shown in the diagram. As long as there is a physical partition on the flash drive corresponding to the logical partition that can be written to, it is fine. The same principle applies to Block SSDs.
[0060] Table 1
[0061] Logical block identifier First flash drive identifier and logical block address Second flash drive identifier and logical block address 0 ZNS SSD1: LABA200-LBA204 Block SSD1: LABA205-LBA209 1 ZNS SSD2: LABA210-LBA214 Block SSD2: LABA215-LBA219 2 ZNS SSD3:LABA220-LBA2224 Block SSD3: LABA225-LBA229 3 ZNS SSD4: LABA230-LBA234 4 ZNS SSD5: LABA235-LBA239 5 ZNS SSD6: LABA240-LBA244
[0062] The above example of constructing a RAID group is merely an illustration. In actual implementation, other methods can be used, as long as the capacity of the constructed logical blocks remains consistent. This application embodiment does not impose any limitations. The RAID group constructed above is the smallest allocation unit of the storage pool. When the storage service layer requests storage space from the storage pool, the storage pool can provide one or more logical block groups to the storage service layer. Simultaneously, the RAID group includes multiple data stripes, and each data stripe includes multiple data stripe units. The storage space of a data stripe unit in one of the multiple data stripes comes from at least two different types of flash drives, namely ZNS SSDs and Block SSDs. The storage space of data stripe units in other data stripes can come from one type of flash drive, such as either a ZNS SSD or a Block SSD.
[0063] In practice, within the same RAID group, a flash drive typically participates in the construction of only one logical block. This means that the storage space in each logical block comes from different flash drives. For example, referring to the RAID group construction process above, logical block 0's storage space comes from ZNS SSD1 and Block SSD1, logical block 1's storage space comes from ZNS SSD2 and Block SSD2, logical block 2's storage space comes from ZNS SSD3 and Block SSD3, logical block 3's storage space comes from ZNS SSD4, logical block 4's storage space comes from ZNS SSD5, and logical block 5's storage space comes from ZNS SSD6. Each logical block's storage space comes from a different flash drive. On the other hand, in the embodiments of this application, the storage space of different logical blocks can be from a single source or from multiple sources. In other words, the storage space of logical block 0 comes from ZNS SSD and Block SSD, while the storage space of logical blocks 3, 4 and 5 comes from ZNS SSD. In some possible implementations, the storage space of logical blocks can also come from Block SSD.
[0064] In this embodiment, a first logical block is constructed using two flash drives with different granularities, ensuring that the capacity of the first logical block is the same as that of other logical blocks within the RAID group. This means that only one type of logical block capacity needs to be recorded in the RAID group, reducing metadata management complexity. Furthermore, the RAID group can include multiple data stripes. Since the logical blocks within the RAID group have the same capacity, each data stripe in the RAID group contains the same number of data stripe units. Therefore, only one data stripe length and one parity stripe storage location need to be recorded in the RAID group, further simplifying metadata management.
[0065] Furthermore, if the logical blocks in a RAID group have the same capacity, then logical blocks of the same capacity can ensure that the calculation mode is the same when calculating parity data, thus facilitating acceleration using software or hardware. Also, with logical blocks of the same capacity in a RAID group, each logical block can participate in data striping, thereby effectively utilizing the storage space of the logical blocks in the RAID group and avoiding space waste.
[0066] In the above description, the first granularity is a page, and the second granularity is a partition. Of course, in some possible implementations, both the first and second granularities can be partitions, but the partition at the first granularity is much smaller than the partition at the second granularity. For example, the partition size of flash drive 702 is 8MB, while the partition size of flash drive 703 is 120MB. It should also be noted that the first and second granularities do not only refer to absolute values in the specific implementation. For example, 8KB and 80MB mentioned above can also refer to page-level and partition-level values, respectively. For example, page-level values include 4KB, 8KB, or 12KB, while partition-level values include 80MB, 100MB, or 120MB.
[0067] The above describes the process of building a RAID group. After the RAID group is built, the following section describes its usage. Please refer to [link / reference]. Figure 11 The RAID group is located in disk enclosure 1100, which includes flash drives 1101, including ZNS SSDs and Block SSDs. The disk enclosure also includes:
[0068] Control unit 1102 is communicatively connected to flash drive 1101, and is used to perform the above-mentioned tasks. Figure 10 The diagram illustrates the construction structure of a RAID group. After constructing the RAID group, a mapping relationship needs to be established between logical block 0, ZNS SSD1, and Block SSD1, as detailed in Table 1. In the implementation process, the control unit 1101 determines the source of storage space for each logical block based on the acquired information about the logical blocks to be constructed, such as the size of the logical block being 5MB and the number of logical blocks being 6, as well as its own record of the available storage space for each flash drive. To ensure that there is always sufficient available space in the storage system 700 for creating the RAID group, the control unit 1102 can monitor the available storage space of each flash drive in real time, thereby obtaining the available storage space of the entire storage system 700.
[0069] The control unit 1102 can have various forms:
[0070] Configuration 1: Generally, the control unit 1102 includes a central processing unit (CPU) and memory. The CPU is used to perform address translation and read / write data operations. The memory is used to temporarily store data to be written to the flash drive, or data to be read from the flash drive and sent to the host.
[0071] In this context, RAM refers to internal memory that directly exchanges data with the processor. It can read and write data at any time at high speed, serving as temporary data storage for the operating system or other running programs. RAM includes at least two types of memory, such as random access memory (RAM) or read-only memory (ROM). For example, RAM can be Dynamic Random Access Memory (DRAM) or Storage Class Memory (SCM). DRAM is a semiconductor memory, and like most RAM, it is a volatile memory device. SCM is a hybrid storage technology that combines the characteristics of traditional storage devices and RAM. SCM offers faster read and write speeds than hard drives, but slower access speeds than DRAM, and is also cheaper. However, DRAM and SCM are merely illustrative examples in this embodiment; RAM can also include other types of RAM, such as Static Random Access Memory (SRAM). For read-only memory, examples include programmable read-only memory (PROM) and erasable programmable read-only memory (EPROM). Additionally, memory can also be a dual in-line memory module (DIMM), i.e., a module composed of dynamic random access memory (DRAM), or a solid-state drive (SSD). In practical applications, the control unit can be configured with multiple memory modules and different types of memory. This embodiment does not limit the number or type of memory. Furthermore, the memory can be configured to have a power-saving function. A power-saving function means that when the system experiences a power outage and is then powered on again, the data stored in the memory will not be lost. Memory with a power-saving function is called non-volatile memory.
[0072] Type 2: The control unit 1102 is a programmable electronic component, such as a Data Processing Unit (DPU). The DPU possesses the versatility and programmability of a CPU, but is more specialized, capable of efficiently operating on network packets, storage requests, or analysis requests. The DPU differs from the CPU through a high degree of parallelism (handling a large number of requests). In some possible implementations, the DPU can also be replaced by a Graphics Processing Unit (GPU), an embedded Neural-Network Processing Unit (NPU), or other processing chips. Typically, the number of control units 1102 can be one, two, or more. When the storage system 700 contains at least two control units 1102, the two control units 1102 can serve as backups for each other, preventing the failure of one control unit 1102 from rendering all flash drives under that control unit unusable. When the number of control units 1102 is two or more, there can be a hierarchical relationship between the flash drives and the control units 1102, meaning each control unit 1102 can only access the flash drives belonging to it.
[0073] Format 3: The functions of the control unit 1102 can be offloaded to the network interface card (NIC). In other words, in this implementation, the storage system 700 does not have a control unit; instead, the NIC performs data reading / writing, address translation, and other computational functions. In this case, the NIC is a smart NIC. It can include a CPU and memory. In some applications, the NIC may also have persistent memory media, such as persistent memory (PM), non-volatile random access memory (NVRAM), or phase-change memory (PCM). The CPU performs address translation and data reading / writing operations. There is no ownership relationship between the NIC and flash drives in the storage system; the NIC can access any one of the multiple flash drives.
[0074] In some possible implementations, the storage device that provides storage space to the logical block at the first granularity is, in addition to the flash drive that supports the block interface protocol as described above, the storage medium can also be storage class memory (SCM), magnetic random access memory (MRAM), or HDD, or other hard disks that can support the block interface protocol. There are no restrictions here.
[0075] Furthermore, the control unit 1102 is also used to receive a write request, the request carrying data, the data having a first logical block address.
[0076] RAID group 701 may contain one or more stripes. The data fragments and parity fragments contained within a stripe can both be referred to as stripe units. In this embodiment, a stripe unit size of 1MB is used as an example, but it is not limited to 1MB. Continuing with the above example, the created logical block group includes logical blocks 0-5, where logical blocks 0, 1, 2, and 3 are data block groups, and logical blocks 4 and 5 are parity block groups. In the specific implementation process, if the received data cannot fill a data strip, the received data can be temporarily stored in memory. When the data stored in memory reaches a certain value, such as 8MB, the data is divided into two groups of data fragments. Each group includes four data fragments (group 1 includes data fragments 00, 01, 02, and 03; group 2 includes data fragments 10, 11, 12, and 13). The size of each data fragment is 1MB. Then, the check fragments for each group of data fragments are calculated. Two check fragments are calculated for each group (the check fragments for group 1 are P00 and Q00; the check fragments for group 2 are P10 and Q10). The size of each check fragment is also 1MB. The data in data fragments 00 and 10 are the data carried in the receiving request.
[0077] Before sending data to the flash drive, the control unit 1102 needs to determine whether there is an allocated logical block group. If so, and the logical block group still has enough space to accommodate the data, the control unit 1102 can instruct the flash drive to write the data into the allocated logical block group. In a specific implementation, taking data fragment 00 and data fragment 10 as examples, since the data carries the address of the first logical block, the control unit 1102 first determines the logical block 0 corresponding to the first logical block address based on the address of the first logical block (e.g., LBA200-LBA209). According to the mapping relationship in Table 1 above, the control unit 1102 can determine the flash drive 702 and flash drive 703 corresponding to logic block 0, and then store data fragment 00 and data fragment 10 into the pages included in the first physical erase block of flash drive 702 and the first physical partition included in flash drive 703, respectively. Here, the first physical erase block can be any physical erase block in flash drive 702 that can be written to, and the first physical partition can be any partition in flash drive 703 that can be written to.
[0078] Before the control unit 1102 writes data fragment 00 and data fragment 10 to the first physical erase block and the first physical partition respectively, the control unit 1102 also sends a first write command to the flash drive 703. The first write command carries logical block addresses LBA200-LBA204 (first sub-logical block address), data fragment 00, and the length of data fragment 00 (i.e., first data). After ZNS SSD1 receives the first write command, ZNS SSD1 determines the first physical partition where data can be written based on the logical block addresses LBA200-LBA204, data fragment 00, and the data length of data fragment 00 carried in the first write command, and then writes data fragment 00 to the first physical partition. As mentioned in the description above, only sequential writes are supported in the write operation within the partition. The LBAs within a single partition are distributed continuously, and the write pointer (WP) always points to the next LAB position for sequential writing. In order to remember where this "sequential write" has been written to, if repeated writing is required within a single partition, a partition reset operation is required first.
[0079] Each partition has a series of states, which are used for decision-making. All partitions are in an Empty state before use. If writing is required, the partition must be changed from the Empty state to the Open state. Each partition has a capacity limit. Once the write volume reaches the partition's capacity limit, the partition will be in a Full state. If the ZNS SSD has a maximum number of partitions, and the number of Open partitions reaches the limit, and you want to open a new partition, one of these limited partitions needs to be switched to the Closed state. A Closed partition can still be written to, but it must first enter the Open state.
[0080] After the first data is written, a mapping relationship is established between the logical block addresses of the logical partition and the physical block addresses of the physical partition. This mapping relationship is the logical block address corresponding to the physical partition identifier and the offset within the physical partition. Taking logical block 0 as an example, a mapping relationship is established between logical block addresses LBA200-LBA204 in logical block 0's logical block addresses LBA200-LBA209 and the first physical partition 01 and the offset within the first physical partition 01. For details, please refer to Table 2 below.
[0081] Table 2
[0082] Logical block address Physical partition identifier and offset within the physical partition LBA200-LBA204 01: PAB0-PAB4
[0083] In this embodiment, the control unit 1102 also sends a second write command to the flash drive Block SSD1. The second write command carries logical block addresses LBA205-LBA209 (second sub-logical block addresses), data fragment 10 (second data), and the length of data fragment 10. After receiving the second write command, Block SSD1 determines the first physical erase block that can be written to based on the length of data fragment 10 carried in the second write command, and then stores data fragment 10 in the page included in the first physical erase block. After the second data is written, Block SSD1 needs to establish a mapping relationship between the physical block address of the page where data fragment 10 is located and the logical block addresses LBA205-LBA209, as detailed in Table 3 below. This mapping relationship will be added to (first write) or changed (overwrite write) the FTL. With such a mapping, whenever a certain data is read, the SSD first looks up the PBA corresponding to the LBA of the data in the FTL, and then reads the corresponding data according to the PBA.
[0084] Table 3
[0085] Logical block address Physical block address LAB205 PBA1 LBA206 PBA4 LBA207 PBA7 LBA208 PBA3 LBA209 PBA9
[0086] The above describes the process of writing data to the flash drive; the following describes the process of reading data from the flash drive. After receiving a read request, the control unit 1102, based on the logical block addresses LBA200-LBA209 carried in the read request and the correspondence in Table 1 above, determines that logical block addresses LBA200-LBA204 correspond to ZNS SSD1. It then sends a first read instruction to ZNS SSD1, which carries the logical block addresses LBA200-LBA204. Upon receiving the first read instruction, ZNS SSD1, based on the mapping relationship in Table 2 above, determines the physical partition identifier corresponding to logical block addresses LBA200-LBA204, as well as the start and end physical block addresses within the physical partition. Based on the physical partition identifier, the start and end physical addresses, it reads data with an offset corresponding to the specified length from the physical partition corresponding to the physical partition identifier and sends it to the host.
[0087] Correspondingly, the control unit 1102 also determines the Block SSD1 corresponding to the logical block address LBA205-LBA209 based on the logical block address LBA205-LBA209, and then sends a second read instruction to the flash drive Block SSD1. The second read instruction carries the logical block address LBA205-LBA209. After receiving the second read instruction, Block SSD1 maps the logical block address LBA205-LBA209 carried in the second read instruction to a physical block address based on the mapping relationship in Table 3 above, and then reads the corresponding data according to the physical block address and sends it to the host.
[0088] Continuing with the example above, during the data reading process, if the data in data shard 00, data shard 01, data shard 03, and check shards 00 and 01 can all be read normally, but the flash drive containing the data in data shard 02 malfunctions and cannot be read normally, in this case, data shards 00, 01, 02, and 03 can be reconstructed from the data in data shards 00, 01, 03, and check shards 00 and 01. This ensures that the damaged data shards can be reconstructed, thereby significantly improving data reliability.
[0089] Users access data through applications, and the computers running these applications are typically referred to as "application servers" or "hosts." Therefore, please see... Figure 12 , Figure 11The storage system shown further includes hosts, and in specific implementations, the number of hosts can be one or more. Hosts can be physical machines or virtual machines. Hosts include, but are not limited to, desktop computers, servers, laptops, and mobile devices. Hosts access the storage system to retrieve data via a fiber optic switch. However, the switch is only an optional device; application servers can also communicate directly with the storage system via the network. Here, the network can refer to a Local Area Network (LAN), which can be implemented using various structures, devices, and protocols. For example, LAN structures can include Ethernet, wireless, etc. Data communication protocols used in a LAN can include Transmission Control Protocol (TCP), User Datagram Protocol (UDP), Internet Protocol (IP), Hypertext Transfer Protocol (HTTP), Wireless Access Protocol (WAP), Handheld Device Transport Protocol (HDTP), Session Initiation Protocol (SIP), etc.; or the fiber optic switch can be replaced with an Ethernet switch, InfiniBand switch, RoCE (RDMA over Converged Ethernet) switch, etc. Figure 12 In the storage system shown, the above Figure 10 The RAID group construction process shown can be completed by the host.
[0090] In some possible implementations, the RAID group is located in the disk enclosure, and the storage system also includes:
[0091] The controller is connected to the disk enclosure and is used to execute the construction of the RAID group. After the RAID group is constructed, the controller establishes the mapping relationship between the first logical block, the first physical partition, and the first physical erase block.
[0092] Please see below. Figure 13A controller can be a component included in the engine. Taking an engine with two controllers as an example, controller 0 and controller 1 have a mirror channel. When controller 0 writes data to its memory, it can send a copy of the data to controller 1 via the mirror channel. Controller 1 then stores the copy in its local memory. Thus, controller 0 and controller 1 act as backups for each other. When controller 0 fails, controller 1 can take over its operations, and vice versa, preventing hardware failures from causing the entire storage system to become unavailable. When the engine has four controllers, any two controllers have a mirror channel, thus any two controllers act as backups for each other.
[0093] In terms of hardware, the controller includes at least a processor and memory. The processor is a CPU used to process data access requests from outside the storage system, as well as requests generated within the storage system. For example, when the processor receives write data requests from the host through the front-end port, it temporarily stores the data in these requests in memory. When the total amount of data in memory reaches a certain threshold, the processor sends the data stored in memory to the flash drive for persistent storage through the back-end port.
[0094] The engine also includes front-end and back-end interfaces. The front-end interface communicates with the host to provide storage services. The back-end interface communicates with the flash drives. In practice, these multiple flash drives can be in the form of hard drive enclosures, communicating with the engine through the back-end ports. The back-end interfaces exist within the engine as adapter cards, and an engine can simultaneously use two or more back-end interfaces to connect multiple hard drive enclosures. Alternatively, the adapter card can be integrated onto the motherboard, in which case it can communicate with the processor via the PCIe bus.
[0095] Depending on the communication protocol between the engine and the disk enclosure, the disk enclosure may be a Serial Attached SCSI (SAS) disk enclosure, a Non-Volatile Memory Express (NVMe) disk enclosure, an Internet Protocol (IP) disk enclosure, or other types of disk enclosures. SAS disk enclosures use the SAS 3.0 protocol, and each enclosure supports 25 SAS disks. The application server connects to the disk enclosure via an onboard SAS interface or a SAS interface module. NVMe disk enclosures are more like a complete computer system, with NVMe disks plugged into the NVMe disk enclosure. The NVMe disk enclosure then connects to the application server via a RAMA port.
[0096] In some possible implementations, the engine can have hard drive slots, and flash drives can be deployed directly on the engine. In this case, the engine can also have a back-end port, through which a hard drive enclosure can be connected when the flash drive space is insufficient, in order to achieve the purpose of expansion.
[0097] For further information, please see [link / reference]. Figure 14 , Figure 13 The storage system also includes a host computer, which is communicatively connected to the front-end port of the engine. Figure 14 In the storage system shown, the above Figure 10 The RAID group construction process shown can be performed by the host.
[0098] In some possible implementations, Figure 7 Based on the storage system shown, the storage system further includes: one or more hosts, the hosts of which can be found in [reference needed]. Figure 12 The description of the host in this scenario will not be repeated here. In this scenario, the communication connection between multiple flash drives (ZNS SSD1-ZNS SSD6 and Block SSD1-Block SSD1) and the host includes, but is not limited to, the following methods, which will be described below.
[0099] Method 1, please refer to Figure 15 Multiple flash drives can be directly plugged into the host's slots. The slot interface type can be Serial ATA (Serial Advanced Technology Attachment, SATA), SAS, PCIe, or NVMe, etc. The SAS interface adds SCSI technology to the SATA interface, which is mainly used to improve the stability and security of data transmission.
[0100] In some possible implementations, the flash drive can be placed in a server rack containing a hard drive tray. The tray can have a plastic or metal frame and include multiple hard drive slots—specifically, 4, 8, 16, or 32, or more. The tray includes an interface for communication with the host, such as via a fiber optic switch, Ethernet switch, InfiniBand switch, or RoCE (RDMA over Converged Ethernet) switch. Alternatively, communication with the host can be via a network, specifically a Local Area Network (LAN). LANs can be implemented using various architectures, devices, and protocols, such as Ethernet and wireless. Data communication protocols used in LANs can include Transmission Control Protocol (TCP), User Datagram Protocol (UDP), Internet Protocol (IP), Hypertext Transfer Protocol (HTTP), Wireless Access Protocol (WAP), Handheld Device Transport Protocol (HDTP), and Session Initiation Protocol (SIP), among others. In practice, several hard drives can be located in the same rack or in different racks, distributed across various locations, and remotely connected to the host via a gateway or router.
[0101] In this implementation, the RAID group construction process can be performed by the host, or it can be offloaded to another dedicated computer device. This dedicated computer can be a device designed or programmed solely to perform the RAID group construction function, such as a DPU, NPU, or smart network interface card, etc. No specific limitations are imposed in this embodiment. The host's RAID group construction process has already been described in detail above and will not be repeated here. It should be further noted that when constructing a RAID group, a single flash drive can only participate in the construction of one logical block within the same RAID group. This ensures that if one or two flash drives fail, the original data can be recovered from the data stored on the remaining flash drives, thus preventing data loss.
[0102] After the RAID group is built, only one logical block in the RAID group needs to have its storage space composed of ZNS SSD storage space and Block SSD storage space. It is not required that the storage space of all logical blocks be multi-source, that is, all of them need to be composed of ZNS SSD storage space and Block SSD storage space.
[0103] In method two, the ZNS SSDs among the multiple flash drives communicate with the host via the same connection method as in method one, while the Block SSDs are interconnected with the host via a drive enclosure. (See also...) Figure 16 Depending on the type of communication protocol between the host and the disk enclosure, the disk enclosure may be a SAS disk enclosure, an NVMe disk enclosure, an IP disk enclosure, or other types of disk enclosures.
[0104] When the second hard drive group communicates with the host via a disk enclosure, the disk enclosure may include one or more control units. A description of these control units can be found in [link to relevant documentation]. Figure 11 The control unit 1102 shown is described in detail here, and will not be repeated. It should be noted that in this scenario, the user mainly uses ZNS SSDs to build RAID groups, while Block SSDs are only used to fill in the capacity of logical blocks and do not need to be particularly large. In order to save costs and ensure the independence between hard drives, small-capacity pluggable storage blocks can be used to replace ordinary Block SSDs.
[0105] Among some possible implementations, Figure 11 The storage system shown can also be applied in distributed scenarios. Please refer to [link / reference]. Figure 17In this scenario, there are compute node clusters and storage node clusters. A compute node cluster includes one or more compute nodes that can communicate with each other. A compute node is a computing device, such as a server, desktop computer, or the controller of a storage array. Hardware-wise, a compute node includes at least a processor, memory, and a network interface card (NIC). The processor is a central processing unit (CPU) used to process data access requests from outside the compute node or requests generated internally within the compute node. For example, when the processor receives a write data request from a user, it temporarily stores the data in the write data request in memory. When the total amount of data in memory reaches a certain threshold, the processor sends the data stored in memory to the storage node for persistent storage. In addition, the processor is also used for data computation or processing, such as metadata management, deduplication, data compression, virtualization of storage space, and address translation. In practical applications, there are often multiple CPUs, and each CPU has one or more CPU cores. This application does not limit the number of CPUs or the number of CPU cores. For a description of memory, please refer to [link to relevant documentation]. Figure 9 The memory shown will not be described in detail here.
[0106] Any compute node can access any storage node in the storage node cluster via the network. The storage node cluster consists of multiple storage nodes, and these storage nodes can be... Figure 11 The storage system shown illustrates this. In a distributed scenario, the flash drives used to build a RAID group can come from different flash drives within a single storage node, or they can come from different storage nodes. When the flash drives in a RAID group come from different storage nodes, if one storage node fails, data can still be recovered using the data stored on the flash drives in the remaining healthy storage nodes, thus improving the reliability of the storage system.
[0107] Furthermore, in this embodiment, different types of flash drives can be used to construct RAID groups. In other words, the storage system provided in this embodiment is compatible with different types of flash drives. Thus, if a flash drive in the storage system fails, any type of flash drive can be used as a replacement. For example, if a ZNS SSD in the storage system fails, it can be replaced with a ZNS SSD or a Block SSD without affecting the construction of logical blocks, thereby improving the system's compatibility.
[0108] Among some possible implementations, Figure 11 The storage system shown can also be applied to in-store computing scenarios in distributed environments. Please refer to [link / reference]. Figure 18In in-memory computing scenarios, Figure 11 The control unit 1102 shown can be a processor included in a server or desktop computer. A server is a computing power unit; for example, an Advanced Reduced Instruction Set Machine (ARM) server or an x86 server can be used as the server. In terms of hardware, a server includes other components besides the processor, such as memory, a network interface card (NIC), and a hard drive. The processor, memory, NIC, and hard drive are connected via a bus. The processor and memory provide computing resources. Specifically, the processor is a central processing unit used to handle data access requests from outside the server and also to handle requests generated internally by the server. For example, when the processor receives a write data request, it temporarily stores the data in the write data request in memory. When the total amount of data in memory reaches a certain threshold, the processor sends the data stored in memory to the hard drive for persistent storage. Here, "persistent storage" refers to the ability of a flash drive to retain the recorded data after power failure. In addition, the processor is also used for data computation or processing, such as metadata management, deduplication, data compression, data verification, virtualization of storage space, and address translation. In practical applications, there can often be multiple CPUs, with each CPU having one or more CPU cores. This application does not limit the number of CPUs or the number of CPU cores. For an introduction to memory, please refer to... Figure 1 The details of the memory included will not be repeated here.
[0109] In this embodiment, the server in this application scenario can be installed in a server rack, which can have multiple slots. The number of slots can be 4, 8, 16, 32, or other suitable numbers, with each slot accommodating one server. In specific implementations, the storage system is scalable. In some possible implementations, servers can be inserted into or removed from a server rack, depending on the actual situation. The storage capacity of each server can be any integer multiple of 4TB, such as 8TB, 12TB, 16TB, 32TB, etc.
[0110] Secondly, embodiments of this application also provide a data storage method, which can be applied to the first aspect. Figure 7 The storage system shown herein, specifically the data storage method, can be referred to in the first aspect as the process by which the control unit 1102 receives a write request and writes data to the flash drive. To avoid redundancy, it will not be described further here. It should be noted that, depending on the specific architecture of the storage system, the executing entity of this data storage method may differ slightly. In the actual implementation, the executing entity of this data storage method can be the entity described in the first aspect. Figure 12The control unit shown can also be Figure 12 The host shown can, of course, also be Figure 13 The engine shown.
[0111] Thirdly, embodiments of this application also provide a computer-readable medium including instructions that, when executed on a computer, cause the computer to perform the data storage method as described in the second aspect above.
[0112] Fourthly, embodiments of this application also provide a computer program product that, when run on a computer, causes the computer to execute the data storage method as described in the second aspect above.
[0113] The methods provided in this application can be implemented entirely or partially through software, hardware, firmware, or any combination thereof. When implemented in software, they can be implemented entirely or partially in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of this application are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, a network device, a user equipment, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., coaxial cable, fiber optic, digital subscriber line) or wireless means. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that integrates one or more available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., digital video disc (DVD)), or a semiconductor medium, etc.
[0114] After introducing a storage system provided by the embodiments of this application, the application scenarios of the storage system provided by the embodiments of this application will be introduced, including but not limited to the following:
[0115] The first type: The storage system provided in this application embodiment can support artificial intelligence applications, machine learning applications, big data analytics applications, and many other types of applications. The rapid growth of such applications is driven by three technologies: deep learning (DL), graphics processing units (GPUs), and big data. Deep learning is a computational model that utilizes massively parallel neural networks inspired by the human brain. GPUs are modern processors with thousands of cores, well-suited for running algorithms that roughly match the parallelism of the human brain.
[0116] The second approach: The storage system provided in this application embodiment can be used in a neuromorphic computing environment. Neuromorphic computing is a form of computing that mimics brain cells. To support neuromorphic computing, the architecture of interconnected "neurons" replaces the traditional computing model with low-power signals transmitted directly between neurons, achieving more efficient computation.
[0117] Thirdly, the storage system provided in this application embodiment can also be configured to support the storage or use of blockchain. Such a blockchain can be represented as a continuously growing list of records, called blocks, which are linked and protected using cryptography. Each block in the blockchain may contain a hash pointer as a link to the previous block, a timestamp, transaction data, etc. This structure makes data modification and tampering extremely difficult. With the continuous development of technology, blockchain has been widely applied in finance, logistics, healthcare, public services, and other fields, providing new solutions for data security and trustworthiness.
[0118] Of course, the storage system provided in this application can also be applied to other application scenarios, such as big data analysis, edge computing, etc., which are not specifically limited in the embodiments of this application.
[0119] In summary, the above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions within the technical scope disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
Claims
1. A storage system, characterized in that, include: Independent disk redundant array RAID group; The RAID group comprises multiple logical blocks, each with the same capacity; The storage space of the first logical block among the plurality of logical blocks comes from at least the first flash drive and the second flash drive; Wherein, the first flash drive provides storage space for the first logical block at a first granularity, and the second flash drive provides storage space for the first logical block at a second granularity, wherein the first granularity and the second granularity are different.
2. The storage system according to claim 1, characterized in that, The first granularity is a page, wherein the first flash drive includes multiple erase blocks, and each erase block includes multiple pages.
3. The storage system according to claim 1 or 2, characterized in that, The second granularity is a zone, wherein the storage space of the second flash drive is divided into several zones, and a zone includes one or more erase blocks.
4. The storage system according to claim 3, characterized in that, The number of the second flash drives is at least two, wherein the partitions contained in the at least two second flash drives are of different sizes.
5. The storage system according to claim 1, characterized in that, The first flash drive and the second flash drive are physically independent flash drives.
6. The storage system according to claim 1, characterized in that, The first flash drive and the second flash drive are different parts of a hybrid flash drive; wherein, The first flash drive is configured with a first portion of flash memory chips included in the hybrid flash drive; The second flash drive is configured with a second portion of flash memory chips included in the hybrid flash drive.
7. The storage system according to claim 3, characterized in that, The RAID group is located in a disk enclosure, which also includes: The control unit is used to establish a first mapping relationship between the first logic block and the first flash drive and the second flash drive.
8. The storage system according to claim 7, characterized in that, The control unit is also used for: Receive a write request, the write request carrying data and the address of the first logical block pointing to the first logical block; Based on the first mapping relationship, the first flash drive and the second flash drive corresponding to the first logical block address are determined; The data is stored in the pages included in the first erase block of the first flash drive and the first partition of the second flash drive, respectively.
9. The storage system according to claim 8, characterized in that, The control unit is also used for: Send a first write instruction to the first flash drive. The first write instruction carries the address of the first sub-logical block and the length of the first data. A second write instruction is sent to the second flash drive. The second write instruction carries the address of a second sub-logical block and the length of second data. The data includes the first data and the second data, and the first logical block address includes the address of the first sub-logical block and the address of the second sub-logical block.
10. The storage system according to claim 3, characterized in that, The RAID group is located in the disk enclosure, and the storage system also includes: The controller is communicatively connected to the disk frame; The controller is used to establish a first mapping relationship between the first logic block and the first flash drive and the second flash drive.
11. The storage system according to any one of claims 1-5, characterized in that, The first flash drive is a flash drive that supports the ZNS partition namespace interface protocol, and the second flash drive is a flash drive that supports the block interface protocol.
12. A data storage method, characterized in that, Applied to a storage system, the storage system including a Redundant Array of Independent Disks (RAID) group, the RAID group comprising multiple logical blocks of equal capacity, the method comprising: Obtain a write request, which carries data; The data is stored in a first flash drive and a second flash drive corresponding to a first logical block among the plurality of logical blocks. The first flash drive provides storage space for the first logical block with a first granularity, and the second flash drive provides storage space for the first logical block with a second granularity. The first granularity and the second granularity are different.
13. The method according to claim 12, characterized in that, The first flash drive includes multiple erase blocks, and each erase block includes multiple pages; The storage space of the second flash drive is divided into several zones, and each zone includes one or more erase blocks; The first granularity is a page, and the second granularity is a partition.
14. The method according to claim 13, characterized in that, The request carries the address of a first logical block pointing to the first logical block, and storing the data into the first flash drive and the second flash drive corresponding to the first logical block among the plurality of logical blocks includes: Based on the mapping relationship between the first logical block and the first and second flash drives, determine the first and second flash drives corresponding to the addresses of the first logical block; The data is stored in the first partition of the second flash drive and the pages included in the first erase block of the first flash drive, respectively.
15. The method according to claim 14, characterized in that, The method further includes: Before storing the data into the first partition of the second flash drive and the pages included in the first erase block of the first flash drive, a first write instruction is sent to the first flash drive, the first write instruction carrying the address of the first sub-logical block and the length of the first data; A second write instruction is sent to the second flash drive. The second write instruction carries the address of a second sub-logical block and the length of second data. The data includes the first data and the second data, and the first logical block address includes the address of the first sub-logical block and the address of the second sub-logical block.
16. A computer-readable storage medium, characterized in that, Includes instructions that, when executed on a computer, cause the computer to perform the method as described in any one of claims 12-15 above.
17. A computer program product, characterized in that, When it is run on a computer, it causes the computer to perform the method as described in any one of claims 12-15 above.
Citation Information
Patent Citations
Dynamically resizing logical storage blocks
CN109117084A
Sequential data optimized sub-regions in storage devices
CN113396383A
Providing Raid-10 with a configurable Raid width using a mapped raid group
US10229022B1
Using flash storage devices with different sized erase blocks
US10496330B1
Memory system having an unequal number of memory die
US20140189210A1