Redundant array management method and system based on ZNS solid state disk

By using ZRWA features on ZNS SSD for on-site update of verification blocks, the problem of difficulty in updating verification blocks on ZNS solid-state drives is solved, and efficient write performance and low latency control are achieved.

CN119960673AActive Publication Date: 2025-05-09TSINGHUA UNIVERSITY

Patent Information

Application Number
CN202411955134.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-27
Publication Date
2025-05-09
Estimated Expiration
2044-12-27

AI Technical Summary

Technical Problem

In the prior art, it is difficult to update the verification blocks on ZNS solid-state drives in situ, resulting in low throughput and slow recovery speed.

Method used

By leveraging the Zone Random Write Area (ZRWA) feature of ZNS SSD, the on-site update of the verification block is realized, that is, data overwrite directly at the original verification block location without additional write operations.

Benefits of technology

This method not only reduces the number of write operations, but also effectively utilizes logical area resources, improves write performance and reduces latency, while reducing the need for garbage collection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119960673A_ABST
    Figure CN119960673A_ABST
Patent Text Reader

Abstract

The invention provides a redundant array management method and system based on a ZNS solid state disk, and the method comprises the steps: sequentially writing to-be-written data into a corresponding data block for each strip under the condition of receiving a write request, and carrying out the coverage updating of verification data of a verification block; positioning a logic actual write pointer to the tail position of the data block in which the data is continuously and successfully written, generating metadata corresponding to each data block, and writing the metadata into a standby area; the logic actual writing pointer is used for pointing to a next writable position in the logic area; under the condition that a fault occurs, repositioning the logic actual write pointer based on the metadata in the standby area; and on the basis of repositioning of the logic actual write pointer after the fault, discarding the data blocks outside the logic actual write pointer, recovering the data blocks inside the logic actual write pointer, and recalculating to generate verification data, thereby completing redundant array management of the ZNS solid state disk.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of solid state drive management, and in particular to a redundant array management method and system based on a ZNS solid state drive. Background Art

[0002] ZNS is a new type of solid-state drive device. Compared with existing commercial solid-state drives, it has the advantages of low write amplification, controllable garbage collection, and stable performance. ZNS has the following three characteristics. 1) ZNS abstracts the flash memory resources in the SSD into multiple zones, and a single zone only supports appending. Each zone maintains a write pointer that points to the end of the data currently written in the zone. The data of each write operation only supports appending after the write pointer, and the data before any write pointer in the zone can be read. A single zone also provides a reset interface, which supports erasing all data in the zone and placing the write pointer at the start of the zone. 2) ZNS has multiple zones, and users can run applications in multiple zones exclusively. The zones between applications are not shared with each other, and performance isolation can be achieved. 3) ZNS provides both write and append interfaces for appending. The write interface requires the host to specify the address to be written, and the address to be written must be the address pointed to by the write pointer of the current zone; when executing the append interface, only the zone to which the request is sent needs to be specified. After writing the data, ZNS will return the address where the data is written.

[0003] The redundant array of independent disks (RAID) mechanism is aimed at the scenario of single server with multiple disks. Through redundant disk resources and checksum calculation, the storage system can still provide data read and write services to users when some disks are damaged. With the widespread application of ZNS and the increasing number of solid-state disks that can be accommodated by a single server, it is necessary to design a dedicated redundant array mechanism for ZNS.

[0004] The prior art RAIZN provides a RAID mechanism that adapts to the zone interface. In order to avoid in-place updates of PartialParity Update (PPU), RAIZN allocates a physical metadata zone to each ZNS SSD to store the parity log entries generated by all logical zones to temporarily store the update data of these partial checks. When all the data of a stripe is written, the check is integrated and written back to the corresponding position of the check block. When the newly opened zone is full, another zone needs to be opened to assist in garbage collection. The introduction of metadata zones solves the problem of in-place updates of check data, but it also brings three major performance problems, resulting in low throughput and slow recovery. 1) Contention for PPU aggregation between multiple zones: As shown in the attached Figure 1 It is shown that when multiple logical zones in RAIZN receive write requests at the same time and generate PPUs for a single metadata zone, there will be contention for the aggregation of multi-zone PPUs. Different applications update data on the same zone, resulting in bandwidth contention for their zones. 2) Garbage collection: In traditional RAID systems, the space ratio of data to parity is k:m (recall that a stripe consists of k data blocks and m parity blocks). Considering the append-write semantics of ZNS, RAIZN uses offline updates to store PPUs in the metadata zone, which introduces additional space overhead. In order to maintain the same parity space ratio as traditional RAID, RAIZN must perform garbage collection (GC) periodically to clean up invalid PPUs. 3) Write amplification problem of log header: If ZNS RAID uses log records to buffer PPUs, log headers are necessary because information such as parity ranges must be persisted on the device to prevent machine failures. The existence of log headers introduces additional space and write bandwidth overhead, especially under small request workloads. Therefore, it is necessary to study a control method that can update in place on ZNS SSDs and ensure certain write performance and low latency. Summary of the invention

[0005] The present invention provides a redundant array management method and system based on a ZNS solid state drive, which is used to solve the problem in the prior art that the check block on the ZNS solid state drive is difficult to update in situ, and can not only efficiently utilize the logic area resources, but also improve the writing performance and reduce the delay.

[0006] The present invention provides a redundant array management method based on a ZNS solid state drive, wherein physical areas belonging to different ZNS solid state drives are combined into corresponding logical areas, wherein the logical areas include multiple stripes, and each stripe includes at least one check block and multiple data blocks belonging to different ZNS solid state drives;

[0007] The method comprises:

[0008] For each stripe, when a write request is received, the data to be written is written into the corresponding data block in sequence, and the check data of the check block is overwritten and updated;

[0009] Positioning the logical actual write pointer to the tail position of the data block to which the data is continuously and successfully written, generating metadata corresponding to each data block, and writing the metadata into the spare area; wherein the logical actual write pointer is used to point to the next writable position in the logical area;

[0010] In the event of a failure, repositioning the logical actual write pointer based on the metadata in the spare area;

[0011] Based on the repositioning of the logical actual write pointer after the failure, the data blocks outside the logical actual write pointer are discarded, the data blocks within the logical actual write pointer are restored, and the verification data is recalculated to complete the redundant array management of the ZNS solid state drive.

[0012] According to the redundant array management method of the ZNS solid state drive provided by the present invention, the method further comprises: upon receiving a read request, directly reading corresponding data according to a data block address in a specified logical area in the read request;

[0013] In the event of a disk failure, data required for a read request is rebuilt based on the data and the check data in the same stripe, so that the failed data can be recovered based on the data and the check data in the same stripe.

[0014] According to the redundant array management method of the ZNS solid state drive provided by the present invention, in the event of a failure, the logical actual write pointer is relocated based on the metadata in the spare area, specifically comprising:

[0015] In the event of a machine failure, query the status information of each physical area in the ZNS solid state drive and determine the logical area status; wherein the logical area status is used to describe the status of the corresponding logical area as full, empty, available, closed, read-only or abnormal;

[0016] Based on the state of the logical area and the metadata in the spare area, the write pointer information of each physical area is determined, and the logical actual write pointer position of the corresponding logical area is relocated.

[0017] According to the redundant array management method of the ZNS solid state drive provided by the present invention, the querying of the status information of each physical area in the ZNS solid state drive and determining the status of the logical area specifically includes:

[0018] By sending a query command to the solid-state drive, receiving a descriptor containing the physical area status and the write pointer of the physical area, determining the status information of each physical area, and determining the corresponding logical area status and logical write pointer position; wherein the status information of each physical area is used to describe the status of the corresponding physical area as full, empty, available, closed, read-only or abnormal.

[0019] According to the redundant array management method of the ZNS solid state drive provided by the present invention, the write pointer information of each physical area is determined based on the state of the logical area and the metadata in the spare area, and the logical actual write pointer position of the corresponding logical area is relocated, specifically including:

[0020] Based on the logical area state and the logical write pointer position, query the metadata in the spare area to determine whether each data block has data written; wherein, if the data block has data written, the corresponding metadata is 1, and if the data block has no data written, the corresponding metadata is 0;

[0021] According to the write pointer information of each data block, the logical actual write pointer is relocated to the tail position of the data block where data is continuously and successfully written.

[0022] According to the redundant array management method of the ZNS solid state drive provided by the present invention, the repositioning of the logical actual write pointer after the failure and discarding the data blocks outside the logical actual write pointer specifically include:

[0023] Based on the repositioning of the logical actual write pointer after the failure, the data in the data block outside the logical actual write pointer is treated as invalid data and overwritten with zeros, so as to discard the invalid data outside the logical actual write pointer.

[0024] According to the redundant array management method of the ZNS solid state drive provided by the present invention, the data block within the logical actual write pointer is recovered, and the verification data is recalculated to generate the redundant array management of the ZNS solid state drive, which specifically includes:

[0025] Based on the recovered data blocks within the logical actual write pointer, the check block data is recalculated and generated to determine the remaining continuous data blocks of the stripe after the invalid data is discarded. When there is a new write request, the remaining continuous data blocks are directly written to complete the redundant array management of the ZNS solid state drive.

[0026] The present invention provides a redundant array management system based on a ZNS solid state drive, wherein physical areas belonging to different ZNS solid state drives are combined into corresponding logical areas, wherein the logical areas include multiple stripes, and each stripe includes at least one check block and multiple data blocks belonging to different ZNS solid state drives. The system includes:

[0027] A data writing module is used for, for each stripe, when receiving a write request, writing the data to be written into the corresponding data block in sequence, and overwriting and updating the check data of the check block;

[0028] A pointer generation module, used to position the logical actual write pointer to the tail position of the data block to which the data is continuously and successfully written, and to generate metadata corresponding to each data block, and to write the metadata into the spare area; wherein the logical actual write pointer is used to point to the next writable position in the logical area;

[0029] A pointer positioning module, used for repositioning the logical actual write pointer based on the metadata in the spare area in the event of a failure;

[0030] The data recovery module is used to discard the data blocks outside the logical actual write pointer based on the repositioning of the logical actual write pointer after the failure, recover the data blocks within the logical actual write pointer, and recalculate and generate verification data to complete the redundant array management of the ZNS solid state drive.

[0031] According to the redundant array management system of the ZNS solid state drive provided by the present invention, the system further includes:

[0032] A data reading module, configured to directly read corresponding data according to a data block address in a specified logical area in the read request when a read request is received;

[0033] In the event of a disk failure, data required for a read request is rebuilt based on the data and the check data in the same stripe, so that the failed data can be recovered based on the data and the check data in the same stripe.

[0034] The present invention also provides an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, a redundant array management method based on a ZNS solid state drive as described above is implemented.

[0035] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, and when the computer program is executed by a processor, the redundant array management method based on a ZNS solid state drive as described in any one of the above is implemented.

[0036] The present invention also provides a computer program product, including a computer program, wherein when the computer program is executed by a processor, the redundant array management method based on a ZNS solid state drive as described above is implemented.

[0037] The present invention provides a redundant array management method and system based on a ZNS solid-state drive. By utilizing the Zone Random Write Area (ZRWA) feature of the ZNS SSD, the in-situ update of the check block is realized, that is, data is directly overwritten at the original check block location without the need for additional write operations. The realization of this overwrite update benefits from the ability of ZRWA to allow random writing within a certain range. When a data block is updated, the related check information will also change accordingly. ZRWA provides a mechanism that allows new check data to be written directly to the original check block location without allocating new storage space. In this way, not only the number of write operations is reduced, the logical area resources can be efficiently utilized, but also the write performance can be improved and the delay can be reduced.

[0038] In addition, the ability to overwrite and update check blocks also helps reduce the need for garbage collection. In traditional ZNS SSD management, garbage collection is a resource-intensive process that involves cleaning up invalid data and reallocating storage space. By reducing the amount of data that requires garbage collection, the method of the present invention further optimizes the performance and efficiency of the storage system.

[0039] Secondly, by storing lightweight metadata in the spare area, the present invention can accurately locate the logical real write pointer (Real WP) after a machine failure. This design not only ensures data consistency, but also avoids the risk of data loss due to write pointer loss.

[0040] Thirdly, in the process of fault recovery, the present invention achieves zero storage overhead recovery by discarding invalid data outside the actual write pointer and overwriting it with zeros. This method avoids the loss of additional storage space and improves the recovery efficiency. BRIEF DESCRIPTION OF THE DRAWINGS

[0041] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.

[0042] Figure 1 This is a schematic diagram of RAZIN opening an additional zone and updating verification data in the prior art.

[0043] Figure 2 It is a schematic diagram of the working principle of ZRWA provided by the present invention.

[0044] Figure 3 It is a schematic diagram of the overall architecture of Zebra provided by the present invention.

[0045] Figure 4 It is one of the flow charts of the redundant array management method based on the ZNS solid state drive provided by the present invention.

[0046] Figure 5 It is a schematic diagram of the execution process of processing a write request provided by the present invention.

[0047] Figure 6 This is the second flow chart of the redundant array management method based on the ZNS solid state drive provided by the present invention.

[0048] Figure 7 It is a schematic diagram of the execution process of processing a write request provided by the present invention.

[0049] Figure 8 It is a schematic diagram illustrating the fault recovery storage overhead provided by the present invention.

[0050] Fig. 9 It is a schematic diagram of module connection of a redundant array management system based on a ZNS solid state drive provided by the present invention.

[0051] Fig.10 It is a structural schematic diagram of the electronic device provided by the present invention. DETAILED DESCRIPTION

[0052] In order to make the purpose, technical solution and advantages of the present invention clearer, the technical solution of the present invention will be clearly and completely described below in conjunction with the drawings of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0053] First, in order to facilitate understanding of the solution of this embodiment, the terms involved in the present invention are schematically explained.

[0054] Physical region: A specific storage space on a single ZNS SSD. These regions are the hardware abstraction of the ZNS SSD, and each physical region represents a zone for storing data blocks and check blocks. The physical region is the underlying storage unit of the Zebra framework and is determined by the physical characteristics of the ZNS SSD.

[0055] Logical area: It is the virtual storage space provided by the Zebra framework to the application. It is composed of physical areas from different ZNS SSDs to achieve uniform distribution of data blocks. The logical area is transparent to the application and provides an interface similar to that of traditional storage devices, so that the application can access the logical area just like accessing traditional storage devices.

[0056] Stripe: In a RAID configuration, a stripe is the basic unit of data distribution. Each stripe includes at least one parity block and multiple data blocks belonging to different ZNS SSDs. The stripe design allows data to be written to multiple physical areas in parallel, thereby improving write performance and data redundancy.

[0057] Check block: A key component used for error detection and correction in RAID systems. It contains check information calculated based on the data block, such as the results of parity, CRC or other check algorithms. During the in-place update of the check block, the new check data will directly overwrite the old check data to reflect the latest status of the data.

[0058] Data block: The basic unit for storing actual user data. In a RAID system, data blocks are distributed to different physical areas to achieve data redundancy and improve read and write performance. The size of the data block can be adjusted according to the RAID configuration and application requirements.

[0059] RAID (Redundant Array of Independent Disks): RAID is a data storage virtualization technology that combines multiple physical disk drives (or solid-state drives SSD) into a logical unit to improve data reliability, fault tolerance, and performance. RAID achieves these advantages by distributing data and verification information across multiple disks. There are multiple levels of RAID, such as RAID0, RAID1, RAID5, RAID6, etc. Each level has its own specific data distribution and verification strategy.

[0060] ZNS: is a new type of solid-state drive (SSD) architecture that divides the storage space of the SSD into multiple zones, each of which can be managed independently. In a ZNS SSD, data is written to a zone in an append-write manner, which means that once data is written to a zone, it cannot be overwritten unless the entire zone is erased. The design of the ZNS SSD helps reduce write amplification, improve garbage collection efficiency, and achieve better performance isolation.

[0061] ZRWA (Zone Random Write Area): For a logical zone, ZRWA is a fixed-size continuous logical block address area (LBA) that can receive random, concurrent overwrite write requests. In ZRWA mode, the write pointer (WP) of a zone points to the starting position of ZRWA. Figure 2Shows how ZRWA works. An application or file system can be aware of a set of logical zones or LBAs. Since the size of a single flash page is usually larger than the size of an LBA, ZNS SSDs usually use a RAM-based write buffer to collect enough LBA data and then write it to the flash page. The write buffer allows repeated overwrite operations, and a zone can expose the write buffer to the host through ZRWA. Commitment of ZRWA can be done implicitly (i.e., data is moved to other LBAs when it exceeds the range of the zone). Figure 2 In the example, assuming that the size of a flash page is 4 LBAs and the size of the write buffer is 8 LBAs, then a write to LBA11 will cause the data of the first 4 LBAs to be flushed to the flash page. After that, overwrite operations will no longer be allowed in the LBA range [0,3]. It should be noted that overwrite operations within ZRWA will not move WP. ZNS SSD uses the ZRWA size field to define the size of the write buffer and the ZRWAFG field to define the minimum number of LBAs to write to the flash page.

[0062] Zebra: is a redundant array management system based on ZNS SSD. Zebra uses the features of ZNS SSD, such as ZRWA (Zone Random Write Area), to implement in-place updates of parity blocks, which is difficult to achieve in traditional RAID configurations. Zebra's design allows for efficient management of data redundancy and fault recovery on ZNS SSDs while improving write performance and reducing latency. Zebra optimizes the functionality of implementing RAID on ZNS SSDs through a lightweight metadata scheme and a zero storage overhead fault-tolerant recovery method.

[0063] Specifically, see Figure 3 The redundant array management system mentioned in the embodiment of the present invention includes:

[0064] Host: A host is a computer system that initiates requests to store, retrieve, or manage data. In a storage system, a host can be a server, a personal computer, a workstation, or any other system that needs to access a storage device. The host is responsible for running applications and operating systems that need to store data and therefore send read and write commands to the storage device.

[0065] Device side: A device refers to a physical storage unit in a storage system, such as a hard disk drive (HDD), solid-state drive (SSD), optical drive, USB flash drive, or other forms of non-volatile storage media. In the context of ZNS SSD, a device refers to an SSD with a partitioned namespace that is configured in ZNS mode to support more efficient data management and storage operations.

[0066] The relationship between the host side and the device side is as follows:

[0067] The host will send data write, read or other management commands to the RAID controller or Zebra system as needed.

[0068] The RAID controller or Zebra system receives these commands and distributes the data across the attached storage devices, which can be multiple physical hard drives or SSDs.

[0069] The device side is responsible for actually storing the data and returning it to the host when needed.

[0070] Logical Zone: The logical zone is the virtual storage space provided by the Zebra system for applications. Each logical zone consists of multiple physical zones (from different ZNS SSDs) and contains multiple stripes.

[0071] Stripe: A stripe is the basic unit of data distribution in RAID and consists of multiple data chunks and a parity chunk.

[0072] Physical Zone: A physical zone is the actual storage area on a ZNS SSD. Each physical zone can contain multiple zones.

[0073] The physical areas belonging to different ZNS solid state drives are combined into corresponding logical areas, and the logical area includes a plurality of stripes.

[0074] Host-side data structures: Similar to existing ZNSRAID systems, Zebra maintains write pointers (WPs) and the state of active logical zones in host memory. In addition, Zebra maintains an in-memory stripe cache (stripecache) on the host side, which contains the stripes of all active logical zones. This design avoids the need to read the ZNS SSD for the parity unit (PPU) calculation and reduces the request dependency of RAID write operations.

[0075] Device-side data layout: In Zebra, all physical zones are opened in ZRWA mode and can store parity blocks in any way. The blocks of the stripe are evenly distributed on all ZNS SSDs, which means that data blocks and parity blocks coexist in each zone. Zebra supports different RAID configurations (such as RAID5 and RAID6).

[0076] Figure 3In FIG, the logical zone Logical Zone 0 includes the physical zone of ZNS 0, the physical zone of ZNS1 and the physical zone of ZNS2. The data blocks D2, D3 and the check block P1 form a stripe.

[0077] When data is appended to the stripe, it causes repeated overwriting of the parity block. For example, when data block D2 is appended, parity block P1 is calculated as D2 XOR 0 (because this is the first write, it is XORed with 0). Subsequently, when data block D3 is appended, parity block P1 is updated to D2 XOR D3.

[0078] In order to realize the ZRWA mechanism of ZNS in the embodiment of the present invention, the problem of in-place update of verification data on ZNS SSD device is effectively solved, which can not only efficiently utilize Zone resources, but also significantly improve write performance and reduce latency. Figure 4 An embodiment of the present invention provides a redundant array management method based on a ZNS solid state drive, wherein for ease of description, the method is performed in each stripe.

[0079] The method comprises:

[0080] 401. For each stripe, when a write request is received, the data to be written is written into the corresponding data block in sequence, and the check data of the check block is overwritten and updated.

[0081] While the data block is being written, the system calculates new parity block data. The parity block calculation is based on RAID algorithms (such as XOR, P+Q, Reed-Solomon, etc.) to ensure data redundancy and error detection capabilities.

[0082] By utilizing the ZRWA feature of ZNS SSD, the update of the check block is performed in situ, that is, the newly calculated check data is directly overwritten on the original check block location without the need for additional write operations.

[0083] After the data block and check block are updated, the logical actual write pointer is updated to point to the next writable location in the logical area. This step ensures the sequential writing of data and the efficient use of ZNS SSD.

[0084] The following is a schematic illustration of the process flow when each stripe receives a write request using a specific example. In a specific example, the process flow of appending writes to some stripes is as follows: Figure 5 The execution process of the Zebra system processing five write requests and the resulting ZRWA movement are shown.

[0085] (1) Initial state: The write pointer (WP) of the logical zone is located at the start of D3 (i.e., the start of the second stripe).

[0086] (2) Process W0. Zebra appends the data of W0 to D3 and writes part of the check data into the first half of P1. Then, the tail of ZRWA of ZNS0 and ZNS3 moves to the middle of D3 and P1 respectively.

[0087] (3) Process W1. Zebra continues to append the data of W1 to D3 and appends part of the check data to P1. At this time, the tail of ZRWA of ZNS0 and ZNS3 moves further.

[0088] (4) Process W2. The amount of data in W2 is equal to the block size, and Zebra appends its data to D4 and implicitly moves ZNS1's ZRWA. At the same time, Zebra writes the checksum update triggered by W2 to P1, but since P1 is within the ZRWA range, ZNS3's ZRWA is not moved.

[0089] (5) and (6) process W3 and W4. Zebra writes the data of W3 and W4 to D5, and overwrites the corresponding two check blocks (PPU) in P1 in sequence. Finally, the tails of all ZRWAs are moved to the end of their corresponding blocks in the stripe.

[0090] 402. Position the logical actual write pointer at the tail position of the data block to which data is continuously and successfully written, generate metadata corresponding to each data block, and write the metadata into the spare area; wherein the logical actual write pointer is used to point to the next writable position in the logical area.

[0091] After receiving a write request and processing a data block write, the system needs to determine the location of the logical real write pointer (Real WP). This is usually located at the end of the data block where the data is written successfully. Once the final location of the written data is determined, the logical real write pointer will be updated to point to the next writable location in the logical area. This location is prepared for subsequent write operations.

[0092] To record write operations and provide information for failure recovery, the system generates metadata for each successfully written data block. The metadata may include the identity of the written data block, the write time, the write size, and the parity block information related to the data block.

[0093] The generated metadata is written to a spare area of ​​the ZNS SSD (such as the OOB area). This area is a portion of each flash page that is dedicated to storing metadata associated with data blocks.

[0094] Ensure that the written metadata remains consistent with the actual written data blocks, which means that if a data block is written, the corresponding metadata should also be updated to reflect this change.

[0095] After the metadata is written, the system will check whether a failure has occurred. If the system is stable, the metadata will be used for future failure recovery processes, such as repositioning the write pointer and recovering data.

[0096] 403. In case of a failure, reposition the logical actual write pointer based on the metadata in the spare area.

[0097] Among them, failures are mainly divided into two categories: disk failure and machine failure. Disk failure occurs when some ZNSSSDs in the RAID cannot be accessed due to firmware problems or end of life. In this case, Zebra uses redundant parity data to recover the lost data on the failed device. The other is a machine failure, usually caused by a power outage event. Assuming that the data in the host-side DRAM is lost, Zebra must ensure that the write request that has been marked as completed to the application is not lost. In addition, Zebra also needs to recover the request prefix part that has been successfully written but not yet completed.

[0098] In the event of a disk failure, Zebra uses a general recovery method for disk failure recovery. For write requests sent to the failed ZNS SSD, they can be skipped directly because no response can be obtained; while read requests are completed through "reconstructed reads", that is, the data required for the read request is rebuilt from the data and check data in the same stripe, resulting in increased read traffic. When a new ZNS SSD is ready to replace the failed device, Zebra starts the recovery process. Specifically, Zebra first reads the data of all stripes and calculates the check data on the failed ZNS SSD, and then writes these check data and data to the newly added ZNS SSD.

[0099] In the event of a machine failure, the data structures stored in the host memory are lost. Zebra must recover these data structures from the device after a failure. These data structures include the in-memory stripe cache, logical region status, and WP (write pointer).

[0100] Assume that a write request includes data blocks D2 and D3, which need to be written to different physical areas of the ZNS SSD. In step 401, these data blocks have been written, and the check block P1 has also been updated.

[0101] In step 402, the system first determines the write position of D2 and D3, and updates the logical actual write pointer to point to the position after D3. Then, the system generates metadata for D2 and D3, including their write position, size and verification information. These metadata are written into the OOB area, corresponding to the data blocks of D2 and D3.

[0102] In step 403, if the system fails after writing D2 and D3, the system will use the metadata in the OOB area to reposition the logical actual write pointer and determine which data blocks are valid. The system can then discard the invalid data and recalculate the check blocks using the valid data blocks to restore the consistency of the redundant array.

[0103] 404. Based on the repositioning of the logical actual write pointer after the failure, the data blocks outside the logical actual write pointer are discarded, the data blocks within the logical actual write pointer are restored, and the verification data is recalculated to complete the redundant array management of the ZNS solid state drive.

[0104] After the logical real write pointer is repositioned, the system identifies data blocks outside the Real WP, which are considered invalid because they may not have been correctly written or synchronized when the failure occurred.

[0105] To ensure data consistency and prevent potential data corruption, the system overwrites the data blocks outside the Real WP with zeros. This step effectively discards invalid data and prepares storage space for new write operations.

[0106] For the data blocks within the Real WP, the system performs recovery operations, which may include re-reading the data blocks, verifying the integrity of the data, and ensuring the consistency of the data blocks with the check blocks.

[0107] Based on the restored data blocks, the system recalculates the parity block data. This step ensures that the redundancy and error detection capabilities of the RAID array are restored. After recalculating the parity block, the system verifies the consistency between the data block and the parity block. This may involve additional verification operations on the data block to ensure the accuracy of the data.

[0108] After completing data recovery and parity block recalculation, the system updates its internal state, including the logical area state and write pointer position, to reflect the latest data state. The system then sends a recovery completion notification to the host, indicating that the storage system has been restored to a consistent state and is ready to accept new write requests.

[0109] The embodiment of the present invention provides a redundant array management method based on a ZNS solid-state drive. By utilizing the Zone Random Write Area (ZRWA) feature of the ZNS SSD, the in-situ update of the check block is realized, that is, the data is directly overwritten at the original check block location without the need for additional write operations. The realization of this overwrite update benefits from the ability of ZRWA to allow random writing within a certain range. When a data block is updated, the related check information will also change accordingly. ZRWA provides a mechanism so that new check data can be written directly to the original check block location without allocating new storage space. In this way, not only the number of write operations is reduced, the logical area resources can be efficiently utilized, but also the write performance can be improved and the delay can be reduced.

[0110] Further, see Figure 6 , step 403 specifically includes:

[0111] 601. In the event of a machine failure, query the status information of each physical area in the ZNS solid state drive and determine the status of the logical area.

[0112] Among them, the status information of each physical area is used to describe the status of the corresponding physical area as full, empty, available, closed, read-only or abnormal; the logical area status is used to describe the status of the corresponding logical area as full, empty, available, closed, read-only or abnormal.

[0113] Specifically:

[0114] Full: The zone is full;

[0115] Empty: The state where the zone data is empty;

[0116] Opened: The status after the open zone command is successfully executed on the zone (available status);

[0117] Closed: The zone is not full yet, and is in this state after the close zone command succeeds.

[0118] Read Only: a zone in read-only state;

[0119] Offline: The zone is in an abnormal state, which may be caused by media abnormality or other problems.

[0120] Specifically, querying the status information of each physical area in the ZNS solid state drive and determining the status of the logical area specifically includes: sending a query instruction to the solid state drive, receiving a descriptor including the status of the physical area and the write pointer of the physical area, determining the status information of each physical area, and determining the corresponding logical area status and logical write pointer position.

[0121] 602. Based on the state of the logical area and the metadata in the spare area, determine the write pointer information of each physical area, and relocate the logical actual write pointer position of the corresponding logical area.

[0122] Specifically, step 602 includes: based on the logical area state and the logical write pointer position, querying metadata in the spare area to determine whether each data block has data written; wherein, if the data block has data written, the corresponding metadata is 1, and if the data block has no data written, the corresponding metadata is 0;

[0123] According to the write pointer information of each data block, the logical actual write pointer is relocated to the tail position of the data block where data is continuously and successfully written.

[0124] In the embodiment of the present invention, the movement granularity of the ZRWA zone and the non-ZRWA zone is different, which makes it difficult to recover the write pointer (WP) after a machine failure. If the write pointer is not stored or recovered properly, it may cause data loss and data persistence problems. The present invention proposes a lightweight metadata scheme for locating the write pointer to ensure the correctness of Zebra.

[0125] In a non-ZRWA zone, newly written data will move the write pointer (WP), and the minimum moving unit is 1LBA (i.e. 4KB), which is consistent with the write granularity of ZNS SSD. In a ZRWA zone, the moving unit of the write pointer is ZRWAFG (usually equal to the size of the SSD flash page, for example, usually greater than 16KB in MLC / TLC SSD). This is significantly different from the moving granularity of a non-ZRWA zone.

[0126] like Figure 7 (a) and Figure 7 As shown in the change of (b), assuming that ZRWAFG is 4LBA, there is currently an additional write request of size 2LBA, and the target address is LBA16. The start position of ZRWA moves to LBA4, and the tail position moves to LBA20. Even if no data is written in the range of LBA [18, 20), the movement of ZRWA will still update the start and tail positions of the entire range. This difference in movement granularity further increases the complexity of positioning the write pointer during fault recovery.

[0127] Zebra makes ZRWA transparent to applications and provides traditional zone abstraction (i.e., non-ZRWA zone) to upper-layer applications. In a non-ZRWA zone, the write pointer (WP) points to the tail of the successfully written data. However, in a ZRWA zone, WP points to the starting position of ZRWA, which is different from the semantics of WP in a non-ZRWA zone and cannot directly locate the tail LBA of the most recent append write. To eliminate this semantic difference, Zebra defines a logical actual write pointer RealWP, which points to the tail position of consecutively successfully written data within a ZRWA zone. It should be noted that due to the coarse movement granularity of ZRWA, the position of the logical actual write pointer RealWP is not always equal to the tail position of ZRWA.

[0128] In the event of a sudden power outage, the cached WP of all zones stored in DRAM will be lost. Although the WP is maintained on the ZNS device, the ZNS SSD does not record which areas of the ZRWA have been written, so the RealWP cannot be accurately restored during the recovery process. Zebra solved this problem with the help of the spare area (OOB), which is a proven technology supported by modern NVMe SSDs and widely used in industrial systems. (Concept of OOB area: The OOB area is a spare area for each flash page, which is used to store the metadata of the page. The size of the OOB area is usually 8B to 64B per LBA. When writing data to a flash page, the user can attach metadata to the OOB area. The metadata written to the OOB area is consistent with the data written to the corresponding flash page).

[0129] Zebra uses the metadata in the OOB area to determine whether the data has been written. The OOB area is initially filled with 0. When a 2LBA write request is issued to LBA16 (see Figure 7 (c)), the corresponding OOB area is written as 1, indicating that a write request has occurred to the flash page. If a machine failure occurs and restarts at this time, Zebra needs to query WP from ZNSSSD. ZNS SSD maintains the logical write pointer WP position and area status in hardware. By using OOB, Zebra traverses the OOB area metadata of the ZRWA flash page and places the logical actual write pointer RealWP at the tail position of the data block where data is written continuously and successfully, and its metadata value is 1 (such as Figure 7 (c) LBA18). Therefore, Zebra successfully recovers the logical real write pointer RealWP of the zone after the power failure event.

[0130] Furthermore, the method further comprises:

[0131] When a read request is received, directly reading corresponding data according to the data block address in the logical area specified in the read request;

[0132] In the event of a disk failure, data required for a read request is rebuilt based on the data and the check data in the same stripe, so that the failed data can be recovered based on the data and the check data in the same stripe.

[0133] When processing a read request:

[0134] When a host application or operating system issues a read request, the request contains the address of a specific data block in the logical area. The Zebra system parses the address in the read request and maps it to the corresponding physical area and data block.

[0135] The system retrieves the data block from the corresponding physical area. If the data block is distributed across multiple ZNS SSDs, the system reads the data from these SSDs in parallel. If the data block spans multiple physical areas, the system merges these scattered data blocks to reconstruct the original data.

[0136] After the data is merged, the system may use the check block to verify the integrity of the data. After the verification is correct, the system returns the data to the host.

[0137] Disk failure data recovery:

[0138] The system detects a disk failure, which may be due to hardware failure or other reasons that make data inaccessible. The system identifies the physical area of ​​the failure and determines which data blocks are affected.

[0139] For data blocks in the failed area, the system uses other data blocks and parity blocks in the same stripe to reconstruct the missing data. The system uses the parity information in the parity block to verify and reconstruct the data block. The reconstructed data is written back to the corresponding location in the logical area to maintain data consistency. If possible, the system replaces the failed physical area with a new healthy area.

[0140] The system updates its internal state, including logical region state and physical region state, to reflect data recovery and region replacement. During data recovery, the system monitors performance metrics such as read latency and throughput to ensure that the recovery process does not negatively impact system performance.

[0141] Assume that the host sends a read request for data block D2 in the logical area. The Zebra system parses the address and finds that D2 is distributed on two different ZNS SSDs. The system reads data from the two SSDs in parallel and returns it to the host after merging.

[0142] In another scenario, if a physical area on ZNS SSD#1 fails, causing data block D3 to be inaccessible, the system will use other data blocks and parity block P1 in the same stripe to rebuild D3. After the reconstruction is complete, the system will write the reconstructed data back to the corresponding location of the logical area and may replace the failed physical area with a new healthy area.

[0143] The above embodiments demonstrate how the Zebra system processes read requests under normal circumstances and how to use RAID technology to recover data when a disk fails, thereby ensuring system reliability and data integrity.

[0144] In the event of a machine failure, the data structures stored in the host memory are lost. Zebra must recover these data structures from the device after the failure. These data structures include the stripe cache in memory (the stripe cache on the host side, which contains the stripes of all active logical areas), the logical area status, and WP. The recovery of the logical area status and WP (write pointer) has been mentioned above. For the stripe cache, it includes the restored data blocks and the check block data.

[0145] Furthermore, in step 404, based on the repositioning of the logical actual write pointer after the failure, the data blocks outside the logical actual write pointer are discarded, specifically including: based on the repositioning of the logical actual write pointer after the failure, the data in the data blocks outside the logical actual write pointer are treated as invalid data and overwritten with zeros to discard the invalid data outside the logical actual write pointer.

[0146] The method of recovering the data blocks within the logical actual write pointer and recalculating and generating verification data to complete the redundant array management of the ZNS solid state drive specifically includes: recalculating and generating verification block data based on the recovered data blocks within the logical actual write pointer, determining the remaining continuous data blocks of the stripe after discarding invalid data, and directly writing the remaining continuous data blocks when there is a new write request to complete the redundant array management of the ZNS solid state drive.

[0147] Assume that after a failure, the system finds that the Real WP is located after data block D5 through the metadata of the OOB area. Data blocks D6 and D7 had started writing but not completed before the failure, so they were marked as invalid. The system performs a zero overwrite operation on D6 and D7 to clear the invalid data in these data blocks. Then, the system recalculates the check block P1 to ensure consistency with data block D5. After completing these steps, the system identifies new continuous data blocks to prepare for new write requests.

[0148] Figure 8 An example is shown. Figure 8In Figure 1, a full stripe write request (green) W is split into multiple sub-I / Os and distributed to multiple physical regions simultaneously. If a machine failure occurs before W is completed, some sub-I / Os are completed while others are not.

[0149] exist Figure 8 In (b), in Zebra, if a crash occurs, the unfinished write operation will be marked as invalid. Zebra recovers the logical area WP to the WP of Zone0 and cleans up the invalid data that was not completed on Zone1 and Zone2. After recovery, the new write operation (New write after recovery) (purple) will be directly performed in stripe#0, using the in-place overwrite capability to directly overwrite the data block marked as invalid on stripe 0.

[0150] Zebra uses the ZRWA (Zone Random Write Area) feature to allow in-place overwriting on the check block, which reduces write losses caused by crashes. In this way, Zebra avoids skipping stripe #0 after recovery, and continues writing directly on stripe #0, effectively utilizing storage space and reducing wasted space.

[0151] In contrast, RAIZN cannot overwrite dirty data blocks of stripe 0 on Zone1 and Zone2 due to write order restrictions, such as Figure 8 (a). To this end, RAIZN must skip stripe 0 and write the new request to the next stripe (i.e., stripe 1), resulting in permanent storage space loss of stripe 0. In order to record the information of lost space, RAIZN needs to consume additional memory overhead to maintain the mapping table.

[0152] After experiments, Zebra performed significantly better than RAIZN under small request write loads. For 4KB load, Zebra achieved a 30% throughput improvement, requiring only 5 open zones to reach peak throughput (1178MB / s), while RAIZN required 9 open zones. Under larger requests (8KB / 16KB / 32KB), Zebra's throughput was 34%, 17%, and 14% higher than RAIZN on average. In addition, Zebra eliminates the need for metadata areas and GC through in-place verification updates, reduces write amplification, and further optimizes performance and latency (by 19%-20% respectively). In contrast, RAIZN's latency increases due to metadata area contention in multi-zone PPU aggregation, becoming a performance bottleneck.

[0153] The embodiments of the present invention can also achieve the following beneficial effects:

[0154] 1) The ability to overwrite and update check blocks also helps reduce the need for garbage collection. In traditional ZNS SSD management, garbage collection is a resource-intensive process that involves cleaning up invalid data and reallocating storage space. By reducing the amount of data that requires garbage collection, the method of the present invention further optimizes the performance and efficiency of the storage system.

[0155] 2) By storing lightweight metadata in the spare area, the present invention can accurately locate the logical real write pointer (Real WP) after a machine failure. This design not only ensures data consistency, but also avoids the risk of data loss due to write pointer loss.

[0156] 3) During the fault recovery process, the present invention achieves zero storage overhead recovery by discarding invalid data outside the actual write pointer and overwriting it with zeros. This method avoids the loss of additional storage space and improves recovery efficiency.

[0157] refer to Fig. 9 The present invention also discloses a redundant array management system based on a ZNS solid state drive, wherein physical areas belonging to different ZNS solid state drives are combined into corresponding logical areas, wherein the logical areas include multiple stripes, and each stripe includes at least one check block and multiple data blocks belonging to different ZNS solid state drives; the system includes:

[0158] The data writing module 901 is used for, for each stripe, when receiving a write request, writing the data to be written into the corresponding data block in sequence, and overwriting and updating the check data of the check block;

[0159] Pointer generation module 902, used to locate the logical actual write pointer to the tail position of the data block to which the data is continuously and successfully written, and generate metadata corresponding to each data block, and write the metadata into the spare area; wherein the logical actual write pointer is used to point to the next writable position in the logical area;

[0160] A pointer positioning module 903, configured to reposition the logical actual write pointer based on the metadata in the spare area in the event of a failure;

[0161] The data recovery module 904 is used to discard the data blocks outside the logical actual write pointer based on the repositioning of the logical actual write pointer after the failure, recover the data blocks within the logical actual write pointer, and recalculate and generate verification data to complete the redundant array management of the ZNS solid state drive.

[0162] Optionally, the system further comprises:

[0163] A data reading module, configured to directly read corresponding data according to a data block address in a specified logical area in the read request when a read request is received;

[0164] In the event of a disk failure, data required for a read request is rebuilt based on the data and the check data in the same stripe, so that the failed data can be recovered based on the data and the check data in the same stripe.

[0165] Fig.10 An example of a physical structure diagram of an electronic device is shown in FIG. Fig.10 As shown, the electronic device may include: a processor (processor) 1010, a communication interface (Communications Interface) 1020, a memory (memory) 1030 and a communication bus 1040, wherein the processor 1010, the communication interface 1020, and the memory 1030 communicate with each other through the communication bus 1040. The processor 1010 can call the logic instructions in the memory 1030 to execute a redundant array management method based on a ZNS solid-state hard drive, the method comprising: for each stripe, when a write request is received, the data to be written is written into the corresponding data block in sequence, and the check data of the check block is overwritten and updated; the logical actual write pointer is positioned to the tail position of the data block to which the data is continuously and successfully written, and metadata corresponding to each data block is generated, and the metadata is written into the spare area; wherein the logical actual write pointer is used to point to the next writable position in the logical area; in the event of a failure, the logical actual write pointer is repositioned based on the metadata in the spare area; based on the repositioning of the logical actual write pointer after the failure, the data blocks outside the logical actual write pointer are discarded, the data blocks in the logical area are restored and the check data is recalculated to complete the redundant array management of the ZNS solid-state hard drive.

[0166] In addition, the logic instructions in the above-mentioned memory 1030 can be implemented in the form of software functional units and can be stored in a computer-readable storage medium when sold or used as an independent product. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art or the part of the technical solution, can be embodied in the form of a software product, which is stored in a storage medium and includes several instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes: various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a disk or an optical disk.

[0167] On the other hand, the present invention also provides a computer program product, which includes a computer program, which can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the redundant array management method based on the ZNS solid-state hard disk provided by the above methods, the method comprising: for each stripe, when a write request is received, the data to be written is written into the corresponding data block in sequence, and the check data of the check block is overwritten and updated; the logical actual write pointer is positioned to the tail position of the data block where the data is successfully written continuously, and metadata corresponding to each data block is generated, and the metadata is written into a spare area; wherein the logical actual write pointer is used to point to the next writable position in the logical area; in the event of a failure, the logical actual write pointer is repositioned based on the metadata in the spare area; based on the repositioning of the logical actual write pointer after the failure, the data blocks outside the logical actual write pointer are discarded, the data blocks in the logical area are restored and the check data is recalculated to complete the redundant array management of the ZNS solid-state hard disk.

[0168] On the other hand, the present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which is implemented when the processor executes the redundant array management method based on the ZNS solid-state hard disk provided by the above-mentioned methods, the method comprising: for each stripe, when a write request is received, the data to be written is written into the corresponding data block in sequence, and the check data of the check block is overwritten and updated; the logical actual write pointer is positioned to the tail position of the data block where the data is continuously and successfully written, and metadata corresponding to each data block is generated, and the metadata is written into a spare area; wherein the logical actual write pointer is used to point to the next writable position in the logical area; in the event of a failure, the logical actual write pointer is repositioned based on the metadata in the spare area; based on the repositioning of the logical actual write pointer after the failure, the data blocks outside the logical actual write pointer are discarded, the data blocks in the logical area are restored and the check data is recalculated to complete the redundant array management of the ZNS solid-state hard disk.

[0169] Through the description of the above implementation methods, those skilled in the art can clearly understand that each implementation method can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solution is essentially or the part that contributes to the prior art can be embodied in the form of a software product, and the computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a disk, an optical disk, etc., including a number of instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.

[0170] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A redundant array management method based on a ZNS solid state drive, characterized in that: Physical areas belonging to different ZNS solid state drives are combined into corresponding logical areas, wherein the logical area includes a plurality of stripes, and each stripe includes at least one check block and a plurality of data blocks belonging to different ZNS solid state drives; The method comprises: For each stripe, when a write request is received, the data to be written is written into the corresponding data block in sequence, and the check data of the check block is overwritten and updated; Positioning the logical actual write pointer to the tail position of the data block to which the data is continuously and successfully written, generating metadata corresponding to each data block, and writing the metadata into the spare area; wherein the logical actual write pointer is used to point to the next writable position in the logical area; In the event of a failure, repositioning the logical actual write pointer based on the metadata in the spare area; Based on the repositioning of the logical actual write pointer after the failure, the data blocks outside the logical actual write pointer are discarded, the data blocks within the logical actual write pointer are restored, and the verification data is recalculated to complete the redundant array management of the ZNS solid state drive.

2. The redundant array management method of a ZNS solid state drive according to claim 1, characterized in that: The method further comprises: When a read request is received, directly reading corresponding data according to the data block address in the logical area specified in the read request; In the event of a disk failure, data required for a read request is rebuilt based on the data and the check data in the same stripe, so that the failed data can be recovered based on the data and the check data in the same stripe.

3. The redundant array management method of a ZNS solid state drive according to claim 1, characterized in that: In the event of a failure, repositioning the logical actual write pointer based on the metadata in the spare area specifically includes: In the event of a machine failure, query the status information of each physical area in the ZNS solid state drive and determine the logical area status; wherein the logical area status is used to describe the status of the corresponding logical area as full, empty, available, closed, read-only or abnormal; Based on the state of the logical area and the metadata in the spare area, the write pointer information of each physical area is determined, and the logical actual write pointer position of the corresponding logical area is relocated.

4. The redundant array management method of a ZNS solid state drive according to claim 3, characterized in that: The querying of the status information of each physical area in the ZNS solid state drive and determining the status of the logical area specifically includes: By sending a query command to the solid-state drive, receiving a descriptor containing the physical area status and the write pointer of the physical area, determining the status information of each physical area, and determining the corresponding logical area status and logical write pointer position; wherein the status information of each physical area is used to describe the status of the corresponding physical area as full, empty, available, closed, read-only or abnormal.

5. The redundant array management method of a ZNS solid state drive according to claim 4, characterized in that: The step of determining the write pointer information of each physical area based on the state of the logical area and the metadata in the spare area and relocating the logical actual write pointer position of the corresponding logical area specifically includes: Based on the logical area state and the logical write pointer position, query the metadata in the spare area to determine whether each data block has data written; wherein, if the data block has data written, the corresponding metadata is 1, and if the data block has no data written, the corresponding metadata is 0; According to the write pointer information of each data block, the logical actual write pointer is relocated to the tail position of the data block where the data is continuously and successfully written.

6. The redundant array management method of a ZNS solid state drive according to claim 1, characterized in that: The repositioning of the logical actual write pointer after the failure and discarding the data blocks outside the logical actual write pointer specifically includes: Based on the repositioning of the logical actual write pointer after the failure, the data in the data block outside the logical actual write pointer is treated as invalid data and overwritten with zeros, so as to discard the invalid data outside the logical actual write pointer.

7. The redundant array management method of a ZNS solid state drive according to claim 1, characterized in that: The data block within the logical actual write pointer is recovered, and the verification data is recalculated to complete the redundant array management of the ZNS solid state drive, which specifically includes: Based on the recovered data blocks within the logical actual write pointer, the check block data is recalculated and generated to determine the remaining continuous data blocks of the stripe after the invalid data is discarded. When there is a new write request, the remaining continuous data blocks are directly written to complete the redundant array management of the ZNS solid state drive.

8. A redundant array management system based on ZNS solid state drive, characterized in that: Physical areas belonging to different ZNS solid state drives are combined into corresponding logical areas, the logical area includes a plurality of stripes, each stripe includes at least one check block and a plurality of data blocks belonging to different ZNS solid state drives, the system includes: A data writing module is used for, for each stripe, when receiving a write request, writing the data to be written into the corresponding data block in sequence, and overwriting and updating the check data of the check block; A pointer generation module, used to position the logical actual write pointer to the tail position of the data block to which the data is continuously and successfully written, and to generate metadata corresponding to each data block, and to write the metadata into the spare area; wherein the logical actual write pointer is used to point to the next writable position in the logical area; A pointer positioning module, used for repositioning the logical actual write pointer based on the metadata in the spare area in the event of a failure; The data recovery module is used to discard the data blocks outside the logical actual write pointer based on the repositioning of the logical actual write pointer after the failure, recover the data blocks within the logical actual write pointer, and recalculate and generate verification data to complete the redundant array management of the ZNS solid state drive.

9. The redundant array management system of a ZNS solid state drive according to claim 8, characterized in that: The system further comprises: A data reading module, configured to directly read corresponding data according to a data block address in a specified logical area in the read request when a read request is received; In the event of a disk failure, data required for a read request is rebuilt based on the data and the check data in the same stripe, so that the failed data can be recovered based on the data and the check data in the same stripe.

10. An electronic device comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that: When the processor executes the computer program, the redundant array management method based on the ZNS solid state drive as described in any one of claims 1 to 7 is implemented.

Citation Information

Patent Citations

  • Solid state disk array construction method, electronic device and storage medium

    CN110297601A

  • Data management method based on ZRWA function of ZNS solid state disk

    CN115686372A

  • Solid state disk data processing method and device, electronic equipment and storage medium

    CN117406933A

  • Testing method, device and equipment for solid state disk and medium

    CN117690474A

  • Methods and systems for raid protection in zoned solid-state drives

    US11340987B1

Cited By

  • High-reliability high-capacity hard disk implementation method and system based on fine-grained ZNS partition

    CN120179183A