ZNS SSD data distribution method based on adaptive partition size configuration
By adopting adaptive variable partition size configuration method and dynamic mapping technology on ZNS SSD, the shortcomings in performance and capacity of traditional SSDs and the problems of parallel utilization of ZNS SSDs are solved, efficient data allocation and recycling are achieved, and write bandwidth is improved.
Patent Information
- Application Number
- CN202510034116.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-09
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2045-01-09
AI Technical Summary
Traditional block flash-based SSDs have additional space reservations and complex flash conversion layer problems when meeting high performance and high capacity requirements. Although ZNS SSDs provide better storage interfaces, they lack effective methods to leverage device parallelism, resulting in latency and poor performance of partition recycling.
Adaptive variable partition size configuration method is adopted, and the partition size and parallelism are dynamically adjusted in combination with data access characteristics and device parallelism requirements to achieve maximum parallelism while reducing recycling latency and overhead. The specific solutions include setting up four levels of partitions, using linked list management, designing new ZNS commands and dynamic partition mapping methods based on chip conflict perception.
It realizes the reduction or even elimination of garbage collection overhead under maximum parallelism, improves the write bandwidth of the device, and meets the needs of high performance and high capacity.
Smart Images

Figure CN119937930A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of computers, and in particular relates to a ZNS SSD data allocation method based on adaptive partition size configuration. Background Art
[0002] With the exponential growth of data, modern web-scale systems and applications place higher demands on the capacity and performance of storage systems. Traditional block flash-based SSDs, which are widely used in modern data, often require additional reserved space and cumbersome flash conversion layers, which are not conducive to meeting performance and capacity requirements. The emerging storage device ZNS SSD presents a new storage interface that divides the entire logical address space into fixed-size partitions. Each partition must be written sequentially and cannot be overwritten. It must be reset before it can be rewritten. At the same time, ZNS SSD allows the host to directly manage these partitions, which can minimize garbage collection overhead and write amplification.
[0003] Although ZNS is a good interface that can draw a clear boundary between the host and the flash firmware through partitions, it lacks hardware abstraction that explicitly exploits the internal parallelism of the SSD from the host. For this reason, many studies have explored the size and mapping method of the partitions in the hope of achieving a mapping scheme that can achieve the best performance of the device. Usually a partition is converted into many flash blocks across multiple flash chips, and each partition has a large writable capacity, which is the commonly used large partition. The larger the partition, the more flash chips can be mapped to the partition, which means it has higher internal parallelism, so that multiple blocks can be written and erased in parallel, showing high performance.
[0004] Unfortunately, however, it may extend the delay time of partition recovery, thereby blocking subsequent services. In order to solve this problem, the concept of small partitions, that is, partitions with smaller write capacity, is introduced. Small partitions are generally only mapped to flash blocks spanning one or several flash chips. Small partitions can solve the problem of long delays in partition recovery, but its problems are also obvious. The fewer the number of chips mapped, the lower the internal parallelism, resulting in lower performance. It can be seen that a fixed partition size often cannot meet both high parallelism and low recovery delay. Therefore, the present invention proposes a configuration of an adaptive variable partition size and a corresponding data allocation method based on ZNS SSD, so that the system can fully utilize the full parallelism of the device while reducing garbage collection delays and overhead. Summary of the invention
[0005] In view of this, the purpose of the present invention is to provide a ZNS SSD data allocation method based on adaptive partition size configuration. In combination with the access characteristics of the data and the full parallelism writing requirements of the device, partitions of different sizes and parallelisms are matched for data with different characteristics, so as to achieve the maximum parallelism of the device while reducing or even eliminating the recycling overhead.
[0006] The specific plan includes the following steps:
[0007] S1. Based on the logical space of ZNS SSD, set up 4 levels of partitions and use 4 linked lists to manage the 4 levels of partitions respectively;
[0008] S2. If the written data is generated by the refresh operation of the LSM tree, a partition with maximum parallelism is directly assigned to it; if the written data is generated by the compression operation of the LSM tree, a partition is assigned to the written data through a variable partition combination allocation method;
[0009] S3. Design a new ZNS command and implement partition mapping through a partition dynamic mapping method based on chip conflict awareness.
[0010] Beneficial effects of the present invention:
[0011] The present invention designs a variable-size partition configuration based on ZNS SSD and proposes a corresponding data allocation method, which enables the system to reduce or even eliminate garbage collection overhead while achieving the maximum parallelism of the device, greatly improving the write bandwidth of the device. BRIEF DESCRIPTION OF THE DRAWINGS
[0012] Figure 1 is a flow chart of the method of the present invention;
[0013] Figure 2 It is a schematic diagram of the present invention managing the logical partition space based on the linked list. DETAILED DESCRIPTION
[0014] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0015] The present invention provides a ZNS SSD data allocation method based on adaptive partition size configuration, such as Figure 1 As shown, the following steps are included:
[0016] S1. Based on the logical space of ZNS SSD, 4 levels of partitions are set and 4 linked lists are used to manage the 4 levels of partitions respectively.
[0017] Specifically, Figure 2 As shown, the present invention defines SZ as the partition size that can be mapped to a parallel block group. First, the logical space of the ZNS SSD system is partitioned in units of SZ. That is, the partition of SZ size is used as the base partition, and the logical space of the ZNS SSD is divided into multiple base partitions. In the specific use process, in order to realize the flexible and variable partition size, four levels of partitions are obtained by further dividing and merging according to the base partition; the size of the partition belonging to the first level is equal to the size of one base partition, that is, a single base partition can be regarded as a first-level partition; the size of the partition belonging to the second level is equal to the size of two base partitions, that is, two base partitions can be merged to obtain a second-level partition; the size of the partition belonging to the third level is equal to the size of four base partitions, and the size of the partition belonging to the fourth level is equal to the size of eight base partitions. The number of partitions of the four levels is randomly initialized.
[0018] In the present invention, the SZ size is 64M, and from the first level to the fourth level, the size of the corresponding partition increases by 2 times, which are 64M, 128M, 256M, and 512M respectively; at the same time, from the first level to the fourth level, the parallelism of the corresponding partition also increases by 2 times, among which the partition of the 4th level with a size of 512M has the maximum parallelism of the device.
[0019] The logical space of the ZNS SSD system is divided into partitions of different sizes, and four types of linked lists are used to manage the four different sizes of partitions. When a partition of a certain size is needed, first check the linked list corresponding to that level to see if there is an allocable partition. If not, merge or split the free partitions; when a new partition of a certain size is allocated, add it to the linked list corresponding to the level for unified management; when a partition of a certain size is recycled, remove it from the linked list corresponding to the level.
[0020] S2. If the written data is a single file generated by the LSM tree refresh operation, a partition with maximum parallelism is directly allocated to it (that is, a fourth-level partition is directly allocated); if the written data is multiple files generated by the LSM tree compression operation, partitions are allocated to the written data through a variable partition combination allocation method.
[0021] Specifically, in order to ensure that the maximum parallelism of the device can be achieved at every moment of writing in the ZNS SSD system, it is necessary to give different writing characteristics to the data in combination with the structural characteristics of the LSM tree. With respect to the writing characteristics of the data, the present invention mainly considers the refresh operation and compression operation of the LSM tree.
[0022] The refresh operation of the LSM tree, that is, writing files from memory to disk, generally only generates one SSTable file and expires very quickly, and the data in a file often expires at the same time. Therefore, for this type of data generated by the LSM tree refresh operation, it can be directly assigned to a partition with maximum parallelism to ensure that all chips of the device can be fully utilized during writing. At the same time, it expires very quickly and does not bring any recovery overhead.
[0023] The number of files and expiration time generated by the compression operation of the LSM tree are uncertain, so the data generated by the compression operation needs to be classified according to the specific number of files generated each time. Since the number of files generated by the compression operation is unknown, only by knowing the specific number of files generated can the full parallelism-aware partition allocation be performed. Therefore, the compression results are predicted and counted based on each actual compression process. Specifically, based on the number of files involved in each compression and the merged and discarded expired key-value pairs, the specific number of files generated by this compression can be reasonably inferred.
[0024] Specifically, after obtaining the number of files generated by each compression operation, a partition of the required size is allocated to the data file through a variable partition combination allocation method under the premise of ensuring that the writing of this compression process can achieve full device parallelism, including:
[0025] S21. Get the number of files generated by the LSM tree compression operation;
[0026] S22. Obtain the required number of partitions and the size of each required partition according to the number of files; initialize the allocated set;
[0027] S23. Determine whether the number of partitions stored in the allocated set is equal to the required number of partitions. If so, end the process to obtain all the partitions required for writing data; if not, start a new round of allocation and execute step S24;
[0028] S24. To avoid conflicts between partitions during parallel writing, the partition size required for this round of allocation is determined based on the maximum parallelism of device writing and the storage conditions in the allocated set;
[0029] S25. A partition is allocated using the partition space management method, and the partition does not conflict with any partition stored in the allocated set; the partition allocated in this round is placed in the allocated set, and the process returns to step S23.
[0030] Specifically, a partition space management method is designed based on the buddy algorithm to achieve efficient allocation and recovery of variable logical partitions. Step S25 The process of allocating a partition through the partition space management method includes:
[0031] S251. Determine the target level corresponding to the required partition size and traverse the linked list corresponding to the target level to determine whether there is an allocatable partition at the target level (i.e., an idle partition with a lifespan close to that of the data to be written). If so, allocate it and add it to the linked list corresponding to the target level; if not, execute step S252;
[0032] S252. Determine whether the sum of the sizes of all currently free partitions is greater than the required partition size. If so, merge or split the free partitions to obtain free partitions of the target level, allocate the obtained free partitions and add them to the linked list corresponding to the target level; if not, execute step S253;
[0033] Specifically, when the total size of all current free partitions is greater than the required partition size, there are two situations. The first is that there are free partitions at a higher level than the target level. In this case, you can choose to split the free partitions at the higher level to obtain free partitions at the target level. The second is that there are many free partitions at a lower level than the target level. In this case, you can choose to merge these low-level free partitions to obtain free partitions at the target level.
[0034] S253. The allocated partitions recovered by the space recovery method are used as new free partitions, and the free partitions are then merged or split to obtain free partitions of the target level, and the obtained free partitions are allocated and added to the linked list corresponding to the target level.
[0035] Specifically, the process of reclaiming the allocated partition by the space reclaiming method in step S253 includes:
[0036] S2531. According to the required partition size, determine the target level corresponding to the required partition size, traverse each allocated partition in the linked list corresponding to the target level, recycle the allocated partition that meets the recycling conditions and remove it from the linked list;
[0037] S2532. Determine whether there is an allocated partition of the target level that has been recycled. If so, the recycling ends; if not, traverse each allocated partition in the linked list corresponding to the remaining levels, recycle the allocated partition that meets the recycling conditions and remove it from the linked list.
[0038] The recycling condition is that the ratio of the effective data size in the allocated partition to the size of the allocated partition is less than 30%.
[0039] Specifically, the present invention will automatically determine whether the sum of the sizes of all currently free partitions is less than the recycling threshold at regular intervals. If so, it will determine whether each allocated partition meets the recycling conditions, and recycle the allocated partitions that meet the recycling conditions as reserve partitions. Among them, the present invention sets the recycling threshold to 30% of the total space size. If the ratio of the sum of the sizes of all currently free partitions to the total space size is less than 30%, the allocated partition needs to be recycled.
[0040] In the allocation process of step S25, when the free partitions are insufficient to meet the allocation requirements, the reserve partitions are first extracted as new free partitions for merging or splitting. If they are still insufficient for allocation, step S253 is executed to reclaim the allocated partitions in time through the space recovery method.
[0041] S3. Design a new ZNS command and implement partition mapping through a partition dynamic mapping method based on chip conflict awareness.
[0042] Specifically, in order to fully utilize the maximum parallelism of the device, it is necessary to ensure that all chips are utilized at each write moment of the device. Therefore, all chips and blocks on the device are first grouped according to the number of chips required by the benchmark partition. The benchmark partition needs to map 4 blocks from different chips. Therefore, every four chips form a parallel chipset, and the blocks with the same offset in the parallel chipset form a parallel block group, where the size of each parallel block group is equal to the size of a benchmark partition. Parallel chipsets do not conflict with each other, parallel block groups in different parallel chipsets do not conflict with each other, and parallel block groups in the same parallel chipset conflict with each other.
[0043] The present invention proposes a new ZNS command, namely the Merge command. The specific format of the Merge command is (src, size), where src represents the starting logical address of the partition to be merged and mapped (i.e., the partition to be mapped), and size represents the size of the partition to be mapped. The Merge command is dedicated to device-side partition mapping. When the host side needs to allocate a partition of a certain size, the Merge command is sent to the device side. The device side maps a corresponding number of non-conflicting parallel block groups to the partition to be mapped according to the starting logical address of the partition to be mapped and the size of the partition to be mapped passed in the Merge command.
[0044] Specifically, under the premise of avoiding conflicts between the mapped parallel block groups, based on the Merge command, the specific process of dynamically mapping parallel block groups for partitions through the partition dynamic mapping method is as follows:
[0045] S31. Confirm the number of partitions to be mapped according to the Merge command and initialize the mapped chipset set;
[0046] S32. For each partition to be mapped, dynamically select a parallel chipset to perform a mapping operation, and add the selected parallel chipset to the mapped chipset set; when all partitions to be mapped are mapped, the process ends.
[0047] Specifically, in step 32, the process of any partition to be mapped dynamically selecting a parallel chipset to perform a mapping operation includes:
[0048] S321. Confirm the number of parallel block groups z required for the partition to be mapped, and determine whether z is equal to the total number of parallel chipsets on the device. If so, select all parallel chipsets on the device, add all parallel chipsets to the mapped chipset set, and then execute step S324. If not, execute step S322.
[0049] S322. Calculate the number of idle parallel block groups of each parallel chipset on the device side, sort all parallel chipsets in descending order according to the number of idle parallel block groups to obtain a first sequence, remove the parallel chipsets in the mapped chipset set from the first sequence, and then execute step S323;
[0050] S323. Select the first z parallel chipsets and add these z parallel chipsets to the mapped chipset set;
[0051] S324. Map the idle parallel block groups in the selected parallel chipset.
[0052] Specifically, before executing step S323, a preliminary judgment needs to be performed, including:
[0053] S3231. Determine whether the mapped chipset set includes all parallel chipsets on the device side. If so, clear the mapped chipset set; if not, execute step S3232;
[0054] S3232. Determine whether the number of parallel chipsets included in the first sequence is less than z. If so, clear the mapped chipset set; otherwise, keep it unchanged.
[0055] In the present invention, unless otherwise clearly stipulated and limited, the terms such as "installation", "setting", "connection", "fixation" and "rotation" should be understood in a broad sense. For example, it can be a fixed connection, a detachable connection, or an integral one; it can be a mechanical connection or an electrical connection; it can be directly connected or indirectly connected through an intermediate medium; it can be the internal connection of two elements or the interaction relationship between two elements. Unless otherwise clearly defined, ordinary technicians in this field can understand the specific meanings of the above terms in the present invention according to the specific circumstances.
[0056] Although embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions and variations may be made to the embodiments without departing from the principles and spirit of the present invention, and that the scope of the present invention is defined by the appended claims and their equivalents.
Claims
1. A ZNS SSD data allocation method based on adaptive partition size configuration, characterized in that: The following steps are involved: S1. Based on the logical space of ZNS SSD, set up 4 levels of partitions and use 4 linked lists to manage the 4 levels of partitions respectively; S2. If the written data is generated by the refresh operation of the LSM tree, a partition with maximum parallelism is directly assigned to it; if the written data is generated by the compression operation of the LSM tree, a partition is assigned to the written data through a variable partition combination allocation method; S3. Design a new ZNS command and implement partition mapping through a partition dynamic mapping method based on chip conflict awareness.
2. According to claim 1, a ZNS SSD data allocation method based on adaptive partition size configuration is characterized in that: Step S1 specifically includes: First, the logical space of the ZNS SSD system is divided into base partitions; Classify the benchmark partitions to obtain four levels of partitions, wherein the size of a partition belonging to the first level is equal to the size of one benchmark partition, the size of a partition belonging to the second level is equal to the size of two benchmark partitions, the size of a partition belonging to the third level is equal to the size of four benchmark partitions, and the size of a partition belonging to the fourth level is equal to the size of eight benchmark partitions; For each level of partition, when it is allocated, the partition is added to the linked list of the corresponding level; when it is recycled, the partition is removed from the linked list of the corresponding level.
3. According to claim 1, a ZNS SSD data allocation method based on adaptive partition size configuration is characterized in that: Step S2 allocates partitions for write data using a variable partition combination allocation method, including: S21. Get the number of files generated by the LSM tree compression operation; S22. Obtain the required number of partitions according to the number of files; initialize the allocated set; S23. Determine whether the number of partitions stored in the allocated set is equal to the required number of partitions. If so, end the process to obtain all the partitions required for writing data; if not, start a new round of allocation and execute step S24; S24. Determine the partition size required for this round of allocation based on the maximum parallelism of device writes and the storage situation in the allocated set; S25. A partition is allocated using the partition space management method, and the partition does not conflict with any partition stored in the allocated set; the partition allocated in this round is placed in the allocated set, and the process returns to step S23.
4. A ZNS SSD data allocation method based on adaptive partition size configuration according to claim 3, characterized in that: The process of step S25 allocating a partition through the partition space management method includes: S251. Determine the target level corresponding to the required partition size and traverse the linked list corresponding to the target level to determine whether there is an allocatable partition at the target level. If so, allocate it and add it to the linked list corresponding to the target level; if not, execute step S252; S252. Determine whether the sum of the sizes of all currently free partitions is greater than the required partition size. If so, merge or split the free partitions to obtain free partitions of the target level, allocate the obtained free partitions and add them to the linked list corresponding to the target level; if not, execute step S253; S253. Use the allocated partitions recovered by the space recovery method as new free partitions, merge or split the free partitions to obtain free partitions of the target level, allocate the obtained free partitions and add them to the linked list corresponding to the target level.
5. A ZNS SSD data allocation method based on adaptive partition size configuration according to claim 4, characterized in that: The process of step S253 of reclaiming the allocated partition by using the space reclaiming method includes: S2531. According to the required partition size, determine the target level corresponding to the required partition size, traverse each allocated partition in the linked list corresponding to the target level, recycle the allocated partition that meets the recycling conditions and remove it from the linked list; S2532. Determine whether there is an allocated partition of the target level to be recycled. If so, the recycling ends. If not, traverse each allocated partition in the corresponding linked list of the remaining levels, recycle the allocated partition that meets the recycling conditions and remove it from the linked list; The recycling condition is that the ratio of the effective data size in the allocated partition to the size of the allocated partition is less than 30%.
6. According to the ZNS SSD data allocation method based on adaptive partition size configuration as described in claim 1, the new ZNS command is a Merge command for device-side partition mapping. When the host side needs to allocate a partition of a certain size, the Merge command is sent to the device side. The device side maps a parallel block group for the partition to be mapped according to the starting logical address and size of the partition to be mapped in the Merge command. The specific format of the Merge command is (src, size), where src represents the starting logical address of the partition to be mapped, and size represents the size of the partition to be mapped.
7. According to the ZNS SSD data allocation method based on adaptive partition size configuration according to claim 6, based on the Merge command, the process of implementing partition mapping through the partition dynamic mapping method includes: S31. Confirm the number of partitions to be mapped according to the Merge command and initialize the mapped chipset set; S32. For each partition to be mapped, dynamically select a parallel chipset for mapping operation, and add the selected parallel chipset to the mapped chipset set; When all the partitions to be mapped are mapped, the process ends.
8. According to the ZNS SSD data allocation method based on adaptive partition size configuration according to claim 7, in step 32, the process of dynamically selecting a parallel chipset for mapping operation for any partition to be mapped comprises: S321. Confirm the number of parallel block groups z required for the partition to be mapped, and determine whether z is equal to the total number of parallel chipsets on the device. If so, select all parallel chipsets on the device, add all parallel chipsets to the mapped chipset set, and then execute step S324. If not, execute step S322. S322. Calculate the number of idle parallel block groups of each parallel chipset on the device side, sort all parallel chipsets in descending order according to the number of idle parallel block groups to obtain a first sequence, remove the parallel chipsets in the mapped chipset set from the first sequence, and then execute step S323; S323. Select the first z parallel chipsets and add these z parallel chipsets to the mapped chipset set; S324. Map the idle parallel block groups in the selected parallel chipset.
Citation Information
Patent Citations
Intelligent write allocation method and device based on ZNS solid state disk
CN114546295A
ZNS-SSD storage system-oriented garbage collection method
CN116301576A
ZNS SSD management method, data writing method, storage device and controller
CN117215485A
Method for reducing invalid erasure of ZNS SSD device end based on strips and groups
CN117891404A
Method for optimizing reset operation in ZenFS to prolong service life of ZNS-SSD
CN117950590A