Data space recovery method and device, computer equipment and storage medium
By aggregating and batching data blocks in distributed file systems, the method addresses storage space fragmentation caused by small-granularity Unmap requests, improving data recovery efficiency and reliability.
Patent Information
- Application Number
- CN202510806266.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-17
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2045-06-17
AI Technical Summary
In distributed file systems, frequent small-grain Unmap requests lead to fragmentation of storage space release, increasing backend load, and affecting subsequent allocation efficiency.
By judging the length threshold of the data block to be recycled, split the large-grained data blocks and aggregate the small-grained data blocks, batch process Unmap requests, adapt to the requirements of the storage area network array, and reduce the number of recycling times.
The pressure on the backend load of multiple data recovery is reduced, the reliability and subsequent allocation efficiency of data space recovery is improved, and the network array environment is adapted to different storage areas.
Smart Images

Figure CN120316079A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer technology, and in particular, to a data space recycling method, apparatus, computer device, and storage medium. Background Art
[0002] In a distributed file system, space recycling is a key mechanism for efficiently managing storage resources, avoiding space waste, and ensuring the performance and stability of the storage system. In a SAN (Storage Area Network) storage environment, when a file system deletes a file, traditional SAN storage does not automatically release the underlying physical space (the logical volume is still marked as "used"). The role of Unmap (block device unloading) is to notify the storage array which data blocks are no longer in use, allowing the array to recycle this space for other volumes.
[0003] In related technologies, when a distributed file system issues an Unmap instruction, it generally issues the instruction to the kernel driver directly through its own block device management interface or through a virtual machine SCSI command (Small Computer System Interface command), and a large number of small-grained Unmap requests will be sent, with each request corresponding to one or more logical block addresses. Frequent execution of small-grained Unmap requests will cause frequent metadata updates, increase the backend load, and lead to fragmentation of the released storage space. Summary of the Invention
[0004] This application provides a data space recycling method, apparatus, computer device, and storage medium to at least solve the problem in related technologies that directly processing block device unloading requests will cause fragmentation of the released storage space and affect subsequent allocation efficiency.
[0005] This application provides a data space recycling method, including: In response to the total length of the data blocks to be recycled in the first recycling queue being greater than a first threshold, determining whether there are data blocks to be recycled in the first recycling queue with a data length greater than a second threshold; In response to there being data blocks to be recycled in the first recycling queue with a data length greater than a second threshold, splitting the data blocks to be recycled with a data length greater than the second threshold to obtain a plurality of split data blocks, and storing the plurality of split data blocks in a second recycling queue; Selecting split data blocks from the second recycling queue and adding them to a third recycling queue, and in response to the number of split data blocks in the third recycling queue being greater than a third threshold, merging the split data blocks in the third recycling queue and issuing them to the storage area network array for data space recycling.
[0006] This application also provides a data space recycling apparatus, including: A judgment module, configured to respond that the total length of the data blocks to be recycled in the first recycling queue is greater than a first threshold; and judge whether the first recycling queue includes data blocks to be recycled with a data length greater than a second threshold; A splitting module, configured to respond that the first recycling queue includes data blocks to be recycled with a data length greater than the second threshold, split the data blocks to be recycled with a data length greater than the second threshold to obtain a plurality of split data blocks, and store the plurality of split data blocks in a second recycling queue; A recycling module, configured to select split data blocks from the second recycling queue and add them to a third recycling queue, and respond that the number of split data blocks in the third recycling queue is greater than a third threshold, merge the split data blocks in the third recycling queue and send them to a storage area network array for data space recycling.
[0007] This application also provides a computer device, including: a memory, configured to store a computer program; a processor, configured to implement the steps of the data space recycling method in the following embodiments when executing the computer program.
[0008] Respond that the total length of the data blocks to be recycled in the first recycling queue is greater than the first threshold, and judge whether the first recycling queue includes data blocks to be recycled with a data length greater than the second threshold; Respond that the first recycling queue includes data blocks to be recycled with a data length greater than the second threshold, split the data blocks to be recycled with a data length greater than the second threshold to obtain a plurality of split data blocks, and store the plurality of split data blocks in a second recycling queue; Select split data blocks from the second recycling queue and add them to a third recycling queue, and respond that the number of split data blocks in the third recycling queue is greater than the third threshold, merge the split data blocks in the third recycling queue and send them to a storage area network array for data space recycling.
[0009] This application also provides a computer-readable storage medium, in which a computer program is stored. When the computer program is executed by a processor, the steps of the data space recycling method in any of the following embodiments are implemented.
[0010] Respond that the total length of the data blocks to be recycled in the first recycling queue is greater than the first threshold, and judge whether the first recycling queue includes data blocks to be recycled with a data length greater than the second threshold; Respond that the first recycling queue includes data blocks to be recycled with a data length greater than the second threshold, split the data blocks to be recycled with a data length greater than the second threshold to obtain a plurality of split data blocks, and store the plurality of split data blocks in a second recycling queue; Select split data blocks from the second recovery queue and add them to the third recovery queue. In response to the number of split data blocks in the third recovery queue being greater than the third threshold, merge the split data blocks in the third recovery queue and send them to the storage area network array for data space recovery.
[0011] Through the data space recovery method provided by this application, since it is possible to determine whether there are to-be-recovered data blocks with a data length greater than the second threshold in the first recovery queue when the total length of the to-be-recovered data blocks in the first recovery queue is greater than the first threshold; when there are to-be-recovered data blocks with a data length greater than the second threshold in the first recovery queue, split the to-be-recovered data blocks with a data length greater than the second threshold to obtain multiple split data blocks, store the multiple split data blocks in the second recovery queue; select split data blocks from the second recovery queue and add them to the third recovery queue, and when the number of split data blocks in the third recovery queue is greater than the third threshold, merge the split data blocks in the third recovery queue and send them to the storage area network array for data space recovery.
[0012] In this way, multiple small-grained to-be-recovered data blocks can be aggregated, and the data space corresponding to the to-be-recovered data blocks can be recovered in batches, reducing the number of recoveries, which can reduce the pressure on the backend load caused by multiple data recoveries. It also considers setting the data length threshold of the to-be-recovered data blocks and the data volume threshold of the split data blocks to adapt to the data length requirements and the number of data of the block device unloaded by the storage area network array, so as to solve the technical problem that directly processing the block device unloading request in the related art will cause fragmentation of the storage space release and affect the subsequent allocation efficiency, and achieve the technical effect of improving the reliability of data space recovery. BRIEF DESCRIPTION OF THE DRAWINGS
[0013] To more clearly illustrate the embodiments of the present application, the following will briefly introduce the drawings required in the embodiments. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0014] Figure 1 It is a schematic flowchart of the data space recovery method provided by an embodiment of the present application; Figure 2 It is a schematic diagram of data block aggregation provided by an embodiment of the present application; Figure 3 It is a schematic flowchart of the data space recovery method provided by another embodiment of the present application; Figure 4 It is a schematic flowchart of the data space recovery method provided by yet another embodiment of the present application; Figure 5Schematic diagram of splitting data blocks to be recycled provided by an embodiment of the present application; Figure 6 Schematic flow chart of a data space recycling method provided by another embodiment of the present application; Figure 7 Block diagram of the structure of a data space recycling device provided by an embodiment of the present application; Figure 8 Internal structure diagram of a computer device provided by an embodiment of the present application. Detailed implementation manners
[0015] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.
[0016] It should be noted that in the description of the present application, the terms "including", "comprising" or any other variation thereof are intended to cover a non-exclusive inclusion, such that a process, method, article or device including a series of elements not only includes those elements but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or device. The terms "first", "second", etc. in the present application are used to distinguish similar objects and are not used to describe a specific order or sequence.
[0017] To enable those skilled in the art of the present technology to better understand the solution of the present application, the present application will be further described in detail below in conjunction with the accompanying drawings and specific implementation manners.
[0018] As Figure 1 shown, an embodiment of the present application provides a data space recycling method, which specifically includes the following steps: Step 101: In response to the total length of the data blocks to be recycled in the first recycling queue being greater than the first threshold, determine whether the first recycling queue includes data blocks to be recycled with a data length greater than the second threshold.
[0019] Step 102: In response to the first recycling queue including data blocks to be recycled with a data length greater than the second threshold, split the data blocks to be recycled with a data length greater than the second threshold to obtain a plurality of split data blocks, and store the plurality of split data blocks in the second recycling queue.
[0020] Specifically, before the total length of the data blocks to be recycled in the first recycling queue is greater than the first threshold, it includes: determining whether the block device unloading instruction has been completed; in response to the block device unloading instruction not being completed, sequentially selecting data blocks to be recycled from the cache queue and adding them to the first recycling queue until the total length of the data blocks to be recycled in the first recycling queue is greater than the first threshold.
[0021] The block device unloading instruction (Unmap RBD) is a key command for the client to notify the storage system to release unused disk blocks, which is used for efficient storage space management. Especially in a thin-provisioned storage system, its core function is to release the occupancy of the RBD device in the kernel, and release the network connection and memory cache between the client and the OSD.
[0022] The data block to be recycled refers to the data block whose data space is about to be recycled, and the data length refers to the size of the data volume contained in the data block.
[0023] The total length of the data blocks to be recycled in the first recycling queue refers to the sum of the data volumes contained in all the data blocks to be recycled in the first recycling queue.
[0024] First, when the block device unloading instruction has not been completed, sequentially select data blocks to be recycled from the cache queue (the queue of data blocks released from the metadata cache) and add them to the first recycling queue until the total length of the data blocks to be recycled in the first recycling queue is greater than the first threshold.
[0025] Among them, sequentially selecting data blocks to be recycled from the cache queue and adding them to the first recycling queue includes: sequentially selecting any data block to be recycled in the cache queue as the target data block; obtaining the target data block capacity corresponding to the target data block and the size of the used space of the target data block, and calculating the target recycling benefit of the target data block based on the target data block capacity and the size of the used space of the target data block; obtaining the target data block access frequency, access decay factor, and new access count of the target data block within the historical period, and calculating the target access frequency of the target data block based on the target data block access frequency, access decay factor, and new access count within the historical period; obtaining the current time corresponding to the target data block and the last access time of the target data block, and calculating the target data age of the target data block based on the current time corresponding to the target data block and the last access time of the target data block; calculating the recycling priority of the target data block based on the target recycling benefit, target access frequency, target data age, and the priority calculation formula, and sequentially selecting the target data block with the largest recycling priority according to the size of the recycling priority and adding it to the first recycling queue.
[0026] Among them, the priority calculation formula is as follows: ; Among them, P represents the recycling priority of the target data block, a represents the recycling benefit coefficient, X represents the capacity of the target data block, Y represents the size of the used space of the target data block, b represents the data age coefficient, T1 represents the current time, T2 represents the time of the last access to the target data block, c represents the access frequency coefficient, N represents the access frequency of the target data block within the historical period, A represents the access attenuation factor, and C represents the new access count.
[0027] The capacity of the target data block refers to the maximum amount of data that can be accommodated in the target data block. The size of the used space of the target data block refers to the amount of data stored in the current target data block. The access frequency of the target data block within the corresponding historical period refers to the number of times the target data block is accessed within a certain historical period. The access attenuation factor is a constant and can be set according to actual experience. The access attenuation factor is used to characterize the impact of the number of accesses on performance attenuation. The new access count refers to the current number of new accesses to the data block. The current access time of the target data block, the time of the last access to the target data block refers to the time of the last access to the target data block.
[0028] The higher the recycling benefit, the better the execution of the data space recycling operation for the target data block. The higher the data age, the more likely it is that the target data block is no longer used, the target data block is inactive, and it is more optimal to execute the data space recycling operation. The lower the access frequency, the more likely it is that the target data block is cold data, and it is more optimal to execute the data space recycling operation.
[0029] The recycling benefit coefficient, data age coefficient, and access frequency coefficient here can be set respectively according to the influence of the recycling benefit, data age, and access frequency on data space recycling. For example, if the recycling benefit has the greatest impact on data space recycling, the data age has the second greatest impact, and the access frequency has little impact, then set the value of the recycling benefit coefficient to be the largest, the value of the data age coefficient to be the second largest, and the value of the access frequency coefficient to be the smallest. Here, the sum of the recycling benefit coefficient, data age coefficient, and access frequency coefficient is 1.
[0030] In this way, by calculating the recycling priority of each data block to be recycled through the target recycling benefit, access frequency, data age, and priority calculation formula of the data block to be recycled, and preferentially adding the data blocks to be recycled with high recycling priority to the first recycling queue for processing, it optimizes the multi-dimensional quantitative evaluation of the performance, revenue efficiency, and resource utilization rate of the storage system, and can improve the reliability of data space recycling for the data blocks to be recycled.
[0031] The first threshold here is a pre-set threshold for the total length of aggregated data blocks. When the total length of the data blocks to be recycled in the first recycling queue is greater than the first threshold, the data blocks to be recycled in the first recycling queue can be considered for aggregation operations. The first threshold here can be set according to actual experience and is used to represent the critical value of the data length of the data blocks for data aggregation.
[0032] After the total length of the data blocks to be recycled in the first recycling queue is greater than the first threshold, it can be judged whether there are data blocks to be recycled in the first recycling queue with a data length greater than the second threshold.
[0033] The second threshold here refers to the upper limit threshold for the length of a single data block set. The second threshold is used to distinguish whether the data block to be processed is a large-grained data block or a small-grained data block.
[0034] When the first recycling queue does not include data blocks to be recycled with a data length greater than the second threshold, it is considered that all the data blocks to be recycled in the first recycling queue are small-grained data blocks, and the data blocks to be recycled in the first recycling queue can be directly aggregated and sent to the storage area network array for data space recycling. In this way, the scattered data blocks to be recycled can be aggregated, and the space corresponding to the data blocks can be recycled as a whole, avoiding the problem of frequent metadata updates caused by multiple small-grained data blocks performing multiple space recycles and increasing the backend load.
[0035] When the first recycling queue includes data blocks to be recycled with a data length greater than the second threshold, the data blocks to be recycled with a data length greater than the second threshold are split to obtain multiple split data blocks, that is, the large-grained data blocks are converted into small-grained data blocks. The data block lengths of the multiple split data blocks meet the requirements of the SAN array for the length of a single data block, and the multiple split data blocks are stored in the second recycling queue.
[0036] In one embodiment, splitting the data blocks to be recycled with a data length greater than the second threshold to obtain multiple split data blocks includes: obtaining the data block processing length threshold set by the storage area network array; using the data block processing length threshold as the splitting unit, splitting the data blocks to be recycled with a data length greater than the second threshold to obtain multiple split data blocks.
[0037] The data block processing length threshold set by the storage area network array here refers to the maximum length threshold for processing a single data block set by the storage area network array. Splitting the data blocks to be recycled with a data length greater than the second threshold using the data block processing length threshold can split the large-grained data blocks to be recycled into multiple split data blocks adapted to the storage area network array, flexibly adapting to the requirements of different storage area network array environments, and improving the reliability of data space recycling.
[0038] Step 103: Select split data blocks from the second recovery queue and add them to the third recovery queue. In response to the number of split data blocks in the third recovery queue being greater than the third threshold, merge the split data blocks in the third recovery queue and send them to the storage area network array for data space recovery.
[0039] Specifically, obtain the total amount of data block recovery for a single block device unloading operation; calculate the number of data block recoveries for a single block device unloading operation based on the total amount of data block recovery for a single block device unloading operation and the data block processing length threshold; in response to the number of split data blocks in the third recovery queue being greater than the number of data block recoveries for a single block device unloading operation, merge the split data blocks in the third recovery queue and send them to the storage area network array for data space recovery.
[0040] The total amount of data block recovery for a single block device unloading operation here refers to the maximum capacity of data space that can be recovered by executing the block device unloading operation command once. Taking the total amount of data block recovery for a single block device unloading operation as the dividend and the maximum length threshold for processing a single data block set by the storage area network array as the divisor, the quotient obtained by calculation is used as the number of data block recoveries for a single block device unloading operation. The number of data block recoveries for a single block device unloading operation here is the third threshold. The third threshold is the maximum data block processing quantity threshold for the set split data blocks, which is used to determine whether the number of split data blocks in the third recovery queue meets the requirement for the number of data blocks for a single data space recovery process in the SAN array.
[0041] Select split data blocks from the second recovery queue and add them to the third recovery queue until the number of split data blocks in the third recovery queue is greater than the third threshold. When the number of split data blocks in the third recovery queue is greater than the third threshold, merge the split data blocks in the third recovery queue and send them to the storage area network array for data space recovery.
[0042] In one embodiment, the data space recovery method further includes: after aggregating and sending the split data blocks in the third recovery queue to the storage area network array for data space recovery, it further includes judging to obtain the return data of the data space recovery in the storage area network array, and judging whether the split data blocks in the third recovery queue are successfully recovered based on the return data; in response to the unsuccessful recovery of the split data blocks in the third recovery queue, re-aggregating the split data blocks in the third recovery queue and sending them to the storage area network array for data space recovery; in response to the successful recovery of the split data blocks in the third recovery queue, clearing the data in the third recovery queue, and continuing to select split data blocks from the second recovery queue to join the third recovery queue until all the split data blocks in the second recovery queue are added to the third recovery queue. In response to all the split data blocks in the second recovery queue being added to the third recovery queue, clearing the data in the first recovery queue, and continuing to select data blocks to be recovered from the cache queue to join the first recovery queue.
[0043] This application also sets that the return data after the data space recovery of the SAN array can be obtained, and it is judged whether the split data blocks in the third recovery queue are successfully recovered according to the return data. When the split data blocks in the third recovery queue are not successfully recovered, re-aggregate the split data blocks in the third recovery queue and send them to the storage area network array for data space recovery. When the split data blocks in the third recovery queue are successfully recovered, clear the data in the third recovery queue, and continue to select split data blocks from the second recovery queue to join the third recovery queue until all the split data blocks in the second recovery queue are added to the third recovery queue. Similarly, when all the split data blocks in the second recovery queue are added to the third recovery queue, clear the data in the first recovery queue, and continue to select data blocks to be recovered from the cache queue to join the first recovery queue until all the data blocks to be recovered in the cache queue are added to the first recovery queue.
[0044] In this way, it is verified whether the data space of the data blocks to be recovered sent to each recovery queue is successfully released, and the problem of missed space recovery is avoided by the retry method, improving the reliability of data space recovery.
[0045] In this application, by actively aggregating the small-granularity block device unmap requests (Unmap requests), the pressure on the backend load caused by the small-granularity Unmap requests is reduced, the fragmentation problem of storage space release is solved, and the subsequent allocation efficiency is improved. And by setting the data block processing length threshold and the number of data blocks recovered in a single block device unload operation, it flexibly adapts to the requirements of different storage area network array environments to achieve a better data aggregation effect.
[0046] In a feasible implementation, before performing a data space recovery operation on the data blocks to be recycled / split data blocks based on the SAN array, the available space size corresponding to the data blocks to be recycled / split data blocks can be recorded. After performing the data space recovery operation on the data blocks to be recycled / split data blocks based on the SAN array, obtain the available space size of the data blocks after recovery, and compare whether the available space sizes of the data blocks before and after the data space recovery operation are consistent. If the available space sizes of the data blocks before and after the data space recovery operation are consistent, it is considered that the data space recovery operation of the data blocks to be recycled / split data blocks is successful. If the available space sizes of the data blocks before and after the data space recovery operation are inconsistent, it is considered that the data space recovery operation of the data blocks to be recycled / split data blocks is unsuccessful, and an alarm operation is performed. In this way, the effectiveness of the data space recovery operation can be ensured.
[0047] In a specific implementation, the data space recovery method can be as follows: S1: Set the total length (the first threshold) of the data block aggregation data blocks, such as 512k.
[0048] In the Ceph system (an open-source, distributed, and scalable storage system), files are split into multiple Blobs (binary large objects) for storage. A Blob is the physical storage unit of file data in the file system, and it can be managed and stored independently of the file system. A data block (extent) represents a continuous byte range within a file. When a file is split into multiple Blobs, each Blob corresponds to one or more extents in the file. Extent is used to describe the layout of file data in physical storage, specifying the starting position (Lba offset) and length (length) of the file data in the Blob, as shown in the Figure 2 data structure shown. Common length specifications of Extent are: 4k, 128K, 2M, 4M, etc.
[0049] S2: Create a thread (discard_thread) to process the data block information released from the cache in the MDS (metadata server) after deleting a file. Each data block is recorded by Lba offset and length.
[0050] S3: When the block device unmount operation sent by the metadata is not completed, take a data block to be recycled from the extent queue (cache queue) released from the MDS cache, and add the data block to be recycled to the first recovery queue.
[0051] S4: Determine whether the total length of all the data blocks to be recycled in the first recovery queue is greater than the first threshold.
[0052] If it exceeds the first threshold, perform an aggregation operation on the data blocks to be recycled in the first recycling queue and send them to the SAN array for data space recycling. Otherwise, execute step S3, continue to take the next data block to be recycled and put it into the first recycling queue until the total length of all the data blocks to be recycled in the first recycling queue is greater than the first threshold. Assume that all the data blocks to be recycled taken out from the first recycling queue are small-granularity ones with a length of 4k. Only when the number of data blocks to be recycled accumulates to 128 can the first threshold be exceeded to meet the total length requirement for data block aggregation, and then a batch process will be triggered, so as to achieve the effect of aggregating frequent small-granularity Unmap requests. The processing flow of steps S1 - S4 is as shown in the appendix Figure 3 as follows
[0053] Assume that the first recycling queue includes data blocks to be recycled with a data length greater than the second threshold, that is, large-granularity data blocks to be recycled, such as a length of 2M, which exceeds the set upper limit length threshold of a single data block (the second threshold). Then execute steps S5 - S8, specifically refer to Figure 4 .
[0054] S5: Set the upper limit of the length of a single data block EXTENT_MAX_SIZE, such as 2M
[0055] S6: Traverse the first recycling queue in step S4. Take out a data block to be recycled from the first recycling queue. If the data block length of the data block to be recycled is greater than the second threshold, split this data block to be recycled into multiple split data blocks; otherwise, do not split
[0056] Since most SAN arrays can set the length threshold of a single data block to be processed (the second threshold), such as set to 2MB. Split the large-granularity data blocks to be recycled into multiple split data blocks adapted to the storage area network array, flexibly adapting to the requirements of different storage area network array environments. The process of splitting the data blocks can be as shown in the appendix Figure 5 as follows. Take the data block to be recycled as the division unit with the data block processing length threshold, and split it into multiple split data blocks (as shown in the figure, split data block 0, split data block 1, split data block 2,... split data block N). Record the starting position and length of the data block to be recycled and the split data blocks. The starting length of the split data block is the starting position of the previous split data block plus the data block processing length threshold. The length of the split data block is usually equal to the data block processing length threshold. The length of the data block to be recycled is the sum of the lengths of its corresponding multiple split data blocks. For example, a data block to be recycled with a length of 4MB can be split into two 2MB split data blocks, and a data block to be recycled with a length of 2.128MB can be split into a 2MB split data block and a 128kB split data block
[0057] S7: Store the multiple split data blocks obtained in step S6 into the second recycle queue.
[0058] S8: Perform a batch process on the multiple split data blocks in the second recycle queue.
[0059] S9: Set the upper limit (the third threshold) of the number of data blocks released by a single block device unloading operation, such as 64.
[0060] S10: Traverse the second recycle queue, take out a split data block from the second recycle queue and insert it into the third recycle queue.
[0061] S11: Determine whether the number of split data blocks in the current third recycle queue is greater than the third threshold. If it is greater than the third threshold, send the split data blocks in the third recycle queue to the SAN array for space recycling of these split data blocks. If it is less than the third threshold, return to step S10 to continue taking the next split data block and insert it into the third recycle queue until the number of split data blocks in the third recycle queue is greater than the third threshold.
[0062] Since most SAN arrays can set the total amount of data block recycling for a single block device unloading operation, the present invention calculates the number of data blocks recycled by a single block device unloading operation (the third threshold) based on the total amount of data block recycling and the data block processing length threshold of a single block device unloading operation. The third threshold is used to determine whether the number of split data blocks in the third recycle queue meets the requirement of the number of data blocks for a single data space recycling process of the SAN array, so as to flexibly adapt to the configurations of different SAN array environments.
[0063] S12: Obtain the return data after the SAN array performs data space recycling, and determine whether the split data blocks in the third recycle queue are successfully recycled according to the return data. If it fails, re-send the split data blocks in the third recycle queue in step 11, and avoid the problem of missed space recycling through the retry method.
[0064] S13: If it is determined in step S12 that the split data blocks in the third recycle queue are successfully recycled according to the return data, clear the third recycle queue, return to step 10, and process the next batch of data blocks. Steps S9 - S13 are as shown in the appendix Figure 6 shown.
[0065] S14: If the processing of multiple split data blocks to be recycled in the second recycling queue in step S13 is completed, it indicates that there is no need for the first recycling queue to split the data blocks to be recycled with a data length greater than the second threshold again to obtain multiple split data blocks and send the multiple split data blocks to the first recycling queue. At this time, clear the first recycling queue. Then return to step S3 to continue to take new data blocks to be recycled from the queue of cached releases in the MDS, and continue to perform data space recycling processing on the data blocks to be recycled that have not been processed in the cache queue according to the above steps.
[0066] In this application, by intelligently identifying the unmap requests sent by the MDS, it is distinguished whether they are small-grained Unmap requests or large-grained Unmap requests, that is, small-grained data blocks to be recycled or large-grained data blocks to be recycled, and the small-grained Unmap requests are actively aggregated to reduce the pressure of small-grained Unmap requests on the backend load, solve the fragmentation problem of storage space release, and improve the subsequent allocation efficiency. And by setting the data block processing length threshold and the number of data blocks recycled in a single block device unloading operation, it flexibly adapts to the requirements of different storage area network array environments to achieve a better data aggregation effect.
[0067] An embodiment of this application provides a data space recycling device, and the data space recycling device is specifically as Figure 7 shown. The data space recycling device includes: a judgment module 20, a splitting module 21, and a recycling module 22.
[0068] The judgment module is used to respond that the total length of the data blocks to be recycled in the first recycling queue is greater than the first threshold; judge whether there are data blocks to be recycled with a data length greater than the second threshold in the first recycling queue; The splitting module is used to respond that there are data blocks to be recycled with a data length greater than the second threshold in the first recycling queue, split the data blocks to be recycled with a data length greater than the second threshold to obtain multiple split data blocks, and store the multiple split data blocks in the second recycling queue; The recycling module is used to select split data blocks from the second recycling queue and add them to the third recycling queue. In response to the number of split data blocks in the third recycling queue being greater than the third threshold, aggregate the split data blocks in the third recycling queue and send them to the storage area network array for data space recycling.
[0069] For the description of the features in the corresponding embodiment of the data space recycling device, reference can be made to the relevant description of the corresponding embodiment of the data space recycling method, which will not be elaborated here one by one.
[0070] An embodiment of this application also provides a computer device, such as Figure 8As shown, it includes a memory and a processor. A computer program is stored in the memory, and the processor is configured to run the computer program to execute the steps in any of the above-described embodiments of the data space recovery method.
[0071] An embodiment of the present application further provides a computer-readable storage medium in which a computer program is stored. The computer program is configured to execute the steps in any of the above-described embodiments of the data space recovery method when running.
[0072] In an exemplary embodiment, the above computer-readable storage medium may include, but is not limited to: various media such as USB flash drives, read-only memories (ROMs for short), random access memories (RAMs for short), mobile hard disks, magnetic disks, or optical discs that can store computer programs.
[0073] Those skilled in the art can further realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of the two. To clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described according to functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present application.
[0074] The above has introduced in detail a data space recovery method provided by the present application. Specific examples are used herein to elaborate on the principle and implementation manner of the present application. The description of the above embodiments is only used to help understand the method and its core idea of the present application. It should be noted that for those of ordinary skill in the art in the technical field, without departing from the principle of the present application, several improvements and modifications can be made to the present application, and these improvements and modifications also fall within the protection scope of the claims of the present application.
Claims
1. A data space recycling method, characterized in that, Including: In response to the total length of the data blocks to be recycled in the first recycling queue being greater than a first threshold, determining whether there are data blocks to be recycled with a data length greater than a second threshold in the first recycling queue; In response to there being data blocks to be recycled with a data length greater than the second threshold in the first recycling queue, splitting the data blocks to be recycled with a data length greater than the second threshold to obtain a plurality of split data blocks, and storing the plurality of split data blocks in a second recycling queue; Selecting split data blocks from the second recycling queue and adding them to a third recycling queue, and in response to the number of split data blocks in the third recycling queue being greater than a third threshold, merging the split data blocks in the third recycling queue and sending them to a storage area network array for data space recycling.
2. The data space recycling method according to claim 1, wherein Before the step of "In response to the total length of the data blocks to be recycled in the first recycling queue being greater than the first threshold", it includes: Determining whether the block device unloading instruction has been completed; In response to the block device unloading instruction not being completed, sequentially selecting data blocks to be recycled from the cache queue and adding them to the first recycling queue until the total length of the data blocks to be recycled in the first recycling queue is greater than the first threshold.
3. The data space recycling method according to claim 1, wherein The step of "splitting the data blocks to be recycled with a data length greater than the second threshold to obtain a plurality of split data blocks" includes: Obtaining the data block processing length threshold set by the storage area network array; Taking the data block processing length threshold as the splitting unit, splitting the data blocks to be recycled with a data length greater than the second threshold to obtain a plurality of split data blocks.
4. The data space recycling method according to claim 1, wherein The step of "in response to the number of split data blocks in the third recycling queue being greater than the third threshold, merging the split data blocks in the third recycling queue and sending them to the storage area network array for data space recycling" includes: Obtaining the total amount of data blocks recycled in a single block device unloading operation; Calculating the number of data blocks recycled in a single block device unloading operation based on the total amount of data blocks recycled in the single block device unloading operation and the data block processing length threshold; In response to the number of split data blocks in the third recycling queue being greater than the number of data blocks recycled in a single block device unloading operation, merging the split data blocks in the third recycling queue and sending them to the storage area network array for data space recycling.
5. The data space recycling method according to claim 1, wherein The method further includes: Obtaining the return data for data space recycling in the storage area network array, and based on the return data, determining whether the split data blocks in the third recycling queue are recycled successfully; In response to the split data blocks in the third recycling queue not being recycled successfully, re-merging the split data blocks in the third recycling queue and sending them to the storage area network array for data space recycling; In response to the split data blocks in the third recycling queue being recycled successfully, clearing the data in the third recycling queue, and continuing to select split data blocks from the second recycling queue and adding them to the third recycling queue until all the split data blocks in the second recycling queue are added to the third recycling queue.
6. The data space recycling method according to claim 1, wherein The method further includes: In response to all the split data blocks in the second recycling queue being added to the third recycling queue, clearing the data in the first recycling queue, and continuing to select data blocks to be recycled from the cache queue and adding them to the first recycling queue.
7. The data space recycling method according to claim 2, wherein Sequentially selecting data blocks to be recycled from the cache queue and adding them to the first recycling queue includes: Sequentially selecting any data block to be recycled in the cache queue as the target data block; Obtaining the capacity of the target data block corresponding to the target data block and the size of the used space of the target data block, and calculating the target recycling benefit of the target data block based on the capacity of the target data block and the size of the used space of the target data block; Obtaining the access frequency of the target data block, the access attenuation factor, and the new access count of the target data block within the historical period corresponding to the target data block, and calculating the target access frequency of the target data block based on the access frequency of the target data block within the historical period, the access attenuation factor, and the new access count; Obtaining the current time corresponding to the target data block and the last access time of the target data block, and calculating the target data age of the target data block based on the current time and the last access time; Calculating the recycling priority of the target data block based on the target recycling benefit, the target access frequency, the target data age, and the priority calculation formula, and sequentially selecting the target data block with the largest recycling priority and adding it to the first recycling queue according to the size of the recycling priority; Among them, the priority calculation formula is as follows: ; Among them, P represents the recycling priority of the target data block, a represents the recycling benefit coefficient, X represents the capacity of the target data block, Y represents the size of the used space of the target data block, b represents the data age coefficient, T1 represents the current time, T2 represents the last access time, c represents the access frequency coefficient, N represents the access frequency of the target data block within the historical period, A represents the access attenuation factor, and C represents the new access count.
8. A data space recovery device, characterized in that, Including: A judgment module for responding that the total length of the data blocks to be recycled in the first recycling queue is greater than the first threshold; Judging whether there is a data block to be recycled with a data length greater than the second threshold in the first recycling queue; A splitting module for, in response to the first recycling queue including a data block to be recycled with a data length greater than the second threshold, splitting the data block to be recycled with a data length greater than the second threshold to obtain a plurality of split data blocks, and storing the plurality of split data blocks in the second recycling queue; A recycling module for selecting split data blocks from the second recycling queue and adding them to the third recycling queue, and in response to the number of split data blocks in the third recycling queue being greater than the third threshold, merging the split data blocks in the third recycling queue and sending them to the storage area network array for data space recycling.
9. A computer device, characterized in that, Including: A memory for storing a computer program; A processor for implementing the steps of the data space recycling method according to any one of claims 1 to 7 when executing the computer program.
10. A computer-readable storage medium, characterized in that, A computer program is stored in the computer-readable storage medium, wherein the computer program implements the steps of the data space recycling method according to any one of claims 1 to 7 when executed by a processor.
Citation Information
Patent Citations
Swarm intelligence based spatial data copy self-adapting distribution method
CN101504663A
Method and device for realizing space release in solid-state disk array
CN108304139A
Storage space recovery method, device and equipment and computer storage medium
CN112162701A
Method and apparatus for managing storage system
US20180210798A1