A method for reorganizing distributed object storage space

By migrating small objects to new empty large objects in a distributed object storage system, the problem of storage space fragmentation is solved, and higher storage space utilization and lower storage costs are achieved.

CN117827100BActive Publication Date: 2025-06-27CHINA TELECOM CLOUD TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202311709521.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-12-13
Publication Date
2025-06-27
Estimated Expiration
2043-12-13

AI Technical Summary

Technical Problem

The existing distributed object storage system leaves discontinuous empty space after small objects is deleted, resulting in fragmentation of storage space, reducing data read and write efficiency, and increasing storage costs.

Method used

Eliminate empty holes and improve storage space utilization by migrating undeleted small objects from the original large objects to the new empty large objects, establishing a target object mapping relationship table, and performing data and metadata migration.

Benefits of technology

It effectively reduces the fragmentation and waste of storage space, improves storage space utilization, reduces storage costs, and improves the efficiency of data reading and writing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117827100B_ABST
    Figure CN117827100B_ABST
Patent Text Reader

Abstract

The present invention relates to a method for reorganizing a distributed object storage space, belonging to the field of data storage. The method includes: obtaining an aggregated large object that meets the hole rate requirement as a source object, and inserting the source object whose fragmented space meets the minimum hole threshold into a hash list; selecting the source object with the highest fragmentation rate from the hash list for space reorganization, and creating a target aggregated large object as a target object according to requirements; generating a target object mapping relationship table; migrating the small object data on the source object to the target object according to the target object mapping relationship table; after the migration is successful, updating the layout information corresponding to the metadata of the small object itself; and performing a cleaning operation to end the task. The method of the present invention effectively reduces the fragmentation and waste of the storage space and better improves the storage space utilization rate by migrating the undeleted small objects from the original large object to a new empty large object.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of data storage, and particularly relates to a method for reorganizing a distributed object storage space. Background Art

[0002] (1) In existing distributed object storage, the definitions of the sizes of large objects and small objects are different, and the upload processing methods for large objects and small objects are also different.

[0003] For small object upload, it is basically supported that small objects are aggregated into a large object on an SSD disk and then stored on a backend HDD disk, which can not only improve the disk space utilization rate, but also enhance the performance of small object upload. However, when a small object is deleted, many discontinuous empty spaces will be left in the space of the aggregated large object, forming fragmented data blocks, resulting in an increase in latency overhead when randomly reading the undeleted objects, bringing a bad user experience to the upper-layer applications. Deleting and releasing the storage space occupied by the aggregated small objects may cause the storage space to become discontinuous, that is, a large number of non-adjacent free spaces appear, which may lead to fragmentation of the storage system, reduce the data reading and writing efficiency, and affect the system availability. To solve this problem, a common optimization method is to read the empty spaces and valid data on each aggregated large object, and then centrally migrate all the valid data to form continuous data blocks for reading.

[0004] (2) Although the existing fragmented space optimization method solves the problem of the continuity of valid data on the aggregated large object, there will be a large memory overhead when processing all the aggregated objects, and there is still storage space waste at the end of the aggregated object. When performing space reorganization, it will instead cause resource competition with other processes. The scattered empty spaces may increase the complexity of metadata management because the system needs to maintain the metadata information of non-empty objects to ensure correct object positioning and management.

[0005] (3) For the storage of video surveillance data, after small objects such as logs, texts, and pictures smaller than 1M are deleted, empty spaces of different sizes will be left in the storage space of the aggregated large object. These empty spaces may lead to waste of storage space because the system cannot effectively reuse these spaces, thus increasing the storage cost. Summary of the Invention

[0006] In view of the above deficiencies of the prior art, the purpose of the invention is to provide a method for reorganizing a distributed object storage space, which effectively reduces the fragmentation and waste of the storage space and better improves the storage space utilization rate by migrating the undeleted small objects from the original large object to a new empty large object.

[0007] The present invention provides a method for reorganizing a distributed object storage space, including:

[0008] S1, Obtain aggregated large objects that meet the requirements of the hole rate as source objects, and insert the source objects whose fragmented space meets the minimum hole threshold into the hash list;

[0009] S2, Select the source object with the highest fragmentation rate from the hash list for space arrangement, and create a target aggregated large object as the target object according to the requirements;

[0010] S3, Traverse the object table of the source object batch by batch, obtain the offset information of the small object on the source object, generate an ordered object table after sorting, obtain the interval information of each small object on the large object according to the ordered object table, add the interval information to the interval mapping table of the source object, allocate a corresponding target object interval for the source object interval, and then generate a target object mapping relationship table;

[0011] S4, Migrate the small object data on the source object to the target object according to the target object mapping relationship table, where,

[0012] First, perform data migration, including: reading the interval data corresponding to the small object in the source object, and then writing it to the new interval marked by the target object. After receiving the write success returned by the data disk write, it indicates that the data migration is successful;

[0013] Then, perform metadata and object table migration, including: batch inserting the small object list into the object table LSM tree corresponding to the target object, and at the same time inserting the metadata after interval arrangement of the source object into the metadata LSM tree of the target object;

[0014] S5, After the migration is successful, batch read the small objects in the source object, and update the interval layout information corresponding to the small object's own metadata according to the interval layout information corresponding to the small object on the target object;

[0015] S6, Perform a cleanup operation to end the task.

[0016] Further, the minimum hole threshold is dynamically adjusted according to the cluster capacity utilization rate.

[0017] Further, each source object can only select one target object, and each target object corresponds to multiple source objects.

[0018] Further, in S2, when the space arrangement starts to execute, the target object is automatically created in the storage bucket, and the states of the source object and the target object are set to the space arrangement state, where the sum of the effective capacities of the source objects cannot exceed the capacity of the target object.

[0019] Further, in the step S3, the object table information on the source object is stored in the LSM tree, and the object table and the memory of the relevant interval on the source object are obtained through the LSM enumeration interface.

[0020] Further, in the step S3, the offset information of the small object on the source object is obtained, the offset information of the small object on the source object is obtained according to the object name and the segment number, sorted from large to small, and an ordered object table is generated.

[0021] Further, in the step S3, after reading the interval mapping table of the source object, the offset information and the interval layout of the small object are modified to form continuous addresses, and the target object mapping relationship table is generated.

[0022] Further, in the step S5, before updating the layout information of the small information object, first check whether the small object still exists. If it does not exist, it means that it has been deleted during the space reorganization, and there is no need to update the layout information.

[0023] Further, in the step S5, when updating the layout information of the small information object, the update is performed in the order of the offset information of the small object, and it is judged whether the update process is normal. If it is normal, the memory is solidified after the update to ensure that the subsequent reorganization task can continue from the failed offset in case of an exception; if it is abnormal, it means that the layout information of some objects has been successfully updated, and the source object and the target object with valid data are used as the source object for the next space reorganization.

[0024] Further, in the step S6, the source object is cleaned, the memory structure solidified by the space reorganization is cleaned, and the solidified space reorganization task is cleaned.

[0025] The beneficial effects of the present invention are as follows:

[0026] (1) Small object dynamic migration strategy: By continuously integrating the valid data segments in large objects with a large gap between the valid space and the allocated space into a new empty large object, the purpose of eliminating holes can be achieved, and at the same time, the impact on the system load during the execution of space reorganization can also be reduced. This dynamic adjustment and migration strategy can dynamically adjust the threshold according to the cluster capacity usage and system load, and automatically start the execution of the background space reorganization task. This optimization strategy can enable the system to still maintain high efficiency execution under different loads, and solve the problems of high complexity of metadata management and impact on object reading and writing efficiency caused by the increase in storage space fragmentation due to repeated writing and deletion or overwrite writing of a large number of small objects.

[0027] (2) Object metadata maintenance method: The maintenance methods of metadata and object tables ensure the consistency and integrity of data and metadata before and after migration by establishing an interval mapping table. The data migration is performed by establishing an interval mapping table, and the metadata and object tables are migrated synchronously, ensuring data consistency and integrity.

[0028] (3) Space reorganization control structure: It is adaptively and dynamically adjusted according to the cluster capacity utilization rate and load, including the number of concurrent tasks, the minimum space threshold, the task execution interval, etc., improving the space reorganization efficiency and enhancing the continuity of data storage. By migrating small undeleted objects from the original large object to a new empty large object, the fragmentation and waste of storage space are effectively reduced, and the storage space utilization rate is better improved.

[0029] (4) Storage space optimization: The existing technology eliminates holes by migrating data within the space range of the original large object itself, while the present invention effectively reduces the fragmentation and waste of storage space and better improves the storage space utilization rate by migrating small undeleted objects from the original large object to a new empty large object. The present invention can help users and enterprises use storage space more effectively, reduce fragmentation problems, and lower storage costs.

[0030] (5) Adaptive space reorganization control structure: According to the cluster capacity usage and system load, various parameters of background tasks are adaptively adjusted, including the number of concurrent tasks, task intervals, etc.

[0031] (6) Data continuity: Optimizing the storage space helps to create continuous data storage blocks, thereby reducing the overhead of random access and accelerating read and write operations.

[0032] (7) Simplifying operation and maintenance operations: It supports starting and stopping background space reorganization tasks at any time. After starting, various parameters are adaptively adjusted according to the current situation of the cluster to improve the hole elimination efficiency. There is no need for complex operations to monitor and manage tasks, and only simple query commands are needed to view the intermediate process and results. Description of the Drawings

[0033] The drawings are only for the purpose of illustrating specific embodiments and are not considered as a limitation of the present invention. Throughout the drawings, the same reference signs denote the same components. Obviously, the drawings described below are only some embodiments recorded in the embodiments of the present invention, and those of ordinary skill in the art can obtain other drawings based on these drawings.

[0034] Figure 1 It is a flowchart of the distributed object storage space reorganization method according to the embodiment of the present invention;

[0035] Figure 2Schematic diagram of the distributed object storage space reorganization method according to an embodiment of the present invention;

[0036] Figure 3 Schematic diagram of the reorganization result of the source object space according to an embodiment of the present invention;

[0037] Figure 4 Flowchart of the interval reorganization work according to an embodiment of the present invention;

[0038] Figure 5 Flowchart of the object migration work according to an embodiment of the present invention;

[0039] Figure 6 Flowchart of the layout update work according to an embodiment of the present invention;

[0040] Figure 7 Flowchart of the space cleaning work according to an embodiment of the present invention. Detailed implementation manners

[0041] In order to enable those skilled in the art to better understand the technical solutions in the embodiments of the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. It should be understood that these descriptions are merely exemplary and are not intended to limit the scope of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0042] In addition, in the following description, the descriptions of well-known structures and technologies are omitted to avoid unnecessarily obscuring the concepts disclosed in the present invention.

[0043] In the description of the present invention, it should be noted that unless otherwise clearly defined and limited, the terms "center", "upper", "lower", "left", "right", "vertical", "horizontal", "inner", "outer", etc. indicate the orientation or positional relationship based on the orientation or positional relationship shown in the drawings. It is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and thus cannot be understood as a limitation of the present invention. In addition, the terms "first", "second", "third" are only used for descriptive purposes and cannot be understood as indicating or implying relative importance. The terms "installation", "connection", "connection" should be understood in a broad sense. For example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be directly connected or indirectly connected through an intermediate medium, and it can be the communication inside two elements. For those of ordinary skill in the art, the specific meanings of the above terms in the present invention can be understood according to specific situations.

[0044] Exemplary embodiments will be described in detail herein, and examples thereof are shown in the accompanying drawings. When the following description refers to the accompanying drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present invention. On the contrary, they are merely examples of methods and systems consistent with some aspects of the present invention as detailed in the appended claims.

[0045] The following describes the technical terms involved in the distributed object storage space reorganization method according to the embodiments of the present invention:

[0046] GC: Garbage Collection. In distributed object storage, GC generally refers to some asynchronous disk space reclamation or space reorganization operations. When a client performs an operation to delete an object or overwrite an object, the disk space occupied by the corresponding object is not immediately released, but is handed over to the background GC module for processing.

[0047] LSM tree: The full name is Log-Structured Merge Tree, which is a data structure used to implement a key-value storage system and is widely used in fields such as distributed databases and distributed file systems. Its core idea is to append key-value pairs to multiple storage files at different levels on the disk in sequence.

[0048] As Figure 1 and Figure 2 shown, the distributed object storage space reorganization method according to the embodiments of the present invention includes:

[0049] S1, obtain aggregated large objects that meet the hole rate requirement as source objects, and insert the source objects whose fragmented space meets the minimum hole threshold into the hash list.

[0050] For a distributed object storage cluster, the client uploads objects of different sizes to the cluster. Objects with a size less than 1M are defined as small objects. After the small objects are uploaded, they are aggregated into large objects in the cache space composed of SSD disks and stored on the backend HDD disks. The size of the aggregated large objects is defined as a maximum of 96M or the number of aggregated objects does not exceed 36,000. As the cluster running time increases, the number of uploaded small objects and generated aggregated large objects also increases. When the client performs an operation to delete an object, discontinuous hole spaces will be generated on the corresponding large object accordingly, and the space of the deleted object forms fragmented space. The space reorganization thread continuously scans in the background. When the ratio of the hole space on the aggregated large object to the effective capacity reaches a certain threshold, combined with the usage capacity ratio of the cluster storage pool, the space reorganization execution is automatically triggered.

[0051] To avoid the need to read the metadata information of the source object from disk every time when selecting the source object for space arrangement, the metadata management module inserts large objects with the status of "aggregated" and whose fragmented space meets the minimum threshold of the hole into a hash table. When selecting the source object, only the hash table needs to be traversed, and a batch of objects with the highest hole rate are selected for arrangement.

[0052] In the embodiment of the present invention, the minimum threshold of the hole is dynamically adjusted according to the utilization rate of the cluster capacity. The higher the utilization rate, the lower the minimum threshold, ensuring that the front-end service is minimally affected by the background space arrangement when the cluster is in a space-available state.

[0053] S2. Select the source object with the highest fragmentation rate from the hash list for space arrangement, and create a target aggregated large object as the target object according to the requirements.

[0054] Specifically, when the space arrangement starts to execute, the target object is automatically created in the storage bucket, and the statuses of the source object and the target object are set to the space arrangement status. The background thread ensures that the source object does not accept other migration operations, and the target object does not accept new object aggregation writes. During the space arrangement, the source object can be read and deleted but not written.

[0055] As Figure 3 shown, each source object can only select one target object, and each target object corresponds to multiple source objects. That is, multiple source objects can select the same target object. The sum of the valid capacities on multiple source objects cannot exceed the capacity of the target object.

[0056] S3. Traverse the object table of the source object batch by batch, obtain the offset information of the small object on the source object, generate an ordered object table after sorting, obtain the interval (layout) information of each small object on the large object according to the ordered object table, add the interval information to the interval mapping table of the source object, allocate the corresponding target object interval for the source object interval, and then generate the target object mapping relationship table.

[0057] As Figure 4 shown, first, the object table information on the source object is saved in the LSM tree, and the object table and the memory of the relevant interval on the source object are obtained through the LSM enumeration interface. Specifically, read the metadata of the source object and the corresponding object table, batch-read the layout information of the object metadata recorded on the object table, and store it in a predefined interval mapping table. Since the number of small objects in the object table may be very large, to avoid memory shortage, the object table is traversed batch by batch, with a maximum of N objects enumerated each time, and then the offset information of the small object on the source object is obtained according to the object name and the segment number and sorted from large to small to generate an ordered object table.

[0058] Then, according to the object information in the ordered object table, obtain the interval information of each small object on the large object. That is, obtain the interval (layout) information of each small object on the large object.

[0059] If the acquisition is successful, add the interval information to the interval mapping table of the source object, and then allocate the corresponding target large object interval according to the interval, and allocate the corresponding target object interval for the source object interval; otherwise, judge whether the object is deleted and whether there is a segment number record.

[0060] In this step, traverse the small objects on the selected source object batch by batch to avoid memory shortage caused by full traversal, and then add the interval information to the interval mapping table of the source object.

[0061] Finally, read the interval mapping table of the source object, modify the offset offset information and layout of the object to form a continuous address, and generate a new target object mapping relationship table.

[0062] Specifically, modify the offset information of the small object, that is, modify the offset information of the small object on the target large object, and then read the layout information or length information of the small object. In this way, multiple small objects are stored on the target large object in an orderly manner.

[0063] S4. Migrate the small object data on the source object to the target object according to the target object mapping relationship table.

[0064] As Figure 5 shown, after interval arrangement, the source interval of the source object to be migrated and the target interval mapping table of the target object have been saved. According to the generated target object mapping relationship table, realize the mapping migration from the source object to the corresponding interval of the target object. Migrate the data of each object from the object interval table of the source object to the continuous address space defined in the target object table.

[0065] Target interval mapping table: It refers to the corresponding interval (layout) of the small object to be migrated on the target large object, and an ordered list formed by the intervals of multiple small objects.

[0066] Object interval table: It refers to the interval where the small object on the source object (aggregated large object) is located on the source object, including offset information and length (layout) information.

[0067] Target object table: It refers to multiple target objects recorded in the hash list, and is also managed by a key-value table.

[0068] The relationships among the above three types of tables are as follows: The object interval table is the interval corresponding to small objects on large objects. The target object table records the metadata-related information of the target large object. The target interval mapping table refers to the interval mapping relationship of small objects on the source object to those on the target large object.

[0069] During the migration process, if it fails, a retry operation will be performed. If the retry still fails, it will be determined whether to execute destruction or only update the status of the target object based on whether there is already valid data on the target object. At the same time, subsequent resource cleanup work should be carried out in a timely manner.

[0070] The migration in this step includes: data migration, metadata and object table migration.

[0071] (1) Data migration

[0072] First, perform data migration, convert the source / target intervals of the objects to be migrated into the intervals corresponding to the underlying disk data distribution, and start the data migration process.

[0073] Specifically, the data migration task is executed by the background migration thread, reads the interval data corresponding to the small objects in the source object, and then writes it to the new interval marked by the target object. After receiving the write success return when the data is successfully written to the disk, it indicates that the data migration is successful.

[0074] (2) Metadata and object table migration

[0075] Then, after the data migration is completed, perform metadata and object table migration, update the metadata of the target large object, and insert the small object list into the object table LSM tree corresponding to the target object in batches.

[0076] Specifically, after the data migration of the small objects corresponding to the source object is successful, insert the small object list into the object table LSM tree corresponding to the target object in batches, and at the same time insert the metadata of the source object after interval sorting into the metadata LSM tree of the target object.

[0077] S5. After the migration is successful, batch-read the small objects in the source object, and update the layout information corresponding to the metadata of the small objects themselves according to the layout information corresponding to the small objects on the target object.

[0078] As Figure 6 shown, after the data and metadata migration are completed, it is necessary to synchronously update the layout information of the small objects. That is, read the small objects in the source object, and update the layout information corresponding to the metadata of the small objects themselves according to the layout information corresponding to the small objects on the target object to ensure the consistency of the metadata.

[0079] Before updating the layout information of the small information object, read the object table of the source object. First, check whether the small object still exists. If it does not exist, it means that it was deleted during space reorganization, and there is no need to update the layout information.

[0080] Read the target object range table and batch obtain objects for update. When updating the layout information of the small information object, update it in the order of the offset information of the small object, and determine whether the update process is normal. If it is normal, solidify the memory after the update to ensure that the subsequent reorganization task can continue from the failed offset in case of an exception; if it is abnormal, it means that the layout information of some objects is updated successfully. The source object and the target object both have valid data, which is used as the source object for the next space reorganization. This is because the currently created target large object will become the source object for the next space reorganization after this successful execution or failure interruption, which is a repeated process.

[0081] After the metadata migration and update are completed, it means that the space reorganization task is successful. At this time, the target object can provide services externally.

[0082] S6. Perform a cleanup operation to end the task.

[0083] After the layout information is updated, the target object can provide object read, write, and delete services, but the status needs to be updated first. If the target object is not full, space reorganization can continue; if it is full, notify the space reorganization to be released and the status is restored to normal. The client can perform read, write, and delete operations on the target object.

[0084] Specifically, as Figure 7 shown, the background starts the cleanup activities of the space reorganization task, including: updating the status of the target object, cleaning the source object, cleaning the memory structure solidified by the space reorganization, and cleaning the solidified space reorganization task. Then start a new round of space reorganization work.

[0085] (1) Update the status of the target object: Delete the space reorganization mark and update the status of the target object to normal.

[0086] (2) Clean the source object: Clean the metadata and object table of the source object and destroy the source object.

[0087] (3) Clean the task record: Clean the GC memory structure resources and clean the solidified task.

[0088] After the above steps, the valid data on one or more source objects is centrally migrated to the continuous space of one or more target large objects, and the metadata corresponding to the valid data is also updated synchronously, thus ensuring data consistency and integrity. It should be noted that the actual capacity of the target large object is not necessarily a complete 96M. When the size of the last object cannot be fully inserted into the large object space, a certain amount of space is allowed to be reserved at the end.

[0089] The distributed object storage space reorganization method according to the embodiment of the present invention can be applied to the fields of video surveillance, streaming media data processing, and target object state update transmission. It not only needs to transmit a large amount of media data such as audio and video, but also needs to record and analyze a large number of small objects such as logs, texts, and pictures. The distributed object storage space reorganization method according to the embodiment of the present invention can improve data storage and space utilization for media storage with a large number of delete and overwrite write operations.

[0090] In a cloud storage environment, there are also a large number of large and small objects. The present invention can optimize the storage space in the cloud storage system, improve storage efficiency, and reduce storage costs.

[0091] The distributed object storage space reorganization method according to the embodiment of the present invention has the following beneficial effects:

[0092] (1) Small object dynamic migration strategy: By continuously integrating the valid data segments in large objects with a large gap between the valid space and the allocated space into a new empty large object, the purpose of eliminating holes can be achieved, and at the same time, the impact on the system load during space reorganization can also be reduced. This dynamic adjustment and migration strategy can dynamically adjust the threshold according to the cluster capacity usage and system load, and automatically start the background space reorganization task execution. This optimization strategy can enable the system to still maintain high efficiency execution under different loads, and solve the problems of increased storage space fragmentation caused by repeated writing and deletion, or overwrite writing of a large number of small objects, resulting in high complexity of metadata management and affecting object read and write efficiency.

[0093] (2) Object metadata maintenance method: The method for maintaining metadata and object tables ensures the consistency and integrity of data and metadata before and after migration by establishing an interval mapping table. By establishing an interval mapping table to perform data migration, the metadata and object tables are also migrated synchronously, ensuring data consistency and integrity.

[0094] (3) Spatial organization control structure: It makes adaptive dynamic adjustments based on the cluster capacity utilization rate and load, including the number of concurrent tasks, the minimum space threshold, the task execution interval, etc., which improves the spatial organization efficiency and enhances the continuity of data storage. By migrating the undeleted small objects from the original large object to a new empty large object, it effectively reduces the fragmentation and waste of storage space and better improves the storage space utilization rate.

[0095] (4) Storage space optimization: The prior art eliminates holes by migrating data within the space range of the original large object itself, while the present invention migrates the undeleted small objects from the original large object to a new empty large object, effectively reducing the fragmentation and waste of storage space and better improving the storage space utilization rate. The present invention can help users and enterprises use storage space more effectively, reduce fragmentation problems, and lower storage costs.

[0096] (5) Adaptive spatial organization control structure: It adaptively adjusts various parameters of background tasks according to the cluster capacity usage and system load, including the number of concurrencies, task intervals, etc.

[0097] (6) Data continuity: Optimizing the storage space helps create continuous data storage blocks, thereby reducing the overhead of random access and accelerating read and write operations.

[0098] (7) Simplifying operation and maintenance operations: It supports enabling and disabling background spatial organization tasks at any time. After enabling, it adaptively adjusts various parameters according to the current situation of the cluster to improve the hole elimination efficiency. There is no need for complex operations to monitor and manage tasks, and only simple query commands are needed to view the intermediate process and results.

[0099] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the embodiments of the present invention, and not to limit them. Although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention. Any changes or replacements that can be easily thought of by those skilled in the art within the technical scope disclosed by the present invention should be covered within the protection scope of the present invention.

Claims

1. A method for organizing a distributed object storage space, characterized in that, Including: S1. Obtain an aggregated large object that meets the hole rate requirement as the source object, and insert the source object whose fragmented space meets the minimum hole threshold into the hash list; S2. Select the source object with the highest fragmentation rate from the hash list for space arrangement, and create a target aggregated large object as the target object according to requirements; S3. Traverse the object table of the source object in batches, obtain the offset information of the small object on the source object, generate an ordered object table after sorting, obtain the interval information where each small object is located on the large object according to the ordered object table, add the interval information to the interval mapping table of the source object, allocate a corresponding target object interval for the source object interval, and then generate a target object mapping relationship table; S4. Migrate the small object data on the source object to the target object according to the target object mapping relationship table, where First, perform data migration, including: reading the interval data corresponding to the small object in the source object, and then writing it to the new interval marked by the target object. After receiving the write success returned by the data landing on disk, it indicates that the data migration is successful; Then, perform metadata and object table migration, including: batch inserting the small object list into the object table LSM tree corresponding to the target object, and at the same time inserting the metadata after interval arrangement of the source object into the metadata LSM tree of the target object; S5. After the migration is successful, batch read the small objects in the source object, and update the interval layout information corresponding to the metadata of the small object according to the interval layout information corresponding to the small object on the target object; when updating the layout information of the small object, update it in the order of the offset information of the small object, and judge whether the update process is normal. If it is normal, solidify the memory after the update to ensure that the subsequent arrangement task can continue from the failed offset in case of an exception; if it is abnormal, it means that the layout information update of some objects is successful, and the valid data existing on both the source object and the target object is used as the source object for the next space arrangement; S6. Perform a cleanup operation to end the task.

2. The method for organizing a distributed object storage space according to claim 1, wherein The minimum hole threshold is dynamically adjusted according to the cluster capacity utilization rate.

3. A method for organizing a distributed object storage space according to claim 1, characterized in that, Each source object can only select one target object, and each target object corresponds to multiple source objects.

4. A distributed object storage space arrangement method according to claim 1, characterized in that, In S2, when the space arrangement starts to execute, the target object is automatically created in the storage bucket, and the states of the source object and the target object are set to the space arrangement state, where the sum of the effective capacities of the source objects cannot exceed the capacity of the target object.

5. A distributed object storage space arrangement method according to claim 1, characterized in that In S3, the object table information on the source object is saved in the LSM tree, and the object table and the memory of the relevant interval on the source object are obtained through the LSM tree listing interface.

6. A distributed object storage space arrangement method according to claim 1, characterized in that, In S3, obtain the offset information of the small object on the source object, obtain the offset information of the small object on the source object according to the object name and segment number, and sort it from large to small to generate an ordered object table.

7. A distributed object storage space sorting method according to claim 1, characterized in that, In the step S3, after reading the interval mapping table of the source object, modify the offset information and layout information of the small object to form continuous addresses, and generate the target object mapping relationship table.

8. A method for organizing a distributed object storage space according to claim 1, characterized in that, In the step S5, before updating the layout information of the small object, first check whether the small object still exists. If it does not exist, it means that it has been deleted during the space organization, and there is no need to update the layout information.

9. A distributed object storage space sorting method according to claim 1, characterized in that In the step S6, clean up the source object, clean up the memory structure solidified by the space organization, and clean up the solidified space organization tasks.

Citation Information

Patent Citations

  • Method, device and equipment for aggregating multi-version small objects

    CN113821166A

  • Storage object processing method and device, terminal and storage medium

    CN115878027A