A method for cleaning residual data in radosgw sharding

By managing the residual data of radosgw shards through an independent cleaning thread, the database read amplification and user request blocking problems caused by sharding in the Ceph cluster are solved, achieving more efficient data cleaning and service stability.

CN117033357BActive Publication Date: 2025-09-23QINGDAO INSPUR HAIRUO ARTIFICIAL INTELLIGENCE CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310921080.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-07-26
Publication Date
2025-09-23
Estimated Expiration
2043-07-26

AI Technical Summary

Technical Problem

In the existing technology, radosgw sharding causes read amplification and user request blocking problems in the rocksdb database of the ceph cluster osd. In particular, during the sharding process of bucket index data with a large amount of data, the ceph cluster is overloaded when cleaning the index data.

Method used

The cleaning of residual data in shards is separated from the sharding thread and is handled by a separate cleaning thread. By controlling the working time and number of cleaning threads, the cleaning process is flexibly managed, a distributed lock mechanism is used to prevent lock conflicts, and cleaning tasks are performed during low-load periods.

Benefits of technology

This reduces the pressure on the Ceph cluster caused by cleaning up residual index data in shards, avoids the read amplification problem in the RocksDB database, and improves the stability and response speed of the service.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117033357B_ABST
    Figure CN117033357B_ABST
Patent Text Reader

Abstract

The present invention relates to the field of object storage technology, specifically a radosgw shard residual data cleaning method, comprising the following steps: judging whether the working time is met; judging whether the time interval from the last cleaning exceeds the set time interval; judging whether the memory list to be cleaned is empty; judging whether the file list to be cleaned is empty; the beneficial effect is: the radosgw shard residual data cleaning method proposed by the present invention separates the cleaning of shard residual index data from the shard thread, and a separate cleaning thread is responsible for the cleaning. The working time and cleaning quantity of the cleaning thread can be controlled, and whether to clean can also be flexibly controlled according to the cluster pressure. The present invention can significantly reduce the pressure on the cluster caused by the simultaneous cleaning of a large amount of shard residual index data, thereby causing the problem of decreased or peak service capacity, and also avoids the problem of read amplification of the rocksdb database of the backend osd caused by the mixing of a large amount of reading and deleting omap.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of object storage technology, and in particular to a method for cleaning residual data of a radosgw shard. Background Art

[0002] Radosgw is an object storage gateway component provided by Ceph for distributed storage. It uses the librados interface to operate Ceph cluster objects within the cluster, converting the Rados protocol to standard object storage protocols such as S3 or Swift for external use. Bucket index data must be stored for object storage listings. Radosgw stores this index data in the omap entries of Ceph cluster objects.

[0003] In the existing technology, under the default configuration, each bucket in radosgw has only one index object. If the index object contains too many entries, it will cause problems in Ceph cluster balance. Therefore, radosgw provides a dynamic sharding function for bucket index data containing large amounts of data. During sharding, the old Ceph cluster index object is read, its index data is rehashed, and written to the new Ceph cluster index object. After the sharding is completed, the old Ceph cluster index object is deleted.

[0004] However, deletion also requires reading the omap list of the Ceph cluster index objects and deleting each entry one by one. This operation can cause read amplification in the rocksdb database of the Ceph cluster's OSDs. In severe cases, it can cause requests that originally only needed to read tens or hundreds of bytes of data to read tens of megabytes of data. Furthermore, sharding can block requests from object storage users. Summary of the Invention

[0005] The purpose of the present invention is to provide a radosgw shard residual data cleaning method to solve the problems raised in the above background technology.

[0006] To achieve the above object, the present invention provides the following technical solution: a method for cleaning residual data of radosgw shards, the method comprising the following steps:

[0007] Step 1: Determine whether the working time is met; if so, jump to step 2; if not, return to exit the current work and wait for the next wake-up;

[0008] Step 2: Determine whether the time interval since the last cleanup exceeds the set time interval; if so, skip to step 3; if not, return and wait for the next wake-up;

[0009] Step 3: Determine whether the memory list to be cleaned is empty; if the list is empty, jump to step 4; if the list is not empty, jump to step 12;

[0010] Step 4: Determine whether the list of files to be cleaned is empty; if the list is empty, jump to step 5; if the list is not empty, jump to step 12;

[0011] Step 5: Get the list of index objects to be cleaned and proceed to step 6;

[0012] Step 6: Determine whether the list of index objects to be cleaned is empty; if the list is empty, jump to step 12; if the list is not empty, jump to step 7;

[0013] Step 7: The current object is the first object in the list of index objects to be cleaned, so proceed to step 8;

[0014] Step 8: Get the index entry of the current object from the Ceph cluster and proceed to step 9;

[0015] Step 9: Count the total index entries to be cleaned in this cleanup thread workflow and proceed to step 10;

[0016] Step 10: Determine whether the total number of index entries to be cleaned exceeds the memory limit; if so, save the index objects and index entries to be cleaned to the file list to be cleaned; if not, save the index objects and index entries to be cleaned to the memory list to be cleaned;

[0017] Step 11: Delete the current object from the list of index objects to be cleaned up, and return to step 6;

[0018] Step 12: Determine whether the memory list to be cleared is empty; if the list is not empty, jump to step 13; if the list is empty, jump to step 17;

[0019] Step 13: The current object is the first object in the memory list to be cleaned; proceed to step 14;

[0020] Step 14: Determine whether the total number of index entries cleaned exceeds the single cleanup limit; if so, return and wait for the next wake-up; if not, delete the index data of the current object's object index entry cleanup quota, obtain the index list from the memory list, delete the corresponding index and index list corresponding to the index data from the Ceph cluster, and delete it from the memory list; proceed to step 15;

[0021] Step 15: Count the total number of index entries cleaned by this cleaning thread workflow; proceed to step 16;

[0022] Step 16: Determine whether the index data of the current object is empty; if so, delete the current object from the memory list to be cleaned and return to step 12; if not, jump to step 14;

[0023] Step 17: Determine whether the list of files to be cleaned is empty; if the list is not empty, jump to step 18; if the list is empty, return and wait for the next wake-up;

[0024] Step 18: The current object is the first object in the list of files to be cleaned; proceed to step 19;

[0025] Step 19: Determine whether the total number of index entries cleaned exceeds the single cleanup limit; if so, return and wait for the next wake-up; if not, delete the index data of the current object's object index entry cleanup quota, obtain the index list from the file list, delete the corresponding index from the Ceph cluster, and delete the index list corresponding to the index data from the file list; then proceed to step 20;

[0026] Step 20: Count the total number of index entries cleaned by this cleaning thread workflow; proceed to step 21;

[0027] Step 21: Determine whether the index data of the current object is empty; if so, delete the current object from the list of files to be cleaned up, delete the current object from the Ceph cluster, and return to step 12; if not, jump to step 19.

[0028] Preferably, the memory limit is configurable, including selecting the number of entries or the memory size.

[0029] Preferably, radosgw decides whether to start the cleanup thread based on the configuration items, including selecting some radosgw external services that are not affected by this thread; and also including specifying a fixed radosgw to perform this cleanup task exclusively and not provide external services.

[0030] Preferably, the working hours of the radosgw cleanup thread are optional and default to 00:00 to 06:00 in the local time zone.

[0031] Preferably, multiple radosgw cleaning threads use distributed locks to achieve thread exclusivity, and multiple radosgws use the same ceph object as a lock object.

[0032] Preferably, when the cleaning thread starts, it obtains the specified object through the Ceph cluster and detects whether it is locked by other radosgw cleaning threads. If it is locked, it waits for the next wake-up. If it is not locked, it locks the object, enters the workflow, and unlocks the object after the work is completed.

[0033] Preferably, during the cleanup process, the lock added to the specified object has a timeout release mechanism to prevent radosgw from crashing and causing the lock to be unable to be released, and the timeout period is configurable.

[0034] Preferably, during the cleaning process, the radosgw cleaning thread periodically refreshes the lock of the specified object to prevent the work from being unfinished, the lock being released in advance, and being acquired by other radosgw cleaning threads.

[0035] Compared with the prior art, the present invention has the following beneficial effects:

[0036] The Radosgw shard residual data cleaning method proposed in this invention separates the cleaning of shard residual index data from the sharding thread, and assigns it to a separate cleaning thread. The working time and cleaning quantity of the cleaning thread can be controlled, and whether to perform cleaning can also be flexibly controlled according to cluster pressure. This method can significantly reduce the pressure on the cluster caused by the simultaneous cleaning of a large amount of shard residual index data, which can lead to reduced or peak server capacity. It also avoids the problem of read amplification of the backend OSD's RocksDB database caused by the mixing of large amounts of read and delete OMAP. BRIEF DESCRIPTION OF THE DRAWINGS

[0037] Figure 1 Flow chart of the method of the present invention. DETAILED DESCRIPTION

[0038] In order to clearly and completely describe the objectives and technical solutions of the present invention and make the advantages more clearly understood, the embodiments of the present invention are further described in detail below with reference to the accompanying drawings. It should be understood that the specific embodiments described herein are part of the embodiments of the present invention, not all of them, and are only used to explain the embodiments of the present invention, not to limit the embodiments of the present invention. All other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.

[0039] Example 1

[0040] See also Figure 1 The present invention provides a technical solution: a method for cleaning residual data of radosgw fragments, the method comprising the following steps:

[0041] Step 1: Determine whether the working time is met; if so, jump to step 2; if not, return to exit the current work and wait for the next wake-up;

[0042] Step 2: Determine whether the time interval since the last cleanup exceeds the set time interval; if so, jump to step 3; if not, return and wait for the next wake-up;

[0043] Step 3: Determine whether the memory list to be cleaned is empty; if the list is empty, jump to step 4; if the list is not empty, jump to step 12;

[0044] Step 4: Determine whether the list of files to be cleaned is empty; if the list is empty, jump to step 5; if the list is not empty, jump to step 12;

[0045] Step 5: Get the list of index objects to be cleaned and proceed to step 6;

[0046] Step 6: Determine whether the list of index objects to be cleaned is empty; if the list is empty, jump to step 12; if the list is not empty, jump to step 7;

[0047] Step 7: The current object is the first object in the list of index objects to be cleaned, so proceed to step 8;

[0048] Step 8: Get the index entry of the current object from the Ceph cluster and proceed to step 9;

[0049] Step 9: Count the total index entries to be cleaned in this cleanup thread workflow and proceed to step 10;

[0050] Step 10: Determine whether the total number of index entries to be cleaned exceeds the memory limit; if so, save the index objects and index entries to be cleaned to the file list to be cleaned; if not, save the index objects and index entries to be cleaned to the memory list to be cleaned;

[0051] Step 11: Delete the current object from the list of index objects to be cleaned up, and return to step 6;

[0052] Step 12: Determine whether the memory list to be cleared is empty; if the list is not empty, jump to step 13; if the list is empty, jump to step 17;

[0053] Step 13: The current object is the first object in the memory list to be cleaned; proceed to step 14;

[0054] Step 14: Determine whether the total number of index entries cleaned exceeds the single cleanup limit; if so, return and wait for the next wake-up; if not, delete the index data of the current object's object index entry cleanup quota, obtain the index list from the memory list, delete the corresponding index and index list corresponding to the index data from the Ceph cluster, and delete it from the memory list; proceed to step 15;

[0055] Step 15: Count the total number of index entries cleaned by this cleaning thread workflow; proceed to step 16;

[0056] Step 16: Determine whether the index data of the current object is empty; if so, delete the current object from the memory list to be cleaned and return to step 12; if not, jump to step 14;

[0057] Step 17: Determine whether the list of files to be cleaned is empty; if the list is not empty, jump to step 18; if the list is empty, return and wait for the next wake-up;

[0058] Step 18: The current object is the first object in the list of files to be cleaned; proceed to step 19;

[0059] Step 19: Determine whether the total number of index entries cleaned exceeds the single cleanup limit; if so, return and wait for the next wake-up; if not, delete the index data of the current object's object index entry cleanup quota, obtain the index list from the file list, delete the corresponding index from the Ceph cluster, and delete the index list corresponding to the index data from the file list; then proceed to step 20;

[0060] Step 20: Count the total number of index entries cleaned by this cleaning thread workflow; proceed to step 21;

[0061] Step 21: Determine whether the index data of the current object is empty; if so, delete the current object from the list of files to be cleaned up, delete the current object from the Ceph cluster, and return to step 12; if not, jump to step 19.

[0062] Example 2

[0063] Based on the first embodiment, the memory limit is configurable, including the number of items or memory size; radosgw decides whether to start the cleaning thread according to the configuration items, including selecting some radosgw external services that are not affected by this thread; it also includes specifying a fixed radosgw to perform this cleaning task exclusively and not provide external services; the working time of the radosgw cleaning thread is optional, and the default is 00:00 to 06:00 in the local time zone; multiple radosgw cleaning threads use distributed locks to achieve thread exclusivity, and multiple radosgws use the same ceph object as the lock object; cleaning When the cleaning thread starts, it obtains the specified object through the Ceph cluster and checks whether it is locked by other radosgw cleaning threads. If it is locked, it waits for the next wake-up. If it is not locked, it locks the object, enters the workflow, and unlocks the object after the work is completed. During the cleaning process, the lock added to the specified object has a timeout release mechanism to prevent radosgw from crashing and causing the lock to be unable to be released, and the timeout period is configurable. During the cleaning process, the radosgw cleaning thread regularly refreshes the lock of the specified object to prevent the work from being unfinished, the lock is released early, and it is acquired by other radosgw cleaning threads.

[0064] While embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions, and variations may be made to these embodiments without departing from the principles and spirit of the invention, and that the scope of the invention is defined by the appended claims and their equivalents.

Claims

1. A radosgw sharding residual data cleaning method, characterized by: The method comprises the following steps: Step 1: Determine whether the working time is met; if so, jump to step 2; if not, return to exit the current work and wait for the next wake-up; Step 2: Determine whether the time interval since the last cleanup exceeds the set time interval; if so, jump to step 3; if not, return and wait for the next wake-up; Step 3: Determine whether the memory list to be cleaned is empty; if the list is empty, jump to step 4; if the list is not empty, jump to step 12; Step 4: Determine whether the list of files to be cleaned is empty; if the list is empty, jump to step 5; if the list is not empty, jump to step 12; Step 5: Get the list of index objects to be cleaned and proceed to step 6; Step 6: Determine whether the list of index objects to be cleaned is empty; if the list is empty, jump to step 12; if the list is not empty, jump to step 7; Step 7: The current object is the first object in the list of index objects to be cleaned, so proceed to step 8; Step 8: Get the index entry of the current object from the Ceph cluster and proceed to step 9; Step 9: Count the total index entries to be cleaned in this cleanup thread workflow and proceed to step 10; Step 10: Determine whether the total number of index entries to be cleaned exceeds the memory limit; if so, save the index objects and index entries to be cleaned to the file list to be cleaned; if not, save the index objects and index entries to be cleaned to the memory list to be cleaned; Step 11: Delete the current object from the list of index objects to be cleaned up, and return to step 6; Step 12: Determine whether the memory list to be cleared is empty; if the list is not empty, jump to step 13; if the list is empty, jump to step 17; Step 13: The current object is the first object in the memory list to be cleaned; proceed to step 14; Step 14: Determine whether the total number of index entries cleaned exceeds the single cleanup limit; if so, return and wait for the next wake-up; if not, delete the index data of the current object's object index entry cleanup quota, obtain the index list from the memory list, delete the corresponding index and index list corresponding to the index data from the Ceph cluster, and delete it from the memory list; proceed to step 15; Step 15: Count the total number of index entries cleaned by this cleaning thread workflow; proceed to step 16; Step 16: Determine whether the index data of the current object is empty; if so, delete the current object from the memory list to be cleaned and return to step 12; if not, jump to step 14; Step 17: Determine whether the list of files to be cleaned is empty; if the list is not empty, jump to step 18; if the list is empty, return and wait for the next wake-up; Step 18: The current object is the first object in the list of files to be cleaned; proceed to step 19; Step 19: Determine whether the total number of index entries cleaned exceeds the single cleanup limit; if so, return and wait for the next wake-up; if not, delete the index data of the current object's object index entry cleanup quota, obtain the index list from the file list, delete the corresponding index from the Ceph cluster, and delete the index list corresponding to the index data from the file list; then proceed to step 20; Step 20: Count the total number of index entries cleaned by this cleaning thread workflow; proceed to step 21; Step 21: Determine whether the index data of the current object is empty; if so, delete the current object from the list of files to be cleaned up, delete the current object from the Ceph cluster, and return to step 12; if not, jump to step 19.

2. A radosgw shard residual data cleaning method according to claim 1, characterized in that: The memory limit is configurable, including selecting the number of entries or the memory size.

3. A radosgw shard residual data cleaning method according to claim 1, characterized in that: radosgw decides whether to start the cleanup thread based on the configuration items, including selecting some radosgw external services that are not affected by this thread; it also includes specifying a fixed radosgw to perform this cleanup task exclusively and not provide external services.

4. A radosgw shard residual data cleaning method according to claim 1, characterized in that: The working hours of the radosgw cleanup thread are optional and default to 00:00 to 06:00 in the local time zone.

5. A radosgw shard residual data cleaning method according to claim 1, characterized in that: Multiple radosgw cleaning threads use distributed locks to achieve thread exclusivity, and multiple radosgws use the same ceph object as the lock object.

6. A radosgw shard residual data cleaning method according to claim 1, characterized in that: When the cleaning thread starts, it obtains the specified object through the Ceph cluster and checks whether it is locked by other radosgw cleaning threads. If it is locked, it waits for the next wake-up. If it is not locked, it locks the object, enters the workflow, and unlocks the object after the work is completed.

7. A radosgw shard residual data cleaning method according to claim 6, characterized in that: During the cleanup process, the locks added to the specified objects have a timeout release mechanism to prevent radosgw from crashing and causing the locks to be unable to be released. The timeout period is configurable.

8. A radosgw shard residual data cleaning method according to claim 6, characterized in that: During the cleanup process, the radosgw cleanup thread periodically refreshes the lock of the specified object to prevent the work from being incomplete, the lock being released prematurely, and being acquired by other radosgw cleanup threads.

Citation Information

Patent Citations

  • Method for automatically cleaning and maintaining ElasticSearch log index file

    CN106649461A

  • Data sanitization

    US20160300069A1