A copy-based shard non-inductive capacity expansion implementation method
By adopting the replicated Shard-free expansion method in CEPH object storage, creating a new bucket instance and synchronizing data, and using errorlog to record failed operations, the problem of users being unable to operate during Shard expansion is solved, improving user experience and read and write performance.
Patent Information
- Application Number
- CN202311712645.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-13
- Publication Date
- 2025-10-14
- Estimated Expiration
- 2043-12-13
AI Technical Summary
CEPH object storage cannot perform user operations during Shard expansion, which affects the user experience. The existing Reshard method cannot allow users to upload, delete, and other operations during the expansion process.
A replication-based Shard-free expansion method is used. By creating a new bucket instance, synchronizing data to the new Shard, using the error log to record failed operations, and performing a full comparison to ensure data consistency, users are allowed to perform operations during the expansion process.
This enables user-unnoticed operations during Shard expansion, improving user experience. It also improves read and write performance after expansion and ensures data consistency.
Smart Images

Figure CN117851402B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of computer storage technology, and in particular relates to a method for realizing Shard seamless expansion based on replication. Background Art
[0002] CEPH is currently the most widely used distributed storage system, capable of providing read and write services through parallel processing on multiple servers. CEPH boasts excellent scalability and reliability, and its applications include block storage (RBD, Rados Block Device), file storage (CephFS, Ceph file system), and object storage (RADOSGW, Reliable Autonomic Distributed Object Storage Gateway). CEPH's object storage uses a data and metadata separation approach when providing read and write services. Data is generally stored in the data storage pool, and metadata is stored in the index storage pool. When using object storage for data read and write, users upload or download data files to their corresponding buckets through REST API requests. Buckets store metadata (index data) in the bucket's shards. When a user stores too many objects in a bucket, the bucket's shards may become overloaded with indexes, impacting read and write performance. Shard expansion is necessary to reduce the number of metadata entries per shard.
[0003] Currently, CEPH supports two methods to modify the number of shards in a bucket: dynamic resharding and static resharding. Dynamic resharding means that after a user uploads an object to object storage, the object storage will first determine whether the number of metadata on the shard exceeds the preset maximum number of metadata before storing the metadata in the bucket. If it does, the reshard task will be added to the queue, and then resharding will be triggered by a background scheduled task to reduce the number of metadata on the shard. Static resharding, on the other hand, means resharding by issuing the radosgw-admin command through the monitor node of the cluster. This has the advantage of manually controlling the resharding time and the specific number of shards after resharding.
[0004] However, there is a problem with both dynamic and static resharding: the shard of the bucket cannot be operated during the reshard expansion, resulting in users being unable to upload, delete, or perform other operations involving object metadata modification during the bucket reshard period. When the number of metadata records in the user's bucket is too large, the shard expansion time is very long, which will seriously affect the user experience. Therefore, the need to design a method for seamless shard expansion is becoming increasingly urgent. Summary of the Invention
[0005] The purpose of the present invention is to provide a method for realizing Shard seamless expansion based on replication, aiming to solve the problems raised in the above background technology.
[0006] To achieve the above object, the present invention provides the following technical solutions:
[0007] A method for implementing Shard seamless expansion based on replication includes the following steps:
[0008] Step S1: Create a new bucket instance for the bucket that needs to be resharded and start preparing for reshard expansion;
[0009] Step S2: Set the bucket status to Resharding, indicating that the bucket in step S1 has started the Reshard operation;
[0010] Step S3: Add the ID of the newly created bucket instance to the bucket information, indicating that if the user uploads or deletes data later, the new bucket instance needs to be operated synchronously;
[0011] Step S4: Traverse each shard of the old bucket and rebalance the data on the shard to the new bucket shard instance;
[0012] Step S5: Traverse the error log and query the reshard process. Find any operations that failed on the new bucket shard due to concurrent data synchronization from the old shard to the new shard, and re-synchronize and apply them to the new bucket shard.
[0013] Step S6: Use a full comparison tool to ensure data consistency between the new shard and the old shard.
[0014] Step S7: Delete the old bucket shard to complete the seamless expansion of the bucket.
[0015] As a preferred solution of the present invention, in step S4, the data is balanced to the new bucketshard instance through a hash algorithm.
[0016] As a preferred solution of the present invention, it further includes adding a data structure, wherein the added data structure is new_bucket_instance_id and errorlog.
[0017] As a preferred solution of the present invention, the new_bucket_instance_id is used to mark the newly generated bucket id in the Reshard process.
[0018] As a preferred solution of the present invention, the errorlog is used to record the success of the operation on the original Shard and the failure of the operation on the new Shard during the Reshard process, and synchronously update the records to the new Shard.
[0019] As a preferred solution of the present invention, the structure of the errorlog is:
[0020]
[0021] As a preferred solution of the present invention, in step S5, the business process and Reshard will be concurrent, causing the two processes to overlap each other, and then modification is performed. The specific steps of modification include adding metadata records, modifying metadata records, deleting metadata records, errorlog records and consumption, and full comparison.
[0022] As a preferred solution of the present invention, the steps of recording the newly added metadata are as follows:
[0023] I. When a new object is uploaded during the Reshard process, the Reshard entry is obtained at time t1, written to the old bucket shard at time t2, and written to the new bucket shard at time t3;
[0024] Ⅱ. If writing fails, it will be recorded in the errorlog and background synchronization will be performed through the errorlog.
[0025] As a preferred solution of the present invention, the modified metadata is recorded as follows:
[0026] Ⅰ. During the resharding process, metadata modification operations are performed. First, the index on the original shard is updated, then the index on the original shard is read and synchronized to the new shard.
[0027] Ⅱ. If synchronization fails, it needs to be recorded in the errorlog and synchronized later.
[0028] As a preferred solution of the present invention, the full comparison is used to ensure strong consistency of data. The full comparison is performed twice.
[0029] Compared with the prior art, the present invention has the following beneficial effects:
[0030] 1. The present invention can allow users to delete, upload, and modify metadata records during the Shard expansion process, so that Shard expansion can be performed in the background without affecting user business, and can improve user read and write performance after the Shard expansion is completed.
[0031] 2. In the scenario where users modify index metadata, the present invention uses errorlog to save the index records of the new Shard operation and performs background consumption to ensure data consistency. A Sharding_status is added to the metadata of each object to indicate whether the object is being synchronized. If it is being synchronized, index modification, deletion, or deletion operations are not allowed, reducing the number of records in the errorlog.
[0032] 3. In order to ensure strong data consistency in the present invention, a full comparison is required before Reshard is completed to ensure that the data on the new Shard is completely consistent with the original Shard. BRIEF DESCRIPTION OF THE DRAWINGS
[0033] The accompanying drawings are used to provide a further understanding of the present invention and constitute a part of the specification. Together with the embodiments of the present invention, they are used to explain the present invention and do not constitute a limitation of the present invention. In the accompanying drawings:
[0034] Figure 1 This is an overall flow chart of a Shard senseless expansion method based on replication in the present invention;
[0035] Figure 2 This is a Reshard replication flow chart in a Shard senseless expansion implementation method based on replication of the present invention;
[0036] Figure 3 This is a schematic diagram of the timing of adding metadata records in a Shard seamless expansion method based on replication of the present invention;
[0037] Figure 4 This is a flowchart of modifying metadata records in a Shard senseless expansion method based on replication of the present invention;
[0038] Figure 5This is a timing diagram of Reshard modification in a Shard senseless expansion method based on replication of the present invention;
[0039] Figure 6 This is a flowchart of a method for implementing Shard seamless expansion based on replication in the present invention;
[0040] Figure 7 This is a timing diagram of deleting metadata records in a Shard senseless expansion method based on replication of the present invention;
[0041] Figure 8 This is a flowchart of errorlog background consumption in a Shard senseless expansion method based on replication of the present invention;
[0042] Figure 9 This is a full comparison flowchart in a method for implementing Shard seamless expansion based on replication in the present invention. DETAILED DESCRIPTION
[0043] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.
[0044] Example 1
[0045] See also Figures 1-9 , the present invention provides the following technical solutions:
[0046] A method for implementing Shard seamless expansion based on replication includes the following steps:
[0047] Step S1: Create a new bucket instance for the bucket that needs to be resharded and start preparing for reshard expansion;
[0048] Step S2: Set the bucket status to Resharding, indicating that the bucket in step S1 has started the Reshard operation;
[0049] Step S3: Add the ID of the newly created bucket instance to the bucket information, indicating that if the user uploads or deletes data later, the new bucket instance needs to be operated synchronously;
[0050] Step S4: Traverse each shard of the old bucket and rebalance the data on the shard to the new bucket shard instance;
[0051] Step S5: Traverse the error log and query the reshard process. Find any operations that failed on the new bucket shard due to concurrent data synchronization from the old shard to the new shard, and re-synchronize and apply them to the new bucket shard.
[0052] Step S6: Use a full comparison tool to ensure data consistency between the new shard and the old shard.
[0053] Step S7: Delete the old bucket shard to complete the seamless expansion of the bucket.
[0054] In a specific embodiment of the present invention, the present invention addresses the issue of data on a shard being unable to be modified during CEPH object storage bucket shard expansion, resulting in users being unable to upload or delete the bucket. This method proposes a Shard-free expansion method based on replication, enabling users to upload and delete objects simultaneously during bucket shard expansion, thereby improving the user experience. Furthermore, after resharding, the number of index entries on each shard in the user's bucket can be reduced, thereby improving the speed at which users can read, write, or delete objects.
[0055] Specifically, the main process of the present invention is as follows Figure 1As shown in the figure, when expanding Shard capacity, you first need to create a bucket instance and set the Resharding bucket status to Rsharding, indicating that the bucket is undergoing Resharding operations; add a new flag bit -new_bucket_instance_id in bucket info, indicating that if the user performs upload and download operations at this time, the bucket instance needs to be operated synchronously at the same time; during the Reshard process, the metadata on the original bucket Shard needs to be migrated. During the migration process, if the metadata record exists on the new bucket Shard, the metadata on the original bucket Shard will not be overwritten. If the user fails to upload or download the new bucket instance during the migration process, the corresponding information will be recorded in the errorlog; during the migration process, the metadata record in the errorlog on the new bucket Shard needs to be synchronized, and the Resharding bucket status needs to be changed to Complete; finally, change the new_bucket_instance_id in bucket info to the new bucket instance id, and delete the old bucket instance to complete the seamless Reshard expansion;
[0056] The main data structures added in this invention are: (1) new_bucket_instance_id, which marks the newly generated bucket id during the reshard process; (2) errorlog, which mainly records the records of successful operations on the original shard but failed operations on the new shard during the reshard process. Therefore, the background thread needs to synchronize these operations to the new shard; at the same time, in order to ensure strong data consistency, after all index data is migrated to reshard, it is necessary to perform a full comparison of the indexes on the original shard and the new shard, and then update the bucket id (rgw_link_bucket);
[0057] The errorlog structure is as follows (ErrorOmapKey):
[0058] OMAP<string,errorkey>
[0059] The key is a string, and the key is the currently obtained timestamp. The time when the operation on the new Shard failed can be obtained through the key value. The structure of errorkey is:
[0060]
[0061] The errorkey mainly records the bucket_name, shard_id, and corresponding obj_key of the failed operation. Due to issues such as multiple versions, multiple obj_keys may exist in the same operation. These different obj_keys are a transaction operation and must ensure consistency. There are two types of op_type: one is set, which means reading the corresponding obj from the original shard and then synchronizing it to the new shard; the other is delete, which means deleting the corresponding obj from the new shard.
[0062] Specifically, such as Figure 2 As shown in the figure, if the current bucket is resharding, writing the bucket index will no longer be blocked and will become a double write. That is, the original shard is written first, and then the new shard is written. The subsequent process will not proceed until both are written successfully. Writing to the shard includes modifying the omap header, adding the omap key, and deleting the omap key. The business process includes uploading objects, deleting objects, and modifying object attributes.
[0063] Due to the concurrency of business processes and Reshard, there may be a problem of two processes overlapping each other for the same omapkey. To modify the process specifically:
[0064] (1) Add metadata records
[0065] like Figure 3 As shown in the figure, if a new object is uploaded during the reshard process, the reshard entry is obtained at time t1, written to the old bucket shard at time t2, and written to the new bucket shard at time t3. If the write fails, it is recorded in the error log and synchronized in the background through the error log. In this scenario, whether the write process or the reshard process writes the entry first will not affect the data and will not cause data inconsistency.
[0066] (2) Modify metadata records
[0067] like Figure 4As shown in the figure, during the resharding process, metadata modification operations must first update the index on the original shard, then read the index on the original shard and synchronize it to the new shard. If the synchronization fails, it needs to be recorded in the error log and synchronized later.
[0068] like Figure 5 However, in the following scenario, at time t1, the reshard retrieves the entry, at time t2, the write process updates the old shard, at time t3, the write process writes to the new shard, and at time t4, the reshard process synchronizes the old entry with the same key to the new shard, overwriting the original entry. In this case, the metadata record is directly read from the old shard, modified, and written to the new shard. If an entry with the same key already exists in the new shard during the reshard process, the key is ignored.
[0069] (3) Deleting metadata records
[0070] like Figure 6 As shown in the figure, during the resharding process, the deletion process, during the resharding process, if a request to delete an object is made at the same time, the index metadata on the original shard is deleted first, and then the index metadata on the new shard is deleted. If the deletion of the index metadata on the new shard fails, an error log is recorded; if the deletion of the index on the original shard is successful, a success message is returned; otherwise, a failure message is returned.
[0071] like Figure 7 If the value is not displayed, it indicates that a deletion operation is being performed during the resharding process. This may result in the same situation as modifying metadata records: the user request is successful, but the metadata is not deleted, that is, the object is not deleted successfully. In this case, the redundant index metadata on the new shard is deleted through a subsequent full comparison.
[0072] (4) Errorlog recording and consumption
[0073] like Figure 8As shown, errorlog is mainly recorded in the new rados object of the log pool, and double writing is performed in the uploading, modifying and deleting processes, i.e., the original Shard is written first, and then the new Shard is written; if the original Shard is successfully written and the new Shard fails to be written, the bucket_name, shard_id of the original Shard and the key value of the index metadata of the op operation are required; and the errorlog is consumed by a background thread, and the index of the errorlog record is synchronized in the background when the cluster starts; it is worth noting that when the errorlog is obtained, it needs to be locked, so as to prevent multiple threads from synchronizing the same record when consuming.
[0074] (5) Full comparison tool
[0075] Full comparison is mainly to ensure the strong consistency of data, and is a bottom-up means for possible problems in the Reshard process; full comparison mainly has two times: (1) comparing the new Shard with the original Shard, first listing all the index metadata on the original Shard, if there are some index metadata on the original Shard but no index metadata on the new Shard, the index metadata on the original Shard is synchronized to the new Shard; (2) comparing the index data on the original Shard with the new Shard, if there are some index metadata on the new Shard but no index metadata on the original Shard, the index data on the new Shard is deleted.
[0076] Specifically, the present invention is applicable to object storage services developed based on CEPH. Too many objects in an object storage bucket will affect the user's read and write performance, and Shard expansion is required. This method can achieve user-unaware Shard expansion. It can allow users to delete, upload, and modify metadata records during the Shard expansion process, so that Shard expansion can be performed in the background without affecting user services. After the Shard expansion is completed, the user's read and write performance can be improved. For scenarios where users modify index metadata, the errorlog is used to save the index records of the new Shard and perform background consumption to ensure data consistency. A Sharding_status is added to the metadata of each object to indicate whether the object is being synchronized. If it is being synchronized, index modification, deletion, or deletion operations are not allowed, thereby reducing the number of records in the errorlog. In order to ensure strong data consistency, a full comparison is required before Reshard is completed to ensure that the data on the new Shard is completely consistent with the data on the original Shard.
[0077] Finally, it should be noted that the above descriptions are merely preferred embodiments of the present invention and are not intended to limit the present invention. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art will be able to modify the technical solutions described in the aforementioned embodiments or substitute equivalents for some of the technical features. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention shall be included within the scope of protection of the present invention.
Claims
1. A Shard seamless expansion method based on replication, characterized in that: The steps include: Step S1: Create a new bucket instance for the bucket that needs to be resharded and start preparing for reshard expansion; Step S2: Set the bucket status to Resharding, indicating that the bucket in step S1 has started the Reshard operation; Step S3: Add the ID of the newly created bucket instance to the bucket information, indicating that if the user uploads or deletes data later, the operation needs to be performed on the new bucket instance simultaneously; Step S4: Traverse each shard of the old bucket and rebalance the data on the shard to the new bucket shard instance; Step S5: Traverse the error log and query the reshard process. Find any operations that failed on the new bucket shard due to concurrent data synchronization from the old shard to the new shard, and re-synchronize and apply them to the new bucket shard. Step S6: Use a full comparison tool to ensure data consistency between the new shard and the old shard. Step S7: Delete the old bucket shard to complete the seamless expansion of the bucket.
2. A Shard seamless expansion method based on replication according to claim 1, characterized in that: In step S4, the data is balanced to the new bucket shard instance through a hash algorithm.
3. The method for realizing Shard capacity expansion based on replication according to claim 2 is characterized in that: It also includes adding data structures, where the added data structures are new_bucket_instance_id and errorlog.
4. A Shard seamless expansion method based on replication according to claim 3, characterized in that: The new_bucket_instance_id is used to mark the newly generated bucket ID during the Reshard process.
5. The method for realizing Shard capacity expansion based on replication according to claim 4 is characterized in that: The errorlog is used to record the success of the original shard operation and the failure of the new shard operation during the reshard process, and is synchronously updated to the new shard.
6. A Shard seamless expansion method based on replication according to claim 5, characterized in that: The structure of the errorlog is: struct errorkey{ string bucket_name; int shard_id; map<obj_key,op_type> }。 7. The method for realizing Shard capacity expansion based on replication according to claim 6 is characterized in that: In step S5, the business process and Reshard will run concurrently, causing the two processes to overlap each other, so they are modified. The specific steps of modification include adding metadata records, modifying metadata records, deleting metadata records, error log records and consumption, and full comparison.
8. The method for realizing Shard capacity expansion based on replication according to claim 7 is characterized in that: The steps for recording new metadata are as follows: I. When a new object is uploaded during the Reshard process, the Reshard entry is obtained at time t1, written to the old bucket shard at time t2, and written to the new bucket shard at time t3; Ⅱ. If writing fails, it will be recorded in the errorlog and background synchronization will be performed through the errorlog.
9. The method for realizing Shard capacity expansion based on replication according to claim 8 is characterized in that: The modified metadata is recorded as follows: Ⅰ. During the resharding process, metadata modification operations are performed. First, the index on the original shard is updated, then the index on the original shard is read and synchronized to the new shard. Ⅱ. If synchronization fails, it needs to be recorded in the errorlog and synchronized later.
10. A Shard seamless expansion method based on replication according to claim 9, characterized in that: The full comparison is used to ensure strong consistency of data, and the full comparison occurs twice.
Citation Information
Patent Citations
Method for achieving system dynamic expansion in shared-nothing database cluster
CN102521297A
Method and system for online migration of DRDS sub-libraries and sub-tables based on canal
CN116303348A