Data synchronization method and device, computer equipment and storage medium
By synchronizing metadata in TiDB database, the performance problems caused by too many or too few shards in the bucket are solved, ensuring data consistency and integrity.
Patent Information
- Application Number
- CN202510513073.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-03
- Publication Date
- 2025-08-05
- Estimated Expiration
- 2045-01-03
AI Technical Summary
In traditional technology, too many or too few shards in the bucket lead to insufficient concurrent storage capacity, delay in enumerating object data, and business interruption when the number of shards is dynamically adjusted.
By obtaining the first metadata corresponding to each shard of the target bucket in the object index pool during the target migration stage, and using forward and reverse comparison results, the first metadata is synchronized to the TiDB database, which solves the problem of too many or too few shards in the bucket and ensures that the metadata has no dirty data in the database.
It realizes efficient storage of metadata in TiDB database, avoids performance problems caused by too many or too few shards, and ensures data consistency and integrity.
Smart Images

Figure CN120429362A_ABST
Abstract
Description
[0001] This application is a divisional application for application number: 2025100072652 (data synchronization method, device, computer equipment and storage medium, application date: January 3, 2025). Technical Field
[0002] The present application relates to the technical field of data storage, and in particular to a data synchronization method, apparatus, computer equipment, and storage medium. Background Art
[0003] Metadata can be understood as data that describes the attributes of data. It has functions such as recording storage location, recording historical data, and facilitating data retrieval. With the rapid development of computer technology, various types of data are experiencing explosive growth, and managing and storing metadata for these types of data is facing new challenges.
[0004] Traditionally, a new bucket is created to store object data. This bucket is then divided into a specified number of shards to store metadata related to the object data. For example, when an object is uploaded, a consistent hashing algorithm is used to calculate the target shard for the object's metadata, and the metadata is stored in the target shard.
[0005] However, if there are too few shards in a bucket, the concurrent data storage capacity will be insufficient; if there are too many shards, the object data listing delay will be long; and if the number of shards is adjusted dynamically, business interruption will occur. Summary of the Invention
[0006] Based on this, it is necessary to provide a data synchronization method, device, computer equipment and storage medium to address the above technical problems, which can solve the problems caused by too many or too few shards in the bucket and dynamic adjustment of the number of shards.
[0007] In a first aspect, the present application provides a data synchronization method, the method comprising:
[0008] Obtaining first metadata corresponding to each shard of a target bucket in an object index pool in a target migration phase; wherein the target migration phase is a first migration phase, a second migration phase, or a third migration phase, the end time of the first migration phase is the start time of the second migration phase, and the end time of the second migration phase is the start time of the third migration phase;
[0009] Metadata synchronization is performed on the database based on a forward comparison result and a reverse comparison result between the first metadata and second metadata stored in the database; wherein the forward comparison result is a result of comparing whether the first metadata exists in the second metadata; and the reverse comparison result is a result of comparing whether the second metadata exists in the first metadata.
[0010] In one embodiment, when the target migration phase is the first migration phase, the first metadata is metadata stored in the object index pool before the start time of the first migration phase;
[0011] The step of synchronizing metadata on the database according to a forward comparison result and a reverse comparison result between the first metadata and second metadata stored in the database includes:
[0012] Comparing the first metadata with metadata stored in the database before the start of the first migration phase to obtain a forward comparison result;
[0013] If the forward comparison result is that the first target metadata in the first metadata does not exist in the database, synchronizing the first target metadata to the database;
[0014] Comparing all metadata stored in the database with the first metadata to obtain a reverse comparison result;
[0015] If the reverse comparison result shows that the second target metadata exists in the database and the second target metadata does not exist in the first metadata, the second target metadata is deleted from the database.
[0016] In one embodiment, when the target migration phase is the second migration phase, the first metadata is metadata stored in the object index pool between the start time of the first migration phase and the end time of the first migration phase;
[0017] The step of synchronizing metadata on the database according to a forward comparison result and a reverse comparison result between the first metadata and second metadata stored in the database includes:
[0018] Comparing the first metadata with metadata stored in the database between the start time and the end time of the first migration phase to obtain a forward comparison result;
[0019] If the forward comparison result is that the third target metadata in the first metadata does not exist in the database, synchronizing the third target metadata to the database;
[0020] Comparing the metadata stored in the database between the start time and the end time of the first migration phase with the first metadata to obtain a reverse comparison result;
[0021] If the reverse comparison result shows that the fourth target metadata exists in the database and the fourth target metadata does not exist in the first metadata, deleting the fourth target metadata from the database;
[0022] After the metadata for non-read operations in the target bucket is locked, the header data of each shard of the target bucket is migrated to the database.
[0023] In one embodiment, during the process of synchronizing metadata on the database based on a forward comparison result and a reverse comparison result between the first metadata and second metadata stored in the database, the method further includes:
[0024] In the second migration phase, a deletion mark is added to the metadata to be deleted;
[0025] The metadata to be deleted is metadata in the database corresponding to the data deletion operation, and the metadata to be deleted is metadata shared by the database and the object index pool.
[0026] In one embodiment, when the target migration phase is the third migration phase, the first metadata is all metadata stored in the object index pool up to the end time of the second migration phase;
[0027] The step of synchronizing metadata on the database according to a forward comparison result and a reverse comparison result between the first metadata and second metadata stored in the database includes:
[0028] Comparing the first metadata with all metadata stored in the database to obtain a positive comparison result;
[0029] If the forward comparison result shows that the fifth target metadata in the first metadata does not exist in the database, synchronizing the fifth target metadata to the database; wherein the update time of the fifth target metadata is less than the start time of the third migration phase;
[0030] Comparing all metadata stored in the database with the first metadata to obtain a reverse comparison result;
[0031] If the reverse comparison result shows that the sixth target metadata exists in the database and does not exist in the first metadata, deleting the sixth target metadata; wherein the update time of the sixth target metadata is less than the start time of the third migration phase;
[0032] Delete the metadata to be deleted that is marked for deletion in the database.
[0033] In one embodiment, the method further comprises:
[0034] If the forward comparison result shows that the first metadata exists in the second metadata, the metadata with the latest update time between the first metadata and the second metadata is synchronized to the database.
[0035] In one embodiment, the method further comprises:
[0036] Obtaining the number of threads used for comparing the first metadata with the second metadata;
[0037] Determining, according to the number of shards of the target bucket in the object index pool and the number of threads, a shard allocated to each thread for comparison with the second metadata;
[0038] Controlling each thread to perform a forward comparison operation on the first metadata in the corresponding shard and the second metadata to obtain a forward comparison result;
[0039] determining, according to the amount of the second metadata and the number of threads, second metadata allocated to each thread for comparison with the first metadata;
[0040] Each thread is controlled to perform a reverse comparison operation on the corresponding second metadata and the first metadata to obtain a reverse comparison result.
[0041] In a second aspect, the present application further provides a data synchronization device, the device comprising:
[0042] An acquisition module is configured to obtain first metadata corresponding to each shard of a target bucket in an object index pool during a target migration phase; wherein the target migration phase is the first migration phase, the second migration phase, or the third migration phase, the end time of the first migration phase is the start time of the second migration phase, and the end time of the second migration phase is the start time of the third migration phase;
[0043] A synchronization module is configured to synchronize metadata in the database based on a forward comparison result and a reverse comparison result between the first metadata and second metadata stored in the database; wherein the forward comparison result is a result of comparing whether the first metadata exists in the second metadata; and the reverse comparison result is a result of comparing whether the second metadata exists in the first metadata.
[0044] In a third aspect, the present application further provides a computer device, comprising a memory and a processor, wherein the memory stores a computer program, and when the processor executes the computer program, the following steps are implemented:
[0045] Obtaining first metadata corresponding to each shard of a target bucket in an object index pool in a target migration phase; wherein the target migration phase is a first migration phase, a second migration phase, or a third migration phase, the end time of the first migration phase is the start time of the second migration phase, and the end time of the second migration phase is the start time of the third migration phase;
[0046] Metadata synchronization is performed on the database based on a forward comparison result and a reverse comparison result between the first metadata and second metadata stored in the database; wherein the forward comparison result is a result of comparing whether the first metadata exists in the second metadata; and the reverse comparison result is a result of comparing whether the second metadata exists in the first metadata.
[0047] In a fourth aspect, the present application further provides a computer-readable storage medium having a computer program stored thereon, wherein the computer program, when executed by a processor, implements the following steps:
[0048] Obtaining first metadata corresponding to each shard of a target bucket in an object index pool in a target migration phase; wherein the target migration phase is a first migration phase, a second migration phase, or a third migration phase, the end time of the first migration phase is the start time of the second migration phase, and the end time of the second migration phase is the start time of the third migration phase;
[0049] Metadata synchronization is performed on the database based on a forward comparison result and a reverse comparison result between the first metadata and second metadata stored in the database; wherein the forward comparison result is a result of comparing whether the first metadata exists in the second metadata; and the reverse comparison result is a result of comparing whether the second metadata exists in the first metadata.
[0050] In a fifth aspect, the present application further provides a computer program product, comprising a computer program, which, when executed by a processor, implements the following steps:
[0051] Obtaining first metadata corresponding to each shard of a target bucket in an object index pool in a target migration phase; wherein the target migration phase is a first migration phase, a second migration phase, or a third migration phase, the end time of the first migration phase is the start time of the second migration phase, and the end time of the second migration phase is the start time of the third migration phase;
[0052] Metadata synchronization is performed on the database based on a forward comparison result and a reverse comparison result between the first metadata and second metadata stored in the database; wherein the forward comparison result is a result of comparing whether the first metadata exists in the second metadata; and the reverse comparison result is a result of comparing whether the second metadata exists in the first metadata.
[0053] The above-mentioned data synchronization method, apparatus, computer equipment and storage medium obtain the first metadata corresponding to each shard of the target bucket in the object index pool during the target migration phase; wherein the target migration phase is the first migration phase, the second migration phase or the third migration phase, the end time of the first migration phase is the start time of the second migration phase, and the end time of the second migration phase is the start time of the third migration phase; based on the forward comparison results and reverse comparison results between the first metadata and the second metadata stored in the database, the database is synchronized with the metadata; wherein the forward comparison result is the result of comparing whether the first metadata exists in the second metadata; the reverse comparison result is the result of comparing whether the second metadata exists in the first metadata. The above-mentioned scheme compares the first metadata in the object index pool with the second metadata in the database in stages and synchronizes the first metadata to the database, that is, uses the database to store metadata without involving bucket storage, thereby solving the problem caused by too many or too few shards in the bucket; and, through the forward comparison and reverse comparison, it can be ensured that the first metadata in the object index pool is all stored in the database, and there is no dirty data in the database. BRIEF DESCRIPTION OF THE DRAWINGS
[0054] Figure 1 1 is a flow chart of a data synchronization method according to an embodiment;
[0055] Figure 2 A schematic diagram of storing metadata in an object index pool in one embodiment;
[0056] Figure 3 A schematic diagram of a process for synchronizing metadata on a database in one embodiment;
[0057] Figure 4 Schematic diagram of forward alignment in the first migration stage in one embodiment;
[0058] Figure 5 Schematic diagram of reverse alignment in the first migration stage in one embodiment;
[0059] Figure 6 1 is a schematic diagram of another process for synchronizing metadata on a database in one embodiment;
[0060] Figure 7 Schematic diagram of the forward alignment in the second migration stage in one embodiment;
[0061] Figure 8 Schematic diagram of reverse alignment in the second migration phase in one embodiment;
[0062] Figure 9 A schematic diagram of locking and migrating header data in one embodiment;
[0063] Figure 10 1 is a schematic diagram of a process for synchronizing metadata on a database in accordance with another embodiment;
[0064] Figure 11 Schematic diagram of forward alignment in the third migration stage in one embodiment;
[0065] Figure 12 Schematic diagram of reverse alignment in the third migration stage in one embodiment;
[0066] Figure 13 A schematic diagram of cleaning metadata marked for deletion in one embodiment;
[0067] Figure 14 A schematic diagram of a process for executing a data comparison task using multiple threads in one embodiment;
[0068] Figure 15 A schematic diagram of allocating slices to each thread in one embodiment;
[0069] Figure 16 A schematic diagram of allocating second metadata to each thread in one embodiment;
[0070] Figure 17 is a schematic diagram of another data synchronization process in one embodiment;
[0071] Figure 18 is a structural block diagram of a data synchronization device in one embodiment;
[0072] Figure 19 FIG. 1 is a diagram showing the internal structure of a computer device in one embodiment. DETAILED DESCRIPTION
[0073] In order to make the purpose, technical solutions and advantages of this application more clear, the following further describes this application in detail with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain this application and are not intended to limit this application.
[0074] The data synchronization method provided in the embodiment of the present application can be applied to the application scenario of migrating metadata stored in the object index pool to a distributed database. The method can be executed by a server or a terminal with a certain computing power.
[0075] The server can be implemented as a standalone server or a server cluster consisting of multiple servers. Terminals can include, but are not limited to, various personal computers, laptops, smartphones, tablets, IoT devices, and portable wearable devices. IoT devices can include smart speakers, smart TVs, smart air conditioners, and smart car devices. Portable wearable devices can include smart watches, smart bracelets, and head-mounted devices.
[0076] Against the backdrop of the rapid development of the internet, the Internet of Things, cloud computing, and big data, data is experiencing explosive growth. The data generated by various applications, such as social networking, mobile communications, online video and audio, e-commerce, sensor networks, and scientific experiments, not only has large storage capacities but also exhibits diverse data types, large size fluctuations, and rapid data flow. This often generates massive volumes of files, often numbering in the billions, even billions, or even tens of billions, creating significant challenges in metadata management. In the current storage process for object storage services, object metadata management faces the following three issues.
[0077] 1. When a new bucket is created for data storage, it is divided into a specified number of shards to store object-related metadata. This information is stored in an ObjectMap Key-Value (omap(kv)) structure. Each shard corresponds to a Reliable Autonomic Distributed ObjectStore (rados) object. A rados object consists of three parts: an object identifier, object data, and object metadata. The object identifier uniquely identifies an object; the object data corresponds to a file in the local file system, where the object data is stored; and the object metadata is stored as key-value pairs.
[0078] When uploading a file object, the consistent hashing algorithm can be used to calculate the target shard of the object metadata and store the metadata information of the file object. The consistent hashing algorithm ensures that each shard can carry the file object metadata as evenly as possible. However, as the number of objects in the bucket increases, the metadata entries stored on each shard will increase accordingly. When the alarm threshold specified by the cluster is reached, a large number of object metadata mapping (large omap) alarms will appear. At this time, a large number of concurrent requests for the bucket will increase the number of requests falling on the same shard. The consistency of Rados object data is guaranteed by a single thread, so it will directly affect the execution efficiency of operations such as uploading, downloading and deleting object data.
[0079] 2. The cluster supports resharding of existing buckets. This increases the number of shards within a bucket and simultaneously rebalances the data, reducing the number of metadata (omap) entries stored on a single shard. When concurrent requests arrive, data is balanced across more shards, reducing blocking on requests to a single shard, thereby improving concurrent performance. However, resharding involves data migration, which may cause some service interruptions.
[0080] 3. Increasing the number of shards when creating a new bucket will also introduce new problems. When the number of files in a bucket reaches the hundreds of millions, the number of omap entries stored on each shard increases accordingly. For a single shard, the object metadata key value (omap_key) managed on it is ordered, but the data between shards is not. Therefore, when a request to list 1,000 objects in batches is issued, the processing logic is to obtain data for a certain number of entries from each shard, then merge and sort all the listed data, and select the top 1,000 entries in lexicographic order as the response result. Therefore, when the number of shards is too large, the listing performance will be significantly reduced.
[0081] Based on this, the present embodiment introduces the Advanced Distributed SQL Database (TiDB) to manage object metadata, migrating the metadata in the object index pool to the TiDB database. TiDB is an open source distributed relational database that supports both online transaction processing and online analytical processing. It has the following advantages:
[0082] 1. TiDB database is a distributed storage database that can implement replica-level disaster recovery backup, with 3 replicas by default.
[0083] 2. The TiDB database supports massive data and highly concurrent online transaction processing (OLTP) scenarios, eliminating the need to worry about shard resource preemption and competition leading to decreased concurrent performance.
[0084] 3. The TiDB database's storage and computing separation architecture design allows for online expansion or reduction of computing and storage capacity on demand, and the expansion or reduction process is transparent to application operation and maintenance personnel.
[0085] 4. The key values of the TiDB database are strictly ordered globally. Based on this feature, the concept of sharding can be completely abandoned, and the performance of data enumeration will be significantly improved.
[0086] The following describes in detail the various implementation steps of the data synchronization method provided in the embodiments of the present application.
[0087] In one embodiment, Figure 1 As shown, a data synchronization method is provided, which is described by taking the method applied to a server as an example, and includes the following steps:
[0088] S101: Obtain first metadata corresponding to each shard of a target bucket in an object index pool during a target migration phase.
[0089] For example, see Figure 2 , Figure 2A schematic diagram of metadata storage in an object index pool is provided. The object index pool includes multiple buckets, each of which can be divided into multiple shards. The target bucket is any bucket in the object index pool. Metadata is stored in the omap(kv) structure. Taking shard 1 and shard 2 as examples, omap_key is the metadata key value, and type is the type of omap_key. Shard 1 stores three single-version object metadata, namely single_obj1, single_obj2, and single_obj3; the single-version object metadata types are all plain; while shard 2 stores two versions of a file, of which ver_obj is of plain type and has no practical use but exists as a placeholder. Ver_obj_1 and ver_obj_2 are also of plain type, representing the placeholders of the two versions respectively. When listing all the version information of this object, ver_obj_1 and ver_obj_2 will be listed. The ver_obj_1_1000 and ver_obj_2_1000 types are instance types, representing the instance metadata identifier corresponding to a specific version. For example, the specific instance corresponding to ver_obj_1 is ver_obj_1_1000. The last ver_obj_1_1001 type is the latest version (olh) type; it points to the latest version of the object. If no version number is specified, the latest version of the object is found based on the olh information. Statistics for each shard are stored in the shard's header attributes. For example, statistics for shard 1 are stored in the shard1_header attribute, while statistics for shard 2 are stored in the shard2_header attribute. The header records statistical information such as the number of object data stored on the shard and its capacity. The number of objects in a bucket and the capacity occupied are summarized using the collected shard header information.
[0090] The target migration phase is the first migration phase, the second migration phase, or the third migration phase. The end time of the first migration phase is the start time of the second migration phase, and the end time of the second migration phase is the start time of the third migration phase. For example, assuming that time T0 is the start time of the first migration phase and time T1 is the end time of the first migration phase, the time period between time T0 and time T1 is the first migration phase; assuming that time T1 is the start time of the second migration phase and time T2 is the end time of the second migration phase, the time period between time T1 and time T2 is the second migration phase; the start time of the third migration phase is time T2, and the time period after time T2 is the third migration phase.
[0091] Optionally, a comparison object called compare_obj can be introduced. It records the start and end comparison times, last_compare_time and current_compare_time, as the two ends of the migration phase, in microseconds (μs). A new flag, is_suspend(bool), is added to indicate whether the comparison is suspended and the duration of the suspension (in seconds), respectively. If the comparison task puts pressure on the cluster and increases the latency of normal operations, the comparison task can be suspended at any time and resumed when the cluster is no longer under pressure.
[0092] For example, compare_obj is initialized as follows:
[0093]
[0094] Exemplarily, the first metadata corresponding to each shard of the target bucket in the object index pool during the target migration phase is stored in the disk, and the first metadata corresponding to each shard of the target bucket in the object index pool during the target migration phase can be read into the memory for data comparison operations.
[0095] S102 : Synchronize metadata on the database according to a forward comparison result and a reverse comparison result between the first metadata and the second metadata stored in the database.
[0096] The forward comparison result is the result of comparing whether the first metadata exists in the second metadata; the reverse comparison result is the result of comparing whether the second metadata exists in the first metadata.
[0097] For example, the database can be synchronized with metadata based on the forward comparison results. For example, the first metadata not included in the second metadata of the database can be synchronized to the database. Then, the database can be synchronized with metadata based on the reverse comparison results. For example, the second metadata not included in the first metadata of the object index pool can be treated as dirty data and deleted from the database. The database can be a TiDB database.
[0098] The above-mentioned data synchronization method obtains the first metadata corresponding to each shard of the target bucket in the object index pool during the target migration phase; wherein the target migration phase is the first migration phase, the second migration phase or the third migration phase, the end time of the first migration phase is the start time of the second migration phase, and the end time of the second migration phase is the start time of the third migration phase; based on the forward comparison result and the reverse comparison result between the first metadata and the second metadata stored in the database, the database is synchronized with the metadata; wherein the forward comparison result is the result of comparing whether the first metadata exists in the second metadata; the reverse comparison result is the result of comparing whether the second metadata exists in the first metadata. The above-mentioned scheme compares the first metadata in the object index pool and the second metadata in the database in stages, and synchronizes the first metadata to the database, that is, uses the database to store metadata, does not involve bucket storage, and thus solves the problem caused by too many or too few shards in the bucket; and, through the forward comparison and reverse comparison, it can be ensured that the first metadata in the object index pool is stored in the database, and there is no dirty data in the database.
[0099] In some optional implementations, the embodiments of the present application are described using the target migration phase as the first migration phase. Assuming the first migration phase is the time period between time T0 and time T1, the first metadata to be compared is the metadata stored in the object index pool before the start time of the first migration phase, that is, the metadata before time T0.
[0100] Based on this, see Figure 3 , Figure 3 A schematic diagram of the process of synchronizing database metadata is provided, which specifically includes the following steps:
[0101] S301 : Compare the first metadata with metadata stored in a database before the start time of the first migration phase to obtain a forward comparison result.
[0102] For example, the first migration phase starts at compare_obj as follows:
[0103]
[0104] In this case, the second metadata is the metadata stored in the database before the start of the first migration phase. The metadata stored in the database before the start of the first migration phase can be read into memory, and the first metadata can be compared with the metadata stored in the database before the start of the first migration phase to obtain a positive comparison result.
[0105] S302: When the forward comparison result shows that the first target metadata in the first metadata does not exist in the database, the first target metadata is synchronized to the database.
[0106] Furthermore, if the forward comparison result shows that the first target metadata in the first metadata does not exist in the database, the full data of the first target metadata is copied to the database. If the forward comparison result shows that the first target metadata in the first metadata exists in the database, the update time mtime (recorded in the metadata attribute field) of the metadata in the object index pool and the database can be compared, and the metadata with the latest update time mtime is used as the standard, and then the metadata with the latest update time mtime is stored in the database.
[0107] For example, see Figure 4 , Figure 4 A schematic diagram of the forward comparison of the first migration phase is provided. Among them, the data update time of the metadata key values Obj1, Obj4 and Obj801 included in shard 1 is less than T0, and the data update time of the metadata key values Obj3, Obj6, and Obj800 included in shard 2 is less than T0, and they can all be used as the first metadata; when the forward comparison result shows that these data do not exist in the database, these data are synchronized to the database. Among them, Obj2 and Obj5 are deleted data and do not need to be synchronized to the database. The update time of Obj7 and Obj8 is greater than T0, and they are changed data. Obj7 and Obj8 do not need to be processed in the first migration phase. The update time of Obj804 and Obj802 is greater than T0, and they are written data. Since the first migration phase supports writing data into the object index pool and the database at the same time, the metadata corresponding to Obj804 and Obj802 will be written into the database; and the read data is only read in the object index pool.
[0108] S303: Compare all metadata stored in the database with the first metadata to obtain a reverse comparison result.
[0109] Furthermore, all metadata stored in the database may be extracted into the memory, and all metadata stored in the database may be compared with the first metadata to obtain a reverse comparison result.
[0110] S304: If the reverse comparison result shows that the second target metadata exists in the database and does not exist in the first metadata, the second target metadata is deleted from the database.
[0111] Furthermore, if the reverse comparison result shows that the second target metadata exists in the database and does not exist in the first metadata, the second target metadata is considered to be dirty data and can be deleted from the database.
[0112] For example, see Figure 5 , Figure 5A schematic diagram of the reverse comparison in the first migration phase is provided. If the update time of Obj34, Obj41, and Obj502 is less than T0 and they exist in the first metadata, no processing will be performed. If the update time of Obj43 is greater than T0, Obj43 will not be processed in the first migration phase. If the update time of Obj107 and Obj504 is less than T0 and they do not exist in the first metadata, they are considered dirty data and can be deleted from the database.
[0113] After traversing all secondary metadata in the database, the last_compare_time (T0) is assigned to the current_compare_time and persisted. This means that all data compared during the first migration phase was updated before T0. However, new data generated during the comparison process, with an update time later than T0, will be processed in the second migration phase. This completes the comparison process for the first migration phase.
[0114] Then compare_obj can be updated as follows:
[0115]
[0116] In an embodiment of the present application, in the first migration phase, all first metadata stored in the object index pool before the start of the first migration phase is migrated to the database, completing the migration of the existing data before the start of the first migration phase.
[0117] In some optional implementations, the embodiments of the present application are described using the second migration phase as an example. Assuming the second migration phase is the time period between time T1 and time T2, the first metadata to be compared is the metadata stored in the object index pool between the start time and the end time of the first migration phase, that is, the metadata between time T1 and time T2.
[0118] Based on this, see Figure 6 , Figure 6 A flowchart for synchronizing database metadata is provided, which includes the following steps:
[0119] S601 : Compare the first metadata with metadata stored in a database between the start time and the end time of the first migration phase to obtain a forward comparison result.
[0120] For example, the second migration phase starts at compare_obj as follows:
[0121]
[0122] At the start of the second migration phase, data can still be written to both the object index pool and the database simultaneously, while data is read only from the object index pool. Furthermore, the target bucket's mark-delete switch is enabled. This means adding a marker field, mark_del, to the target bucket's attribute data. This field, with a default value of false, is changed to true and persisted to the underlying layer. Deletion operations on objects within the target bucket will delete the actual data, but the mark_del attribute in the TiDB database's metadata will be marked as true. This means that the metadata in the TiDB database is only marked for deletion, not actually deleted, while the metadata in the object index pool is actually deleted.
[0123] In this case, the second metadata is the metadata stored in the database between the start and end of the first migration phase. The metadata stored in the object index pool between the start and end of the first migration phase can be extracted into memory, and the metadata stored in the database between the start and end of the first migration phase can be extracted into memory. The first metadata can then be compared with the metadata stored in the database between the start and end of the first migration phase to obtain a forward comparison result.
[0124] S602: If the forward comparison result is that the third target metadata in the first metadata does not exist in the database, synchronize the third target metadata to the database.
[0125] For example, if the forward comparison result indicates that the third target metadata in the first metadata does not exist in the database, the third target metadata can be synchronized to the database. If the forward comparison result indicates that the third target metadata in the first metadata exists in the database, the update time mtime of the metadata in the object index pool and the database can be compared, and the metadata with the latest update time mtime is used as the standard, and then the metadata with the latest update time mtime is stored in the database.
[0126] For example, see Figure 7 , Figure 7A schematic diagram of the forward comparison in the second migration phase is provided. Among them, Obj5 is the existing data before T0, and it was deleted after T1, so the actual deletion was completed in the object index pool, but the deletion was marked in the database and the mtime was updated to be greater than T1, and it was not actually deleted. Obj7 is an update of the existing data, so the update of the data in the database and the object index pool was completed during the comparison process, and the update time mtime was greater than T1, so Obj7 does not belong to the comparison metadata of this phase. Obj800 is the new data added to the business, and it is in dual-write mode at this time, so an object metadata record is added to both the database and the object index pool. Therefore, only Obj4 is the metadata uploaded after the start of the first migration phase (after T0), and if the write fails when writing to the database, the metadata can be migrated from the object index pool to the database. Among them, Obj1, Obj3, Obj803 and Obj804 are all existing data in the database.
[0127] S603: Compare the metadata stored in the database between the start time and the end time of the first migration phase with the first metadata to obtain a reverse comparison result.
[0128] Furthermore, the metadata stored in the database between the start time of the first migration phase and the end time of the first migration phase may be reversely compared with the first metadata to obtain a reverse comparison result.
[0129] S604: If the reverse comparison result shows that the fourth target metadata exists in the database and does not exist in the first metadata, the fourth target metadata is deleted from the database.
[0130] Furthermore, if the reverse comparison result shows that the fourth target metadata exists in the database and does not exist in the first metadata, the fourth target metadata is considered to be dirty data and can be deleted from the database.
[0131] See also Figure 8 , Figure 8 A schematic diagram of the reverse comparison of the second migration phase is provided. Among them, the metadata Obj41 is deleted after T1, the database marks Obj41 as deleted and updates the time to be greater than T1, while Obj41 in the object index pool is actually deleted. The newly added metadata Obj43 for the business is still added to the database and the object index pool at the same time; the metadata (Obj107, Obj504) that does not exist in the object index pool but exists in the database, that is, the fourth target metadata, is regarded as dirty data and is cleared from the database; the updated metadata Obj741 is still updated simultaneously in the database and the object index pool. Obj34 and Obj502 are metadata that exist in both the database and the object index pool and can be left unprocessed.
[0132] S605: After locking the metadata for non-read operations in the object index pool, migrate the header data of each shard of the target bucket in the object index pool to the database.
[0133] For example, since the metadata is not locked, a phased comparison migration method is adopted during the metadata migration process, taking into account the final data consistency; however, the target bucket must be locked when migrating the shard header, so that before the header migration is completed, all add, delete, and modify operations on the target bucket are not allowed. This is because the header statistics information read into the memory will synchronously modify the underlying header information during the add, delete, and modify operations of the object, resulting in a mismatch between the header information in the memory and the disk. Direct migration to the database without locking will cause the header statistics in the database to be inconsistent with the actual situation.
[0134] Furthermore, the header data of each shard in the target bucket can be further migrated. Before migrating the header, the metadata for non-read operations in the target bucket needs to be locked, that is, adding, deleting, and modifying objects in the target bucket are not allowed during the header migration process. After locking the metadata for non-read operations in the object index pool, the header data of each shard of the target bucket in the object index pool is migrated to the database. Since the number of shards in a bucket is usually 2048, each shard corresponds to a header statistical information, the number is very small, and the migration can be completed within 1 second, so in order to ensure the consistency of the header data, it is necessary to sacrifice short-term business.
[0135] See also Figure 9 , Figure 9 A schematic diagram of locking the migration header data is provided. During the header migration process, the dual-write, single-read mode is maintained, meaning metadata is written to the database and object index pool, and only read from the object index pool. Reading metadata from the object index pool is unaffected. However, operations to add, delete, and modify metadata in the object index pool and database will fail due to the lock on the target bucket.
[0136] Because header migration is a short process, a retry mechanism has been added to the add, delete, and modify processes. After the header migration is complete and the target bucket is unlocked, the retry succeeds, with no real impact on normal user services. While the bucket is locked and the header is migrated, user read requests are not affected.
[0137] For example, during the migration of the header data of shards 1 to shard 9, add, delete, and modify operations on the object index pool and metadata in the database are not allowed, but read operations on the data in the object index pool are allowed.
[0138] Furthermore, after migrating the header, you should first switch the target bucket to single-write, single-read. This means that data addition, deletion, modification, and query operations will be performed on the metadata in the database, not on the data in the object index pool. Only then can you unlock the target bucket and resume normal addition, deletion, modification, and query operations. From then on, addition, deletion, modification, and query operations for the target bucket will be handled by the database, and no addition, deletion, modification, and query operations will be performed on the metadata in the object index pool.
[0139] Then compare_obj can be updated as follows:
[0140]
[0141] In an embodiment of the present application, in the second migration phase, the incremental metadata generated by the object index pool in the second migration phase is migrated to the database, and the migration of the header data of each shard in the target bucket is completed.
[0142] Furthermore, in the above embodiment, based on the forward comparison results and the reverse comparison results between the first metadata and the second metadata stored in the database, during the process of synchronizing the metadata of the database, that is, in the second migration phase, it is necessary to add a deletion mark to the metadata to be deleted; wherein the metadata to be deleted is the metadata in the database corresponding to the data deletion operation, and the metadata to be deleted is the metadata shared by the database and the object index pool.
[0143] For example, in the second migration phase, the reasons for adding a deletion mark to the metadata to be deleted in the database are as follows:
[0144] At the start of the second migration phase, mark-for-delete is enabled for bucket operations on the database side. That is, after the start of the second migration phase, all delete operations on the target bucket are still double-deletes, that is, the corresponding metadata in the object index pool and the database are deleted simultaneously. However, the metadata attributes in the database are marked for deletion instead of being deleted immediately. The corresponding metadata in the object index pool is actually deleted.
[0145] This is because the data is not locked when the metadata is compared. If 1,000 pieces of data are read from the object index pool and are about to be compared with the metadata in the database, and the user issues a delete operation, if the data is not marked for deletion, there will be two scenarios:
[0146] Scenario 1: In dual-write, single-read mode, metadata is written to the database and object index pool, and metadata is read only from the object index pool. If metadata is successfully deleted from both the object index pool and the database, but 1,000 metadata entries read from the object index pool are in memory, the metadata will be migrated from memory to the database. This will cause dirty data to be present in the database that should have been deleted, and users will be aware of this dirty data. This dirty data will be cleaned up in the third migration phase, preventing invalid dirty data from being generated. However, users may still be aware of dirty data from the start of the second migration phase to the end of the third migration phase.
[0147] In scenario 2, after the target bucket switches to a single-read, single-write database mode, data addition, deletion, modification, and query operations all act on the metadata in the database. Data in the object index pool is no longer added, deleted, modified, or queried. All businesses access the database. Therefore, deletions are changed to deleting metadata in the database only, while the same metadata in the object index pool is not deleted. If marked deletion is not enabled in the database at this stage, the metadata in the database will be deleted, while the metadata in the object index pool will not be deleted. At this time, there is more metadata in the object index pool than in the database, and it is impossible to determine whether it should be migrated to the database in the third migration phase.
[0148] Therefore, if marked-for-deletion metadata is enabled in the database, after switching to a read-only, write-only database, delete operations will mark the hit metadata attributes in the database as deleted and update the mtime of the metadata. This makes the metadata of the marked-for-deletion objects invisible to users. Therefore, after the metadata comparison is completed in the third migration phase, the metadata marked for deletion in the marked-for-deletion database can be directly cleared. This solves both of the above problems.
[0149] In an embodiment of the present application, by adding a deletion mark to the metadata to be deleted in the second migration phase, it is possible to avoid the problem of deleting data in the database without deleting data in the object index pool after the single-read single-write database mode is enabled, but the comparison result is unreliable.
[0150] In some optional implementations, the embodiments of the present application are described using the target migration phase as the third migration phase. Assuming the third migration phase is the time period after time T2, the first metadata to be compared is all metadata stored in the object index pool up to the end of the second migration phase, that is, all metadata stored before time T2.
[0151] Based on this, see Figure 10 , Figure 10 A schematic diagram of another process for synchronizing metadata for a database is provided, which specifically includes the following steps:
[0152] S1001: Compare the first metadata with all metadata stored in a database to obtain a positive comparison result.
[0153] For example, the second migration phase has already completed the switch to the single-read, single-write database mode, but the third migration phase is still necessary. In this phase, the second metadata is the full metadata in the database. By comparing the full metadata in the object index pool with the full metadata in the database, the difference metadata is further processed to ensure that the metadata in the database is up to date and free of residual dirty data.
[0154] After the target bucket's mark-delete status is changed to actual deletion, the database's marked-delete metadata is cleaned up. During the second migration phase, the time T2 for switching to single-read, single-write database mode was recorded and persisted in the current_compare_time variable. During the third migration phase, the database determines whether inconsistent metadata should be deleted based on whether its mtime is after T2.
[0155] Exemplarily, all first metadata in the object index pool may be extracted into memory, and all metadata stored in the database may be extracted into memory, and the first metadata may be compared with all metadata stored in the database to obtain a positive comparison result.
[0156] S1002: If the forward comparison result is that the fifth target metadata in the first metadata does not exist in the database, the fifth target metadata is synchronized to the database.
[0157] If the forward comparison result shows that the fifth target metadata in the first metadata does not exist in the database, and the update time of the fifth target metadata is less than the start time of the third migration phase, that is, it represents metadata that failed to be written to the database in the second migration phase, then the fifth target metadata can be synchronized to the database.
[0158] For example, see Figure 11 , Figure 11 A schematic diagram of the forward comparison in the third migration phase is provided. The mtimes of Obj800 and Obj804 in the object index pool are both less than T2 and are not found in the database. Therefore, they are considered valid metadata and can be migrated to the database. Obj4 and Obj803 exist in both the database and the object storage pool and can be left untouched. Obj1, Obj3, Obj5, and Obj7 all represent operations performed on the database after T2.
[0159] S1003: Compare all metadata stored in the database with the first metadata to obtain a reverse comparison result.
[0160] Furthermore, all metadata stored in the database may be compared with the first metadata to obtain a reverse comparison result.
[0161] S1004: If the reverse comparison result shows that the sixth target metadata exists in the database and does not exist in the first metadata, the sixth target metadata is deleted.
[0162] For example, when the reverse comparison result shows that the sixth target metadata exists in the database and does not exist in the first metadata, and the update time of the sixth target metadata is less than the start time of the third migration phase, the sixth target metadata is considered to be dirty data and can be deleted.
[0163] See also Figure 12 , Figure 12 A schematic diagram of the reverse comparison in the third migration phase is provided. Obj741 is dirty data and can be directly cleared. However, if the mtime is after T2, such as Obj43, no action is taken. If both sides have the same metadata, such as Obj34 and Obj502, the data is considered consistent and no action is taken.
[0164] S1005: Delete the metadata to be deleted that is marked for deletion in the database.
[0165] Furthermore, you can remove the target bucket's deletion mark and change it to actual deletion, for example, by setting the field mark_del to false. After that, all deletion operations on the target bucket will be processed in the database and will actually delete the metadata, and the metadata marked for deletion in the database will be deleted.
[0166] See also Figure 13 , Figure 13 This diagram shows how to clean up metadata marked for deletion. Obj5, Obj133, Obj121, and Obj412 are metadata marked for deletion and are actually deleted from the database. Obj41 and Obj743 are the actual deletion operations after the target bucket's deletion mark is removed. Obj41 and Obj743 are actually deleted from the database.
[0167] Furthermore, you can initialize compare_obj to allow other buckets in the object index pool to migrate data.
[0168]
[0169] In an embodiment of the present application, in the third migration phase, the incremental metadata generated by the object index pool before the third migration phase is migrated to the database, and the dirty data and metadata marked for deletion in the database are cleared, thereby achieving the purpose of migrating the metadata in the object storage pool to the database, and there is no missing metadata or redundant dirty data in the database.
[0170] In some optional implementations, in the above embodiment, if the forward comparison result indicates that the first metadata exists within the second metadata, the metadata with the latest update time between the first and second metadata is synchronized to the database. For example, if the forward comparison result indicates that the first metadata exists within the second metadata, the update time mtime of the metadata in the object index pool and the database can be compared, and the metadata with the latest update time mtime is used as the standard, and then the metadata with the latest update time mtime is stored in the database.
[0171] In an embodiment of the present application, when the object index pool and the database both have the same metadata, the metadata with the latest update time is stored in the database, which can ensure that the metadata version stored in the database is the latest.
[0172] In some optional implementations, in order to improve data comparison efficiency, multiple concurrent threads may be used to execute the data comparison task.
[0173] Based on this, see Figure 14 , Figure 14 A flowchart for executing a data comparison task using multiple threads is provided, which specifically includes the following steps:
[0174] S1401: Obtain the number of threads used to compare the first metadata and the second metadata.
[0175] Exemplarily, the number of threads used for comparing the first metadata and the second metadata, that is, the number of threads used for forward comparison, can be first obtained.
[0176] S1402: Determine, based on the number of shards of the target bucket in the object index pool and the number of threads, a shard allocated to each thread for comparison with the second metadata.
[0177] Furthermore, the number of shards allocated to each thread for comparison with the second metadata can be determined based on the number of shards in the target bucket in the object index pool and the number of threads. For example, the number of shards in the target bucket and the number of threads can be rounded to an integer, so that the number of shards allocated to each thread is as balanced as possible.
[0178] For example, see Figure 15 , Figure 15This article provides a schematic diagram for allocating shards to each thread. Assuming there are 128 shards and 13 threads, an equal-sharding algorithm can be used to calculate a task list for each thread, storing the shard identities (IDs). For example, 10 shards can be allocated to the first 12 threads and 8 shards to the last thread, ensuring a balanced number of shards across all threads. Thread tasks are then created, and each thread concurrently processes the shard tasks in its task list.
[0179] S1403: Control each thread to perform a forward comparison operation on the first metadata and the second metadata in the corresponding slice to obtain a forward comparison result.
[0180] Furthermore, each thread is controlled to perform a forward comparison operation on the first metadata and the second metadata in the respective allocated slices to obtain a forward comparison result, thereby improving the efficiency of the forward comparison.
[0181] S1404 : Determine, based on the amount of second metadata and the number of threads, the second metadata allocated to each thread for comparison with the first metadata.
[0182] For example, the number of threads used to compare the second metadata with the first metadata is the number of threads used for the reverse comparison. Similarly, the amount of second metadata allocated to each thread for comparison with the first metadata can be determined based on the amount of second metadata and the number of threads. For example, the quotient of the amount of second metadata and the number of threads can be rounded to an integer, so that the amount of second metadata allocated to each thread is as balanced as possible.
[0183] For example, see Figure 16 , Figure 16 This article provides a schematic diagram for assigning secondary metadata to each thread. Since the database lacks the concept of sharding, it's impossible to partition the subtask list by shard. Instead, the subtasks are partitioned using the [start_key, end_key] scheme. These are assembled into multiple [start_key, end_key] combinations, each of which is assigned to the corresponding thread task. The start_key and end_key values are the index values for the partition interval, and the corresponding secondary metadata can be retrieved based on the interval index values.
[0184] The first set of start_keys is empty, indicating that the query begins at interval index 0, and the last set of end_keys is empty, indicating that the query ends at the end. Within a thread task, data is enumerated based on the given start_key and end_key. If the last enumerated key is lexicographically greater than or equal to the end_key, the thread task terminates. This ensures that each concurrent thread processes the same second metadata entry as much as possible.
[0185] Assuming that there are 10,000 pieces of second metadata and 10 threads, an equal distribution algorithm can be used to allocate 1,000 pieces of second metadata to each thread, so as to make the amount of second metadata allocated to each thread as balanced as possible.
[0186] S1405 , controlling each thread to perform a reverse comparison operation on the corresponding second metadata and the first metadata to obtain a reverse comparison result.
[0187] Furthermore, each thread can be controlled to concurrently perform a reverse comparison operation on the corresponding second metadata and the first metadata to obtain a reverse comparison result, thereby improving the efficiency of the reverse comparison.
[0188] In the embodiment of the present application, by using multiple concurrent threads to execute the forward comparison task and the reverse comparison task, the efficiency of the forward comparison and the reverse comparison is improved while ensuring that the tasks assigned to each thread are balanced.
[0189] In some optional implementations, see Figure 17 , Figure 17 Another data synchronization process diagram is provided, which includes the following steps:
[0190] S1701: Obtain first metadata corresponding to each shard of a target bucket in an object index pool during a target migration phase, and second metadata corresponding to the database during a target migration phase.
[0191] S1702: Obtain the number of threads used to compare the first metadata and the second metadata.
[0192] S1703: Determine, based on the number of shards of the target bucket in the object index pool and the number of threads, a shard allocated to each thread for comparison with the second metadata.
[0193] S1704: Control each thread to perform a forward comparison operation on the first metadata and the second metadata in the corresponding shard to obtain a forward comparison result.
[0194] S1705 : Determine, based on the amount of second metadata and the number of threads, the second metadata allocated to each thread for comparison with the first metadata.
[0195] S1706 , controlling each thread to perform a reverse comparison operation on the corresponding second metadata and the first metadata to obtain a reverse comparison result.
[0196] S1707: Synchronize metadata on the database according to a forward comparison result and a reverse comparison result between the first metadata and the second metadata stored in the database.
[0197] This embodiment of the present application no longer relies on the concept of sharded metadata storage when uploading large amounts of file data. This design avoids the bottleneck of large omap alerts and significantly reduced enumeration efficiency when the number of objects in a single bucket increases, and also improves the response efficiency of services such as lifecycle operations. Furthermore, this embodiment of the present application utilizes the current phased, bidirectional, multi-concurrent comparison and migration solution, which not only ensures metadata consistency but also significantly improves metadata comparison and migration efficiency, preventing the potential impact of a lengthy migration process on the cluster's existing services.
[0198] It should be understood that, although the various steps in the flowcharts involved in the various embodiments described above are displayed in sequence according to the instructions of the arrows, these steps are not necessarily executed in sequence in the order indicated by the arrows. Unless otherwise specified herein, there is no strict order restriction on the execution of these steps, and these steps can be executed in other orders. Moreover, at least a portion of the steps in the flowcharts involved in the various embodiments described above can include multiple steps or multiple stages, and these steps or stages are not necessarily executed and completed at the same time, but can be executed at different times, and the execution order of these steps or stages is not necessarily to be carried out in sequence, but can be executed in turn or alternately with other steps or at least a portion of steps or stages in other steps.
[0199] Based on the same inventive concept, embodiments of the present application also provide a data synchronization device for implementing the aforementioned data synchronization method. The implementation solution provided by this device is similar to the implementation solution described in the aforementioned method. Therefore, the specific limitations of one or more data synchronization device embodiments provided below can be found in the above-mentioned limitations of the data synchronization method and will not be repeated here.
[0200] In one embodiment, Figure 18 As shown, a data synchronization device is provided, comprising:
[0201] An acquisition module 10 is configured to obtain first metadata corresponding to each shard of a target bucket in an object index pool during a target migration phase; wherein the target migration phase is the first migration phase, the second migration phase, or the third migration phase, the end time of the first migration phase is the start time of the second migration phase, and the end time of the second migration phase is the start time of the third migration phase;
[0202] The synchronization module 20 is used to synchronize metadata in the database based on the forward comparison results and reverse comparison results between the first metadata and the second metadata stored in the database; wherein the forward comparison result is the result of comparing whether the first metadata exists in the second metadata; the reverse comparison result is the result of comparing whether the second metadata exists in the first metadata.
[0203] The above-mentioned data synchronization device obtains the first metadata corresponding to each shard of the target bucket in the object index pool during the target migration phase; wherein the target migration phase is the first migration phase, the second migration phase or the third migration phase, the end time of the first migration phase is the start time of the second migration phase, and the end time of the second migration phase is the start time of the third migration phase; based on the forward comparison result and the reverse comparison result between the first metadata and the second metadata stored in the database, the database is synchronized with the metadata; wherein the forward comparison result is the result of comparing whether the first metadata exists in the second metadata; the reverse comparison result is the result of comparing whether the second metadata exists in the first metadata. The above-mentioned scheme compares the first metadata in the object index pool and the second metadata in the database in stages, and synchronizes the first metadata to the database, that is, uses the database to store metadata, does not involve bucket storage, and thus solves the problem caused by too many or too few shards in the bucket; and, through the forward comparison and reverse comparison, it can be ensured that the first metadata in the object index pool is stored in the database, and there is no dirty data in the database.
[0204] In one embodiment, when the target migration phase is the first migration phase, the first metadata is metadata stored in the object index pool before the start time of the first migration phase; the synchronization module 20 is specifically configured to:
[0205] The first metadata is compared with metadata stored in the database before the start of the first migration phase to obtain a forward comparison result; if the forward comparison result is that the first target metadata in the first metadata does not exist in the database, the first target metadata is synchronized to the database; all metadata stored in the database are compared with the first metadata to obtain a reverse comparison result; if the reverse comparison result is that the second target metadata exists in the database and does not exist in the first metadata, the second target metadata is deleted from the database.
[0206] In one embodiment, when the target migration phase is the second migration phase, the first metadata is metadata stored in the object index pool between the start time of the first migration phase and the end time of the first migration phase; the synchronization module 20 is specifically configured to:
[0207] Comparing the first metadata with metadata stored in the database between the start time of the first migration phase and the end time of the first migration phase to obtain a forward comparison result;
[0208] If the forward comparison result shows that the third target metadata in the first metadata does not exist in the database, the third target metadata is synchronized to the database; the metadata stored in the database between the start time and the end time of the first migration phase is compared with the first metadata to obtain a reverse comparison result; if the reverse comparison result shows that the fourth target metadata exists in the database and does not exist in the first metadata, the fourth target metadata is deleted from the database; after locking the metadata for non-read operations in the target bucket, the header data of each shard of the target bucket is migrated to the database.
[0209] In one embodiment, the device further includes an adding module for:
[0210] During metadata synchronization of the database based on forward and reverse comparison results between the first metadata and the second metadata stored in the database, in the second migration phase, a deletion marker is added to the metadata to be deleted; the metadata to be deleted is metadata in the database corresponding to the data deletion operation, and the metadata to be deleted is metadata shared by the database and the object index pool.
[0211] In one embodiment, when the target migration phase is the third migration phase, the first metadata is all metadata stored in the object index pool up to the end of the second migration phase; the synchronization module 20 is specifically configured to:
[0212] The first metadata is compared with all metadata stored in the database to obtain a forward comparison result; if the forward comparison result shows that the fifth target metadata in the first metadata does not exist in the database, the fifth target metadata is synchronized to the database; wherein the update time of the fifth target metadata is less than the start time of the third migration phase; all metadata stored in the database are compared with the first metadata to obtain a reverse comparison result; if the reverse comparison result shows that the sixth target metadata exists in the database and does not exist in the first metadata, the sixth target metadata is deleted; wherein the update time of the sixth target metadata is less than the start time of the third migration phase; and the metadata to be deleted marked in the database is deleted.
[0213] In one embodiment, the synchronization module 20 is further configured to:
[0214] When the forward comparison result shows that the first metadata exists in the second metadata, the metadata with the latest update time between the first metadata and the second metadata is synchronized to the database.
[0215] In one embodiment, the device further includes a comparison module for:
[0216] Obtain the number of threads for comparing the first metadata with the second metadata; determine the shards allocated to each thread for comparison with the second metadata based on the number of shards of the target bucket in the object index pool and the number of threads; control each thread to perform a forward comparison operation on the first metadata and the second metadata in the corresponding shard to obtain a forward comparison result; determine the second metadata allocated to each thread for comparison with the first metadata based on the number of second metadata and the number of threads; control each thread to perform a reverse comparison operation on the corresponding second metadata and the first metadata to obtain a reverse comparison result.
[0217] Each module in the above-mentioned data synchronization device can be implemented in whole or in part through software, hardware, or a combination thereof. Each of the above-mentioned modules can be embedded in or independent of the processor of the computer device in hardware form, or can be stored in the memory of the computer device in software form, so that the processor can call and execute the corresponding operations of each of the above modules.
[0218] In one embodiment, a computer device is provided. The computer device may be a server, and its internal structure diagram may be as follows: Figure 19 As shown. The computer device includes a processor, a memory, and a network interface connected via a system bus. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and computer program in the non-volatile storage medium. The database of the computer device is used to store object data and object metadata. The network interface of the computer device is used to communicate with an external terminal via a network connection. When the computer program is executed by the processor, a data synchronization method is implemented.
[0219] Those skilled in the art will understand that Figure 19 The structure shown in the figure is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than shown in the figure, or combine certain components, or have a different component arrangement.
[0220] In one embodiment, a computer device is provided, including a memory and a processor. The memory stores a computer program, and the processor implements the steps of the data synchronization method described in any of the above embodiments when executing the computer program.
[0221] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the steps of the data synchronization method described in any of the above embodiments are implemented.
[0222] In one embodiment, a computer program product is provided, comprising a computer program, which implements the steps of the data synchronization method described in any of the above embodiments when executed by a processor.
[0223] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties.
[0224] Those skilled in the art will appreciate that all or part of the processes in the above-mentioned embodiment methods can be implemented by instructing the relevant hardware through a computer program, and the computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, database or other media used in the embodiments provided in this application may include at least one of non-volatile and volatile memory. Non-volatile memory may include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory may include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM). The database involved in the various embodiments provided herein may include at least one of a relational database and a non-relational database. Non-relational databases may include, but are not limited to, distributed databases based on blockchains. The processor involved in the various embodiments provided herein may be, but are not limited to, a general-purpose processor, a central processing unit, a graphics processing unit, a digital signal processor, a programmable logic unit, a data processing logic unit based on quantum computing, and the like.
[0225] The technical features of the above embodiments can be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0226] The above-described embodiments merely represent several implementation methods of the present application. While the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the present application. It should be noted that a person of ordinary skill in the art may make various modifications and improvements without departing from the spirit of the present application, and these modifications and improvements fall within the scope of protection of the present application. Therefore, the scope of protection of the present application shall be determined by the appended claims.
Claims
1. A data synchronization method, characterized in that: The method comprises: Obtaining first metadata corresponding to each shard of a target bucket in an object index pool in a target migration phase; wherein the target migration phase is a first migration phase, a second migration phase, or a third migration phase, the end time of the first migration phase is the start time of the second migration phase, and the end time of the second migration phase is the start time of the third migration phase; Performing metadata synchronization on the database based on a forward comparison result and a reverse comparison result between the first metadata and second metadata stored in the database; wherein the forward comparison result is a result of comparing whether the first metadata exists in the second metadata; and the reverse comparison result is a result of comparing whether the second metadata exists in the first metadata; When the target migration phase is the third migration phase, the first metadata is all metadata stored in the object index pool up to the end of the second migration phase; the second metadata is all metadata stored in the database; The step of synchronizing metadata on the database according to a forward comparison result and a reverse comparison result between the first metadata and second metadata stored in the database includes: Comparing the first metadata with all metadata stored in the database to obtain a positive comparison result; If the forward comparison result shows that the fifth target metadata in the first metadata does not exist in the database, synchronizing the fifth target metadata to the database; wherein the update time of the fifth target metadata is less than the start time of the third migration phase; Comparing all metadata stored in the database with the first metadata to obtain a reverse comparison result; If the reverse comparison result shows that the sixth target metadata exists in the database and does not exist in the first metadata, deleting the sixth target metadata; wherein the update time of the sixth target metadata is less than the start time of the third migration phase; Delete the metadata to be deleted that is marked for deletion in the database.
2. The method according to claim 1, characterized in that When the target migration phase is the second migration phase, the first metadata is metadata stored in the object index pool between the start time and the end time of the first migration phase; The second metadata is metadata stored in the database between the start time of the first migration phase and the end time of the first migration phase; The step of synchronizing metadata on the database according to a forward comparison result and a reverse comparison result between the first metadata and second metadata stored in the database includes: Comparing the first metadata with metadata stored in the database between the start time and the end time of the first migration phase to obtain a forward comparison result; If the forward comparison result is that the third target metadata in the first metadata does not exist in the database, synchronizing the third target metadata to the database; Comparing the metadata stored in the database between the start time and the end time of the first migration phase with the first metadata to obtain a reverse comparison result; If the reverse comparison result shows that the fourth target metadata exists in the database and the fourth target metadata does not exist in the first metadata, deleting the fourth target metadata from the database; After the metadata for non-read operations in the target bucket is locked, the header data of each shard of the target bucket is migrated to the database.
3. The method according to claim 2, characterized in that In the process of synchronizing metadata on the database based on a forward comparison result and a reverse comparison result between the first metadata and second metadata stored in the database, the method further includes: In the second migration phase, a deletion mark is added to the metadata to be deleted; The metadata to be deleted is metadata in the database corresponding to the data deletion operation, and the metadata to be deleted is metadata shared by the database and the object index pool.
4. The method according to claim 1, wherein The method further comprises: If the forward comparison result shows that the first metadata exists in the second metadata, the metadata with the latest update time between the first metadata and the second metadata is synchronized to the database.
5. The method according to claim 1, wherein The method further comprises: Obtaining the number of threads used for comparing the first metadata with the second metadata; Determining, according to the number of shards of the target bucket in the object index pool and the number of threads, a shard allocated to each thread for comparison with the second metadata; Controlling each thread to perform a forward comparison operation on the first metadata in the corresponding shard and the second metadata to obtain a forward comparison result; determining, according to the amount of the second metadata and the number of threads, second metadata allocated to each thread for comparison with the first metadata; Each thread is controlled to perform a reverse comparison operation on the corresponding second metadata and the first metadata to obtain a reverse comparison result.
6. The method according to claim 1, characterized in that When the target migration phase is the first migration phase, the first metadata is metadata stored in the object index pool before the start of the first migration phase; the second metadata is metadata stored in the database before the start of the first migration phase.
7. A data synchronization device, characterized in that: The device comprises: An acquisition module is configured to obtain first metadata corresponding to each shard of a target bucket in an object index pool during a target migration phase; wherein the target migration phase is the first migration phase, the second migration phase, or the third migration phase, the end time of the first migration phase is the start time of the second migration phase, and the end time of the second migration phase is the start time of the third migration phase; a synchronization module, configured to synchronize metadata with the database based on a forward comparison result and a reverse comparison result between the first metadata and second metadata stored in the database; wherein the forward comparison result is a result of comparing whether the first metadata exists in the second metadata; and the reverse comparison result is a result of comparing whether the second metadata exists in the first metadata; When the target migration phase is the third migration phase, the first metadata is all metadata stored in the object index pool up to the end of the second migration phase; the second metadata is all metadata stored in the database; and the synchronization module is specifically configured to: The first metadata is compared with all metadata stored in the database to obtain a forward comparison result; if the forward comparison result shows that the fifth target metadata in the first metadata does not exist in the database, the fifth target metadata is synchronized to the database; wherein the update time of the fifth target metadata is earlier than the start time of the third migration phase; all metadata stored in the database are compared with the first metadata to obtain a reverse comparison result; if the reverse comparison result shows that the sixth target metadata exists in the database and does not exist in the first metadata, the sixth target metadata is deleted; wherein the update time of the sixth target metadata is earlier than the start time of the third migration phase; and the metadata to be deleted marked in the database is deleted.
8. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 6 are implemented.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 6 are implemented.
10. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 6 are implemented.
Citation Information
Patent Citations
Data migration method, device and equipment, computer storage medium and program product
CN118210782A
Data migration method and device, equipment, storage medium and program product
CN118467502A
Object Format and Upload Process for Archiving Data in Cloud / Object Storage
US20190220198A1
Data lifetime-aware migration
US20190384525A1