Data synchronization method and migration device

By dividing the data to be migrated into non-idempotent and idempotent, and adopting the corresponding fault recovery method, the problems of data consistency and fault recovery efficiency during the data migration process are solved, and an efficient and reliable data migration process is achieved.

CN120067073APending Publication Date: 2025-05-30HUAWEI CLOUD COMPUTING TECHNOLOGIES CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410284686.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-11-30
Filing Date
2024-03-12
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

During the data migration process, how to ensure the comprehensive performance of failure recovery, ensure data consistency between the source database and the target database, and improve the failure recovery efficiency.

Method used

By dividing the data to be migrated in the source database into non-idempotent data and idempotent data, and using different fault recovery methods respectively, we ensure the consistency and efficient recovery of the data during the migration process.

Benefits of technology

It realizes data consistency during data migration, while improving the efficiency of failure recovery, reducing the amount of data that is deleted and resynchronized, and avoiding repeated writing of non-idempotent data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120067073A_ABST
    Figure CN120067073A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a data synchronization method and a data migration device, and relates to the technical field of databases. In the embodiment of the invention, based on the idempotent / non-idempotent characteristics of the data, the to-be-migrated data in the source database is divided, and different fault recovery modes are set respectively, so that the consistency of the to-be-migrated data in the source database and the data migrated to the target database can be ensured; and meanwhile, relatively high fault recovery efficiency can be ensured.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This disclosure claims the priority of a Chinese patent application with the application number 202311653823.X and the invention title "Method, System, and Device for Data Storage" filed on November 30, 2023, the entire content of which is incorporated herein by reference. Technical Field

[0002] This disclosure relates to the field of database technologies, and particularly to a data synchronization method and a migration device. Background Art

[0003] With the development of computer technologies, cloud databases are becoming increasingly popular due to their advantages such as good elastic scaling ability and high availability. Based on the business development needs of users, it is often necessary to use migration tools to migrate data from a source database to a target database. When an exception occurs during the data migration process and fault recovery is required, how to ensure the comprehensive performance of fault recovery is also an important consideration for major cloud providers to improve the competitiveness of database products. For example, after fault recovery, it is necessary to ensure the data consistency between the source database and the target database before and after migration as much as possible. At the same time, it is hoped that the efficiency of fault recovery is as high as possible. Summary of the Invention

[0004] This disclosure provides a data synchronization method and a migration device, which can not only ensure the consistency between the data to be migrated in the source database and the data migrated to the target database, but also ensure a high fault recovery efficiency.

[0005] In a first aspect, a data synchronization method is provided. The method includes: obtaining the data to be migrated in the source database, and dividing the data to be migrated into two groups, namely the first data to be migrated and the second data to be migrated, where the first data to be migrated has non-idempotent characteristics and the second data to be migrated has idempotent characteristics. Then, the first data to be migrated and the second data to be migrated are respectively synchronized to the target database. When the synchronization of the first data to be migrated fails, a fault recovery method corresponding to the first data to be migrated is selected for fault recovery; when the synchronization of the second data to be migrated fails, a fault recovery method corresponding to the second data to be migrated is selected for fault recovery. Among them, the fault recovery methods for the first data to be migrated and the second data to be migrated are different.

[0006] First, the data is classified based on idempotent and non-idempotent characteristics, and then different fault recovery methods are adopted for idempotent data and non-idempotent data respectively. In this way, not only can the consistency between the data to be migrated in the source database and the data migrated to the target database be ensured, but also a high fault recovery efficiency can be ensured.

[0007] In a possible implementation, when the synchronization of the first data to be migrated fails, the target database is notified to delete the first data to be migrated that has been synchronized to the target database, and then the first data to be migrated is synchronized to the target database again. This processing method can be called the first fault recovery method.

[0008] In this way, when performing fault recovery on non-idempotent data, it is possible to prevent data chaos caused by the repeated execution of the same write command multiple times. Thus, the consistency between the data synchronized to the target database and the data to be migrated in the source database can be ensured.

[0009] In a possible implementation, when the synchronization of the target data in the second data to be migrated fails, the target data is synchronized to the target database again. This processing method can be called the second fault recovery method.

[0010] In this way, when performing fault recovery on idempotent data, only the target data that fails during the data synchronization process needs to be re-migrated, which can ensure a relatively high fault recovery efficiency to a certain extent.

[0011] Moreover, adopting the first fault recovery method for non-idempotent data and the second fault recovery method for idempotent data as described above can reduce the amount of data to be deleted and re-synchronized compared to adopting the first fault recovery method without distinguishing all data, improving the fault recovery efficiency. Compared to adopting the second fault recovery method without distinguishing all data, there will be no repeated writing of non-idempotent data, ensuring the consistency before and after data migration.

[0012] In a possible implementation, the write command of the first data to be migrated or the second data to be migrated is sent to the target database for execution to write the first data to be migrated or the second data to be migrated into the target database.

[0013] In a possible implementation, the second data to be migrated includes multiple types of data. In this case, each type of the second data to be migrated is grouped, and each type of the second data to be migrated is synchronized to the target database separately; when the synchronization of the target type data in the second data to be migrated fails, the target type of the second data to be migrated is synchronized to the target database again.

[0014] In this way, parallel synchronization processing can be performed on multiple types of data, thereby improving the data synchronization efficiency.

[0015] In a possible implementation, the first data to be migrated or the second data to be migrated is synchronized to the target database in multiple batches.

[0016] In this way, compared to synchronizing one by one, the processing volume related to the transmission process of the write instructions can be reduced, improving the efficiency of the synchronization process.

[0017] In a possible implementation, synchronization site identification information is recorded, and the synchronization site identification information indicates the latest batch of the second data to be migrated that has been successfully synchronized to the target database. When the synchronization of the second data to be migrated in the target batch fails, the second data to be migrated in the target batch is resynchronized to the target database according to the synchronization site identification information.

[0018] In this way, when performing fault recovery, the data with synchronization failures can be directly determined based on the synchronization site identification information, improving the efficiency of data synchronization.

[0019] In a possible implementation, the first data to be migrated includes list-type data.

[0020] In a possible implementation, the second data to be migrated includes at least one of string-type data, hash-type data, set-type data, and sorted set-type data (zset).

[0021] In a second aspect, a data migration system is provided. The data migration system includes a source database, a target database, and a migration tool. The migration tool is used to execute the methods provided in the first aspect and its possible implementations above.

[0022] In a third aspect, a migration device is provided. The device includes at least one module, and the at least one module is used to implement the methods provided in the first aspect and its possible implementations above.

[0023] In a fourth aspect, a computing device cluster is provided, including at least one computing device. Each computing device includes a processor and a memory; the processor of at least one computing device is used to execute instructions stored in the memory of at least one computing device, so that the computing device cluster executes the methods provided in the first aspect and its possible implementations above.

[0024] In a fifth aspect, a computer device is provided. The computer device includes a memory and a processor. The memory is used to store computer instructions. The processor executes the computer instructions stored in the memory so that the computer device executes the methods provided in the first aspect and its possible implementations above.

[0025] In a sixth aspect, a computer-readable storage medium is provided. The computer-readable storage medium includes computer program instructions. When the computer program instructions are executed by the computing device cluster, the computing device cluster executes the methods provided in the first aspect and its possible implementations above.

[0026] In a seventh aspect, there is provided a computer program product including instructions, which, when run by a cluster of computing devices, cause the cluster of computing devices to execute the method provided in the first aspect and its possible implementations as described above. Description of the Drawings

[0027] Figure 1 FIG. is a schematic structural diagram of a data migration system provided by an embodiment of the present disclosure;

[0028] Figure 2 FIG. is a flowchart of a data synchronization method provided by an embodiment of the present disclosure;

[0029] Figure 3 FIG. is a flowchart of a data synchronization method provided by an embodiment of the present disclosure;

[0030] Figure 4 FIG. is a schematic diagram of data stored in a source database provided by an embodiment of the present disclosure;

[0031] Figure 5 FIG. is a schematic diagram of data partitioning provided by an embodiment of the present disclosure;

[0032] Figure 6 FIG. is a flowchart of a data synchronization method provided by an embodiment of the present disclosure;

[0033] Figure 7 FIG. is a flowchart of a data synchronization method provided by an embodiment of the present disclosure;

[0034] Figure 8 FIG. is a schematic diagram of synchronization site identification information provided by an embodiment of the present disclosure;

[0035] Figure 9 FIG. is a flowchart of a data synchronization method provided by an embodiment of the present disclosure;

[0036] Figure 10 FIG. is a flowchart of a data synchronization method provided by an embodiment of the present disclosure;

[0037] Figure 11 FIG. is a flowchart of a data synchronization method provided by an embodiment of the present disclosure;

[0038] Figure 12 FIG. is a schematic structural diagram of a migration device provided by an embodiment of the present disclosure;

[0039] Figure 13 FIG. is a schematic diagram of a computing device provided by an embodiment of the present disclosure;

[0040] Figure 14 FIG. is a schematic diagram of a cluster of computing devices provided by an embodiment of the present disclosure;

[0041] Figure 15It is a schematic diagram of a computing device cluster provided by an embodiment of the present disclosure. Detailed implementation manners

[0042] To make the objectives, technical solutions, and advantages of the present disclosure clearer, the following will further describe the embodiments of the present disclosure in detail with reference to the accompanying drawings.

[0043] The following explains the concepts related to the present disclosure:

[0044] Redis database

[0045] The Redis database is a widely used non-relational database. The data in the Redis database can be stored distributively. At the same time, the Redis database can maintain multiple data sets (also called sub-databases) simultaneously, and each data set can store different data based on actual business requirements.

[0046] The data types corresponding to the data in the Redis database mainly include five types, namely character type (which can be called string type), dictionary type (which can be called hash type), set type (which can be called set type), sorted set type (which can be called zset type), and list type (which can be called list type). The following briefly introduces each data type:

[0047] 1. string-type data

[0048] The form of string-type data is "key - value". For example, string-type data is used to store a username, and the corresponding form is: "user1: Jack", where "user1" is the key and "Jack" is the value. Moreover, for string-type data, the order of arrangement of each "key - value" data in the database has no special meaning. Whether any number of data are written first or later, the final writing result is the same.

[0049] The write command corresponding to string-type data is "SET key value". This write command consists of a fixed field and two parameter fields. Among them, "SET" is the fixed field and the command name, and "key" and "value" are the two parameter fields. "key" is the key parameter field, and "value" is the value parameter field. It can be easily obtained that the write command for the data in the above example is "SET user1 Jack". In addition, the write command corresponding to string-type data can also be "MSET key1 value1 key2 value2...", which means writing multiple values corresponding to multiple keys. The embodiments of the present disclosure will not elaborate on this in detail.

[0050] The mechanism for executing the write command corresponding to string-type data is to check whether the key of the write command already exists. If it exists, the value of the write command will overwrite the stored value corresponding to the key. If it does not exist, the key and value of the write command will be stored in the database correspondingly. It can be obtained that for a write command corresponding to string-type data, compared with executing it once and executing it multiple times repeatedly, the final write result of the data is the same. For example, for the write command "SET user1 Jack", no matter how many times it is executed, the final write result of the data is that the username of user1 is recorded as JACK.

[0051] 2. hash-type data

[0052] The form of hash-type data is "key-key value pair set" (for the sake of distinction, the first key is called the primary key, and the second key is called the secondary key). For example, hash-type data is used to store the basic information of employees, and the corresponding form is: "person1: {name: Jack, age: 25, sex: male}", where "person1" is the primary key, "name", "age", and "sex" are secondary keys, and "Jack", "25", and "male" are values. Moreover, for hash-type data, the arrangement order of each "key-key value pair set" and each "secondary key-value" data in the key-value pair set in the database has no special meaning. No matter which data is written first and which is written later among any number of data, the final write result is the same.

[0053] The write command corresponding to hash-type data is "HSET key field value". This write command consists of a fixed field and three parameter fields. Among them, "HSET" is the fixed field and the command name, and "key", "field", and "value" are the three parameter fields. "key" is the primary key parameter field, "field" is the secondary key parameter field, and "value" is the value parameter field. It is easy to conclude that the data in the above example can be written through three write commands, namely "HSET person1 name Jack", "HSET person1 age 25", and "HSET person1 sex male". In addition, the write command corresponding to hash-type data can also be "HMSET key field value field value...", which means writing multiple values of multiple secondary keys corresponding to a primary key. This embodiment of the present disclosure will not elaborate on it in detail.

[0054] The mechanism for executing the write command corresponding to hash-type data is similar to that of string-type. First, check whether the primary key of the write command already exists. The following explains the cases where the primary key exists and does not exist respectively: In the first case, when the primary key exists, further check whether the secondary key of the write command already exists in the key-value pair set corresponding to the primary key. If the secondary key exists, overwrite the already stored value corresponding to the secondary key with the value of the write command. If the secondary key does not exist, store the secondary key and value of the write command correspondingly in the key-value pair set corresponding to the primary key. In the second case, when the primary key does not exist, store the primary key, secondary key, and value of the write command correspondingly in the database. It can be obtained that for a write command corresponding to hash-type data, the final write result of the data is the same whether it is executed once or repeatedly. For example, for the write command "HSET person1 name Jack", no matter how many times it is executed, the final write result of the data is that the name of person1 is recorded as JACK.

[0055] 3. set-type data

[0056] The form of set-type data is "key - set". For example, set-type data is used to store the clubs established by the school, and the corresponding form is: "club: {dancing, music, painting, football}", where "club" is the key, "{dancing, music, painting, football}" is the set, and "dancing", "music", "painting", and "football" are the elements in the set. Moreover, for set-type data, the arrangement order of each "key - set" and each "element" data in the set in the database has no special meaning. No matter which data is written first and which is written later, the final write result is the same.

[0057] The write command corresponding to set-type data is "SADD key member". This write command consists of a fixed field and two parameter fields. Among them, "SADD" is the fixed field and the command name, and "key" and "member" are the two parameter fields. "key" is the key parameter field, and "member" is the element parameter field, representing an element in the set. It is easy to conclude that the data in the above example can be written through four write commands, namely "SADD club dancing", "SADD club music", "SADD club painting", and "SADD club football". In addition, the write command corresponding to set-type data can also be "SADD key member1 member2 member3...", indicating writing multiple elements into the set corresponding to this key. This will not be elaborated in detail in the embodiments of the present disclosure.

[0058] The mechanism for executing the write command corresponding to set-type data is to first check whether the key of the write command already exists. The following explains the cases where the key exists and does not exist respectively: The first case is that the key exists. Further check whether the element in the write command already exists in the set corresponding to this key. If the element exists, use this element to overwrite the already stored element, and the same element will not be written repeatedly. If the element does not exist, store this element in the set corresponding to the key. The second case is that the key does not exist. Store the key and element corresponding to this write command in the database. It can be obtained that for a write command corresponding to set-type data, compared with executing it once and repeatedly executing it multiple times, the final write result of the data is the same. For example, for the write command "SADD club dancing", no matter how many times it is executed, the final write result of the data is to record the element dancing in club.

[0059] 4. zset-type data

[0060] The form of zset-type data is similar to that of set-type data, also in the form of "key - set". The difference is that the elements in the set corresponding to the key in zset-type data can be automatically sorted based on the sorting parameter. Therefore, for zset-type data, it doesn't matter which data is written first and which is written later. The final write result is the same.

[0061] The write command corresponding to zset type data is "ZADD key score member". This write command consists of a fixed field and three parameter fields. Among them, "ZADD" is the fixed field and the command name, and "key", "score", and "member" are the three parameter fields. Among them, "key" and "member" are the same as the parameter fields of the write command corresponding to set type data, which are the key parameter field and the element parameter field respectively. "Score" is the sorting parameter field, usually a numerical value, used to sort the element in the set. For example, there are four write commands, namely "ZADD club 3 dancing", "ZADD club 1 music", "ZADD club 2 painting", and "ZADD club 4 football". After executing these four write commands in sequence, we can get "club: {music, painting, dancing, football}". In addition, the write command corresponding to zset type data can also be "ZADD key score1 member1 score2 member2...", which means writing multiple elements into the set corresponding to this key. At the same time, these multiple elements can be automatically sorted. This embodiment of the present disclosure will not elaborate on this in detail.

[0062] The mechanism for executing the write command corresponding to zset type data is similar to that of set type data. The difference is that after writing each element, it can be sorted according to the sorting parameter corresponding to the element. It can be obtained that for a write command corresponding to zset type data, the final write result of the data is the same whether it is executed once or repeatedly. For example, for the write command "ZADD club 3 dancing", no matter how many times it is executed, the final write result of the data is to record the element dancing in club and sort the element dancing based on the sorting parameter "3".

[0063] 5. list type data

[0064] The form of list type data is "key - list". Here is an example of list type data. For example, list type data is used to store scores, and the corresponding form is: "score: [66, 52, 77, 88]", where "score" is the key, "[66, 52, 77, 88]" is the list, and "66", "52", "77", and "88" are the ordered elements in the list. For list type data, the arrangement order of each "element" data in the list in the database has meaning, and there are differences in the final write results for writing one or more data first or later.

[0065] There are two forms of write commands corresponding to list-type data, namely "LPUSH key member" and "RPUSH key member". The first write command means inserting an element from the left side of the list, and the second write command means inserting an element from the right side of the list. Both of these write commands consist of a fixed field and two parameter fields. Among them, "LPUSH" and "RPUSH" are fixed fields and are command names, "key" is the key parameter field, and "member" is the element parameter field, representing an element in the list. For the data "score: [66, 52, 77, 88]", it can be written by executing four write commands successively, namely "LPUSH score 88", "LPUSH score 77", "LPUSH score 52", and "LPUSH score 66". In addition, the write commands corresponding to list-type data can also be "LPUSH key member1 member2 member3..." and "RPUSH key member1 member2 member3...", which means inserting multiple elements from the left or right side of the list. This will not be elaborated in detail in the embodiments of the present disclosure.

[0066] The mechanism for executing the write command corresponding to list-type data is to check whether the key corresponding to the write command already exists. If it exists, insert the element in the write command from the left or right side according to the command name in the list corresponding to the key. If it does not exist, store the key and element of the write command in the database correspondingly. It can be obtained that for a write command corresponding to list-type data, the final write result of the data is different when executed once and when executed repeatedly multiple times. For example, after executing the write command "LPUSH score 88" once, the write result of the data is "score:

[88] ". After repeating the execution of the write command "LPUSH score 88" once again, the write result of the data will become "score: [88, 88]".

[0067] Idempotent property / Non-idempotent property

[0068] Both the idempotent property and the non-idempotent property are properties of data types. For a certain data type to have the idempotent property means that for a write command corresponding to the data of that data type, the final write result of the data is the same when executed once and when executed repeatedly multiple times. For a certain data type to have the non-idempotent property means that for a write command corresponding to that data type, the final write result of the data is different when executed once and when executed repeatedly multiple times.

[0069] The following is an example. For instance, when the database is a Redis database, the data types include five types: string, hash, set, zset, and list. From the above introduction of the five data types and their corresponding write commands, it can be obtained that the string, hash, set, and zset data types have the idempotency property, while the list data type has the non - idempotency property.

[0070] Ordered write property / Unordered write property

[0071] The ordered write property and the unordered write property are properties of data types. If a data type has the ordered write property, it means that the write result of the data of this data type is related to the write order. If a data type has the unordered write property, it means that the write result of the data of this data type is not related to the write order. The following is an example. For instance, when the database is a Redis database, the data types include five types: string, hash, set, zset, and list. From the above introduction of the five data types and their corresponding write commands, it can be obtained that the string, hash, set, and zset data types have the unordered write property, while the list data type has the ordered write property.

[0072] Redis database backup (RDB) file (which can be simply referred to as RDB file)

[0073] The RDB file is a binary file, which can be used to record the data and data attributes in the Redis database. The data attributes include storage location, the identifier of the data set to which it belongs, data type, etc. There can be multiple data sets maintained in the Redis database, and each data set has a pre - set identifier.

[0074] The embodiments of the present disclosure provide a data synchronization method, which is applied to a data migration system. The data migration system may include a source database 110, a target database 120, and a migration tool 130. The structure of the data migration system is as Figure 1 shown. The following is an introduction to each component in the data migration system respectively:

[0075] The source database 110 is generally user - facing and can be a data management system that processes users' read requests, write requests, or other computing operations (such as aggregation calculations, etc.). For example, it can be a Redis database. The source database 110 (abbreviated as the source library) can be deployed in an independent computer device, or a virtual machine, or a computer cluster (i.e., a distributed database system), etc. During the data migration process, the data in the source database 110 is the data that needs to be migrated.

[0076] Similar to the source database 110, the target database 120 is generally also user-oriented and can be a data management system that processes users' read requests, write requests, or other computing operations (such as aggregation calculations, etc.) for data. For example, it can be a Redis database. The target database 120 (abbreviated as the target library) can be deployed in an independent computer device, or a virtual machine, or a computer cluster (i.e., a distributed database system), etc. During the data migration process, the target database 120 is mainly used to receive and store the data in the source database 110. In this way, the business of the source database 110 can be migrated to the target database 120. Or, when an exception occurs in the source database 110, the target database 120 can be used as a backup business system to ensure that the business can continue.

[0077] The migration tool 130 is used to migrate the data in the source database to the target database and can be deployed in an independent computer device, or a virtual machine, or a computer cluster. The migration tool 130 can obtain the data in the source database 110 and then synchronize this data to the target database 120.

[0078] The migration tool 130 can be deployed separately from the source database 110 and the target database 120 on different devices, or it can be deployed on the same device as the target database.

[0079] When the data volume of the database for the user is relatively small, the user often chooses to directly build and maintain the database by themselves (which can also be called the source database). As the data increases, the difficulty and cost of managing and maintaining the database will also increase accordingly. At this time, the user hopes to move the database to the cloud and rent the database built by the cloud provider in the cloud (which can also be called the target database). The target database is managed and maintained by the cloud provider, and the user can save a lot of manpower and material resources. At this time, a migration tool is needed to migrate the data in the source database to the target database.

[0080] In the related technology, the steps for the migration tool to migrate the data in the source database to the target database are as follows: First, the migration tool obtains the data to be migrated in the source database, and then synchronizes all the data to be migrated to the target database in batches. Specifically, the migration tool obtains the RDB file of the data to be migrated in the source database, and the data in the RDB file is a mixture of various data types. Then the migration tool generates the write commands corresponding to the data to be migrated based on the RDB file, and finally sends the write commands to the target database in batches for execution to write the data into the target database. During the data migration process, it is inevitable that exceptions will occur and fault recovery is required. When performing fault recovery, on the one hand, it is necessary to ensure the consistency of the data to be migrated in the source database and the data migrated to the target database, and on the other hand, it is hoped that the efficiency of fault recovery is as high as possible.

[0081] The following introduces two ways of fault recovery in related technologies.

[0082] In the first way, when there is an execution error in the write commands of the current batch, the migration tool resends the write commands of the current batch to the target database for execution. Generally, when there is an execution error in the write commands of a certain batch, only a part of the write commands in that batch has an execution error. At this time, if all the write commands of the current batch are repeatedly executed, those write commands that have been successfully executed will be repeatedly executed. If the data type corresponding to the write command has a non-idempotent property, it will cause the written data to be disordered, that is, the data to be migrated in the source database is inconsistent with the data migrated to the target database. The following is an example. As Figure 2 shown, the data corresponding to the write commands of a certain batch includes list-type data and set-type data. When an execution error occurs in the execution of the write commands of this batch, and this method is used for fault recovery, the data migrated to the target database is inconsistent with the data in the source database.

[0083] In the second way, the data already written in the target database is cleared, and then the migration tool re-executes the migration operation. By adopting this processing method, although the consistency between the data to be migrated in the source database and the data migrated to the target database is ensured, it will lead to low efficiency of fault recovery and increase costs, such as increasing the load of the source database, requiring manual intervention in resource consumption, blocking services, etc., which affects the business of customers.

[0084] The embodiments of the present disclosure provide a data synchronization method. Based on the idempotent / non-idempotent properties of data, the data to be migrated in the source database is divided, and different fault recovery methods are respectively set. In this way, not only can the consistency between the data to be migrated in the source database and the data migrated to the target database be ensured, but also a high fault recovery efficiency can be ensured. The processing flow of this method is as Figure 3 shown, and it includes the following steps:

[0085] Step 301, the migration tool obtains the data to be migrated in the source database.

[0086] A user or an operation and maintenance personnel of the migration tool can operate related devices to send a migration notification to the migration tool. After receiving the migration notification, the migration tool sends a fetch command to the source database. This fetch command can be used to obtain the data to be migrated in the source database. The data to be migrated can be all the data in the source database or part of the data in the source database. For example, when the source database is a Redis database, this fetch command can be "PSYNC", and this fetch command can also be set by technicians themselves. The embodiments of the present disclosure do not make any limitations in this regard.

[0087] After receiving the acquisition command, the source database sends the data to be migrated to the migration tool. The specific sending method can be to generate an RDB file based on the data to be migrated in the source database, and then send the RDB file to the migration tool. When the source database uses distributed storage, the data in the source database can be stored on multiple storage nodes. For a data set, it can be stored on one storage node or on different storage nodes. As Figure 4 shown, the data in a certain database is stored on three storage nodes. db1 and db2 are data set identifiers, representing two different data sets. The entire data set db1 is stored in storage node 1. The data set db2 is divided into db2_1, db2_2, and db2_3, and stored in storage node 1, storage node 2, and storage node 3 respectively. After receiving the RDB file, the migration tool can divide the data based on the storage location of the data and the data set identifier of the data. The corresponding processing can be: for all the data in the RDB file, perform the first division based on the storage node where the data is located, and then for the data in each storage node, perform a further division based on the data set identifier of the data, so as to obtain the data in each data set in each storage node. The data obtained by this division can be regarded as a unit of data. The processing of this process can be performed separately on each unit of data.

[0088] In this way, when synchronizing data to the target database later, the storage architecture of the data in the target database is the same as that of the data in the source database. For example, when the storage of the data in the source database is as Figure 4 shown, the division of the data in the RDB file by the migration tool is as Figure 5 shown. The first division respectively obtains the data of storage node 1, storage node 2, and storage node 3, and then performs a second division to obtain the data of different data sets in each storage node.

[0089] Step 302, the migration tool divides the data to be migrated into two groups: the first data to be migrated and the second data to be migrated.

[0090] The first data to be migrated has non-idempotent characteristics and can be simply referred to as non-idempotent data. The second data to be migrated has idempotent characteristics and can be simply referred to as idempotent data.

[0091] Step 303, the migration tool synchronizes the first data to be migrated and the second data to be migrated to the target database respectively.

[0092] The migration tool sends the write command of the first data to be migrated or the second data to be migrated to the target database for execution to write the first data to be migrated or the second data to be migrated into the target database.

[0093] The synchronization processing of the first data to be migrated and the second data to be migrated can be carried out batch by batch. Correspondingly, the migration tool can synchronize the first data to be migrated to the target database in multiple batches, or synchronize the second data to be migrated to the target database in multiple batches. The number of data in each batch can be set based on actual needs.

[0094] Step 304, when the synchronization of the first data to be migrated or the second data to be migrated fails, select the corresponding fault recovery method for the first data to be migrated or the second data to be migrated to perform fault recovery.

[0095] The fault recovery methods for the first data to be migrated and the second data to be migrated are different. For the first data to be migrated (i.e., non-idempotent data), the processing flow of this step is as Figure 6 shown; for the second data to be migrated (i.e., idempotent data), the processing flow of this step is as Figure 9 shown.

[0096] For data with non-idempotent characteristics, the processing flow of the data synchronization method is as Figure 6 shown, including the following steps:

[0097] Step 601, the migration tool obtains the first data to be migrated from the data to be written into the target database.

[0098] Among them, the first data to be migrated is data with non-idempotent characteristics. Non-idempotent data has the following characteristics: for a write command corresponding to non-idempotent data, the final written result of the data is different when executed once and when executed multiple times. When the database is a Redis database, non-idempotent data includes list-type data.

[0099] For the data in the source database, there may be one or multiple types of non-idempotent data. The following is an explanation of different situations: Situation 1, there is only one type of non-idempotent data, and this non-idempotent data can be directly obtained for subsequent step processing; Situation 2, there are multiple types of non-idempotent data, and each type of non-idempotent data can be divided into a group, and each type of non-idempotent data is synchronized to the target database respectively (that is, each type of non-idempotent data can be obtained respectively, and subsequent steps are processed for each type of non-idempotent data); Situation 3, there are multiple types of non-idempotent data, and all non-idempotent data can be obtained, and these data are uniformly processed in subsequent steps without distinction.

[0100] Step 602, the migration tool synchronizes the first data to be migrated to the target database.

[0101] The migration tool sends the write command of the first data to be migrated to the target database for execution, and can send it to the target database for execution in multiple batches.

[0102] The migration tool can be pre-set with a specified number. The data of the specified number forms a batch of data. The specified number can be arbitrarily set by technicians according to relevant requirements. For example, it can be set by considering the maximum data volume of a single data packet and the data volume of a single write instruction. The embodiments of the present disclosure do not limit this. For example, the specified number can be set to 1000 or 2000, etc.

[0103] The migration tool can generate corresponding write commands based on the type of data. All the write commands corresponding to the data of the specified number form a batch of write commands. The migration tool sends the write commands to the target database in batches. The target database executes the received write commands in batches and feeds back the execution results of this batch of write commands to the migration tool after executing a batch of write commands.

[0104] Step 603, when the synchronization of the first data to be migrated fails, the migration tool notifies the target database to delete the first data to be migrated that has been synchronized to the target database, and resynchronizes the first data to be migrated to the target database.

[0105] For the three cases of step 601, the processing methods of step 603 are also different. For case one and case three, the corresponding processing is as follows: when the migration tool determines that there is an execution error in the execution result feedback by the target database, delete all non-idempotent data that has been synchronized to the target database, and re-execute step 602. For case two, the corresponding processing is as follows: for a certain type of non-idempotent data, when the migration tool determines that there is an execution error in the execution result feedback by the target database, delete the non-idempotent data of this type that has been synchronized to the target database, and re-execute step 602.

[0106] For example, there is only one type of non-idempotent data. After the migration tool obtains the message that there is an execution error in the 8001-9000th write commands, clear the non-idempotent data of this type that has been written to the target database, and re-execute step 602.

[0107] When executing the above steps 602 and 603, the migration tool can record the synchronization site identification information. The synchronization site identification information indicates the latest batch of the first data to be migrated that has been successfully synchronized to the target database. When the synchronization of the first data to be migrated in the target batch fails, notify the target database to delete the first data to be migrated that has been synchronized to the target database, and reset the site identification information to the initial value, and resynchronize the first data to be migrated to the target database. The corresponding processing can be as Figure 7 shown, including the following steps:

[0108] Step 701, the migration tool records the first data as the reference data.

[0109] The migration tool can apply for a cache space to record the synchronization site identification information, which can be simply referred to as site information. The site information can be the identification information of the reference data, and this identification information can be the offset of the reference data in the RDB file, such as Figure 8 shown. The reference data can be the first data in the batch data that the migration tool is writing.

[0110] Step 702: The migration tool obtains the write commands corresponding to a specified number of data starting from the reference data, encapsulates the write commands corresponding to the specified number of data, and sends them to the target database.

[0111] After determining the reference data, the migration tool can generate write commands for the data one by one starting from the reference data and cache the generated write commands. When the write commands corresponding to a specified number of data are generated, these write commands can be encapsulated and sent to the target database.

[0112] If, during the process of generating the write commands, when generating the write command corresponding to the last data, the specified number has not been reached, the processing of generating the current write commands can be ended, and the current batch of write commands can be encapsulated and sent to the target database.

[0113] Step 703: When it is determined that the write commands are successfully executed based on the execution results feedback by the target database, the migration tool records the data after the specified number of data as the reference data and proceeds to execute Step 702.

[0114] After receiving the encapsulated write commands, the target database first unpacks them and then executes the write commands in sequence. Each write command corresponds to an execution result, and the execution result can include successful execution or execution failure. For the case of execution failure, it can also include the reason for failure. For example, the reason for failure can be that the write command received by the target database has an error, which may be an error during generation or an error during transmission (such as the write command being partially or completely lost due to a receive queue overflow), and the reason for failure can also be that the target database has an error in executing the write command, etc. After all the write commands in this batch are executed, the target database encapsulates and sends the execution results of this batch of write commands to the migration tool. The migration tool unpacks the execution results of this batch of write commands. When it is determined that all the execution results of this batch of write commands are successful executions, the reference data is updated to the data after this batch of data, that is, the site information in the cache space is updated to the identification information of the data after this batch of data, and then proceeds to execute Step 702. If there is no other data after this batch of data, the process is ended.

[0115] The following is an example of the above steps 701-703. For example, 1000 data items are considered as a batch of data, and each data item corresponds to a write command, that is, 1000 write commands are a batch of write commands. The migration tool records the first data item as the reference data, and then encapsulates and sends the 1-1000 write commands corresponding to the 1-1000 data items to the target database. After receiving the message that all the 1-1000 write commands have been successfully executed, the migration tool records the 1001st data item as the reference data, and then encapsulates and sends the 1001-2000 write commands corresponding to the 1001-2000 data items to the target database. After receiving the message that all the 1001-2000 write commands have been successfully executed, the migration tool records the 2001st data item as the reference data, and repeats the above process until all the write commands for the data to be written to the target database have been successfully executed.

[0116] Step 704, when it is determined based on the execution result feedback from the target database that there are write commands with execution errors, the migration tool clears the non-idempotent data that has been written to the target database and proceeds to execute step 701.

[0117] For data with idempotent characteristics, the processing flow of the data synchronization method is as Figure 9 shown, including the following steps:

[0118] Step 901, the migration tool obtains the second data to be migrated from the data to be written to the target database.

[0119] Among them, the second data to be migrated is data with idempotent characteristics. Idempotent data has the following characteristics: for a write command corresponding to idempotent data, the final written result of the data is the same whether it is executed once or repeatedly. When the database is a Redis database, idempotent data includes string-type data, hash-type data, set-type data, and zset-type data, and the subsequent steps are processed for each type of idempotent data respectively.

[0120] Step 902, the migration tool synchronizes the second data to be migrated to the target database.

[0121] Step 903, when the synchronization of the target data in the second data to be migrated fails, the migration tool synchronizes the target data to the target database again.

[0122] When there are multiple types of the second data to be migrated (for example, multiple types among string-type data, hash-type data, set-type data, and zset-type data), each type of the second data to be migrated can be grouped, and each type of the second data to be migrated is synchronized to the target database respectively. When the synchronization of the target type of data in the second data to be migrated fails, the target type of data is synchronized to the target database again.

[0123] Each group of data can be synchronized batch by batch, that is, a group of data is synchronized to the target database in multiple batches. When a failure occurs in synchronizing the data of the current batch of a certain group of data, the data of this batch can be resynchronized to the target database.

[0124] The migration tool can record the synchronization site identification information, which has been introduced in the previous content. When a failure occurs in synchronizing the second data to be migrated in the target batch, according to the synchronization site identification information, the second data to be migrated in the target batch is resynchronized to the target database. The corresponding processing procedure is similar to that Figure 7 shown in the process, the only difference is that step 704 is changed to: when it is determined based on the execution result feedback by the target database that there is a write command with an execution error, the migration tool keeps the recorded reference data unchanged and proceeds to execute step 702. The target data is all the data corresponding to the write command of the batch (i.e., the target batch) where the current synchronization fails.

[0125] The following is an example to illustrate the processing procedure of data synchronization, as Figure 10 and Figure 11 shown, the database is a Redis database, the non-idempotent data includes list-type data, and the idempotent data includes string-type data, hash-type data, set-type data, and zset-type data. The corresponding processing procedure is as follows:

[0126] Step 1, the source database sends the data to be migrated to the migration tool.

[0127] Step 2, the migration tool divides the received data into non-idempotent data and idempotent data.

[0128] For the idempotent data (list-type data), the following processing steps 3(A)-10(A) are taken:

[0129] Step 3(A), the migration tool generates a write command based on the list-type data, where every 1000 write commands are the write commands of one batch.

[0130] Step 4(A), the migration tool sends the write commands to the target database batch by batch.

[0131] Step 5(A), the target database executes the write commands batch by batch.

[0132] Step 6(A), the target database feeds back the execution result to the migration tool batch by batch.

[0133] Step 7(A), when the migration tool receives an execution error in the execution result of the Xth batch, it sends a deletion notice of the list-type data to the target database.

[0134] Step 8(A), the target database deletes the list-type data written in each batch.

[0135] Step 9(A), the migration tool resends the write command for the list-type data to the target database in batches.

[0136] Step 10(A), the target database executes the write command in batches again, and all executions are successful. (Here, it is illustrated by taking that all the resubmitted write commands are successfully executed. Additionally, if there is an execution error for the write command of a certain batch in the subsequent process, it can be processed in a similar manner to steps 3(A)-10(A).)

[0137] For each type of idempotent data, the following processing steps 3(B)-8(B) are taken, that is, for string-type data, hash-type data, set-type data, and zset-type data, the following processing steps 3(B)-8(B) are respectively taken (here, it is illustrated by taking string-type data as an example, and the same applies to other types without repeated description):

[0138] Step 3(B), the migration tool generates a write command based on the string-type data, where every 1000 write commands form a batch of write commands.

[0139] Step 4(B), the migration tool sends the write command to the target database in batches.

[0140] Step 5(B), the target database executes the write command in batches.

[0141] Step 6(B), the target database feeds back the execution result to the migration tool in batches.

[0142] Step 7(B), when the migration tool receives that the execution result of the Y-th batch is an execution error, it resends the write command of the Y-th batch and continues to send the subsequent write commands in batches.

[0143] Step 8(B), the target database executes the write commands of the Y-th batch and the subsequent batches, and all executions are successful. (Here, it is illustrated by taking that all the write commands of the Y-th batch and the subsequent batches are successfully executed. Additionally, if there is an execution error for the write command of a certain batch in the subsequent process, it can be processed in a similar manner to steps 7(B)-8(B).)

[0144] Based on the same technical concept, an embodiment of the present disclosure provides a migration device, which can be applied to the above-mentioned migration tool, as Figure 12 shown, the device includes:

[0145] An acquisition module 1210, configured to acquire data to be migrated in a source database. Specifically, it can implement the processing functions of step 301 above, as well as other implicit steps.

[0146] A partitioning module 1220, configured to divide the data to be migrated into two groups: the first data to be migrated and the second data to be migrated. Among them, the first data to be migrated has non-idempotent characteristics, and the second data to be migrated has idempotent characteristics. Specifically, it can implement the processing functions of step 302 above, as well as other implicit steps.

[0147] A processing module 1230, configured to synchronize the first data to be migrated and the second data to be migrated to a target database respectively. When the synchronization of the first data to be migrated or the second data to be migrated fails, a failure recovery method corresponding to the first data to be migrated or the second data to be migrated is selected for failure recovery, where the failure recovery methods for the first data to be migrated and the second data to be migrated are different. Specifically, it can implement the processing functions of steps 303 - 304 above, as well as other implicit steps.

[0148] In a possible implementation manner, the processing module 1230 is configured to, when the synchronization of the first data to be migrated fails, notify the target database to delete the first data to be migrated that has been synchronized to the target database, and then re-synchronize the first data to be migrated to the target database. Specifically, it can implement the processing functions of steps 601 - 603 above, as well as other implicit steps.

[0149] In a possible implementation manner, the processing module 1230 is configured to, when the synchronization of the target data in the second data to be migrated fails, re-synchronize the target data to the target database. Specifically, it can implement the processing function of step 903 above, as well as other implicit steps.

[0150] In a possible implementation manner, the processing module 1230 is configured to send a write command for the first data to be migrated or the second data to be migrated to the target database for execution, so as to write the first data to be migrated or the second data to be migrated into the target database. Specifically, it can implement the processing function of step 303 above, as well as other implicit steps.

[0151] In a possible implementation manner, the second data to be migrated includes multiple types of data. The processing module 1230 is configured to divide each type of the second data to be migrated into a group, and synchronize each type of the second data to be migrated to the target database respectively. When the synchronization of the target type of data in the second data to be migrated fails, re-synchronize the target type of the second data to be migrated to the target database. Specifically, it can implement the processing functions of steps 902 - 903 above, as well as other implicit steps.

[0152] In a possible implementation, the processing module 1230 is configured to synchronize the first data to be migrated or the second data to be migrated to the target database in multiple batches. Specifically, it can implement the processing functions of steps 902-903 above, as well as other implicit steps.

[0153] In a possible implementation, the processing module 1230 is further configured to record the synchronization site identification information, where the synchronization site identification information indicates the latest batch of the second data to be migrated that has been successfully synchronized to the target database. When the synchronization of the target batch of the second data to be migrated fails, the target batch of the second data to be migrated is resynchronized to the target database according to the synchronization site identification information. Specifically, it can implement the processing functions of steps 902-903 above, as well as other implicit steps.

[0154] In a possible implementation, the first data to be migrated includes list-type data.

[0155] In a possible implementation, the second data to be migrated includes at least one of string-type data, hash-type data, set-type data, and sorted set-type data.

[0156] The embodiments of the present disclosure provide a data synchronization method. Based on the idempotent / non-idempotent characteristics of the data, the data to be migrated in the source database is divided, and different fault recovery methods are set respectively. In this way, it can not only ensure the consistency between the data to be migrated in the source database and the data migrated to the target database, but also ensure a high fault recovery efficiency.

[0157] Among them, the acquisition module 1210, the division module 1220, and the processing module 1230 can all be implemented by software or by hardware. Exemplarily, next, taking the acquisition module 1210 as an example, the implementation manner of the acquisition module 1210 is introduced. Similarly, the implementation manners of the division module 1220 and the processing module 1230 can refer to the implementation manner of the acquisition module 1210.

[0158] As an example of a software functional unit, the obtaining module 1210 may include code running on a computing instance. The computing instance may include at least one of a physical host (computing device), a virtual machine, and a container. Further, the above computing instance may be one or more. For example, the obtaining module 1210 may include code running on multiple hosts / virtual machines / containers. It should be noted that the multiple hosts / virtual machines / containers for running the code may be distributed in the same region or in different regions. Further, the multiple hosts / virtual machines / containers for running the code may be distributed in the same availability zone (AZ) or in different AZs, and each AZ includes one data center or multiple geographically proximate data centers. Usually, one region may include multiple AZs.

[0159] Similarly, the multiple hosts / virtual machines / containers for running the code may be distributed in the same virtual private cloud (VPC) or in multiple VPCs. Usually, one VPC is set within one region. For cross-region communication between two VPCs within the same region and between VPCs in different regions, a communication gateway needs to be set in each VPC, and the interconnection between VPCs is realized through the communication gateway.

[0160] As an example of a hardware functional unit, the obtaining module 1210 may include at least one computing device, such as a server. Alternatively, the obtaining module 1210 may also be a device implemented by an application-specific integrated circuit (ASIC) or a programmable logic device (PLD). Among them, the above PLD may be implemented by a complex programmable logic device (CPLD), a field-programmable gate array (FPGA), a generic array logic (GAL), or any combination thereof.

[0161] The multiple computing devices included in the acquisition module 1210 may be distributed in the same region or in different regions. The multiple computing devices included in the acquisition module 1210 may be distributed in the same availability zone (AZ) or in different AZs. Similarly, the multiple computing devices included in the acquisition module 1210 may be distributed in the same virtual private cloud (VPC) or in multiple VPCs. Among them, the multiple computing devices may be any combination of computing devices such as servers, application-specific integrated circuits (ASICs), programmable logic devices (PLDs), complex programmable logic devices (CPLDs), field-programmable gate arrays (FPGAs), and generic array logic (GALs).

[0162] It should be noted that in other embodiments, the acquisition module 1210, the partitioning module 1220, and the processing module 1230 may be used in any step of the data synchronization method in the migration tool. The steps to be implemented by the acquisition module 1210, the partitioning module 1220, and the processing module 1230 can be specified as needed. The entire function of the device for the migration tool to perform data synchronization is realized by the acquisition module 1210, the partitioning module 1220, and the processing module 1230 respectively implementing different steps in the data synchronization method.

[0163] The present disclosure also provides a data migration system, as Figure 1 shown, including a source database, a target database, and a migration tool. The functions of the source database, the target database, and the migration tool have been introduced above and will not be elaborated here.

[0164] The source database, the target database, and the migration tool can all be implemented by software or can be implemented by hardware. Exemplarily, the implementation manner of the migration tool will be introduced next. Similarly, the implementation manners of the source database and the target database can refer to the implementation manner of the migration tool.

[0165] As an example of a software functional unit, the migration tool may include code running on a computing instance. Among them, the computing instance may be at least one of computing devices such as a physical host (computing device), a virtual machine, and a container. Further, the above-mentioned computing devices may be one or more. For example, the device for performing data synchronization may include code running on multiple hosts / virtual machines / containers. It should be noted that the multiple hosts / virtual machines / containers for running the application may be distributed in the same region or in different regions. The multiple hosts / virtual machines / containers for running the code may be distributed in the same AZ or in different AZs, and each AZ includes one data center or multiple geographically proximate data centers. Among them, generally, one region may include multiple AZs.

[0166] Similarly, multiple hosts / virtual machines / containers for running the code can be distributed in the same VPC or in multiple VPCs. Usually, one VPC is set up in one region. For cross-region communication between two VPCs within the same region and between VPCs in different regions, communication gateways need to be set up in each VPC, and the interconnection between VPCs is achieved through the communication gateways.

[0167] As an example of a hardware functional unit, the migration tool can be deployed on at least one computing device, such as a server. Alternatively, the device for synchronizing data can also be a device implemented by ASIC or PLD. Among them, the above PLD can be implemented by CPLD, FPGA, GAL or any combination thereof.

[0168] Multiple computing devices deployed with the migration tool can be distributed in the same region or in different regions. Multiple computing devices deployed with the migration tool can be distributed in the same AZ or in different AZs. Similarly, multiple computing devices deployed with the migration tool can be distributed in the same VPC or in multiple VPCs. Among them, the multiple computing devices can be any combination of computing devices such as servers, ASICs, PLDs, CPLDs, FPGAs, and GALs.

[0169] The present disclosure also provides a computing device 100. The computing device 100 can be applied to the above source database, target database, and migration tool. As Figure 13 shown, the computing device 100 includes: a bus 102, a processor 104, a memory 106, and a communication interface 108. The processor 104, the memory 106, and the communication interface 108 communicate with each other through the bus 102. The computing device 100 can be a server or a terminal device. It should be understood that the present disclosure does not limit the number of processors and memories in the computing device 100.

[0170] The bus 102 can be a peripheral component interconnect (PCI) bus or an extended industry standard architecture (EISA) bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc. For the sake of representation, Figure 8 only one line is used herein, but it does not mean that there is only one bus or one type of bus. The bus 102 can include a path for transmitting information between various components (for example, the memory 106, the processor 104, the communication interface 108) of the computing device 100.

[0171] The processor 104 may include any one or more of processors such as a central processing unit (CPU), a graphics processing unit (GPU), a microprocessor (MP), or a digital signal processor (DSP).

[0172] The memory 106 may include volatile memory, such as random access memory (RAM). The memory 106 may also include non-volatile memory, such as read-only memory (ROM), flash memory, a hard disk drive (HDD), or a solid state drive (SSD).

[0173] The memory 106 stores executable program code, and the processor 104 executes the executable program code to implement the functions of the foregoing acquisition module 1210, division module 1220, and processing module 1230 respectively, thereby implementing the method for data synchronization. That is, the memory 106 stores instructions for the method of data synchronization.

[0174] Alternatively, the memory 106 stores executable code, and the processor 104 executes the executable code to implement the functions of the foregoing acquisition module 1210, division module 1220, and processing module 1230 respectively, thereby implementing the method for data synchronization. That is, the memory 106 stores instructions for the method of data synchronization.

[0175] The communication interface 108 uses a transceiver module such as, but not limited to, a network interface card or a transceiver to implement communication between the computing device 100 and other devices or a communication network.

[0176] The embodiments of the present disclosure further provide a computing device cluster. The computing device cluster includes at least one computing device. The computing device may be a server, such as a central server, an edge server, or a local server in a local data center. In some embodiments, the computing device may also be a terminal device such as a desktop computer, a laptop computer, or a smart phone.

[0177] As Figure 14 shown, the computing device cluster includes at least one computing device 100. The memory 106 in one or more of the computing devices 100 in the computing device cluster may store the same instructions for the method of data synchronization.

[0178] In some possible implementations, partial instructions for the data synchronization method may also be stored respectively in the memories 106 of one or more computing devices 100 in the computing device cluster. In other words, the combination of one or more computing devices 100 can jointly execute the instructions for the data synchronization method.

[0179] It should be noted that the memories 106 in different computing devices 100 in the computing device cluster may store different instructions for partial functions of the data synchronization apparatus. That is, the instructions stored in the memories 106 of different computing devices 100 can implement the functions of one or more of the foregoing obtaining module 1210, partitioning module 1220, and processing module 1230.

[0180] In some possible implementations, one or more computing devices in the computing device cluster can be connected via a network. Among them, the network can be a wide area network or a local area network, etc., and the network can be a Transmission Control Protocol (TCP) network or a Remote Direct Memory Access (RDMA) network, etc. Figure 15 A possible implementation is shown. As Figure 15 shown, two computing devices 100A and 100B are connected via a network. Specifically, they are connected to the network through the communication interfaces in each computing device. In this type of possible implementation, the memory 106 in the computing device 100A stores instructions for executing the functions of the obtaining module 1210 and the partitioning module 1220. At the same time, the memory 106 in the computing device 100B stores instructions for executing the function of the processing module 1230.

[0181] Figure 15 The connection manner between the computing device clusters shown can be considered that since the method for providing data synchronization in the present disclosure requires a large amount of data storage, the function implemented by the processing module 1230 is considered to be executed by the computing device 100B.

[0182] It should be understood that Figure 15 the function of the computing device 100A shown in

[0183] can also be completed by multiple computing devices 100. Similarly, the function of the computing device 100B can also be completed by multiple computing devices 100. Figure 14 and Figure 15The connection method of the computing device cluster. Differently, instructions of the same method for data synchronization may be stored in the memory 106 of one or more computing devices 100 in the computing device cluster.

[0184] In some possible implementation manners, partial instructions of the method for data synchronization may also be separately stored in the memory 106 of one or more computing devices 100 in the computing device cluster. In other words, a combination of one or more computing devices 100 may jointly execute the instructions of the method for data synchronization.

[0185] The embodiments of the present disclosure also provide a computer program product including instructions. The computer program product may be software or a program product including instructions that can run on a computing device or be stored in any available medium. When the computer program product runs on at least one computing device, at least one computing device is caused to execute the method for data synchronization.

[0186] The embodiments of the present disclosure also provide a computer-readable storage medium. The computer-readable storage medium may be any available medium that a computing device can store or a data storage device such as a data center including one or more available media. The available medium may be a magnetic medium (for example, a floppy disk, a hard disk, a magnetic tape), an optical medium (for example, a digital video disk (DVD)), or a semiconductor medium (for example, a solid-state drive), etc. The computer-readable storage medium includes instructions that instruct the computing device to perform the method for data synchronization or instruct the computing device to execute the method for data synchronization.

[0187] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present disclosure and are not intended to limit them; although the present disclosure has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the protection scope of the technical solutions of the embodiments of the present disclosure.

Claims

1. A data synchronization method, characterized in that: The method comprises: Obtain the data to be migrated from the source database; Dividing the data to be migrated into two groups: first data to be migrated and second data to be migrated, wherein the first data to be migrated has a non-idempotent property, and the second data to be migrated has an idempotent property; Synchronizing the first data to be migrated and the second data to be migrated to a target database respectively; When synchronization of the first data to be migrated or the second data to be migrated fails, a fault recovery method corresponding to the first data to be migrated or the second data to be migrated is selected for fault recovery, wherein the first data to be migrated and the second data to be migrated have different fault recovery methods.

2. The method according to claim 1, characterized in that When synchronization of the first data to be migrated fails, selecting a failure recovery method corresponding to the first data to be migrated to perform failure recovery includes: When synchronization of the first data to be migrated fails, notifying the target database to delete the first data to be migrated that has been synchronized to the target database; Resynchronize the first data to be migrated to the target database.

3. The method according to claim 1, characterized in that When synchronization of the second data to be migrated fails, selecting a failure recovery method corresponding to the second data to be migrated to perform failure recovery includes: When synchronization of target data in the second data to be migrated fails, the target data is synchronized to the target database again.

4. The method according to any one of claims 1 to 3, characterized in that Synchronizing the first to-be-migrated data or the second to-be-migrated data to the target database includes: A write command of the first data to be migrated or the second data to be migrated is sent to the target database for execution, so as to write the first data to be migrated or the second data to be migrated into the target database.

5. The method according to any one of claims 1 to 4, characterized in that The second data to be migrated includes multiple types of data; The synchronizing the second data to be migrated to the target database includes: grouping each type of the second data to be migrated into a group, and synchronizing each type of the second data to be migrated to the target database respectively; When synchronization of the second data to be migrated fails, selecting a failure recovery method corresponding to the second data to be migrated to perform failure recovery includes: When synchronization of the target type of data in the second data to be migrated fails, the target type of the second data to be migrated is synchronized to the target database again.

6. The method according to any one of claims 1 to 5, characterized in that Synchronizing the first to-be-migrated data or the second to-be-migrated data to the target database includes: The first data to be migrated or the second data to be migrated is synchronized to the target database in multiple batches.

7. The method according to claim 6, characterized in that The method further includes: recording synchronization site identification information, the synchronization site identification information indicating the latest batch of second data to be migrated that has been successfully synchronized to the target database; When the synchronization of the second data to be migrated fails, a fault recovery method corresponding to the second data to be migrated is selected for fault recovery, including: when the synchronization of the second data to be migrated of the target batch fails, re-synchronizing the second data to be migrated of the target batch to the target database according to the synchronization site identification information.

8. The method according to any one of claims 1 to 7, characterized in that The first data to be migrated includes list-type data.

9. The method according to any one of claims 1 to 7, characterized in that The second data to be migrated includes at least one of string type data, dictionary hash type data, set type data and sorted set zset type data.

10. A migration device, characterized in that: The device comprises: An acquisition module, used to acquire the data to be migrated in the source database; A division module, used for dividing the data to be migrated into two groups of first data to be migrated and second data to be migrated, wherein the first data to be migrated has a non-idempotent property, and the second data to be migrated has an idempotent property; A processing module is used to synchronize the first data to be migrated and the second data to be migrated to the target database respectively; when the synchronization of the first data to be migrated or the second data to be migrated fails, select a fault recovery method corresponding to the first data to be migrated or the second data to be migrated for fault recovery, wherein the fault recovery methods of the first data to be migrated and the second data to be migrated are different.

11. The device according to claim 10, characterized in that The processing module is used for: When synchronization of the first data to be migrated fails, notifying the target database to delete the first data to be migrated that has been synchronized to the target database; Resynchronize the first data to be migrated to the target database.

12. The device according to claim 10, characterized in that The processing module is used for: When synchronization of target data in the second data to be migrated fails, the target data is synchronized to the target database again.

13. The device according to any one of claims 10 to 12, characterized in that The processing module is used for: A write command of the first data to be migrated or the second data to be migrated is sent to the target database for execution, so as to write the first data to be migrated or the second data to be migrated into the target database.

14. The device according to any one of claims 10 to 13, characterized in that The second data to be migrated includes multiple types of data; The processing module is used to group each type of second data to be migrated into one group, and synchronize each type of second data to be migrated to the target database respectively; When synchronization of the target type of data in the second data to be migrated fails, the target type of the second data to be migrated is synchronized to the target database again.

15. The device according to any one of claims 10 to 14, characterized in that The processing module is used for: The first data to be migrated or the second data to be migrated is synchronized to the target database in multiple batches.

16. The device according to claim 15, characterized in that The processing module is further used for: Recording synchronization site identification information, where the synchronization site identification information indicates the latest batch of second data to be migrated that has been successfully synchronized to the target database; When synchronization of the second data to be migrated of the target batch fails, the second data to be migrated of the target batch is resynchronized to the target database according to the synchronization site identification information.

17. The device according to any one of claims 10 to 16, characterized in that The first data to be migrated includes list-type data.

18. The device according to any one of claims 10 to 16, characterized in that The second data to be migrated includes at least one of string type data, dictionary hash type data, set type data and sorted set zset type data.

19. A computing device cluster, characterized in that: comprising at least one computing device, each computing device comprising a processor and a memory; The processor of the at least one computing device is configured to execute instructions stored in the memory of the at least one computing device, so that the computing device cluster executes the method according to any one of claims 1 to 9.

20. A computer-readable storage medium, characterized in that: The method comprises computer program instructions, and when the computer program instructions are executed by a computing device cluster, the computing device cluster performs the method according to any one of claims 1 to 9.

21. A computer program product comprising instructions, characterized in that When the instructions are executed by a computing device cluster, the computing device cluster is caused to perform the method according to any one of claims 1 to 9.