A Method and Device for Initializing Parallel Synchronization of Full Data and Incremental Data

By creating multiple packets on the target end to process initializing full data and incremental data in parallel, the problem of lack of incremental data parallel cache and synchronization in the process of initializing full data synchronization in heterogeneous database data migration is solved, and the efficiency of data migration and synchronization is improved.

CN115470295BActive Publication Date: 2025-05-30WUHAN DAMENG DATABASE
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211051430.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-08-31
Publication Date
2025-05-30
Estimated Expiration
2042-08-31

AI Technical Summary

Technical Problem

In the process of heterogeneous database data migration, there is a lack of a method to perform incremental data cache and synchronization in parallel during the initialization of full data synchronization, resulting in inefficient data migration.

Method used

By creating three groups on the target end: one is used to cache the initialized full data, and two are used to cache the incremental data before the initialized full data is completed and the incremental data after the synchronization is completed. After the initialized full data is completed, all incremental data are synchronized to achieve parallel synchronization between the incremental data and the initialized full data.

Benefits of technology

The efficiency of data migration and synchronization of heterogeneous databases is improved, and the data synchronization time is shortened by initializing full and incremental data in parallel processing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115470295B_ABST
    Figure CN115470295B_ABST
Patent Text Reader

Abstract

The present invention provides a method and device for parallel synchronization of initialization full amount data and incremental data. During the process of data migration from the source end to the target end, the initialization full amount data of the synchronization table and the incremental data of the operation log are sent in parallel, and three different groups are created at the target end. One group is used to cache the initialization full amount data and perform synchronization, and the other two groups respectively cache the incremental data migrated before the synchronization of the initialization full amount data is completed, and cache the incremental data migrated after the synchronization of the initialization full amount data is completed, and all the incremental data is synchronized after the synchronization of the initialization full amount data is completed, realizing the parallel synchronization of the incremental data and the initialization full amount data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of database synchronization, and particularly to a method and device for parallel synchronization of initial full data and incremental data. Background Art

[0002] When performing real-time synchronization of heterogeneous database data, in order to ensure the consistency of the initial full data between the source end and the target end, generally, the initial full data of the source-end database is first migrated to the target-end database to establish an initial point for data synchronization, and then, based on this, the real-time synchronization function of incremental data is started to achieve real-time synchronization of data between the source-end and target-end databases.

[0003] During the process of heterogeneous data synchronization, the initial full data synchronization generally adopts database query extraction technology, directly queries the data of the synchronization table through the database interface, and then sends it to the target end for storage synchronization. The real-time synchronization of incremental data is generally based on the capture and parsing technology of database archive logs, and restores the transaction operation logic of the source-end database at the target end. Based on this, the initial full data and incremental data have different characteristics. For example, incremental data generally belongs to a specific synchronization transaction and has corresponding transaction IDs, operation timestamps, commit LSNs, etc., while the initial full data does not have these characteristics. Therefore, when synchronizing and storing at the target end, the initial full data and incremental data cannot be processed in parallel according to a unified transaction operation logic. When performing synchronization, generally, the initial full data is first cached and synchronized, and after the caching and synchronization process of the full data is completed, the caching and synchronization of incremental data are performed. This process is a serial synchronization process.

[0004] Currently, there is a lack of a method for parallelly caching incremental data during the caching and synchronization process of initial full data and synchronizing incremental data after the synchronization of initial full data is completed.

[0005] In view of this, overcoming the above-mentioned defects of the prior art is an urgent problem to be solved in this technical field. Summary of the Invention

[0006] The technical problem to be solved by the present invention is to provide a method for parallelly performing initial full data synchronization and incremental data synchronization when migrating data in a heterogeneous database.

[0007] The embodiments of the present invention adopt the following technical solutions:

[0008] In a first aspect, a method for parallel synchronization of initial full data and incremental data is provided, including:

[0009] The target end creates a first group, a second group, and a third group;

[0010] Receive the initial full - volume data and incremental data of the table to be synchronized from the source end;

[0011] Cache the initial full - volume data into the third group for synchronizing the initial full - volume data of the table to be synchronized;

[0012] Cache the first incremental data into the first group and cache the second incremental data into the second group, where the first incremental data includes the incremental data generated after the synchronization of the initial full - volume data is completed, and the second incremental data includes the incremental data generated before the synchronization of the initial full - volume data of the table to be synchronized is completed;

[0013] After completing the synchronization of the initial full - volume data of the table to be synchronized, synchronize the incremental data in the first group and the second group to achieve parallel synchronization of the incremental data and the initial full - volume data.

[0014] Preferably, the synchronizing the incremental data in the first group and the second group includes:

[0015] For the incremental data in the first group, after completing the synchronization of the initial full - volume data of the table to be synchronized and after receiving the commit operation of the transaction corresponding to the incremental data, perform a synchronization operation on the corresponding incremental data.

[0016] Preferably, the synchronizing the incremental data in the first group and the second group includes:

[0017] For the incremental data in the second group, after the synchronization of the initial full - volume data of the table to be synchronized is completed, determine whether there is third incremental data in the second group that belongs to the table to be synchronized;

[0018] If there is third incremental data in the second group that belongs to the table to be synchronized, determine whether there is a commit operation of the transaction corresponding to the third incremental data in the second group;

[0019] If there is a commit operation of the transaction corresponding to the third incremental data in the second group, perform a synchronization operation on the third incremental data.

[0020] Preferably, after determining whether there is a commit operation of the transaction corresponding to the third incremental data in the second group, it further includes:

[0021] If there is no commit operation of the transaction corresponding to the third incremental data in the second group, migrate the third incremental data to the first group.

[0022] Preferably, the method further includes:

[0023] After receiving the commit operation of a transaction, determine whether there is incremental data corresponding to the transaction in the second group;

[0024] If it exists, add the commit operation of the transaction to the second group.

[0025] Preferably, receiving the initialization full amount data and incremental data of the table to be synchronized from the source end includes:

[0026] Query the source end, and migrate the queried initialization full amount data to the target end;

[0027] Capture the operation log of the source end, and synchronously migrate the captured incremental data to the target end.

[0028] Preferably, the method further includes:

[0029] After the target end receives the operation message indicating the end of the synchronization of the initialization full amount data of the table to be synchronized, mark the state of the table to be synchronized as the completion of the synchronization of the initialization full amount data.

[0030] Preferably, during the process that the target end receives the migration data from the source end, the target end extracts the control information in the received message, and obtains the commit operation message of the transaction and the operation message indicating the end of the synchronization of the initialization full amount data of the table to be synchronized according to the control information.

[0031] In a second aspect, a device for parallel synchronization of initialization full amount data and incremental data includes at least one processor, and a memory communicatively connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor, and the instructions are executed by the processor for executing the method for parallel synchronization of initialization full amount data and incremental data as described above.

[0032] In a third aspect, a non-volatile computer storage medium is characterized in that the computer storage medium stores computer-executable instructions, and the computer-executable instructions are executed by one or more processors for completing the method for parallel synchronization of initialization full amount data and incremental data provided by the present invention.

[0033] The present invention provides a method and apparatus for parallel synchronization of initial full data and incremental data. During the process of migrating data from a source end to a target end, the initial full data of a synchronization table and the incremental data of operation logs are sent in parallel, and three different groups are created at the target end. One group is used to cache the initial full data and perform synchronization, and the other two groups respectively cache the incremental data migrated before the synchronization of the initial full data is completed, and the incremental data migrated after the synchronization of the initial full data is completed, and all the incremental data is synchronized after the synchronization of the initial full data is completed, realizing the parallel synchronization of the incremental data and the initial full data. Brief Description of the Drawings

[0034] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings required for the embodiments of the present invention will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention, and those of ordinary skill in the art can obtain other drawings based on these drawings without creative efforts.

[0035] Figure 1 It is a flowchart of a method for parallel synchronization of initial full data and incremental data provided by an embodiment of the present invention;

[0036] Figure 2 It is a flowchart of a method for grouped caching according to the received operation message type in a method for parallel synchronization of initial full data and incremental data provided by an embodiment of the present invention;

[0037] Figure 3 It is a simple schematic diagram of a device for parallel synchronization of initial full data and incremental data provided by an embodiment of the present invention;

[0038] Figure 4 It is a flowchart of another method for parallel synchronization of initial full data and incremental data provided by an embodiment of the present invention;

[0039] Figure 5 It is a schematic diagram of a device for parallel synchronization of initial full data and incremental data provided by an embodiment of the present invention. Detailed Description of the Embodiments

[0040] In order to make the purpose, technical solutions and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0041] In the description of the present invention, the orientation or positional relationship indicated by the terms "inner", "outer", "longitudinal", "transverse", "upper", "lower", "top", "bottom", etc. is based on the orientation or positional relationship shown in the drawings. It is only for the convenience of describing the present invention and does not require the present invention to be constructed and operated in a specific orientation. Therefore, it should not be construed as a limitation to the present invention.

[0042] In addition, the technical features involved in the various embodiments of the present invention described below can be combined with each other as long as they do not conflict with each other.

[0043] Embodiment 1:

[0044] Embodiment 1 of the present invention provides a method for parallel synchronization of initialized full - volume data and incremental data, including:

[0045] The target end creates a first group, a second group, and a third group;

[0046] Receives the initialized full - volume data and incremental data of the table to be synchronized from the source end;

[0047] Caches the initialized full - volume data into the third group for synchronizing the initialized full - volume data of the table to be synchronized;

[0048] Caches the first incremental data into the first group and the second incremental data into the second group, where the first incremental data includes the incremental data generated after the synchronization of the initialized full - volume data is completed, and the second incremental data includes the incremental data generated before the synchronization of the initialized full - volume data of the table to be synchronized is completed;

[0049] Synchronizes the incremental data in the first group and the second group to achieve parallel synchronization of the incremental data and the initialized full - volume data.

[0050] Among them, the grouping at the target end is performed before its own synchronization software is started; when the source end migrates and synchronizes data with the target end, it is divided into two stages. One stage is the stage where the synchronization of the initialized full - volume data of the table is not completed, and the other stage is the stage where the synchronization of the initialized full - volume data of the table is completed;

[0051] The first group is used to cache the incremental data generated after the synchronization of the initialized full - volume data is completed, the second group is used to cache the incremental data when the synchronization of the initialized full - volume data is not completed, and the third group is used to cache at least one batch of initialized full - volume data.

[0052] Among them, the incremental data in the first group, the second group, and the third group are all stored in the order of the time of generation.

[0053] In this embodiment, the initialized full amount of data is divided into at least one batch, and then the initialized full amount of data is migrated to the target end according to the batches. During the process of migrating in batches, the source end will perform operations such as adding, deleting, or modifying the synchronization table. The operation logs generated during this period are called incremental data. The source end needs to send both the initialized full amount of data and the incremental data to the target end. The target end first performs full amount data synchronization based on the initialized full amount of data. After completing the full amount data synchronization, based on the initialized full amount of data, incremental synchronization is performed according to the incremental data.

[0054] In this embodiment, after the full amount of data corresponding to the incremental data is synchronized and after the commit operation of the transaction corresponding to the incremental data is captured, the corresponding incremental data synchronization operation is executed.

[0055] Since the incremental data in the first group is generated after the full amount of data synchronization is completed, after the commit operation of the transaction corresponding to the incremental data is captured, the corresponding incremental data synchronization operation can be executed.

[0056] The incremental data in the second group is generated before the full amount of data synchronization is completed. Before performing the synchronization operation on the incremental data in the second group, it is necessary to determine whether "the full amount of data corresponding to the incremental data has been synchronized" (condition 1), and it is also necessary to determine whether "the transaction corresponding to the incremental data has been committed" (condition 2). If both condition 1 and condition 2 are satisfied, the corresponding incremental data synchronization operation is executed.

[0057] In order to facilitate the target end to determine whether the incremental data meets condition 1 and condition 2, in this embodiment, the first group is also used to store the incremental data migrated from the second group, and the second group is also used to store the commit operations of the transactions, specifically as follows:

[0058] In one application scenario, some incremental data is generated before the full amount of data synchronization is completed, but the commit operation of the transaction corresponding to this incremental data occurs after the full amount of data synchronization is completed. In this case, it is necessary to transfer this incremental data from the second group to the first group.

[0059] Specifically, the first group is also used to cache the incremental data migrated from the second group. When the incremental data is generated when the initialization of the full amount of data is not synchronized, this incremental data will be first stored in the second group. When the commit operation of the transaction corresponding to this incremental data occurs after the initialization of the full amount of data is synchronized, this incremental data is migrated to the first group. The aforementioned migration means deleting the corresponding incremental data from the second group and adding it to the first group.

[0060] That is, the first group is used to cache the incremental data generated after the initialization full - volume data synchronization is completed, and the incremental data generated before the initialization full - volume data synchronization is completed and whose corresponding transaction commit operation occurs after the full - volume data synchronization is completed.

[0061] In another application scenario, some incremental data is generated before the full - volume data synchronization is completed. This incremental data is cached in the second group, but the commit operation of the transaction corresponding to this incremental data occurs before the full - volume data synchronization is completed. In this case, since the full - volume data is not synchronized yet, even if the commit operation of the transaction corresponding to the incremental data is captured, the synchronization operation cannot be executed. At this time, the commit operation of the corresponding transaction needs to be added to the second group so that the target end can execute the corresponding synchronization operation after detecting that the full - volume data synchronization is completed.

[0062] Specifically, when the target end detects the commit operation of a transaction, it determines whether there is incremental data with the same transaction ID in the second group according to the transaction ID. If there is, the commit operation of this transaction is stored in the second group; if not, there is no need to store it.

[0063] That is, the second group is used to cache the incremental data when the initialization full - volume data synchronization is not completed, and the commit operations of transactions with the same transaction ID as the previously cached incremental data. Among them, the aforementioned "previously cached incremental data" is in relation to the second group and does not include the incremental data migrated to the first group.

[0064] In this implementation, during the data migration from the source end to the target end, the initialization full - volume data of the synchronization table and the incremental data of the operation log are sent in parallel, greatly increasing the data migration efficiency. And considering that the synchronization of incremental data requires the use of the initialization full - volume data and needs to wait until all the initialization full - volume data is migrated and synchronized before the synchronization of incremental data can be carried out, three different groups are created at the target end. One group is used to cache and synchronize the initialization full - volume data, and the other two groups respectively cache the incremental data migrated before the initialization full - volume data synchronization is completed and the incremental data migrated after the initialization full - volume data synchronization is completed, and then synchronize all the incremental data after the initialization full - volume data synchronization is completed. Each group processes the synchronization operations at different stages, and the multiple groups are effectively connected and operate in parallel, greatly improving the efficiency of heterogeneous data migration and synchronization.

[0065] Synchronizing the incremental data in the first group and the second group includes:

[0066] For the incremental data in the first group, after the synchronization of the initial full amount of data in the table to be synchronized is completed and after the commit operation of the transaction corresponding to the incremental data is received, a synchronization operation is performed on the corresponding incremental data.

[0067] After the target end receives the operation message indicating the end of the synchronization of the initial full amount of data in the table to be synchronized, it marks the status of the table to be synchronized as the completion of the synchronization of the initial full amount of data.

[0068] In this embodiment, after the target end receives the commit operation of a transaction, it determines whether there is incremental data corresponding to this transaction in the second group;

[0069] If there is, the commit operation of this transaction is added to the second group. As shown in Table 1 below, when the commit operation of Transaction 1 is received and there is incremental data W1(A) and W1(C) corresponding to Transaction 1 in the second group, the commit operation of Transaction 1 is added to the second group;

[0070] If not, the commit operation of this transaction is not added to the second group. As shown in Table 1 below, when the commit operation of Transaction 2 is received in the second group and there is no incremental data corresponding to Transaction 2 (W2(C) has been migrated to the first group), there is no need to add the commit operation of Transaction 2 to the second group.

[0071] Further, the synchronization of the incremental data in the first group and the second group includes:

[0072] For the incremental data in the second group, after the synchronization of the initial full amount of data in the table to be synchronized is completed, it is determined whether there is third incremental data subordinate to the table to be synchronized in the second group;

[0073] The third incremental data is the incremental data corresponding to the synchronization of the initial full amount of data in the corresponding synchronization table in the second group and is subordinate to the second incremental data.

[0074] If there is third incremental data subordinate to the table to be synchronized in the second group, it is determined whether there is a commit operation of the transaction corresponding to the third incremental data in the second group;

[0075] If there is a commit operation of the transaction corresponding to the third incremental data in the second group, a synchronization operation is performed on the third incremental data.

[0076] If there is no commit operation of the transaction corresponding to the third incremental data in the second group, the third incremental data is migrated to the first group.

[0077] Taking the following Table 1 as an example, at time T11, when the full amount of data in Table C is initialized and the synchronization is completed, it is judged whether there is a third incremental data subordinate to Table C in the second group. In Table 1, there are incremental data W1(C) and W2(C) subordinate to Table C, and then it is continued to judge the commit operations of the transactions corresponding to the incremental data W1(C) and W2(C) existing in the second group. If there is a commit operation of Transaction 1 corresponding to the incremental data W1(C) in the second group, then a synchronization operation is performed on the incremental data W1(C); if there is no commit operation of Transaction 2 corresponding to the incremental data W2(C) in the second group, then the incremental data W2(C) is migrated to the first group.

[0078] In this embodiment, after the synchronization of the full amount of data corresponding to the incremental data is completed and after the commit operation of the transaction corresponding to the incremental data is captured, the corresponding synchronization operation of the incremental data is performed.

[0079] Since the incremental data in the first group is generated after the synchronization of the full amount of data is completed, after the commit operation of the transaction corresponding to the incremental data is captured, the corresponding synchronization operation of the incremental data can be performed.

[0080] However, the incremental data in the first group is generated before the synchronization of the full amount of data is completed. Before performing the synchronization operation on the incremental data in the first group, it is necessary to judge whether "the full amount of data corresponding to the incremental data is synchronized" (Condition 1), and it is also necessary to judge whether "the transaction corresponding to the incremental data has been committed" (Condition 2). If both Condition 1 and Condition 2 are satisfied, then the corresponding synchronization operation of the incremental data is performed.

[0081] In order to facilitate the target end to determine whether the incremental data meets Condition 1 and Condition 2, in this embodiment, the first group is also used to store the incremental data migrated from the second group, and the second group is also used to store the commit operations of the transactions, specifically as follows:

[0082] In one application scenario, some incremental data is generated before the synchronization of the full amount of data is completed, but the commit operation of the transaction corresponding to this incremental data occurs after the synchronization of the full amount of data is completed. In this case, it is necessary to transfer this incremental data from the second group to the first group.

[0083] Specifically, the first group is also used to cache the incremental data migrated from the second group. When the incremental data is generated when the initialization of the full amount of data synchronization is not completed, this incremental data will be stored in the second group first. When the commit operation of the transaction corresponding to this incremental data occurs after the initialization of the full amount of data synchronization is completed, then this incremental data is migrated to the first group. The aforementioned migration means deleting the corresponding incremental data from the second group and adding it to the first group.

[0084] That is, the first group is used to cache the incremental data generated after the initialization full - volume data synchronization is completed, as well as the incremental data that is generated before the initialization full - volume data synchronization is completed and the commit operation of the transaction corresponding to the incremental data occurs after the full - volume data synchronization is completed.

[0085] The receiving of the initialization full - volume data and incremental data of the table to be synchronized from the source end includes:

[0086] Query the source end, and migrate the queried initialization full - volume data to the target end;

[0087] Capture the operation logs of the source end, and synchronously migrate the captured incremental data to the target end.

[0088] Among them, since the initialization full - volume data is the data type that originally exists in the source - end database, the source - end database is directly queried, and then the queried initialization full - volume data is sent; while the incremental data is generated by the operation logs during the data migration of the initialization full - volume data, and it needs to be continuously tracked and captured during the data migration of the initialization full - volume data. During this process, the captured incremental data is continuously sent to the target end until all the initialization full - volume data is migrated.

[0089] Embodiment 2:

[0090] Embodiment 2 of the present invention provides a method for parallel synchronization of initialization full - volume data and incremental data;

[0091] As Figure 1 shown, the method includes:

[0092] In step 101, the target end creates a first group, a second group, and a third group.

[0093] Among them, the grouping of the target end is created before its own synchronization software is started; when the source end and the target end perform data migration and synchronization, it is divided into two stages. One stage is the stage where the initialization full - volume data synchronization of the synchronized table is not completed, and the other stage is the stage where the initialization full - volume data synchronization of the synchronized table is completed;

[0094] The first group is used to cache the incremental data generated after the initialization full - volume data synchronization is completed, the second group is used to cache the incremental data when the initialization full - volume data synchronization is not completed, and the third group is used to cache at least one batch of initialization full - volume data.

[0095] Among them, the incremental data in the first group, the second group, and the third group are all stored in the order of the time of generation.

[0096] In this embodiment, after the synchronization of the full amount of data corresponding to the incremental data is completed and after the commit operation of the transaction corresponding to the incremental data is captured, the synchronization operation of the corresponding incremental data is executed.

[0097] Since the incremental data in the first group is generated after the synchronization of the full amount of data is completed, after the commit operation of the transaction corresponding to the incremental data is captured, the synchronization operation of the corresponding incremental data can be executed.

[0098] For the incremental data in the first group that is generated before the synchronization of the full amount of data is completed, before performing the synchronization operation on the incremental data in the first group, it is necessary to determine whether "the full amount of data corresponding to the incremental data has been synchronized" (condition 1), and it is also necessary to determine whether "the transaction corresponding to the incremental data has been committed" (condition 2). If both condition 1 and condition 2 are satisfied, the synchronization operation of the corresponding incremental data is executed.

[0099] In order to facilitate the target end to determine whether the incremental data meets condition 1 and condition 2, in this embodiment, the first group is also used to store the incremental data migrated from the second group, and the second group is also used to store the commit operation of the transaction, specifically as follows:

[0100] In one application scenario, some incremental data is generated before the synchronization of the full amount of data is completed, but the commit operation of the transaction corresponding to this incremental data occurs after the synchronization of the full amount of data is completed. In this case, it is necessary to transfer this incremental data from the second group to the first group.

[0101] Specifically, the first group is also used to cache the incremental data migrated from the second group. When the incremental data is generated when the initialization of the full amount of data synchronization is not completed, this incremental data will be stored in the second group first. When the commit operation of the transaction corresponding to this incremental data occurs after the initialization of the full amount of data synchronization is completed, this incremental data is migrated to the first group. The aforementioned migration means deleting the corresponding incremental data from the second group and adding it to the first group.

[0102] That is, the first group is used to cache the incremental data generated after the initialization of the full amount of data synchronization is completed, and the incremental data that is generated before the initialization of the full amount of data synchronization is completed and the commit operation of the transaction corresponding to the incremental data occurs after the synchronization of the full amount of data is completed.

[0103] In another application scenario, some incremental data is generated before the completion of the full data synchronization. This incremental data is cached in the second group, but the commit operation of the transaction corresponding to this incremental data occurs before the completion of the full data synchronization. In this case, since the full data has not been synchronized, even if the commit operation of the transaction corresponding to the incremental data is captured, the synchronization operation cannot be executed. At this time, the commit operation of the corresponding transaction needs to be added to the second group so that the target end can execute the corresponding synchronization operation after detecting the completion of the full data synchronization.

[0104] Specifically, when the target end detects the commit operation of a transaction, it determines whether there is incremental data with the same transaction ID in the second group according to the transaction ID. If it exists, the commit operation of this transaction is stored in the second group; if it does not exist, there is no need to store it.

[0105] That is, the second group is used to cache the incremental data when the initialization of the full data synchronization is not completed, and the commit operations of the transactions with the same transaction ID as the previously cached incremental data. Among them, the aforementioned "previously cached incremental data" is for the second group and does not include the incremental data migrated to the first group.

[0106] In step 102, the source end starts data synchronization and adds a synchronization table, and migrates the initialization full data of the synchronization table and the incremental data of the operation log in parallel, and then performs step 103 and step 104 simultaneously.

[0107] When the source end migrates data to the target end, it migrates the initialization full data with the synchronization table. During the migration, incremental data based on the operation log is generated. The incremental data also needs to be migrated to the target end and synchronized. In this embodiment, the initialization full data and the incremental data are migrated in parallel and stored in different target end groups.

[0108] In this embodiment, the initialization full data is divided into at least one batch, and then the initialization full data is migrated to the target end according to the batch. During the process of migrating in batches, the source end will perform operations such as adding, deleting, or modifying the synchronization table. The operation log generated during this period is called incremental data. The source end needs to send both the initialization full data and the incremental data to the target end. The target end first performs full data synchronization according to the initialization full data. After completing the full data synchronization, it performs incremental synchronization according to the incremental data based on the initialization full data.

[0109] In step 103, the source end migrates the initialization full data to the target end, and the target end caches the initialization full data in the third group.

[0110] Since incremental data will be generated during the process of migrating and initializing all data, and the synchronization of incremental data must be carried out after the synchronization of all data is completed, the source end needs to send a control message to the target end to notify the target end that the initialization of all data has been transmitted.

[0111] Among them, the control message can be a specific character or other end flags. The control message can be set after the last batch of initialized all data. After receiving the control message, the target end can know that the initialization of all data in the synchronization table has been transmitted.

[0112] Among them, the initialized all data is cached in the third group throughout the process. After receiving the control message indicating that the transmission of the initialized all data is completed, the synchronization of incremental data is then carried out.

[0113] In step 104, the source end migrates the incremental data to the target end; among them, before the synchronization of the initialized all data is completed, the target end caches the received incremental data in the second group;

[0114] The migration of incremental data by the source end is carried out in parallel with the migration of initialized all data. Before the synchronization of initialized all data is completed, the incremental data is cached in the second group of the target end. After the synchronization of initialized all data is completed and after receiving the commit operation of the transaction corresponding to the incremental data, the synchronization of incremental data is then carried out.

[0115] In an actual application scenario, if the commit operation of the transaction corresponding to the incremental data has been received before the synchronization of the initialized all data is completed, in this case, the synchronization of the incremental data is not carried out first and needs to wait until the synchronization of the initialized all data is completed before the synchronization of the incremental data can be carried out.

[0116] Here it should be noted that the reason for waiting until the synchronization of the initialized all data is completed before carrying out the synchronization of incremental data is that the synchronization of incremental data needs to use the initialized all data. It is necessary to wait until the synchronization of the initialized all data is completed so that the target end can allocate and use the initialized all data and use this as a benchmark for the synchronization of incremental data.

[0117] In step 105, after the synchronization of the initialized all data is completed, the target end caches the subsequent received incremental data in the first group and starts to synchronize the incremental data in the first group and the second group to achieve the parallel synchronization of the incremental data and the initialized all data.

[0118] After all the initialized full - volume data is migrated to the second group at the target end and synchronization is completed, it is necessary to inform the target end. At the same time, the target end caches the newly received incremental data in the first group. Since the initialized full - volume data can be used at this time, after receiving the commit operation of the transaction corresponding to the incremental data, the incremental data cached in the second group and the first group is synchronized.

[0119] In this embodiment, during the data migration from the source end to the target end, the initialized full - volume data of the synchronization table and the incremental data of the operation log are sent in parallel, which greatly improves the efficiency of data migration. And considering that the synchronization of incremental data requires the use of initialized full - volume data and needs to wait until all the initialized full - volume data is migrated and synchronized, three different groups are created at the target end. One group is used to cache and synchronize the initialized full - volume data, and the other two groups respectively cache the incremental data migrated before the synchronization of the initialized full - volume data is completed and the incremental data migrated after the synchronization of the initialized full - volume data is completed. And after the synchronization of the initialized full - volume data is completed, all the incremental data is synchronized. Each group processes the synchronization operations at different stages, and the multiple groups are effectively connected and operated in parallel, which greatly improves the efficiency of heterogeneous data migration and synchronization.

[0120] Since the synchronization of incremental data needs to wait for the completion of the synchronization of the initialized full - volume data, after the target end receives the initialized full - volume data of the synchronization table, it is necessary to track and follow up the synchronization status of the initialized full - volume data, and mark its synchronization table before the completion of the initialized full - volume data to inform the target end of the status of the initialized full - volume data at this time.

[0121] In an alternative embodiment, when the source end starts data synchronization and adds a synchronization table, it further includes:

[0122] After the source end starts data synchronization and adds a synchronization table, it notifies the target end. The target end receives the synchronization table from the source end and marks the received synchronization table as the first state before the synchronization of the initialized full - volume data in the synchronization table is completed.

[0123] When the source end starts data migration and synchronizes the initialized full - volume data, it sends a control message to the target end, and when there is still un - migrated initialized full - volume data at the source end subsequently, it continuously sends control messages to the target end to inform the target end that the migration of the initialized full - volume data is still not completed. The target end marks the received synchronization table as the first state. The first state represents that the synchronization of the initialized full - volume data is not completed. Taking the first state as the judgment criterion, only the received incremental data is cached in the second group, and no synchronization operation is performed on the incremental data. In this way, the reliability is higher, but the target end will continuously receive control messages, and the transmission of control messages is too redundant.

[0124] In a preferred embodiment, the source end needs to send a control message to the target end to notify the target end that the transmission of the initial full amount of data has been completed. Among them, the control message can be a specific character or other end flag. The control message can be set after the last batch of initial full amount of data. After receiving the control message, the target end can know that the transmission of the initial full amount of data of the synchronization table has been completed. In this way, only one control message needs to be sent after the full amount of data is transmitted, and the control message can be continuously sent. Due to the different attribute characteristics of the full amount of data and the incremental data, the incremental data is generated from the operation logs during the migration of the full amount of data. Therefore, the methods for querying and migrating the two different data types are also different.

[0125] Parallelly migrate the initial full amount of data of the synchronization table and the incremental data of the operation logs, specifically including:

[0126] Query the source end database, and migrate the queried initial full amount of data to the target end;

[0127] At the same time, capture and parse the operation logs of the source end database in real time, and synchronously migrate the captured incremental data to the target end.

[0128] Among them, since the initial full amount of data is a data type that originally existed in the source end database, directly query the source end database and then send the queried initial full amount of data; while the incremental data is generated from the operation logs during the data migration of the initial full amount of data, and it is necessary to continuously track and capture during the data migration process of the initial full amount of data. During this process, continuously send the captured incremental data to the target end until all the initial full amount of data is migrated.

[0129] When all the initial full amount of data is migrated to the target end and the synchronization is completed, it is necessary to change the marked state of the synchronization table of the initial full amount of data, thereby changing the judgment result of whether the synchronization of the initial full amount of data is completed, thereby changing the cache location of the incremental data and starting to synchronize the incremental data.

[0130] The target end caches the initial full amount of data in the third group and performs synchronization, and further includes:

[0131] When the source end migrates all the initial full amount of data to the target end and completes the synchronization of the initial full amount of data, the target end marks all the received synchronization tables as the second state.

[0132] When all the initial full - volume data at the source end has been migrated to the target end and the synchronization is completed, the source end sends a control message indicating the completion of the initial full - volume data synchronization to the target end. After receiving the notification, the target end changes the synchronization table flag to the second state, where the second state represents that the initial full - volume data has been synchronized. Taking the second state as the judgment criterion, first, it queries whether all the incremental data at the source end has been migrated. If so, it immediately synchronizes the incremental data for which the transaction to which the synchronization table operation in the second group belongs has been committed, transfers the incremental data for which the transaction to which the synchronization table belongs has not been committed to the first group, and caches the subsequent incremental data migrated from the source end to the target end in the first group. After receiving the commit operation for the transaction corresponding to the incremental data, it immediately performs the synchronization. If not, it continues to cache the received incremental data in the second group.

[0133] As Figure 2 shown, when migrating heterogeneous data, the source end needs to continuously send control messages to the target end. According to the data messages, the target end changes the status of the synchronization table and operation log of the received migration data, changes the cache location of the received migration data, and determines the start time for synchronizing different types of migration data. The method flow is as follows:

[0134] In step 201, during the process of the source end migrating data to the target end, the target end receives the message migrated from the source end, extracts the control information of the message, and simultaneously judges the operation message type of the control information.

[0135] The operation message type includes one or more of a data synchronization message, a COMMIT operation message, and an initial full - volume data synchronization end operation message.

[0136] In step 202, when the operation type is a data synchronization message, it further includes:

[0137] When the operation message is an incremental data synchronization, it judges the marked status of the synchronization table received by the target end at this time. When the synchronization table is marked as the first state, the target end caches the received incremental data in the second group. When the synchronization table is marked as the second state, the target end caches the received incremental data in the first group;

[0138] When the data synchronization message is an initial full - volume data synchronization, the target end caches the initial full - volume data of the received synchronization table in the third group.

[0139] Among them, the data synchronization message includes initializing full - volume data synchronization and incremental data synchronization, which is used to determine the data type of the received migration data, and the received migration data is cached into different groups by judging the data type; it should be noted that when the data synchronization message is incremental data synchronization, it is necessary to judge whether the initialization of full - volume data synchronization is completed according to the marked state of the current synchronization table. When the synchronization table is marked as the first state, the target end caches the received incremental data into the second group; when the synchronization table is marked as the second state, the target end caches the received incremental data into the first group.

[0140] In step 203, when the operation type is a COMMIT operation message, it further includes:

[0141] The target end finds all the operations corresponding to the transaction ID according to the transaction ID of the COMMIT operation message, and the group includes one or more of the first group, the second group, and the third group;

[0142] When the transaction ID corresponds to the first group and / or the third group, all the data corresponding to the transaction ID cached in the corresponding group is synchronized;

[0143] When the transaction ID corresponds to the second group, the incremental data of the operation log corresponding to this transaction in the second group is continuously cached and waits to be synchronized, and at the same time, the transaction ID is marked as committed.

[0144] The COMMIT operation message is used to judge which group the current cached data is in, and decides whether to synchronize the cached data according to the group; among them, the third group contains the initialization full - volume data and can be directly synchronized. The first group contains the incremental data cached after the initialization full - volume data synchronization is completed. At this time, the target end can directly call the initialization full - volume data, so the incremental data in the first group can also be directly synchronized. Therefore, when the transaction ID of the COMMIT operation message corresponds to the first group and / or the third group, the corresponding cached data can be directly synchronized; while the second group is used to cache the incremental data before the initialization full - volume data synchronization is completed. At this time, the target end cannot directly call the initialization full - volume data, so the data cached in the second group cannot be synchronized, and the subsequent received incremental data is still cached in the second group until the marked state of the synchronization table is changed to the second state.

[0145] In step 204, when the operation type is an initialization full - volume data synchronization end operation message, it further includes:

[0146] Synchronize the incremental data of the transactions to which the synchronization tables cached in the second group belong and which have been committed, and mark the corresponding synchronization tables as the second state;

[0147] Transfer the incremental data of the transactions to which the synchronization tables cached in the second group belong and which have not been committed to the first group, and perform synchronization immediately after receiving the commit operation message subsequently; at the same time, cache the incremental data migrated from the source end to the target end subsequently into the first group and perform synchronization immediately.

[0148] When receiving the initialization full - volume data synchronization end operation message, immediately deliver the incremental data that has been cached in the second group and the transactions to which it belongs and have been committed in the synchronization table to an idle execution thread for storage execution to complete the synchronization. At this time, the initialization full - volume data has been fully synchronized, and the associated full - volume data required for the incremental data generated during this period has been correctly received and can be called, so the synchronization operation of the incremental data in the second group can be immediately executed; at the same time, the target end changes the mark of the received synchronization table to the second state; secondly, determine whether there is still un - migrated and un - committed incremental data. If so, transfer the remaining incremental data to the first group for caching, and at the same time, immediately execute the synchronization operation for the incremental data cached in the first group subsequently; and at this time, the state of the synchronization table has been updated, and the synchronization table should be corresponding to the first group, so transfer the synchronization table to the first group for storage.

[0149] Embodiment 3:

[0150] Embodiment 3 of the present invention provides a method for parallel synchronization of initialization full - volume data and incremental data. Based on the above - mentioned embodiments, the execution process of the method for parallel synchronization of initialization full - volume data and incremental data is presented in a more specific scenario.

[0151] As Figure 3 shown, the source end includes an initialization full - volume data extraction module and an incremental data capture module. The initialization full - volume data extraction module is used to extract the initialization full - volume data and perform migration and synchronization; the incremental data capture module is used to capture incremental data and perform migration and synchronization; the target end includes a message receiving and classification module, a first group, a second group, a third group, and multiple synchronization and storage execution threads.

[0152] In this embodiment, Table A, Table B, and Table C are synchronization tables for addition.

[0153] The synchronization operation sequence received by the target end in this embodiment is as follows:

[0154] L1(A)L1(B)L1(C)W1(A)W1(C)Comimt(1)L2(A)L2(B)W2(C)L2(C)X(C)L3(B)L3(A)X(B)L4(A)W2(B)Commit(2)L5(A)W3(A)W3(C)Commit(3)L6(A)X(A)W4(B)

[0155] Among them, the Li(Z) represents the i-th batch of initialization full-volume synchronization operations for table Z. This operation includes the commit operation for initialization synchronization, indicating that this batch of operations can be executed immediately; the i represents any batch of initialization full-volume synchronization operations with ascending serial numbers in table Z; the Z represents any one of A, B, and C, representing any one of the synchronization tables of table A, table B, and table C.

[0156] The X(Z) represents the end of the initialization full-volume data synchronization for table Z; the Z represents any one of A, B, and C, representing any one of the synchronization tables of table A, table B, and table C.

[0157] The Wj(Z) represents the incremental operation of transaction j on table Z; the j represents the ID number of the transaction; the Z represents any one of A, B, and C, representing any one of the synchronization tables of table A, table B, and table C.

[0158] The Commit(j) represents the commit operation of transaction j, and the j represents the ID number of the transaction.

[0159] The content of the method is described using a timing operation table, and the transaction synchronization model includes the following 4 types:

[0160] (1) Transactions that start during the synchronization operation of the initialization full-volume data in the synchronization table and are committed before the completion of the initialization full-volume data synchronization operation; for example, transaction 1 in this embodiment, the operations are: W1(A)W1(C)Comimt(1).

[0161] (2) Transactions that start during the synchronization operation of the initialization full-volume data in the synchronization table and are committed after the completion of the initialization full-volume data synchronization operation; for example, transaction 2 in this embodiment, the operations are: W2(C)W2(B)Commit(2).

[0162] (3) Transactions that start after the completion of the synchronization operation of the initialization full-volume data in the synchronization table and the initialization full-volume data synchronization of the synchronization table is completed; for example, transaction 4 in this embodiment, the operations are: W4(B).

[0163] (4) Other transaction synchronization models other than the above three transaction synchronization models; for example, Transaction 3 in this embodiment, which operates on Table A and Table C. At the T22 moment when Transaction 3 completes the commit, the initialization full amount data of Table A has not been synchronized, and the initialization full amount data of Table C has been synchronized. The operations are: W3(A)W3(C)Commit(3).

[0164] The timing operation table is shown in Table 1 as follows:

[0165]

[0166]

[0167] Table 1

[0168] As Figure 4 shown, the method for parallel synchronization of initialization full amount data and incremental data in this embodiment is described according to Table 1 above, and the method flow is as follows:

[0169] In step 401, the target end creates a first group, a second group, and a third group;

[0170] Among them, the first group is used to cache the incremental data of the synchronization tables for which the initialization full amount data synchronization is completed; the second group is used to cache the incremental data of the synchronization tables for which the initialization full amount data has not been synchronized yet; the third group is used to cache the initialization full amount data of all synchronization tables.

[0171] In step 402, the source end synchronization service adds synchronization Table A, synchronization Table B, and synchronization Table C, and then simultaneously starts the data migration of the initialization full amount data and the data migration of the incremental data.

[0172] In step 403, the target end synchronization service receives the synchronization operation sequence and processes it.

[0173] The processing is as follows:

[0174] At the T1 - T3 moment, the target end receives the initialization full amount data synchronization operations of synchronization Table A, synchronization Table B, and synchronization Table C, and adds them to the third group. At the same time, in this embodiment, the initialization full amount data synchronization operation is set as a batch commit operation, so L1(A), L1(B), and L1(C) can be immediately executed for storage and synchronization.

[0175] At the T4 - T5 moment, Transaction 1 operates on synchronization Table A and synchronization Table C. After the target end receives the operation sequence W1(A)W1(C), since the initialization full amount data of synchronization Table A and synchronization Table C has not been synchronized yet at this time, it is added to the second group.

[0176] At time T6, the target end receives the commit operation of transaction 1, locates the operations W1(A) and W1(C) of transaction 1 in the second group. Since the initialization of the full amount of data in synchronization tables A and C has not been completed yet, the operations W1(A) and W1(C) cannot be immediately synchronized and need to be cached in the second group continuously. Meanwhile, mark that the transaction to which this operation belongs has been committed.

[0177] At times T7 - T10, the target end receives the operation sequence L2(A)L2(B)W2(C)L2(C), adds the operations L2(A)L2(B)L2(C) to the third group and immediately performs the synchronization operation; adds the operation W2(C) to the second group and waits to perform the synchronization operation.

[0178] At time T11, the target end receives the operation X(C) indicating the end of the initialization of the full amount of data in synchronization table C. It searches for the committed transaction operation W1(C) on synchronization table C in the second group and delivers it to the execution thread to immediately perform the synchronization; then searches for the uncommitted transaction operation W2(C) on synchronization table C and transfers it to the first group.

[0179] At time T15, the target end receives the operation X(B) indicating the end of the initialization of the full amount of data in synchronization table B.

[0180] At time T17, the target end receives the operation W2(B) of transaction 2 on synchronization table B. Since the initialization of the full amount of data in synchronization table B has been completed, it is added to the first group.

[0181] At time T18, the target end receives the commit operation of transaction 2, locates the operations W2(B)W2(C) of transaction 2 in the first group and immediately performs the synchronization operation.

[0182] At time T20, the target end receives the operation W3(A) of transaction 3. Since the initialization of the full amount of data in synchronization table A has not been completed yet, it is cached in the second group.

[0183] At time T21, the target end receives the operation W3(C) of transaction 3. Since the initialization of the full amount of data in synchronization table C has been completed, it is cached in the first group.

[0184] At time T22, the target end receives the commit operation of transaction 3. The operation W3(A) needs to be cached in the second group continuously because the initialization of the full amount of data in synchronization table A has not been completed yet, and mark that transaction 3 has been committed; the operation W3(C) can be immediately synchronized because the initialization of the full amount of data in synchronization table C has been completed.

[0185] At time T24, the target end receives the operation of completing the initialization full-volume data synchronization of synchronization table A, finds the committed transaction operations W1(A) and W3(A) on synchronization table A in the second group, and immediately performs the synchronization operation on them.

[0186] At time T25, the target end receives the operation W4(B) of transaction 4. Since the initialization full-volume data synchronization of synchronization table B is completed, it is added to the first group and can be immediately executed after receiving the commit operation of transaction 4.

[0187] Embodiment 3:

[0188] As Figure 5 shown, it is a schematic diagram of the device for parallel synchronization of initialization full-volume data and incremental data according to an embodiment of the present invention. The device for parallel synchronization of initialization full-volume data and incremental data in this embodiment includes one or more processors 51 and a memory 52. Among them, Figure 5 one processor 51 is taken as an example.

[0189] The processor 51 and the memory 52 can be connected through a bus or other means, Figure 5 and taking the connection through a bus as an example.

[0190] The memory 52, as a non-volatile computer-readable storage medium, can be used to store non-volatile software programs and non-volatile computer-executable programs, such as the method for parallel synchronization of initialization full-volume data and incremental data in Embodiment 1. The processor 51 executes the method for parallel synchronization of initialization full-volume data and incremental data by running the non-volatile software programs and instructions stored in the memory 52.

[0191] The memory 52 may include a high-speed random access memory, and may also include non-volatile memory, such as at least one magnetic disk storage device, a flash memory device, or other non-volatile solid-state storage devices. In some embodiments, the memory 52 may optionally include a memory remotely set relative to the processor 51, and these remote memories can be connected to the processor 51 through a network. Examples of the above network include but are not limited to the Internet, an enterprise intranet, a local area network, a mobile communication network, and combinations thereof.

[0192] The program instructions / modules are stored in the memory 52, and when executed by the one or more processors 51, execute the method for parallel synchronization of initialization full-volume data and incremental data in Embodiment 1 and Embodiment 2 above. For example, execute each step described above Figures 1 - 5 as shown.

[0193] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, and improvements made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.

Claims

1. A method for parallel synchronization of initial full - volume data and incremental data, characterized in that, it includes: The target end creates a first group, a second group and a third group; Receives the initial full - volume data and incremental data of the table to be synchronized from the source end; Caches the initial full - volume data into the third group for synchronizing the initial full - volume data of the table to be synchronized; Caches the first incremental data into the first group and the second incremental data into the second group, where the first incremental data includes the incremental data generated after the synchronization of the initial full - volume data is completed, and the second incremental data includes the incremental data generated before the synchronization of the initial full - volume data of the table to be synchronized is completed; After the synchronization of the initial full - volume data of the table to be synchronized is completed, synchronizes the incremental data in the first group and the second group to achieve parallel synchronization of the incremental data and the initial full - volume data; For the incremental data in the second group, after the synchronization of the initial full - volume data of the table to be synchronized is completed, determines whether there is third incremental data subordinate to the table to be synchronized in the second group; If there is third incremental data subordinate to the table to be synchronized in the second group, determines whether there is a commit operation of the transaction corresponding to the third incremental data in the second group; If there is a commit operation of the transaction corresponding to the third incremental data in the second group, performs a synchronization operation on the third incremental data.

2. The method for parallel synchronization of initial full - volume data and incremental data according to claim 1, characterized in that, The synchronization of the incremental data in the first group and the second group includes: For the incremental data in the first group, after the synchronization of the initial full - volume data of the table to be synchronized is completed and after receiving the commit operation of the transaction corresponding to the incremental data, performs a synchronization operation on the corresponding incremental data.

3. The method for parallel synchronization of initial full - volume data and incremental data according to claim 1, characterized in that, After determining whether there is a commit operation of the transaction corresponding to the third incremental data in the second group, it further includes: If there is no commit operation of the transaction corresponding to the third incremental data in the second group, migrates the third incremental data to the first group.

4. The method for parallel synchronization of initial full - volume data and incremental data according to claim 1, characterized in that, The method further includes: After receiving the commit operation of the transaction, determines whether there is incremental data corresponding to the transaction in the second group; If it exists, adds the commit operation of the transaction to the second group.

5. The method for parallel synchronization of initial full - volume data and incremental data according to any one of claims 1 to 4, characterized in that, Receiving the initial full - volume data and incremental data of the table to be synchronized from the source end includes: Queries the source end and migrates the queried initial full - volume data to the target end; Captures the operation log of the source end and synchronously migrates the captured incremental data to the target end.

6. The method for parallel synchronization of initialized full - volume data and incremental data according to any one of claims 1 to 4, characterized in that, the method further includes: after the target end receives the operation message indicating the end of the initialization full - volume data synchronization of the table to be synchronized, it marks the state of the table to be synchronized as the completion of the initialization full - volume data synchronization.

7. The method for parallel synchronization of initialized full - volume data and incremental data according to claim 6, characterized in that, during the process that the target end receives the migration data from the source end, the target end extracts the control information in the received message, and obtains the commit operation message of the transaction and the operation message indicating the end of the initialization full - volume data synchronization of the table to be synchronized according to the control information.

8. An apparatus for parallel synchronization of initialized full - volume data and incremental data, characterized in that, it includes at least one processor and a memory communicatively connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor, and the instructions, when executed by the processor, are used to execute the method for parallel synchronization of initialized full - volume data and incremental data according to any one of claims 1 - 7.

9. A non - volatile computer storage medium, characterized in that, the computer storage medium stores computer - executable instructions, and the computer - executable instructions, when executed by one or more processors, are used to execute the method for parallel synchronization of initialized full - volume data and incremental data according to any one of claims 1 - 7.

Citation Information

Patent Citations

  • Data migration method and device, storage medium and electronic equipment

    CN112286905A

  • Parallel high-performance increment synchronization method

    CN113918657A