Data synchronization method, electronic device and program product

By managing different versions of source data and utilizing data snapshot identifiers, the problem of low data synchronization efficiency is solved, achieving efficient and stable data synchronization and backup.

CN120910155APending Publication Date: 2025-11-07KE COM (BEIJING) TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511010734.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-22
Publication Date
2025-11-07

AI Technical Summary

Technical Problem

In data analysis scenarios, when synchronizing data from a data warehouse to multiple target devices, existing technologies require multiple data computations, resulting in low synchronization efficiency, high resource consumption, and poor offline scheduling timeliness.

Method used

By acquiring and storing different versions of source data, including initial full data and incremental data, and using data snapshots to manage synchronization progress, synchronization data is directly obtained from the stored source data and synchronized to the target end, avoiding redundant calculation logic.

Benefits of technology

It improves the efficiency of data synchronization to multiple target devices, enhances the stability and timeliness of data backup, reduces resource consumption, and lowers data latency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120910155A_ABST
    Figure CN120910155A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses a data synchronization method, electronic equipment and a program product, aiming at any data synchronization task in at least one data synchronization task, registration information corresponding to the data synchronization task is acquired, and the registration information comprises data source end information and target end information corresponding to the data synchronization task; on the basis of the data source end information, source data of different versions are obtained from the data source end and stored, and the source data of different versions comprise initially obtained full data and incremental data based on the full data; acquiring synchronous data of an unsynchronized version corresponding to the data synchronization task from the stored source data; and the synchronization data is synchronized to the target end based on the target end information, so that the synchronization efficiency of data synchronization can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present disclosure relates to data management technology, and in particular, to a data synchronization method, an electronic device, and a program product. BACKGROUND

[0002] In a data analysis scenario, data is aggregated in a data warehouse through various links, and the data in the data warehouse is synchronized to an engine or a cluster for data analysis.

[0003] Generally, in order to support various business scenarios and data backup, the data in the data warehouse needs to be synchronized to multiple target ends, such as multiple engines or multiple clusters. When synchronizing the data in the data warehouse to each target end respectively, multiple data calculation logics need to be executed, and the synchronization efficiency is low. SUMMARY

[0004] Embodiments of the present disclosure provide a data synchronization method, an electronic device, and a program product to solve the above technical problems to some extent.

[0005] In one aspect of the embodiments of the present disclosure, a data synchronization method is provided, and the method comprises:

[0006] For any data synchronization task in at least one data synchronization task, registration information corresponding to the data synchronization task is obtained, and the registration information includes data source end information and target end information corresponding to the data synchronization task;

[0007] Based on the data source end information, different versions of source data are obtained and stored from the data source end, wherein the different versions of source data include full data initially obtained and incremental data based on the full data;

[0008] From the stored source data, synchronization data of an unsynchronized version corresponding to the data synchronization task is obtained;

[0009] Based on the target end information, the synchronization data is synchronized to the target end.

[0010] In one exemplary embodiment, after the different versions of source data are obtained and stored from the data source end, the method further comprises:

[0011] A data snapshot identifier of each obtained source data is recorded, and the data snapshot identifier is used to indicate version information of the source data;

[0012] The synchronization data of the unsynchronized version corresponding to the data synchronization task is obtained from the stored source data, comprising:

[0013] In the process of executing the data synchronization task, synchronization progress information of the data synchronization task is acquired, the synchronization progress information including a latest synchronization snapshot identifier, the latest synchronization snapshot identifier being used to indicate version information of source data that is synchronized by the data synchronization task last time;

[0014] Based on the latest synchronization snapshot identifier, source data corresponding to a data snapshot identifier recorded after the latest synchronization snapshot identifier is acquired from the stored source data, to obtain the unsynchronized version of the synchronization data.

[0015] In an exemplary embodiment, the acquiring and storing of different versions of source data from the data source end includes:

[0016] In response to the absence of a data snapshot identifier of the recorded source data, full-amount data corresponding to the current version of the source data is acquired from the data source end;

[0017] In response to the presence of a data snapshot identifier of the recorded source data, incremental data between a latest data snapshot identifier corresponding to the latest source data in the data source end and a last recorded data snapshot identifier is acquired.

[0018] In an exemplary embodiment, the acquiring of source data corresponding to a data snapshot identifier recorded after the latest synchronization snapshot identifier from the stored source data based on the latest synchronization snapshot identifier includes:

[0019] In response to the latest synchronization snapshot identifier being an initial state value, full-amount data corresponding to a last recorded full-amount data identifier is acquired from the stored source data, the full-amount data identifier being a data snapshot identifier corresponding to the full-amount data;

[0020] In response to the latest synchronization snapshot identifier being a recorded identifier value, incremental data corresponding to an incremental data identifier recorded after the latest synchronization snapshot identifier is acquired from the stored source data, the incremental data identifier being a data snapshot identifier corresponding to the incremental data.

[0021] In an exemplary embodiment, after the acquiring and storing of different versions of source data from the data source end, the method further includes:

[0022] At a preset target time as a period, merged full-amount data is obtained by merging the historically stored full-amount data and incremental data, and a data snapshot identifier corresponding to the merged full-amount data is recorded;

[0023] The historically stored full-amount data, incremental data, and corresponding data snapshot identifiers are deleted.

[0024] In an exemplary embodiment, the method further includes:

[0025] In response to the synchronization data synchronization success, the synchronization progress information is updated based on the latest data identifier corresponding to the synchronization data.

[0026] In one example embodiment, the target end data is stored according to synchronization time corresponding to the synchronization data, and the method further comprises:

[0027] For any time partition to be detected, the target end data amount corresponding to the time partition is compared with the data source end data amount that should be synchronized in the time partition, to obtain a data comparison result.

[0028] In response to the data comparison result indicating that the data amounts are inconsistent, the target end data corresponding to the time partition is data-completed based on the data source end data corresponding to the time partition in the data source end.

[0029] In one example embodiment, the data-completing of the target end data corresponding to the time partition based on the data source end data corresponding to the time partition in the data source end comprises:

[0030] Generating replacement data based on the data source end data corresponding to the time partition in the data source end;

[0031] Replacing the target end data corresponding to the time partition with the replacement data.

[0032] Another aspect of the embodiment provides a data synchronization device, comprising:

[0033] An information acquisition module, configured to acquire, for any data synchronization task in at least one data synchronization task, registration information corresponding to the data synchronization task, the registration information including data source end information and target end information corresponding to the data synchronization task;

[0034] A data storage module, configured to acquire and store different versions of source data from a data source end based on the data source end information, wherein the different versions of source data include full data initially acquired and incremental data based on the full data;

[0035] A data acquisition module, configured to acquire, from the stored source data, synchronization data of an unsynchronized version corresponding to the data synchronization task;

[0036] A data synchronization module, configured to synchronize the synchronization data to the target end based on the target end information.

[0037] Another aspect of the embodiment provides an electronic device, comprising:

[0038] A memory, configured to store a computer program;

[0039] A processor is configured to execute a computer program stored in the memory, and the computer program, when executed, implements the data synchronization method of any of the above embodiments.

[0040] In another aspect of the embodiment, a computer readable storage medium is provided, and the computer readable storage medium stores a computer program. The computer program, when executed by a processor, implements the data synchronization method of any of the above embodiments.

[0041] In another aspect of the embodiment, a computer program product is provided, and the computer program product includes a computer program. The computer program, when executed by a processor, implements the data synchronization method of any of the above embodiments.

[0042] In the embodiment of the present disclosure, the source data to be synchronized corresponding to any data synchronization task can be managed, the source data of different versions can be obtained and stored from the source end based on the data source end information included in the registration information of the data synchronization task, and then in the data synchronization process, the synchronization data of the unsynchronized version corresponding to the data synchronization task can be obtained from the stored source data, so as to synchronize the synchronization data to the corresponding target end. By making the full data and the incremental data for the source data to be synchronized, the full data and the incremental data can be directly used as the original data to be directly imported into the target end when the data to be synchronized is needed, that is, the required synchronization data can be directly obtained from the stored source data of different versions, without the need to repeatedly execute the data calculation logic, which helps to improve the synchronization efficiency of synchronizing the source data from the source end to multiple target ends, helps to improve the data backup efficiency, and improves the stability of the business. In addition, the source data of different versions includes the full data and the incremental data initially obtained, and by storing the source data of different versions, the source data of different versions can be respectively obtained for data synchronization in the data synchronization process, which helps to quickly synchronize the updated source data to the target end and improves the timeliness of data synchronization.

[0043] The technical solutions of the present disclosure will be further described in detail below with reference to the accompanying drawings and embodiments. BRIEF DESCRIPTION OF DRAWINGS

[0044] The accompanying drawings, which form a part of the specification, illustrate embodiments of the present disclosure and, together with the description, serve to explain the principles of the present disclosure.

[0045] The present disclosure can be more clearly understood and appreciated from the following detailed description, taken in conjunction with the accompanying drawings, in which:

[0046] Figure 1 A flowchart of a data synchronization method provided for an exemplary embodiment of the present disclosure is shown in FIG. 1.

[0047] Figure 2 A flowchart of obtaining synchronization data provided for an exemplary embodiment of the present disclosure is shown in FIG. 2.

[0048] Figure 3 State transition diagram of synchronization state of data synchronization task provided for an exemplary embodiment of the present disclosure;

[0049] Figure 4 Flowchart of source data acquisition process provided for an exemplary embodiment of the present disclosure;

[0050] Figure 5 Result diagram of generating full data and incremental data provided for an exemplary embodiment of the present disclosure;

[0051] Figure 6 Diagram of full data and incremental data corresponding to different periods provided for an exemplary embodiment of the present disclosure;

[0052] Figure 7 Structural diagram of data synchronization system provided for an exemplary embodiment of the present disclosure;

[0053] Figure 8 Flowchart of data synchronization process provided for an exemplary embodiment of the present disclosure;

[0054] Figure 9 Flowchart of data detection process provided for an exemplary embodiment of the present disclosure;

[0055] Figure 10 Structural diagram of data synchronization device provided for an exemplary embodiment of the present disclosure;

[0056] Figure 11 Structural diagram of electronic device provided for an exemplary embodiment of the present disclosure. DETAILED DESCRIPTION

[0057] Various exemplary embodiments of the present disclosure will now be described in detail below with reference to the accompanying drawings. Note that the relative arrangement, numerical expressions, and numerical values of components and steps set forth in these embodiments do not limit the scope of the present disclosure unless otherwise specifically stated.

[0058] Those skilled in the art can understand that the terms “first”, “second”, and the like in the embodiments of the present disclosure are only used to distinguish different steps, devices, or modules, and do not represent any specific technical meaning, nor do they represent the inevitable logical order between them.

[0059] It should also be understood that in the embodiments of the present disclosure, “a plurality of” can refer to two or more, and “at least one” can refer to one, two, or more.

[0060] It should also be understood that for any component, data, or structure mentioned in the embodiments of the present disclosure, unless specifically limited or given a contrary implication by the context or prior art, it can generally be understood as one or more.

[0061] In addition, the term "and / or" in the present disclosure is merely a description of the association relationship of the associated objects, which means that there can be three relationships, for example, A and / or B can represent the following three cases: A exists alone, A and B exist together, and B exists alone. In addition, the character " / " in the present disclosure generally represents an "or" relationship between the front and rear associated objects.

[0062] It should also be understood that the description of the various embodiments of the present disclosure focuses on the differences between the various embodiments, and the same or similar parts can be referred to each other, and for the sake of brevity, will not be repeated.

[0063] At the same time, it should be understood that, for the convenience of description, the size of each part shown in the drawings is not drawn according to the actual proportional relationship.

[0064] The following description of at least one example embodiment is merely illustrative in nature and does not in any way limit the disclosure and its application or uses.

[0065] Techniques, methods, and devices known to those of ordinary skill in the relevant art can not be discussed in detail, but should be considered part of the specification where appropriate.

[0066] It should be noted that similar reference numbers and letters refer to similar items in the following drawings, so once an item is defined in one drawing, it does not need to be further discussed in subsequent drawings.

[0067] Embodiments of the present disclosure can be applied to terminal devices, computer systems, servers and other electronic devices, which can operate with many other general or special computing system environments or configurations. Examples of well-known terminal devices, computing systems, environments, and / or configurations suitable for use with terminal devices, computer systems, servers and other electronic devices include, but are not limited to, personal computer systems, server computer systems, thin clients, thick clients, handheld or laptop devices, microprocessor-based systems, set-top boxes, programmable consumer electronics, network personal computers, small computer systems, mainframe computer systems, and distributed cloud computing technology environments including any of the above systems.

[0068] Electronic devices such as terminal devices, computer systems, and servers can be described in the general context of computer system executable instructions (such as program modules) executed by a computer system. Typically, program modules can include routines, programs, object programs, components, logic, data structures, etc., which perform specific tasks or implement specific abstract data types. Computer systems / servers can be implemented in distributed cloud computing environments, where tasks are executed by remote processing devices linked through communication networks. In distributed cloud computing environments, program modules can reside on local or remote computing system storage media, including storage devices.

[0069] In related technologies, when synchronizing data from a data warehouse (such as Hive) to multiple engines or clusters, independent synchronization tasks are created to execute corresponding data computation logic, thereby synchronizing the data from the data warehouse to each engine or cluster separately. This process requires repeated calculations, and if the amount of data to be synchronized is large, it takes a considerable amount of time, resulting in low execution efficiency and high consumption of data computing resources. Furthermore, data synchronization often involves data scheduling in offline scenarios, and offline scheduling tasks based on data warehouses are slow and have poor timeliness.

[0070] In this embodiment of the disclosure, by managing the source data, different versions of source data are periodically acquired and stored, so that the stored different versions of source data can be directly imported into the target end that needs to be synchronized. This avoids the repetitive task calculation process when synchronizing to multiple target ends and improves the data synchronization efficiency.

[0071] like Figure 1 The diagram illustrates a flowchart of a data synchronization method provided by an exemplary embodiment of this disclosure. This method can be used in the aforementioned electronic device and includes steps 101-104:

[0072] Step 101: For any data synchronization task in at least one data synchronization task, obtain the registration information corresponding to the data synchronization task.

[0073] The registration information includes the data source and target information for the data synchronization task.

[0074] A data synchronization task is a synchronization task that synchronizes data from the data source to the target. You can create a corresponding data synchronization task according to the synchronization requirements.

[0075] When creating a data synchronization task, you can register the relevant information for that task, including the data source information and the target information. The data source information indicates the data source corresponding to the source data to be synchronized. Optionally, the data source information includes the storage location of the source data on the data source side; for example, the data source information includes the database name and table name corresponding to the source data.

[0076] The target end information is used to indicate a target end to which the source data is to be synchronized. Optionally, the target end information includes type information of the target end and identification information of the target end, and the identification information of the target end is used to uniquely indicate the target end. Illustratively, when the source data is to be synchronized to a target cluster, the target end information includes a type of the target cluster and a cluster identification (ID) of the target cluster.

[0077] Different data synchronization tasks can be created according to different synchronization requirements, wherein different data synchronization tasks correspond to different registration information. Optionally, one data synchronization task can be to synchronize data from one data source end to one target end. Different data synchronization tasks correspond to different data source ends and / or target ends. For example, when the same source data is to be synchronized to different target ends, different data synchronization tasks can be registered.

[0078] When the registration signal corresponding to any data synchronization task is received, the corresponding registration information can be acquired, so as to perform data synchronization based on the registration information.

[0079] In step 102, different versions of source data are acquired and stored from the data source end based on the data source end information.

[0080] The different versions of source data include full data initially acquired and incremental data based on the full data.

[0081] When the data source end information corresponding to the data synchronization task is acquired, the source data to be synchronized can be acquired from the data source end, and the acquired source data can be stored for data synchronization.

[0082] In one possible implementation, when the source data of the data source end is initially acquired, the full data corresponding to the current data source end needs to be acquired first, wherein the full data includes all the data currently stored in the data source end.

[0083] Since the data of the data source end is updated at irregular intervals, in the embodiments of the present disclosure, the latest source data can be acquired and stored from the data source end according to the data source end information. The latest source data acquired at regular intervals is incremental data based on the full data. When the source data of the data source end is updated for multiple versions, the incremental data acquired includes incremental data of different versions, so as to synchronize the source data of the latest version to the target end. Optionally, a first time interval for acquiring data can be preset, and the incremental data is acquired from the data source end at a preset first time interval as a period. Illustratively, the preset time interval can be 10 minutes.

[0084] In one possible implementation, the different versions of source data can be arranged and stored in order of different versions of update time, wherein the source data of an old version is arranged more forwardly than the source data of a new version.

[0085] Step 103, obtaining the synchronization data of the unsynchronized version corresponding to the data synchronization task from the stored source data.

[0086] In the process of executing a data synchronization task, first, the synchronization data of the unsynchronized version corresponding to the data synchronization task is obtained from the stored source data, wherein the synchronization data of the unsynchronized version is the source data of different versions which has not been synchronized to the corresponding target end.

[0087] Optionally, the data synchronization task can be executed in a time manner. In a possible implementation, the synchronization data of the unsynchronized version is obtained from the stored source data in a preset second time interval as a period to perform data synchronization. Optionally, the second time interval can be the same as or different from the first time interval, and the present embodiment does not limit this.

[0088] In a possible implementation, in the first synchronization of the data synchronization task, the first synchronization can be performed from the first version of the stored source data (i.e. full data), and then in the synchronization process, the difference data between the source data version of the last synchronization completion and the latest stored source data version (i.e. incremental data of different versions based on full data) can be determined, and the difference data is obtained to perform data synchronization.

[0089] Step 104, synchronizing the synchronization data to the target end based on the target end information.

[0090] After obtaining the synchronization data corresponding to the data synchronization task, the synchronization data is synchronized to the corresponding target end according to the target end information corresponding to the data synchronization task. For example, when the target end information indicates that the target end is a target cluster, the synchronization data can be synchronized to the corresponding target cluster.

[0091] After synchronizing the synchronization data to the target end, the synchronization progress information corresponding to the data synchronization task can be updated. The updated synchronization progress information can be used to indicate that the source data that has been completed synchronization is the data version corresponding to the synchronization data.

[0092] In the embodiments of the present disclosure, the source data to be synchronized corresponding to any data synchronization task can be managed, the source data of different versions can be obtained and stored from the source end based on the source end information included in the registration information of the data synchronization task, and then in the data synchronization process, the synchronization data corresponding to the data synchronization task and not synchronized is obtained from the stored source data, so as to synchronize the synchronization data to the corresponding target end. By making the full data and the incremental data for the source data to be synchronized, the full data and the incremental data can be directly used as the original data to be directly imported into the target end when the data to be synchronized is required, that is, the required synchronization data can be directly obtained from the stored source data of different versions, without repeatedly performing multiple data calculation logics, which helps to improve the synchronization efficiency of synchronizing the source data from the source end to multiple target ends, helps to improve the data backup efficiency, and improves the stability of the business. In addition, the source data of different versions includes the full data and the incremental data obtained initially, and by storing the source data of different versions, the source data of different versions can be obtained for data synchronization in the data synchronization process, which helps to quickly synchronize the updated source data to the target end and improves the timeliness of data synchronization.

[0093] Optionally, the data storage end where the data source end is located has a data version management capability, and the stored data can be managed based on the version of data update. For example, the data storage end can be a data lake, and each time the data is updated, a version identifier can be generated to identify the current version information, and the version identifier can be recorded. Optionally, the version identifier can be a snapshot identifier (snapshot ID).

[0094] In a possible implementation, after obtaining and storing the source data of different versions from the data source end, the snapshot identifier of the source data obtained each time can be recorded, and the snapshot identifier is used to indicate the version information of the source data currently obtained.

[0095] In a possible implementation, as shown in Figure 2 The step 103 can include the following steps:

[0096] In step 1031, in the process of executing the data synchronization task, the synchronization progress information of the data synchronization task is obtained, and the latest synchronization snapshot identifier is included in the synchronization progress information, and the latest synchronization snapshot identifier is used to indicate the version information of the source data synchronized by the data synchronization task last time.

[0097] The data synchronization task corresponds to the synchronization progress information, wherein the synchronization progress information is used to indicate the data progress information of the data synchronized to the target end by the data synchronization task, and according to the synchronization progress information, the synchronization data in the stored source data that has not been synchronized to the target end can be determined, so as to obtain the corresponding synchronization data.

[0098] In a possible implementation, the synchronization progress can be identified based on the synchronized source data corresponding data snapshot identification, that is, the latest synchronization snapshot identification is included in the synchronization progress information.

[0099] Optionally, the latest synchronization snapshot identification included in the synchronization progress information corresponds to the recorded data snapshot identification, and after one synchronization is completed, the latest synchronization snapshot identification can be updated based on the recorded synchronization data corresponding data snapshot identification. The latest synchronization snapshot identification can be the data snapshot identification corresponding to the last data version synchronized in the last data synchronization process.

[0100] In step 1032, based on the latest synchronization snapshot identification, the source data corresponding to the data snapshot identification recorded after the latest synchronization snapshot identification is obtained from the stored source data, and the unsynchronized version of the synchronization data is obtained.

[0101] In a possible implementation, the data snapshot identification corresponds to the update time of the data version, the storage order of the source data of different versions is arranged according to the update time of each version in chronological order, and correspondingly, the data snapshot identification of the recorded source data of different versions is also arranged according to the update time of each version in chronological order, and the data snapshot identification corresponding to the source data of an earlier version is arranged at a more forward position.

[0102] In the data synchronization process, for the source data of different versions, synchronization is performed based on the arrangement order.

[0103] In a possible implementation, when each synchronization corresponding to the data synchronization task is performed, the source data corresponding to the data snapshot identification recorded after the latest synchronization snapshot identification can be obtained from the stored source data according to the obtained latest synchronization snapshot identification, and the source data corresponding to the data snapshot identification recorded after the latest synchronization snapshot identification is the unsynchronized version of the synchronization data.

[0104] Optionally, the synchronization progress information further includes synchronization state information of the current synchronization task, and the synchronization state information is used to indicate the execution state of the current data synchronization task, wherein the synchronization state can include an initial state (NEW), a synchronization execution state (RUNNING), a synchronization completion state (FINISHED) and a synchronization failure state (FAILED).

[0105] Illustratively, the state transition diagram of the data synchronization task can be as shown in Figure 3 When the data synchronization task has not started, the synchronization state is the initial state, after the first synchronization is started, the synchronization execution state is entered, when the synchronization is successful, the synchronization success state is updated, when the next synchronization is performed, the synchronization execution state is updated again, and when the synchronization fails, the synchronization failure state is updated, and the synchronization is attempted again, and the synchronization execution state is updated.

[0106] In a possible implementation, the registration information of the data synchronization task and the synchronization progress information can be stored in a data source management table for management, and different data synchronization tasks correspond to different data source management tables. Illustratively, a data source management table (t_datasource_management) corresponding to a data synchronization task is shown in Table 1 as follows:

[0107] Table 1

[0108]

[0109] Among them, the database name and the data table name correspond to the data source side information, the cluster type and the cluster identifier correspond to the target side information, the latest synchronization snapshot identifier and the synchronization state correspond to the synchronization progress information, and the table identifier is used to identify the data source management table.

[0110] In a possible implementation, when the latest synchronization snapshot identifier is obtained, the latest synchronization snapshot identifier can be obtained in a non-executing state (including an initial state, a success state and a failure state), so as to avoid errors in obtaining.

[0111] In the embodiments of the present disclosure, the source data is managed and synchronized based on the data snapshot identifiers corresponding to each version of the source data, and the source data is managed based on the version of the data, so that the latest updated data of the source data can be quickly synchronized to the target side, the data synchronization delay is reduced, and the timeliness of data synchronization is improved.

[0112] In a possible implementation, as shown in Figure 4 The step 102 includes the following steps.

[0113] In step 1021, in response to the absence of the recorded data snapshot identifier of the source data, full data corresponding to the current version of the source data is obtained from the data source side.

[0114] After obtaining and storing one version of the source data each time, the corresponding data snapshot identifier is recorded. In a possible implementation, when the source data corresponding to the data source side is obtained for the first time, that is, there is no recorded data snapshot identifier, the full data needs to be obtained from the data source side for the first time.

[0115] In a possible implementation, when the full data is acquired, the latest data snapshot identifier corresponding to the source data in the data source end can be acquired first, and all data corresponding to the current version of the source data, i.e., the full data of the source data in the current version, is acquired according to the latest data snapshot identifier. Optionally, the latest data snapshot identifier can be queried from the data management table corresponding to the data source end by using a structured query language (SQL) statement. For example, the latest data snapshot identifier can be queried from the iceberg table of the data lake, and the query statement can be SQL1, as shown in the following:

[0116] [SQL1]SELECT snapshot_id FROM my_iceberg_table.snapshots ORDER BYcommitted_at DESC LIMIT1.

[0117] In the foregoing, FROM my_iceberg_table.snapshots indicates that data is queried from the metadata table of the iceberg table, ORDER BY committed_at DESC LIMIT 1 is used to return the snapshot record of the first row arranged in descending order according to the snapshot commit time, and SELECT snapshot_id indicates that the snapshot identifier is selected.

[0118] After the latest data snapshot identifier is queried, a first acquisition statement can be executed to acquire and store all data corresponding to the current version of the source data. For example, the first acquisition statement can be SQL2, as shown in the following:

[0119] [SQL2]SELECT*FROM my_iceberg_table VERSION AS OF<snapshot_id>.

[0120] In the foregoing, VERSION AS OF is used to read the associated data file according to the specified data snapshot identifier.

[0121] After the full data is acquired, a full data identifier corresponding to the full data can be generated and recorded. The full data identifier includes the data snapshot identifier of the full data and the type identifier corresponding to the full data type. For example, the full data identifier can be full-<snapshot_id>.data, where full is used to indicate that the data type is full data, and snapshot_id is the data snapshot identifier corresponding to the full data.

[0122] At step 1022, in response to the existence of the recorded data snapshot identifier of the source data, the incremental data between the latest data snapshot identifier corresponding to the latest source data in the data source end and the last recorded data snapshot identifier is obtained.

[0123] When the data snapshot identifier corresponding to the recorded source data exists, it indicates that the historical version of the source data has been obtained, and in this case, the incremental data corresponding to the updated version can be obtained. In a possible implementation, the last recorded data snapshot identifier can be compared with the recorded data snapshot identifier in the data source end to obtain the data between the last recorded data snapshot identifier and the latest data snapshot identifier corresponding to the latest version of the source data in the data source end, and the data is the incremental data.

[0124] Optionally, the latest data snapshot identifier can be obtained by using the above SQL1, and a second obtaining statement can be executed to obtain and store the incremental data, where the second obtaining statement can be SQL3, as shown below:

[0125] [SQL3] SELECT * FROM my_iceberg_table FOR CHANGES BETWEEN <snapshot_id_1> AND <snapshot_id_2>.

[0126] The SQL3 is used to query and obtain the incremental data between the two data snapshot identifiers snapshot_id_1 and snapshot_id_2 in the iceberg table (my_iceberg_table).

[0127] After the incremental data is obtained, an incremental data identifier corresponding to the incremental data can be generated and recorded, where the incremental data identifier includes a data snapshot identifier of the incremental data and a type identifier corresponding to the type of the incremental data. For example, the incremental data identifier can be Incremental-<snapshot_id>.data, where Incremental is used to indicate that the data type is incremental data, and snapshot_id is the data snapshot identifier corresponding to the incremental data.

[0128] For example, the process of obtaining and storing the source data is as shown in FIG. 2B. Figure 5 As shown in FIG. 2B, the full data and the incremental data can be obtained from the data lake, and the corresponding full data file and incremental data file can be generated. The full data file includes full data with a full data identifier full-0001.data, and the incremental data file includes incremental data with incremental data identifiers Incremental-0002.data and Incremental-0003.data. 0001, 0002 and 0003 are data snapshot identifiers corresponding to different versions of the source data.

[0129] In the embodiments of the present disclosure, when the source data is obtained for the first time, the full data corresponding to the current version of the source data can be obtained, and when the source data is obtained subsequently, the incremental data can be obtained on the basis of the full data. The source data is managed and stored based on the full data and the incremental data, the source data can be synchronized to the target end in small batches, and the data synchronization efficiency can be improved.

[0130] In a possible implementation, as shown in FIG. 10, after step 1022, the method further includes the following steps: Figure 4

[0131] Step 1023: merging the historical stored full data and incremental data to obtain merged full data, and recording the data snapshot identifier corresponding to the merged full data.

[0132] When the source data update frequency is high, more incremental data files can be generated, occupying more storage space. In a possible implementation, the historical obtained and stored full data and incremental data can be merged to obtain merged full data, and the merged full data is the full data corresponding to the current period. For example, the preset target time can be one day, that is, the data is merged in a day as a period.

[0133] After the merged full data is generated, the merged full data can be stored, and the data snapshot identifier corresponding to the merged full data is recorded, wherein the data snapshot identifier corresponding to the merged full data is the last data snapshot identifier recorded in the last period.

[0134] Optionally, the full data identifier corresponding to the merged full data can be generated, and the full data identifier includes the data snapshot identifier corresponding to the merged full data, so as to identify the merged full data.

[0135] Step 1024: deleting the historical stored full data, incremental data and corresponding data snapshot identifier.

[0136] When the historical stored full data and incremental data are merged, the corresponding data can be deleted, including the corresponding data snapshot identifier.

[0137] For example, as shown in FIG. 10, after step 1024, the method further includes the following steps: Figure 6 ​As shown, the full data file generated in the first period includes full data identified as full-0001.data, and the incremental data file includes incremental data identified as Incremental-0002.data and Incremental-0003.data. When the second period is reached, the full data identified as full-0001.data and the incremental data identified as Incremental-0002.data and Incremental-0003.data can be merged to obtain merged full data, and the data of the first period can be deleted. The data snapshot identifier included in the full data identifier corresponding to the merged incremental data is the last data snapshot identifier recorded in the previous period, that is, 0003. In the second period, incremental data of versions 0004 and 0005 compared with version 0003 (namely, Incremental-0004.data and Incremental-0005.data) can be obtained. When the third period is reached, the full data identified as full-0003.data and the incremental data identified as Incremental-0004.data and Incremental-0005.data can be merged to obtain merged full data (full-0005.data), and the data of the second period can be deleted. In the second period, incremental data of versions 0006 and 0007 compared with version 0005 (Incremental-0006.data and Incremental-0007.data) can be obtained.

[0138] In the embodiments of the present disclosure, the source data obtained and stored can be stored by period to avoid occupying too much memory space by storing too many files, thereby saving the use of memory space and reducing resource waste.

[0139] In a possible implementation, the source data can be synchronized based on the stored full data and incremental data during the data synchronization process. The step 1032 further includes the following steps.

[0140] In step 1321, in response to the latest synchronization snapshot identifier being an initial state value, the full data corresponding to the last recorded full data identifier is obtained from the stored source data. The full data identifier includes the data snapshot identifier corresponding to the full data.

[0141] The data synchronization task corresponds to the data source management table through the record of the latest synchronization snapshot identifier to record the synchronization progress. In a possible implementation, when performing the first synchronization of the data synchronization task, the value corresponding to the latest synchronization snapshot identifier is the initial state value, i.e., the default value, indicating that the synchronization has not been performed. In this case, the latest full data can be obtained from the stored source data, i.e., the full data corresponding to the latest record of the full data identifier, to synchronize the latest full data first.

[0142] Illustratively, Table 2 shows the values corresponding to the fields in the data source management table of the three data synchronization tasks, and Table 2 is as follows:

[0143] Table 2

[0144] db_name table_name cluster_type cluster_id latest_snapshot_id status db1 tb1 StarRocks 10001 -1 NEW db1 tb1 StarRocks 10002 -1 NEW db1 tb1 ClickHouse 20001 -1 NEW

[0145] In Table 2, the latest synchronization snapshot identifiers corresponding to the three data synchronization tasks for synchronizing the source data in tb1 to three different clusters are shown, which are all -1, indicating the initial state value, and the three data synchronization tasks have not been performed.

[0146] Illustratively, in combination with Figure 5 As shown, the stored different versions of the source data include the full data of the full data full-0001.data and the incremental data of the incremental data Incremental-0002.data and Incremental-0003.data. When the obtained latest synchronization snapshot identifier is the initial state value, the latest stored full data full-0001.data can be obtained.

[0147] In step 1322, in response to the latest synchronization snapshot identifier being the record identifier value, the incremental data corresponding to the incremental data identifier recorded after the latest synchronization snapshot identifier is obtained from the stored source data. The incremental data identifier includes the data snapshot identifier corresponding to the incremental data.

[0148] When the value corresponding to the latest synchronization snapshot identifier is the record identifier value, it indicates that the synchronization has been performed. The incremental data corresponding to the incremental data identifier recorded after the latest synchronization snapshot identifier can be obtained from the stored source data, wherein the data version update time indicated by the data snapshot identifier included in the incremental data identifier recorded after the latest synchronization snapshot identifier is later than the data version corresponding to the data snapshot identifier included in the latest synchronization snapshot identifier.

[0149] Illustratively, in combination with Figure 5As shown, when the latest synchronization snapshot identifier of the record is full-0001.data, the incremental data corresponding to the incremental data identifiers recorded after the latest synchronization snapshot identifier can be obtained from the stored source data, including Incremental-0002.data and Incremental-0003.data, and the incremental data can be executed as the synchronization data for synchronization.

[0150] In a possible implementation, in response to successful synchronization of the synchronization data, the synchronization progress information is updated based on the latest data identifier corresponding to the synchronization data. Optionally, the latest synchronization snapshot identifier can be updated to the latest data identifier corresponding to the synchronization data.

[0151] In combination with the above example, when the full-0001.data is successfully synchronized, the latest synchronization snapshot identifier can be updated to full-0001.data, and when the incremental data Incremental-0002.data and Incremental-0003.data are successfully synchronized, the latest synchronization snapshot identifier can be updated to Incremental-0003.data.

[0152] Optionally, after the synchronization data is successfully synchronized, the synchronization state can be updated to FINISHED, if the synchronization fails, the synchronization is retried, if the number of retries reaches a preset threshold, the synchronization state can be updated to FAILED, and an alarm information is generated to prompt timely processing, and illustratively, the preset threshold can be 1.

[0153] In the embodiments of the present disclosure, in the initial synchronization of the data synchronization task, the latest full-amount data can be obtained from the stored source data for data synchronization, and the data synchronization of large data volume is completed. Since the full-amount data of the source data has been generated, the target end can be directly imported without complex calculation, which helps to improve the data synchronization efficiency. After the full-amount data is synchronized, the incremental data can be synchronized in the subsequent synchronization process. Through the synchronization of the incremental data, the data delay can be reduced, and the timeliness of data synchronization can be greatly improved.

[0154] In a possible implementation, after the data synchronization is completed, the synchronization data is also detected to ensure the integrity of the synchronization data, wherein the data of the target end is stored according to the synchronization time corresponding to the synchronization data, and illustratively, the data of the target end can be stored in units of days, and the data synchronized yesterday is stored in the same area, and the data synchronized today is stored in another area. Optionally, the synchronization data can be detected according to the time partition, and after step 104, the following steps are further included.

[0155] Step 105, for any time partition to be detected, the target end data volume corresponding to the time partition is compared with the data source end data volume that should be synchronized with the time partition, to obtain a data comparison result.

[0156] For any time partition, data detection can be performed. In a possible implementation, since the data synchronization of the time partition corresponding to the historical time has been completed, data detection can be performed on the data of the time partition corresponding to the historical time. For example, data detection can be performed on the data corresponding to yesterday.

[0157] In the data detection process, the target end data volume corresponding to the time partition can be compared with the data source end data volume that should be synchronized with the time partition. Both the target end data volume and the data source end data volume refer to the total data volume. The data source end data volume that should be synchronized with the time partition refers to the total data volume of all data contained in the data source end within the time end corresponding to the time partition. For example, when the time period corresponding to the time partition is yesterday, the total data volume of yesterday stored in the target end can be compared with the total data volume of the data source end corresponding to yesterday, to obtain a data comparison result.

[0158] The data comparison result is used to indicate whether the total data volume of the synchronization double-end is consistent. When consistent, it indicates that the synchronization is correct and the data has been completely synchronized. When inconsistent, it indicates that the data synchronization has missed some data and the data needs to be completed.

[0159] In a possible implementation, when the data source end data volume is counted, counting can be performed based on the data snapshot identifier. Optionally, the last data snapshot identifier of the data source end within the time end corresponding to the time partition can be obtained, and counting can be performed based on the data snapshot identifier. Optionally, the data source end data volume can be counted by executing a first statistical statement. For example, the first statistical statement can be SQL4, as shown below:

[0160] [SQL4] SELECT pt, count(0) FROM my_iceberg_table VERSION AS OF <snapshot_id> group by pt

[0161] Wherein, pt represents the partition field of the table, count(0) is used to count the number of records in the group, that is, the data volume, and group by pt indicates grouping according to the value of the field pt. The data volume is calculated for each group.

[0162] Optionally, the data volume of the target end can be counted by executing a second statistical statement. The second statistical statement can be SQL5, as shown below:

[0163] [SQL5] SELECT pt, count(0) FROM my_iceberg_table group by pt

[0164] In response to the data comparison result indicating that the data volumes are inconsistent, the target-end data corresponding to the time partition is complemented based on the data source-end data corresponding to the time partition.

[0165] When the data comparison result indicates inconsistency, data complementation is required. In one possible implementation, the data corresponding to the time partition in the target end can be deleted, the data corresponding to the time partition is obtained from the data source end, and the data is synchronized to the target end.

[0166] In another possible implementation, data complementation can be performed in the form of data replacement to reduce the user's perception of use in the data complementation process. The process includes the following steps:

[0167] In step 1061, replacement data is generated based on the data source-end data corresponding to the time partition.

[0168] First, the replacement data table can be generated based on the data source-end data corresponding to the time partition. For example, when the data volume corresponding to yesterday is inconsistent, the data of the data source end of yesterday is obtained to generate the replacement data table.

[0169] In step 1062, the replacement data is imported into the target end, and the target-end data corresponding to the time partition in the target end is deleted.

[0170] After the replacement data table is generated, the replacement data table can be imported into the target end, and the original data stored in the target end corresponding to the time partition is deleted to shorten the data complementation time and improve the data complementation efficiency, which helps to reduce the user's perception of data complementation in the process of using data.

[0171] In the embodiments of the present disclosure, through data detection, whether the data in the downstream target end is abnormal, missing or damaged can be detected. When the data in the downstream target end is abnormal and damaged, it can be self-perceived and repaired, which helps to improve the robustness of the business.

[0172] In one possible implementation, the structure of the data synchronization system is as follows Figure 7As shown, the system includes a data source management unit, a snapshot management unit, a data synchronization unit, and a data detection unit. The data source management unit is used to manage data sources, specifically, to manage registration information of data synchronization tasks. The snapshot management unit is used to obtain source data from the data source end located in the data lake and generate corresponding data files (including full data files and incremental data files). The data synchronization unit is used to execute data synchronization tasks, can obtain synchronization progress information, obtain unsynchronized version of synchronization data from the stored data files based on the synchronization progress information, and execute data synchronization. The data detection unit is used to detect data and complete data. In the case that the data volume of the synchronized data at the data source end and the target end is inconsistent, data completion can be performed.

[0173] As shown, the data synchronization process can be performed after the snapshot management unit first obtains full data and periodically obtains incremental data. The data synchronization unit can obtain synchronization progress information from the data source management unit to perform data synchronization according to the synchronization progress information. After completing data synchronization, the synchronization progress information can be updated. Figure 8

[0174] As shown, the data detection process includes the following steps. Figure 9 As shown, first, a statistical statement is executed at the data source end and the target end to obtain the data volume of the data source end and the data volume of the target end corresponding to the same time partition, determine whether the data volume of the data source end and the data volume of the target end are consistent, if consistent, end, if not consistent, generate a complete data file based on the data source end data, and synchronize the complete data file to the target end.

[0175] In the embodiments of the present disclosure, for the problem of synchronizing data among multiple target ends, based on the version management capability of the storage end (such as the data lake) corresponding to the data source end, full data and incremental data can be made for the source data to be synchronized at regular intervals. When data needs to be synchronized, full data and incremental data can be directly used as raw data to import into the target end, solving the problem of repeated calculation of multiple synchronization tasks, achieving rapid cluster backup or fault recovery, and more efficiently ensuring the stability of the business. In addition, each small batch of data import generates a snapshot version, and based on the version management capability, incremental data can be quickly captured and quickly synchronized to the downstream target end, greatly improving the timeliness of data synchronization, reducing data delay from hours to minutes. At the same time, a data detection module is added, which can perceive and repair itself when the data in the downstream target end is abnormal or damaged, improving the robustness of the business.

[0176] ​Specifically, for multiple data synchronization tasks, the calculation of generating full data and incremental data with complex calculation logic is only performed once, and the full data and the incremental data can be directly imported into the target end for use, greatly reducing repeated calculation. For the same data source end, each target end can generate the same file data, and the data in the target end is unified, and the data is stored in the form of a local table in each target end, which improves the data synchronization efficiency and also uses the various index optimization capabilities of the target end, facilitating use.

[0177] In addition, through data snapshot management, the perception interval of data synchronization can be reduced, and through unified incremental data capture, incremental data can be synchronized to multiple target ends that need to be synchronized, ensuring the timeliness of the data. In addition, through the data detection module, inconsistent data in the downstream target end can be automatically checked and repaired, ensuring the consistency of data query.

[0178] As shown in Figure 10 The device includes:

[0179] The information acquisition module 1001 is configured to acquire registration information corresponding to any data synchronization task in at least one data synchronization task, wherein the registration information includes data source end information and target end information corresponding to the data synchronization task;

[0180] The data storage module 1002 is configured to acquire and store different versions of source data from the data source end based on the data source end information, wherein the different versions of source data include full data initially acquired and incremental data based on the full data;

[0181] The data acquisition module 1003 is configured to acquire synchronization data of an unsynchronized version corresponding to the data synchronization task from the stored source data;

[0182] The data synchronization module 1004 is configured to synchronize the synchronization data to the target end based on the target end information.

[0183] In one example embodiment, the device further includes:

[0184] The snapshot recording module is configured to record a data snapshot identifier of each acquired source data, wherein the data snapshot identifier is used to indicate version information of the source data;

[0185] The data acquisition module 1003 is further configured to:

[0186] In the process of executing the data synchronization task, synchronization progress information of the data synchronization task is acquired, the synchronization progress information including a latest synchronization snapshot identifier, the latest synchronization snapshot identifier being used to indicate version information of source data of a latest synchronization of the data task;

[0187] Based on the latest synchronization snapshot identifier, source data corresponding to a data snapshot identifier recorded after the latest synchronization snapshot identifier is acquired from stored source data, to obtain the unsynchronized version of the synchronization data, wherein source data of different versions are arranged in order of version update time.

[0188] In an exemplary embodiment, the data storage module 1002 is further configured to:

[0189] In response to the absence of a recorded data snapshot identifier of source data, full-amount data corresponding to current version source data is acquired from the data source end;

[0190] In response to the presence of a recorded data snapshot identifier of source data, incremental data between a latest data snapshot identifier corresponding to latest source data in the data source end and a last recorded data snapshot identifier is acquired.

[0191] In an exemplary embodiment, the data acquisition module 1003 is further configured to:

[0192] In response to the latest synchronization snapshot identifier being an initial state value, full-amount data corresponding to a last recorded full-amount data identifier is acquired from the stored source data, the full-amount data identifier including a data snapshot identifier corresponding to the full-amount data;

[0193] In response to the latest synchronization snapshot identifier being a recorded identifier value, incremental data corresponding to an incremental data identifier recorded after the latest synchronization snapshot identifier is acquired from the stored source data, the incremental data identifier including a data snapshot identifier corresponding to the incremental data.

[0194] In an exemplary embodiment, the apparatus further includes:

[0195] A data merging module is configured to merge historical stored full-amount data and incremental data to obtain merged full-amount data and record a data snapshot identifier corresponding to the merged full-amount data at a preset target time as a period.

[0196] A data deletion module is configured to delete historical stored full-amount data, incremental data, and corresponding data snapshot identifiers.

[0197] In an exemplary embodiment, the apparatus further includes:

[0198] The progress updating module is configured to update the synchronization progress information based on the latest data identifier corresponding to the synchronization data in response to the synchronization data being successfully synchronized.

[0199] In one example embodiment, the apparatus further comprises:

[0200] The data detecting module is configured to, for any time partition to be detected, compare the target end data amount corresponding to the time partition with the data source end data amount that should be synchronized for the time partition to obtain a data comparison result.

[0201] The data supplementing module is configured to, in response to the data comparison result indicating that the data amounts are inconsistent, perform data supplementing on the target end data corresponding to the time partition based on the data source end data corresponding to the time partition in the data source end.

[0202] In one example embodiment, the data supplementing module is further configured to:

[0203] generate replacement data based on the data source end data corresponding to the time partition in the data source end;

[0204] import the replacement data into the target end and delete the target end data corresponding to the time partition in the target end.

[0205] The data synchronization apparatus in the embodiments of the present disclosure and the embodiments of the data synchronization method described above correspond to each other, and related content can be referred to each other, which will not be described herein again. The beneficial technical effects of the data synchronization apparatus in the embodiments of the present disclosure can be referred to the corresponding beneficial technical effects of the example method part described above, which will not be described herein again.

[0206] In addition, the embodiments of the present disclosure further provide an electronic device, comprising:

[0207] a memory configured to store a computer program;

[0208] a processor configured to execute the computer program stored in the memory, and the computer program, when executed, implements the data synchronization method described in any of the embodiments of the present disclosure.

[0209] Figure 11 is a structural schematic diagram of one application embodiment of the electronic device of the present disclosure. Hereinafter, the electronic device according to the embodiments of the present disclosure will be described with reference to Figure 11 As shown in Figure 11 , the electronic device comprises one or more processors and a memory.

[0210] The processor can be a central processing unit (CPU) or other form of processing unit having data processing and / or instruction execution capabilities and can control other components in the electronic device to perform desired functions.

[0211] The memory can include one or more computer program products that can include various forms of computer-readable storage media, such as volatile memory and / or non-volatile memory. The volatile memory, for example, can include random access memory (RAM), cache memory, and / or the like. The non-volatile memory, for example, can include read only memory (ROM), hard disk, flash memory, and / or the like. One or more computer program instructions can be stored on the computer-readable storage media, and the processor can execute the program instructions to implement the data synchronization method of various embodiments of the disclosure described above and / or other desired functions.

[0212] In one example, the electronic device can further include an input device and an output device, which are interconnected through a bus system and / or other forms of connection mechanisms (not shown).

[0213] In addition, the input device can further include, for example, a keyboard, a mouse, and / or the like.

[0214] The output device can output various information, including the determined distance information, direction information, and / or the like, to the outside. The output device can include, for example, a display, a speaker, a printer, a communication network and a remote output device connected thereto, and / or the like.

[0215] Of course, in order to simplify, Figure 11 Only some of the components related to the present disclosure among the electronic device are shown in FIG. 1, and components such as a bus, an input / output interface, and / or the like are omitted. In addition, the electronic device can further include any other appropriate components according to a specific application.

[0216] In addition to the above-described method and device, an embodiment of the present disclosure can be a computer program product including computer program instructions that, when executed by a processor, cause the processor to perform the steps of the data synchronization method according to various embodiments of the present disclosure described in the above parts of the specification.

[0217] The computer program product can be written in any combination of one or more programming languages, including an object oriented programming language such as Java, C++, or the like, and conventional procedural programming languages, such as the "C" programming language or similar programming languages. The program code can execute entirely on the user's computing device, partly on the user's device, as a stand-alone software package, partly on the user's computing device and partly on a remote computing device or entirely on the remote computing device or server.

[0218] Furthermore, embodiments of the present disclosure can also be a computer readable storage medium, having stored thereon computer program instructions which, when executed by a processor, cause the processor to perform the steps described in the foregoing disclosure of various embodiments of the present disclosure.

[0219] The computer readable storage medium can be a combination of one or more computer readable media. The computer readable media can be a computer readable signal medium or a computer readable storage medium. The computer readable storage medium can include, for example, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples (a non-exhaustive list) of the computer readable storage medium include an electrical connection having one or more wires, a portable disc, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0220] Those skilled in the art can understand that all or part of the steps of the above-mentioned method embodiments can be completed by program instruction related hardware, and the foregoing program can be stored in a computer readable storage medium, and the program executes the steps of the above-mentioned method embodiments when executed; and the foregoing storage medium includes ROM, RAM, magnetic disc or optical disc and various storage medium that can store program code.

[0221] The above describes the basic principles of the present disclosure in combination with specific embodiments, but it should be noted that the advantages, advantages, effects and the like mentioned in the present disclosure are only examples and are not limited, and these advantages, advantages, effects and the like cannot be considered as the various embodiments of the present disclosure must have. In addition, the above specific details are only for the purpose of example and for the purpose of understanding, and the above details do not limit the present disclosure to the above specific details.

[0222] The various embodiments described in this specification are intended to be illustrative only and in no way limit the scope of the application. One skilled in the art will readily recognize from the disclosure herein, possible alternative techniques within the scope of the application. Accordingly, the examples are not to be regarded as limiting, but rather are to be understood to be illustrative of the possible aspects of the application. The various embodiments described in this specification are described in the context of a system. As such, the system embodiments are described in relatively greater detail than the method embodiments, with the understanding that the method embodiments are substantially analogous to the system embodiments.

[0223] The block diagrams of devices, apparatuses, equipment, systems referred to in this disclosure are merely illustrative examples and are not intended to require or imply that the connection, arrangement, configuration must be as shown in the block diagrams. These devices, apparatuses, equipment, systems can be connected, arranged, configured in any manner as will be appreciated by those skilled in the art. Words such as "include," "contain," "have," etc. are open-ended words that are to be interpreted to mean "including but not limited to," and are to be taken in a non-limiting sense. The words "or" and "and" as used herein are to be interpreted as the word "and / or," and are to be taken in a non-limiting sense. The word "such as" as used herein is to be interpreted as the phrase "such as but not limited to," and is to be taken in a non-limiting sense.

[0224] The methods and apparatuses of the present disclosure can be implemented in a number of ways. For example, the methods and apparatuses of the present disclosure can be implemented using software, hardware, firmware, or any combination of these. The above described order of steps for the methods is merely illustrative, and the steps of the methods of the present disclosure are not limited to the order specifically described above unless otherwise specifically stated. Furthermore, in some embodiments, the present disclosure can also be implemented as a program recorded in a recording medium, which includes machine readable instructions for implementing the methods according to the present disclosure. Thus, the present disclosure also covers a recording medium storing a program for executing the methods according to the present disclosure.

[0225] It is also important to note that the apparatuses, equipment and methods of the present disclosure can be embodied in a variety of ways. These are to be considered as equivalent to one another.

[0226] The above description of the disclosed aspects is intended to be illustrative only and not limiting of the scope of the application. Various modifications to these aspects will be readily apparent to those skilled in the art, and the generic principles defined herein can be applied to other aspects without departing from the scope of the application. Thus, the present disclosure is not intended to be limited to the aspects shown herein but is to be accorded the widest scope consistent with the claims and the principles and novel features disclosed herein.

[0227] The foregoing description has been presented for the purposes of illustration and description. Furthermore, the description is not intended to limit the embodiments of the disclosure to the forms disclosed herein. Although the various example aspects and embodiments have been described herein with regard to particular aspects and embodiments, those skilled in the art will recognize that certain modifications, changes, substitutions, additions and sub-combinations can be made without departing from the spirit of the disclosure.

Claims

1. A method of data synchronization, the method comprising: The method comprises the following steps: For any data synchronization task in at least one data synchronization task, obtain the registration information corresponding to the data synchronization task, wherein the registration information comprises data source end information and target end information corresponding to the data synchronization task; Based on the data source end information, obtain and store different versions of source data from the data source end, wherein different versions of source data include full data initially obtained and incremental data based on the full data; Obtain the synchronization data of the unsynchronized version corresponding to the data synchronization task from the stored source data; Based on the target end information, synchronize the synchronization data to the target end.

2. The method of claim 1, wherein, After obtaining and storing different versions of source data from the data source end, the method further comprises: Record the data snapshot identifier of the source data obtained each time, which is used to indicate the version information of the source data; The method further comprises: In the process of executing the data synchronization task, obtain the synchronization progress information of the data synchronization task, wherein the synchronization progress information comprises the latest synchronization snapshot identifier, which is used to indicate the version information of the source data synchronized by the data synchronization task last time; Based on the latest synchronization snapshot identifier, obtain the source data corresponding to the data snapshot identifier recorded after the latest synchronization snapshot identifier from the stored source data to obtain the synchronization data of the unsynchronized version, wherein different versions of source data are arranged in order of version update time.

3. The method of claim 2, wherein, The method further comprises: In response to the absence of the recorded data snapshot identifier of the source data, obtain the full data corresponding to the current version of source data from the data source end; In response to the presence of the recorded data snapshot identifier of the source data, obtain the incremental data between the latest data snapshot identifier corresponding to the latest source data in the data source end and the last recorded data snapshot identifier.

4. The method of claim 2, wherein, The method further comprises: In response to the latest synchronization snapshot identifier being an initial state value, obtain the full data corresponding to the last recorded full data identifier from the stored source data, wherein the full data identifier comprises the data snapshot identifier corresponding to the full data; In response to the latest synchronization snapshot identifier being a recorded identifier value, obtain the incremental data corresponding to the incremental data identifier recorded after the latest synchronization snapshot identifier from the stored source data, wherein the incremental data identifier comprises the data snapshot identifier corresponding to the incremental data.

5. The method of claim 3, wherein, After obtaining and storing different versions of source data from the data source end, the method further comprises: Merge the historical stored full data and incremental data to obtain merged full data at a preset target time, and record the data snapshot identifier corresponding to the merged full data; Delete the historical stored full data, incremental data and corresponding data snapshot identifier.

6. The method of claim 2, wherein, The method further comprises: In response to the synchronization data synchronization success, updating the synchronization progress information based on the latest data identification corresponding to the synchronization data.

7. The method according to any one of claims 1 to 6, characterized in that, The data of the target end is stored according to synchronization time partition corresponding to the synchronization data, and the method further comprises: For any time partition to be detected, comparing the target end data amount corresponding to the time partition with the data source end data amount that should be synchronized in the time partition to obtain a data comparison result; In response to the data comparison result indicating that the data amounts are inconsistent, performing data completion on the target end data corresponding to the time partition based on the data source end data corresponding to the time partition in the data source end.

8. The method of claim 7, wherein, The data completion on the target end data corresponding to the time partition based on the data source end data corresponding to the time partition in the data source end comprises: generating replacement data based on the data source end data corresponding to the time partition in the data source end; importing the replacement data into the target end and deleting the target end data corresponding to the time partition in the target end.

9. An electronic device, comprising: comprise: a memory for storing a computer program; a processor for executing the computer program stored in the memory, and when the computer program is executed, the data synchronization method of any one of claims 1-8 is implemented.

10. A computer program product comprising a computer program, characterized in that, When the computer program is executed by the processor, the data synchronization method of any one of claims 1-8 is implemented.