Data synchronization method and device, equipment, storage medium and computer program product

By dividing the data synchronization task of large data into multiple subtasks corresponding to the file partition path and performing these subtasks in parallel, the problem of inefficiency of traditional data synchronization methods is solved and efficient data synchronization is achieved.

CN120179728APending Publication Date: 2025-06-20SF TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311768886.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-12-20
Publication Date
2025-06-20

AI Technical Summary

Technical Problem

Traditional data synchronization methods are inefficient when processing large amounts of data, and cannot effectively use multithreads to synchronize data.

Method used

By obtaining the data synchronization task and determining its data volume, if it exceeds the preset value, multiple file partition paths of the data table are obtained, and the data synchronization task is divided into multiple data synchronization subtasks corresponding to each file partition path, and these subtasks are executed in parallel.

Benefits of technology

It improves the efficiency of data synchronization and can effectively handle data synchronization tasks with large amounts of data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120179728A_ABST
    Figure CN120179728A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a data synchronization method and device, computer equipment, a storage medium and a computer program product, and relates to the technical field of big data. According to the invention, the data synchronization efficiency can be improved. The method comprises the following steps: acquiring a data synchronization task, determining the data volume of a data table corresponding to the data synchronization task, and determining a plurality of file partition paths of the data table under the condition that the data volume exceeds a preset data volume, and then segmenting the data synchronization task into a plurality of data synchronization sub-tasks respectively corresponding to the plurality of file partition paths, and executing the plurality of data synchronization sub-tasks in parallel.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of big data technology, and particularly to a data synchronization method, apparatus, computer device, storage medium, and computer program product. Background Art

[0002] The field of data processing involves data synchronization between databases, and there is a scenario of data synchronization processing between databases that are not in the same internal network.

[0003] In the above synchronization scenario, according to the traditional data synchronization processing method, a query statement is used for query, and usually only single-threaded data synchronization can be performed, resulting in the technical problem of low data synchronization efficiency. Summary of the Invention

[0004] Based on this, it is necessary to provide a data synchronization method, apparatus, computer device, storage medium, and computer program product for the above technical problems.

[0005] In a first aspect, this application provides a data synchronization method. The method includes:

[0006] Obtain a data synchronization task, and determine the data volume of the data table corresponding to the data synchronization task;

[0007] In the case where the data volume exceeds a preset data volume, determine multiple file partition paths of the data table;

[0008] Split the data synchronization task into multiple data synchronization subtasks respectively corresponding to the multiple file partition paths;

[0009] Execute the multiple data synchronization subtasks in parallel.

[0010] In one embodiment, the determining the multiple file partition paths of the data table includes: obtaining the abstract syntax tree corresponding to the data synchronization task; based on the abstract syntax tree, obtaining the multiple file partition paths corresponding to the data table.

[0011] In one embodiment, the obtaining the multiple file partition paths corresponding to the data table based on the abstract syntax tree includes: identifying the data address information of the data table included in the abstract syntax tree; based on the data address information, obtaining the multiple file partition paths corresponding to the data table.

[0012] In one embodiment, the determining the data volume of the data table corresponding to the data synchronization task includes: determining the number of rows of the data table; determining the data volume of the data table according to the number of rows.

[0013] In one embodiment, determining the number of rows in the data table includes: executing a query plan execution statement on the data synchronization task to obtain the number of rows in the data table.

[0014] In one embodiment, the method further includes: executing the data synchronization task when the data volume does not exceed a preset data volume.

[0015] In a second aspect, the present application provides a data synchronization device. The device includes:

[0016] An acquisition module, configured to acquire a data synchronization task and determine the data volume of the data table corresponding to the data synchronization task;

[0017] A determination module, configured to determine multiple file partition paths of the data table when the data volume exceeds a preset data volume;

[0018] A task splitting module, configured to split the data synchronization task into multiple data synchronization subtasks respectively corresponding to the multiple file partition paths;

[0019] A task execution module, configured to execute the multiple data synchronization subtasks in parallel.

[0020] In a third aspect, the present application further provides a computer device. The computer device includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the following steps are implemented:

[0021] Acquire a data synchronization task and determine the data volume of the data table corresponding to the data synchronization task;

[0022] When the data volume exceeds a preset data volume, determine multiple file partition paths of the data table;

[0023] Split the data synchronization task into multiple data synchronization subtasks respectively corresponding to the multiple file partition paths;

[0024] Execute the multiple data synchronization subtasks in parallel.

[0025] In a fourth aspect, the present application further provides a computer-readable storage medium. The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the following steps are implemented:

[0026] Acquire a data synchronization task and determine the data volume of the data table corresponding to the data synchronization task;

[0027] When the data volume exceeds a preset data volume, determine multiple file partition paths of the data table;

[0028] Split the data synchronization task into multiple data synchronization subtasks respectively corresponding to the multiple file partition paths;

[0029] Execute the multiple data synchronization subtasks in parallel.

[0030] In a fifth aspect, the present application also provides a computer program product. The computer program product includes a computer program, and when the computer program is executed by a processor, the following steps are implemented:

[0031] Obtain a data synchronization task, and determine the data volume of the data table corresponding to the data synchronization task;

[0032] In the case that the data volume exceeds a preset data volume, determine multiple file partition paths of the data table;

[0033] Split the data synchronization task into multiple data synchronization subtasks respectively corresponding to the multiple file partition paths;

[0034] Execute the multiple data synchronization subtasks in parallel.

[0035] The above data synchronization method, device, computer device, storage medium and computer program product obtain a data synchronization task and determine the data volume of the data table corresponding to the data synchronization task. In the case that the data volume exceeds a preset data volume, determine multiple file partition paths of the data table, and then split the data synchronization task into multiple data synchronization subtasks respectively corresponding to the multiple file partition paths, and execute the multiple data synchronization subtasks in parallel. This solution can divide the data synchronization task into multiple data synchronization subtasks corresponding to each file partition path according to the data volume of the data table corresponding to the data synchronization task, and perform data synchronization by executing the multiple data synchronization subtasks in parallel, thereby improving the efficiency of data synchronization. Description of the Drawings

[0036] Figure 1 It is an application environment diagram of a data synchronization method provided by an embodiment of the present application;

[0037] Figure 2(a) is a schematic flowchart of a data synchronization method in the prior art;

[0038] Figure 2(b) is a schematic flowchart of the data synchronization method according to an embodiment of the present application;

[0039] Figure 3 It is a schematic flowchart of a data synchronization method provided by an embodiment of the present application;

[0040] Figure 4 It is a schematic flowchart of a method for determining multiple file partition paths of a data table provided by an embodiment of the present application;

[0041] Figure 5Schematic flowchart of another data synchronization method provided by an embodiment of the present application;

[0042] Figure 6 Block diagram of the structure of a data synchronization device provided by an embodiment of the present application;

[0043] Figure 7 Internal structure diagram of a computer device provided by an embodiment of the present application. Detailed implementation manners

[0044] In order to make the objectives, technical solutions and advantages of the present application clearer and more understandable, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.

[0045] First, the application environment of the data synchronization method in the embodiments of the present application will be introduced. As Figure 1 shown, the data synchronization server communicates with the server to be data synchronized through the internal network; the data synchronization server communicates with the target server through the public network; the server to be data synchronized communicates with the target server through the public network. The data synchronization server, the server to be data synchronized, and the target server can be implemented by independent servers or a server cluster composed of multiple servers.

[0046] Based on the data synchronization processing between the target server and the server to be data synchronized in this application environment, the traditional method is shown in Figure 2(a). There may be a firewall between the internal network and the public network. When the data synchronization server executes the data synchronization task for the target database of the target server, the data synchronization task can be an SQL query task, that is, an SQL query statement. The SQL query statement can be:

[0047]

[0048] The data synchronization server can directly execute this SQL query task by using the Java Database Connectivity (JDBC) interface, read the synchronization data of the data table corresponding to this SQL query task in the target database in a single-threaded manner, and send the synchronization data to the data synchronization server to be synchronized through the internal network. The data synchronization server to be synchronized can write the synchronization data into the database to be synchronized. The JDBC is a set of application programming interfaces (APIs) provided by the Java platform for interacting with databases. The JDBC allows Java application programs to establish connections with various relational databases and execute database operations such as querying data, inserting, updating, and deleting data. In this method, when the amount of data in the data table of this SQL query task is large, there will be a problem of low data synchronization efficiency. In addition, the available processing methods also include: Similarly, when the data synchronization server executes the data synchronization task for the target database of the target server, the data synchronization task can be an SQL query task, that is, an SQL query statement. The SQL query statement can be:

[0049]

[0050] Select a certain field of the data table of this SQL query task as the splitting key. The field value of this field is usually numeric, such as values of int and long types; furthermore, the maximum field value and the minimum field value of this field can be obtained, and according to the maximum field value and the minimum field value, this SQL query task can be split into multiple SQL query subtasks. For example, this SQL query task can be:

[0051]

[0052] These multiple SQL query subtasks can be respectively:

[0053]

[0054] In this method, when the amount of data in the data table of this SQL query task is large, there will still be a problem of low data synchronization efficiency.

[0055] In this regard, for the data synchronization method provided in the embodiments of the present application, referring to FIG. 2(b), this solution can determine whether to split a data synchronization task according to the data volume of the data table corresponding to the data synchronization task; and in the case where the data volume exceeds a preset data volume, obtain the abstract syntax tree corresponding to the data synchronization task, and determine multiple file partition paths of the data table based on the abstract syntax tree; further, according to the multiple file partition paths, split the data synchronization task into multiple data synchronization subtasks respectively corresponding to the multiple file partition paths; and perform data synchronization by parallelly executing the multiple data synchronization subtasks, thereby improving the efficiency of data synchronization.

[0056] Based on the application environment as follows Figure 1 shown, in combination with each embodiment and the corresponding drawings, the data synchronization method of the present application will be specifically described.

[0057] In one embodiment, as Figure 3 shown, a data synchronization method is provided. Taking the data synchronization server in Figure 1 as an example for illustration, the method includes the following steps:

[0058] Step S301: Obtain a data synchronization task and determine the data volume of the data table corresponding to the data synchronization task.

[0059] Among them, the data synchronization task is used to synchronize the new data and / or updated data in the target database, as well as the data that meets the data query conditions in the data synchronization task, to the database to be data synchronized. In a possible implementation manner, the target database and the database to be data synchronized may be a distributed file system (Hadoop Distributed File System, Hadoop). Among them, a data warehouse analysis system Hive can be built based on the Hadoop distributed file system. Further, the Hive can be used to execute the data synchronization task, and the data synchronization task may be an sql query task, that is, an sql query statement. It should be understood that generally, it can be determined whether to split the sql query task according to the data volume of the data table corresponding to the data synchronization task. In the case where the data volume of the data table exceeds the preset data volume, the sql query task can be split to obtain multiple sql query subtasks; in the case where the data volume of the data table does not exceed the preset data volume, the sql query task can be directly executed. In a possible implementation manner, the data volume of the data table can be determined based on the number of rows of the data table corresponding to the sql query task.

[0060] Step S302: In the case where the data volume exceeds the preset data volume, determine multiple file partition paths of the data table.

[0061] Among them, the data volume of the data table of the data synchronization task can be measured by the number of rows of the data table. First, a preset number of rows can be set, for example, 1 million. Furthermore, the number of rows of the data table can be identified, and the number of rows of the data table can be compared with the preset number of rows. When the number of rows of the data table exceeds the preset number of rows, it can be determined that the data volume of the data table exceeds the preset data volume. In this case, the data synchronization task can be split into multiple data synchronization subtasks. Specifically, the abstract syntax tree of the data synchronization task can be obtained. The abstract syntax tree can be a tree-like data structure generated by the compiler when parsing the source code of the data synchronization task, used to represent the abstract syntax structure of the source code. Compared with the source code, the abstract syntax tree is easier to analyze and operate. Therefore, by parsing the abstract syntax tree of the data synchronization task, the file partition list corresponding to the data table can be identified. Further, combined with the data address information corresponding to the data table included in the abstract syntax tree, the data address information can be used to represent the storage address of the data corresponding to the data table, and the file partition paths of multiple file partitions corresponding to the file partition list are traversed and queried. The file partition path can be used to represent the data storage path of the corresponding file partition. For example, the file partition paths corresponding to the file partition list can be:

[0062]

[0063] The file partition list can contain multiple file partitions. The file partitions divide the data of the data table in the form of partitions. The file partition paths of the multiple file partitions can be respectively:

[0064]

[0065] Step S303: Split the data synchronization task into multiple data synchronization subtasks respectively corresponding to multiple file partition paths.

[0066] Among them, the data synchronization task can be split based on the multiple file partition paths. In a possible implementation, the data synchronization task can be split into multiple data synchronization subtasks respectively corresponding to the multiple file partition paths, that is, the data corresponding to each file partition belongs to the same data synchronization subtask. For example, if the first file partition among the multiple file partitions corresponds to the first data synchronization subtask among the multiple data synchronization subtasks, then the first data synchronization subtask can be used to synchronize the synchronization data of the first file partition included in the target database to the database to be data synchronized; if the second file partition among the multiple file partitions corresponds to the second data synchronization subtask among the multiple data synchronization subtasks, then the second data synchronization subtask can be used to synchronize the synchronization data of the second file partition included in the target database to the database to be data synchronized. For example, the multiple data synchronization subtasks can be:

[0067]

[0068] Step S304, execute multiple data synchronization subtasks in parallel.

[0069] Among them, the multiple data synchronization subtasks can be assigned to different processing nodes of the data warehouse analysis system Hive for execution, and the multiple processing nodes execute the multiple data synchronization subtasks simultaneously. Specifically, the synchronization data of the file partition paths corresponding to the multiple data synchronization subtasks can be read in parallel in the target database, and the synchronization data corresponding to the multiple file partition paths is sent to the database server to be data synchronized, and the database server to be data synchronized can import the synchronization data into the database to be data synchronized.

[0070] The method of this embodiment obtains a data synchronization task and determines the data volume of the data table corresponding to the data synchronization task. When the data volume exceeds the preset data volume, the multiple file partition paths of the data table are determined, and then the data synchronization task is split into multiple data synchronization subtasks respectively corresponding to the multiple file partition paths, and the multiple data synchronization subtasks are executed in parallel. This solution can divide the data synchronization task into multiple data synchronization subtasks corresponding to each file partition path according to the data volume of the data table corresponding to the data synchronization task, and perform data synchronization by executing the multiple data synchronization subtasks in parallel, thereby improving the efficiency of data synchronization.

[0071] In one embodiment, as Figure 4 shown, determining the multiple file partition paths of the data table in step S302 may include the following steps:

[0072] Step S401, obtain the abstract syntax tree corresponding to the data synchronization task.

[0073] Step S402: Based on the abstract syntax tree, obtain multiple file partition paths corresponding to the data table.

[0074] Among them, based on the abstract syntax tree corresponding to the data synchronization task (for example, an SQL query task), the data synchronization task can be split to obtain multiple data synchronization subtasks. Specifically, the abstract syntax tree of the data synchronization task can be obtained. The abstract syntax tree is a tree-like data structure generated by the compiler when parsing the source code of the data synchronization task, and is used to represent the abstract syntax structure of the source code. Compared with the source code, the abstract syntax tree is easier to analyze and operate. Therefore, by parsing the abstract syntax tree of the data synchronization task, the file partition list corresponding to the data table can be identified. Further, in combination with the data address information of the data table included in the abstract syntax tree, the data address information can be used to represent the storage address of the data corresponding to the data table, and the file partition paths of multiple file partitions corresponding to the file partition list are traversed and queried. The file partition path can be used to represent the data storage path of the corresponding file partition. For example, the file partition paths corresponding to the file partition list can be:

[0075]

[0076] The file partition list can include multiple file partitions. The file partition divides the data of the data table in the form of partitions. The file partition paths of the multiple file partitions can be respectively:

[0077]

[0078] In one embodiment, step S402 may include:

[0079] Identify the data address information of the data table included in the abstract syntax tree; based on the data address information, obtain multiple file partition paths corresponding to the data table.

[0080] The abstract syntax tree is a tree-like data structure generated by the compiler when parsing the source code of the data synchronization task. The abstract syntax tree includes a field for representing the storage address of the data of the data table, that is, the data address information. Based on the data address information, the file partition paths respectively corresponding to the multiple file partitions of the data table can be obtained. The file partition path can be used to represent the data storage path of the corresponding file partition.

[0081] The method of this embodiment can, based on the data address information included in the abstract syntax tree of the data table, and further, based on the data address information, obtain multiple file partition paths corresponding to the data table. Further, based on the multiple file partition paths, the data synchronization task can be split, which can improve the efficiency of data synchronization.

[0082] In one embodiment, determining the data volume of the data table corresponding to the data synchronization task in step S301 may include:

[0083] Determining the number of rows in the data table; determining the data volume of the data table based on the number of rows.

[0084] Among them, the data volume of the data table of the data synchronization task can be measured by the number of rows in the data table. First, a preset number of rows can be set, for example, 1 million. Furthermore, the number of rows in the data table can be identified, and the number of rows in the data table is compared with the preset number of rows. In the case where the number of rows in the data table exceeds the preset number of rows, it can be determined that the data volume of the data table exceeds the preset data volume; in the case where the number of rows in the data table does not exceed the preset number of rows, it can be determined that the data volume of the data table does not exceed the preset data volume.

[0085] The method of this embodiment can determine the data volume of the data table according to the number of rows in the data table. Further, according to the data volume of the data table, different task execution methods can be determined, which can improve the efficiency of data synchronization.

[0086] In one embodiment, determining the number of rows in the data table may include:

[0087] Executing a query plan execution statement for the data synchronization task to obtain the number of rows in the data table.

[0088] Among them, the query plan execution statement may be:

[0089]

[0090] Executing the query plan execution statement for the data synchronization task to obtain the number of rows in the data table.

[0091] In one embodiment, the method of the embodiment of the present application further includes the following steps for providing a data synchronization method when the data volume of the data table does not exceed the preset data volume. The specific steps may include:

[0092] When the data volume does not exceed the preset data volume, execute the data synchronization task.

[0093] Among them, generally, it is possible to determine whether to perform splitting processing on the data synchronization task based on the data volume of the data table corresponding to the data synchronization task. That is, for the task execution method of the data synchronization task, which can also be understood as the data extraction method, when the data volume of the data table exceeds the preset data volume, the data synchronization task can be split to obtain multiple data synchronization subtasks; when the data volume of the data table does not exceed the preset data volume, the data synchronization task can be directly executed. Specifically, the synchronization data of the data table corresponding to the data synchronization task can be read in the target database, and the synchronization data can be sent to the server to be data synchronized, and the server to be data synchronized can import the synchronization data into the corresponding database to be data synchronized.

[0094] The method of this embodiment can determine the data volume of the data table according to the number of rows of the data table. Further, according to the data volume of the data table, different task execution methods can be determined, which can improve the efficiency of data synchronization.

[0095] In another embodiment, as Figure 5 shown, a data synchronization method is provided, and the method may include:

[0096] Step S501, obtain a data synchronization task, and determine the data volume of the data table corresponding to the data synchronization task.

[0097] Step S502, determine whether the data volume of the data table exceeds the preset data volume.

[0098] Among them, when the data volume of the data table does not exceed the preset data volume, step S503 is executed;

[0099] When the data volume of the data table exceeds the preset data volume, step S504 is executed.

[0100] Step S503, execute the data synchronization task.

[0101] Step S504, obtain the abstract syntax tree corresponding to the data synchronization task.

[0102] Step S505, based on the abstract syntax tree, obtain multiple file partition paths corresponding to the data table.

[0103] Among them, the data synchronization task (for example, an SQL query task) can be segmented based on the corresponding abstract syntax tree to obtain multiple data synchronization subtasks. Specifically, the abstract syntax tree of the data synchronization task can be obtained. The abstract syntax tree is a tree-like data structure generated by a compiler when parsing the source code of the data synchronization task, used to represent the abstract syntax structure of the source code. Compared with the source code, the abstract syntax tree is easier to analyze and operate. Therefore, by parsing the abstract syntax tree of the data synchronization task, the list of file partitions corresponding to the data table can be identified. Further, in combination with the data address information corresponding to the data table included in the abstract syntax tree, the data address information can be used to represent the storage address of the data corresponding to the data table, traverse and query the file partition paths of multiple file partitions corresponding to the list of file partitions. The file partition path can be used to represent the data storage path of the corresponding file partition. For example, the file partition paths corresponding to the list of file partitions can be:

[0104]

[0105] The list of file partitions can include multiple file partitions. The file partition is to divide the data of the data table in the form of partitions. The file partition paths of the multiple file partitions can be respectively:

[0106]

[0107] Step S506: Segment the data synchronization task into multiple data synchronization subtasks respectively corresponding to multiple file partition paths.

[0108] Among them, the data synchronization task can be segmented based on the multiple file partition paths. In a possible implementation, the data synchronization task can be segmented into multiple data synchronization subtasks respectively corresponding to multiple file partition paths, that is, the data corresponding to each file partition is attributed to the same data synchronization subtask. For example, if the first file partition among the multiple file partitions corresponds to the first data synchronization subtask among the multiple data synchronization subtasks, then the first data synchronization subtask can be used to synchronize the synchronization data of the first file partition included in the target database to the database to be data synchronized; if the second file partition among the multiple file partitions corresponds to the second data synchronization subtask among the multiple data synchronization subtasks, then the second data synchronization subtask can be used to synchronize the synchronization data of the second file partition included in the target database to the database to be data synchronized. For example, the multiple data synchronization subtasks can be:

[0109]

[0110] Step S507: Execute multiple data synchronization subtasks in parallel.

[0111] The method of this embodiment obtains a data synchronization task and determines the data volume of the data table corresponding to the data synchronization task, determines multiple file partition paths of the data table when the data volume exceeds the preset data volume, and then divides the data synchronization task into multiple data synchronization subtasks corresponding to the multiple file partition paths respectively, and executes the multiple data synchronization subtasks in parallel. This solution can divide the data table corresponding to the data synchronization task into multiple data synchronization subtasks corresponding to each file partition path according to the data volume, and synchronizes data by executing the multiple data synchronization subtasks in parallel, thereby improving the efficiency of data synchronization.

[0112] It should be understood that, although the various steps in the flowcharts involved in the above-mentioned embodiments are displayed in sequence according to the indication of the arrows, these steps are not necessarily executed in sequence according to the order indicated by the arrows. Unless there is a clear explanation in this article, the execution of these steps does not have a strict order restriction, and these steps can be executed in other orders. Moreover, at least a part of the steps in the flowcharts involved in the above-mentioned embodiments can include multiple steps or multiple stages, and these steps or stages are not necessarily executed at the same time, but can be executed at different times, and the execution order of these steps or stages is not necessarily to be carried out in sequence, but can be executed in turn or alternately with other steps or at least a part of the steps or stages in other steps.

[0113] Based on the same inventive concept, the embodiment of the present application also provides a data synchronization device for implementing the data synchronization method involved above. The implementation solution provided by the device to solve the problem is similar to the implementation solution recorded in the above method, so the specific limitations in one or more data synchronization device embodiments provided below can refer to the limitations on the data synchronization method above, and will not be repeated here.

[0114] In one embodiment, Figure 6 As shown, a data synchronization device is provided, including: an acquisition module 601, a determination module 602, a task segmentation module 603 and a task execution module 604, wherein:

[0115] An acquisition module 601 is used to acquire a data synchronization task and determine the amount of data in a data table corresponding to the data synchronization task;

[0116] A determination module 602, configured to determine a plurality of file partition paths of the data table when the data amount exceeds a preset data amount;

[0117] A task division module 603, used for dividing the data synchronization task into a plurality of data synchronization subtasks corresponding to the plurality of file partition paths respectively;

[0118] The task execution module 604 is used to execute the multiple data synchronization subtasks in parallel.

[0119] In addition, the determination module 602 is further used to: obtain the abstract syntax tree corresponding to the data synchronization task; based on the abstract syntax tree, obtain multiple file partition paths corresponding to the data table.

[0120] The determination module 602 is further used to: identify the data address information of the data table included in the abstract syntax tree; based on the data address information, obtain multiple file partition paths corresponding to the data table.

[0121] The acquisition module 601 is further used to: determine the number of rows of the data table; determine the data volume of the data table according to the number of rows of the table.

[0122] The acquisition module 601 is further used to: execute a query plan execution statement on the data synchronization task to obtain the number of rows of the data table.

[0123] The determination module 602 is further used to: execute the data synchronization task when the data volume does not exceed a preset data volume.

[0124] Each module in the above data synchronization device can be implemented in whole or in part by software, hardware, and their combination. The above modules can be embedded in the processor of the computer device in hardware form or independent of it, or stored in the memory of the computer device in software form, so that the processor can call and execute the operations corresponding to the above modules.

[0125] In one embodiment, a computer device is provided. The computer device can be a server, and its internal structure diagram can be as Figure 7 shown. The computer device includes a processor, a memory, and a network interface connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store data synchronization-related data. The network interface of the computer device is used to communicate with an external terminal through a network connection. The computer program, when executed by the processor, implements a data synchronization method.

[0126] Those skilled in the art can understand, Figure 7The structure shown is only a block diagram of some structures related to the solution of this application, and does not constitute a limitation on the computer device to which the solution of this application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements.

[0127] In one embodiment, a computer device is further provided, including a memory and a processor. A computer program is stored in the memory, and when the processor executes the computer program, the steps in the above method embodiments are implemented.

[0128] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored, and when the computer program is executed by a processor, the steps in the above method embodiments are implemented.

[0129] In one embodiment, a computer program product is provided, including a computer program, and when the computer program is executed by a processor, the steps in the above method embodiments are implemented.

[0130] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in this application are all information and data that have been authorized by the user or fully authorized by all parties.

[0131] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, database, or other medium used in the embodiments provided in the present application can include at least one of non-volatile and volatile memories. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetoresistive random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc. The databases involved in the embodiments provided in the present application can include at least one of relational databases and non-relational databases. Non-relational databases can include distributed databases based on blockchain, etc., without limitation. The processors involved in the embodiments provided in the present application can be general-purpose processors, central processors, graphics processors, digital signal processors, programmable logic devices, data processing logics based on quantum computing, etc., without limitation.

[0132] The technical features of the above embodiments can be combined arbitrarily. For the sake of concise description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope recorded in this specification.

[0133] The above-described embodiments merely represent several implementation manners of the present application. Their descriptions are relatively specific and detailed, but they should not be construed as limiting the patent scope of the present application. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all belong to the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the appended claims.

Claims

1. A data synchronization method, characterized in that, The method includes: Obtain a data synchronization task and determine the data volume of the data table corresponding to the data synchronization task; In the case where the data volume exceeds a preset data volume, determine multiple file partition paths of the data table; Split the data synchronization task into multiple data synchronization subtasks respectively corresponding to the multiple file partition paths; Execute the multiple data synchronization subtasks in parallel.

2. The method according to claim 1, characterized in that, The determining of the multiple file partition paths of the data table includes: Obtain the abstract syntax tree corresponding to the data synchronization task; Based on the abstract syntax tree, obtain multiple file partition paths corresponding to the data table.

3. The method according to claim 2, characterized in that, The obtaining of the multiple file partition paths corresponding to the data table based on the abstract syntax tree includes: Identify the data address information of the data table included in the abstract syntax tree; Based on the data address information, obtain multiple file partition paths corresponding to the data table.

4. The method according to claim 1, characterized in that, The determining of the data volume of the data table corresponding to the data synchronization task includes: Determine the number of rows in the data table; Determine the data volume of the data table according to the number of rows.

5. The method according to claim 4, characterized in that, The determining of the number of rows in the data table includes: Execute a query plan execution statement on the data synchronization task to obtain the number of rows in the data table.

6. The method according to any one of claims 1-5, characterized in that, The method further includes: In the case where the data volume does not exceed the preset data volume, execute the data synchronization task.

7. A data synchronization device, characterized in that, The apparatus includes: An obtaining module, configured to obtain a data synchronization task and determine the data volume of the data table corresponding to the data synchronization task; A determining module, configured to determine multiple file partition paths of the data table in the case where the data volume exceeds a preset data volume; A task splitting module, configured to split the data synchronization task into multiple data synchronization subtasks respectively corresponding to the multiple file partition paths; A task execution module, configured to execute the multiple data synchronization subtasks in parallel.

8. A computer device, comprising a memory and a processor, the memory storing a computer program, characterized in that, When the processor executes the computer program, the steps of the method according to any one of claims 1-6 are implemented.

9. A computer-readable storage medium, having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, the steps of the method according to any one of claims 1-6 are implemented.

10. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, the steps of the method according to any one of claims 1-6 are implemented.