Data synchronization method and device, equipment and storage medium

By reading and synchronizing data chunking according to the preset chunk size during the data synchronization process, the problem of low data synchronization efficiency in the existing technology is solved, efficient and stable data synchronization is achieved, user experience is improved and breakpoint continuous transmission is supported.

CN119938788APending Publication Date: 2025-05-06BEIJING BAIDU NETCOM SCI & TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411999398.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-31
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

When the prior art synchronizes the data of a relational database to the target end, there is a problem of low data synchronization efficiency, especially when processing large tables, performance degradation, high resource usage and low query efficiency.

Method used

Responsive to the data synchronization instruction for the target data table in the source relational database, the data in the target data table is read according to the preset chunk size, and the data blocks are obtained, and the data blocks are synchronized from the source relational database to the target end.

Benefits of technology

Automatic splitting of target data tables is achieved, reducing the amount of data obtained by each query statement, thereby improving the overall efficiency and fault tolerance of data synchronization, improving user experience, and supporting breakpoint continuous transmission.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119938788A_ABST
    Figure CN119938788A_ABST
Patent Text Reader

Abstract

The invention provides a data synchronization method and device, equipment and a storage medium, and relates to the field of big data, in particular to the technical field of databases, computers, data processing and the like. The data synchronization method comprises the following steps: in response to a data synchronization instruction for a target data table in a source relational database, reading data in the target data table according to a preset block size to obtain data blocks; and synchronizing the data blocks from the source relational database to a target end. According to the data synchronization method and device, the target data table can be automatically split, it is guaranteed that the query statement corresponding to each data block obtains a small amount of data, and therefore the overall efficiency and fault tolerance of data synchronization can be effectively improved, and the user experience is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to technical fields such as database, computer and data processing in the field of big data, and in particular to a data synchronization method, device, equipment and storage medium. Background Art

[0002] In the information age, large amounts of data are generated every day and stored in different systems, such as relational databases. After the original data is stored, it is necessary to synchronize the large amounts of data to other systems or engines for further development or analysis, thereby mining the value of the data.

[0003] Currently, for relational databases, the method of reading all table data from the relational database at one time is usually adopted to synchronize data from the relational database to the target end, such as a data warehouse or a data lake. However, the data synchronization through the above method has the problem of low data synchronization efficiency. Summary of the invention

[0004] The present disclosure provides a data synchronization method, apparatus, device and storage medium for improving data synchronization efficiency.

[0005] According to a first aspect of the present disclosure, a data synchronization method is provided, comprising:

[0006] In response to a data synchronization instruction for a target data table in a source relational database, data in the target data table is read according to a preset block size to obtain data blocks;

[0007] Synchronize data chunks from the source relational database to the target.

[0008] According to a second aspect of the present disclosure, there is provided a data synchronization device, comprising:

[0009] A reading unit, configured to respond to a data synchronization instruction for a target data table in a source relational database, read the data in the target data table according to a preset block size, and obtain data blocks;

[0010] The synchronization unit is used to synchronize data blocks from the source relational database to the target end.

[0011] According to a third aspect of the present disclosure, an electronic device is provided, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the data synchronization method described in the first aspect.

[0012] According to a fourth aspect of the present disclosure, a non-transitory computer-readable storage medium storing computer instructions is provided, wherein the computer instructions are used to enable the computer to execute the data synchronization method described in the first aspect.

[0013] According to a fifth aspect of the present disclosure, a computer program product is provided, comprising: a computer program, wherein the computer program is stored in a readable storage medium, at least one processor of an electronic device can read the computer program from the readable storage medium, and the at least one processor executes the computer program so that the electronic device executes the data synchronization method described in the first aspect.

[0014] The technology disclosed in the present invention solves the problem of low data synchronization efficiency when performing data synchronization in the current manner. The present invention responds to a data synchronization instruction for a target data table in a source relational database, reads the data in the target data table according to a preset block size, and obtains data blocks; synchronizes the data blocks from the source relational database to the target end, realizes automatic splitting of the target data table, reads the data in the target data table in sequence in the form of data blocks, and synchronizes the data blocks from the source relational database to the target end, which can ensure that the query statement corresponding to each data block obtains a smaller amount of data, thereby effectively improving the overall efficiency and fault tolerance of data synchronization, improving user experience, and further helping to improve the product strength of data synchronization services.

[0015] It should be understood that the content described in this section is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it intended to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] The accompanying drawings are used to better understand the present solution and do not constitute a limitation of the present disclosure.

[0017] Figure 1 is a schematic diagram of an application scenario to which the data synchronization method disclosed herein is applicable;

[0018] Figure 2 is a schematic diagram according to a first embodiment of the present disclosure;

[0019] Figure 3 is a schematic diagram according to a second embodiment of the present disclosure;

[0020] Figure 4 is a schematic diagram according to a third embodiment of the present disclosure;

[0021] Figure 5 is a schematic diagram according to a fourth embodiment of the present disclosure;

[0022] Figure 6FIG. 6 is a schematic block diagram of an example electronic device 600 that may be used to implement embodiments of the present disclosure. DETAILED DESCRIPTION

[0023] The following is a description of exemplary embodiments of the present disclosure in conjunction with the accompanying drawings, including various details of the embodiments of the present disclosure to facilitate understanding, which should be considered as merely exemplary. Therefore, it should be recognized by those of ordinary skill in the art that various changes and modifications may be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, for the sake of clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.

[0024] The present disclosure provides a data synchronization method, device, equipment and storage medium, which are applied to technical fields such as database, computer and data processing in the field of big data to achieve the purpose of improving data synchronization efficiency.

[0025] In the technical solution of the present disclosure, the collection, storage, use, processing, transmission, provision and disclosure of user personal information involved are in compliance with the provisions of relevant laws and regulations and do not violate public order and good morals.

[0026] In the information age, large amounts of data are generated every day and stored in different systems, such as relational databases. After the original data is stored, it is necessary to synchronize the large amounts of data to other systems or engines for further development or analysis, thereby mining the value of the data.

[0027] At present, for relational databases, the method of reading the entire table data from the relational database at one time is usually adopted to synchronize the data from the relational database to the target end, such as a data warehouse or a data lake. However, when synchronizing data through the above method, it is difficult to avoid full table scanning in some cases by reading the entire table data at one time or by index query, resulting in low data synchronization efficiency. In addition, the commonly used relational database synchronization methods, such as JDBC (an application programming interface for database access), have problems such as performance degradation, high resource usage and low query efficiency when processing large tables.

[0028] In addition, in the related technologies, the entire table data can be read by default for data synchronization, but breakpoint resumption is not supported. Some database reading plug-ins support users to manually specify fields to split a single table; or, breakpoint resumption is not supported, but tables can be split according to user-specified fields for parallel reading; or, users can be supported to specify fields to split a single table to perform synchronization tasks in parallel, or users can be supported to specify partition reading to split the table. The above-mentioned related technologies mainly have the following problems: (1) Reading the entire table data at one time may face problems such as memory limitations and network transmission pressure; for databases with frequent incremental data, synchronization failure may also occur due to problems such as the rollback segment being too old; in addition, if data synchronization supports breakpoint resumption synchronization, query operations such as primary key sorting may be introduced, which may lead to problems such as insufficient temporary database space; (2) Partitions are isolated from splitting tables according to fields, and are not considered together to reduce the amount of data in a single query; (3) Users are required to manually specify field names, which increases the complexity of the operation and requires users to have certain professional knowledge; (4) Breakpoint resumption is not supported. Therefore, there is an urgent need for an effective data synchronization solution for large-scale data in relational databases.

[0029] In order to solve the above problems, the present disclosure provides a data synchronization method. For the target data table in the source relational database, the data in the target data table is read according to the preset block size to obtain data blocks, and the data blocks are synchronized from the source relational database to the target end to realize automatic splitting of the target data table, so as to ensure that the query statement corresponding to each data block obtains a smaller amount of data, thereby effectively improving the overall efficiency and fault tolerance of data synchronization and improving user experience. In addition, during the data synchronization process, the location information corresponding to the synchronized data can be recorded for breakpoint resumption when an interruption occurs during the data synchronization process.

[0030] It can be seen that the data synchronization method provided by the present disclosure is user-imperceptible and does not require user intervention. It is a pure backend optimization and has the advantages of table splitting ease of use without user intervention, strong stability for large tables, and default support for breakpoint continuation. The data synchronization method provided by the present disclosure can be applied to synchronize data in a relational database to a target end offline.

[0031] Figure 1It is a schematic diagram of an application scenario to which the data synchronization method of the present disclosure is applicable. In this application scenario, the user to which the source relational database belongs triggers a data synchronization instruction for the target data table in the source relational database through the client 101. The user may be, for example, an enterprise user, an enterprise user on a public cloud, or an individual user. Accordingly, the server 102 responds to the data synchronization instruction and synchronizes the data of the target data table in the source relational database on the server 103 to the target end on the server 104 according to the data synchronization method provided by the present disclosure. The source relational database may include, for example, MySQL and Oracle, and the target end may include, for example, data warehouses or data lakes such as Hive and Doris.

[0032] It should be noted that Figure 1 This is only a schematic diagram of an application scenario provided by the embodiment of the present disclosure. Figure 1 does not limit the equipment included in Figure 1 The positional relationship between the devices is limited.

[0033] The following specific embodiments are used to describe in detail the technical solution of the present invention and how the technical solution of the present invention solves the above-mentioned technical problems. The following specific embodiments can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments. The embodiments of the present invention will be described below in conjunction with the accompanying drawings.

[0034] Figure 2 Schematic diagram of the first embodiment of the present disclosure. Figure 2 As shown, the data synchronization method provided in the first embodiment of the present disclosure can be applied to an electronic device, which can be a server or a server cluster. The data synchronization method provided in the first embodiment of the present disclosure includes:

[0035] S201 . In response to a data synchronization instruction for a target data table in a source relational database, read data in the target data table according to a preset block size to obtain data blocks.

[0036] In the disclosed embodiment, the source relational database is, for example, MySQL or Oracle. The data synchronization instruction may be input by the user of the source relational database to the electronic device executing the embodiment of the method, or may be sent by other devices to the electronic device executing the embodiment of the method. The preset block size may be set as needed, for example, 10,000 data items are used as one data block. Exemplarily, if a user wants to synchronize the data of the target data table in his source relational database to the target end, he may create a data synchronization task through the front-end page of the client, and specify the data source end and the target end in the data synchronization task. Among them, the data source end is the user's own source relational database, and the target end is the target end of the data to be synchronized. After the user creates the data synchronization task, the data synchronization instruction for the target data table in the source relational database is triggered. Accordingly, the electronic device executing the embodiment of the method receives the data synchronization instruction for the target data table in the source relational database. Assuming that the data in the target data table in the source relational database is 14,500 items, and the preset block size is, for example, 10,000 items, the electronic device executing the embodiment of the method reads 10,000 items of data in the target data table to obtain data blocks, so as to synchronize the data blocks from the source relational database to the target end. The query condition in the structured query language (SQL) statement for reading the data block adds the judgment of the lower limit and upper limit of the data read in the target data table. Specifically, for example, in the where statement (used to specify the filter condition) of the SQL statement, the judgment of greater than or equal to the lower limit and less than or equal to the upper limit is added to accurately obtain the size of the data block.

[0037] S202: Synchronize data blocks from the source relational database to the target end.

[0038] In this step, the target end is, for example, a data warehouse or a data lake. After obtaining the data blocks, the data blocks can be synchronized from the source relational database to the target end. Exemplarily, assuming that the data in the target data table in the source relational database is 14,500, and the preset block size is 10,000 data, then first obtain a data block of 10,000 data, and synchronize the data block from the source relational database to the target end; then obtain a data block of 4,500 data, and synchronize the data block from the source relational database to the target end, thereby completing the data synchronization of the target data table.

[0039] It can be understood that for the data in the target data table in the source relational database, the data in the target data table can be synchronized from the source relational database to the target end by executing steps S201 and S202. For each target data table, the data in the target data table is read sequentially in the form of data blocks, and the data blocks are synchronized serially from the source relational database to the target end, which can ensure that the query statement corresponding to each data block obtains a small amount of data, thereby improving the overall efficiency and fault tolerance of data synchronization.

[0040] Optionally, the data of each target data table may be synchronized from the source relational database to the target end by executing steps S201 and S202 in parallel between multiple target data tables.

[0041] In the disclosed embodiment, in response to a data synchronization instruction for a target data table in a source relational database, data in the target data table is read according to a preset block size to obtain data blocks; the data blocks are synchronized from the source relational database to the target end to automatically split the target data table, and the data in the target data table is read sequentially in the form of data blocks, and the data blocks are synchronized from the source relational database to the target end, which can ensure that the query statement corresponding to each data block obtains a smaller amount of data, thereby effectively improving the overall efficiency and fault tolerance of data synchronization, improving user experience, and further helping to improve the product strength of the data synchronization service.

[0042] On the basis of the above embodiments, the data synchronization method provided by the embodiment of the present disclosure can perform full data synchronization on the data in the source relational database, and can also read incremental data by specifying a where condition. For example, if the target data table contains a time field (time), you can specify "time> target start time" in the where condition to read incremental data, so that the incremental data read can be synchronized by the data synchronization method provided by the embodiment of the present disclosure.

[0043] On the basis of the above embodiment, optionally, the step S201 of reading the data in the target data table according to the preset block size to obtain the data blocks may further include: determining whether the target data table is stored in partitions; if the target data table is stored in partitions, then for each partition, reading the data in the corresponding partition of the target data table according to the preset block size to obtain the data blocks; if the target data table is not stored in partitions, then reading the data in the target data table according to the preset block size to obtain the data blocks.

[0044] Exemplarily, the data of the target data table can be stored in different partitions. In this embodiment, the partitions of the target data table are automatically queried to determine whether the target data table is partitioned and stored without the intervention of user-specified methods. If the target data table is partitioned and stored, the data in the corresponding partition of the target data table can be read according to the preset block size for each partition to obtain data blocks. Optionally, the partitions can also be sorted according to the partition names of the string type to obtain sorted partitions, so that the data in the corresponding partition of the target data table can be read according to the preset block size for each sorted partition to obtain data blocks. If the target data table is not partitioned and stored, the entire target data table can be taken as a single partition, and the data in the target data table can be read according to the preset block size to obtain data blocks. After obtaining the data blocks, the data blocks can be synchronized from the source relational database to the target end.

[0045] Figure 3 FIG. 1 is a schematic diagram of a second embodiment of the present disclosure. Based on the above embodiment, the present disclosure further describes the data synchronization method. Figure 3 As shown, the data synchronization method provided by the second embodiment of the present disclosure may include:

[0046] S301 . Responding to a data synchronization instruction for a target data table in a source relational database, obtaining a target data table.

[0047] Exemplarily, the data synchronization instruction for the target data table in the source relational database is triggered by a user to which the source relational database belongs. Accordingly, the electronic device executing the embodiment of the method receives the data synchronization instruction and obtains the target data table.

[0048] S302: Determine whether the target data table is stored in partitions.

[0049] Exemplarily, the partition of the target data table is automatically queried to determine whether the target data table is stored in partitions.

[0050] If the target data table is not stored in partitions, then execute S303, determine that the target data table is a single partition, and continue to execute step S305; if the target data table is stored in partitions, then execute step S304.

[0051] S304: Sort the partitions to obtain sorted partitions.

[0052] In this step, after determining the target data table partition storage, the partitions of the target data table can be sorted to obtain sorted partitions. Exemplarily, for example, the partitions can be sorted according to the partition names of the string type to obtain sorted partitions, so that step S305 can be performed for each sorted partition to read the data in the corresponding partition of the target data table according to the preset block size to obtain data blocks.

[0053] S305: Determine whether the target data table has a primary key or a unique index.

[0054] In this step, it can be automatically determined whether the target data table has a primary key or a unique index without the intervention of user specification or the like.

[0055] Further, optionally, determining whether the target data table has a primary key or a unique index may include: determining whether the target data table has a primary key or a unique index through a preset query method; and / or determining whether the target data table has a primary key or a unique index based on the configured target field.

[0056] Exemplarily, the preset query method can also be understood as a primary key detection method, and the preset query method is, for example, a preset query SQL statement, which is used to automatically query whether the target data table has a primary key, or automatically query whether the target data table has a unique index, without user intervention. Alternatively, it is possible to determine whether the target data table has a primary key based on a pre-configured target field, such as the primary key of the target data table; or to determine whether the target data table has a unique index based on the target field, such as the unique index of the target data table.

[0057] If it is determined that the target data table has a primary key or a unique index, steps S306 to S308 are executed; if it is determined that the target data table has no primary key and no unique index, step S309 is executed.

[0058] S306. If the target data table has a primary key, the data in the corresponding partition of the target data table is sorted according to the primary key to obtain the sorted data corresponding to the partition; or, if the target data table has a unique index, the data in the corresponding partition of the target data table is sorted according to the unique index to obtain the sorted data corresponding to the partition.

[0059] Exemplarily, when the target data table has a primary key, the data in the corresponding partition of the target data table can be sorted according to the primary key to obtain sorted data corresponding to the partition. Alternatively, when the target data table has a unique index, the data in the corresponding partition of the target data table can be sorted according to the unique index to obtain sorted data corresponding to the partition. It can be understood that the method of sorting the data in the corresponding partition of the target data table can be sorting according to the primary key or according to the unique index, either of which can be selected.

[0060] S307: Read the sorted data according to the preset block size to obtain data blocks.

[0061] Exemplarily, the preset block size, for example, takes 10,000 data as one data block, and the sorted data can be read according to the preset block size to obtain data blocks, so as to realize data reading according to the combination of partitions and blocks, and ensure that the query statement corresponding to each data block obtains a smaller amount of data, thereby improving the overall efficiency and fault tolerance of data synchronization. Among them, each time the sorted data is read according to the preset block size, the read sorted data is divided by the boundary query method, the upper limit of the boundary is continuously obtained, and the lower limit and upper limit of the data reading in the target data table are added to the query conditions. Specifically, for example, a judgment greater than or equal to the lower limit and less than or equal to the upper limit is added to the where statement of the SQL statement to accurately obtain the size of the data block.

[0062] S308: Synchronize the data blocks from the source relational database to the target end.

[0063] The implementation principle and technical effect of S308 can be referred to the above Figure 2 The embodiment of step S202 is not described in detail.

[0064] S309: Synchronize the data in the partition from the source relational database to the target end.

[0065] It can be understood that after determining that the target data table has no primary key and no unique index, there is no need to split the data in the partition, and all data in the partition can be synchronized directly from the source relational database to the target end in a partitioned manner.

[0066] In the disclosed embodiment, in response to a data synchronization instruction for a target data table in a source relational database, the target data table is obtained, and it is determined whether the target data table is stored in partitions; in the case where the target data table is stored in partitions, the partitions are sorted to obtain sorted partitions; it is determined whether the target data table has a primary key or a unique index, and in the case where it is determined that the target data table has a primary key or a unique index, the data in the corresponding partition of the target data table is sorted according to the primary key or the unique index to obtain sorted data corresponding to the partition, and the sorted data is read according to a preset block size to obtain data blocks; the data blocks are synchronized from the source relational database to the target end, and sequential reading and synchronization of data according to a combination of partitions and blocks are realized, which can ensure that the query statement corresponding to each data block obtains a smaller amount of data, thereby improving the overall efficiency and fault tolerance of data synchronization, improving user experience, and further helping to improve the product strength of data synchronization services.

[0067] Considering that data synchronization in blocks may be interrupted, Figure 4 is a schematic diagram according to the third embodiment of the present disclosure. Figure 4As shown, the data synchronization method provided by the third embodiment of the present disclosure may include:

[0068] S401 . In response to a data synchronization instruction for a target data table in a source relational database, obtain a target data table.

[0069] S402: Determine whether the target data table is stored in partitions.

[0070] If the target data table is not stored in partitions, then execute S403, determine that the target data table is a single partition, and continue to execute S405 and steps after S405; if the target data table is stored in partitions, then execute step S404.

[0071] S404: Sort the partitions to obtain sorted partitions.

[0072] For each sorted partition, execute S405 and steps after S405.

[0073] S405: Determine whether the target data table has a primary key or a unique index.

[0074] S406. If the target data table has a primary key, the data in the corresponding partition of the target data table is sorted according to the primary key to obtain the sorted data corresponding to the partition; or, if the target data table has a unique index, the data in the corresponding partition of the target data table is sorted according to the unique index to obtain the sorted data corresponding to the partition.

[0075] S407: Read the sorted data according to the preset block size to obtain data blocks.

[0076] S408: Synchronize the data blocks from the source relational database to the target end.

[0077] The implementation principles and technical effects of S401 to S408 can be referred to in the above Figure 3 The embodiments of the related steps in will not be repeated here.

[0078] S409. During the data synchronization process, the location information corresponding to the synchronized data is recorded. The location information includes the partition identifier of the partition to which the synchronized data belongs, the block identifier of the data block to which the synchronized data belongs, and the target primary key or target unique index corresponding to the synchronized data.

[0079] Exemplarily, the partition identifier of the partition to which the synchronized data belongs is, for example, the partition name, and the block identifier of the data block to which the synchronized data belongs is, for example, the block name. Assume that the target data table is used to store student information, wherein the student number is the primary key of the target data table, and the data of the target data table is stored in three partitions, and the partition names of the three partitions are A, B, and C respectively; partition A can be split into three data blocks for reading, and the block names of the three data blocks are chunk1, chunk2, and chunk3 respectively. Then, during the data synchronization process, the location information corresponding to the synchronized data is recorded, and the location information is, for example: A-chunk1-target student number, chunk1 contains information of students with student numbers from 1 to 20, and the target student number is specifically, for example, 10, indicating the information corresponding to the student with student number 10. The location information is recorded by combining the partition identifier, the block identifier, and the primary key, or by combining the partition identifier, the block identifier, and the unique index, wherein the partition identifier and the block identifier are both in a fixed order, so that the location information can be used to accurately resume the transmission from a breakpoint.

[0080] Optionally, the location information corresponding to the most recently successfully synchronized data may be recorded to save storage space.

[0081] S410: During the data synchronization process, monitor whether an interruption occurs.

[0082] In this step, the data synchronization process is monitored to promptly detect whether an interruption occurs. The interruption may be caused by the user pausing the data synchronization, or may be caused by an abnormality in the data synchronization process, which is not limited in the embodiments of the present disclosure.

[0083] It should be noted that the embodiment of the present disclosure does not limit the order in which steps S409 and S410 are executed.

[0084] S411. If an interruption occurs, when a data synchronization instruction for a target data table is detected, the target data table is resumed according to the recorded location information.

[0085] Exemplarily, after an interruption occurs, such as an interruption caused by a user pausing data synchronization, after the user starts the data synchronization task again, the electronic device executing the embodiment of the method detects a data synchronization instruction for the target data table, and according to the recorded location information, performs breakpoint-resume transmission of the target data table. Referring to the example of step S409, assuming that the recorded location information is: A-chunk1-10, when a data synchronization instruction for the target data table is detected, the target data table is resumed from A-chunk1-11.

[0086] Optionally, if an interruption occurs, a prompt message indicating the data synchronization interruption is output.

[0087] Exemplarily, when an interruption occurs during the data synchronization process, the electronic device executing the embodiment of the method may output a prompt message of the data synchronization interruption, such as outputting the prompt message to the client, and the client displays the data synchronization interruption (failure) on the front page, so that the user can handle the interruption problem in time. The prompt message of the data synchronization interruption may also be output through the log corresponding to the data synchronization.

[0088] Optionally, when it is determined that the target data table does not have a primary key and a unique index, the target data table is synchronized according to partitions. If an interruption occurs during the data synchronization process, when a data synchronization instruction for the target data table is detected, the target data table can be resumed according to the recorded location information, wherein the location information includes the partition identifier of the partition to which the synchronized data belongs.

[0089] In the disclosed embodiment, during the data synchronization process, the location information corresponding to the synchronized data is recorded, and the location information includes the partition identifier of the partition to which the synchronized data belongs, the block identifier of the data block to which the synchronized data belongs, and the target primary key or target unique index corresponding to the synchronized data; during the data synchronization process, it is monitored whether an interruption occurs; if an interruption occurs, when a data synchronization instruction for the target data table is detected, the target data table is resumed according to the recorded location information. Since the location information is recorded by combining the partition identifier, the block identifier and the primary key, or combining the partition identifier, the block identifier and the unique index, wherein the partition identifier and the block identifier are in a fixed order, the breakpoint resume can be accurately performed according to the recorded location information, so that the user experience is more efficient and more stable data synchronization service, thereby achieving efficient, convenient, stable and high fault tolerance of data synchronization, which can greatly improve the user experience.

[0090] Based on the above embodiments, the technical solution provided by the present disclosure has at least the following advantages:

[0091] (1) After testing and verification, compared with the existing data synchronization service, the data synchronization efficiency of the technical solution provided by the present disclosure is improved by approximately 20%, which greatly improves the synchronization rate, thereby helping to improve the product strength of the data synchronization service.

[0092] (2) The technical solution provided by the present disclosure supports breakpoint resume by default. The breakpoint resume mechanism is imperceptible to users. Users can create data synchronization tasks on the front-end page. If the data synchronization task fails due to various reasons, the data synchronization service will automatically record and store the latest location information on the back-end. At this time, the user clicks to start the data synchronization task again to continue synchronizing data from the latest location.

[0093] (3) The technical solution provided by the present disclosure can realize the splitting of the data table of the source relational database without the intervention of the user, and the usability and processing stability for large-scale data are significantly enhanced.

[0094] In summary, the technical solution provided by the present disclosure can achieve efficient, convenient, stable and high fault tolerance of data synchronization, and can greatly improve the user experience. The technical solution provided by the present disclosure can be applied to the data synchronization tasks of the big data platform, used to serve a large number of users, and can effectively improve the synchronization efficiency and service quality of the relational database. Among them, the technical solution provided by the present disclosure can support relational databases such as MySQL and Oracle on the data source side, and can support data warehouses or data lakes such as Hive or Doris on the target side. In terms of application, the technical solution provided by the present disclosure is user-imperceptible and does not require user intervention. It is a pure back-end optimization. Specifically, the user can create a data synchronization task on the front-end page as usual. If the data synchronization task fails due to various reasons, the data synchronization service will automatically record and store the latest site information on the back end. At this time, the user clicks to start the data synchronization task again, and the data can continue to be synchronized from the latest site according to the latest site information. Users can experience more efficient and stable data synchronization services.

[0095] The following are embodiments of the device disclosed herein, which can be used to execute the method embodiments disclosed herein. For details not disclosed in the device embodiments disclosed herein, please refer to the method embodiments disclosed herein.

[0096] Figure 5 is a schematic diagram according to the fourth embodiment of the present disclosure. Figure 5 As shown, the data synchronization device 500 provided in the fourth embodiment of the present disclosure includes: a reading unit 501 and a synchronization unit 502. Among them:

[0097] The reading unit 501 is used to respond to a data synchronization instruction for a target data table in a source relational database, read the data in the target data table according to a preset block size, and obtain data blocks.

[0098] The synchronization unit 502 is used to synchronize data blocks from the source relational database to the target end.

[0099] In some embodiments, the reading unit 501 may include: a determination module (not shown in the figure), used to determine whether the target data table is stored in partitions; a first reading module (not shown in the figure), used to read the data in the corresponding partition of the target data table according to a preset block size for each partition to obtain data blocks if the target data table is stored in partitions; a second reading module (not shown in the figure), used to read the data in the target data table according to a preset block size to obtain data blocks if the target data table is not stored in partitions.

[0100] In some embodiments, the first reading module may include: a determination submodule (not shown in the figure) for determining whether the target data table has a primary key or a unique index; a first sorting submodule (not shown in the figure) for sorting the data in the corresponding partition of the target data table according to the primary key if the target data table has a primary key, and obtaining the sorted data corresponding to the partition; or, if the target data table has a unique index, sorting the data in the corresponding partition of the target data table according to the unique index, and obtaining the sorted data corresponding to the partition; the first reading submodule (not shown in the figure) is used to read the sorted data according to a preset block size to obtain data blocks.

[0101] In some embodiments, where the target data table is stored in partitions, the first reading module may include: a second sorting submodule (not shown in the figure), used to sort the partitions to obtain sorted partitions; a second reading submodule (not shown in the figure), used to read the data in the corresponding partition of the target data table according to a preset block size for each partition in the sorted partitions to obtain data blocks.

[0102] In some embodiments, the determination submodule is specifically used to: determine whether the target data table has a primary key or a unique index through a preset query method; and / or determine whether the target data table has a primary key or a unique index based on the configured target field.

[0103] In some embodiments, the synchronization unit 502 may further include: a synchronization module (not shown in the figure) for synchronizing the data in the partition from the source relational database to the target end if the target data table has no primary key and no unique index.

[0104] In some embodiments, the data synchronization device 500 may also include a recording unit (not shown in the figure) for recording the location information corresponding to the synchronized data during the data synchronization process, the location information including the partition identifier of the partition to which the synchronized data belongs, the block identifier of the data block to which the synchronized data belongs, and the target primary key or target unique index corresponding to the synchronized data.

[0105] In some embodiments, the data synchronization device 500 may also include a processing unit (not shown in the figure) for: monitoring whether an interruption occurs during the data synchronization process; if an interruption occurs, when a data synchronization instruction for the target data table is detected, the target data table is resumed according to the recorded location information.

[0106] In some embodiments, the processing unit may further include: an output module (not shown in the figure) for outputting prompt information of data synchronization interruption if an interruption occurs.

[0107] Figure 5The provided data synchronization device can execute the steps in the method embodiment corresponding to the above-mentioned data synchronization method. Its implementation principle and technical effects are similar and will not be repeated here.

[0108] According to an embodiment of the present disclosure, the present disclosure also provides an electronic device, which includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the solution provided by any of the above embodiments.

[0109] According to an embodiment of the present disclosure, the present disclosure further provides a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to enable a computer to execute a solution provided by any of the above embodiments.

[0110] According to an embodiment of the present disclosure, the present disclosure also provides a computer program product, which includes: a computer program, the computer program is stored in a readable storage medium, at least one processor of an electronic device can read the computer program from the readable storage medium, and at least one processor executes the computer program so that the electronic device executes the solution provided by any of the above embodiments.

[0111] Figure 6 6 is a schematic block diagram of an example electronic device 600 that can be used to implement an embodiment of the present disclosure. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital assistants, cellular phones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present disclosure described and / or claimed herein.

[0112] like Figure 6 As shown, the electronic device 600 includes a computing unit 601, which can store data stored in a read-only memory (ROM) ( Figure 6 A computer program in ROM 602 is loaded from storage unit 608 to random access memory (RAM) ( Figure 6The computer programs in the RAM 603 are used to perform various appropriate actions and processes. The RAM 603 may also store various programs and data required for the operation of the electronic device 600. The computing unit 601, the ROM 602, and the RAM 603 are connected to each other via a bus 604. The input / output (I / O) interface ( Figure 6 An I / O interface 605 (for example) is also connected to the bus 604 .

[0113] Multiple components in the electronic device 600 are connected to the I / O interface 605, including: an input unit 606, such as a keyboard, a mouse, etc.; an output unit 607, such as various types of displays, speakers, etc.; a storage unit 608, such as a disk, an optical disk, etc.; and a communication unit 609, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 609 allows the electronic device 600 to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks.

[0114] The computing unit 601 may be a variety of general and / or special processing components with processing and computing capabilities. Some examples of the computing unit 601 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, digital signal processors (DSP), and any appropriate processors, controllers, microcontrollers, etc. The computing unit 601 performs the various methods and processes described above, such as a data synchronization method. For example, in some embodiments, the data synchronization method may be implemented as a computer software program, which is tangibly included in a machine-readable medium, such as a storage unit 608. In some embodiments, part or all of the computer program may be loaded and / or installed on the electronic device 600 via the ROM 602 and / or the communication unit 609. When the computer program is loaded into the RAM 603 and executed by the computing unit 601, one or more steps of the data synchronization method described above may be performed. Alternatively, in other embodiments, the computing unit 601 may be configured to execute the data synchronization method in any other appropriate manner (eg, by means of firmware).

[0115] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard parts (ASSPs), systems on chips (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: being implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.

[0116] The program code for implementing the method of the present disclosure may be written in any combination of one or more programming languages. These program codes may be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device, so that the program code, when executed by the processor or controller, enables the functions / operations specified in the flow chart and / or block diagram to be implemented. The program code may be executed entirely on the machine, partially on the machine, partially on the machine and partially on a remote machine as a stand-alone software package, or entirely on a remote machine or server.

[0117] In the context of the present disclosure, a machine-readable medium may be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, device, or equipment. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or equipment, or any suitable combination of the foregoing. A more specific example of a machine-readable storage medium may include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0118] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).

[0119] The systems and techniques described herein can be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer with a graphical user interface or a web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: Local Area Networks (LANs), Wide Area Networks (WANs), and the Internet.

[0120] A computer system may include a client and a server. The client and the server are generally remote from each other and usually interact through a communication network. The relationship between the client and the server is generated by computer programs running on the corresponding computers and having a client-server relationship with each other. The server may be a cloud server, also known as a cloud computing server or cloud host, which is a host product in the cloud computing service system to solve the defects of difficult management and weak business scalability in traditional physical hosts and VPS services ("Virtual Private Server", or "VPS" for short). The server may also be a server of a distributed system, or a server combined with a blockchain.

[0121] It should be understood that the various forms of processes shown above can be used to reorder, add or delete steps. For example, the steps recorded in this disclosure can be executed in parallel, sequentially or in different orders, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved, and this document does not limit this.

[0122] The above specific implementations do not constitute a limitation on the protection scope of the present disclosure. It should be understood by those skilled in the art that various modifications, fusions, sub-fusions and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions and improvements made within the spirit and principles of the present disclosure should be included in the protection scope of the present disclosure.

Claims

1. A data synchronization method, comprising: In response to a data synchronization instruction for a target data table in a source relational database, data in the target data table is read according to a preset block size to obtain data blocks; The data blocks are synchronized from the source relational database to the target end.

2. The data synchronization method according to claim 1, wherein: The step of reading the data in the target data table according to the preset block size to obtain data blocks includes: Determine whether the target data table is stored in partitions; If the target data table is stored in partitions, for each partition, the data in the target data table corresponding to the partition is read according to the preset block size to obtain data blocks; If the target data table is not stored in partitions, the data in the target data table is read according to a preset block size to obtain data blocks.

3. The data synchronization method according to claim 2, wherein: For each partition, reading the data in the partition corresponding to the target data table according to the preset block size to obtain the data block includes: Determine whether the target data table has a primary key or a unique index; If the target data table has a primary key, the data in the partition corresponding to the target data table is sorted according to the primary key to obtain sorted data corresponding to the partition; or, if the target data table has a unique index, the data in the partition corresponding to the target data table is sorted according to the unique index to obtain sorted data corresponding to the partition; The sorted data is read according to the preset block size to obtain data blocks.

4. The data synchronization method according to claim 2, wherein: The target data table is stored in partitions, and for each partition, the data in the target data table corresponding to the partition is read according to the preset block size to obtain data blocks, including: Sorting the partitions to obtain sorted partitions; For each partition in the sorted partitions, the data in the partition corresponding to the target data table is read according to the preset block size to obtain a data block.

5. The data synchronization method according to claim 3, wherein: Determining whether the target data table has a primary key or a unique index includes: Determine whether the target data table has a primary key or a unique index by using a preset query method; And / or, based on the configured target field, determine whether the target data table has a primary key or a unique index.

6. The data synchronization method according to claim 3, wherein: The method further comprises: If the target data table does not have a primary key and a unique index, the data in the partition is synchronized from the source relational database to the target end.

7. The data synchronization method according to any one of claims 1 to 6, wherein: The method further comprises: During the data synchronization process, the location information corresponding to the synchronized data is recorded, and the location information includes the partition identifier of the partition to which the synchronized data belongs, the block identifier of the data block to which the synchronized data belongs, and the target primary key or target unique index corresponding to the synchronized data.

8. The data synchronization method according to claim 7, wherein: The method further comprises: During data synchronization, monitor whether there are interruptions; If an interruption occurs, when a data synchronization instruction for the target data table is detected, the target data table is resumed according to the recorded location information.

9. The data synchronization method according to claim 8, wherein: The method further comprises: If an interruption occurs, a prompt message indicating data synchronization interruption is output.

10. A data synchronization device, comprising: A reading unit, configured to respond to a data synchronization instruction for a target data table in a source relational database, read the data in the target data table according to a preset block size, and obtain data blocks; A synchronization unit is used to synchronize the data blocks from the source relational database to the target end.

11. An electronic device, comprising: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the data synchronization method according to any one of claims 1 to 9.

12. A non-transitory computer-readable storage medium storing computer instructions, wherein: The computer instructions are used to enable the computer to execute the data synchronization method according to any one of claims 1 to 9.

13. A computer program product, comprising a computer program, wherein when the computer program is executed by a processor, the steps of the data synchronization method according to any one of claims 1 to 9 are implemented.