Data Synchronization Method, Device and Electronic Device Based on Distributed Database

By establishing a batch transaction operation number queue in a distributed transaction database cluster and using streaming merging algorithm for data synchronization, data synchronization performance bottlenecks and transaction consistency problems are solved, and efficient and reliable data synchronization is achieved.

CN119474220BActive Publication Date: 2025-06-24TIANJIN NANKAI UNIV GENERAL DATA TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510038936.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-01-10
Publication Date
2025-06-24
Estimated Expiration
2045-01-10

AI Technical Summary

Technical Problem

In a distributed transaction database cluster, there are performance bottlenecks in the data synchronization process, and in high concurrency and high throughput scenarios, it is difficult to ensure transaction consistency and reduce redundant synchronization operations.

Method used

By establishing a batch transaction operation number queue, responding to batch transaction processing requests, filtering the data update records corresponding to the target batch transaction operation number, and using streaming merging algorithm and parallel loading threads to merge and synchronize data, ensuring transaction logic consistency and efficient performance.

Benefits of technology

It realizes the consideration of database performance and transaction logic consistency in distributed databases, improves data synchronization performance, reduces redundant operations, and ensures data consistency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119474220B_ABST
    Figure CN119474220B_ABST
Patent Text Reader

Abstract

The present invention provides a data synchronization method, apparatus and electronic device based on a distributed database, which can be applied to the field of electronic digital data processing. The method includes: in response to a batch transaction processing request, reading a target batch transaction operation number corresponding to the currently to-be-processed batch transaction from a batch transaction operation number queue; determining change data queues respectively corresponding to L source data nodes; screening N target data update records with transaction operation numbers smaller than the target batch transaction operation number from the L change data queues, and the N target data update records are associated with M source data tables; based on a streaming merge algorithm, using M merge threads to parallelly perform merge processing on the target data update records respectively associated with the M source data tables to obtain M data merge hash tables corresponding to the M source data tables; using M loading threads to parallelly synchronize the merged data of the M source data tables to a target database.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of electronic digital data processing, and more particularly to a data synchronization method, apparatus, and electronic device based on a distributed database. Background Art

[0002] With the rise and development of distributed transaction database cluster technology, a cluster supports hundreds of nodes to provide transaction services externally. When using a unified transaction log of the cluster to merge transactions accessed by hundreds of nodes during the data synchronization process of the cluster database, there is a performance bottleneck; if each node separately performs data change merging and synchronization, a transaction often involves multiple data nodes, and the consistency of the transaction cannot be guaranteed, resulting in data errors; in addition, in high-concurrency and high-throughput scenarios, there will be many duplicate transaction operations, resulting in unnecessary synchronization operations, seriously affecting the performance of synchronization. Therefore, how to balance database performance and transaction logic consistency in cluster database data synchronization is an urgent problem to be solved. Summary of the Invention

[0003] In view of the above problems, the present invention provides a data synchronization method, apparatus, and electronic device based on a distributed database.

[0004] According to a first aspect of the present invention, there is provided a data synchronization method based on a distributed database. The distributed database includes a source database and a target database for which data synchronization is to be performed. The source database includes L source data nodes. The method includes: in response to a batch transaction processing request, reading a target batch transaction operation number corresponding to the currently pending batch transaction from a batch transaction operation number queue; determining change data queues corresponding to the L source data nodes respectively, where the change data queues store multiple groups of data update records generated by the source data nodes through transaction operations during the data synchronization process, and the multiple groups of data update records correspond to multiple transaction operation numbers; screening N target data update records with transaction operation numbers less than the target batch transaction operation number from the L change data queues, the N target data update records are associated with M source data tables, and the L source data nodes jointly complete the data storage of the M source data tables; based on a streaming merge algorithm, using M merge threads to parallelly merge the target data update records associated with the M source data tables respectively to obtain M data merge hash tables corresponding to the M source data tables, where the m-th source data table is associated with a target data update record, and the data merge hash table corresponding to the m-th source data table stores associated with primary keys, the primary keys are derived from the target data update records, and each merge result includes at most one merged data, , ; Use M loading threads to synchronize the merged data of M source data tables to the target database in parallel.

[0005] The second aspect of the present invention provides a data synchronization device based on a distributed database, including: a reading module for reading the target batch transaction operation number corresponding to the currently pending batch transaction from the batch transaction operation number queue in response to a batch transaction processing request; a determination module for determining the change data queues corresponding to L source data nodes respectively, wherein the change data queues store multiple groups of data update records generated by the source data nodes through transaction operations during the data synchronization process, and the multiple groups of data update records correspond to multiple transaction operation numbers; a screening module for screening N target data update records with transaction operation numbers smaller than the target batch transaction operation number from the L change data queues, the N target data update records are associated with M source data tables, and the L source data nodes jointly complete the data storage of the M source data tables; a merging module for performing a merging process on the target data update records associated with each of the M source data tables in parallel by using M merging threads based on a streaming merging algorithm to obtain M data merging hash tables corresponding to the M source data tables, wherein the m-th source data table is associated with a target data update record, and the data merging hash table corresponding to the m-th source data table stores associated with merged results, the primary keys are derived from the target data update records, and each merged result contains at most one merged data, , ; a synchronization module for using M loading threads to synchronize the merged data of M source data tables to the target database in parallel.

[0006] The third aspect of the present invention provides an electronic device, including: one or more processors; a memory for storing one or more programs, wherein when the one or more programs are executed by the one or more processors, the one or more processors execute the above-mentioned data synchronization method based on a distributed database.

[0007] According to the data synchronization method, device and electronic device based on a distributed database provided by the present invention, by establishing a batch transaction operation number queue to periodically read batch transaction operation numbers, the transaction submission boundaries between multiple batch transaction operation numbers in the batch transaction operation number queue are clear. Then, based on the target batch transaction operation number, a batch of target data update records associated with L source data nodes are filtered out, and batch transaction processing is performed according to the transaction operation numbers in the target data update records, ensuring the consistency of the in-batch transaction logic and the inter-batch transaction boundaries; by establishing respective change data queues for the source data nodes to continuously collect data update records on the source data nodes, and directly reading from the change data queues when batch transaction processing is involved, without the need for the source data nodes to operate, the pressure on the source data nodes is alleviated, and the database performance is improved; the respective change data queues of the L source data nodes are collected in parallel, the M source data tables are merged in batch within the batch, and the M loading threads are loaded in parallel, ensuring parallel synchronization processing and effectively utilizing the central processing unit resources; in addition, two batch processes of in-batch merging of the obtained batch of data update records and batch synchronization of the M source data tables are performed to remove redundant transaction operations and synchronization operations, greatly improving the distributed database synchronization performance. BRIEF DESCRIPTION OF THE DRAWINGS

[0008] Through the following description of the embodiments of the present invention with reference to the accompanying drawings, the above and other objects, features and advantages of the present invention will become more apparent.

[0009] Figure 1 FIG. shows an application scenario of the data synchronization method and device based on a distributed database according to an embodiment of the present invention.

[0010] Figure 2 FIG. shows a flowchart of the data synchronization method based on a distributed database according to an embodiment of the present invention.

[0011] Figure 3 FIG. shows a schematic diagram of a distributed database according to an embodiment of the present invention.

[0012] Figure 4 FIG. shows a structural block diagram of the data synchronization device based on a distributed database according to an embodiment of the present invention.

[0013] Figure 5 FIG. shows a block diagram of an electronic device suitable for implementing the data synchronization method based on a distributed database according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0014] Hereinafter, embodiments of the present invention will be described with reference to the accompanying drawings. However, it should be understood that these descriptions are merely exemplary and are not intended to limit the scope of the present invention. In the following detailed description, for the sake of explanation, numerous specific details are set forth in order to provide a thorough understanding of the embodiments of the present invention. However, it is obvious that one or more embodiments can also be implemented without these specific details. In addition, in the following description, descriptions of well-known structures and technologies are omitted to avoid unnecessarily obscuring the concepts of the present invention.

[0015] The terms used herein are merely for describing specific embodiments and are not intended to limit the present invention. The terms "comprising", "including", etc. used herein indicate the presence of features, steps, operations, and / or components, but do not exclude the presence or addition of one or more other features, steps, operations, or components.

[0016] All terms used herein (including technical and scientific terms) have the meanings commonly understood by those skilled in the art, unless otherwise defined. It should be noted that the terms used herein should be interpreted as having a meaning consistent with the context of this specification and should not be interpreted in an idealized or overly rigid manner.

[0017] In the case of using expressions such as "at least one of A, B, and C, etc.", generally, it should be interpreted according to the meaning commonly understood by those skilled in the art (for example, "a system having at least one of A, B, and C" should include, but not be limited to, a system having only A, having only B, having only C, having A and B, having A and C, having B and C, and / or having A, B, and C, etc.).

[0018] In the process of implementing the present invention, it is found that with the rise and development of the distributed transaction database cluster technology, a cluster supports hundreds of nodes to provide transaction services externally. When using the unified transaction log of the cluster to merge the transactions accessed by hundreds of nodes during the data synchronization process of the cluster database, there is a performance bottleneck; if each node separately performs data change merging and synchronization, a transaction often involves multiple data nodes, and the consistency of the transaction cannot be guaranteed, resulting in data errors; in addition, in high-concurrency and high-throughput scenarios, there will be many duplicate transaction operations, thus causing unnecessary synchronization operations to occur (for example: updating the value of a certain field of the same record multiple times, only the last update needs to be synchronized), seriously affecting the performance of synchronization. Therefore, how to balance database performance and transaction logic consistency in the data synchronization of the cluster database is an urgent problem to be solved.

[0019] In view of this, embodiments of the present invention provide a data synchronization method, apparatus, and electronic device based on a distributed database. The method includes: in response to a batch transaction processing request, reading a target batch transaction operation number corresponding to the current batch transaction to be processed from a batch transaction operation number queue; determining change data queues corresponding to L source data nodes respectively; screening N target data update records with transaction operation numbers less than the target batch transaction operation number from the L change data queues, where the N target data update records are associated with M source data tables; based on a streaming merge algorithm, using M merge threads to parallelly merge the target data update records associated with each of the M source data tables to obtain M data merge hash tables corresponding to the M source data tables; using M loading threads to parallelly synchronize the merged data of the M source data tables to a target database.

[0020] In the technical solution of the present invention, the involved user information (including but not limited to user personal information, user image information, user device information, such as location information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) are all information and data authorized by the user or fully authorized by all parties. Moreover, the processing of relevant data, such as collection, storage, use, processing, transmission, provision, disclosure, and application, complies with relevant laws, regulations, and standards, takes necessary confidentiality measures, does not violate public order and good customs, and provides corresponding operation entrances for users to choose to authorize or reject.

[0021] Figure 1 FIG. shows an application scenario of a data synchronization method and apparatus based on a distributed database according to an embodiment of the present invention.

[0022] As Figure 1 shown, the service system architecture 100 according to this embodiment may include a first terminal device 101, a second terminal device 102, a third terminal device 103, a network 104, and a server 105. The network 104 is used to provide a medium for communication links between the first terminal device 101, the second terminal device 102, the third terminal device 103, and the server 105. The network 104 may include various connection types, such as wired, wireless communication links, or fiber optic cables, etc.

[0023] Users can use at least one of the first terminal device 101, the second terminal device 102, and the third terminal device 103 to interact with the server 105 through the network 104 to receive or send messages, etc. Various communication client applications may be installed on the first terminal device 101, the second terminal device 102, and the third terminal device 103, such as shopping applications, web browser applications, search applications, instant messaging tools, email clients, social platform software, etc. (only for example).

[0024] The first terminal device 101, the second terminal device 102, and the third terminal device 103 may be various electronic devices with a display screen and supporting web browsing, including but not limited to smartphones, tablet computers, laptop computers, desktop computers, and the like.

[0025] The server 105 may be a server that provides various services. For example, it may be a background management server (only for example) that supports the websites browsed by users using the first terminal device 101, the second terminal device 102, and the third terminal device 103. The cluster distributed database is deployed on the cluster server 105. The server 105 includes multiple server nodes that can analyze and process data such as user requests received, and feedback the processing results (such as web pages, information, or data obtained or generated according to user requests) to the terminal devices.

[0026] It should be noted that the data synchronization method based on the distributed database provided by the embodiments of the present invention can generally be executed by the server 105. Correspondingly, the data synchronization device based on the distributed database provided by the embodiments of the present invention can generally be set in the server 105. The data synchronization method based on the distributed database provided by the embodiments of the present invention can also be executed by a server or a server cluster different from the server 105 and capable of communicating with the first terminal device 101, the second terminal device 102, the third terminal device 103, and / or the server 105. Correspondingly, the data synchronization device based on the distributed database provided by the embodiments of the present invention can also be set in a server or a server cluster different from the server 105 and capable of communicating with the first terminal device 101, the second terminal device 102, the third terminal device 103, and / or the server 105.

[0027] Alternatively, the data synchronization method based on the distributed database provided by the embodiments of the present invention can also be executed by the first terminal device 101, the second terminal device 102, or the third terminal device 103, or can also be executed by other terminal devices different from the first terminal device 101, the second terminal device 102, or the third terminal device 103. Correspondingly, the data synchronization device based on the distributed database provided by the embodiments of the present invention can also be set in the first terminal device 101, the second terminal device 102, or the third terminal device 103, or set in other terminal devices different from the first terminal device 101, the second terminal device 102, or the third terminal device 103.

[0028] It should be understood that Figure 1 the numbers of terminal devices, networks, and servers in

[0029] It should be noted that the sequence numbers of the respective operations in the following methods are only used as representations of the operations for description purposes and should not be regarded as indicating the execution order of the respective operations. Unless explicitly stated, the method does not need to be executed exactly in the order shown.

[0030] Figure 2 The flowchart of a data synchronization method based on a distributed database according to an embodiment of the present invention is shown.

[0031] As Figure 2 shown, the method 200 includes operation S210 to operation S250.

[0032] In operation S210, in response to a batch transaction processing request, the target batch transaction operation number corresponding to the currently pending batch transaction is read from the batch transaction operation number queue.

[0033] Optionally, the distributed database includes a source database and a target database for which data synchronization is to be performed. The source database includes L source data nodes. The source data table to be synchronized is stored in the source database, and the target table mapped to the source data table is stored in the target database.

[0034] Optionally, a batch transaction processing request is created. The batch transaction processing request is a batch data synchronization task. Configuration parameters are configured when creating the batch transaction processing request. For example, the configuration parameters include source database connection parameters (such as Uniform Resource Locator (Java Database Connectivity URL, JDBC URL), connection configuration file (properites), etc.), target database connection parameters, batch transaction operation number reading interval parameters, etc. After creation, the system saves the configuration parameters and starts the batch transaction processing request.

[0035] Optionally, a Kafka component can be used to create a batch transaction operation number queue. The batch transaction operation number queue is used to store the batch transaction operation numbers periodically read from the global transaction manager based on the batch transaction operation number reading interval parameter. Transactions with transaction operation numbers less than or equal to the batch transaction operation number in the global transaction manager have been committed, and thus data synchronization operations can be performed.

[0036] Optionally, multiple batch transaction operation numbers are stored in the batch transaction operation number queue, and the smallest batch transaction operation number is read from the batch transaction operation number queue based on the response moment as the target batch transaction operation number corresponding to the currently pending batch transaction.

[0037] In operation S220, the change data queues corresponding to each of the L source data nodes are determined.

[0038] Optionally, a Kafka component can be used to create a change data queue for each source data node. L is a positive integer, and L change data queues corresponding to the L source data nodes are determined.

[0039] Optionally, a two-layer producer-consumer model is adopted to form an efficient pipeline operation; at the same time, the introduction of the Kafka component effectively improves the reliability and high availability of the queue.

[0040] Optionally, the change data queue stores multiple groups of data update records generated by the source data node through transaction operations during the data synchronization process, and the multiple groups of data update records correspond to multiple transaction operation numbers.

[0041] Optionally, the transaction operation represents transaction operations such as querying, deleting, and modifying the source data table, and the data update record is the data change information generated by executing the transaction operation on the source data table. In response to a transaction operation, a globally unique transaction operation number is applied for from the global transaction manager.

[0042] Optionally, the data update record generated by the source data node executing the transaction operation contains the transaction operation number corresponding to this transaction operation, and the data update record is stored in the change data queue corresponding to this source data node.

[0043] Optionally, in the case where a transaction operation involves multiple source data nodes, the multiple groups of data update records generated correspond to the same transaction operation number; the multiple groups of data update records generated for multiple transaction operations correspond to multiple transaction operation numbers.

[0044] In operation S230, N target data update records with transaction operation numbers less than the target batch transaction operation number are filtered from the L change data queues.

[0045] Optionally, the source database includes M source data tables, M is a positive integer, the N target data update records are the batch data to be synchronized filtered based on the target batch transaction operation number, N is a positive integer, and the N target data update records are associated with the M source data tables.

[0046] Optionally, L source data nodes jointly complete the data storage of M source data tables. Multiple source data nodes jointly complete the storage of a source data table. A transaction operation may involve operating on part of the data in the source data table. In the case where part of the data is stored in multiple source data nodes, multiple source data nodes need to jointly complete the execution of a transaction operation.

[0047] Optionally, the target data update records with transaction operation numbers less than the target batch transaction operation number are sequentially read from each change data queue to obtain N target data update records.

[0048] In operation S240, based on the streaming merge algorithm, M merge threads are used to parallelly merge the target data update records associated with each of the M source data tables to obtain M data merge hash tables corresponding to the M source data tables.

[0049] Optionally, during the reading process, determining the source data table corresponding to the filtered target data update record may include: determining the associated source data table according to the table name information in the target data update record.

[0050] Optionally, based on the streaming merge algorithm, a merge thread performs a merge process on the target data update records associated with a source data table to obtain a data merge hash table corresponding to the source data table.

[0051] Optionally, the data merge hash table is a hash (Hashtable) storage structure. The m-th source data table is associated with a number of target data update records, and each target data update record contains a primary key. In the case where there are no identical primary keys among the number of target data update records, ; in the case where there are identical primary keys among the number of target data update records, it means that the data records corresponding to this primary key have undergone multiple transaction operations, . The sum of the multiple target data update records associated with each of the M source data tables is the total number N of the batch data to be synchronized, .

[0052] Optionally, traverse the number of target data update records in ascending order of the transaction operation number, perform a merge process on the target data update records associated with each primary key to obtain a merge result, and each merge result contains 1 or 0 merge data. The data merge hash table corresponding to the m-th source data table stores the number of merge results associated with primary keys.

[0053] Optionally, the multiple target data update records associated with a source data table may come from the change data queues of multiple source data nodes respectively. Therefore, it is possible to first sort the multiple target data update records based on the transaction operation number order, and then perform merge and deduplication on the multiple target data update records corresponding to each primary key to obtain the merge result for each primary key.

[0054] For example, the merge process includes: converting an update operation into a delete-then-add operation; for duplicate update operations and delete operations, only keep the last operation; for operations that can be merged for consecutive operations before and after, Example 1: Inserted a piece of data, then updated a field, and finally deleted the data. After merging, these 3 operations will be deleted and regarded as not having occurred; Example 2: Updated a certain field to value 1 and then to value 2. After merging, only keep the latter.

[0055] In operation S250, the merged data of M source data tables is synchronously transferred to the target database in parallel by M loading threads.

[0056] Optionally, a dynamic data operation statement (Structured Query Language, SQL) for the merged data of the source data table is created, and each merged data variable of the source data table is assigned to the dynamic data operation statement according to the transaction operation type and stored in the operation queue.

[0057] Optionally, the configuration parameters include the parameter of the amount of data submitted in a single batch. A loading thread is started for each source data table to batch-read the dynamic data operation statements to be executed from the operation queue and execute them in the target database, so as to synchronize the merged data to the target table mapped to the source data table in the target database. After the reading of the operation queue ends, a batch transaction submission operation is executed.

[0058] Optionally, batch processing is used in multiple places to improve performance. When collecting the updated records of the target data, the processing is carried out according to the target batch transaction operation number; when analyzing, merging, and deduplicating the updated records of the target data, the in-batch processing method is adopted; when synchronizing the merged data to the target database, it is also batch processing.

[0059] For example, a total of multiple data update records are generated by 3 source data nodes. In response to a batch transaction processing request, the target batch transaction operation number 1000 is obtained, and multiple target data update records with a transaction operation number less than 1000 are filtered from the change data queues of the 3 source data nodes respectively, and a total of 10,000 target data update records are obtained. Similarly, in response to a batch transaction processing request, the target batch transaction number 2001 is obtained, and a total of 30,000 target data update records are obtained from the 3 source data nodes. In response to a batch transaction processing request, the target batch transaction number 3001 is obtained, and a total of 40,000 target data update records are obtained from the 3 source data nodes. Thus, multiple data update records from multiple source data nodes are synchronized in multiple batches. For each batch, the in-batch merging of multiple data update records across source data nodes and under a controllable quantity is realized through the target batch transaction operation number, ensuring the transaction consistency of multiple source data nodes within a single batch and consuming fewer resources to realize the in-batch merging of data update records.

[0060] Optionally, by establishing a batch transaction operation number queue to periodically read batch transaction operation numbers, the transaction submission boundaries between multiple batch transaction operation numbers in the batch transaction operation number queue are clear. Then, based on the target batch transaction operation number, a batch of target data update records associated with L source data nodes are filtered out, and batch transaction processing is performed according to the transaction operation numbers in the target data update records, ensuring the consistency of the intra-batch transaction logic and the inter-batch transaction boundaries; establish a change data queue for each source data node to continuously collect data update records on the source data node, and directly read from the change data queue when it comes to batch transaction processing, without the need for the source data node to operate, alleviating the pressure on the source data node and improving the database performance; the change data queues of L source data nodes collect in parallel, M source data tables are merged within the batch in parallel, and M loading threads load in parallel, ensuring parallel synchronous processing and effectively utilizing the central processing unit resources; in addition, two batch processes of in-batch merging of the obtained batch of data update records and batch synchronization of M source data tables are performed to remove redundant transaction operations and synchronization operations, greatly improving the synchronization performance of the distributed database.

[0061] Optionally, the data synchronization method based on the distributed database further includes: in response to a transaction operation request for each source data table, determining a transaction operation number and i target source data table sub-tables to be operated, where the transaction operation number is obtained by allocation from the global transaction manager, the source data table includes I source data table sub-tables, the i target source data table sub-tables are derived from the I source data table sub-tables, the source data nodes correspond to the source data table sub-tables one by one, and i ≤ I; performing a data change operation on the i target source data table sub-tables to obtain i data update records corresponding to the i target source data table sub-tables, where each data update record includes a transaction operation number, a transaction operation type, a target source data table sub-table name, a primary key, and a changed column; synchronously storing the i data update records in i logical replication slots corresponding to the target source data nodes where the i target source data table sub-tables are located respectively; using i reading threads to parallelly store the data update records in the i logical replication slots into i change data queues corresponding to the i target source data nodes.

[0062] Optionally, in response to a transaction operation request for the source data table, the server applies for a transaction operation number from the global transaction manager, determines i target source data table sub-tables to be operated from the I source data table sub-tables, and further determines the target source data nodes where the i target source data table sub-tables are located respectively.

[0063] Optionally, the source data nodes correspond one-to-one with the logical replication slots. A distributed database will have multiple source data nodes for storing and processing transactions. Each source data node will provide a logical replication slot to externally provide a data change capture service. The changed data is read in parallel from these logical replication slots and stored in the kafka changed data queue. Each source data node corresponds to a changed data queue.

[0064] Optionally, the target source data node performs a data change operation on the sub-table of the target source data table to obtain a data update record. The logical replication slot records the data update record on the target source data node. The reading thread can be a logical replication slot reading thread, and the logical replication slot reading thread then reads the data update record into the changed data queue corresponding to the target source data node.

[0065] For example, for each source data node in the distributed database, start a logical replication slot data reading thread (e.g., data thread 1, which reads the data changes on the source data node node1). After reading the data update record, synchronously add it to the changed data queue node1_db1 of the source data node node1.

[0066] Optionally, the transaction operation type can be an insert type, a delete type, an update type, etc. The changed column is the changed field in the sub-table of the target source data table.

[0067] Optionally, synchronously store the i data update records in the i logical replication slots corresponding to the i target source data nodes where the i sub-tables of the target source data table are located respectively; use the i logical replication slot reading threads to parallelly store the data update records in the i logical replication slots into the i changed data queues corresponding to the i target source data nodes.

[0068] Optionally, the logical replication slot reading threads parallelly capture the data changes of each source data node in the distributed database, which grows linearly with the increase of the source data nodes. Therefore, it is not affected by the scale of the distributed database and does not affect the performance, and can match the throughput of the distributed database.

[0069] Optionally, before the merge processing, it further includes: for any m-th source data table, based on the transaction operation number, perform a sorting process on the target data update records associated with the m-th source data table to obtain the data sorting hash table corresponding to the m-th source data table, where the data sorting hash table corresponding to the m-th source data table stores the sorting results associated with the primary keys, and the sorting result corresponding to the h-th primary key contains the ordered target data update records. , ; among them, for any m-th source data table, the merging process includes: for the sorting result corresponding to the h-th primary key in the sorting hash table, based on the streaming merging algorithm, using merging threads to merge the ordered target data update records to obtain the merging result corresponding to the h-th primary key; according to the number of primary keys and the merging result corresponding to each primary key, obtain the data merging hash table corresponding to the m-th source data table.

[0070] Optionally, during the process of respectively reading the target data update records from the change data queues corresponding to each source data node, determine the source data table corresponding to the target data update record. Start a sorting thread for each source data table, and based on the transaction operation number, use M sorting threads to parallelly sort the multiple target data update records associated with the M source data tables to obtain the data sorting hash tables respectively corresponding to the M source data tables.

[0071] For example, read the target data update records from the change data queues node1_db1, node2_db1, node3_db1 corresponding to the source data nodes node1, node2, node3 respectively. The target data update records a1 (transaction operation number 001, xxx), a2 (transaction operation number 005, xxx) read from node1_db1; the target data update records b1 (transaction operation number 004, xxx), b2 (transaction operation number 006, xxx) read from node2_db1; the target data update records c1 (transaction operation number 003, xxx), c2 (transaction operation number 008, xxx) read from node3_db1. Source data table 1 is associated with a1, b2, c1, and source data table 2 is associated with a2, b1, c2. Based on the ascending order rule of the transaction operation number, use two sorting threads to sort the target data update records associated with source data table 1 and source data table 2 respectively, and the data sorting hash table corresponding to source data table 1 includes the ordered a1, c1, b2, and the data sorting hash table corresponding to source data table 2 includes the ordered b1, a2, c2.

[0072] Optionally, the data sorting hash table corresponding to the m-th source data table includes the primary keys from the target data update records associated with the m-th source data table, and the sorting result corresponding to each primary key contains multiple ordered target data update records, and the sorting results of different primary keys are also ordered.

[0073] For example, in the data sorting hash table corresponding to the source data table 3, it includes d1 (transaction operation number 010, xxx), d2 (transaction operation number 011, xxx), d3 (transaction operation number 012, xxx), d3 (transaction operation number 014, xxx). The sorting result corresponding to the primary key d1 contains a target data update record d1 (transaction operation number 010, xxx); the sorting result corresponding to the primary key d2 contains a target data update record d2 (transaction operation number 011, xxx); the sorting result corresponding to the primary key d3 contains two ordered target data update records d3 (transaction operation number 012, xxx), d3 (transaction operation number 014, xxx).

[0074] Optionally, h, 、 、 are all positive integers. The sum of the ordered target data update records corresponding to each of the primary keys is the total number of target data update records associated with the m-th source data table , 。

[0075] Optionally, the configuration parameter includes a merging parameter. When the merging parameter is 0, all target data update records are recorded and no merging operation is performed; when the merging parameter is 1, a merging operation is performed.

[0076] Optionally, the merging and deduplication processing of the target data update records within a batch is performed at the dimension of the source data table. For each source data table, a merging thread is started. While sorting each source data table, the obtained data sorting hash table is handed over to the merging thread for merging processing. For the sorting result corresponding to each primary key in the sorting hash table, based on the streaming merging algorithm, the merging thread is used to merge at least one ordered target data update record corresponding to each primary key to obtain the merging result corresponding to each primary key. The merging result corresponding to each primary key contains 1 or 0 merged data.

[0077] Optionally, according to the primary keys and the merging result corresponding to each primary key, the data merging hash table corresponding to the m-th source data table is obtained.

[0078] Optionally, target data update records are read from the change data queues corresponding to each source data node based on the target batch transaction operation number. Since the target data update records read from a change data queue are sorted by the transaction operation number, when sorting the target data update records from different change data queues in a source data table, the merge sort algorithm can be directly used to sort only the target data update records between the change data queues, and the target data update records in the same change data queue do not need to be sorted again. The sorting time complexity is O(N), which greatly improves the database performance.

[0079] Optionally, for the sorting result corresponding to the h-th primary key, based on the streaming merge algorithm, the merge thread is used to merge the ordered target data update records to obtain the merge result corresponding to the h-th primary key, including: traversing the ordered target data update records; for each target data update record, updating the data change operations in the linked list corresponding to the h-th primary key according to the h-th primary key and the transaction operation type; obtaining the merge result corresponding to the h-th primary key according to the data change operations in the linked list corresponding to the h-th primary key.

[0080] Optionally, the data merge hash table contains the primary key and the linked list corresponding to the primary key, and the linked list stores the data change operations of the primary key in the order of the transaction operation number.

[0081] Optionally, for the sorting result corresponding to the h-th primary key, traverse the ordered target data update records, and for each target data update record, update the data change operations in the linked list corresponding to the h-th primary key according to the h-th primary key and the transaction operation type in the target data update record.

[0082] Optionally, after traversing the ordered target data update records, obtain the merge result corresponding to the h-th primary key according to the data change operations in the linked list corresponding to the h-th primary key.

[0083] Optionally, for each target data update record, according to the h-th primary key and the transaction operation type, the data change operations for updating the data in the linked list corresponding to the h-th primary key include: for each target data update record, when the h-th primary key does not exist among multiple candidate target data update records in the data sorting hash table, adding a data change operation to the linked list corresponding to the h-th primary key, where the transaction operation number in the candidate target data update record is less than the transaction operation number in the target data update record; for each target data update record, when the h-th primary key exists among multiple candidate target data update records and the transaction operation type is the insert type, clearing the linked list corresponding to the h-th primary key and adding the current data insertion operation; for each target data update record, when the h-th primary key exists among multiple candidate target data update records and the transaction operation type is the update type, traversing the linked list corresponding to the h-th primary key and adding the current data update operation when there is no historical data update operation for the changed columns in the linked list; for each target data update record, when the h-th primary key exists among multiple candidate target data update records and the transaction operation type is the delete type, clearing the linked list corresponding to the h-th primary key and adding the current data deletion operation when there is no historical data insertion operation for the h-th primary key in the linked list.

[0084] Optionally, for a certain target data update record, the target data update records with a transaction operation number less than that of this target data update record in the data sorting hash table are determined as candidate target data update records. This target data update record includes the h-th primary key, and the non-existence of the h-th primary key among multiple candidate target data update records indicates that this target data update record is generated based on the first data change operation for the h-th primary key. For example, the first data insertion operation for the h-th primary key.

[0085] Optionally, when the h-th primary key does not exist among multiple candidate target data update records, adding a data change operation corresponding to this target data update record to the linked list corresponding to the h-th primary key, and the data change operation is an operation for changing the data in the sub-table of the target source data table based on the transaction operation request for the source data table.

[0086] Optionally, the existence of the h-th primary key among multiple candidate target data update records indicates that this target data update record is generated based on a non-first data change operation for the h-th primary key.

[0087] Optionally, when the h-th primary key exists among multiple candidate target data update records and the transaction operation type is the insert type, clearing the linked list corresponding to the h-th primary key and adding the current data insertion operation. At this time, it is equivalent to clearing all previous operations for the h-th primary key and performing a merge.

[0088] Optionally, when the h-th primary key exists in multiple candidate target data update records and the transaction operation type is the update type, traverse the linked list corresponding to the h-th primary key. If there is no historical data update operation for the changed columns in the linked list, add the current data update operation and merge the historical data update operations for other columns. At this time, it is equivalent to merging all update results and will not affect the final data result.

[0089] Optionally, if there is a historical data update operation for the changed columns in the linked list, remove the historical data update operation and add the current data update operation.

[0090] Optionally, when the h-th primary key exists in multiple candidate target data update records and the transaction operation type is the delete type, clear the linked list corresponding to the h-th primary key. If there is no historical data insertion operation for the h-th primary key in the linked list, add the current data delete operation. At this time, it is equivalent to ignoring all previous operations for the h-th primary key and only retaining the data delete operation, which does not affect the final data result.

[0091] Optionally, if there is a historical data insertion operation for the h-th primary key in the linked list, do not add the current data delete operation. This is to avoid deleting the historical data insertion operation when there is only a data delete operation, resulting in an error when deleting.

[0092] Optionally, merge and deduplicate multiple target data update records corresponding to the primary key. According to the characteristics of the changed data involved in the synchronization process, in the linked list of the primary key, only retain the last operation for duplicate data update operations and data delete operations; merge the operations that can be merged for consecutive operations before and after, and then remove the redundant data change operations with low importance without affecting the final result, thereby effectively improving the synchronization performance.

[0093] Optionally, the data synchronization method based on a distributed database further includes: determining the execution status of a transaction operation request corresponding to a transaction operation number; reading interval parameters based on a batch transaction operation number, and reading multiple transaction operation numbers corresponding to the execution status representing the uncommitted state from a global transaction manager; selecting the smallest transaction operation number from the multiple transaction operation numbers as the target transaction operation number; determining the previous transaction operation number of the target transaction operation number as the batch transaction operation number, where the execution status of the batch transaction operation number and the transaction operation numbers before the batch transaction operation number both represent the committed state; storing the batch transaction operation number in a batch transaction operation number queue.

[0094] Optionally, the execution status of a transaction operation request represents the committed and uncommitted states of a transaction.

[0095] Optionally, based on the batch transaction operation number reading interval parameter, read multiple transaction operation numbers corresponding to the execution status representing the uncommitted state from the global transaction manager, and select the smallest transaction operation number from the multiple transaction operation numbers as the target transaction operation number.

[0096] For example, the batch transaction operation number reading interval parameter is 10 minutes, and the target transaction operation number is read from the global transaction manager every 10 minutes. In one reading, the execution status corresponding to the transaction operation request number in the global transaction manager is (transaction operation number 0000: committed status, xxx, transaction operation number 1000: committed status, transaction operation number 1001: uncommitted status, transaction operation number 1002: committed status, transaction operation number 1003: uncommitted status, xxx). The smallest transaction operation number among the multiple transaction operation numbers corresponding to the uncommitted execution status is 1001, so the target transaction operation number is determined as transaction operation number 1001, and then transaction operation number 1000 is determined as the batch transaction operation number.

[0097] Optionally, the execution statuses of the batch transaction operation number and the transaction operation numbers before the batch transaction operation number both represent the committed status. Based on the batch transaction operation number reading interval parameter, read one batch transaction operation number periodically each time and store it in the batch transaction operation number queue. When the synchronization of the previous batch of transactions is completed, read the target batch transaction number from the batch transaction operation number queue to complete the synchronization of the next batch of transactions.

[0098] Optionally, the batch transaction operation number is the maximum global security transaction boundary, and all transaction operation requests smaller than this batch transaction operation number have been committed. There will be many committed transactions between this batch transaction operation number and the previous batch transaction operation number read, thus ensuring that there is no duplicate submission or omission of unsubmitted data between the target data update records processed by each reading of the target batch transaction operation number. Therefore, data synchronization based on the target batch transaction operation number in the transaction operation number queue guarantees the ultimate consistency of the transaction boundary and will not result in logical errors in the data.

[0099] Optionally, using M loading threads to synchronize the merged data of M source data tables to the target database in parallel includes: for the merged data in each source data table, perform variable assignment processing on the changed columns according to the primary key and transaction operation type to obtain dynamic data operation statements associated with each source data table; use M loading threads to synchronize the dynamic data operation statements associated with the M source data tables to the target database in parallel.

[0100] Optionally, after the merging process of each source data table is completed, traverse the linked list of each primary key in the data merging hash table corresponding to the source data table, and according to the transaction operation type corresponding to the data change operation in the linked list, assign the changed columns to the variables of the dynamic data operation statement to obtain the dynamic data operation statement associated with each source data table.

[0101] Optionally, start a loading thread for each source data table, and use M loading threads to batch synchronize the dynamic data operation statements associated with M source data tables into the target tables mapped to the source data tables in the target database in parallel.

[0102] Optionally, the method of parallel synchronous loading of multiple source data tables ensures parallel synchronous processing and effectively utilizes the central processing unit resources.

[0103] Optionally, the dynamic data operation statements include dynamic data insert operation statements, dynamic data update operation statements, and dynamic data delete operation statements; among them, for the merged data in each source data table, variable assignment processing is performed on the changed columns according to the primary key and the transaction operation type, and the dynamic data operation statements associated with each source data table are obtained as follows: for each piece of merged data, when the transaction operation type is the insert type, update the primary key and the changed columns into the dynamic data insert operation statement; for each piece of merged data, when the transaction operation type is the update type, based on the merge clause processing threshold, update the primary key and the changed columns into the dynamic data update operation statement; for each piece of merged data, when the transaction operation type is the delete type, based on the merge clause processing threshold, update the primary key into the dynamic data delete operation statement.

[0104] Optionally, traverse the linked list of each primary key in the data merging hash table corresponding to the source data table to generate dynamic data insert operation statements, dynamic data update operation statements, and dynamic data delete operation statements.

[0105] Optionally, for each piece of merged data, update the merged data into the dynamic data operation statement corresponding to the transaction operation type according to the primary key in the merged data and the transaction operation type of the data change operation in the linked list corresponding to the primary key.

[0106] Optionally, when the transaction operation type is the insert type, assign the primary key and the changed columns to the primary key variable and the column variable in the dynamic data insert operation statement.

[0107] Optionally, when the transaction operation type is the update type, when there is no same column variable value in the dynamic data update operation statement, assign the primary key and the changed columns to the primary key variable and the column variable in the dynamic data update operation statement.

[0108] Optionally, the configuration parameters include a merge clause parameter and a merge clause processing threshold. The merge clause processing threshold represents the upper limit for processing merge clauses. In the case where there are identical column variable values in the dynamic data update operation statement, if the merge clause processing threshold is not reached, the primary key is assigned to the primary key variable in the dynamic data update operation statement based on the merge clause. The merge clause can be an IN clause.

[0109] For example, a merge clause processing threshold of 1000 means that at most 1000 data update operations can be merged in a dynamic data update operation statement. For example, in the case where the column variable value corresponding to the column variable "score" in the dynamic data update operation statement is 80, if the merge clause processing threshold is not reached, the primary key is assigned to the primary key variable in the dynamic data update operation statement based on the merge clause. For example, the dynamic data update operation statement: UPDATE source data table 1 SET score = 80 WHERE department_id IN (10, 15, 20), where the primary key variable is department_id and the column variable score is "score". It can be seen that an IN clause merges three statements with primary key variables department_id being 10, 15, and 20.

[0110] Optionally, in the case where the transaction operation type is the delete type, if the merge clause processing threshold is not reached, the primary key is assigned to the primary key variable in the dynamic data delete operation statement based on the merge clause; if the merge clause processing threshold is reached, the primary key is directly assigned to the primary key variable in the dynamic data delete operation statement.

[0111] Optionally, if the merge clause parameter is set, when assigning variables, if the same transaction operation makes data changes to different primary keys, they will be merged into the merge clause.

[0112] For example, the dynamic data delete operation statement: DELETE FROM source data table 1 WHERE customer_id IN (21, 22, 23), where the primary key variable is customer_id. It can be seen that an IN clause merges three statements with primary key variables customer_id being 21, 22, and 23.

[0113] Optionally, dynamic variable assignment assigns the value of the data change operation when traversing each data change operation. When the set single-batch commit data volume parameter is reached, the dynamic data operation statement is executed.

[0114] Optionally, update the data change operation to a dynamic data operation statement, and put the dynamic data operation statement into a dynamic operation queue. When the set single-batch commit data volume parameter is reached, read and execute the dynamic data operation statement from the dynamic operation queue. The memory pool for data buffering in the database can allocate a plan for executing the dynamic data operation statement based on the single-batch commit data volume parameter, effectively utilizing the fragmented execution plan cache of the database for batch loading synchronization. In addition, optimize the dynamic data operation statement using the merge clause to batch-merge multiple original statements into one statement, effectively improving the synchronization performance and the performance of the target database.

[0115] Figure 3 FIG. shows a schematic diagram of a distributed database according to an embodiment of the present invention.

[0116] As Figure 3 shown, the distributed database includes source data nodes Node1, Node2, Node3, and GTM. GTM can be understood as a Global Transaction Manager (GTM). Logical replication slots record data update records on the source data nodes. Using logical replication slot reading threads: data thread 1, data thread 2, and data thread 3, logically replicate the data update records on each source data node to the KAFKA change data queues Node1_db1, Node2_db1, and Node3_db1 corresponding to the source data nodes Node1, Node2, and Node3 respectively. Use the batch transaction operation number reading thread (GTM thread) to periodically read a batch transaction operation number from the Global Transaction Manager (GTM) based on the batch transaction operation number reading interval parameter and store it in the batch transaction operation number queue Batch_GTM.

[0117] Optionally, read the target batch transaction operation number from the batch transaction operation number queue, and then screen N target data update records with transaction operation numbers smaller than the target batch transaction operation number from multiple change data queues. Based on the streaming merge algorithm, use 3 merge threads to parallelly perform data merging processing on the target data update records associated with the source data table 1 (or db1_table1), source data table 2 (or db1_table2), and source data table 3 (or db1_table3) respectively, and then generate dynamic data operation statements (or dynamic SQL). Assign the variables in the dynamic data operation statements corresponding to each source data table (db1_table1 dynamic SQL + value, db1_table2 dynamic SQL + value, db1_table3 dynamic SQL + value) to the KAFKA dynamic operation queue corresponding to each source data table. Use 3 loading threads (db1_table1 loading thread 1, db1_table2 loading thread 2, db1_table3 loading thread 3) to read and execute the dynamic data operation statements from the KAFKA dynamic operation queue, and then synchronize the merged data of each source data table to the target database.

[0118] Based on the above data synchronization method based on a distributed database, the present invention also provides a data synchronization device based on a distributed database. The following will be combined with Figure 4 to describe this device in detail.

[0119] Figure 4 Shows a structural block diagram of a data synchronization device based on a distributed database according to an embodiment of the present invention.

[0120] As Figure 4 shown, the data synchronization device 400 based on a distributed database in this embodiment includes a reading module 410, a determination module 420, a screening module 430, a merging module 440, and a synchronization module 450.

[0121] The reading module 410 is configured to read the target batch transaction operation number corresponding to the current batch transaction to be processed from the batch transaction operation number queue in response to a batch transaction processing request. In one embodiment, the reading module 410 can be used to execute the operation S210 described above, which will not be elaborated here.

[0122] The determination module 420 is configured to determine the change data queues corresponding to L source data nodes respectively. Among them, the change data queues store multiple groups of data update records generated by the source data nodes through transaction operations during the data synchronization process, and the multiple groups of data update records correspond to multiple transaction operation numbers. In one embodiment, the determination module 420 can be used to execute the operation S220 described above, which will not be elaborated here.

[0123] A screening module 430 is configured to screen, from L change data queues, N target data update records whose transaction operation numbers are less than a target batch transaction operation number. The N target data update records are associated with M source data tables, and L source data nodes jointly complete the data storage of the M source data tables. In one embodiment, the screening module 430 may be configured to perform the operation S230 described above, which will not be elaborated herein.

[0124] A merging module 440 is configured to, based on a streaming merge algorithm, use M merge threads to perform a parallel merging process on the target data update records respectively associated with the M source data tables, so as to obtain M data merge hash tables corresponding to the M source data tables. Among them, the m-th source data table is associated with a target number of data update records, and the data merge hash table corresponding to the m-th source data table stores records associated with a primary key. The primary key is derived from the target data update records, and each merge result includes at most one merged data. , . In one embodiment, the merging module 440 may be configured to perform the operation S240 described above, which will not be elaborated herein.

[0125] A synchronization module 450 is configured to use M loading threads to synchronize the merged data of the M source data tables into the target database in parallel. In one embodiment, the synchronization module 450 may be configured to perform the operation S250 described above, which will not be elaborated herein.

[0126] Optionally, the data synchronization device 400 based on a distributed database further includes a response module, a change module, a synchronization module, and a storage module.

[0127] The response module is configured to, in response to a transaction operation request for each source data table, determine a transaction operation number and i target source data table sub-tables to be operated, where the transaction operation number is allocated from a global transaction manager. The source data table includes I source data table sub-tables, the i target source data table sub-tables are derived from the I source data table sub-tables, the source data nodes correspond to the source data table sub-tables one by one, and i ≤ I.

[0128] The change module is configured to perform a data change operation on the i target source data table sub-tables to obtain i data update records corresponding to the i target source data table sub-tables. Each data update record includes a transaction operation number, a transaction operation type, a target source data table sub-table name, a primary key, and a changed column.

[0129] The synchronization module is configured to synchronously store the i data update records in i logical replication slots corresponding to the i target source data nodes where the i target source data table sub-tables are located respectively.

[0130] A storage module, configured to use i read threads to parallelly store data update records in i logical replication slots into i change data queues corresponding to i target source data nodes.

[0131] Optionally, the data synchronization device 400 based on a distributed database further includes a sorting module, a processing module, and an obtaining module.

[0132] The sorting module is configured to, for any m-th source data table, perform a sorting process on the target data update records associated with the m-th source data table based on a transaction operation number, to obtain a data sorting hash table corresponding to the m-th source data table, where the data sorting hash table corresponding to the m-th source data table stores records associated with primary keys, and the sorting result corresponding to the h-th primary key includes ordered target data update records. , .

[0133] The processing module is configured to, for the sorting result corresponding to the h-th primary key in the sorting hash table, based on a streaming merge algorithm, use merge threads to perform a merge process on the ordered target data update records, to obtain a merge result corresponding to the h-th primary key.

[0134] The obtaining module is configured to obtain a data merge hash table corresponding to the m-th source data table according to primary keys and the merge result corresponding to each primary key.

[0135] Optionally, the processing module includes a first processing sub-module, a second processing sub-module, and a third processing sub-module.

[0136] The first processing sub-module is configured to traverse the ordered target data update records.

[0137] The second processing sub-module is configured to, for each target data update record, update the data change operation in the linked list corresponding to the h-th primary key according to the h-th primary key and the transaction operation type.

[0138] The third processing sub-module is configured to obtain the merge result corresponding to the h-th primary key according to the data change operation in the linked list corresponding to the h-th primary key.

[0139] Optionally, any number of the reading module 410, the determining module 420, the filtering module 430, the merging module 440, and the synchronization module 450 may be combined and implemented in one module, or any one of them may be split into multiple modules. Alternatively, at least part of the functions of one or more of these modules may be combined with at least part of the functions of other modules and implemented in one module. Optionally, at least one of the reading module 410, the determining module 420, the filtering module 430, the merging module 440, and the synchronization module 450 may be at least partially implemented as a hardware circuit, such as a field programmable gate array (FPGA), a programmable logic array (PLA), a system on chip, a system on substrate, a system on package, an application specific integrated circuit (ASIC), or any other reasonable way of integrating or packaging circuits, etc., implemented by hardware or firmware, or implemented in any one of the three implementation manners of software, hardware, and firmware, or in any appropriate combination of several of them. Alternatively, at least one of the reading module 410, the determining module 420, the filtering module 430, the merging module 440, and the synchronization module 450 may be at least partially implemented as a computer program module, which can execute corresponding functions when the computer program module is run.

[0140] Figure 5 FIG. shows a block diagram of an electronic device suitable for implementing a data synchronization method based on a distributed database according to an embodiment of the present invention.

[0141] Figure 5 The illustrated electronic device is only an example and should not impose any limitation on the functions and usage scope of the embodiments of the present invention.

[0142] As Figure 5 shown, the computer electronic device 500 according to an embodiment of the present invention includes a processor 501, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 502 or a program loaded from a storage section 508 into a random access memory (RAM) 503. The processor 501 may include, for example, a general microprocessor (such as a CPU), an instruction set processor, and / or a related chipset, and / or a dedicated microprocessor (such as an application specific integrated circuit (ASIC)), and so on. The processor 501 may also include on-board memory for caching purposes. The processor 501 may include a single processing unit or multiple processing units for performing different actions of the method flow according to an embodiment of the present invention.

[0143] In the RAM 503, various programs and data required for the operation of the electronic device 500 are stored. The processor 501, the ROM 502, and the RAM 503 are connected to each other via the bus 504. The processor 501 performs various operations of the method flow according to the embodiments of the present invention by executing the programs in the ROM 502 and / or the RAM 503. It should be noted that the programs can also be stored in one or more memories other than the ROM 502 and the RAM 503. The processor 501 can also perform various operations of the method flow according to the embodiments of the present invention by executing the programs stored in one or more memories.

[0144] Optionally, the electronic device 500 may further include an input / output (I / O) interface 505, and the input / output (I / O) interface 505 is also connected to the bus 504. The electronic device 500 may further include one or more of the following components connected to the input / output (I / O) interface 505: an input portion 506 including a keyboard, a mouse, etc.; an output portion 507 including, for example, a cathode ray tube (CRT), a liquid crystal display (LCD), etc. and a speaker, etc.; a storage portion 508 including a hard disk, etc.; and a communication portion 509 including a network interface card such as a LAN card, a modem, etc. The communication portion 509 performs communication processing via a network such as the Internet. The drive 510 is also connected to the input / output (I / O) interface 505 as needed. A removable medium 511, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 510 as needed so that a computer program read from it can be installed into the storage portion 508 as needed.

[0145] Optionally, the method flow according to the embodiments of the present invention can be implemented as a computer software program. For example, an embodiment of the present invention includes a computer program product, which includes a computer program carried on a computer-readable storage medium, and the computer program includes program codes for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from the network through the communication portion 509, and / or installed from the removable medium 511. When the computer program is executed by the processor 501, the above functions defined in the system according to the embodiments of the present invention are executed. Optionally, the systems, devices, apparatuses, modules, units, etc. described above can be implemented by computer program modules.

[0146] The present invention also provides a computer-readable storage medium, which can be included in the device / device / system described in the above embodiments; or can exist separately without being assembled into the device / device / system. The above computer-readable storage medium carries one or more programs, and when the one or more programs are executed, a data synchronization method based on a distributed database according to the embodiments of the present invention is implemented.

[0147] Optionally, the computer-readable storage medium may be a non-volatile computer-readable storage medium. For example, it may include, but is not limited to: portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the above. In the present invention, the computer-readable storage medium may be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, apparatus, or device.

[0148] For example, optionally, the computer-readable storage medium may include one or more memories other than the above-described ROM 502 and / or RAM 503 and / or ROM 502 and RAM 503.

[0149] An embodiment of the present invention further includes a computer program product, which includes a computer program that contains program code for executing the method provided by the embodiment of the present invention. When the computer program product runs on an electronic device, the program code is used to cause the electronic device to implement the data synchronization method based on a distributed database provided by the embodiment of the present invention.

[0150] When the computer program is executed by the processor 501, the above functions defined in the system / apparatus of the embodiment of the present invention are executed. Optionally, the above-described systems, apparatuses, modules, units, etc. may be implemented by computer program modules.

[0151] In one embodiment, the computer program may rely on tangible storage media such as optical storage devices, magnetic storage devices, etc. In another embodiment, the computer program may also be transmitted and distributed in the form of a signal on a network medium and downloaded and installed through the communication part 509, and / or installed from the removable medium 511. The program code contained in the computer program may be transmitted using any suitable network medium, including but not limited to: wireless, wired, etc., or any suitable combination of the above.

[0152] Optionally, the program code for executing the computer program provided by the embodiments of the present invention can be written in any combination of one or more programming languages. Specifically, these computing programs can be implemented using high-level procedural and / or object-oriented programming languages, and / or assembly / machine languages. Programming languages include, but are not limited to, such as Java, C++, Python, the "C" language, or similar programming languages. The program code can be executed entirely on the user's computing device, partially on the user's device, partially on a remote computing device, or entirely on a remote computing device or server. In the case of a remote computing device, the remote computing device can be connected to the user's computing device through any type of network, including a local area network (LAN) or a wide area network (WAN), or it can be connected to an external computing device (for example, by using an Internet service provider to connect through the Internet).

[0153] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in the flowchart or block diagram may represent a module, a program segment, or a part of code that contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than that marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram or flowchart, and the combination of blocks in the block diagram or flowchart, can be implemented by a dedicated hardware-based system for performing the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions. Those skilled in the art can understand that the features described in the various embodiments of the present invention can be combined and / or combined in various ways, even if such combinations or combinations are not explicitly described in the present invention. In particular, without departing from the spirit and teachings of the present invention, the features described in the various embodiments of the present invention can be combined and / or combined in various ways. All such combinations and / or combinations fall within the scope of the present invention.

[0154] The above describes the embodiments of the present invention. However, these embodiments are only for illustrative purposes and not for limiting the scope of the present invention. Although the embodiments are described separately above, this does not mean that the measures in each embodiment cannot be used advantageously in combination. Without departing from the scope of the present invention, those skilled in the art can make various substitutions and modifications, and all such substitutions and modifications should fall within the scope of the present invention.

Claims

1. A data synchronization method based on a distributed database, characterized in that: The distributed database includes a source database and a target database to be synchronized with each other, the source database includes L source data nodes; the method includes: In response to a transaction operation request for each source data table, determine a transaction operation number and i target source data table sub-tables to be operated, wherein the transaction operation number is allocated from a global transaction manager, the source data table includes I source data table sub-tables, the i target source data table sub-tables are derived from the I source data table sub-tables, the source data nodes correspond to the source data table sub-tables one by one, and i≤I; Performing data modification operations on i sub-tables of the target source data table to obtain i data update records corresponding to the i sub-tables of the target source data table; Synchronously storing i data update records in i change data queues corresponding to i target source data nodes; Based on a batch transaction operation number reading interval parameter, read from a global transaction manager a plurality of transaction operation numbers corresponding to an execution state representing an uncommitted state; Selecting the smallest transaction operation number from the plurality of transaction operation numbers as the target transaction operation number; Determine a transaction operation number preceding the target transaction operation number as a batch transaction operation number, wherein the execution status of the batch transaction operation number and the transaction operation number preceding the batch transaction operation number both represent a commit status; Storing the batch transaction operation number in a batch transaction operation number queue; In response to a batch transaction processing request, reading a target batch transaction operation number corresponding to a current batch transaction to be processed from the batch transaction operation number queue, wherein the target batch transaction operation number represents a minimum batch transaction operation number read from the batch transaction operation number queue based on a response time; Determine a changed data queue corresponding to each of the L source data nodes; Filtering N target data update records whose transaction operation numbers are smaller than the target batch transaction operation numbers from the L change data queues, wherein the N target data update records are associated with the M source data tables; Based on the stream merging algorithm, M merging threads are used to merge the target data update records associated with the M source data tables in parallel to obtain M data merging hash tables corresponding to the M source data tables, wherein the mth source data table is associated with The target data update record is stored in the data merge hash table corresponding to the mth source data table. Primary key associated merge results, the primary key is derived from the target data update record, each of the merge results contains at most one merged data, , ; The merged data of the M source data tables are synchronized to the target database in parallel using M loading threads.

2. The method according to claim 1, characterized in that: Each of the data update records includes the transaction operation number, transaction operation type, target source data table sub-table name, the primary key, and the changed column; The step of synchronously storing the i data update records in the i change data queues corresponding to the i target source data nodes includes: Synchronously storing the i data update records in the i logical replication slots corresponding to the target source data nodes where the i target source data table sub-tables are located; The data update records in the i logical replication slots are stored in parallel by using i reading threads into the i change data queues corresponding to the i target source data nodes.

3. The method according to claim 2, characterized in that Before the merging process, the method further includes: For any m-th source data table, based on the transaction operation number, The target data update records are sorted to obtain a data sorting hash table corresponding to the m-th source data table, wherein the data sorting hash table corresponding to the m-th source data table stores Primary key associated The sorting result corresponding to the hth primary key contains the ordered the target data update record, , ; Wherein, for any of the m-th source data tables, the merging process includes: For the sorting result corresponding to the hth primary key in the sorted hash table, based on the stream merge algorithm, the merge thread is used to merge the ordered merging the target data update records to obtain a merging result corresponding to the hth primary key; according to The primary keys and the merging results corresponding to each of the primary keys are used to obtain the data merging hash table corresponding to the mth source data table.

4. The method according to claim 3, characterized in that: For the sorting result corresponding to the hth primary key, based on the stream merging algorithm, the merging thread is used to merge the ordered The target data update records are merged to obtain the merge result corresponding to the hth primary key, including: Traverse the ordered update record of the target data; For each target data update record, according to the hth primary key and the transaction operation type, update the data change operation in the linked list corresponding to the hth primary key; According to the data modification operation in the linked list corresponding to the hth primary key, the merging result corresponding to the hth primary key is obtained.

5. The method according to claim 4, characterized in that For each target data update record, according to the hth primary key and the transaction operation type, the data modification operation in the linked list corresponding to the hth primary key is updated, including: For each of the target data update records, if the h-th primary key does not exist in the multiple candidate target data update records in the data sorting hash table, add a data change operation in the linked list corresponding to the h-th primary key, wherein the transaction operation number in the candidate target data update record is smaller than the transaction operation number in the target data update record; For each of the target data update records, if the hth primary key exists in the multiple candidate target data update records and the transaction operation type is an insert type, clear the linked list corresponding to the hth primary key and add a current data insert operation; For each target data update record, if the hth primary key exists in the multiple candidate target data update records and the transaction operation type is an update type, traverse the linked list corresponding to the hth primary key, and add a current data update operation if there is no historical data update operation for the changed column in the linked list; For each of the target data update records, if the hth primary key exists in the multiple candidate target data update records and the transaction operation type is a deletion type, clear the linked list corresponding to the hth primary key, and add a current data deletion operation if there is no historical data insertion operation for the hth primary key in the linked list.

6. The method according to claim 1, characterized in that The method further comprises: Determine the execution status of the transaction operation request corresponding to the transaction operation number.

7. The method according to claim 2, characterized in that: Using M loading threads to synchronize the merged data of the M source data tables to the target database in parallel includes: For the merged data in each of the source data tables, variable assignment processing is performed on the change column according to the primary key and the transaction operation type to obtain a dynamic data operation statement associated with each of the source data tables; The dynamic data operation statements associated with the M source data tables are synchronized to the target database in parallel using M loading threads.

8. The method according to claim 7, characterized in that The dynamic data operation statements include dynamic data insert operation statements, dynamic data update operation statements, and dynamic data delete operation statements; Among them, for the merged data in each of the source data tables, performing variable assignment processing on the change column according to the primary key and the transaction operation type, and obtaining the dynamic data operation statement associated with each of the source data tables includes: For each piece of the merged data, when the transaction operation type is an insert type, the primary key and the change column are updated into the dynamic data insert operation statement; For each of the merged data, when the transaction operation type is an update type, based on a merge clause processing threshold, updating the primary key and the change column into the dynamic data update operation statement; For each piece of the merged data, when the transaction operation type is a deletion type, based on the merge clause processing threshold, the primary key is updated into the dynamic data deletion operation statement.

9. A data synchronization device based on a distributed database, characterized in that: The device comprises: A response module, configured to respond to a transaction operation request for each source data table, determine a transaction operation number and i target source data table sub-tables to be operated, wherein the transaction operation number is allocated from a global transaction manager, the source data table includes I source data table sub-tables, the i target source data table sub-tables are derived from the I source data table sub-tables, and source data nodes correspond to the source data table sub-tables one by one, i≤I; A modification module, used for performing data modification operations on i sub-tables of the target source data table to obtain i data update records corresponding to the i sub-tables of the target source data table; A storage module, used for synchronously storing i data update records in i change data queues corresponding to i target source data nodes; A first reading module, configured to read, based on a batch transaction operation number reading interval parameter, a plurality of transaction operation numbers corresponding to an execution state representing an uncommitted state from a global transaction manager; A selection module, used for selecting a minimum transaction operation number from the plurality of transaction operation numbers as a target transaction operation number; A first determining module is used to determine a transaction operation number preceding the target transaction operation number as a batch transaction operation number, wherein the execution status of the batch transaction operation number and the transaction operation number preceding the batch transaction operation number both represent a commit status; A sending module, used for storing the batch transaction operation number into a batch transaction operation number queue; A reading module, configured to read, in response to a batch transaction processing request, a target batch transaction operation number corresponding to a current batch transaction to be processed from the batch transaction operation number queue, wherein the target batch transaction operation number represents a minimum batch transaction operation number read from the batch transaction operation number queue based on a response time; A determination module, used to determine the changed data queues corresponding to the L source data nodes respectively; A screening module, used for screening N target data update records whose transaction operation numbers are smaller than the target batch transaction operation numbers from the L change data queues, wherein the N target data update records are associated with the M source data tables; A merging module is used to merge the target data update records associated with the M source data tables in parallel using M merging threads based on a stream merging algorithm to obtain M data merging hash tables corresponding to the M source data tables, wherein the mth source data table is associated with The target data update record is stored in the data merge hash table corresponding to the mth source data table. Primary key associated merge results, the primary key is derived from the target data update record, each of the merge results contains at most one merged data, , ; A synchronization module is used to synchronize the merged data of the M source data tables to the target database in parallel using M loading threads.

10. An electronic device, characterized in that: include: one or more processors; a memory for storing one or more instructions, Wherein, when the one or more instructions are executed by the one or more processors, the one or more processors are enabled to implement the method according to any one of claims 1 to 8.

Citation Information

Patent Citations

  • Database synchronization method, device and equipment and computer storage medium

    CN114490865A

  • Method and device for executing database transaction

    CN115827172A