A distributed database data synchronization method and device

By adopting a multi-node concurrent synchronization architecture and table lock mechanism in a distributed database cluster, the DML and DDL operation timing problems are solved, data synchronization performance is improved, and low performance under massive data is avoided.

CN115422286BActive Publication Date: 2025-05-23WUHAN DAMENG DATABASE
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211003580.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-08-19
Publication Date
2025-05-23
Estimated Expiration
2042-08-19

AI Technical Summary

Technical Problem

In distributed database clusters, timing problems of DML and DDL operations lead to data synchronization timing errors. The existing technology has low performance under massive data conditions and high cost of capturing and sorting logs.

Method used

A multi-node concurrent synchronization architecture is adopted, each node is synchronized independently, DDL operations are obtained from metadata management nodes, and a table lock mechanism is used to avoid repeated DDL entry into the library.

Benefits of technology

It ensures the timing and consistency of data operations on each node table, improves the overall performance of data synchronization, and reduces the problem of low performance under massive data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115422286B_ABST
    Figure CN115422286B_ABST
Patent Text Reader

Abstract

The present invention relates to a data synchronization method and device for a distributed database. The method part mainly includes: deploying a source-end data synchronization system on a source-end database and deploying a target-end data synchronization system on a target-end; in a metadata management node operation mode, the source-end data synchronization system reads, parses, and caches logs from the source-end database; in a data node operation mode, the source-end data synchronization system reads, parses logs from the source-end database and packages the logs and sends them to a target-end data synchronization system; the target-end data synchronization system unpacks the message packets sent by the source-end data synchronization system after receiving them, and applies these unpacking operations to the target-end database. The present invention realizes data parallel synchronization by independently deploying synchronization software on each node of a distributed database, thereby greatly improving data synchronization performance.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of database data processing, and in particular to a data synchronization method and device for a distributed database. Background Art

[0002] Data Manipulation Language (DML) uses the "Select", "Insert", "Update" and "Delete" keywords to manipulate data. Data Definition Language (DDL) is used to create and manipulate table structures.

[0003] In a distributed database cluster system, due to the characteristics of distributed transactions, transaction logs are scattered on different nodes, and at the same time, DDL operations may be processed on another metadata management node. In other words, the operation logs of table data are scattered on various nodes of the distributed cluster, and DDL operations are distributed on other dedicated nodes. In this case, if data synchronization and DDL synchronization need to be supported, then the timing of DDL and DML operations, that is, the order of precedence, needs to be considered. For example, the table creation operation is on the distributed node EP0, and the table insertion operation is on the distributed node EP1. During synchronization, if the insertion operation is synchronized first and then the table creation operation, there will be a timing error and the insertion operation will report an error. Therefore, how to correctly synchronize table DML and DDL operations has become an urgent problem to be solved.

[0004] To address this problem, the current common solution is to merge the logs of all nodes, then sort the logs, and finally determine the operation sequence of DDL and DML, so as to synchronize DML and DDL normally. The disadvantage of this solution is that it captures the logs of all nodes and sorts and merges the logs at the same time. In the case of massive data, this is very costly and has low performance.

[0005] In view of this, how to overcome the defects of the existing technology and how to solve the above-mentioned technical problems have become important technical problems that the industry needs to solve urgently. Summary of the invention

[0006] In view of the above defects or improvement needs of the prior art, the present invention provides a data synchronization method and device for a distributed database. Based on the distributed database, the DML and DDL operation logs are distributed on different nodes. The present invention adopts a multi-node concurrent synchronization architecture. The DDL operations involved in each node are obtained from the metadata management node, so that the timing and consistency of the table data operations of each node can be guaranteed. When the DDL on the target side is stored in the warehouse, a lock table mechanism is used to ensure that each DDL will not be stored repeatedly. In the synchronization architecture of the present invention, the synchronization of each node is independent, the synchronized transactions are dispersed, and the synchronized data is complete and consistent as a whole, so that concurrent processing can be performed to improve the overall performance of synchronization.

[0007] The embodiment of the present invention adopts the following technical solution:

[0008] In a first aspect, the present invention provides a data synchronization method for a distributed database, comprising:

[0009] A source-side data synchronization system is deployed on the source-side database, and a target-side data synchronization system is deployed on the target-side. The source-side data synchronization system includes a metadata management node operation mode and a data node operation mode.

[0010] In the metadata management node operation mode, the source-end data synchronization system initializes a log reading thread, a log parsing thread, and a log caching thread to read, parse, and cache logs from the source-end database;

[0011] In the data node operation mode, the source-end data synchronization system initializes a log reading thread, a log parsing thread, and a log sending thread, which are used to read and parse logs from the source-end database and package and send the logs to the target-end data synchronization system;

[0012] The target-side data synchronization system unpacks the message package after receiving it from the source-side data synchronization system, and applies these unpacking operations to the target-side database. For the unpacking operation, if it is a DDL operation, the target-side synchronization system uses the DDL collaborative warehousing strategy to perform DDL warehousing; if it is a non-DDL operation, it is directly warehousing.

[0013] Furthermore, the DDL collaborative storage strategy of the target-side synchronization system includes:

[0014] The target-side data synchronization system initializes a DDL collaborative storage auxiliary table DDL_SYNC on the target-side database; the STATE field of the DDL_SYNC table indicates the status information, and the specific values ​​include: "over" indicates that the corresponding DDL has been stored; "ready" indicates that the corresponding DDL synchronization is ready; the total number of data nodes is set to M, and the variable i is set to 0;

[0015] Query the DDL_SYNC table to obtain the number of data rows with STATE value "over" n1;

[0016] If n1 is not greater than 0, the DDL operation type is determined and different procedures are used to set an X lock on the DDL_SYNC table according to the operation type.

[0017] Get the number of data rows n3 whose STATE value is "over" in the DDL_SYNC table;

[0018] If n3 is not greater than 0, the DDL storage operation is executed and the X lock of the DDL_SYNC table is released.

[0019] Furthermore, if n1 is not greater than 0, the DDL operation type is determined and different processes are used to lock the X on the DDL_SYNC table according to different operation types. Specifically, the following steps are performed:

[0020] If n1 is not greater than 0, determine whether the DDL operation is an ALTER or TRUNCATE operation;

[0021] If it is not an ALTER or TRUNCATE operation, an X lock is directly set on the DDL_SYNC table;

[0022] If it is an ALTER or TRUNCATE operation, obtain the number of data rows n2 with a STATE value of "ready" from the DDL_SYNC table, and determine whether n2+1 is equal to M; if n2+1 is equal to M, set an X lock on the DDL_SYNC table; if n2+1 is not equal to M, determine whether i is equal to 0; if i is equal to 0, insert the current DDL information with a STATE value of "ready" into the DDL_SYNC table, set i=i+1, wait for 1 second and then re-enter the n1 acquisition step; if i is not equal to 0, directly wait for 1 second and then re-enter the n1 acquisition step.

[0023] Furthermore, when the number of data rows n1 with a STATE value of "over" is obtained by querying from the DDL_SYNC table, if n1 is greater than 0, the DDL operation is skipped, that is, the DDL storage operation is not performed, and the DDL synchronization is directly ended.

[0024] Furthermore, if n3 is not greater than 0, executing the DDL storage operation and releasing the X lock of the DDL_SYNC table specifically includes:

[0025] If n3 is not greater than 0, insert the current DDL information with STATE value "over" into the DDL_SYNC table;

[0026] Execute the storage operation of the current DDL and release the X lock of the DDL_SYNC table;

[0027] End this DDL synchronization.

[0028] Furthermore, when obtaining the number of data rows n3 whose STATE value of the DDL_SYNC table is "over", if n3 is greater than 0, the X lock of the DDL_SYNC table is directly released, and then the DDL operation is skipped, that is, the DDL storage operation is not executed, and the DDL synchronization is directly ended.

[0029] Furthermore, the source-end database includes a distributed database cluster; the target end includes one or more of other data sources, a single-node database, and a multi-node cluster system.

[0030] Furthermore, in the metadata management node operation mode, the source-side data synchronization system initializes a log reading thread, a log parsing thread, and a log caching thread, which are used to read, parse, and cache logs from the source-side database, and specifically includes:

[0031] The source data synchronization system corresponding to the metadata management node initializes a log reading thread, a log parsing thread, and a log caching thread;

[0032] The log reading thread is used to read database logs and add the read logs to the queue to be parsed;

[0033] The log parsing thread is used to obtain logs from the queue to be parsed and parse them into transaction information to be processed, and add them to the log cache queue;

[0034] The log cache thread is used to obtain log information from the cache queue and cache it according to the transaction ID classification.

[0035] Furthermore, in the data node operation mode, the source-end data synchronization system initializes a log reading thread, a log parsing thread, and a log sending thread, which are used to read and parse logs from the source-end database and package and send the logs to the target-end data synchronization system, specifically including:

[0036] The source data synchronization system corresponding to the data node initializes a log reading thread, a log parsing thread, and a log sending thread. The log parsing thread includes a DDL log request module, which is used to obtain the relevant DDL operation log from the metadata management node according to the transaction ID.

[0037] The log reading thread is used to read logs from the corresponding data node and add the read logs to the log parsing queue;

[0038] The log parsing thread is used to parse the logs in the queue to be parsed. If a DDL request log is encountered, the DDL log request module is used to obtain the DDL-related logs from the source data synchronization system corresponding to the metadata management node and add them to the queue to be parsed. The log parsing thread is also used to package the parsed logs into the internal message format of the synchronization system and add them to the queue to be sent.

[0039] The log sending thread is used to send the messages in the queue to be sent to the target data synchronization system.

[0040] On the other hand, the present invention provides a data synchronization device for a distributed database, specifically: comprising at least one processor and a memory, the at least one processor and the memory are connected via a data bus, the memory stores instructions that can be executed by the at least one processor, and after the instructions are executed by the processor, they are used to complete the data synchronization method for the distributed database in the first aspect.

[0041] Compared with the prior art, the beneficial effects of the present invention are as follows: the present invention provides a data synchronization method and device for a distributed database. Based on the distributed database, the DML and DDL operation logs are distributed on different nodes. The present invention adopts a multi-node concurrent synchronization architecture. The DDL operations involved in each node are obtained from the metadata management node, so that the timing and consistency of the table data operations of each node can be guaranteed. When the DDL on the target side is stored in the warehouse, a lock table mechanism is used to ensure that each DDL will not be stored repeatedly. In the synchronization architecture of the present invention, the synchronization of each node is independent, the synchronized transactions are dispersed, and the synchronized data is complete and consistent as a whole, so that concurrent processing can be performed to improve the overall performance of synchronization. BRIEF DESCRIPTION OF THE DRAWINGS

[0042] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for use in the embodiments of the present invention. Obviously, the drawings described below are only some embodiments of the present invention, and for ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.

[0043] Figure 1 A flow chart of a distributed database data synchronization method provided in Example 1 of the present invention;

[0044] Figure 2 A specific flow chart of step 200 provided in Example 1 of the present invention;

[0045] Figure 3 A specific flow chart of step 300 provided in Example 1 of the present invention;

[0046] Figure 4 A flowchart of the DDL collaborative warehousing strategy provided in Example 1 of the present invention;

[0047] Figure 5 A specific flow chart of step 430 provided in Example 1 of the present invention;

[0048] Figure 6 A specific flow chart of step 450 provided in Example 1 of the present invention;

[0049] Figure 7 A processing flow chart of a DDL log request module provided in Example 2 of the present invention;

[0050] Figure 8 A diagram of a distributed database data synchronization architecture based on log parsing provided in Example 3 of the present invention;

[0051] Fig. 9 A target-side data synchronization system DDL collaborative warehousing flow chart provided in Example 3 of the present invention;

[0052] Fig.10 A schematic diagram of the structure of a data synchronization device for a distributed database provided in Example 4 of the present invention. DETAILED DESCRIPTION

[0053] In order to make the purpose, technical solution and advantages of the present invention clearer, the technical solution in the embodiment of the present invention will be clearly and completely described in conjunction with the drawings in the embodiment of the present invention. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention, and the embodiments described herein are part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments in the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention. In addition, the technical features in the various embodiments or single embodiments provided by the present invention can be combined with each other arbitrarily to form a feasible technical solution, but it must be based on the ability of ordinary technicians in this field to achieve it. When the combination of technical solutions is contradictory or cannot be achieved, it should be considered that such a combination of technical solutions does not exist and is not within the scope of protection required by the present invention.

[0054] The present invention is an architecture of a specific functional system, so the specific embodiments mainly illustrate the functional logical relationship between the various structural modules, and do not limit the specific software and hardware implementation methods.

[0055] It should be noted that distributed databases have the following characteristics when recording DML and DDL logs:

[0056] There is only one metadata management master node, and all metadata (DDL) operations are recorded in detail in the log of this node.

[0057] There are multiple data nodes, and there can be multiple copies. The data of DML operations on a data node is only recorded in the log of this node.

[0058] When a transaction spans multiple nodes, the transaction ID recorded on each node is the same; when the transaction ends (committed or rolled back), the relevant nodes record the end information of the transaction in the log.

[0059] A DDL transaction will record a DDL identification log in each data node. The log contains information such as transaction ID, LSN, metadata node number, current node number, and numbers of other data nodes involved.

[0060] The present invention is used to solve the problem of low synchronization performance in the case of massive data according to the characteristics of distributed database clusters, independently deploy synchronization software on each node of the distributed database, realize data parallel synchronization, and thus greatly improve data synchronization performance.

[0061] Based on the above-mentioned actual situation, an embodiment of the present invention provides a method and device for data synchronization of a distributed database. The present invention will be described in detail below with reference to the accompanying drawings and embodiments.

[0062] Embodiment 1:

[0063] like Figure 1 As shown, an embodiment of the present invention provides a data synchronization method for a distributed database, and the specific steps are as follows.

[0064] Step 100: deploy a source-side data synchronization system on the source-side database and deploy a target-side data synchronization system on the target side; wherein the source-side data synchronization system includes a metadata management node operation mode and a data node operation mode.

[0065] Step 200: In the metadata management node operation mode, the source-end data synchronization system initializes a log reading thread, a log parsing thread, and a log caching thread for reading, parsing, and caching logs from a source-end database.

[0066] Step 300: In the data node operation mode, the source-end data synchronization system initializes a log reading thread, a log parsing thread, and a log sending thread for reading and parsing logs from the source-end database and packaging and sending the logs to the target-end data synchronization system.

[0067] Step 400: After receiving the message package sent by the source data synchronization system, the target-side data synchronization system unpacks it and applies these unpacking operations to the target-side database. For the unpacking operation, if it is a DDL operation, the target-side synchronization system uses the DDL collaborative warehousing strategy to execute DDL warehousing; if it is a non-DDL operation, it is directly warehousing.

[0068] The above are the basic steps of this preferred embodiment, which realizes data parallel synchronization by independently deploying synchronization software on each node of the distributed database, thereby greatly improving data synchronization performance. The following describes each step in detail to further explain the solution of this preferred embodiment.

[0069] For step 100 of this preferred embodiment (deploy a source-end data synchronization system on the source-end database, and deploy a target-end data synchronization system on the target-end), the source-end database includes a distributed database cluster; the target-end includes one or more of other data sources, single-node databases, and multi-node cluster systems, that is, the target end can be other data sources, a general single-node database, or a multi-node cluster system. In addition, when the synchronization system is deployed on the source-end database and the target-end data source, a set of source-end data synchronization systems is deployed on each node of the distributed database on the source end, and the source-end data synchronization system is divided into two operating modes according to the mode of the corresponding node: data node (such as a node for DML operations) operating mode and metadata management node (such as a node for DDL operations) operating mode.

[0070] For step 200 of this preferred embodiment (in the metadata management node operation mode, the source-side data synchronization system initializes the log reading thread, the log parsing thread, and the log caching thread for reading, parsing, and caching logs from the source-side database), refer to Figure 2 , which can be specifically expanded to include the following steps.

[0071] Step 201: The source-end data synchronization system corresponding to the metadata management node initializes a log reading thread, a log parsing thread, and a log caching thread.

[0072] Step 202: The log reading thread corresponding to the metadata management node is used to read the database log and add the read log to the queue to be parsed.

[0073] Step 203: The log parsing thread corresponding to the metadata management node is used to obtain logs from the queue to be parsed and parse them into transaction information to be processed, and add them to the log cache queue.

[0074] Step 204: The log cache thread corresponding to the metadata management node is used to obtain log information from the queue to be cached, and cache it according to the transaction ID classification.

[0075] For step 300 of this preferred embodiment (in the data node operation mode, the source-side data synchronization system initializes a log reading thread, a log parsing thread, and a log sending thread for reading and parsing logs from the source-side database and packaging and sending the logs to the target-side data synchronization system), refer to Figure 3 , which can be specifically expanded to include the following steps.

[0076] Step 301: The source data synchronization system corresponding to the data node initializes a log reading thread, a log parsing thread, and a log sending thread. The log parsing thread corresponding to the data node includes a DDL log request module for obtaining the relevant DDL operation log from the metadata management node according to the transaction ID.

[0077] Step 302: The log reading thread corresponding to the data node is used to read the log from the corresponding data node, and add the read log to the log to be parsed queue.

[0078] Step 303: The log parsing thread corresponding to the data node is used to parse the logs in the queue to be parsed. If a DDL request log is encountered, the DDL-related logs are obtained from the source data synchronization system corresponding to the metadata management node through the DDL log request module, and added to the log queue to be parsed; the log parsing thread corresponding to the data node is also used to package the parsed logs into the internal message format of the synchronization system, and add them to the message queue to be sent.

[0079] Step 304: The log sending thread corresponding to the data node is used to send the messages in the queue to be sent to the target end data synchronization system.

[0080] For step 400 of this preferred embodiment, the target-side data synchronization system will first initialize a DDL synchronization auxiliary table, which records the detailed information of the DDL operation, including: transaction number, LSN, operation type, object name, status, number of data nodes, current node number, other node information involved, etc.; the target-side data synchronization system is responsible for unpacking the synchronization message sent by the source end, and applying these unpacking operations to the target-side database. If it is a DDL operation, the target-side synchronization system uses the DDL collaborative storage strategy to perform DDL storage; if it is a non-DDL operation, it is directly stored.

[0081] When this preferred embodiment is run through the above scheme, the source-side data synchronization system runs in the corresponding mode according to the node characteristics of the distributed cluster, and the data node obtains the DDL information from the metadata node to supplement the complete DDL information. Each data node runs independently, synchronizes in parallel, and does not interfere with each other. The target-side synchronization system is responsible for DDL and DML storage.

[0082] In this preferred embodiment, reference Figure 4 The DDL collaborative storage strategy of the target-side synchronization system described in step 400 specifically includes the following steps.

[0083] Step 410: The target-side data synchronization system initializes a DDL collaborative storage auxiliary table DDL_SYNC (also known as the DDL synchronization auxiliary table mentioned above) on the target-side database. The STATE field of the DDL_SYNC table indicates the status information, and the specific values ​​include: "over" indicates that the corresponding DDL has been stored; "ready" indicates that the corresponding DDL synchronization is ready; and then the total number of data nodes needs to be set to M, and the variable i=0.

[0084] Step 420: Query the DDL_SYNC table to obtain the number of data rows n1 whose STATE value is "over".

[0085] Step 430: If n1 is not greater than 0, the DDL operation type is determined and different procedures are used to set an X lock (ie, an exclusive lock) on the DDL_SYNC table according to different operation types.

[0086] Step 440: Get the number of data rows n3 whose STATE value is "over" in the DDL_SYNC table. It should be noted that n3 obtained in this step is different from n1 obtained previously, because new data will be inserted into this table during the concurrent process, so it needs to be queried again.

[0087] Step 450: If n3 is not greater than 0, the DDL storage operation is executed and the X lock of the DDL_SYNC table is released.

[0088] In this preferred embodiment, reference Figure 5 In step 430, if n1 is not greater than 0, the DDL operation type is determined and different processes are used to lock the X on the DDL_SYNC table according to different operation types. Specifically, the steps include the following.

[0089] Step 431: If n1 is not greater than 0, determine whether the DDL operation is an ALTER or TRUNCATE operation.

[0090] Step 432: If it is not an ALTER or TRUNCATE operation, directly set an X lock on the DDL_SYNC table.

[0091] Step 433: If it is an ALTER or TRUNCATE operation, query the DDL_SYNC table to obtain the number of data rows n2 whose STATE value is "ready", and determine whether n2+1 is equal to M.

[0092] Step 434: If n2+1 is equal to M, an X lock is set on the DDL_SYNC table. If n2+1 is not equal to M, a determination is made as to whether i is equal to 0.

[0093] Step 435: If i is equal to 0, insert the current DDL information with STATE value "ready" into the DDL_SYNC table, set i=i+1, wait for 1 second and then re-enter the n1 acquisition step. Re-entering the n1 acquisition step here is to obtain the latest data; if i is not equal to 0, wait directly for 1 second and then re-enter the n1 acquisition step. This is a waiting process, because there will be multiple database connections to operate this DDL_SYNC table, so the value of n1 obtained each time may be different, and n1 is used to determine whether to wait for other connections to operate data.

[0094] Based on the above steps, when the number of data rows n1 with a STATE value of "over" is queried from the DDL_SYNC table, if n1 is greater than 0, the DDL operation is skipped, that is, the DDL storage operation is not executed, and the DDL synchronization is ended directly.

[0095] In this preferred embodiment, reference Figure 6 In step 450, if n3 is not greater than 0, executing the DDL storage operation and releasing the X lock of the DDL_SYNC table specifically includes the following steps.

[0096] Step 451: If n3 is not greater than 0, insert the current DDL information with the STATE value of "over" into the DDL_SYNC table.

[0097] Step 452: Execute the storage operation of the current DDL and release the X lock of the DDL_SYNC table.

[0098] Step 453: End this DDL synchronization.

[0099] Based on the above steps, when the number of data rows n3 with the STATE value of "over" in the DDL_SYNC table is obtained, if n3 is greater than 0, the X lock of the DDL_SYNC table is directly released, and then the DDL operation is skipped, that is, the DDL storage operation is not executed, and the DDL synchronization is directly ended.

[0100] In summary, this preferred embodiment provides a data synchronization method for a distributed database. Based on the distributed DML and DDL operation logs of the distributed database being distributed on different nodes, this preferred embodiment adopts a multi-node concurrent synchronization architecture. The DDL operations involved in each node are obtained from the metadata management node, so that the timing and consistency of the table data operations of each node can be guaranteed. When the DDL on the target side is stored in the warehouse, a table lock mechanism is used to ensure that each DDL will not be stored repeatedly. In the synchronization architecture of this preferred embodiment, the synchronization of each node is independent, the synchronized transactions are dispersed, and the synchronized data is complete and consistent as a whole, so that concurrent processing can be performed to improve the overall synchronization performance.

[0101] Embodiment 2:

[0102] Based on the data synchronization method for a distributed database provided in Example 1, this Example 2 describes in detail the processing flow of the DDL log request module in the log parsing thread of the source data synchronization system. Figure 7 As shown, the following steps are included.

[0103] Step 1: Parse the log. After parsing the log in the log parsing thread, the DDL log request module obtains the log type, timestamp, transaction information, etc., and proceeds to step 2.

[0104] Step 2: Is it a DDL request? Determine whether it is a DDL request log based on the log type. If yes, go to step 3; otherwise, go to step 11.

[0105] Step 3: Create a connection META_CONN with the metadata synchronization system, and proceed to step 4. In the figure, the (source-end) metadata synchronization system is also the source-end data synchronization system in Example 1.

[0106] Step 4: Construct a request message MSG_DDL, including the message type, data node number, timestamp, transaction information, etc., and proceed to step 5.

[0107] Step 5: Get the DDL log, set the total log to M, the current log sequence number CUR_SEQ=i, and go to step 6.

[0108] Step 6: Determine whether i is equal to M. If so, proceed to step 11; otherwise, proceed to step 7.

[0109] Step 7: Get the next log, set CUR_SEQ=i+1, and go to step 8.

[0110] Step 8: Check if there is a communication failure. If yes, go to step 9; otherwise, go to step 10.

[0111] Step 9: Set CUR_SEQ=i, reacquire, and go to step 5.

[0112] Step 10: Cache the current log, set i=i+1, and go to step 6.

[0113] Step 11: This request ends.

[0114] The DDL log request module of this embodiment is included in each data node synchronization system, that is, a DDL operation will be repeated in multiple data node synchronization systems. Then, how to coordinate these DDL synchronizations will be specifically involved in the next embodiment.

[0115] Embodiment 3:

[0116] Based on the distributed database data synchronization method provided in Example 1, this Example 3 provides a distributed database data synchronization architecture diagram based on log analysis to illustrate the present invention in more detail.

[0117] like Figure 8 As shown, it is a distributed database data synchronization architecture diagram based on log parsing provided by this embodiment. Among them, the source database is a distributed database cluster, including a root server (Root Server in the figure) and multiple data nodes (EP0, EP1...EPn in the figure). Among the data nodes, EP0 is a metadata node (or metadata management node), which has DDL operations, while other EP1...Epn are ordinary data nodes with DML operations. Each node of the source corresponds to a source synchronization system (that is, a source data synchronization system), wherein the source synchronization system HS0 corresponding to the metadata node EP0 includes three threads: log reading, log parsing, and log caching. The log reading thread in HS0 is responsible for reading the database log and adding the read log to the queue to be parsed. The parsing thread in HS0 obtains the log from the queue to be parsed and parses it into transaction information to be processed, and adds it to the log to be cached queue. The log cache thread in HS0 obtains the log information from the queue to be cached and caches it according to the transaction ID classification. The source synchronization system HS1...HSn corresponding to other data nodes EP1...Epn includes three threads: log reading, log parsing (DDL request), and log sending. The log reading thread in HS1...HSn is responsible for reading logs from the corresponding data nodes and adding the read logs to the log parsing queue; the log parsing thread in HS1...HSn is responsible for parsing the logs in the parsing queue. If a DDL request log is encountered, the DDL-related logs are obtained from the source synchronization system HS0 corresponding to the metadata management node EP0 and added to the log parsing queue. The log parsing thread in HS1...HSn is also responsible for packaging the parsed logs into the internal message format of the synchronization system and adding them to the message queue to be sent; the log sending thread in HS1...HSn is responsible for sending the messages in the queue to be sent to the target data synchronization system. For the target database, the target synchronization system (i.e., the target data synchronization system) EXEC1...EXECn is set corresponding to the data nodes EP1...Epn and the source synchronization system HS1...HSn, which is used to unpack the synchronization messages sent from the source and apply these unpacking operations to the target database. If it is a DDL operation, the target-side synchronization system EXEC1...EXECn uses the DDL collaborative storage strategy to execute DDL storage; if it is a non-DDL operation, it is directly stored.

[0118] like Fig. 9The figure shows the target end data synchronization system DDL collaborative storage flow chart in this embodiment. The specific process is as follows.

[0119] 101: The target-side data synchronization system initializes a DDL collaborative storage auxiliary table DDL_SYNC on the target database. The STATE field of the table indicates the status information. The specific value is: "over" indicates that the DDL has been stored; "ready" indicates that the DDL synchronization is ready. Set the total number of data nodes to M, set the variable i = 0, and go to step 102.

[0120] 102: Query the DDL_SYNC table to obtain the number of data rows n1 whose STATE value is "over", and go to step 103.

[0121] 103: If n1 is greater than 0, go to step 104; otherwise, go to step 105.

[0122] 104 : Skip the DDL operation, that is, do not execute the DDL storage operation, and go to step 117 .

[0123] 105: Whether the DDL is an ALTER or TRUNCATE operation, if so, proceed to step 106; otherwise, proceed to step 107.

[0124] 106 : Query the DDL_SYNC table to obtain the number of data rows n2 whose STATE value is “ready”, and then go to step 108 .

[0125] 107: Apply an X lock (exclusive lock) to the DDL_SYNC table and proceed to step 112.

[0126] 108: Determine whether n2+1 is equal to M. If so, proceed to step 107; otherwise, proceed to step 109.

[0127] 109: Determine whether i is equal to 0. If yes, proceed to step 110; otherwise, proceed to step 111.

[0128] 110 : Insert the current DDL information with STATE value “ready” into the DDL_SYNC table, set i=i+1; go to step 111 .

[0129] 111: Wait for 1 second and go to step 102.

[0130] 112: Query the DDL_SYNC table to obtain the number of data rows n3 whose STATE value is “over”, and go to step 113.

[0131] 113: Determine whether n3 is greater than 0. If so, proceed to step 114; otherwise, proceed to step 115.

[0132] 114: Release the X lock of the DDL_SYNC table and go to step 104.

[0133] 115 : Insert the current DDL information with the STATE value of “over” into the DDL_SYNC table, and proceed to step 116 .

[0134] 116: Execute the storage operation of the current DDL, release the DDL_SYNC table X lock, and go to step 117.

[0135] 117: End the DDL synchronization.

[0136] In summary, this embodiment is based on the DML and DDL operation logs of the distributed database distributed on different nodes, and adopts a multi-node concurrent synchronization architecture. The DDL operations involved in each node are obtained from the metadata management node, so that the timing and consistency of the table data operations of each node can be guaranteed. When the DDL is stored on the target side, a lock table mechanism is used to ensure that each DDL will not be stored repeatedly. In the synchronization architecture of this embodiment, the synchronization of each node is independent, the synchronized transactions are dispersed, and the synchronized data is complete and consistent as a whole, so that concurrent processing can improve the overall performance of synchronization.

[0137] Embodiment 4:

[0138] Based on the distributed database data synchronization method provided in the above embodiment 1, the present invention also provides a distributed database data synchronization device that can be used to implement the above method, such as Fig.10 , which is a schematic diagram of the device architecture of an embodiment of the present invention. The data synchronization device of the distributed database of this embodiment includes one or more processors 21 and a memory 22. Fig.10 A processor 21 is taken as an example.

[0139] The processor 21 and the memory 22 may be connected via a bus or other means. Fig.10 The example of connecting through bus is taken in the following.

[0140] The memory 22, as a non-volatile computer-readable storage medium, can be used to store non-volatile software programs, non-volatile computer executable programs and modules, such as the distributed database data synchronization method in Example 1. The processor 21 executes various functional applications and data processing of the distributed database data synchronization device by running the non-volatile software programs, instructions and modules stored in the memory 22, that is, implements the distributed database data synchronization method in Example 1.

[0141] The memory 22 may include a high-speed random access memory, and may also include a non-volatile memory, such as at least one disk storage device, a flash memory device, or other non-volatile solid-state storage devices. In some embodiments, the memory 22 may optionally include a memory remotely arranged relative to the processor 21, and these remote memories may be connected to the processor 21 via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.

[0142] The program instructions / modules are stored in the memory 22. When executed by one or more processors 21, the distributed database data synchronization method in the above embodiment 1 is executed. For example, the above described Figure 1-Figure 6 The steps shown.

[0143] A person skilled in the art may understand that all or part of the steps in the various methods of the embodiments may be completed by instructing related hardware through a program, and the program may be stored in a computer-readable storage medium, and the storage medium may include: a read-only memory (ROM), a random access memory (RAM), a disk or an optical disk, etc.

[0144] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions and improvements made within the spirit and principles of the present invention should be included in the protection scope of the present invention.

Claims

1. A data synchronization method for a distributed database, It is characterized in that include: A source-side data synchronization system is deployed on the source-side database, and a target-side data synchronization system is deployed on the target-side. The source-side data synchronization system includes a metadata management node operation mode and a data node operation mode. In the metadata management node operation mode, the source-end data synchronization system initializes a log reading thread, a log parsing thread, and a log caching thread to read, parse, and cache logs from the source-end database; In the data node operation mode, the source-end data synchronization system initializes a log reading thread, a log parsing thread, and a log sending thread, which are used to read and parse logs from the source-end database and package and send the logs to the target-end data synchronization system; The target-side data synchronization system unpacks the message package sent by the source-side data synchronization system after receiving it, and applies these unpacking operations to the target-side database. For the unpacking operation, if it is a DDL operation, the target-side synchronization system uses the DDL collaborative storage strategy to perform DDL storage; if it is a non-DDL operation, it is directly stored; The DDL collaborative storage strategy of the target-side synchronization system includes: the target-side data synchronization system initializes a DDL collaborative storage auxiliary table DDL_SYNC on the target-side database; the STATE field of the DDL_SYNC table represents the state information, and the specific values ​​include: "over" indicates that the corresponding DDL has been stored; "ready" indicates that the corresponding DDL synchronization is ready; the total number of data nodes is set to M, and the variable i is set to 0; Query the DDL_SYNC table to obtain the number of data rows with STATE value "over" n1; If n1 is not greater than 0, the DDL operation type is determined and different procedures are used to set an X lock on the DDL_SYNC table according to the operation type. Get the number of data rows n3 whose STATE value is "over" in the DDL_SYNC table; If n3 is not greater than 0, the DDL storage operation is executed and the X lock of the DDL_SYNC table is released.

2. The data synchronization method of a distributed database according to claim 1, It is characterized in that If n1 is not greater than 0, the DDL operation type is determined and different procedures are used to lock the X lock on the DDL_SYNC table according to the different operation types. Specifically, the following steps are performed: If n1 is not greater than 0, determine whether the DDL operation is an ALTER or TRUNCATE operation; If it is not an ALTER or TRUNCATE operation, an X lock is directly set on the DDL_SYNC table; If it is an ALTER or TRUNCATE operation, obtain the number of data rows n2 with a STATE value of "ready" from the DDL_SYNC table, and determine whether n2+1 is equal to M; if n2+1 is equal to M, set an X lock on the DDL_SYNC table; if n2+1 is not equal to M, determine whether i is equal to 0; if i is equal to 0, insert the current DDL information with a STATE value of "ready" into the DDL_SYNC table, set i=i+1, wait for 1 second and then re-enter the n1 acquisition step; if i is not equal to 0, directly wait for 1 second and then re-enter the n1 acquisition step.

3. The data synchronization method of a distributed database according to claim 2, It is characterized in that When the number of data rows n1 with a STATE value of "over" is obtained by querying from the DDL_SYNC table, if n1 is greater than 0, the DDL operation is skipped, that is, the DDL storage operation is not executed, and the DDL synchronization is directly ended.

4. The data synchronization method of a distributed database according to claim 2, It is characterized in that If n3 is not greater than 0, executing the DDL storage operation and releasing the X lock of the DDL_SYNC table specifically includes: If n3 is not greater than 0, insert the current DDL information with STATE value "over" into the DDL_SYNC table; Execute the storage operation of the current DDL and release the X lock of the DDL_SYNC table; End this DDL synchronization.

5. The data synchronization method of a distributed database according to claim 4, It is characterized in that When obtaining the number of data rows n3 whose STATE value is "over" in the DDL_SYNC table, if n3 is greater than 0, the X lock of the DDL_SYNC table is directly released, and then the DDL operation is skipped, that is, the DDL storage operation is not executed, and the DDL synchronization is directly ended.

6. A distributed database data synchronization method according to any one of claims 1 to 5, It is characterized in that The source-end database includes a distributed database cluster; the target end includes one or more of other data sources, a single-node database, and a multi-node cluster system.

7. A distributed database data synchronization method according to any one of claims 1 to 5, It is characterized in that In the metadata management node operation mode, the source-side data synchronization system initializes the log reading thread, the log parsing thread, and the log caching thread, which are used to read, parse, and cache logs from the source-side database. Specifically, the following are included: The source data synchronization system corresponding to the metadata management node initializes a log reading thread, a log parsing thread, and a log caching thread; The log reading thread corresponding to the metadata management node is used to read the database log and add the read log to the queue to be parsed; The log parsing thread corresponding to the metadata management node is used to obtain logs from the queue to be parsed and parse them into transaction information to be processed, and add them to the log cache queue; The log cache thread corresponding to the metadata management node is used to obtain log information from the queue to be cached and cache it according to the transaction ID classification.

8. A distributed database data synchronization method according to any one of claims 1 to 5, It is characterized in that In the data node operation mode, the source-end data synchronization system initializes a log reading thread, a log parsing thread, and a log sending thread, which are used to read and parse logs from the source-end database and package and send the logs to the target-end data synchronization system, specifically including: The source data synchronization system corresponding to the data node initializes a log reading thread, a log parsing thread, and a log sending thread. The log parsing thread corresponding to the data node includes a DDL log request module, which is used to obtain the relevant DDL operation log from the metadata management node according to the transaction ID; The log reading thread corresponding to the data node is used to read logs from the corresponding data node and add the read logs to the log parsing queue; The log parsing thread corresponding to the data node is used to parse the logs in the queue to be parsed. If a DDL request log is encountered, the DDL-related logs are obtained from the source data synchronization system corresponding to the metadata management node through the DDL log request module, and added to the log queue to be parsed. The log parsing thread corresponding to the data node is also used to package the parsed logs into the internal message format of the synchronization system and add them to the message queue to be sent. The log sending thread corresponding to the data node is used to send the messages in the queue to be sent to the target data synchronization system.

9. A data synchronization device for a distributed database, Features: It includes at least one processor and a memory, the at least one processor and the memory are connected via a data bus, the memory stores instructions that can be executed by the at least one processor, and the instructions, after being executed by the processor, are used to complete the data synchronization method of a distributed database as described in any one of claims 1-8.

Citation Information

Patent Citations

  • Metadata synchronization method and device, electronic equipment and storage medium

    CN114756562A

  • Optimistic concurrency control for database transactions

    US20190179930A1