Method and system for realizing consistent copying of transactions
By using Raft consensus algorithm and CDC technology in distributed databases, transaction consistency replication is achieved, and the problem of transaction atomicity and consistency in the existing technology is solved, and the consistency and integrity of the main and standby database transactions are achieved.
Patent Information
- Application Number
- CN202510139255.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-08
- Publication Date
- 2025-05-16
AI Technical Summary
In the data replication process of primary and secondary clusters, the prior art does not support transaction characteristics and cannot guarantee the atomicity and consistency of transactions.
Using a method based on Raft consensus algorithm, transaction consistency replication is realized in a distributed database. Data changes are captured through CDC and when transaction playback is performed by backup clusters, transaction atomicity is supported to ensure transaction integrity and master-standby transaction consistency.
With the same scale and configuration of the main and standby library cluster, the transaction atomicity of a single transaction and the consistency of the main and standby transactions are ensured, and the integrity of distributed transactions is ensured.
Smart Images

Figure CN120011452A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of distributed databases, and in particular to a method and system for implementing transaction consistency replication. Background Art
[0002] CDC stands for change data capture. It is a way to back up a database and is often used to back up large amounts of data. The CDC mechanism monitors the target data in the primary cluster and captures incremental change data when the data changes. The problem with CDC is that it does not support transactional features and only captures the final result of the data change. There is no transaction association, and data replication can achieve final consistency.
[0003] Currently, the master and standby clusters implement data replication of master and standby distributed clusters based on CDC, capture data changes in the master cluster, launch them to the standby cluster, and replay the data in the standby cluster to achieve the final consistency of the master and standby clusters. Regarding the playback of standby cluster data, transaction features are not supported, and the atomicity and consistency of transactions cannot be guaranteed. Summary of the invention
[0004] The technical task of the present invention is to address the above shortcomings and provide a method and system for implementing transaction consistency replication, optimize distributed database clusters, use data replication based on CDC, and under the master-slave cluster replication scheme, playback of the backup cluster data ensures transaction atomicity and consistency.
[0005] The technical solution adopted by the present invention to solve its technical problem is:
[0006] A method for implementing transaction consistency replication is implemented on a distributed database. Based on the Raft consensus algorithm, the method supports CDC to capture data changes for data replication. When replicating and replaying CDC data, it supports transaction atomicity, ensures data transaction integrity, and ensures master-slave transaction consistency. The method is implemented as follows:
[0007] Distributed databases submit multiple transactions globally (across machines);
[0008] Replay multiple transactions submitted globally;
[0009] When replaying a transaction, you can determine the number of operations contained in the transaction and the transaction order;
[0010] When replaying a transaction, if all operations of the transaction are received, the transaction can be replayed;
[0011] Implicit transactions can be replayed directly.
[0012] This method is based on the Raft consensus protocol in a distributed database, captures data changes and replicates data to ensure transaction integrity and consistency. It is used to ensure that the backup database data has transaction semantics and ensure the consistency of distributed transactions.
[0013] Furthermore, the specific process of implementing this method is as follows:
[0014] 1) CDC captures data changes based on the Raft consensus protocol and captures the consensus KV data; for each data change, the corresponding transaction ID and operation are recorded; when the transaction is committed, the transaction ID and the number of operations involved in the transaction are recorded;
[0015] 2) When the main cluster performs CDC data capture, it records the number of data changes involved in a transaction and the transaction ID of the transaction; when data is transmitted, in addition to the changed data, the transaction ID and the number of changes involved in the transaction are also transmitted;
[0016] 3) After receiving the data, the standby cluster first performs data aggregation based on the transaction ID. If the change data record aggregated by the transaction is the same as the change number of the transaction, it is considered that all operations of the transaction have been sent to the standby cluster, and the data of the transaction can be replayed, thereby ensuring the atomicity of the transaction;
[0017] 4) The Raft consensus protocol ensures the order of data capture changes and the order of transactions on each range;
[0018] 5) During playback, data change operations are read from the CDC queue, and aggregation operations are performed based on the transaction ID. The transaction dependencies are generated based on the order of transactions on each Range, and the transactions are played back in the order of their occurrence, thereby ensuring transaction consistency.
[0019] 6) After successful playback, clear the operation queue of the transaction in the cache.
[0020] Furthermore, the replication playback process of a single transaction is as follows:
[0021] S1: The main cluster performs transaction operations, including data writing and transaction submission;
[0022] S2, CDC captures each operation of a transaction and carries the total number of operations at the end of the transaction;
[0023] S3. After receiving all operations and operands, the standby cluster performs data comparison and replays them after receiving all operations and operands.
[0024] Furthermore, the process of achieving consistency guarantee in this method is as follows:
[0025] Set the distributed database to submit multiple transactions, each transaction contains a transaction ID and a corresponding KV value;
[0026] S1. As time goes by, multiple transactions perform multiple data operations on multiple ranges.
[0027] S2. CDC captures a set of data for each range. This set of data contains the operations of multiple transactions and also describes the execution order of the transactions in the range, forming a conflict queue.
[0028] S3. Generate a waiting queue for transactions according to the transaction conflict queue;
[0029] S4. After the earliest transaction is replayed, the queue is updated and the earliest transaction is executed in a loop to ensure the transaction execution order and transaction consistency.
[0030] The present invention also claims to protect a system for implementing transaction consistency replication, which is applied to a distributed database, based on the Raft consensus algorithm, supports CDC to capture data changes for data replication, supports transaction atomicity when replicating and replaying CDC data, ensures transaction integrity of data, and ensures consistency of master and standby transactions; the implementation method is as follows:
[0031] Distributed databases submit multiple transactions globally (across machines);
[0032] Replay multiple transactions submitted globally;
[0033] When replaying a transaction, you can determine the number of operations contained in the transaction and the transaction order;
[0034] When replaying a transaction, if all operations of the transaction are received, the transaction can be replayed;
[0035] Implicit transactions can be replayed directly.
[0036] Furthermore, the specific process of implementing transaction consistency replication in this system is as follows:
[0037] 1) CDC captures data changes based on the Raft consensus protocol and captures the consensus KV data; for each data change, the corresponding transaction ID and operation are recorded; when the transaction is committed, the transaction ID and the number of operations involved in the transaction are recorded;
[0038] 2) When the main cluster performs CDC data capture, it records the number of data changes involved in a transaction and the transaction ID of the transaction; when data is transmitted, in addition to the changed data, the transaction ID and the number of changes involved in the transaction are also transmitted;
[0039] 3) After receiving the data, the standby cluster first performs data aggregation based on the transaction ID. If the change data record aggregated by the transaction is the same as the change number of the transaction, it is considered that all operations of the transaction have been sent to the standby cluster, and the data of the transaction can be replayed, thereby ensuring the atomicity of the transaction;
[0040] 4) The Raft consensus protocol ensures the order of data capture changes and the order of transactions on each range;
[0041] 5) During playback, data change operations are read from the CDC queue, and aggregation operations are performed based on the transaction ID. The transaction dependencies are generated based on the order of transactions on each Range, and the transactions are played back in the order of their occurrence, thereby ensuring transaction consistency.
[0042] 6) After successful playback, clear the operation queue of the transaction in the cache.
[0043] Furthermore, the replication playback process of a single transaction is as follows:
[0044] S1: The main cluster performs transaction operations, including data writing and transaction submission;
[0045] S2, CDC captures each operation of a transaction and carries the total number of operations at the end of the transaction;
[0046] S3. After receiving all operations and operands, the standby cluster performs data comparison and replays them after receiving all operations and operands.
[0047] Furthermore, the process of achieving consistency assurance in this system is as follows:
[0048] Set the distributed database to submit multiple transactions, each transaction contains a transaction ID and a corresponding KV value;
[0049] S1. As time goes by, multiple transactions perform multiple data operations on multiple ranges.
[0050] S2. CDC captures a set of data for each range. This set of data contains the operations of multiple transactions and also describes the execution order of the transactions in the range, forming a conflict queue.
[0051] S3. Generate a waiting queue for transactions based on the transaction conflict queue;
[0052] S4. After the earliest transaction is replayed, the queue is updated and the earliest transaction is executed in a loop to ensure the transaction execution order and transaction consistency.
[0053] The present invention also claims a device for implementing transaction consistency replication, comprising: at least one memory and at least one processor;
[0054] The at least one memory is used to store a machine-readable program;
[0055] The at least one processor is used to call the machine-readable program to implement the above method.
[0056] The present invention also claims protection for a computer-readable medium having computer instructions stored thereon, which can implement the above method when executed by a processor.
[0057] Compared with the prior art, the method and system for implementing transaction consistency replication of the present invention have the following beneficial effects:
[0058] When the primary and standby database clusters are of the same size and configuration, the data integrity of the standby database can be guaranteed. When a standby database fails, the atomicity of a single transaction and the consistency of the primary and standby transactions can be guaranteed.
[0059] In the master-slave cluster replication solution based on CDC, if some failures occur in the slave cluster, the data changes sent by the master cluster may not be fully applied, which may split multiple operations of a transaction. In this case, the number of transaction operation changes is counted and verified during playback to ensure the atomicity of the transaction and thus the integrity of the data.
[0060] Through transaction dependencies, the playback order of standby database transactions is guaranteed, thereby ensuring the consistency of primary and standby transactions. BRIEF DESCRIPTION OF THE DRAWINGS
[0061] Figure 1 is a flowchart of a single transaction replication playback process provided by an embodiment of the present invention;
[0062] Figure 2 is a flowchart of consistency assurance provided by an embodiment of the present invention;
[0063] Figure 3 is a diagram of a conflict queue provided by an embodiment of the present invention;
[0064] Figure 4 The diagram is a diagram of ensuring the transaction execution order provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0065] The present invention will be further described below in conjunction with specific embodiments.
[0066] The embodiment of the present invention provides a method for implementing transaction consistency replication, which implements transaction consistency replication based on the Raft protocol. The method is implemented on a distributed database, based on the Raft consensus algorithm, supports CDC to capture data changes for data replication, supports transaction atomicity when replicating and replaying CDC data, ensures transaction integrity of data, and ensures consistency of master and standby transactions.
[0067] The implementation of this method is as follows:
[0068] Distributed databases submit multiple transactions globally (across machines).
[0069] Replay multiple transactions submitted globally.
[0070] When replaying a transaction, you can determine the number of operations contained in the transaction and the transaction order.
[0071] When replaying a transaction, if all operations of the transaction are received, the transaction can be replayed.
[0072] Implicit transactions can be replayed directly.
[0073] The specific process of implementing this method is as follows:
[0074] 1. CDC captures data changes based on the Raft consensus protocol and captures the consensus KV data. For each data change, the corresponding transaction ID and operation are recorded; when the transaction is committed, the transaction ID and the number of operations involved in the transaction are recorded.
[0075] 2. When the main cluster performs CDC data capture, it records the number of data changes involved in a transaction and the transaction ID of the transaction; when data is transmitted, in addition to the changed data, the transaction ID and the number of changes involved in the transaction are also transmitted;
[0076] 3. After receiving the data, the standby cluster first performs data aggregation based on the transaction ID. If the change data record aggregated by the transaction is the same as the number of changes in the transaction, it is considered that all operations of the transaction have been sent to the standby cluster, and the data of the transaction can be replayed to ensure the atomicity of the transaction.
[0077] 4. The Raft consensus protocol ensures the order of data capture changes and the order of transactions on each Range;
[0078] 5. During playback, data change operations are read from the CDC queue, and aggregation operations are performed based on the transaction ID. The transaction dependencies are generated based on the order of transactions on each Range, and the transactions are played back in the order of their occurrence, thereby ensuring transaction consistency.
[0079] 6. After the playback is successful, clear the operation queue of the transaction in the cache.
[0080] like Figure 1 As shown, the replication playback process of a single transaction is as follows:
[0081] S1: The main cluster performs transaction operations, including data writing and transaction submission;
[0082] S2, CDC captures each operation of a transaction and carries the total number of operations at the end of the transaction;
[0083] S3. After receiving all operations and operands, the standby cluster performs data comparison and replays them after receiving all operations and operands.
[0084] Combined with Figure 2-4 As shown in the figure, the process of achieving consistency guarantee by this method is as follows:
[0085] Set the distributed database to commit multiple transactions, such as Figure 2 As shown, each color is a transaction, and each transaction includes a transaction ID and a corresponding KV value.
[0086] 1. As time goes by, multiple transactions perform multiple data operations on multiple ranges.
[0087] 2. CDC captures a set of data for each range. This set of data contains the operations of multiple transactions and also explains the execution order of the transactions in the range, forming a conflict queue. The conflict queue diagram is as follows: Figure 3 shown.
[0088] 3. Generate a transaction waiting queue based on the transaction conflict queue.
[0089] 4. After the earliest transaction is replayed, the queue is updated and the earliest transaction is executed in a loop to ensure the order of transaction execution and transaction consistency. Figure 4 shown.
[0090] The Raft protocol is a distributed consistency algorithm (consensus algorithm). Consensus is an algorithm for multiple nodes to reach a consensus on an event. Even if some nodes fail or network delays occur, it will not affect the nodes, thereby improving the overall availability of the system.
[0091] Transactions have ACID properties, and the database always describes the real world. Specifically, it meets some business consistency requirements. For example, when transferring money within the banking system, the total amount remains unchanged, and there are no accounts with negative balances, etc.
[0092] This method is based on the Raft consensus protocol in a distributed database, captures data changes and replicates data to ensure transaction integrity and consistency. It is used to ensure that the backup database data has transaction semantics and ensure the consistency of distributed transactions.
[0093] The embodiment of the present invention further provides a system for implementing transaction consistency replication, which is applied to a distributed database, based on the Raft consensus algorithm, supports CDC to capture data changes for data replication, supports transaction atomicity when replicating and replaying CDC data, ensures transaction integrity of data, and ensures consistency of master and standby transactions; the implementation method is as follows:
[0094] Distributed databases submit multiple transactions globally (across machines);
[0095] Replay multiple transactions submitted globally;
[0096] When replaying a transaction, you can determine the number of operations contained in the transaction and the transaction order;
[0097] When replaying a transaction, if all operations of the transaction are received, the transaction can be replayed;
[0098] Implicit transactions can be replayed directly.
[0099] The specific process of implementing transaction consistency replication in this system is as follows:
[0100] 1. CDC captures data changes based on the Raft consensus protocol and captures the consensus KV data. For each data change, it records the corresponding transaction ID and operation. When a transaction is committed, it records the transaction ID and the number of operations involved in the transaction.
[0101] 2. When the main cluster performs CDC data capture, it records the number of data changes involved in a transaction and the transaction ID of the transaction; when transmitting data, in addition to transmitting the changed data, it also transmits the transaction ID and the number of changes involved in the transaction.
[0102] 3. After receiving the data, the standby cluster first performs data aggregation based on the transaction ID. If the change data record aggregated by the transaction is the same as the number of changes in the transaction, it is considered that all operations of the transaction have been sent to the standby cluster, and the data of the transaction can be replayed to ensure the atomicity of the transaction.
[0103] 4. The Raft consensus protocol ensures the order of data capture changes and the order of transactions on each Range.
[0104] 5. During playback, data change operations are read from the CDC queue, and aggregation operations are performed based on the transaction ID. The transaction dependencies are generated based on the order of transactions on each Range, and the transactions are played back in the order of their occurrence, thereby ensuring transaction consistency.
[0105] 6. After the playback is successful, clear the operation queue of the transaction in the cache.
[0106] The replication playback process of a single transaction is as follows:
[0107] S1: The main cluster performs transaction operations, including data writing and transaction submission;
[0108] S2, CDC captures each operation of a transaction and carries the total number of operations at the end of the transaction;
[0109] S3. After receiving all operations and operands, the standby cluster performs data comparison and replays them after receiving all operations and operands.
[0110] The process of achieving consistency guarantee in this system is as follows:
[0111] Suppose that the distributed database commits multiple transactions, each transaction contains a transaction ID and a corresponding KV value, such as Figure 2 As shown, each color is a transaction.
[0112] 1. As time goes by, multiple transactions perform multiple data operations on multiple ranges.
[0113] 2. CDC captures a set of data for each range. This set of data contains the operations of multiple transactions and also explains the execution order of the range transactions, forming a conflict queue, such as Figure 3 shown.
[0114] 3. Generate a waiting queue for transactions based on the transaction conflict queue.
[0115] 4. After the earliest transaction is replayed, the queue is updated and the earliest transaction is executed in a loop to ensure the order of transaction execution and transaction consistency. Figure 4 shown.
[0116] An embodiment of the present invention also provides a device for implementing transaction consistency replication, comprising: at least one memory and at least one processor;
[0117] The at least one memory is used to store a machine-readable program;
[0118] The at least one processor is used to call the machine-readable program to implement the method for implementing transaction consistency replication described in the above embodiment.
[0119] The embodiment of the present invention further provides a computer-readable medium, on which computer instructions are stored, and when the computer instructions are executed by a processor, the method for implementing transaction consistency replication described in the above embodiment is implemented. Specifically, a system or device equipped with a storage medium can be provided, on which a software program code implementing the functions of any of the above embodiments is stored, and a computer (or CPU or MPU) of the system or device reads and executes the program code stored in the storage medium.
[0120] In this case, the program code itself read from the storage medium can realize the function of any one of the above-mentioned embodiments, and thus the program code and the storage medium storing the program code constitute a part of the present invention.
[0121] The storage medium embodiments for providing the program code include a floppy disk, a hard disk, a magneto-optical disk, an optical disk (such as CD-ROM, CD-R, CD-RW, DVD-ROM, DVD-RAM, DVD-RW, DVD+RW), a magnetic tape, a non-volatile memory card, and a ROM. Alternatively, the program code can be downloaded from a server computer by a communication network.
[0122] In addition, it should be clear that the functions of any of the above embodiments can be implemented not only by executing the program code read by the computer, but also by enabling an operating system operating on the computer to complete part or all of the actual operations based on instructions from the program code.
[0123] In addition, it can be understood that the program code read from the storage medium is written to a memory provided in an expansion board inserted into the computer or written to a memory provided in an expansion unit connected to the computer, and then based on the instructions of the program code, a CPU installed on the expansion board or the expansion unit is enabled to perform part or all of the actual operations, thereby realizing the functions of any of the above-mentioned embodiments.
[0124] The present invention is shown and described in detail above through the accompanying drawings and preferred embodiments. However, the present invention is not limited to these disclosed embodiments. Based on the above multiple embodiments, those skilled in the art can know that the code review methods in the above different embodiments can be combined to obtain more embodiments of the present invention, and these embodiments are also within the protection scope of the present invention.
Claims
1. A method for implementing transaction consistency replication, characterized in that: This method is implemented on a distributed database, based on the Raft consensus algorithm, and supports CDC to capture data changes for data replication. When replicating and replaying CDC data, it supports transaction atomicity, ensures data transaction integrity, and ensures master-slave transaction consistency. The implementation of this method is as follows: The distributed database submits multiple transactions globally; Replay multiple transactions submitted globally; When replaying a transaction, you can determine the number of operations contained in the transaction and the transaction order; When replaying a transaction, if all operations of the transaction are received, the transaction can be replayed; Implicit transactions can be replayed directly.
2. A method for implementing transaction consistency replication according to claim 1, characterized in that: The specific process of implementing this method is as follows: 1) CDC captures data changes based on the Raft consensus protocol and captures the consensus KV data; for each data change, the corresponding transaction ID and operation are recorded; when the transaction is committed, the transaction ID and the number of operations involved in the transaction are recorded; 2) When the primary cluster performs CDC data capture, it records the number of data changes involved in a transaction and the transaction ID of the transaction; When transmitting data, in addition to the changed data, the transaction ID and the number of changes involved in the transaction are also transmitted; 3) After receiving the data, the standby cluster first performs data aggregation based on the transaction ID. If the change data record aggregated by the transaction is the same as the change number of the transaction, it is considered that all operations of the transaction have been sent to the standby cluster, and the data of the transaction can be replayed, thereby ensuring the atomicity of the transaction; 4) The Raft consensus protocol ensures the order of data capture changes and the order of transactions on each range; 5) During playback, data change operations are read from the CDC queue, and aggregation operations are performed based on the transaction ID. The transaction dependencies are generated based on the order of transactions on each Range, and the transactions are played back in the order of their occurrence, thereby ensuring transaction consistency. 6) After successful playback, clear the operation queue of the transaction in the cache.
3. A method for implementing transaction consistency replication according to claim 1 or 2, characterized in that: The replication playback process of a single transaction is as follows: S1: The main cluster performs transaction operations, including data writing and transaction submission; S2, CDC captures each operation of a transaction and carries the total number of operations at the end of the transaction; S3. After receiving all operations and operands, the standby cluster performs data comparison and replays them after receiving all operations and operands.
4. A method for implementing transaction consistency replication according to claim 1 or 2, characterized in that: The process of achieving consistency guarantee in this method is as follows: Set the distributed database to submit multiple transactions, each transaction contains a transaction ID and a corresponding KV value; S1. As time goes by, multiple transactions perform multiple data operations on multiple ranges. S2. CDC captures a set of data for each range. This set of data contains the operations of multiple transactions and also describes the execution order of the transactions in the range, forming a conflict queue. S3. Generate a waiting queue for transactions according to the transaction conflict queue; S4. After the earliest transaction is replayed, the queue is updated and the earliest transaction is executed in a loop to ensure the transaction execution order and transaction consistency.
5. A system for implementing transaction consistency replication, characterized in that: The system is applied to distributed databases. Based on the Raft consensus algorithm, it supports CDC to capture data changes for data replication. When replicating and replaying CDC data, it supports transaction atomicity, ensures data transaction integrity, and ensures master-slave transaction consistency. The implementation method is as follows: The distributed database submits multiple transactions globally; Replay multiple transactions submitted globally; When replaying a transaction, you can determine the number of operations contained in the transaction and the transaction order; When replaying a transaction, if all operations of the transaction are received, the transaction can be replayed; Implicit transactions can be replayed directly.
6. A system for implementing transaction consistency replication according to claim 5, characterized in that: The specific process of implementing transaction consistency replication in this system is as follows: 1) CDC captures data changes based on the Raft consensus protocol and captures the consensus KV data; for each data change, the corresponding transaction ID and operation are recorded; when the transaction is committed, the transaction ID and the number of operations involved in the transaction are recorded; 2) When the primary cluster performs CDC data capture, it records the number of data changes involved in a transaction and the transaction ID of the transaction; When transmitting data, in addition to the changed data, the transaction ID and the number of changes involved in the transaction are also transmitted; 3) After receiving the data, the standby cluster first performs data aggregation based on the transaction ID. If the change data record aggregated by the transaction is the same as the change number of the transaction, it is considered that all operations of the transaction have been sent to the standby cluster, and the data of the transaction can be replayed, thereby ensuring the atomicity of the transaction; 4) The Raft consensus protocol ensures the order of data capture changes and the order of transactions on each range; 5) During playback, data change operations are read from the CDC queue, and aggregation operations are performed based on the transaction ID. The transaction dependencies are generated based on the order of transactions on each Range, and the transactions are played back in the order of their occurrence, thereby ensuring transaction consistency. 6) After successful playback, clear the operation queue of the transaction in the cache.
7. A system for implementing transaction consistency replication according to claim 5, characterized in that: The replication playback process of a single transaction is as follows: S1: The main cluster performs transaction operations, including data writing and transaction submission; S2, CDC captures each operation of a transaction and carries the total number of operations at the end of the transaction; S3. After receiving all operations and operands, the standby cluster performs data comparison and replays them after receiving all operations and operands.
8. A system for implementing transaction consistency replication according to claim 5, characterized in that: The process of achieving consistency guarantee in this system is as follows: Set the distributed database to submit multiple transactions, each transaction contains a transaction ID and a corresponding KV value; S1. As time goes by, multiple transactions perform multiple data operations on multiple ranges. S2. CDC captures a set of data for each range. This set of data contains the operations of multiple transactions and also describes the execution order of the transactions in the range, forming a conflict queue. S3. Generate a waiting queue for transactions based on the transaction conflict queue; S4. After the earliest transaction is replayed, the queue is updated and the earliest transaction is executed in a loop to ensure the transaction execution order and transaction consistency.
9. A device for implementing transaction consistency replication, characterized in that: include: at least one memory and at least one processor; The at least one memory is used to store a machine-readable program; The at least one processor is used to call the machine-readable program to implement the method described in any one of claims 1 to 4.
10. A computer-readable medium, characterized in that The computer readable medium stores computer instructions, which, when executed by a processor, can implement the method according to any one of claims 1 to 4.