A Method and System for Distributed Database Backup Based on LVM Snapshots

By receiving methods such as creating barrier requests, generating barrier points and creating LVM snapshots in a distributed database system, the problem of inconsistent backup data in a distributed database system is solved, and distributed consistent backup and efficient recovery are achieved.

CN119782432BActive Publication Date: 2025-06-13SHUYI TECHNOLOGY (WUHAN) CO LTD +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510251755.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-05
Publication Date
2025-06-13
Estimated Expiration
2045-03-05

AI Technical Summary

Technical Problem

In distributed database systems, the existing technology cannot realize distributed consistent backup, resulting in inconsistent backup data.

Method used

By receiving the creation barrier request sent by the co-ordination point, generate the barrier point, and automatically unlock and create an LVM snapshot after all distributed nodes complete the barrier creation, ensuring that the snapshot is created at a consistent moment for all nodes.

Benefits of technology

The data consistency and integrity of distributed database backup is realized, which avoids the problem of data inconsistency and improves the efficiency of backup and recovery.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119782432B_ABST
    Figure CN119782432B_ABST
Patent Text Reader

Abstract

This application belongs to the field of database technology, and specifically discloses a method and system for distributed database backup based on LVM snapshots. The method includes: receiving a barrier creation request sent by a coordination node, and creating a barrier point based on the barrier creation request; wherein, the generation processes of each barrier point are not carried out simultaneously to achieve the uniqueness and consistency of the barrier points; suspending the commit of the target distributed transaction in the distributed system to ensure that only the distributed transaction for creating the barrier is committed in the distributed system, and the target distributed transaction is any distributed transaction other than the creation of the barrier; automatically unlocking the barrier and creating LVM snapshots when all distributed nodes have completed creating the barrier, and each distributed node independently creates LVM snapshots; dumping the LVM snapshots to a backup storage server, and deleting or retaining the LVM snapshots. Through this application, the integrity and consistency of the backup data can be achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the field of database technology, and more specifically, relates to a method and system for distributed database backup based on LVM snapshots. Background Art

[0002] Current database backup and recovery methods include: full and incremental backups based on WAL and CBM, and full backups based on LVM snapshots. The first method has slow backup speed, slow recovery speed and only supports single-machine backup, not distributed consistency backup, while the second method has fast backup speed and fast recovery speed, but only supports single-machine backup and does not support distributed consistency backup.

[0003] However, in a distributed database system, data is distributed across multiple physical nodes, and backups must ensure that the data is in the same snapshot on all physical nodes. Since LVM runs independently on each physical node, a distributed database system cannot complete consistent backups relying solely on LVM.

[0004] Therefore, how to solve the problem of inconsistent backup data in a distributed database system is a technical problem that urgently needs to be solved currently. Summary of the Invention

[0005] Aiming at the deficiencies of the prior art, the purpose of this application is to provide a method and system for distributed database backup based on LVM snapshots, aiming to solve the problem of inconsistent backup data in a distributed database system.

[0006] To achieve the above purpose, in the first aspect, this application provides a method for distributed database backup based on LVM snapshots, which is applied to distributed nodes of a database. The method includes:

[0007] Receiving a create barrier request sent by a coordination node, and creating a barrier point based on the create barrier request; wherein, the generation processes of each barrier point are not carried out simultaneously to achieve the uniqueness and consistency of the barrier points;

[0008] Pausing the submission of the target distributed transaction of the distributed system to ensure that only the distributed transaction for creating the barrier is submitted in the distributed system, and the target distributed transaction is any distributed transaction other than the create barrier;

[0009] Automatically unlocking the barrier and creating an LVM snapshot when all distributed nodes have completed creating the barrier. Each distributed node independently creates an LVM snapshot;

[0010] Dumping the LVM snapshot to a backup storage server, and deleting or retaining the LVM snapshot.

[0011] Optionally, it further includes:

[0012] Generate a WAL log with special markers based on the barrier creation request, and perform log replay based on the WAL log;

[0013] Among them, the special marker is used to indicate the position of the barrier point, and the special markers corresponding to each distributed node are consistent.

[0014] Optionally, the process of log replay includes:

[0015] Determine the creation time of the barrier point of each distributed node, and use each creation time as the cut-off point, which is used to make the restored database state consistent with the database state during backup;

[0016] During the recovery process, compare the generation time of the data blocks in the database with the cut-off point, and exclude the data blocks after the cut-off point;

[0017] Execute log replay based on the WAL log, and restore the data blocks to the state at the cut-off point.

[0018] Optionally, executing log replay based on the WAL log and restoring the data blocks to the state at the cut-off point includes:

[0019] Receive the backup medium sent by the coordination node, where the backup medium includes the data files corresponding to the LVM snapshot and the WAL log;

[0020] Based on the WAL log and on the basis of the original file blocks of the database, start log replay from the last time point of the data file;

[0021] Replay the WAL log to the cut-off point and stop log replay;

[0022] When the log replay of all distributed nodes is completed, determine that the database recovery is completed.

[0023] Optionally, deleting or retaining the LVM snapshot includes:

[0024] Delete the LVM snapshot from the original storage medium to reduce the occupied snapshot space;

[0025] Or store the LVM snapshot retention to the original storage medium for timely recovery in case of emergency, and delete the LVM snapshot until the snapshot space memory of the original storage medium reaches the storage threshold.

[0026] This application also provides a method for backing up a distributed database based on an LVM snapshot, which is applied to the coordination node of the database. The method includes:

[0027] Send a barrier creation request to the distributed nodes. The barrier creation request is used to create a barrier point and generate a WAL log with a special marker for log replay;

[0028] Poll each distributed node to determine that each distributed node has successfully created a barrier point to start creating an LVM snapshot.

[0029] Optionally, it further includes:

[0030] Obtain the backup media of each distributed node from the LVM snapshot or the dump file of the backup storage server; the backup media includes data files and WAL logs;

[0031] Distribute the backup media to each distributed node to enable the distributed node to perform log replay.

[0032] In a second aspect, the present application further provides a system for distributed database backup based on an LVM snapshot, including:

[0033] A creation module, configured to receive a barrier creation request sent by a coordination node and create a barrier point based on the barrier creation request; wherein, the generation processes of each barrier point are not carried out simultaneously to achieve the uniqueness and consistency of the barrier points;

[0034] A suspension module, configured to suspend the submission of the target distributed transaction of the distributed system to ensure that only the distributed transaction for creating a barrier in the distributed system is submitted, and the target distributed transaction is any distributed transaction other than the creation of the barrier;

[0035] A generation module, configured to automatically unlock the barrier and create an LVM snapshot when all distributed nodes have completed creating the barrier, and each distributed node independently creates an LVM snapshot;

[0036] A dump module, configured to dump the LVM snapshot into a backup storage server and delete or retain the LVM snapshot.

[0037] In a third aspect, the present application provides an electronic device, including: at least one memory for storing a program; at least one processor for executing the program stored in the memory, and when the program stored in the memory is executed, the processor is configured to execute the method described in the first aspect or any possible implementation manner of the first aspect.

[0038] In a fourth aspect, the present application provides a computer-readable storage medium storing a computer program, and when the computer program runs on a processor, it enables the processor to execute the method described in the first aspect or any possible implementation manner of the first aspect.

[0039] Fifth aspect, the present application provides a computer program product, which, when running on a processor, causes the processor to execute the method described in the first aspect or any possible implementation manner of the first aspect.

[0040] It can be understood that for the beneficial effects of the above second aspect to fifth aspect, reference can be made to the relevant descriptions in the first aspect above, which will not be elaborated here.

[0041] Generally speaking, compared with the prior art, the above technical solutions conceived by the present application have the following beneficial effects:

[0042] (1) The present application receives a create barrier request sent by a coordination node and generates a barrier point based on this request. The generation processes of each barrier point are not carried out simultaneously, and the processes of creating barrier points on different distributed nodes are coordinated and synchronized, thereby achieving the consistency of backup operations; since the generation of barrier points is not carried out simultaneously, it can ensure that a synchronized and consistent state can be achieved on all nodes, thereby achieving the integrity and consistency of backup data; and by waiting for all distributed nodes to complete the barrier creation and then creating an LVM snapshot, the creation of the LVM snapshot is carried out at a consistent moment for all nodes, thereby avoiding the problem of data inconsistency, and each distributed node independently creates an LVM snapshot, realizing that the backup states of each node are independent and consistent, thereby improving the consistency of data backup.

[0043] (2) The process of creating a snapshot in the present application ensures that it can be quickly restored to a specific time point, which can significantly reduce the recovery time; using snapshot technology for dumping can collect data without affecting the operation, thereby realizing incremental backup of data and improving the availability and security of data.

[0044] (3) When creating a barrier in the present application, the submission of the target distributed transaction is suspended, which can avoid modifying data during the snapshot creation, thereby reducing the interference with system performance; by quickly creating a barrier and an LVM snapshot, the time for data consistency processing can be shortened, reducing the long-term impact on the overall system performance. BRIEF DESCRIPTION OF THE DRAWINGS

[0045] Figure 1 is one of the flow diagrams of the method for distributed database backup based on LVM snapshots provided by an embodiment of the present application;

[0046] Figure 2 is another flow diagram of the method for distributed database backup based on LVM snapshots provided by an embodiment of the present application;

[0047] Figure 3 is a third flow diagram of the method for distributed database backup based on LVM snapshots provided by an embodiment of the present application;

[0048] Figure 4 It is the fourth schematic flowchart of the method for distributed database backup based on LVM snapshots provided by an embodiment of the present application;

[0049] Figure 5 It is a schematic diagram of the principle of the method for distributed database backup based on LVM snapshots provided by an embodiment of the present application;

[0050] Figure 6 It is a schematic structural diagram of the device for distributed database backup based on LVM snapshots provided by an embodiment of the present application;

[0051] Figure 7 It is a schematic structural diagram of the electronic device provided by an embodiment of the present application. Detailed implementation manners

[0052] In order to make the objectives, technical solutions and advantages of the present application clearer, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.

[0053] The term "and / or" in this document is an association relationship describing associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. The symbol " / " in this document represents an "or" relationship between associated objects. For example, A / B represents A or B.

[0054] The terms "first", "second", etc. in the description and claims of this document are used to distinguish different objects, rather than to describe a specific order of the objects. For example, the first response message and the second response message are used to distinguish different response messages, rather than to describe the specific order of the response messages.

[0055] In the embodiments of the present application, words such as "exemplary" or "for example" are used to represent examples, illustrations or explanations. Any embodiment or design solution described as "exemplary" or "for example" in the embodiments of the present application should not be construed as being more preferred or having more advantages than other embodiments or design solutions. Rather, the use of words such as "exemplary" or "for example" is intended to present related concepts in a specific manner.

[0056] In the description of the embodiments of the present application, unless otherwise specified, the meaning of "a plurality of" refers to two or more. For example, a plurality of processing units refers to two or more processing units, etc.; a plurality of elements refers to two or more elements, etc.

[0057] First, the related terms of the present application will be explained:

[0058] Datanode: That is, the data node of a distributed database; each data node independently manages a data shard. In a distributed database system, there are usually multiple data nodes. The combination of the data managed by these data nodes constitutes the data set managed by the database system.

[0059] Coordinator: That is, the coordination node of a distributed database. The responsibilities of the coordination node are: receiving user input (usually an SQL command), parsing, optimizing, and controlling the execution of user commands. Note that the roles of the Coordinator and Datanode can be independently assumed by two different instances, or can be assumed by one instance (i.e., one instance has both the Coordinator and Datanode roles).

[0060] LVM: That is, the abbreviation of Logical Volume Manager (logical volume management), which is a mechanism for managing disk partitions in a Linux environment.

[0061] Copy-on-write feature: (copy-on-write, abbreviated as COW), its main purpose is to reduce latency and delay memory allocation, increase execution efficiency, and only truly allocate physical resources during the actual write operation. At the same time, it can also protect data from being lost in the event of a system crash. For example, when we modify a file, the file system will first place the data to be modified in another location and then perform the operation.

[0062] Snapshot: A disk snapshot is a quick file system backup of the entire disk volume. The main difference from other backup methods lies in speed. When taking a disk snapshot, no file copying actions are involved. Generally speaking, even if the data volume is large, the backup operation can usually be completed within one second.

[0063] The embodiments of the present application will be described below with reference to the accompanying drawings in the embodiments of the present application.

[0064] Refer to Figure 1 , the present application provides a method for backing up a distributed database based on LVM snapshots, which is applied to distributed nodes of a database. The method includes:

[0065] S101. Receive a create barrier request sent by a coordination node, and create a barrier point based on the create barrier request; wherein, the generation processes of each barrier point are not carried out simultaneously to achieve the uniqueness and consistency of the barrier points;

[0066] S102. Pause the commit of the target distributed transaction in the distributed system to ensure that only the distributed transaction that creates the barrier is committed in the distributed system, where the target distributed transaction is any distributed transaction other than the creation of the barrier;

[0067] S103. Automatically unlock the barrier and create an LVM snapshot when all distributed nodes have completed creating the barrier. Each distributed node independently creates an LVM snapshot;

[0068] S104. Dump the LVM snapshot to the backup storage server and delete or retain the LVM snapshot.

[0069] First, in the above S101, the coordinator node in the distributed system is responsible for sending a barrier creation request, requesting to notify all participating distributed nodes to prepare to create a barrier point. Ensure that all nodes can create a barrier within the same time period, thus providing a basis for subsequent snapshots and data consistency.

[0070] The distributed nodes include a coordinator node CN, a data node DN, and a global transaction management node GTM. After each node receives the barrier creation request, it starts to generate a barrier point. It should be noted that in order to achieve the uniqueness of the barrier point, the generation process of the barrier point is not carried out simultaneously among the nodes. Thus, it is ensured that each barrier point is unique in time, avoiding state conflicts caused by concurrent generation.

[0071] In the above S102, during the creation of the barrier point, the distributed system pauses the commit of all target distributed transactions. Only the distributed transaction corresponding to the creation of the barrier is allowed to be committed, including all transactions other than the creation of the barrier, ensuring that no other transactions modify the data during the creation of the barrier.

[0072] By pausing the commit of other transactions, the distributed system can maintain a consistent data state when creating a snapshot, avoiding data inconsistency during the creation of the snapshot.

[0073] In the above S103, after all distributed nodes have completed the creation of the barrier point, the distributed system automatically unlocks the barrier. After the barrier is unlocked, each distributed node independently creates an LVM snapshot. The LVM snapshot can capture the current state of the data and provide a basis for subsequent data recovery.

[0074] Finally, through S104, the created LVM snapshot is dumped to the backup storage server, ensuring the security and persistence of the data and facilitating subsequent recovery operations. After the snapshot is dumped to the backup storage, it can prevent the loss or damage of the original data. The LVM snapshot can be deleted or retained according to needs. This flexibility makes data management more efficient and can be adjusted according to the current storage requirements.

[0075] This application receives a barrier creation request sent by a coordination node and generates barrier points based on this request. The generation process of each barrier point is not carried out simultaneously, and the process of creating barrier points on different distributed nodes is coordinated and synchronized, thus achieving the consistency of backup operations; since the generation of barrier points is not carried out simultaneously, it can ensure that a synchronized and consistent state can be achieved on all nodes, thus ensuring the integrity and consistency of data; and by waiting for all distributed nodes to complete barrier creation before creating an LVM snapshot, the creation of the LVM snapshot is carried out at a consistent moment on all nodes, thus avoiding the problem of data inconsistency, and each distributed node independently creates an LVM snapshot, achieving that the backup states of each node are independent and consistent.

[0076] Optionally, it further includes:

[0077] Generating a WAL log with a special mark based on the barrier creation request for log replay based on the WAL log;

[0078] Wherein, the special mark is used to indicate the position of the barrier point, and the special marks corresponding to each distributed node are consistent.

[0079] Furthermore, during the process of creating a barrier request, the distributed system generates a WAL log with a special mark. The log records all transaction operations and provides a basis for subsequent replay.

[0080] It should be noted that the special mark is used to indicate the position of the barrier point to ensure that all distributed nodes can accurately restore to a consistent state when recovery is required.

[0081] After backup, the distributed system can perform replay based on the WAL log and restore to the state of the barrier point, ensuring data consistency and integrity.

[0082] Furthermore, the process of the log replay includes:

[0083] Determining the creation time of the barrier points of each distributed node, and using each creation time as a cut-off point, where the cut-off point is used to make the restored database state consistent with the database state during backup;

[0084] During the recovery process, comparing the generation time of the data blocks in the database with the cut-off point and excluding the data blocks after the cut-off point;

[0085] Performing log replay based on the WAL log and restoring the data blocks to the state at the cut-off point.

[0086] Furthermore, performing log replay based on the WAL log and restoring the data blocks to the state at the cut-off point includes:

[0087] Receive the backup medium sent by the coordination node, where the backup medium includes the data files corresponding to the LVM snapshot and the WAL log;

[0088] Based on the WAL log and on the basis of the original file blocks of the database, start performing log replay from the last time point of the data file;

[0089] Replay the WAL log to the cut-off point and stop the log replay;

[0090] When the log replay of all distributed nodes is completed, determine that the database recovery is completed.

[0091] Specifically, when each distributed node in the embodiment of the present application creates a barrier point, it will record the creation timestamp of the barrier point, and this timestamp is used to define the cut-off point in the recovery process.

[0092] During the recovery process, the distributed system traverses all data blocks in the database and compares the generation time of each data block with the cut-off point. If the generation time of a data block is later than the cut-off point, these data blocks will be excluded to ensure that no inconsistent data after the barrier point is introduced during the recovery process.

[0093] The specific process of performing log replay based on the WAL log is as follows:

[0094] First, receive the backup medium sent by the coordination node. The backup medium includes the data files corresponding to the LVM snapshot and the WAL log. The backup medium provides the basic data and operation records required for recovery, ensuring that all important data and operations are completely recorded during the backup process for subsequent recovery.

[0095] Second, according to the time of the cut-off point, determine the database state at this time point to ensure that the database state after recovery is consistent with the database state during backup, and identify the corresponding original file blocks in the snapshot according to the cut-off point.

[0096] Furthermore, based on the WAL log and on the basis of the original file blocks, start performing log replay from the last time point of the data file to ensure that all valid operations are reapplied. According to the records in the WAL log, execute the operations one by one to restore to the state at the cut-off point.

[0097] Finally, during the playback process, monitor the comparison between the current state and the cut-off point. Once the cut-off point is restored, the distributed system will stop the log playback. In this way, it is ensured that no changes after the cut-off point are introduced during the recovery process, guaranteeing data consistency. After all distributed nodes complete the log playback, perform a status confirmation to ensure that the data status of each node is consistent and meets the requirements of the cut-off point. Once the log playback of all nodes is completed and the status is consistent, confirm that the database recovery is completed.

[0098] In the embodiment of the present application, by recording the creation time of the barrier point, a clear time limit is set for the recovery process. By comparing the timestamps, it is ensured that no inconsistent data blocks are introduced during the recovery process, maintaining data integrity; using the records in the WAL log, the system can accurately restore to the state at the cut-off point, ensuring data consistency.

[0099] Optionally, deleting or retaining the LVM snapshot includes:

[0100] Delete the LVM snapshot from the original storage medium to reduce the occupied snapshot space;

[0101] Or retain the storage of the LVM snapshot in the original storage medium for timely recovery in case of emergency, until the LVM snapshot is deleted when the snapshot space memory in the original storage medium reaches the storage threshold.

[0102] Specifically, in this embodiment, over time, unmanaged snapshots may occupy a large amount of storage space, causing tension in storage resources. Therefore, effective deletion or retention management of snapshots is required.

[0103] After determining that certain LVM snapshots are no longer needed, these snapshots can be deleted from the original storage medium. The deletion process should ensure that the deleted snapshots do not affect the current state of the system or the available backups.

[0104] However, in some cases, it may be desirable to retain the LVM snapshot in the original storage medium for quick recovery in case of emergency.

[0105] Furthermore, to avoid occupying too much storage space, a storage threshold can be set. When the storage space approaches this threshold, the system will prompt for snapshot deletion to prevent affecting normal operations. The system needs to have the ability to monitor the usage of storage space, track the storage space occupied by snapshots, and issue warnings in a timely manner. The threshold can be set according to the total storage capacity and the usage of snapshots. For example, it is set to trigger a warning when reaching 80% of the total space.

[0106] When the memory in the snapshot space of the original storage medium reaches the set storage threshold, the system should automatically delete the LVM snapshots that are no longer needed, thus ensuring the efficient use of storage resources and avoiding performance degradation or storage shortage problems caused by excessive snapshots.

[0107] Referring to Figure 2 , this application also provides a method for distributed database backup based on LVM snapshots, which is applied to the coordination node of the database. The method includes:

[0108] S201. Send a create barrier request to the distributed nodes. The create barrier request is used to create a barrier point and generate a WAL log with a special mark for log replay.

[0109] S202. Poll each distributed node to determine that each distributed node has successfully created a barrier point to start creating LVM snapshots.

[0110] Optionally, it further includes:

[0111] Obtain the backup media of each distributed node from the LVM snapshot or the dump file of the backup storage server; the backup media includes data files and WAL logs.

[0112] Distribute the backup media to each distributed node so that the distributed nodes can perform log replay.

[0113] The embodiments of this application are the embodiments of the coordination node. For the specific implementation process and beneficial effects, reference can be made to the embodiments of the distributed node, and no further elaboration will be made here.

[0114] Referring to Figure 3 , Figure 3 is a schematic flowchart of the backup process of the method for distributed database backup based on LVM snapshots of this application.

[0115] 1. Start backup.

[0116] 2. Send Create Barrier to all nodes (CN, DN, GTM distributed transaction nodes). All components involved in the distributed database system need to create a Barrier.

[0117] Create Barrier is a distributed transaction that generates a WAL log with a special mark. The same mark exists in all distributed components.

[0118] At this time, there are two important points: 1. All other distributed transactions need to suspend submission. 2. The system does not allow two Barriers to be generated simultaneously to ensure the uniqueness and consistency of the Barrier.

[0119] 3. All distributed transactions are paused from committing. The distributed system can continue to initiate new distributed transactions, but the commits are paused. The initiation and commit of local transactions are not affected. It is equivalent to only the distributed transaction of Create Barrier being able to commit in the entire system.

[0120] 4. Wait for all nodes to complete Create Barrier. The system polls each node to check if the Barrier point has been successfully created. After all participating nodes have completed, unlock the Barrier, i.e., unlock the barrier, and then the core consistency guarantee process can be completed.

[0121] 5. All nodes create LVM snapshots. At this time, the LVM snapshot process can start. There is no strict requirement for the order in which nodes create snapshots. Since the database recovery process can ensure that new blocks generated after the cut-off point will not be recovered, WAL replay can ensure that blocks will only be recovered to the state at the cut-off point.

[0122] 6. Dump and delete the snapshot. At this time, there are two options. You can choose to dump the snapshot file to a backup storage server and then delete the snapshot. Or you can choose to retain the snapshot for immediate recovery. Note: Keeping a snapshot for a long time will cause the snapshot space to expand as the original file is continuously modified.

[0123] 7. Backup completed.

[0124] Refer to Figure 4 , Figure 4 which is a flowchart of the recovery process of the method for distributed database backup based on LVM snapshots in this application, including the following steps:

[0125] 1. Start recovery.

[0126] 2. Copy the backup medium from the snapshot or dump file.

[0127] 3. Replay the WAL log. All distributed nodes replay their respective WAL logs on the basis of the original file blocks.

[0128] 4. Wait for all nodes to replay the WAL to the Barrier point. When a node replays the WAL to the Barrier point, stop the replay. Poll each node to check if they have all replayed to the Barrier point.

[0129] 5. Recovery completed. The distributed database system can be started.

[0130] Refer to Figure 5 , Figure 5 which is a schematic diagram of the principle of the embodiment of this application.

[0131] In the figure, the original file is divided into multiple blocks, indicating that a parallel processing method can be adopted to improve the computing efficiency. The barrier in the figure indicates that at certain calculation stages, different blocks need to be synchronized.

[0132] The following is described in conjunction with a specific embodiment of the present application:

[0133] 1. User A starts the following two transactions:

[0134] Transaction-1: A single-node transaction;

[0135] Transaction-2: A distributed transaction;

[0136] 2. User B starts a transaction at the primary backup node (any node in the distribution) and executes Create Barrier.

[0137] 3. User A starts the following distributed transaction:

[0138] Transaction-3: A distributed transaction;

[0139] 4. All distributed nodes start to execute Create Barrier;

[0140] 5. All nodes lock the Barrier, prohibiting continued Create Barrier and the submission of distributed transactions; prohibiting the submission of all current distributed transactions (such as Transaction-2 and Transaction-3); however, allowing new distributed transactions to be created until the Barrier is unlocked. Even if these new transactions exist, their submissions will still be blocked to ensure state consistency;

[0141] 6. Transaction-1 is successfully submitted, indicating that the operations of this node have been completed;

[0142] 7. Due to the Barrier being locked, the submissions of Transaction-2 and Transaction-3 are blocked and wait in the background (the waiting time is affected by the cluster scale and generally will be completed within <2s);

[0143] 8. All nodes have completed Create Barrier, ensuring the state consistency of each node;

[0144] 9. The primary backup node issues an unlock command to unlock all Barriers;

[0145] 10. The submission processes of Transaction-2 and Transaction-3 automatically continue;

[0146] 11. User B enables LVM snapshots on all nodes. This operation generally completes within <1s and the service is not affected.

[0147] 12. After snapshots are completed on all nodes, we obtain a complete backup medium at this time.

[0148] 13. The user transfers the backup medium from the snapshot directories of all nodes to other storage media.

[0149] 14. Start recovery.

[0150] 15. Copy the backup media of each node from the backup storage (snapshot directory or backup dump storage).

[0151] 16. Distribute to each node of the distributed cluster (including data files and WAL logs).

[0152] 17. Each node of the distributed cluster starts to replay the WAL log until reaching the Barrier marker point.

[0153] 18. Poll all nodes to check if they have reached the Barrier marker point.

[0154] 19. Recovery is completed.

[0155] Referring to Figure 6 , this application also provides a distributed database backup system based on LVM snapshots, including:

[0156] A creation module 610, configured to receive a create barrier request sent by a coordination node and create a barrier point based on the create barrier request; wherein, the generation processes of each barrier point do not occur simultaneously to achieve the uniqueness and consistency of the barrier points.

[0157] A pause module 620, configured to pause the submission of the target distributed transaction of the distributed system to ensure that only the distributed transaction creating the barrier in the distributed system is submitted, and the target distributed transaction is any distributed transaction other than the create barrier.

[0158] A generation module 630, configured to automatically unlock the barrier and create an LVM snapshot when all distributed nodes have completed creating the barrier, and each distributed node independently creates an LVM snapshot.

[0159] A dump module 640, configured to dump the LVM snapshot to a backup storage server and delete or retain the LVM snapshot.

[0160] Optionally, it further includes a log generation module, configured to:

[0161] Generate a WAL log with special markers based on the create barrier request for log replay based on the WAL log;

[0162] Among them, the special marker is used to indicate the position of the barrier point, and the special markers corresponding to each distributed node are the same.

[0163] Optionally, the process of log replay includes:

[0164] Determine the creation time of the barrier point for each distributed node, and use each creation time as a cut-off point, which is used to keep the restored database state consistent with the database state during backup;

[0165] During the recovery process, compare the generation time of the data blocks in the database with the cut-off point, and exclude the data blocks after the cut-off point;

[0166] Execute log replay based on the WAL log and restore the data blocks to the state at the cut-off point.

[0167] Optionally, executing log replay based on the WAL log and restoring the data blocks to the state at the cut-off point includes:

[0168] Receive the backup medium sent by the coordination node, where the backup medium includes the data files corresponding to the LVM snapshot and the WAL log;

[0169] Based on the WAL log and on the basis of the original file blocks, start log replay from the cut-off point;

[0170] Replay the WAL log to the barrier point and stop log replay;

[0171] When the log replay of all distributed nodes is completed, determine that the database recovery is completed.

[0172] It can be understood that for the detailed function implementation of the above-mentioned each unit / module, reference can be made to the introduction in the foregoing method embodiments, and details are not described herein.

[0173] It should be understood that the above device is used to execute the method in the above embodiment. For the corresponding program modules in the device, their implementation principles and technical effects are similar to the descriptions in the above method. The working process of the device can refer to the corresponding process in the above method, and details are not described herein.

[0174] Refer to Figure 7, based on the method in the above embodiments, an embodiment of the present application provides an electronic device, which may include: a processor 710, a communications interface 720, a memory 730, and a communication bus 740. Among them, the processor 710, the communications interface 720, and the memory 730 communicate with each other through the communication bus 740. The processor 710 can call the logical instructions in the memory 730 to execute the method in the above embodiments.

[0175] In addition, when the logical instructions in the above-mentioned memory 730 are implemented in the form of a software functional unit and sold or used as an independent product, they can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the method described in each embodiment of the present application.

[0176] Based on the method in the above embodiments, an embodiment of the present application provides a computer-readable storage medium. The computer-readable storage medium stores a computer program. When the computer program runs on a processor, it causes the processor to execute the method in the above embodiments.

[0177] Based on the method in the above embodiments, an embodiment of the present application provides a computer program product. When the computer program product runs on a processor, it causes the processor to execute the method in the above embodiments.

[0178] It can be understood that the processor in the embodiments of the present application may be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), or other programmable logic devices, transistor logic devices, hardware components, or any combination thereof. The general-purpose processor may be a microprocessor or any conventional processor.

[0179] The method steps in the embodiments of the present application can be implemented in a hardware manner or by a processor executing software instructions. The software instructions can be composed of corresponding software modules, and the software modules can be stored in a random access memory (RAM), flash memory, read-only memory (ROM), programmable ROM (PROM), erasable PROM (EPROM), electrically erasable PROM (EEPROM), register, hard disk, removable hard disk, CD-ROM, or any other form of storage medium well-known in the art. An exemplary storage medium is coupled to the processor, enabling the processor to read information from the storage medium and write information to the storage medium. Of course, the storage medium can also be a component of the processor. The processor and the storage medium can be located in an ASIC.

[0180] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions described in the embodiments of the present application are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium or transmitted through the computer-readable storage medium. The computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center in a wired manner (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or a wireless manner (such as infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium that the computer can access or a data storage device such as a server or data center that includes one or more integrated available media. The available medium can be a magnetic medium (such as a floppy disk, hard disk, magnetic tape), an optical medium (such as a DVD), or a semiconductor medium (such as a solid state disk (SSD)), etc.

[0181] It can be understood that the various numerical numbers involved in the embodiments of the present application are only for the convenience of description and are not used to limit the scope of the embodiments of the present application.

[0182] Those skilled in the art can easily understand that the above are only the preferred embodiments of the present application and are not intended to limit the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present application shall be included in the protection scope of the present application.

Claims

1. A method for backing up a distributed database based on LVM snapshots, characterized in that: Applied to a distributed node of a database, the method comprises: Receiving a barrier creation request sent by a coordination node, and creating a barrier point based on the barrier creation request; wherein the generation process of each barrier point is not performed simultaneously to achieve uniqueness and consistency of the barrier point; Suspending the submission of a target distributed transaction of the distributed system, ensuring that only a distributed transaction that creates a barrier exists in the distributed system for submission, and the target distributed transaction is any distributed transaction except the one that creates the barrier; When all distributed nodes have completed barrier creation, the barrier is automatically unlocked and an LVM snapshot is created. Each distributed node creates an LVM snapshot independently. The LVM snapshot is dumped to a backup storage server, and the LVM snapshot is deleted or retained.

2. The method for distributed database backup based on LVM snapshot according to claim 1, characterized in that: Also includes: Generate a WAL log with a special mark based on the barrier creation request, so as to perform log playback based on the WAL log; The special mark is used to indicate the position of the barrier point, and the special mark corresponding to each distributed node is consistent.

3. The method for distributed database backup based on LVM snapshot according to claim 2, characterized in that: The log playback process includes: Determine the creation time of the barrier point of each distributed node, and use each creation time as a cutoff point, wherein the cutoff point is used to make the database state after recovery consistent with the database state at the time of backup; During the recovery process, the generation time of the data blocks of the database is compared with the cutoff point, and the data blocks after the cutoff point are excluded; Log playback is performed based on the WAL log to restore the data block to the state at the cutoff point.

4. The method for distributed database backup based on LVM snapshot according to claim 3, characterized in that: Log playback is performed based on the WAL log to restore the data block to the state at the cutoff point, including: Receive the backup medium sent by the coordinating node, where the backup medium includes the data file and WAL log corresponding to the LVM snapshot; Based on the original file blocks of the WAL log in the database, log playback is performed from the last time point of the data file; Replay the WAL log to the cutoff point and stop replaying the log; When the log playback of all distributed nodes is completed, it is determined that the database recovery is complete.

5. The method for distributed database backup based on LVM snapshot according to claim 1, characterized in that: Deleting or retaining LVM snapshots, including: Deleting the LVM snapshot from the original storage medium to reduce the occupied snapshot space; Or the LVM snapshot storage is retained in the original storage medium for timely recovery in case of emergency, until the LVM snapshot is deleted when the snapshot space memory of the original storage medium reaches the storage threshold.

6. A method for backing up a distributed database based on LVM snapshots, characterized in that: Applied to a coordination node of a database, the method comprises: Send a barrier creation request to the distributed node, where the barrier creation request is used to create a barrier point and generate a WAL log with a special mark for log playback; Polling each distributed node to determine that each of the distributed nodes has successfully created a barrier point to start creating an LVM snapshot; Among them, the generation process of each barrier point is not carried out simultaneously to achieve the uniqueness and consistency of the barrier point; during the creation of the barrier point, the distributed system suspends the submission of the target distributed transaction to ensure that only the distributed transaction that creates the barrier exists in the distributed system for submission, and the target distributed transaction is any distributed transaction except the one that creates the barrier.

7. The method for distributed database backup based on LVM snapshot according to claim 6, characterized in that: Also includes: Obtaining a backup medium of each distributed node from the dump file of the LVM snapshot or the backup storage server; the backup medium includes a data file and a WAL log; The backup medium is distributed to each distributed node so that the distributed node can perform log playback.

8. A distributed database backup system based on LVM snapshot, characterized in that: include: A creation module, configured to receive a barrier creation request sent by a coordination node, and to create a barrier point based on the barrier creation request; wherein the generation process of each barrier point is not performed simultaneously to achieve uniqueness and consistency of the barrier point; A suspension module, used for suspending the submission of a target distributed transaction of a distributed system, ensuring that only a distributed transaction that creates a barrier exists in the distributed system for submission, and the target distributed transaction is any distributed transaction except the one that creates the barrier; A generation module is used to automatically unlock the barrier and create an LVM snapshot when all distributed nodes have completed the barrier creation. Each distributed node creates an LVM snapshot independently. The dump module is used to dump the LVM snapshot to a backup storage server, and delete or retain the LVM snapshot.

9. An electronic device, characterized in that: include: at least one memory for storing a computer program; At least one processor is used to execute the program stored in the memory. When the program stored in the memory is executed, the processor is used to execute the method according to any one of claims 1 to 7.

10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program runs on a processor, the processor is caused to execute the method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Data backup method and device for multiple storage engines, electronic equipment and storage medium

    CN115408200A

  • Transaction distributed data synchronization method, device and system and storage medium

    CN116610752A