Data backup, data backup recovery method and device for a distributed database

By using the first and second record files in a distributed database, combined with a multi-level checkpoint mechanism, quickly locate and restore the partitions of data backup, the problem of low data backup efficiency after interruption is solved, and efficient data recovery and system stability are achieved.

CN119847826BActive Publication Date: 2025-07-22BEIJING OCEANBASE TECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510320130.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-18
Publication Date
2025-07-22
Estimated Expiration
2045-03-18

AI Technical Summary

Technical Problem

In distributed systems, the prior art cannot efficiently restore data backups interrupted due to network fluctuations or hardware failures, resulting in the recovery process being time-consuming and inefficient.

Method used

By storing the first record file corresponding to the backup partition group and the second record file corresponding to the backup partition in the target storage medium of the distributed database, the partition of data backup is quickly positioned and restored, and a multi-level checkpoint mechanism is adopted to ensure accuracy and efficiency.

Benefits of technology

It significantly improves the recovery efficiency of data backup, reduces repetitive labor, shortens recovery time after interruption, and improves system stability and user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119847826B_ABST
    Figure CN119847826B_ABST
Patent Text Reader

Abstract

An embodiment of this specification provides a data backup and recovery method for a distributed database, which is executed by a target node in the distributed database, and includes: determining to resume an interrupted backup operation for target data, where the target data includes multiple partitions, and the multiple partitions are divided into several partition groups. Search for a first record file in a target storage medium for data backup, which is used to record the list of partitions in the partition groups that have completed backup. According to the searched first record file, determine multiple candidate partitions to be processed. For each candidate partition, find a second record file corresponding to the candidate partition in the target storage medium, and according to the search result, determine whether to back up the candidate partition, where the second record file is used to mark the corresponding partition as a backed-up partition.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] One or more embodiments of this specification relate to the field of databases, and in particular to a data backup, data backup recovery method, and device for a distributed database. Background Art

[0002] In a distributed system environment, especially in large-scale data storage scenarios, data backup is a crucial task. However, during actual operation, due to various unforeseeable factors such as network fluctuations and hardware failures, the ongoing data backup may be interrupted.

[0003] In traditional techniques, the method for resuming an interrupted data backup is to restart the entire backup process, which is not only time-consuming but also inefficient. Therefore, a more efficient data backup recovery solution is needed. Summary of the Invention

[0004] One or more embodiments of this specification describe a data backup recovery method for a distributed database, which can significantly improve the overall efficiency of data backup.

[0005] In a first aspect, there is provided a data backup recovery method for a distributed database, which is executed by a target node in the distributed database, and includes:

[0006] Determine to resume an interrupted backup operation for target data, where the target data includes multiple partitions, and the multiple partitions are divided into several partition groups;

[0007] Search for a first record file in a target storage medium for data backup, where the first record file is used to record the partition list in the partition groups that have completed backup;

[0008] Determine multiple candidate partitions to be processed according to the searched first record file;

[0009] For each candidate partition, search for a second record file corresponding to the candidate partition in the target storage medium, and determine whether to back up the candidate partition according to the search result, where the second record file is used to mark the corresponding partition as a partition that has been backed up.

[0010] In a second aspect, there is provided a data backup method for a distributed database, which is executed by any first node in the distributed database, and includes:

[0011] Determine several target partition groups to be backed up by the first node; the several target partition groups are at least part of the partition groups obtained by grouping multiple partitions included in the target data to be backed up;

[0012] Back up the several target partition groups, and make the first record and the second record;

[0013] Among them, the first record includes that after any target partition group is backed up, a corresponding first record file is created in the target storage medium for data backup, and the first record file is used to record the partition list in the target partition group;

[0014] The second record includes that after any target partition in any target partition group is backed up, a corresponding second record file is created in the target storage medium, which is used to mark the target partition as a backed-up partition; among them, the first and second record files are used to resume the interrupted backup operation for the target data.

[0015] In a third aspect, a data backup and recovery device for a distributed database is provided, which is set in a target node in the distributed database and includes:

[0016] A determination unit, configured to determine to resume the interrupted backup operation for target data, where the target data includes multiple partitions, and the multiple partitions are divided into several partition groups;

[0017] A search unit, configured to search for a first record file in the target storage medium for data backup, and the first record file is used to record the partition list in the partition group that has been backed up;

[0018] The determination unit is further configured to determine multiple candidate partitions to be processed according to the searched first record file;

[0019] A search unit, configured to, for each candidate partition, search for a second record file corresponding to the candidate partition in the target storage medium, and determine whether to back up the candidate partition according to the search result, where the second record file is used to mark the corresponding partition as a backed-up partition.

[0020] In a fourth aspect, a data backup device for a distributed database is provided, which is set in any first node in the distributed database and includes:

[0021] A determination unit, configured to determine several target partition groups to be backed up by the first node; the several target partition groups are at least part of the partition groups obtained by grouping multiple partitions included in the target data to be backed up;

[0022] A backup unit, configured to back up the several target partition groups, and make the first record and the second record;

[0023] The backup unit includes:

[0024] Create a sub-module for creating a corresponding first record file in the target storage medium for data backup after any target partition group completes backup. The first record file is used to record the partition list in the target partition group;

[0025] After any target partition in any target partition group completes backup, create a corresponding second record file in the target storage medium, which is used to mark the target partition as a backed-up partition; wherein, the first and second record files are used to resume an interrupted backup operation for the target data.

[0026] In a fifth aspect, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed on a computer, the computer is made to execute the method of the first or second aspect.

[0027] In a sixth aspect, a computing device is provided, including a memory and a processor. An executable code is stored in the memory, and when the processor executes the executable code, the method of the first or second aspect is implemented.

[0028] For the data backup and recovery method of the distributed database provided by one or more embodiments of this specification, when the data backup is interrupted, based on the first record file corresponding to the backed-up partition group and the second record file corresponding to the backed-up partition stored in the target storage medium for data backup, the partitions that need to resume the data backup can be quickly determined, and the data backup is resumed for the determined partitions. In other words, this solution can quickly locate the position where the backup fails after the data backup is interrupted and continue to execute the remaining backup process from this point, thereby significantly improving the efficiency of data backup as a whole and further improving the user experience. BRIEF DESCRIPTION OF THE DRAWINGS

[0029] To more clearly illustrate the technical solutions of the embodiments of this specification, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of this specification. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0030] Figure 1 Show a schematic diagram of a distributed database in an example of this specification;

[0031] Figure 2 Is a schematic diagram of an implementation scenario of an embodiment disclosed in this specification;

[0032] Figure 3 Show a flowchart of the data backup method of the distributed database according to an embodiment of this specification;

[0033] Figure 4 Schematic diagram showing the first and second record files in an example of this specification;

[0034] Figure 5 Flowchart of the data backup and recovery method for a distributed database according to an embodiment of this specification;

[0035] Figure 6 Schematic diagram of the data backup and recovery method in an example of this specification;

[0036] Figure 7 Schematic diagram of the data backup and recovery device for a distributed database according to an embodiment of this specification;

[0037] Figure 8 Schematic diagram of the data backup device for a distributed database according to an embodiment of this specification. Detailed implementation manners

[0038] The solutions provided in this specification will be described below with reference to the accompanying drawings.

[0039] As described above, the traditional data backup and recovery solutions have the problems of time-consuming and low efficiency. To solve this problem, there are the following several improvement solutions:

[0040] First, implement resume from breakpoint (i.e., recover data backup) through snapshots and list files. That is, first record the database snapshot information and back up the data file list to the corresponding preset list file. After that, update the file status of the data files in the preset list file after the data transfer is successful, and use this file status to achieve resume from breakpoint of the backup data during resume from breakpoint. However, this method does not consider the complexity in a distributed environment, especially how to accurately determine the uncompleted backup partitions when multiple database nodes execute data backup in parallel, and this method does not propose an effective solution.

[0041] Second, perform backup processing by using a distributed backup system, that is, split the backup task and resume from breakpoint. Although this method can reduce the time for the user to recover the backup in case of backup failure, it is less efficient when dealing with a large number of partitions and lacks fine management of the specific partition status.

[0042] Third, implement resume from breakpoint based on cloud storage technology. Although this method improves the utilization rate of storage space, it has insufficient support for resume from breakpoint at the partition level in a distributed environment.

[0043] Fourth, under the NDMP protocol, according to the context information of the file to be backed up at the transmission interruption point and the first part of the content stored before the transmission interruption, determine the second part of the content that has not been successfully transmitted at the transmission interruption point, and continue to send this second part of the content. However, this method is mainly for data backup under the NDMP protocol and is not applicable to the data backup scenario in a distributed database.

[0044] Generally speaking, the above-mentioned improved solutions have the following disadvantages:

[0045] a. It has weak support for tasks executed concurrently on a large scale;

[0046] b. It lacks the fine-grained management ability at the single partition level;

[0047] c. It performs poorly in error detection and recovery in the face of complex environments.

[0048] For this reason, in the embodiments of this specification, an improved data backup and its recovery method is proposed. First, during the data backup process, each node in the distributed database stores a first record file corresponding to the backed-up partition group and a second record file corresponding to the backed-up partition in the target storage medium for data backup. After that, after the data backup is interrupted, the target node for recovering the data backup reads the first record file and the second record file from the target storage medium, and quickly determines the partitions for which the data backup needs to be recovered based on them, and resumes the execution of the data backup for the determined partitions. In other words, this solution can quickly locate the position where the backup fails after the data backup is interrupted and continue to execute the remaining backup process from this point, thereby significantly improving the efficiency of data backup as a whole and further improving the user experience.

[0049] Figure 1 The schematic diagram of a distributed database shown in an example of this specification. Figure 1 In a distributed database, data is distributedly stored in each node (Node A - Node C). Specifically, a data table can be split into multiple data partitions (hereinafter simply referred to as partitions) according to the partition key. The split partitions are respectively stored in each node. Each node can store one or more partitions, and the partitions stored between nodes can overlap with each other.

[0050] For example, assume that there is a data table t1 in the distributed database, and this data table t1 can be split into 6 partitions, respectively represented as: p1 - p6. Then Node A can store partitions p1, p2, and p3, Node B can store partitions p3, p4, and p5, and Node C can store partitions p4, p5, and p6.

[0051] It should be understood that Figure 1This is just an exemplary illustration. In practice, a distributed database may include a larger number of nodes, and each node may also store one or more partitions split from other data tables.

[0052] Figure 2 Schematic diagram of the implementation scenario of an embodiment disclosed in this specification. Figure 2 In it, when it is necessary to back up the target data table in the distributed database, the multiple partitions split from the target data table can be first divided into several partition groups, and each partition group is respectively assigned to each node for backup. For example, partition group A (including partitions p1 - p2) is assigned to node X, and partition group B (including partitions p3 - p5) is assigned to node Y.

[0053] After that, each node can back up the assigned target partition group to the storage medium shared by each node, and create a first record file corresponding to the backed - up partition group and a second record file corresponding to the backed - up partition in the shared storage medium. Figure 2 In it, assume that all partitions in partition group A have been backed up, so the first record file corresponding to partition group A, as well as the second record files corresponding to partitions p1 and p2 respectively, are stored in the storage medium. In addition, assume that partition p5 in partition group B has not been backed up yet, so only the second record files corresponding to partitions p3 and p4 respectively are stored in the storage medium, and the first record file corresponding to partition group B is not stored.

[0054] It should be understood that for the above - interrupted backup operation, the target node for restoring the data backup reads the above - mentioned first record file and second record file from the shared storage medium, and determines the partitions that need to restore the data backup based on them and performs the backup (the specific determination method will be described later).

[0055] Figure 3 Shows the flowchart of the data backup method for a distributed database according to an embodiment of this specification. This method can be executed by any node (hereinafter referred to as the first node) in the distributed database. As Figure 3 shown, this method may include the following steps:

[0056] Step S302, determine several target partition groups to be backed up by the first node.

[0057] Among them, the above - mentioned several target partition groups can be at least some of the partition groups obtained by grouping the multiple partitions included in the target data to be backed up (including one or more data tables).

[0058] In one example, the above - mentioned multiple partitions can be grouped based on the log stream, that is, each partition in the same partition group belongs to the same log stream.

[0059] Specifically, the root service (RS) node in the distributed database can first receive the SQL statement input by the user, which indicates the target data to be backed up. Then, according to the pre-recorded partition information of the target data (i.e., the identifiers of the nodes where each partition of the target data is located), the root service node assigns corresponding partition groups to each node. Of course, in practice, the corresponding partition groups can also be assigned to each node in combination with the load information (detailed later).

[0060] After that, the root service node can send data backup instructions to each node respectively. The data backup instruction sent to any node can include the identifiers of several target partition groups to be backed up by that node. That is to say, the first node can determine several target partition groups to be backed up by it according to the data backup instruction.

[0061] Take Figure 1 as an example. The target data to be backed up can be, for example, the data table t1, and it is assumed that the following 4 partition groups are obtained for the 6 partitions it includes: Partition group 1: p1, p2; Partition group 2: p3; Partition group 3: p4, p5; Partition group 4: p6. And assume that the load information of node B is relatively large, then node A can back up partition groups 1 and 2, node B can partition partition group 3, and node C can back up partition group 4.

[0062] It should be understood that this is only an example. In practice, a partition group usually can contain 1000 or even more partitions, and each node for data backup can be some nodes in the distributed database. This specification does not make any limitations on this.

[0063] Of course, in practice, the user can also pre-configure several target partition groups for each node to back up respectively, and then each node periodically backs up the corresponding several target partition groups. In other words, the backup operation for the target data is not triggered by the user through the root service node, but is periodically executed by each node.

[0064] Step S304: Back up several target partition groups, and perform the first record and the second record.

[0065] Among them, the first node can start multiple threads and parallelly back up several target partition groups through these multiple threads, so as to improve the backup efficiency for several target partition groups.

[0066] In practice, before a first node backs up a number of target partition groups, a root service node may first perform a partition list backup operation, that is, back up a partition list (also referred to as an initial partition list) containing the identifiers of all partitions to a target storage medium shared by each node for data backup. Of course, each partition in this initial partition list is organized by partition group. For example, the initial partition list records the correspondence between the identifier of the partition group and the identifier of the partition.

[0067] Specifically, the first node can, according to the identifiers of the respective target partition groups, read the identifiers of each target partition included in each target partition group from the target storage medium. Then, the first node reads the data of each target partition based on the identifiers of each target partition, stores it in the above-mentioned target storage medium, and makes a first record and a second record.

[0068] Among them, the first record includes, after any target partition group is backed up, creating a corresponding first record file (also referred to as a partition group checkpoint file) in the above-mentioned target storage medium, and this first record file is used to record the partition list in this target partition group. In one example, this partition list is formed by the identifiers of each target partition in this target partition group.

[0069] The above-mentioned second record includes, after any target partition in any target partition group is backed up, creating a corresponding second record file (also referred to as a partition detection point file) in the above-mentioned target storage medium, and this second record file is used to mark this target partition as a backed-up partition.

[0070] In one embodiment, this second record file is a placeholder file with an empty content, which can thus save storage space.

[0071] More specifically, the above-mentioned second record file can be stored in the partition directory of the corresponding partition that has been backed up in the target storage medium.

[0072] Figure 4 A schematic diagram showing the first and second record files in an example of this specification Figure 4 Among them, the first record file can be created when all partitions in any partition group are backed up, or when any partition group is backed up, and thus can also be expressed as the first record file corresponding to the backed-up partition group. The second record file can be created when any partition is backed up, and thus can also be expressed as the second record file corresponding to the backed-up partition.

[0073] It can be seen that in this solution, during the data backup process, each node in the distributed database stores a first record file corresponding to the backed-up partition group and a second record file corresponding to the backed-up partition in a storage medium shared by multiple nodes for data backup.

[0074] In addition, as can be seen from the above description, the present solution performs data backup with a partition group as the scheduling unit, and each partition in a partition group belongs to a log stream. Therefore, the data backup method of the present solution can improve the recovery efficiency of the backed-up data.

[0075] Finally, as mentioned above, during the actual operation process, due to the influence of various unforeseeable factors such as network fluctuations and hardware failures, the ongoing data backup may be interrupted. The following describes the recovery process after the above data backup is interrupted.

[0076] Figure 5 The flowchart of the data backup recovery method of a distributed database according to an embodiment of the present specification is shown. This method can be executed by a target node in the distributed database.

[0077] In one example, the target node can be selected by the root service node from each node according to the load information of each node. For example, a node with relatively small load information is selected as the target node. Of course, in practice, the target node can also be selected in combination with partition information. For example, the root service RS node can analyze the log stream with backup failure, and then select the node storing each partition belonging to the log stream as the above target node.

[0078] Take Figure 1 as an example. Suppose the backups of partition p4 and partition p5 fail, then node B or node C can be selected as the target node.

[0079] In another example, the target node can also be specified in advance by the user.

[0080] As Figure 5 shown, the method may include the following steps:

[0081] Step S502, determine to resume the interrupted backup operation for the target data.

[0082] Among them, when the target node is dynamically selected by the root service node in combination with partition information and load information, etc., the target node can receive the data backup recovery instruction sent by the root service node, and then determine that data backup recovery needs to be performed according to this data backup recovery instruction.

[0083] When the target node is specified in advance by the user, it can detect the backup operations of each node for the target data, and after detecting that the backup operation of any node is interrupted, determine that data backup recovery needs to be performed.

[0084] Step S504, search for a first record file in a target storage medium for data backup, where the first record file is used to record a list of partitions in a partition group for which backup has been completed.

[0085] In one example, the above partition list is formed by partition identifiers of each partition in the partition group for which backup has been completed.

[0086] As mentioned above, the first record file corresponds to the backed-up partition group, or rather, it is created when a certain partition group has completed backup, so that based on it, the partitions for which backup has been completed can be determined.

[0087] In addition, it should be understood that when multiple partition groups have completed backup, multiple first record files can be searched.

[0088] Step S506, determine multiple candidate partitions to be processed according to the searched first record file.

[0089] Among them, the multiple candidate partitions to be processed here can also be understood as multiple candidate partitions for restoring backup.

[0090] Specifically, an initial partition list (including identifiers of multiple partitions divided for the target data) can be read from the target storage medium, and then by comparing the initial partition list with the partition list in the searched first record file, multiple candidate partitions to be processed are determined. For example, partitions included in the partition list of each first record file can be excluded from the initial partition list, and the remaining partitions are determined as multiple candidate partitions to be processed. Here, excluding the partitions included in the partition list of the first record file is because these partitions have all completed backup.

[0091] Of course, in practice, multiple candidate partitions to be processed can also be determined by other means, which are not limited in this specification.

[0092] Step S508, for each candidate partition, search for a second record file corresponding to the candidate partition in the target storage medium, and determine whether to back up the candidate partition according to the search result, where the second record file is used to mark the corresponding partition as a backed-up partition.

[0093] As mentioned above, the second record file corresponds to the backed-up partition, and it is usually stored in the partition directory where the corresponding partition belongs. Thus, for each candidate partition, the second record file corresponding to the candidate partition can be searched from the partition directory corresponding to the candidate partition.

[0094] In one embodiment, the second record file is a placeholder file with empty content. Thus, for any candidate partition, if the corresponding second record file is found, the candidate partition is skipped; if the corresponding second record file is not found, the candidate partition is backed up.

[0095] In another embodiment, the partition identifier of the corresponding partition is recorded in the second record file. Thus, each candidate partition can be screened based on the partition identifiers recorded in the found second record files, and the screened candidate partitions are backed up.

[0096] Figure 6 Shows a schematic diagram of the data backup and recovery method in an example of this specification. Figure 6 In it, assume that the target data to be backed up includes partitions p1 - p6, and partitions p1, p3, and p5 belong to partition group A, while partitions p2, p4, and p6 belong to partition group B. After backing up the target data, assume that in the target storage medium used for data backup, there are a first record file corresponding to partition group A (recording the identifiers of partitions p1, p3, and p5 respectively), a second record file corresponding to partition p2, and a second record file corresponding to partition p4. Then, based on the first record file, partitions p1, p3, and p5 can be excluded, and based on the two second record files, partitions p2 and p4 can be excluded. Finally, it can be determined that the partition for which the backup needs to be restored is partition p6.

[0097] Generally speaking, in this solution, based on the first record file corresponding to the backed-up partition group, an initial screening of each initial partition can be achieved. And since this initial screening process is carried out in units of partition groups, the efficiency of data backup and recovery can be ensured. Based on the second record file corresponding to the backed-up partition, a secondary screening of the remaining partitions can be achieved, thus avoiding backing up the partitions that have already been backed up, and further saving time and resources.

[0098] It should be noted that in this solution, based on the first record file and the second record file, determining the partition for which the data backup needs to be restored actually adopts a multi-level checkpoint mechanism. By adopting this multi-level checkpoint mechanism, the state changes of each step of the operation can be tracked more accurately, and important information can be ensured not to be lost even in extreme cases. In addition, this solution also takes into account the requirements of cross-node collaboration, so that good consistency and reliability can be maintained even in a distributed deployment mode. The final effect is to significantly shorten the time period required from fault discovery to normal operation restoration, and improve the stability and maintainability of the system.

[0099] This solution can bring the following technical effects compared with the prior art:

[0100] 1. Precise breakpoint resumption: Adopting a multi-level checkpoint mechanism, it can accurately determine which partitions have been backed up and which need to be continuously backed up.

[0101] 2. Efficient state management: Based on the first and second record files, it reduces the amount of state to be verified one by one, improving the efficiency of data backup and recovery.

[0102] 3. Support for parallel processing: Even in a parallel processing environment, it can accurately identify the partitions with incomplete backups and continue the backup.

[0103] 4. Reduction of repetitive labor: Avoids unnecessary duplicate backups, saving time and resources.

[0104] 5. Improvement of backup efficiency: Adopting a multi-level checkpoint mechanism significantly shortens the recovery time after a backup interruption.

[0105] 6. Enhancement of system stability: In a distributed environment, it improves the reliability and stability of backup operations.

[0106] The innovative points of this solution are further described as follows:

[0107] 1. Multi-level checkpoint mechanism: By combining the first record file corresponding to the group of backed-up partitions and the second record file corresponding to the backed-up partitions, it can ensure accurate resumption of backup from the breakpoint after a data backup interruption, and can accurately determine which partitions have not been backed up even when some partitions in the partition group fail to be backed up.

[0108] 2. Flexible task scheduling mechanism: Allows the root service node to dynamically adjust the workloads borne by each node according to the actual situation, optimizing resource utilization.

[0109] 3. Support for parallel processing: In a distributed environment, it can effectively handle parallel backup operations, ensuring that the status of each partition is accurately recorded and managed.

[0110] 4. Efficient recovery process: Adopting a multi-level checkpoint mechanism reduces the time overhead required in the process of recovering from a data backup failure, improving the overall efficiency of data backup.

[0111] Corresponding to the data backup and recovery method of a distributed database, an embodiment of this specification also provides a data backup and recovery device for a distributed database, which is set in a target node in the distributed database. As Figure 7 shown, the device may include:

[0112] A determination unit 702, configured to determine to resume a backup operation for target data whose recovery is interrupted, where the target data includes multiple partitions, and the multiple partitions are divided into several partition groups.

[0113] A search unit 704 is configured to search for a first record file in a target storage medium for data backup, where the first record file is used to record a list of partitions in a partition group for which backup has been completed.

[0114] A determination unit 702 is further configured to determine multiple candidate partitions to be processed according to the searched first record file.

[0115] A search unit 706 is configured to, for each candidate partition, search for a second record file corresponding to the candidate partition in the target storage medium, and determine whether to back up the candidate partition according to the search result, where the second record file is used to mark the corresponding partition as a partition that has been backed up.

[0116] In one embodiment, the determination unit 702 includes:

[0117] A reading sub-module 7022 is configured to read an initial partition list from the target storage medium, where the initial partition list includes multiple partitions;

[0118] A comparison sub-module 7024 is configured to determine multiple candidate partitions to be processed by comparing the initial partition list with the partition list in the searched first record file.

[0119] In one embodiment, the search unit 706 is specifically configured to:

[0120] Search for a second record file corresponding to the candidate partition in the partition directory corresponding to the candidate partition in the target storage medium.

[0121] In one embodiment, the second record file is a placeholder file and its content is empty.

[0122] In one embodiment, each partition in a single partition group belongs to the same log stream.

[0123] In one embodiment, the above target node is selected from other nodes by a root service RS node in a distributed database at least according to the load information of other nodes except the root service RS node.

[0124] The functions of the functional modules of the device in the above embodiments of this specification can be implemented by the steps in the above method embodiments. Therefore, the specific working process of the device provided in an embodiment of this specification will not be repeated here.

[0125] The data recovery device of the distributed database provided in an embodiment of this specification can significantly improve the overall efficiency of data backup.

[0126] Correspondingly to the data backup method of the distributed database, an embodiment of this specification further provides a data backup device for the distributed database, which is set in any first node in the distributed database, such as Figure 8 shown, the device may include:

[0127] A determination unit 802, configured to determine a plurality of target partition groups to be backed up by the first node, and the plurality of target partition groups are at least some of the partition groups obtained by grouping multiple partitions included in the target data to be backed up.

[0128] A backup unit 804, configured to back up the plurality of target partition groups, and perform a first record and a second record.

[0129] The backup unit 804 includes:

[0130] A creation sub-module 8042, configured to create a corresponding first record file in the target storage medium for data backup after any target partition group is backed up, and the first record file is used to record the partition list in the target partition group;

[0131] After any target partition in any target partition group is backed up, create a corresponding second record file in the target storage medium, which is used to mark the target partition as a backed-up partition; wherein, the first and second record files are used to resume the interrupted backup operation for the target data.

[0132] In an embodiment, the backup unit 804 further includes:

[0133] A start sub-module 8044, configured to start a plurality of threads, and back up the plurality of target groups in parallel through the plurality of threads.

[0134] The functions of the functional modules of the device in the above embodiments of this specification can be implemented by the steps of the above method embodiments. Therefore, the specific working process of the device provided by an embodiment of this specification will not be repeated here.

[0135] The data backup device for the distributed database provided by an embodiment of this specification will create the first and second record files during the data backup process, and these two record files are helpful to quickly determine the partitions that need to restore the data backup.

[0136] According to an embodiment of another aspect, there is also provided a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed in a computer, the computer is made to execute the method described in combination with Figure 3 or Figure 5 described.

[0137] According to an embodiment of still another aspect, a computing device is further provided, including a memory and a processor. An executable code is stored in the memory. When the processor executes the executable code, the method described in conjunction with Figure 3 or Figure 5 is implemented.

[0138] Each embodiment in this specification is described in a progressive manner. For the same or similar parts among the embodiments, reference can be made to each other. The key point of each embodiment is to illustrate the differences from other embodiments. In particular, for the medium or device embodiments, since they are basically similar to the method embodiments, the description is relatively simple, and reference can be made to the partial description of the method embodiments for the relevant parts.

[0139] The specific embodiments of the present specification are described above. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims may be performed in a different order than in the embodiments and still achieve the desired results. Additionally, the processes depicted in the drawings do not necessarily require the specific order or sequential order shown to achieve the desired results. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0140] The specific implementation manners described above further elaborate on the purpose, technical solution, and beneficial effects of this specification. It should be understood that the above are only the specific implementation manners of this specification and are not used to limit the protection scope of this specification. Any modifications, equivalent replacements, improvements, etc. made on the basis of the technical solution of this specification should be included within the protection scope of this specification.

Claims

1. A method for data backup and recovery of a distributed database, which is executed by a target node in the distributed database, includes: Determine to resume a backup operation for target data that was interrupted, where the target data includes multiple partitions, and the multiple partitions are divided into several partition groups; wherein, each partition in a single partition group belongs to the same log stream; Search for a first record file in a target storage medium for data backup, where the first record file is used to record the partition list in the partition group that has completed backup; Determine multiple candidate partitions to be processed based on the searched first record file; For each candidate partition, find a second record file corresponding to the candidate partition in the target storage medium, and based on the search result, determine whether to back up the candidate partition, where the second record file is used to mark the corresponding partition as a backed-up partition.

2. The method according to claim 1, wherein The determining multiple candidate partitions to be processed includes: Read an initial partition list from the target storage medium, which includes the multiple partitions; Determine multiple candidate partitions to be processed by comparing the initial partition list and the partition list in the searched first record file.

3. The method according to claim 1, wherein, The finding the second record file corresponding to the candidate partition in the target storage medium includes: Find the second record file corresponding to the candidate partition under the partition directory corresponding to the candidate partition in the target storage medium.

4. The method according to claim 1, wherein, The second record file is a placeholder file with an empty content.

5. The method according to claim 1, wherein The target node is selected from other nodes by the root service RS node in the distributed database, at least based on the load information of other nodes except the root service RS node.

6. A method for data backup of a distributed database, which is executed by any first node in the distributed database, includes: Determine several target partition groups to be backed up by the first node; The several target partition groups are at least part of the partition groups obtained by grouping the multiple partitions included in the target data to be backed up; wherein, each partition in a single partition group belongs to the same log stream; Back up the several target partition groups and perform first recording and second recording; Among them, the first recording includes, after any target partition group is backed up, creating a corresponding first record file in the target storage medium for data backup, and the first record file is used to record the partition list in the target partition group; The second recording includes, after any target partition in any target partition group is backed up, creating a corresponding second record file in the target storage medium, which is used to mark the target partition as a backed-up partition; wherein, the first and second record files are used to resume an interrupted backup operation for the target data.

7. The method according to claim 6, wherein The backing up the several target partition groups includes: Start multiple threads and back up the several target partition groups in parallel through the multiple threads.

8. A data backup and recovery device of a distributed database, which is set in the target node in the distributed database, includes: A determination unit, configured to determine to resume an interrupted backup operation for target data, where the target data includes multiple partitions, and the multiple partitions are divided into several partition groups; wherein, each partition in a single partition group belongs to the same log stream; A search unit, configured to search for a first record file in a target storage medium for data backup, where the first record file is used to record a list of partitions in the partition groups that have completed backup; The determination unit is further configured to determine multiple candidate partitions to be processed according to the searched first record file; A lookup unit, configured to, for each candidate partition, look up a second record file corresponding to the candidate partition in the target storage medium, and determine whether to back up the candidate partition according to the lookup result, where the second record file is used to mark the corresponding partition as a backed-up partition.

9. The device according to claim 8, wherein, The determination unit includes: A reading sub-module, configured to read an initial partition list from the target storage medium, where the initial partition list includes the multiple partitions; A comparison sub-module, configured to determine multiple candidate partitions to be processed by comparing the initial partition list and the partition list in the searched first record file.

10. The device according to claim 8, wherein, The lookup unit is specifically configured to: Look up a second record file corresponding to the candidate partition in the partition directory corresponding to the candidate partition in the target storage medium.

11. The device according to claim 8, wherein, The second record file is a placeholder file, and its content is empty.

12. The device according to claim 8, wherein, The target node is selected from other nodes by the root service RS node in the distributed database according to the load information of at least other nodes except the root service RS node.

13. A data backup device for a distributed database, disposed in any first node in the distributed database, includes: A determination unit, configured to determine several target partition groups to be backed up by the first node; The several target partition groups are at least some of the partition groups obtained by grouping the multiple partitions included in the target data to be backed up; wherein, each partition in a single partition group belongs to the same log stream; A backup unit, configured to back up the several target partition groups, and perform a first record and a second record; The backup unit includes: A creation sub-module, configured to create a corresponding first record file in a target storage medium for data backup after any target partition group is backed up, where the first record file is used to record a list of partitions in the target partition group; After any target partition in any target partition group is backed up, create a corresponding second record file in the target storage medium, which is used to mark the target partition as a backed-up partition; wherein, the first and second record files are used to resume an interrupted backup operation for the target data.

14. The apparatus according to claim 13, wherein, The backup unit further includes: A start sub-module, configured to start multiple threads, and back up the several target groups in parallel through the multiple threads.

15. A computer-readable storage medium having a computer program stored thereon, wherein, When the computer program is executed on a computer, the computer is made to execute the method according to any one of claims 1-7.

16. A computing device, comprising a memory and a processor, wherein, An executable code is stored in the memory, and when the processor executes the executable code, the method according to any one of claims 1-7 is implemented.

Citation Information

Patent Citations

  • Data breakpoint resuming method under distributed backup system and application

    CN116594807A

  • Data backup, recovery and query method and device for distributed database

    CN118519827A