A Backup Data Recovery Method and Device Based on the FUSE System

Through the backup data recovery method based on the FUSE system, the recovery data block information table is generated, which solves the problem of low recovery efficiency in the existing technology, and realizes efficient data recovery, especially in large-scale data processing, which significantly improves the recovery speed.

CN119988103BActive Publication Date: 2025-07-11DISJIE (BEIJING) DATA MANAGEMENT TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510466763.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-15
Publication Date
2025-07-11
Estimated Expiration
2045-04-15

AI Technical Summary

Technical Problem

The prior art has low recovery efficiency when performing full and incremental backups of block equipment, and requires operation of the entire disk, resulting in a long recovery time.

Method used

Through the backup data recovery method based on the FUSE system, a recovery data block information table is generated, and the data block information that needs to be restored is determined based on the backup summary file and version number, which avoids reading and writing operations on the entire disk, and uses chunking compression and differential file recording to optimize the recovery process.

Benefits of technology

It significantly improves data recovery efficiency, especially when processing large-scale data, the recovery time is shortened to seconds, ensuring that the recovery operations of different users do not interfere with each other.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119988103B_ABST
    Figure CN119988103B_ABST
Patent Text Reader

Abstract

The present invention discloses a backup data recovery method based on the FUSE system, comprising the following steps: receiving a data recovery request; when traversing the information of each data block in the backup summary file, obtaining the full amount of information corresponding to the recovery version number in the information of the current data block; determining whether the file capacity of the full amount of information corresponding to the recovery version number is 0; if so, generating a recovery data block information table with empty recovery data information; if not, generating a recovery data block information table according to the latest version backup summary file of the current data block; saving the recovery data block information table until the traversal is completed; in response to a request for reading recovery data, obtaining the recovery data storage path; and sending the read recovery data to the receiving platform according to the recovery data receiving path. The present invention improves the data recovery efficiency and the data processing efficiency. The present invention also discloses an apparatus, an electronic device, and a computer-readable storage medium for implementing the above method.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of server data management, and in particular to a backup data recovery method and device based on the FUSE system. Background Art

[0002] When the prior art performs a full disk backup for a block device, it will generate a dump file by traversing the entire disk. The dump file: is a file composed of data blocks and hole blocks. During incremental backup, a corresponding differential file will be generated according to the difference in this data change. The differential file is used to save the original data overwritten by the new incremental data. Each time a full backup and an incremental backup are performed, a summary file of the differential file and the full file will be generated. One incremental backup corresponds to a set of summary files and differential files. At the same time, the bitmap file also needs to be updated during incremental backup. When recovering data, it is necessary to recover based on the dump file, summary file, differential file, and bitmap file. When recovering, only the entire disk data can be recovered, and the entire disk needs to be operated, resulting in low recovery efficiency. Summary of the Invention

[0003] In order to solve the above problems existing in the prior art, the present invention provides a backup data recovery method and device based on the FUSE system. The technical problems to be solved by the present invention are realized through the following technical solutions:

[0004] A first aspect of an embodiment of the present invention provides a backup data recovery method based on the FUSE system, including the following steps:

[0005] Receiving a data recovery request; wherein, the data recovery request indicates the path of the backup virtual machine, the path for receiving the recovered data, and the recovery version number;

[0006] Traversal step: When traversing the information of each data block in the backup summary file, obtaining the full amount information corresponding to the recovery version number in the information of the current data block;

[0007] Judging whether the file capacity of the full amount information corresponding to the recovery version number is 0;

[0008] If so, generating a recovery data block information table in which the recovery data information except the recovery version number and the data block offset is empty according to the recovery version number;

[0009] If not, generating a recovery data block information table according to the recovery version number and the latest version backup summary file of the current data block; wherein, the recovery data block information table includes recovery data information, and the recovery data information includes: recovery version number, data block offset, file capacity, storage file name, and compression level;

[0010] Save the recovery data block information table, traverse the information of the next data block in the backup summary file, and return to execute the traversal step until the traversal is completed;

[0011] In response to a recovery data reading request, obtain a recovery data storage path according to the recovery data block information table and the path of the backup virtual machine;

[0012] Read the recovery data according to the recovery data storage path, and send the read recovery data to the receiving platform according to the recovery data receiving path.

[0013] In an embodiment of the present invention, generating a recovery data block information table according to the recovery version number and the latest version backup summary file of the current data block includes:

[0014] Judge whether the recovery version number is equal to the version number of the full amount information in the latest version backup summary file of the current data block;

[0015] If they are equal, copy the full amount information in the latest version backup summary file of the current data block, and add the full amount data storage file name as the recovery data block information table;

[0016] If they are not equal, judge whether the recovery version number is equal to the version number of the modified information in the latest version backup summary file of the current data block;

[0017] If they are equal, when the recovery version number is equal to the version number of the last piece of modified information in the latest version backup summary file of the current data block, copy the full amount information in the latest version backup summary file of the current data block, and add the full amount data storage file name as the recovery data block information table;

[0018] When the recovery version number is less than the version number of the last piece of modified information in the latest version backup summary file of the current data block, copy the next piece of modified information with the same version number as the recovery version number in the modified information of the latest version backup summary file of the current data block, and add the modified data storage file name of the next piece of modified information as the recovery data block information table;

[0019] If they are not equal, judge whether the recovery version number is less than the version number of the last piece of modified information in the latest version backup summary file of the current data block;

[0020] If it is less, copy the first piece of modified information with a version number greater than the recovery version number in the modified information of the latest version backup summary file of the current data block, and add the modified data storage file name of the first piece of modified information as the recovery data block information table;

[0021] If it is greater than, copy the full - volume information in the latest version backup summary file of the current data block, and add the full - volume data storage file name to form a recovery data block information table.

[0022] In an embodiment of the present invention, the obtaining the recovery data storage path according to the recovery data block information table and the path of the backup virtual machine includes:

[0023] Obtain the recovery data storage path according to the recovery version number, storage file name in the recovery data block information table and the path of the backup virtual machine.

[0024] In an embodiment of the present invention, the method further includes: creating a cache file for storing the written data under the recovery data receiving path.

[0025] A second aspect of the embodiments of the present invention provides a backup data recovery device based on the FUSE system, including:

[0026] A receiving module, configured to receive a data recovery request; wherein, the data recovery request indicates the path of the backup virtual machine, the recovery data receiving path and the recovery version number;

[0027] A traversing module, configured to execute a traversing step: when traversing the information of each data block in the backup summary file, obtain the full - volume information corresponding to the recovery version number in the information of the current data block;

[0028] A judging module, configured to judge whether the file capacity of the full - volume information corresponding to the recovery version number is 0;

[0029] A first generating module, configured to if so, generate a recovery data block information table in which the recovery data information except the recovery version number and the data block offset is empty according to the recovery version number;

[0030] A second generating module, configured to if not, generate a recovery data block information table according to the recovery version number and the latest version backup summary file of the current data block; wherein, the recovery data block information table includes recovery data information, and the recovery data information includes: recovery version number, data block offset, file capacity, storage file name and compression level;

[0031] A saving module, configured to save the recovery data block information table, traverse the information of the next data block in the backup summary file, and return to execute the traversing step until the traversing is completed;

[0032] An obtaining module, configured to in response to a request for reading recovery data, obtain the recovery data storage path according to the recovery data block information table and the path of the backup virtual machine;

[0033] A reading module, configured to read recovery data according to the recovery data storage path, and send the read recovery data to a receiving platform according to the recovery data receiving path.

[0034] In an embodiment of the present invention, generating a recovery data block information table according to the recovery version number and the latest version backup summary file of the current data block includes:

[0035] Determine whether the recovery version number is equal to the version number of the full amount of information in the latest version backup summary file of the current data block;

[0036] If they are equal, copy the full amount of information in the latest version backup summary file of the current data block, and add the full amount data storage file name as the recovery data block information table;

[0037] If they are not equal, determine whether the recovery version number is equal to the version number of the modified information in the latest version backup summary file of the current data block;

[0038] If they are equal, when the recovery version number is equal to the version number of the last piece of modified information in the latest version backup summary file of the current data block, copy the full amount of information in the latest version backup summary file of the current data block, and add the full amount data storage file name as the recovery data block information table;

[0039] When the recovery version number is less than the version number of the last piece of modified information in the latest version backup summary file of the current data block, copy the next piece of modified information with the same version number as the recovery version number in the modified information of the latest version backup summary file of the current data block, and add the modified data storage file name of the next piece of modified information as the recovery data block information table;

[0040] If they are not equal, determine whether the recovery version number is less than the version number of the last piece of modified information in the latest version backup summary file of the current data block;

[0041] If it is less, copy the first piece of modified information with a version number greater than the recovery version number in the modified information of the latest version backup summary file of the current data block, and add the modified data storage file name of the first piece of modified information as the recovery data block information table;

[0042] If it is greater, copy the full amount of information in the latest version backup summary file of the current data block, and add the full amount data storage file name as the recovery data block information table.

[0043] In an embodiment of the present invention, obtaining a recovery data storage path according to the recovery data block information table and the path of the backup virtual machine includes:

[0044] Obtain the recovery data storage path according to the recovery version number of the recovery data block information table, the storage file name, and the path of the backup virtual machine.

[0045] In an embodiment of the present invention, it further includes: creating a cache file for storing the written data under the recovery data receiving path.

[0046] The third aspect of the embodiments of the present invention provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, it implements a backup data recovery method based on the FUSE system provided by the first aspect of the embodiments of the present invention.

[0047] The fourth aspect of the embodiments of the present invention provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements a backup data recovery method based on the FUSE system provided by the first aspect of the embodiments of the present invention.

[0048] The beneficial effects of the present invention:

[0049] The present invention generates a recovery data block information table according to the information in the backup summary file generated during data backup and the version number of the data to be recovered, integrates the information of the data to be recovered, constructs a complete recovery view, and when reading the recovery data, the storage path of the data to be recovered can be determined by querying the information in the recovery data block information table, so that the data to be recovered can be read. The present invention avoids the inefficient operation of reading and writing the entire disk data in the traditional recovery method, significantly improves the recovery efficiency, especially when processing large-scale data, and improves the processing efficiency.

[0050] At the same time, it ensures that the recovery operations of different users do not interfere with each other. Different users generate their own recovery data block information tables for recovering data from the same database, and the recovery operations of each other are not affected.

[0051] Other features and advantages of the present invention will be described in the following specification, and part of them will become obvious from the specification, or will be understood by implementing the present invention. The objectives and other advantages of the present invention can be realized and obtained through the structures specifically pointed out in the written specification, claims, and drawings.

[0052] The technical solutions of the present invention will be further described in detail below through the drawings and embodiments. Description of the Drawings

[0053] The drawings are used to provide a further understanding of the present invention, and constitute a part of the specification. They are used to explain the present invention together with the embodiments of the present invention, and do not constitute a limitation to the present invention. In the drawings:

[0054] Figure 1 It is a schematic flowchart of a backup data recovery method based on the FUSE system provided by an embodiment of the present invention;

[0055] Figure 2 It is a schematic block diagram of a backup data recovery device based on the FUSE system provided by an embodiment of the present invention. Specific embodiments

[0056] The following further describes the present invention in detail with reference to specific embodiments, but the embodiments of the present invention are not limited thereto.

[0057] As Figure 1 shown, a backup data recovery method based on the FUSE system is provided in the first aspect of the embodiment of the present invention, including the following steps:

[0058] Step 11, receiving a data recovery request.

[0059] Among them, the data recovery request indicates the path of the backup virtual machine, the recovery data reception path, and the recovery version number.

[0060] Step 12, when traversing the information of each data block in the backup summary file, obtaining the full amount of information corresponding to the recovery version number in the information of the current data block.

[0061] Step 13, determining whether the file capacity of the full amount of information corresponding to the recovery version number is 0.

[0062] Step 14, if so, generating a recovery data block information table in which the recovery data information except the recovery version number and the data block offset is empty according to the recovery version number.

[0063] Step 15, if not, generating a recovery data block information table according to the recovery version number and the latest version backup summary file of the current data block.

[0064] Among them, the recovery data block information table includes recovery data information, and the recovery data information includes: recovery version number, data block offset, file capacity, storage file name, and compression level.

[0065] Step 16, saving the recovery data block information table, traversing the information of the next data block in the backup summary file, and returning to execute Step 12 until the traversal is completed.

[0066] Step 17, in response to the read recovery data request, obtaining the recovery data storage path according to the recovery data block information table and the path of the backup virtual machine.

[0067] Step 18: Read the recovery data according to the recovery data storage path, and send the read recovery data to the receiving platform according to the recovery data receiving path.

[0068] In this embodiment, the recovery data block information table is generated according to the information in the backup summary file generated during data backup and the version number of the data to be recovered. The information of the data to be recovered is integrated to construct a complete recovery view. When it is necessary to read the recovery data, the storage path of the data to be recovered can be determined by querying the information in the recovery data block information table, so that the data to be recovered can be read. The present invention avoids the inefficient operation of reading and writing the entire disk data in the traditional recovery method, significantly improves the recovery efficiency, especially when dealing with large-scale data, and improves the processing efficiency.

[0069] At the same time, it is ensured that the recovery operations of different users do not interfere with each other. Different users generate their own recovery data block information tables for recovering data from the same database, and the recovery operations between them are not affected.

[0070] On the basis of the first aspect of the embodiments of the present invention, the second aspect of the embodiments of the present invention further describes the method of the present invention.

[0071] The second aspect of the embodiments of the present invention provides a method for recovering backup data based on the FUSE system, including the following steps:

[0072] Step 20: Perform multiple full and incremental data backups.

[0073] Before data backup, create a backup root directory / storage, and all files generated during the following backup are under this directory. When backing up, create a backup subdirectory and generate directories such as 0, 1, 2... under this directory, representing version 0, version 1, version 2, etc., to store the backup data and other related files of each backup respectively, so that the corresponding version path can be easily obtained by string concatenation during subsequent reading.

[0074] Backup: First, back up the data based on the virtual machine snapshot. Whether it is an incremental or full backup, it is performed by the method of dividing into blocks, and the compression is also based on block compression, which is convenient to obtain data blocks through the count value. For example, the offset of the 0th block is 0 * block_size, and the offset of the first block is 1 * block_size. When backing up, a dump file (full data file) and a digest disk backup summary file (recording information of all blocks) will be generated. If it is an incremental backup, a differential diff file (saving the block data before the change, that is, the data of the previous version) also needs to be generated, and all differential recovery data information should also be recorded in the digest file. The structure of the digest file is as follows:

[0075] Digest:

[0076] 0 -- dump

[0077] --<version, offset, size, compress>

[0078] -- diff

[0079] --size = 2

[0080] --<version1, offset, size, compress>

[0081] --<version2, offset, size, compress>

[0082] 1 -- dump

[0083] --<version, offset, size, compress>

[0084] -- diff

[0085] -- size = 1

[0086] --<version1, offset, size, compress>

[0087] 2 -- dump

[0088] --<version, offset, size, compress>

[0089] -- diff

[0090] -- size = 2

[0091] --<version1, offset, size, compress>

[0092] --<version2, offset, size, compress>

[0093] Among them, in <version1, offset, size, compress>, version represents the version number, offset represents the offset, size represents the file capacity, and compress represents the compression level. 0, 1, 2... respectively represent the block numbers. Each block has only one dump record, but there may be multiple diff file records with a length of size. For example, when size = 2, there are two diff file records. Each backup will update the digest file. If the block has not changed, only the block record needs to be copied and the version number of the dump information modified. If there are changes, the dump block information will be updated and a diff record will be added to indicate that this version of the block has changed compared to the previous version. The advantage of this approach is that regardless of whether the backup data is in raw format, qcow2 format, or any other data structure format, the specific location of any block can be located through the backup information, improving the recovery efficiency. This backup structure is applicable regardless of whether the last version is full data or the first version is full data. Moreover, when recovering, it is not necessary to read the digest files of each version, only the two digest files of the full version and the version to be recovered need to be read.

[0094] Full backup: Create a dump backup file and a digest backup summary file. Read the backup data through the virtualization platform and transfer it to the backup server via the network. Save all valid data blocks in chunks to the dump file, and keep the position of each block the same as that of the source disk. Record the dump block information in the digest file, that is, record a dump information with a structure of <version, offset, size, compress>, and initialize the size of the diff information to 0. For each subsequent incremental backup, the changed block data needs to be overwritten and written into the dump file, and each block information in the digest file needs to be updated, that is, keep the latest version of the dump backup file and the digest file as the latest data.

[0095] Incremental backup: Obtain the changed block information through the cbt changed block table of the virtual platform. Before writing the changed block into the dump file, first copy the data block at the position of this changed block from the dump file to the diff file, that is, copy the data corresponding to the changed block of the previous version to the diff file. Then overwrite the changed data block to the corresponding offset position in the dump file. Add a diff information record of this block to the digest file. For example, after a full backup, the record of this block is:

[0096] Digest:

[0097] 0 -- dump

[0098] --<version = 0, offset = 0, size = 2097152,compress = 0>

[0099] -- diff

[0100] -- size = 0

[0101] Then after the first incremental backup update:

[0102] Digest:

[0103] 0 -- dump

[0104] --<version = 1, offset = 0, size = 2097100, compress = 1>

[0105] -- diff

[0106] -- size = 1

[0107] --<version = 1, offset = 0, size = 2097152, compress = 0>

[0108] That is, add a diff information record <1, 0, 2097152, 0>, indicating that this block has changed in version 1. The data before the change (version 0 data) is saved in the diff file of version 1, with an offset of 0 and a size of 2097152 bytes in the diff file, and it is not compressed. And from the updated dump information, it can be seen that since this block of data has changed, it has been compressed into a data block with a size of 2097100 bytes and stored in the dump file, with a compression level of 1, that is, it needs to be decompressed when this block is restored. Note that the version numbers of each piece of data recorded in the diff are incremented, so that it can be conveniently searched during subsequent traversal. If this block is deleted, similarly record the diff information of this block, but modify the size in the dump information to size = 0, which means that this block is an empty block in the current version. If the block before the change is an empty block, the size in the diff information can be assigned 0. In short, the digest data of each increment is the merged digest data before, ensuring that the block information recorded in the digest file of the last version is the most complete. Subsequently, the specific location of each block of data in any version can be located by reading the block information of the latest digest file. The backed-up data is used for subsequent data recovery.

[0109] Step 21, receive a data recovery request.

[0110] Among them, the data recovery request indicates the path of the backup virtual machine, the data recovery receiving path, and the recovery version number.

[0111] The NFS (Network File System) network file system is adopted to optimize the file network transmission and the time-consuming problem of a large number of data I / O operations during the recovery process. First, create a directory fuse_nfs in the backup environment as the mounting directory of the fuse service. When accessing this directory, the I / O operations will be taken over by the fuse service. At the same time, an actual directory is needed to store the virtual machine disk and related configuration files, so create a directory named fuse_space. The lib-fuse library can export fuse_space to the fuse_nfs directory. When the fuse service is started, when we operate on the fuse_nfs directory, we are actually operating on the fuse_space directory. At the same time, add the fuse_nfs directory in the backup environment as the NFS storage directory of the virtual platform. The virtual platform can export fuse_nfs to the local NFS directory. Now, when the virtual platform operates on the local NFS directory, it is actually operating on the fuse_space directory of the backup server. Based on this NFS storage directory, create a new virtual machine for recovery. At this time, the virtual platform will directly create a virtual machine directory and the most important disk file, etc. in the fuse_space directory. Since the recovery is performed in the backup environment and the fuse needs to take over the I / O operations, all the directories operated during the subsequent recovery are the intermediate mounting directory fuse_nfs.

[0112] Directly restore files such as virtual machine configuration information to the fuse_nfs directory. Then the recovery program communicates with the fuse service. The recovery program transmits information such as the virtual disk path, the path of the backup virtual machine to be recovered, and the version number to be recovered to the fuse program. The fuse program registers this disk. After registration, regardless of whether the original disk is empty or not, the original disk content is invalid, and the disk data is now taken over by the fuse service. At the same time, create a cache file named cache in the recovery file to store the new data blocks after the disk is written with data, that is, after the recovery is completed, only read operations can be performed on the backup file, and when writing, the written data is written to the corresponding position in this file.

[0113] Step 22, after the Fuse program registers the disk, sequentially traverse the information of each data block in the backup summary file, and execute Step 23 - Step 26 for the information of each data block.

[0114] Step 23, obtain the full amount of information corresponding to the recovery version number from the information of the current data block.

[0115] In this step, in the backup summary file of each version, the version number in the full amount of information represents the version number of the data block, and the full amount of information with the same version number as the recovery version number is found in the information of the current data block in the backup summary file.

[0116] Step 24, determine whether the file capacity of the full amount of information corresponding to the recovery version number is 0.

[0117] If the file capacity size = 0 in the full amount of information, it means that the data of the current data block is empty under the version of the recovery version number. If size ≠ 0, it means that the data of the current data block is not empty under the version of the recovery version number.

[0118] Step 25, if so, generate a recovery data block information table with empty recovery data information except the recovery version number according to the recovery version number.

[0119] Among them, the recovery data block information table includes recovery data information, and the recovery data information includes: recovery version number version, data block offset offset, file capacity size, stored file name file, and compression level compress.

[0120] Specifically, when the data block under the recovery version number is empty, except for the version number and the data block offset in the full amount of information with the same version number as the recovery version number, other information is 0. Then copy this full amount of information and add the stored file name file, and file = 0, as the recovery data block information table. For example, in the backup summary file, dump--<version = 1, offset = 5, size = 0, compress = 0>, then the record of the current data block in the recovery data block information table is <version = 1, offset = 5, size = 0, file = 0, compress = 0>. The data block offset is the identifier of the data block.

[0121] Step 26, if not, generate a recovery data block information table according to the recovery version number and the latest version backup summary file of the current data block. When the file capacity of the full amount of information corresponding to the recovery version number is not 0, it means that the data block is not empty.

[0122] When the file capacity of the full amount of information corresponding to the recovery version number is not 0, the specific steps to generate a recovery data block information table according to the recovery version number and the latest version backup summary file of the current data block include steps 261 - 267:

[0123] Step 261, determine whether the recovery version number is equal to the version number of the full amount of information in the latest version backup summary file of the current data block.

[0124] In this step, either the latest version or an older version can be restored. Here, it is necessary to determine whether the version to be restored is the latest version or an older version.

[0125] Step 262, when the restore version number is equal to the version number of the full - volume information in the latest version backup summary file of the current data block, copy the full - volume information in the latest version backup summary file of the current data block, and add the storage file name file to form the restored data block information table, that is, dbt (Data Block Table). Among them, the value of the storage file name is the full - volume data storage file name, that is, file = dump.

[0126] In this step, when the restore version number is equal to the version number of the full - volume information in the latest version backup summary file of the current data block, it means that the data to be restored is the latest version. Therefore, only need to copy the full - volume information in the latest version backup summary file of the current data block and add the storage file name to form the restored data block information table, and then execute step 27.

[0127] For example, if the record in the latest version backup summary file of a data block read is dump--<version =0, offset = 10, size = 2097100,compress = 1>, then the restored data information record in the generated dbt is <version = 0, offset =10, size = 2097100, file = dump, compress = 1>.

[0128] It should be noted that the copy - on - write is not supported by the file system in this step. If the current file system supports copy - on - write and the backup is not compressed, the dump file can be quickly copied to the disk file directly through commands such as cp --reflink=always. This is almost instantaneous, and there is no need to register this disk in fuse. When the virtual platform accesses this disk, it is directly available and does not go through the block mapping of fuse.

[0129] Step 263, when the restore version number is not equal to the version number of the full - volume information in the latest version backup summary file of the current data block, determine whether the restore version number is equal to the version number of the modified information in the latest version backup summary file of the current data block.

[0130] In this step, when the restored version number is not equal to the version number of the full amount of information in the latest version backup summary file of the current data block, the version to be restored is the old version, and the data block may or may not have been modified in the old version. If it has been modified, it will be recorded in the modification information diff during incremental backup (the diff contains the corresponding version number information). If it has not been modified, it will not be recorded in the diff.

[0131] Step 264, when the restored version number is equal to the version number of the modification information in the latest version backup summary file of the current data block, it is divided into two cases:

[0132] First, when the restored version number is equal to the version number of the last piece of modification information in the latest version backup summary file of the current data block, copy the full amount of information in the latest version backup summary file of the current data block, and add the storage file name to form the restored data block information table.

[0133] Here, one of the version numbers in the version number of the modification information is the same as the restored version number, which means that the current data block has been modified in the version to be restored. At the same time, if the restored version number is the same as the version number of the last piece of information in the modification information, it means that this is the last modification of the data block in this version number (no modification in subsequent versions). Therefore, the information of the current data block is recorded in the full amount of information. At this time, since the version number in the full amount of information is the latest version number and is inconsistent with the restored version number, it is necessary to copy the information except the version number, then add the restored version number and the storage file name, write it into the restored data block information table, and then execute step 27. The value of the storage file name is the full amount data storage file name.

[0134] Example, Digest:

[0135] 10 -- dump

[0136] --<version = 5, offset = 10, size = 2097100, compress = 1>

[0137] -- diff

[0138] -- size = 3

[0139] --<version = 2, offset = 10, size = 2097222, compress = 0>

[0140] --<version = 3, offset = 10, size = 2097111, compress = 0>

[0141] --<version = 4, offset = 10, size = 2097333, compress = 0>

[0142] There are a total of six versions from 0 to 5. The restored version number is 4. There are three records with version numbers 2, 3, and 4 in the diff. If the last modification of the current data block is version 4, it means that the data block was not modified in version 5. The latest version backup summary file records dump--<version = 5, offset = 10, size = 2097100, compress = 1>. Then, the data of version 4 of this block is stored in the dump file, and the restored data information record in the generated dbt is <version = 5, offset = 10, size = 2097100, file = dump, compress = 1>.

[0143] Second, when the restored version number is less than the version number of the last modification information in the latest version backup summary file of the current data block, copy the next modification information with the same version number as the restored version number in the modification information of the latest version backup summary file of the current data block, and add the storage file name to form the restored data block information table.

[0144] Here, when the restored version number is less than the version number of the last modification information in the latest version backup summary file of the current data block, it means that the current data block has been modified in the version numbers after the restored version number. At this time, the data corresponding to the restored version number of this data block is stored in the diff file of the next piece of information with the same version number as the restored version number in the modification information. The next piece of information here refers to the modification information of the first modified data after the restored version number. Copy the next modification information with the same version number as the restored version number in the modification information of the latest version backup summary file of the current data block, add the storage file name, write it into the restored data block information table, and then execute step 27. The value of the storage file name is the storage file name of the modified data of this next piece of modification information, that is, file = diff.

[0145] For example, there are a total of six versions numbered from 0 to 5, the recovery version number is 1, and there are three records with version numbers 1, 3, and 4 in the diff, indicating that the data block has been modified twice after restoring version 1. The data of version number 1 is stored in the diff file of version number 3, and the data of version 0 is stored in the diff of version number 1. Then the recovery data information record in the generated dbt is <version = 3, offset = 10, size = 2097100, file = diff, compress = 1>. Version = 3 and file = diff also indicate that the data is stored in the diff file of version 3.

[0146] Step 265, when the recovery version number is not equal to the version number of the modification information in the latest version backup summary file of the current data block, determine whether the recovery version number is less than the version number of the last modification information in the modification information of the latest version backup summary file of the current data block.

[0147] In this step, when there is no version number in the modification information that is the same as the recovery version number, it means that the data block has not been modified in the version of the recovery version number. At this time, it is necessary to judge the size of the recovery version number and the version number of the last modification information in the modification information, so as to find the data of the version to be restored.

[0148] Step 266, when the recovery version number is less than the version number of the last modification information in the modification information of the latest version backup summary file of the current data block, copy the first modification information whose version number in the modification information of the latest version backup summary file of the current data block is greater than the recovery version number, and add the name of the modification data storage file of this first modification information as the recovery data block information table.

[0149] In this step, if the recovery version number is less than the version number of the last one in the modification information, it is necessary to determine that the data recorded in the first modification information after the recovery version number is also the data of the recovery version number. Copy the modification information that is greater than the recovery version number and is the first modification after the recovery version number, and add the storage file name, write it into the recovery data block information table, and then execute step 27. The value of the storage file name is the name of the modification data storage file of the first modification information after the recovery version number.

[0150] For example, there are a total of six versions from 0 to 5, the recovery version number is 3, and there are three records with version numbers 1, 2, and 5 in the diff, indicating that the data block was modified once after restoring version 3, not modified in versions 3 and 4, and modified in version 5. Since the data in version numbers 3 and 4 is the same and stored in the diff file of version number 5, the recovery data information record in the generated dbt is <version = 5, offset = 10, size = 2097100, file = diff, compress = 1>. Version = 5 and file = diff also mean that the data is stored in the diff file of version 5.

[0151] Step 267, when the recovery version number is greater than the version number of the last modification information in the modification information of the latest version backup summary file of the current data block, copy the full amount of information in the latest version backup summary file of the current data block and add the storage file name to form the recovery data block information table.

[0152] In this step, when the recovery version number is greater than the version number of the last one in the modification information, it means that the data block was not modified in the version of the recovery version number and subsequent versions. Therefore, the data record of the recovery version number is in the full amount of information. At this time, copy the full amount of information, add the storage file name, write it into the recovery data block information table, and then execute step 27. The value of the storage file name is the full amount data storage file name.

[0153] For example, there are a total of six versions from 0 to 5, the recovery version number is 4, and there are three records with version numbers 1, 2, and 3 in the diff, indicating that there is no modification in recovery version 4 and version 5. The data in version 4 and 5 is the same and stored in the dump. Then the recovery data information record in the generated dbt is <version = 5, offset = 10, size = 2097100, file = dump, compress = 1>. Version = 5 and file = dump also mean that the data is stored in the dump file of version 5.

[0154] Step 27, save the recovery data block information table, start traversing the information of the next data block in the backup summary file, and execute steps 23 - 26 until the traversal is completed.

[0155] After the traversal is completed, the information of the data blocks of the version to be restored is recorded in the dbt table. At this time, the data recovery is completed, and only the data indicated in the dbt table needs to be read subsequently. Since the fuse library implements a user-space file system and provides various io interfaces, such as read and write system calls and other interfaces, after implementing these interfaces, it can interact with the kernel. When making read and write system calls to the fuse-mounted directory, these system calls will be intercepted by the fuse library, and then these interfaces will be called to perform read or write operations, thus avoiding traversing and reading / writing the entire disk data during recovery and achieving fast backup data recovery.

[0156] Recovery is completed through the above steps. Since only the integration of recovery data information is performed in this process and a large amount of data is not recovered, and there is no operation of decompressing data during the recovery process, the speed is at the second level. At this time, the restored virtual machine starts, and the virtual platform will read the disk. It will call system functions such as wire read to access the data and pass in a parameter structure such as <offset, size> to access the data. Since the disk file was registered before, when the fuse matches that the accessed file is a disk file, it takes over the read and write io operations of the disk file. When reading a file, the corresponding data is found through the dbt and returned. The data read by the virtual platform is the real disk data, that is, the backup data. When writing a file, using the idea of copy-on-write, the data will not be copied to the cache file when reading the file, and only when writing data will the modified data be rewritten to the cache file.

[0157] Through the above steps, a copy-on-write disk file can be restored. Since the original backup data file is not modified, it supports quickly restoring different versions of the disk file of the same backup data at the same time. And the recovery time for each recovery task is at the second level.

[0158] Step 28, in response to a request to read recovery data, obtain the recovery data storage path according to the recovery data block information table and the path of the backup virtual machine.

[0159] In this step, after generating the dbt table, the user can read the data in the dbt table as needed, that is, can read some or all of the data under the recovery version number, so as to quickly obtain the required recovery data. When the user has a request to read recovery data, according to the data block to be read in the request to read recovery data, find the information of the data block in the dbt table. According to the recovery version number and storage file name of the data block, as well as the path of the backup virtual machine, splice the recovery data storage path, read from the path and return it to the user. The path of the backup virtual machine includes the backup root directory and the next-level directory of the root directory.

[0160] For example, / storage / back1 is the backup root directory and the next-level directory. If the record found in the dbt table is <version = 5, offset = 10, size = 2097100, file = diff, compress = 1>, the restored data storage path obtained by concatenation is / storage / back1 / 5 / diff, which means the data to be retrieved is in the diff file of version 5. For the file capacity size of the data block, if size = 0, it indicates that the data block is currently an empty block under this restored version and there is no need to query the subsequent file file. If size > 0, the corresponding file needs to be queried and the data is restored by integrating multiple corresponding files.

[0161] Here, the cache cache file is used to cache the data written by the user after restoration, that is, after generating the dbt table.

[0162] Step 29: Read the restored data according to the restored data storage path and send the read restored data to the receiving platform according to the restored data receiving path.

[0163] In this step, the restored data receiving path is also the path of the restored virtual machine, and the receiving platform is also the restored virtual machine.

[0164] Both the read and write interfaces will have offset and size parameters, and the starting data block number can be located through the offset. For example, the block size is 2, and the user performs a read operation. In the parameters, the offset = 0, and the starting data block is the 0th block. Then, by offset + size, the ending block number can be located. For example, size = 4, which is the size of two blocks, and the ending block is the 1st block. So, it is known that the 0th block and the 1st block need to be read. At this time, by looking up the records of the 0th block and the 1st block in the dbt table, and then normally calling the kernel read function, the storage path is obtained and the block data of the two blocks is read and merged and returned to the user.

[0165] In a feasible implementation manner, after generating the dbt table, the user operates the restored virtual machine and can write data into the restored data block. When performing a write operation, the data will be stored in the cache file cache. After writing the data, the data in the data block is updated to the written data. The written data is stored in the cache file cache, while the backup data will not be modified and will not affect the backup data at the backup end. Specifically, after the data is written, the parameters of the write function are <offset, size>. According to these parameters, the data block and the position of the data block where the data is written can be determined. Then, the corresponding record in the dbt table is modified, that is, the value of file in the record of the corresponding data block in the dbt table is modified to cache, and the compression level is modified to 0. At this time, it is meaningless to update the version number of the data, and the version number is modified to 0.

[0166] For example, taking the block size of 2 as an example, the parameters of the write function are <offset = 1, size = 5>. First, the starting data block to be modified is read as the 0th block through the offset. Through offset + size, it can be known that the last data block is the 2nd block. Then, the 0th, 1st, and 2nd data blocks need to be read out, and the second half of the 0th data block is overwritten with the data to be written. The 1st and 2nd data blocks are all overwritten. Finally, the 0th, 1st, and 2nd blocks are written into the cache file, and the dbt file is updated. The offset of the 0th, 1st, and 2nd blocks is updated to the offset in the cache file, file is modified to cache, and the version number and compression level are 0. When reading these blocks again, the new data can be directly read from the cache file. For example, the modified record of the 0th block is <0, 0, 2097152, cache, 0>. At this time, the data block has been recorded in the cache file, so the value of version is meaningless, and for the compression level, its value must be 0.

[0167] As Figure 2 shown, the third aspect of the embodiment of the present invention provides a backup data recovery device based on the FUSE system, including:

[0168] A receiving module 31, configured to receive a data recovery request; the data recovery request indicates the path of the backup virtual machine, the path for receiving the recovered data, and the recovery version number;

[0169] A traversing module 32, configured to execute a traversing step: when traversing each data block, obtain the full amount of information corresponding to the recovery version number in the backup summary file of the current data block;

[0170] A judging module 33, configured to judge whether the file capacity of the full amount of information corresponding to the recovery version number is 0;

[0171] The first generation module 34, if so, is used to generate a recovery data block information table in which the recovery data information except the recovery version number and the data block offset is empty according to the recovery version number;

[0172] The second generation module 35, if not, is used to generate a recovery data block information table according to the recovery version number and the latest version backup summary file of the current data block; wherein, the recovery data block information table includes recovery data information, and the recovery data information includes: recovery version number, data block offset, file capacity, stored file name, and compression level;

[0173] The saving module 36 is used to save the recovery data block information table, traverse the next data block, and return to execute the traversal step until the traversal is completed;

[0174] The obtaining module 37 is used to respond to a recovery data reading request and obtain a recovery data storage path according to the recovery data block information table and the path of the backup virtual machine;

[0175] The reading module 38 is used to read the recovery data according to the recovery data storage path and send the read recovery data to the receiving platform according to the recovery data receiving path.

[0176] In an embodiment of the present invention, generating a recovery data block information table according to the recovery version number and the latest version backup summary file of the current data block includes:

[0177] Judging whether the recovery version number is equal to the version number of the full amount information in the latest version backup summary file of the current data block;

[0178] If they are equal, copy the full amount information in the latest version backup summary file of the current data block and add the full amount data storage file name as the recovery data block information table;

[0179] If they are not equal, judge whether the recovery version number is equal to the version number of the modified information in the latest version backup summary file of the current data block;

[0180] If they are equal, when the recovery version number is equal to the version number of the last piece of modified information in the latest version backup summary file of the current data block, copy the full amount information in the latest version backup summary file of the current data block and add the full amount data storage file name as the recovery data block information table;

[0181] When the recovery version number is less than the version number of the last piece of modified information in the latest version backup summary file of the current data block, copy the next piece of modified information with the same version number as the modified information of the recovery version number in the latest version backup summary file of the current data block, and add the modified data storage file name of the next piece of modified information as the recovery data block information table;

[0182] If not equal, determine whether the restored version number is less than the version number of the last modification information in the modification information of the latest version backup summary file of the current data block;

[0183] If less, copy the first modification information in the modification information of the latest version backup summary file of the current data block whose version number is greater than the restored version number, and add the storage file name of the modification data of the first modification information as the restored data block information table;

[0184] If greater, copy all the information in the latest version backup summary file of the current data block, and add the all - data storage file name as the restored data block information table.

[0185] In an embodiment of the present invention, obtaining the restored data storage path according to the restored data block information table and the path of the backup virtual machine includes:

[0186] Obtain the restored data storage path according to the restored version number, storage file name of the restored data block information table and the path of the backup virtual machine.

[0187] In an embodiment of the present invention, it further includes: creating a cache file for storing the written data under the restored data receiving path.

[0188] A fourth aspect of the embodiments of the present invention provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, it implements a backup data restoration method based on the FUSE system provided in the above - mentioned embodiments of the present invention.

[0189] A fifth aspect of the embodiments of the present invention further provides a computer - readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the steps of a backup data restoration method based on the FUSE system provided in the above - mentioned embodiments of the present invention.

[0190] Among them, the memory may include a random access memory (RAM), and may also include a non - volatile memory (NVM), such as at least one disk memory. Optionally, the memory may also be at least one storage device located far from the aforementioned processor.

[0191] The above-mentioned processor may be a general-purpose processor, including a Central Processing Unit (CPU), a Network Processor (NP), etc.; it may also be a Digital Signal Processor (DSP), an Application Specific Integrated Circuit (ASIC), a Field-Programmable Gate Array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware systems.

[0192] The method provided by the embodiments of the present invention can be applied to an electronic device. Specifically, the electronic device may be: a desktop computer, a portable computer, a smart mobile terminal, a server, etc. This is not limited herein, and any electronic device that can implement the present invention belongs to the protection scope of the present invention.

[0193] For the device / electronic device embodiments, since they are basically similar to the method embodiments, the description is relatively simple, and for the relevant parts, refer to the partial description of the method embodiments.

[0194] The present invention is described with reference to the flowcharts and / or block diagrams of methods, devices (apparatus), and computer program products according to the embodiments of the present invention. It should be understood that each flow and / or block in the flowchart and / or block diagram, and the combination of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate a device for implementing the functions specified in Figure 1 one or more flows and / or blocks Figure 1 one or more blocks.

[0195] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured article including an instruction device, and the instruction device implements the functions specified in Figure 1 one or more flows and / or blocks Figure 1 one or more blocks.

[0196] These computer program instructions can also be loaded onto a computer or other programmable data processing apparatus, so that a series of operation steps are performed on the computer or other programmable apparatus to generate a computer-implemented process, thereby providing instructions for implementing the steps of the process Figure 1 in one process or a plurality of processes and / or boxes Figure 1 or steps for implementing the functions specified in one box or a plurality of boxes.

[0197] Obviously, those skilled in the art can make various changes and modifications to the present invention without departing from the spirit and scope of the present invention. Thus, if these modifications and variations of the present invention fall within the scope of the claims of the present invention and their equivalent technologies, the present invention is also intended to include these modifications and variations.

Claims

1. A backup data recovery method based on the FUSE system, characterized in that The method includes the following steps: Receiving a data recovery request, where the data recovery request indicates the path of the backup virtual machine, the path for receiving recovered data, and the recovery version number; Traversal step: When traversing the information of each data block in the backup summary file, obtaining the full amount of information corresponding to the recovery version number from the information of the current data block; Determining whether the file capacity of the full amount of information corresponding to the recovery version number is 0; If so, generating a recovery data block information table in which the recovery data information except the recovery version number and the data block offset is empty according to the recovery version number; If not, generating a recovery data block information table according to the recovery version number and the latest version backup summary file of the current data block; where the recovery data block information table includes recovery data information, and the recovery data information includes: recovery version number, data block offset, file capacity, stored file name, and compression level; Saving the recovery data block information table, traversing the information of the next data block in the backup summary file, and returning to execute the traversal step until the traversal is completed; In response to a request for reading recovered data, obtaining a recovery data storage path according to the recovery data block information table and the path of the backup virtual machine; Reading the recovered data according to the recovery data storage path, and sending the read recovered data to the receiving platform according to the recovery data receiving path; The generating of the recovery data block information table according to the recovery version number and the latest version backup summary file of the current data block includes: Determining whether the recovery version number is equal to the version number of the full amount of information in the latest version backup summary file of the current data block; If equal, copying the full amount of information in the latest version backup summary file of the current data block and adding the full amount data storage file name as the recovery data block information table; If not equal, determining whether the recovery version number is equal to the version number of the modified information in the latest version backup summary file of the current data block; If equal, when the recovery version number is equal to the version number of the last piece of modified information in the latest version backup summary file of the current data block, copying the full amount of information in the latest version backup summary file of the current data block and adding the full amount data storage file name as the recovery data block information table; When the recovery version number is less than the version number of the last piece of modified information in the latest version backup summary file of the current data block, copying the next piece of modified information with the same version number as the recovery version number in the modified information of the latest version backup summary file of the current data block and adding the modified data storage file name of the next piece of modified information as the recovery data block information table; If not equal, determining whether the recovery version number is less than the version number of the last piece of modified information in the latest version backup summary file of the current data block; If less, copying the first piece of modified information with a version number greater than the recovery version number in the modified information of the latest version backup summary file of the current data block and adding the modified data storage file name of the first piece of modified information as the recovery data block information table; If it is greater than, copy the full - volume information in the latest version backup summary file of the current data block, and add the full - volume data storage file name as the restored data block information table.

2. The method according to claim 1, wherein The obtaining the restored data storage path according to the restored data block information table and the path of the backup virtual machine includes: Obtaining the restored data storage path according to the restoration version number, storage file name in the restored data block information table, and the path of the backup virtual machine.

3. The method according to claim 1, characterized in that The method further includes: creating a cache file for storing the written data under the restored data receiving path.

4. A backup data recovery device based on the FUSE system, characterized in that, It includes: A receiving module, configured to receive a data restoration request; wherein, the data restoration request indicates the path of the backup virtual machine, the restored data receiving path, and the restoration version number; A traversing module, configured to perform a traversing step: when traversing the information of each data block in the backup summary file, obtain the full - volume information corresponding to the restoration version number in the information of the current data block; A judging module, configured to judge whether the file capacity of the full - volume information corresponding to the restoration version number is 0; A first generating module, configured to if so, generate a restored data block information table in which the restored data information except the restoration version number and the data block offset is empty according to the restoration version number; A second generating module, configured to if not, generate a restored data block information table according to the restoration version number and the latest version backup summary file of the current data block; wherein, the restored data block information table includes restored data information, and the restored data information includes: restoration version number, data block offset, file capacity, storage file name, and compression level; A saving module, configured to save the restored data block information table, traverse the information of the next data block in the backup summary file, and return to execute the traversing step until the traversing is completed; An obtaining module, configured to in response to a request for reading restored data, obtain the restored data storage path according to the restored data block information table and the path of the backup virtual machine; A reading module, configured to read the restored data according to the restored data storage path and send the read restored data to the receiving platform according to the restored data receiving path; The generating the restored data block information table according to the restoration version number and the latest version backup summary file of the current data block includes: Judging whether the restoration version number is equal to the version number of the full - volume information in the latest version backup summary file of the current data block; If they are equal, copy the full - volume information in the latest version backup summary file of the current data block, and add the full - volume data storage file name as the restored data block information table; If they are not equal, judge whether the restoration version number is equal to the version number of the modification information in the latest version backup summary file of the current data block; If they are equal, then when the restoration version number is equal to the version number of the last piece of modification information in the latest version backup summary file of the current data block, copy the full - volume information in the latest version backup summary file of the current data block, and add the full - volume data storage file name as the restored data block information table; When the restored version number is less than the version number of the last piece of modification information in the latest version backup summary file of the current data block, copy the next piece of modification information with the same version number as the modification information of the restored version number in the latest version backup summary file of the current data block, and add the name of the modified data storage file of the next piece of modification information as the restored data block information table; If not equal, determine whether the restored version number is less than the version number of the last piece of modification information in the latest version backup summary file of the current data block; If less, copy the first piece of modification information in the modification information of the latest version backup summary file of the current data block whose version number is greater than the restored version number, and add the name of the modified data storage file of the first piece of modification information as the restored data block information table; If greater, copy all the information in the latest version backup summary file of the current data block, and add the name of the full amount data storage file as the restored data block information table.

5. The device according to claim 4, characterized in that, The obtaining of the restored data storage path according to the restored data block information table and the path of the backup virtual machine includes: Obtain the restored data storage path according to the restored version number, storage file name of the restored data block information table and the path of the backup virtual machine.

6. The device according to claim 4, characterized in that, It also includes: Create a cache file for storing the written data under the restored data receiving path.

7. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the backup data restoration method based on the FUSE system according to any one of claims 1 to 3.

8. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the backup data restoration method based on the FUSE system according to any one of claims 1 to 3.

Citation Information

Patent Citations

  • International freight rate data storage method and system

    CN110928839A