Backup data recovery method and device based on FUSE system
Through the backup data recovery method based on the FUSE system, the recovery data block information table is generated, which solves the problem of low backup and recovery efficiency in the prior art, and realizes efficient data recovery and operation isolation.
Patent Information
- Application Number
- CN202510466763.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-15
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2045-04-15
AI Technical Summary
The prior art has low backup and recovery efficiency when managing server data, especially when processing large-scale data, traditional methods require traversing the entire disk, resulting in low recovery efficiency.
The backup data recovery method based on the FUSE system is adopted, and the information of the required recovery data is integrated by generating the recovery data block information table to build a complete recovery view, avoiding the inefficient operation of reading and writing complete disk data in traditional methods.
It significantly improves data recovery efficiency, especially when processing large-scale data, reduces recovery time and ensures that the recovery operations of different users do not interfere with each other.
Smart Images

Figure CN119988103A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of server data management, and in particular to a backup data recovery method and device based on a FUSE system. Background Art
[0002] When the prior art performs a full disk backup for a block device, a dump file is generated by traversing the entire disk. Dump file: A file composed of data blocks and empty blocks. During incremental backup, a corresponding differential file is generated based on the difference in data changes. The differential file is used to save the original data overwritten by the new incremental data. Differential files and summary files of the full file are generated for each full backup and incremental backup. One incremental backup corresponds to a set of summary files and differential files. At the same time, the bitmap file needs to be updated during incremental backup. When restoring data, it is necessary to restore it based on the dump file, summary file, differential file and bitmap file. During restoration, only the entire disk data can be restored, and the entire disk needs to be operated, resulting in low recovery efficiency. Summary of the invention
[0003] In order to solve the above problems existing in the prior art, the present invention provides a backup data recovery method and device based on the FUSE system. The technical problem to be solved by the present invention is achieved through the following technical solutions: A first aspect of an embodiment of the present invention provides a backup data recovery method based on a FUSE system, comprising the following steps: Receive a data recovery request; wherein the data recovery request indicates a path of the backup virtual machine, a recovery data receiving path, and a recovery version number; Traversal step: when traversing the information of each data block in the backup summary file, obtain the full information corresponding to the recovery version number in the information of the current data block; Determine whether the file capacity of the full amount of information corresponding to the restored version number is 0; If yes, generating a recovery data block information table according to the recovery version number, in which the recovery data information other than the recovery version number and the data block offset is empty; If not, generating a recovery data block information table according to the recovery version number and the latest version backup summary file of the current data block; wherein the recovery data block information table includes recovery data information, and the recovery data information includes: recovery version number, data block offset, file capacity, storage file name and compression level; Saving the restored data block information table, traversing the information of the next data block in the backup summary file, and returning to execute the traversal step until the traversal is completed; In response to a request to read recovery data, acquiring a recovery data storage path according to the recovery data block information table and the path of the backup virtual machine; The recovery data is read according to the recovery data storage path, and the read recovery data is sent to the receiving platform according to the recovery data receiving path.
[0004] In one embodiment of the present invention, generating a recovery data block information table according to the recovery version number and the latest version backup summary file of the current data block includes: Determine whether the recovery version number is equal to the version number of the full information in the latest version backup summary file of the current data block; If they are equal, copy the full information in the latest version backup summary file of the current data block and add the full data storage file name as the recovery data block information table; If not, determining whether the recovery version number is equal to the version number of the modification information in the latest version backup summary file of the current data block; If they are equal, when the recovery version number is equal to the version number of the last modification information in the latest version backup summary file of the current data block, copy the full information in the latest version backup summary file of the current data block, and add the full data storage file name as the recovery data block information table; When the recovery version number is smaller than the version number of the last modification information in the latest version backup summary file of the current data block, copy the next modification information of the modification information with the same version number as the recovery version number in the modification information in the latest version backup summary file of the current data block, and add the modification data storage file name of the next modification information as the recovery data block information table; If not equal, determine whether the recovery version number is less than the version number of the last modification information in the latest version backup summary file of the current data block; If it is smaller, copy the first modification information whose version number in the modification information in the latest version backup summary file of the current data block is greater than the recovery version number, and add the modification data storage file name of the first modification information as the recovery data block information table; If it is greater, the full information in the latest version backup summary file of the current data block is copied, and the full data storage file name is added as the recovery data block information table.
[0005] In one embodiment of the present invention, the acquiring the recovery data storage path according to the recovery data block information table and the path of the backup virtual machine includes: The recovery data storage path is acquired according to the recovery version number of the recovery data block information table, the storage file name and the path of the backup virtual machine.
[0006] In one embodiment of the present invention, the method further includes: creating a cache file for storing write data under the recovery data receiving path.
[0007] A second aspect of an embodiment of the present invention provides a backup data recovery device based on a FUSE system, comprising: A receiving module, configured to receive a data recovery request; wherein the data recovery request indicates a path of a backup virtual machine, a recovery data receiving path, and a recovery version number; A traversal module is used to perform a traversal step: when traversing the information of each data block in the backup summary file, obtaining the full amount of information corresponding to the recovery version number in the information of the current data block; A judgment module, used to judge whether the file capacity of the full information corresponding to the restored version number is 0; A first generating module, configured to generate, according to the recovery version number, a recovery data block information table in which the recovery data information other than the recovery version number and the data block offset is empty; A second generating module is used to generate a recovery data block information table according to the recovery version number and the latest version backup summary file of the current data block if no; wherein the recovery data block information table includes recovery data information, and the recovery data information includes: recovery version number, data block offset, file capacity, storage file name and compression level; A saving module, used for saving the restored data block information table, traversing the information of the next data block in the backup summary file, and returning to execute the traversal step until the traversal is completed; An acquisition module, configured to respond to a request to read recovery data and acquire a recovery data storage path according to the recovery data block information table and the path of the backup virtual machine; The reading module is used to read the recovery data according to the recovery data storage path, and send the read recovery data to the receiving platform according to the recovery data receiving path.
[0008] In one embodiment of the present invention, generating a recovery data block information table according to the recovery version number and the latest version backup summary file of the current data block includes: Determine whether the recovery version number is equal to the version number of the full information in the latest version backup summary file of the current data block; If they are equal, copy the full information in the latest version backup summary file of the current data block and add the full data storage file name as the recovery data block information table; If not, determining whether the recovery version number is equal to the version number of the modification information in the latest version backup summary file of the current data block; If they are equal, when the recovery version number is equal to the version number of the last modification information in the latest version backup summary file of the current data block, copy the full information in the latest version backup summary file of the current data block, and add the full data storage file name as the recovery data block information table; When the recovery version number is smaller than the version number of the last modification information in the latest version backup summary file of the current data block, copy the next modification information of the modification information with the same version number as the recovery version number in the modification information in the latest version backup summary file of the current data block, and add the modification data storage file name of the next modification information as the recovery data block information table; If not equal, determine whether the recovery version number is less than the version number of the last modification information in the latest version backup summary file of the current data block; If it is smaller, copy the first modification information whose version number in the modification information in the latest version backup summary file of the current data block is greater than the recovery version number, and add the modification data storage file name of the first modification information as the recovery data block information table; If it is greater, the full information in the latest version backup summary file of the current data block is copied, and the full data storage file name is added as the recovery data block information table.
[0009] In one embodiment of the present invention, the acquiring the recovery data storage path according to the recovery data block information table and the path of the backup virtual machine includes: The recovery data storage path is acquired according to the recovery version number of the recovery data block information table, the storage file name and the path of the backup virtual machine.
[0010] In one embodiment of the present invention, it also includes: creating a cache file for storing write data under the recovery data receiving path.
[0011] A third aspect of an embodiment of the present invention provides an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, a backup data recovery method based on a FUSE system provided by the first aspect of an embodiment of the present invention is implemented.
[0012] A fourth aspect of an embodiment of the present invention provides a computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, the method for restoring backup data based on a FUSE system provided by the first aspect of an embodiment of the present invention is implemented.
[0013] Beneficial effects of the present invention: The present invention generates a recovery data block information table based on the information in the backup summary file generated during data backup and the version number of the data to be restored, integrates the information of the required recovery data, and constructs a complete recovery view. When it is necessary to read the recovery data, the storage path of the data to be restored can be determined by querying the information in the recovery data block information table, so that the data to be restored can be read. The present invention avoids the inefficient operation of reading and writing complete disk data in the traditional recovery method, significantly improves the recovery efficiency, especially when processing large-scale data, improves the processing efficiency.
[0014] At the same time, it ensures that the recovery operations of different users do not interfere with each other. Different users generate their own recovery data block information tables for the same database recovery data, and the recovery operations between them are not affected.
[0015] Other features and advantages of the present invention will be described in the following description, and partly become apparent from the description, or understood by practicing the present invention. The purpose and other advantages of the present invention can be realized and obtained by the structures particularly pointed out in the written description, claims, and drawings.
[0016] The technical solution of the present invention is further described in detail below through the accompanying drawings and embodiments. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] The accompanying drawings are used to provide a further understanding of the present invention and constitute a part of the specification. Together with the embodiments of the present invention, they are used to explain the present invention and do not constitute a limitation of the present invention. In the accompanying drawings: Figure 1 A flowchart of a backup data recovery method based on a FUSE system provided by an embodiment of the present invention; Figure 2 A block diagram of a FUSE-based backup data recovery device provided in an embodiment of the present invention. DETAILED DESCRIPTION
[0018] The present invention is further described in detail below with reference to specific embodiments, but the embodiments of the present invention are not limited thereto.
[0019] like Figure 1 As shown, a first aspect of an embodiment of the present invention provides a backup data recovery method based on a FUSE system, comprising the following steps: Step 11: Receive a data recovery request.
[0020] The data recovery request indicates the path of the backup virtual machine, the recovery data receiving path and the recovery version number.
[0021] Step 12, when traversing the information of each data block in the backup summary file, obtain the full information corresponding to the recovery version number in the information of the current data block.
[0022] Step 13, determine whether the file capacity of the full information corresponding to the restored version number is 0.
[0023] Step 14: If yes, generate a recovery data block information table according to the recovery version number, in which the recovery data information except the recovery version number and the data block offset is empty.
[0024] Step 15: If not, generate a recovery data block information table according to the recovery version number and the latest version backup summary file of the current data block.
[0025] The restored data block information table includes restored data information, and the restored data information includes: a restored version number, a data block offset, a file capacity, a storage file name, and a compression level.
[0026] Step 16, save the restored data block information table, traverse the information of the next data block in the backup summary file, and return to execute step 12 until the traversal is completed.
[0027] Step 17: In response to the request to read the recovery data, a recovery data storage path is obtained according to the recovery data block information table and the path of the backup virtual machine.
[0028] Step 18, reading the recovery data according to the recovery data storage path, and sending the read recovery data to the receiving platform according to the recovery data receiving path.
[0029] In this embodiment, a recovery data block information table is generated based on the information in the backup summary file generated during data backup and the version number of the data to be restored, and the information of the required recovery data is integrated to construct a complete recovery view. When it is necessary to read the recovery data, the storage path of the data to be restored can be determined by querying the information in the recovery data block information table, so that the data to be restored can be read. The present invention avoids the inefficient operation of reading and writing complete disk data in the traditional recovery method, significantly improves the recovery efficiency, especially when processing large-scale data, improves the processing efficiency.
[0030] At the same time, it ensures that the recovery operations of different users do not interfere with each other. Different users generate their own recovery data block information tables for the same database recovery data, and the recovery operations between them are not affected.
[0031] Based on the first aspect of the embodiment of the present invention, the second aspect of the embodiment of the present invention further illustrates the method of the present invention.
[0032] A second aspect of an embodiment of the present invention provides a backup data recovery method based on a FUSE system, comprising the following steps: Step 20: Perform multiple full and incremental data backups.
[0033] Before backing up data, create a backup root directory / storage. All files generated during the following backup are in this directory. During the backup, create a backup subdirectory and generate directories such as 0, 1, 2... in this directory, representing version 0, version 1, and version 2, respectively storing the backup data and other related files of each backup, so that the corresponding version path can be easily obtained through string concatenation during subsequent reading.
[0034] Backup: First, back up data based on virtual machine snapshots. Whether it is incremental or full backup, it is backed up in blocks. Compression is also based on block compression, which makes it easy to obtain data blocks through count values. For example, the offset of the 0th block is 0* block_size, and the offset of the first block is 1* block_size. During the backup, a dump file (full data file) and a digest disk backup summary file (recording information about all blocks) will be generated. If it is an incremental backup, a differential diff file needs to be generated (to save the block data before the change, that is, the data of the previous version), and all differential recovery data information must be recorded in the differential file to the digest file. The structure of the digest file is as follows: Digest: 0 -- dump --<version, offset, size, compress> --diff --size=2 --<version1, offset, size, compress> --<version2, offset, size, compress> 1 -- dump --<version, offset, size, compress> --diff --size = 1 --<version1, offset, size, compress> 2 -- dump --<version, offset, size, compress> --diff --size = 2 --<version1, offset, size, compress> --<version2, offset, size, compress> ........ in,<version1, offset, size, compress> In the table, version indicates the version number, offset indicates the offset, size indicates the file capacity, and compress indicates the compression level. 0, 1, 2... respectively indicate the block numbers. Each block has only one dump record, but there may be multiple diff file records of length size. For example, if size=2, there are two diff file records. The digest file will be updated for each backup. If the block has not changed, just copy the block record and modify the version number of the dump information. If there is a change, update the dump block information and add a diff record to indicate that this version of the block has changed compared to the previous version. The advantage of this is that no matter whether the backup data is in raw format, qcow2 format, or any other format of data structure, the backup information can be used to locate the specific location of any block, thereby improving recovery efficiency. This backup structure applies to both the last version and the first version of the full data. When restoring, you do not need to read the digest file of each version. You only need to read the two digest files of the full version and the version to be restored.
[0035] Full backup: Create dump backup files and digest backup summary files, read backup data through the virtualization platform and transmit it to the backup server through the network, save all valid data blocks in the dump file, and keep the position of each block the same as the source disk. Record the dump block information in the digest file, that is, record a structure of<version,offset, size,compress> The dump information of the diff file is initialized to 0. Each subsequent incremental backup will overwrite the changed block data into the dump file and update each block information in the digest file, that is, keep the latest version of the dump backup file and digest file with the latest data.
[0036] Incremental backup: Get the changed block information from the cbt changed block table through the virtual platform. Before writing the changed block to the dump file, copy the data block at the changed block location from the dump file to the diff file, that is, copy the corresponding data of the changed block of the previous version to the diff file. Then overwrite the changed data block to the corresponding offset position of the dump file. Add a diff information record of this block to the digest file. For example, after a full backup, this block record is: Digest: 0 -- dump --<version = 0, offset = 0, size = 2097152,compress = 0> --diff --size = 0 Then the first incremental backup is updated as follows: Digest: 0 -- dump --<version = 1, offset = 0, size = 2097100, compress = 1> --diff --size = 1 --<version = 1, offset = 0, size = 2097152, compress = 0> That is, add a diff information <1, 0, 2097152, 0> record, which means that this block has changed in version 1. The data before the change (0 version data) is saved in the diff file of version 1. The offset in the diff file is 0, the size is 2097152 bytes, and it is not compressed. From the updated dump information, we can see that since this block of data has changed, it has been compressed into a data block of 2097100 bytes and stored in the dump file. The compression level is 1, that is, if this block is restored, it needs to be decompressed. Note that the version number of each data recorded in the diff is incremented, so that it can be easily found during subsequent traversal. If this block is deleted, the diff information of this block is also recorded, but the size in the dump information is modified to 0, which means that this block is an empty block in the current version. If the block before the change is an empty block, the size in the diff information can be assigned to 0. In short, each incremental digest data is merged with the previous digest data to ensure that the block information recorded in the digest file of the last version is the most complete. By reading the block information of the latest digest file, the specific location of each block data in any version can be located. The backed-up data is used for subsequent data recovery.
[0037] Step 21, receiving a data recovery request.
[0038] The data recovery request indicates the path of the backup virtual machine, the recovery data receiving path and the recovery version number.
[0039] The nfs (Network File System) network file system is used to optimize the time-consuming file network transmission and large amounts of data IO operations during the recovery process. First, create a directory fuse_nfs in the backup environment as the mount directory of the fuse service. When accessing this directory, the IO operation will be taken over by the fuse service. At the same time, an actual directory is required to store virtual machine disks and related configuration files, so a directory named fuse_space is created. The lib-fuse library can export fuse_space to the fuse_nfs directory. When we start the fuse service, we actually operate the fuse_space directory when we operate the fuse_nfs directory. At the same time, add the fuse_nfs directory of the backup environment as the nfs storage directory of the virtual platform, and the virtual platform can export fuse_nfs to the local nfs directory. Now the virtual platform operates the local nfs directory, which is the fuse_space directory of the backup server. Create a new virtual machine for recovery based on this nfs storage directory. At this time, the virtual platform will directly create the virtual machine directory and the most important disk files in the fuse_space directory. Because the recovery is performed in the backup environment and fuse is required to take over the IO operation, the directory operated during subsequent recovery is the intermediate mount directory fuse_nfs.
[0040] Restore the virtual machine configuration information and other files directly to the fuse_nfs directory, and then the recovery program communicates with the fuse service. The recovery program transmits the virtual disk path, the path of the backup virtual machine to be restored, the version number to be restored and other information to the fuse program. The fuse program registers this disk. After registration, regardless of whether the original disk is empty or not, the original disk content is invalid, and the disk data is now taken over by the fuse service. At the same time, a cache file named cache is created in the recovery file to store the new data blocks after the disk is written. That is, after the recovery is completed, only the backup file can be read, and the written data will be written to the corresponding position in this file.
[0041] Step 22, after the Fuse program registers the disk, it traverses the information of each data block in the backup summary file in turn, and executes steps 23 to 26 for the information of each data block.
[0042] Step 23, obtaining the full amount of information corresponding to the recovery version number from the information of the current data block.
[0043] In this step, in each version of the backup summary file, the version number in the full information indicates the version number of the data block, and the full information with the same version number as the recovery version number is searched in the information of the current data block in the backup summary file.
[0044] Step 24, determine whether the file capacity of the full information corresponding to the restored version number is 0.
[0045] If the file capacity size in the full information is 0, it means that the data of the current data block is empty under the version with the restored version number. If size≠0, it means that the data of the current data block is not empty under the version with the restored version number.
[0046] Step 25: If yes, generate a recovery data block information table according to the recovery version number, in which the recovery data information other than the recovery version number is empty.
[0047] The restored data block information table includes restored data information, and the restored data information includes: a restored version number version, a data block offset offset, a file capacity size, a storage file name file, and a compression level compress.
[0048] Specifically, when the data block under the recovery version number is empty, except for the version number and data block offset, all other information in the full information of the same version number as the recovery version number is 0, then the full information is copied, and the storage file name file is added, and file=0 is used as the recovery data block information table. For example, in the backup summary file dump--<version= 1, offset = 5, size = 0, compress = 0> , then the current data block record in the restored data block information table is<version = 1, offset = 5, size = 0, file = 0, compress = 0> The data block offset is the identifier of the data block.
[0049] Step 26, if not, generate a recovery data block information table according to the recovery version number and the latest version backup summary file of the current data block. When the file capacity of the full information corresponding to the recovery version number is not 0, it means that the data block is not empty.
[0050] When the file capacity of the full information corresponding to the recovery version number is not 0, the specific steps of generating the recovery data block information table according to the recovery version number and the latest version backup summary file of the current data block include steps 261 to 267: Step 261, determine whether the recovery version number is equal to the version number of the full information in the latest version backup summary file of the current data block.
[0051] In this step, you can restore the latest version or the old version. Here you need to determine whether you need to restore the latest version or the old version.
[0052] Step 262, when the recovery version number is equal to the version number of the full information in the latest version backup summary file of the current data block, copy the full information in the latest version backup summary file of the current data block, and add the storage file name file as the recovery data block information table, that is, dbt (Data Block Table). The value of the storage file name is the full data storage file name, that is, file = dump.
[0053] In this step, when the recovery version number is equal to the version number of the full information in the latest version backup summary file of the current data block, it means that the latest version of the data needs to be restored. Therefore, you only need to copy the full information in the latest version backup summary file of the current data block and add the storage file name as the recovery data block information table, and execute step 27.
[0054] For example, read the latest version of a data block and record it in the backup summary file dump--<version =0, offset = 10, size = 2097100,compress = 1> , then the recovery data information in the generated dbt is recorded as<version = 0, offset =10, size = 2097100, file = dump, compress = 1> .
[0055] It should be noted that the file system does not support copy-on-write in this step. If the current file system supports copy-on-write and is backed up without compression, you can directly use cp --reflink=always and other similar commands to quickly copy the dump file to the disk file. This is almost instantaneous, and there is no need to register this disk in fuse. When the virtual platform accesses this disk, it is directly available without going through fuse block mapping.
[0056] Step 263, when the recovery version number is not equal to the version number of the full information in the latest version backup summary file of the current data block, determine whether the recovery version number is equal to the version number of the modified information in the latest version backup summary file of the current data block.
[0057] In this step, when the recovery version number is not equal to the version number of the full information in the latest version backup summary file of the current data block, the old version needs to be restored, and the data block may or may not be modified in the old version. If it is modified, it will be recorded in the modification information diff during the incremental backup (the diff contains the corresponding version number information). If it is not modified, it will not be recorded in the diff.
[0058] Step 264, when the restored version number is equal to the version number of the modification information in the latest version backup summary file of the current data block, there are two cases: First, when the recovery version number is equal to the version number of the last modified information in the latest version backup summary file of the current data block, copy the full information in the latest version backup summary file of the current data block and add the storage file name as the recovery data block information table.
[0059] Here, one of the version numbers in the modified information is the same as the restored version number, which means that the current data block has been modified under the version to be restored. At the same time, if the restored version number is the same as the version number of the last piece of information in the modified information, it means that the data block was last modified under this version number (no modification in subsequent versions). Therefore, the information of the current data block is recorded in the full information. At this time, since the version number in the full information is the latest version number and is inconsistent with the restored version number, it is necessary to copy the information except the version number, and then increase the restored version number and the storage file name, write it into the restored data block information table, and then execute step 27. The value of the storage file name is the full data storage file name.
[0060] For example, Digest: 10 -- dump --<version = 5, offset = 10, size = 2097100, compress = 1> --diff --size = 3 --<version = 2, offset = 10, size = 2097222, compress = 0> --<version = 3, offset = 10, size = 2097111, compress = 0> --<version = 4, offset = 10, size = 2097333, compress = 0> There are six versions in total, 0-5, and the restored version number is 4. There are three records with version numbers 2, 3, and 4 in the diff. The last modification of the current data block is version 4, which means that the data block was not modified in version 5. The latest version backup summary file records dump--<version = 5, offset = 10, size = 2097100,compress = 1> , then the data of version 4 of the block is stored in the dump file, and the recovery data information in the generated dbt is recorded as<version =5, offset =10, size = 2097100, file = dump, compress = 1> .
[0061] The second method is to copy the next modification information with the same version number as the modification information in the latest version backup summary file of the current data block when the recovery version number is smaller than the version number of the last modification information in the latest version backup summary file of the current data block, and add the storage file name as the recovery data block information table.
[0062] Here, when the recovery version number is smaller than the version number of the last modification information in the latest version backup summary file of the current data block, it means that the current data block has been modified in the version number after the recovery version number. At this time, the data corresponding to the recovery version number of the data block is stored in the diff file in the next piece of modification information with the same version number as the recovery version number. The next piece of information here refers to the modification information of the first modified data after the recovery version number. Copy the next piece of modification information with the same version number as the modification information in the latest version backup summary file of the current data block, add the storage file name, write it into the recovery data block information table, and then execute step 27. The value of the storage file name is the name of the modified data storage file of the next piece of modification information, that is, file = diff.
[0063] For example, there are six versions 0-5, the restored version number is 1, and there are three records with version numbers 1, 3, and 4 in the diff file, indicating that the data block has been modified twice after restoring version 1. The data of version number 1 is stored in the diff file of version number 3, and the data of version 0 is stored in the diff file of version number 1. The restored data information record in the generated dbt is<version = 3, offset =10, size = 2097100, file = diff, compress = 1> Version = 3, file = diff means that the data is stored in the diff file of version 3.
[0064] Step 265: when the restored version number is not equal to the version number of the modified information in the latest version backup summary file of the current data block, determine whether the restored version number is less than the version number of the last modified information in the latest version backup summary file of the current data block.
[0065] In this step, if there is no version number in the modification information that is the same as the recovery version number, it means that the data block has not been modified when the recovery version number is the same as the recovery version number. At this time, it is necessary to determine the size of the recovery version number and the version number of the last modification information in the modification information to find the data of the version to be restored.
[0066] Step 266, when the recovery version number is smaller than the version number of the last modification information in the latest version backup summary file of the current data block, copy the first modification information in the latest version backup summary file of the current data block whose version number is greater than the recovery version number, and add the modification data storage file name of the first modification information as the recovery data block information table.
[0067] In this step, if the recovery version number is smaller than the version number of the last modification information, it is necessary to determine that the data recorded in the first modification information after the recovery version number is the data of the recovery version number. Copy the modification information that is larger than the recovery version number and is the first modification after the recovery version number, add the storage file name, write it into the recovery data block information table, and then execute step 27. The value of the storage file name is the modification data storage file name of the modification information that is the first modification after the recovery version number.
[0068] For example, there are six versions 0-5 in total, and the restored version number is 3. There are three records of version numbers 1, 2, and 5 in the diff file, indicating that the data block was modified once after restoring version 3, and was not modified in versions 3 and 4, but was modified in version 5. The data of version numbers 3 and 4 are the same and are stored in the diff file of version number 5. The restored data information record in the generated dbt is<version = 5, offset =10, size = 2097100, file = diff, compress = 1> Version = 5, file = diff means that the data is stored in the diff file of version 5.
[0069] Step 267, when the recovery version number is greater than the version number of the last modification information in the latest version backup summary file of the current data block, copy the full information in the latest version backup summary file of the current data block and add the storage file name as the recovery data block information table.
[0070] In this step, when the recovery version number is greater than the version number of the last item in the modification information, it means that the data block has not been modified in the version of the recovery version number and the subsequent versions. Therefore, the data of the recovery version number is recorded in the full information. At this time, the full information is copied, and the storage file name is added and written into the recovery data block information table, and then step 27 is executed. The value of the storage file name is the full data storage file name.
[0071] In this example, there are six versions 0-5, and the restored version number is 4. There are three records with version numbers 1, 2, and 3 in the diff, indicating that there are no modifications in the restored versions 4 and 5. The data of versions 4 and 5 are the same and are stored in the dump. The restored data information record in the generated dbt is<version = 5, offset =10, size = 2097100, file =dump, compress = 1> Version = 5, file = dump means that the data is stored in the dump file of version 5.
[0072] Step 27, save the restored data block information table, start traversing the information of the next data block in the backup summary file, and execute steps 23 to 26 until the traversal is completed.
[0073] After the traversal is completed, the information of the data block of the version to be restored is recorded in the dbt table. At this point, data recovery has been completed, and subsequent reading only requires reading the data indicated in the dbt table. Since the fuse library is a user space file system that implements various io interfaces, such as read and write system call interfaces, these interfaces can be implemented to interact with the kernel. When read and write system calls are made to the fuse mounted directory, these system calls will be intercepted by the fuse library, and then these interfaces will be called to perform read or write operations, thereby avoiding traversing and reading and writing complete disk data during recovery, and realizing rapid recovery of backup data.
[0074] The above steps complete the recovery. This process is done in seconds because only the restored data information is integrated and no large amount of data is restored. There is no decompression of data during the recovery process, so the speed is in seconds. At this time, the restored virtual machine starts, and the virtual platform reads the disk. It calls system functions such as wire read to access data and passes it in.<offset,size> This parameter structure is used to access data. Since the disk file has been registered before, when the fuse matches the disk file to be accessed, it takes over the read and write io operations of the disk file. When reading the file, it finds the corresponding data through dbt and returns it. The data read by the virtual platform is the real disk data, that is, the backup data. When writing files, the idea of copy-on-write is used. When reading files, the data will not be copied to the cache file. Only when writing data will the modified data be rewritten to the cache file.
[0075] The above steps can be used to restore a copy-on-write disk file. Since the original backup data file is not modified, it supports the simultaneous rapid restoration of different versions of the same backup data disk file. And the recovery time of each recovery task is in seconds.
[0076] Step 28: In response to the request to read the recovery data, a recovery data storage path is obtained according to the recovery data block information table and the path of the backup virtual machine.
[0077] In this step, after the dbt table is generated, the user can read the data in the dbt table as needed, that is, can read part of the data or all of the data under the recovery version number, so as to quickly obtain the required recovery data. When the user has a request to read the recovery data, according to the data block to be read in the read recovery data request, the information of the data block is found in the search dbt table, and the recovery data storage path is spliced according to the recovery version number and storage file name of the data block, as well as the path of the backup virtual machine, and then read from the path and return it to the user. The path of the backup virtual machine includes the backup root directory and the directory under the root directory.
[0078] In this example, / storage / back1 is the backup root directory and the next level directory. The record found in the dbt table is<version = 5, offset =10, size = 2097100, file = diff, compress = 1> , then the restored data storage path obtained by splicing is / storage / back1 / 5 / diff, which means that the data to be obtained is in the diff file of version 5. The file capacity size of the data block. If size=0, it means that the data block is currently an empty block in this recovery version and there is no need to query the subsequent file file. If size>0, it is necessary to query the corresponding file and restore the data by combining the corresponding multiple files.
[0079] Here, the cache file is used to cache the data written by the user after the dbt table is generated.
[0080] Step 29, reading the recovery data according to the recovery data storage path, and sending the read recovery data to the receiving platform according to the recovery data receiving path.
[0081] In this step, the recovery data receiving path is also the path of the restored virtual machine, and the receiving platform is also the restored virtual machine.
[0082] Both read and write interfaces have offset and size parameters, and the offset can be used to locate the number of the starting data block. For example, the block size is 2, and the user performs a read operation. In the parameter, the offset offset=0, and the starting data block is the 0th block. Then the number of the ending block can be located by offset+size. For example, size=4 is the size of two blocks, and the ending block is the 1st block. Therefore, it is known that the 0th block and the 1st block need to be read. At this time, by looking up the records of the 0th block and the 1st block in the dbt table, the kernel read function is called normally to obtain the storage path and read the block data of the two blocks and merge them and return them to the user.
[0083] In a feasible implementation, after the dbt table is generated, the user can write data in the restored data block when operating the restored virtual machine. When performing a write operation, the data will be stored in the cache file cache. After the data is written, the data in the data block is updated to the written data. The written data is stored in the cache file cache, and the backup data will not be modified, which will not affect the backup data on the backup end. Specifically, after the data is written, the parameters of the write function are<offset ,size> According to this parameter, the data block to which the data is written and the location of the data block can be determined, and then the corresponding record in the dbt table is modified, that is, the value of file in the record of the corresponding data block in the dbt table is changed to cache, and the compression level is changed to 0. At this time, the data has been updated and the version number has no practical meaning, so the version number is changed to 0.
[0084] For example, taking the block size as 2, the parameters of the write function are<offset = 1, size = 5> First, read the offset to find that the starting data block to be modified is the 0th block. Through offset+size, we can know that the last data block is the 2nd block. Then we need to read the 0th, 1st and 2nd data blocks, and overwrite the second half of the 0th data block with the data to be written. The 1st and 2nd data blocks are all overwritten. Finally, write the 0th, 1st and 2nd blocks into the cache file, and update the dbt file. Update the offset of the 0th, 1st and 2nd blocks to the offset in the cache file, modify file to cache, and set the version number and compression level to 0. When reading these blocks again, we can read new data directly from the cache file. For example, the modified record of the 0th block is <0, 0, 2097152, cache, 0>. At this time, the data block has been recorded in the cache file, so the version value is meaningless, and its value must be 0 for the compression level.
[0085] like Figure 2 As shown, a third aspect of an embodiment of the present invention provides a backup data recovery device based on a FUSE system, comprising: The receiving module 31 is used to receive a data recovery request; the data recovery request indicates the path of the backup virtual machine, the recovery data receiving path and the recovery version number; The traversal module 32 is used to perform the traversal step: when traversing each data block, the full amount of information corresponding to the recovery version number is obtained in the backup summary file of the current data block; A judgment module 33 is used to judge whether the file capacity of the full amount of information corresponding to the restored version number is 0; A first generating module 34, configured to generate, according to the recovery version number, a recovery data block information table in which the recovery data information other than the recovery version number and the data block offset is empty; The second generating module 35 is used to generate a recovery data block information table according to the recovery version number and the latest version backup summary file of the current data block if no; wherein the recovery data block information table includes recovery data information, and the recovery data information includes: recovery version number, data block offset, file capacity, storage file name and compression level; A saving module 36 is used to save the restored data block information table, traverse the next data block, and return to execute the traversal step until the traversal is completed; An acquisition module 37, configured to respond to a request to read recovery data and acquire a recovery data storage path according to the recovery data block information table and the path of the backup virtual machine; The reading module 38 is used to read the recovery data according to the recovery data storage path, and send the read recovery data to the receiving platform according to the recovery data receiving path.
[0086] In one embodiment of the present invention, a recovery data block information table is generated according to the recovery version number and the latest version backup summary file of the current data block, including: Determine whether the recovery version number is equal to the version number of the full information in the latest version backup summary file of the current data block; If they are equal, copy the full information in the latest version backup summary file of the current data block and add the full data storage file name as the recovery data block information table; If not, determine whether the restored version number is equal to the version number of the modification information in the latest version backup summary file of the current data block; If they are equal, when the recovery version number is equal to the version number of the last modification information in the latest version backup summary file of the current data block, copy the full information in the latest version backup summary file of the current data block, and add the full data storage file name as the recovery data block information table; When the recovery version number is smaller than the version number of the last modification information in the latest version backup summary file of the current data block, copy the next modification information of the modification information with the same version number as the modification information in the latest version backup summary file of the current data block, and add the modification data storage file name of the next modification information as the recovery data block information table; If not equal, determine whether the restored version number is less than the version number of the last modification information in the latest version backup summary file of the current data block; If it is smaller, copy the first modification information whose version number is greater than the recovery version number in the modification information in the latest version backup summary file of the current data block, and add the modification data storage file name of the first modification information as the recovery data block information table; If it is greater, the full information in the latest version backup summary file of the current data block is copied, and the full data storage file name is added as the recovery data block information table.
[0087] In one embodiment of the present invention, obtaining the recovery data storage path according to the recovery data block information table and the path of the backup virtual machine includes: The recovery data storage path is obtained according to the recovery version number, storage file name and backup virtual machine path of the recovery data block information table.
[0088] In one embodiment of the present invention, it also includes: creating a cache file for storing write data under the recovery data receiving path.
[0089] A fourth aspect of an embodiment of the present invention provides an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, a backup data recovery method based on a FUSE system provided in the above embodiment of the present invention is implemented.
[0090] A fifth aspect of an embodiment of the present invention further provides a computer-readable storage medium on which a computer program is stored. When the computer program is executed by a processor, the steps of a backup data recovery method based on a FUSE system provided in the above embodiment of the present invention are implemented.
[0091] The memory may include a random access memory (RAM) or a non-volatile memory (NVM), such as at least one disk memory. Optionally, the memory may also be at least one storage device located away from the aforementioned processor.
[0092] The above-mentioned processor can be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), etc.; it can also be a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware systems.
[0093] The method provided in the embodiment of the present invention can be applied to electronic devices. Specifically, the electronic device can be: a desktop computer, a portable computer, an intelligent mobile terminal, a server, etc. This is not limited here, and any electronic device that can implement the present invention belongs to the protection scope of the present invention.
[0094] As for the device / electronic device embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment.
[0095] The present invention is described with reference to flowcharts and / or block diagrams of methods, devices (apparatus), and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0096] These computer program instructions may also be stored in a computer readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture including an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.
[0097] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.
[0098] Obviously, those skilled in the art can make various changes and modifications to the present invention without departing from the spirit and scope of the present invention. Thus, if these modifications and variations of the present invention fall within the scope of the claims of the present invention and their equivalents, the present invention is also intended to include these modifications and variations.
Claims
1. A backup data recovery method based on FUSE system, characterized in that: The following steps are involved: Receive a data recovery request; wherein the data recovery request indicates a path of the backup virtual machine, a recovery data receiving path, and a recovery version number; Traversal step: when traversing the information of each data block in the backup summary file, obtain the full information corresponding to the recovery version number in the information of the current data block; Determine whether the file capacity of the full amount of information corresponding to the restored version number is 0; If yes, generating a recovery data block information table according to the recovery version number, in which the recovery data information other than the recovery version number and the data block offset is empty; If not, generating a recovery data block information table according to the recovery version number and the latest version backup summary file of the current data block; wherein the recovery data block information table includes recovery data information, and the recovery data information includes: recovery version number, data block offset, file capacity, storage file name and compression level; Saving the restored data block information table, traversing the information of the next data block in the backup summary file, and returning to execute the traversal step until the traversal is completed; In response to a request to read recovery data, acquiring a recovery data storage path according to the recovery data block information table and the path of the backup virtual machine; The recovery data is read according to the recovery data storage path, and the read recovery data is sent to the receiving platform according to the recovery data receiving path.
2. The method according to claim 1, characterized in that The step of generating a recovery data block information table according to the recovery version number and the latest version backup summary file of the current data block includes: Determine whether the recovery version number is equal to the version number of the full information in the latest version backup summary file of the current data block; If they are equal, copy the full information in the latest version backup summary file of the current data block and add the full data storage file name as the recovery data block information table; If not, determining whether the recovery version number is equal to the version number of the modification information in the latest version backup summary file of the current data block; If they are equal, when the recovery version number is equal to the version number of the last modification information in the latest version backup summary file of the current data block, copy the full information in the latest version backup summary file of the current data block, and add the full data storage file name as the recovery data block information table; When the recovery version number is smaller than the version number of the last modification information in the latest version backup summary file of the current data block, copy the next modification information of the modification information with the same version number as the recovery version number in the modification information in the latest version backup summary file of the current data block, and add the modification data storage file name of the next modification information as the recovery data block information table; If not equal, determine whether the recovery version number is less than the version number of the last modification information in the latest version backup summary file of the current data block; If it is smaller, copy the first modification information whose version number in the modification information in the latest version backup summary file of the current data block is greater than the recovery version number, and add the modification data storage file name of the first modification information as the recovery data block information table; If it is greater, copy the full information in the latest version backup summary file of the current data block and add the full data storage file name as the recovery data block information table.
3. The method according to claim 2, characterized in that The obtaining the recovery data storage path according to the recovery data block information table and the path of the backup virtual machine includes: The recovery data storage path is acquired according to the recovery version number of the recovery data block information table, the storage file name and the path of the backup virtual machine.
4. The method according to claim 2, characterized in that The method further includes: creating a cache file for storing write data under the recovery data receiving path.
5. A backup data recovery device based on FUSE system, characterized in that: include: A receiving module, configured to receive a data recovery request, wherein the data recovery request indicates a path of a backup virtual machine, a recovery data receiving path, and a recovery version number; A traversal module is used to perform a traversal step: when traversing the information of each data block in the backup summary file, the full amount of information corresponding to the recovery version number is obtained from the information of the current data block; A judgment module, used to judge whether the file capacity of the full information corresponding to the restored version number is 0; A first generating module, configured to generate, according to the recovery version number, a recovery data block information table in which the recovery data information other than the recovery version number and the data block offset is empty; A second generating module is used to generate a recovery data block information table according to the recovery version number and the latest version backup summary file of the current data block if no; wherein the recovery data block information table includes recovery data information, and the recovery data information includes: recovery version number, data block offset, file capacity, storage file name and compression level; A saving module, used for saving the restored data block information table, traversing the information of the next data block in the backup summary file, and returning to execute the traversal step until the traversal is completed; An acquisition module, configured to respond to a request to read recovery data and acquire a recovery data storage path according to the recovery data block information table and the path of the backup virtual machine; The reading module is used to read the recovery data according to the recovery data storage path, and send the read recovery data to the receiving platform according to the recovery data receiving path.
6. The device according to claim 5, characterized in that The step of generating a recovery data block information table according to the recovery version number and the latest version backup summary file of the current data block includes: Determine whether the recovery version number is equal to the version number of the full information in the latest version backup summary file of the current data block; If they are equal, copy the full information in the latest version backup summary file of the current data block and add the full data storage file name as the recovery data block information table; If not, determining whether the recovery version number is equal to the version number of the modification information in the latest version backup summary file of the current data block; If they are equal, when the recovery version number is equal to the version number of the last modification information in the latest version backup summary file of the current data block, copy the full information in the latest version backup summary file of the current data block, and add the full data storage file name as the recovery data block information table; When the recovery version number is smaller than the version number of the last modification information in the latest version backup summary file of the current data block, copy the next modification information of the modification information with the same version number as the recovery version number in the modification information in the latest version backup summary file of the current data block, and add the modification data storage file name of the next modification information as the recovery data block information table; If not equal, determine whether the recovery version number is less than the version number of the last modification information in the latest version backup summary file of the current data block; If it is smaller, copy the first modification information whose version number in the modification information in the latest version backup summary file of the current data block is greater than the recovery version number, and add the modification data storage file name of the first modification information as the recovery data block information table; If it is greater, copy the full information in the latest version backup summary file of the current data block and add the full data storage file name as the recovery data block information table.
7. The device according to claim 5, characterized in that The obtaining the recovery data storage path according to the recovery data block information table and the path of the backup virtual machine includes: The recovery data storage path is acquired according to the recovery version number of the recovery data block information table, the storage file name and the path of the backup virtual machine.
8. The device according to claim 5, characterized in that Also includes: A cache file for storing write data is created under the recovery data receiving path.
9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the program, the backup data recovery method based on the FUSE system as described in any one of claims 1 to 4 is implemented.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the backup data recovery method based on the FUSE system described in any one of claims 1 to 4 is implemented.
Citation Information
Patent Citations
International freight rate data storage method and system
CN110928839A
Fast fine-grained recovery method and device based on backup data
CN112965856A
Database management method and system, electronic equipment and storage medium
CN114721881A
Metadata backup method and metadata backup server for files on HDD (hard disk drive) disk
CN117076191A
Security protection method and device, storage medium and computer program product
CN117725630A