Disk backup method and device, storage medium and electronic equipment
By performing full and incremental backup of QCOW2 format virtual disk files, and using snapshot and data block update mechanisms, the existing backup methods are solved, and a fast and efficient backup and recovery process is achieved.
Patent Information
- Application Number
- CN202510412399.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-03
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2045-04-03
AI Technical Summary
When performing incremental backup, existing disk backup methods need to update full backup data files, resulting in cumbersome backup and recovery processes and inefficient efficiency.
The virtual disk files in QCOW2 format are fully backup and incremental backup. The data block is determined by generating snapshots, and the data block is written to the full backup data file and the difference file, overwrite the backup data block, and update the full backup data file.
It realizes a fast and efficient disk backup and recovery process, especially when performing the latest version of data recovery, different files that do not require incremental backups. The recovery process is fast and ensures data recovery efficiency.
Smart Images

Figure CN119938409A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data backup, and in particular to a disk backup method, device, storage medium and electronic equipment. Background Art
[0002] When performing disk backup for block devices, full backup data files (such as Dump files) are generally generated by traversing the entire disk to achieve full disk backup; on this basis, incremental disk backup can be further performed as needed. During incremental backup, the Dump file will not be updated, but a corresponding differential file will be generated based on the difference in data transformation. The differential file stores the data overwritten by the new incremental data.
[0003] Generally, a full disk backup is performed initially, and only incremental disk backups are performed periodically thereafter. When restoring data, it is necessary to restore based on dump files, differential files, etc., and the recovery process is relatively cumbersome. Summary of the invention
[0004] To solve the above problems, an object of the embodiments of the present invention is to provide a disk backup method, device, storage medium and electronic device.
[0005] In a first aspect, an embodiment of the present invention provides a disk backup method, comprising: When a full backup of a virtual disk file in QCOW2 format is required, generating a first current snapshot of the virtual disk file; Determine a full amount of data blocks according to the first snapshot, and write each data block into a full backup data file; After the full backup, if it is necessary to perform an incremental backup of the current version of the virtual disk file, generate a current second snapshot of the virtual disk file; Determine, according to the second snapshot and a previous snapshot generated by the previous backup, a first data block that has changed since the previous backup; Writing the backup data block corresponding to the position of the first data block in the full backup data file to the difference file corresponding to the incremental backup of the current version, and then writing the first data block to the full backup data file and overwriting the backup data block; The block information of the first data block is written into the incremental summary file of the current version incremental backup.
[0006] In some optional implementations, determining the full amount of data blocks according to the first snapshot and writing each data block into the full backup data file includes: Reading a file header of the virtual disk file, and saving the read first file header data to a full summary file; the first file header data includes a block size and a first snapshot table offset; searching snapshot information of the first snapshot according to the first snapshot table offset, and determining the first and second level tables according to the snapshot information of the first snapshot; the snapshot information of the first snapshot includes a first level table offset corresponding to the first snapshot, and the first level table offset corresponding to the first snapshot is used to locate the first level table; The data blocks corresponding to the respective table entries of the first secondary table are written into the full backup data file; the table entries of the first secondary table correspond to the offsets of the corresponding data blocks.
[0007] In some optional implementations, determining the first data block that has changed after the last backup according to the second snapshot and the last snapshot generated by the last backup includes: Reading a file header of the virtual disk file, and saving the read second file header data to an incremental summary file of the incremental backup of the current version; the second file header data includes a block size and a second snapshot table offset; searching for snapshot information of the second snapshot according to the second snapshot table offset, and determining the second secondary table according to the snapshot information of the second snapshot; the snapshot information of the second snapshot includes a primary table offset corresponding to the second snapshot, and the primary table offset corresponding to the second snapshot is used to locate the second secondary table; Compare the second secondary table with the previous secondary table corresponding to the previous snapshot generated by the previous backup item by item to determine the target table item that has changed; the target table item corresponds to the offset of the first data block; The first data block of the block size is read according to the offset of the first data block corresponding to the target table entry.
[0008] In some optional embodiments, the method further comprises: After the current version is incrementally backed up, delete the last snapshot generated by the last backup.
[0009] In some optional embodiments, the method further comprises: When the current version incremental backup is performed on the virtual disk file, the block attributes of each storage block are saved to a bitmap file corresponding to the current version incremental backup.
[0010] In some optional embodiments, the method further comprises: When performing data recovery, if the target version incremental backup to be recovered is the latest version incremental backup, determining the block information of the second data block to be recovered according to the incremental summary file of the latest version incremental backup; The second data block is obtained from the full backup data file according to the block information of the second data block, and the second data block is restored to a corresponding position of the virtual disk file.
[0011] In some optional embodiments, the method further comprises: When performing data recovery, if the target version incremental backup to be recovered is not the latest version incremental backup, determining the block information of the third data block to be recovered according to all incremental summary files between the latest version incremental backup and the target version incremental backup; The third data blocks are obtained from the full backup data file or a difference file of other version incremental backups according to the block information of each of the third data blocks, and the third data blocks are restored to corresponding positions of the virtual disk file.
[0012] In a second aspect, an embodiment of the present invention further provides a disk backup device, comprising: A first snapshot module, used for generating a current first snapshot of the virtual disk file when a full backup of the virtual disk file in the QCOW2 format is required; A full backup module, used to determine a full amount of data blocks according to the first snapshot, and write each data block into a full backup data file; A second snapshot module is used to generate a current second snapshot of the virtual disk file after the full backup when it is necessary to perform an incremental backup of the current version of the virtual disk file; a comparison module, configured to determine, based on the second snapshot and a previous snapshot generated by the previous backup, a first data block that has changed since the last backup; The incremental backup module is used to write the backup data block corresponding to the position of the first data block in the full backup data file to the difference file corresponding to the current version incremental backup, and then write the first data block to the full backup data file and overwrite the backup data block; write the block information of the first data block to the incremental summary file of the current version incremental backup.
[0013] In a third aspect, an embodiment of the present invention further provides a computer storage medium, wherein the computer storage medium stores computer executable instructions, and the computer executable instructions are used in any one of the above-mentioned disk backup methods.
[0014] In a fourth aspect, an embodiment of the present invention further provides an electronic device, including: at least one processor; and, a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute any one of the disk backup methods described above.
[0015] In the solution provided in the first aspect of the embodiment of the present invention, full disk backup and incremental disk backup can be realized based on the characteristics of the QCOW2 format file; a full backup data file is generated during full backup, and the full backup data file is iteratively updated during each incremental backup, so that the full backup data file can save all the data of the last backup. When data is subsequently restored, it can be mainly restored based on the full backup data file, especially when restoring the latest version of data, the differential file of the incremental backup is not required, the recovery process is fast, and the data recovery efficiency can be guaranteed.
[0016] In order to make the above-mentioned objects, features and advantages of the present invention more obvious and easy to understand, preferred embodiments are given below and described in detail with reference to the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0018] Figure 1 A flow chart of a disk backup method provided by an embodiment of the present invention is shown; Figure 2 A flow chart of another disk backup method provided by an embodiment of the present invention is shown; Figure 3 A schematic diagram showing the structure of a disk backup device provided by an embodiment of the present invention is shown; Figure 4 A schematic structural diagram of an electronic device for executing a disk backup method provided by an embodiment of the present invention is shown. DETAILED DESCRIPTION
[0019] In the description of the present invention, it should be understood that the terms "center", "longitudinal", "lateral", "length", "width", "thickness", "up", "down", "front", "back", "left", "right", "vertical", "horizontal", "top", "bottom", "inside", "outside", "clockwise", "counterclockwise" and the like indicate orientations or positional relationships based on the orientations or positional relationships shown in the accompanying drawings, and are only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the referred device or element must have a specific orientation, be constructed and operated in a specific orientation, and therefore should not be understood as limiting the present invention.
[0020] In addition, the terms "first" and "second" are used for descriptive purposes only and should not be understood as indicating or implying relative importance or implicitly indicating the number of the indicated technical features. Therefore, the features defined as "first" and "second" may explicitly or implicitly include one or more of the features. In the description of the present invention, the meaning of "plurality" is two or more, unless otherwise clearly and specifically defined.
[0021] In the present invention, unless otherwise clearly specified and limited, the terms "installed", "connected", "connected", "fixed" and the like should be understood in a broad sense, for example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be a direct connection, or it can be an indirect connection through an intermediate medium, or it can be the internal communication of two components. For ordinary technicians in this field, the specific meanings of the above terms in the present invention can be understood according to specific circumstances.
[0022] QCOW2 (QEMU Copy-On-Write 2, second generation QEMU copy-on-write) format is a disk image file format supported by QEMU (QuickEMUlator, quick emulator), supporting snapshot, compression, encryption and other features. Among them, QCOW2 format files use the Copy-On-Write mechanism, that is, the modification of the file will not directly overwrite the original data, but will be written to a new location, and the original data will be retained for subsequent recovery.
[0023] A disk backup method provided by an embodiment of the present invention is shown in Figure 1 As shown, the method includes: Step 101: When a full backup of a virtual disk file in the QCOW2 format is required, a first current snapshot of the virtual disk file is generated.
[0024] In this embodiment, when backing up data of a virtual disk file in QCOW2 format, a full backup is first performed, and when data backup is needed later, one or more incremental backups are performed. In each data backup (including full backup and incremental backup), a snapshot command needs to be executed on the virtual disk file to obtain a corresponding snapshot.
[0025] Specifically, when the virtual disk file is fully backed up, the snapshot generated at this time is called the first snapshot. The first snapshot is hereinafter represented by snapshot_1.
[0026] It should be noted that the first snapshot snapshot_1 is a snapshot generated when the full backup is currently being performed. For ease of distinction, this moment is referred to as moment t0.
[0027] Step 102: Determine full data blocks according to the first snapshot, and write each data block into a full backup data file.
[0028] In this embodiment, the first snapshot snapshot_1 contains relevant information about the data blocks in the virtual disk file. Based on this information, the full data blocks can be located, so that all these data blocks are backed up and written to corresponding files, that is, full backup data files.
[0029] Among them, the full backup data file is a dump file, which is a file composed of data blocks and empty blocks, and is used to save the full disk data. In the traditional full backup process, a full dump file is generally generated, but the dump file is fixed and unchanged. The dump file will not be updated during subsequent incremental updates; in this embodiment, the dump file will also be updated synchronously during each incremental update, that is, the full backup data file will be updated synchronously, which will be explained later.
[0030] The format of the full backup data file may be: disk id.dump; for example, the full backup data file corresponding to disk 0 is, for example, 0.dump. In this embodiment, a dump file is generated during a full backup, and there is no separate corresponding dump file for each incremental backup.
[0031] Optionally, the QCOW2 file uses a two-level table structure to manage data blocks: the first level table (L1 Table): each table entry points to a second level table (L2 Table); the second level table (L2 Table): each table entry points to the offset of a data block (cluster). Based on the above characteristics of the QCOW2 file, it is easy to determine each data block of the virtual disk file.
[0032] Specifically, the above step 102 of "determining full data blocks according to the first snapshot, and writing each data block into the full backup data file" includes steps A1 to A3.
[0033] Step A1: read the file header of the virtual disk file, and save the read first file header data to the full summary file; the first file header data includes a block size and a first snapshot table offset.
[0034] Step A2: searching for snapshot information of the first snapshot according to the first snapshot table offset, and determining the first and second tables according to the snapshot information of the first snapshot; the snapshot information of the first snapshot includes the first-level table offset corresponding to the first snapshot, and the first-level table offset corresponding to the first snapshot is used to locate the first and second tables.
[0035] Step A3: write the data blocks corresponding to the entries of the first and second level tables into the full backup data file; the entries of the first and second level tables correspond to the offsets of the corresponding data blocks.
[0036] In this embodiment, when performing a full backup of a virtual disk file in QCOW2 format, the file header part of the virtual disk file is read to obtain the file header data at this time, that is, the first file header data. The first file header data includes the block size blockSize of the virtual disk file and the current snapshot table offset, that is, the first snapshot table offset. In addition, the read first file header data is backed up and saved in the summary file of the full backup, that is, in addition to generating a full backup data file during the full backup, a full summary file is also generated.
[0037] For example, each backup corresponds to a folder. In the case of a full backup, there is a "file1.digest" digest file in the "full backup" folder, which is the full digest file. After reading the first file header data of the virtual disk file, it is saved to the "file1.digest" digest file.
[0038] The first snapshot table offset is the address offset of the current snapshot table, and the current snapshot information table can be located based on the first snapshot table offset. The snapshot information table is a structure in the QCOW2 file for storing all snapshot metadata, and each snapshot occupies an item in the snapshot information table, that is, the first snapshot is a snapshot in the snapshot information table.
[0039] Specifically, the snapshot information table records the snapshot information of each snapshot, which includes a snapshot ID and an L1 table offset (i.e., a first-level table offset). Based on the snapshot ID, the snapshot information belonging to the first snapshot snapshot_1 can be determined, and the first-level table offset of the first snapshot snapshot_1 can be determined. The first-level table offset points to the location of the first-level table (L1 Table) corresponding to the first snapshot snapshot_1, so that the first-level table corresponding to the first snapshot snapshot_1 can be determined.
[0040] After determining the first level table corresponding to the first snapshot snapshot_1, traversing the first level table, the second level table (L2 Table) corresponding to the first snapshot snapshot_1 can be found, that is, the first second level table. It can be understood that the first second level table contains multiple table entries, each of which points to the offset of a data block.
[0041] For ease of description, L2Table_1 represents the first secondary table; it can be understood that each entry in the first secondary table L2Table_1 records the offset of the data block in the QCOW2 file, and the data block pointed to by the offset is the actual data corresponding to the first snapshot snapshot_1.
[0042] In this embodiment, after the first and second level table L2Table_1 is determined, the data blocks corresponding to the respective table entries of the first and second level table L2Table_1 may be written into the full backup data file (Dump file).
[0043] Specifically, according to the offset of each table entry in the first and second level tables L2Table_1, each table entry is read according to the block size blockSize, and the read data is the corresponding data block, and these data blocks are written in sequence to the corresponding positions of the full backup data file (Dump file).
[0044] For example, each data block corresponds to the index of the corresponding table entry of the first and second level tables L2Table_1, and the data block found from the offset corresponding to the first entry of the first and second level tables L2Table_1 should be written into the interval of the length of blockSize from the beginning of the full backup data file (Dump file). The rest of the data blocks are deduced in the same way.
[0045] In addition, the data blocks read based on the first and second level tables L2Table_1 can be compressed as needed to save backup storage space. The saved full backup data file (Dump file, whose name can be, for example, file1.dump) is an empty file, and the disk space occupied is only the sum of the sizes of all the data blocks written.
[0046] In addition, the block information of each data block and empty block is appended to the "file1.digest" summary file (i.e., the full summary file) in the "full backup" folder, that is, the block information of each data block and empty block in the full backup data file is written to the full summary file; the block information may include, for example, the size, location and other information of the data block and empty block.
[0047] The above steps are a complete full backup. At this time, the full backup data file (such as file1.dump) saves the data of all data blocks of the virtual disk file at time t0. It is the full dump file at time t0. The full summary file (such as file1.digest) saves the file header data of the virtual disk file at time t0, the block information of each block, etc.
[0048] Step 103: After the full backup, if it is necessary to perform an incremental backup of the current version of the virtual disk file, generate a second snapshot of the current virtual disk file.
[0049] In this embodiment, after performing the above-mentioned full backup, if the virtual disk file needs to be backed up again, an incremental backup can be performed. Specifically, at time t1 after time t0, if the virtual disk file needs to be backed up, an incremental backup can be performed. Among them, this embodiment allows multiple incremental backups, each corresponding to a unique version. For the convenience of description, the incremental backup at time t1 is called the current version incremental backup.
[0050] Specifically, after a period of time, the time comes from time t0 to time t1, and data backup is required at this time. Then, similar to the above step 101, a snapshot command is executed on the virtual disk file in the QCOW2 format that needs to be backed up to generate a second snapshot snapshot_2. It can be understood that since there is generally data update from the last backup (for example, time t0) to time t1, some data of the virtual disk file at this time may be different from that of the virtual disk file at time t0.
[0051] Step 104: Determine the first data block that has changed after the last backup according to the second snapshot and the last snapshot generated by the last backup.
[0052] Specifically, before the current incremental backup, there is a previous backup, which is the last data backup before the current incremental backup (i.e., before time t1). For example, if the first incremental backup is performed after the full backup, the previous backup of the current incremental backup is the full backup at time t0; if one or more incremental backups have been performed, the previous backup of the current incremental backup is the incremental backup of the previous version.
[0053] Furthermore, in this embodiment, a corresponding snapshot needs to be generated for each backup, so a snapshot is also generated during the last backup, that is, the last snapshot. For example, the last snapshot is the first snapshot snapshot_1.
[0054] This incremental backup only needs to focus on the changed data. As shown in step 102, the corresponding data blocks can be determined based on the snapshot, so the data blocks at time t1 can be determined based on the second snapshot, and the data blocks at the corresponding time can also be determined based on the previous snapshot.
[0055] Therefore, by comparing the relevant information of the second snapshot with that of the previous snapshot, it is possible to determine which data blocks have changed from the last backup to the current time t1. These changed data blocks are called first data blocks.
[0056] Optionally, the above step 104 of "determining the first data block that has changed after the last backup according to the second snapshot and the last snapshot generated by the last backup" includes steps B1 to B4.
[0057] Step B1, read the file header of the virtual disk file, and save the read second file header data to the incremental summary file of the current version incremental backup; the second file header data includes a block size and a second snapshot table offset.
[0058] Step B2, searching for snapshot information of the second snapshot according to the second snapshot table offset, and determining the second-level table according to the snapshot information of the second snapshot; the snapshot information of the second snapshot includes the first-level table offset corresponding to the second snapshot, and the first-level table offset corresponding to the second snapshot is used to locate the second-level table.
[0059] Similar to the above steps A1 to A2, based on the file header of the virtual disk file at the current moment, the second snapshot table offset can also be determined, and then the corresponding secondary table, that is, the second secondary table L2Table_2, can be determined. Similarly, based on the last snapshot generated by the last backup, the corresponding secondary table, that is, the previous secondary table, can also be determined; for example, if the last backup is a full backup, the previous secondary table is the first secondary table L2Table_1.
[0060] As mentioned above, each data backup generates a corresponding summary file; when performing an incremental backup of the current version, the summary file generated is an incremental summary file, and the second file header data read from the file header is saved to the incremental summary file of the incremental backup of the current version. Therefore, during subsequent backups, the previous level and level 2 tables can be determined based on the summary file of the previous backup.
[0061] Step B3, comparing the second level-2 table with the previous level-2 table corresponding to the previous snapshot generated by the previous backup item by item, and determining the target table item that has changed; the target table item corresponds to the offset of the first data block.
[0062] Step B4: read the first data block of the block size according to the offset of the first data block corresponding to the target table entry.
[0063] In this embodiment, each item in the second level-2 table and the previous level-2 table represents the offset of the corresponding data block. Since the QCOW2 format file adopts the write-time copy method, if a table item changes, it can be determined that the corresponding data block has changed. Therefore, by determining the changed target table item, the changed data block, that is, the first data block, can be located.
[0064] Step 105: write the backup data block corresponding to the position of the first data block in the full backup data file into the differential file corresponding to the incremental backup of the current version, then write the first data block into the full backup data file and overwrite the backup data block.
[0065] In this embodiment, after determining the first data block, the data block corresponding to the position of the first data block in the full backup data file (Dump file) can be determined, that is, the backup data block, and the backup data block is written to the difference file corresponding to the incremental backup of the current version. In addition, the first data block is written to the full backup data file and the original backup data block is overwritten to achieve the update of the full backup data file.
[0066] Among them, the above-mentioned step 105 "writing the backup data block corresponding to the position of the first data block in the full backup data file to the differential file corresponding to the current version incremental backup" can specifically include: locating the backup data block at the corresponding position in the full backup data file according to the position of the target table item in the second-level table; writing the backup data block to the differential file corresponding to the current version incremental backup.
[0067] Step 106: Write the block information of the first data block into the incremental summary file of the current version incremental backup.
[0068] In this embodiment, block information of the first data block is also determined, and the block information may indicate the position and size of the first data block, etc. The block information of the first data block is written to the incremental summary file of the current version incremental backup to facilitate subsequent data recovery.
[0069] Optionally, the method further includes: when performing a current version incremental backup on the virtual disk file, saving the block attributes of each storage block to a bitmap file corresponding to the current version incremental backup.
[0070] In this embodiment, the digest file is used to save the file header and other information of the QCOW2 file, and its format can be: disk id.digest, for example: 0.digest; each backup can generate a corresponding digest file, and one incremental backup corresponds to one version of the incremental digest file. The difference file is used to save the data overwritten by the new incremental data, and its format can be, for example: disk id.diff.backward, for example: 0.diff.backward. One incremental backup corresponds to one version of the difference file.
[0071] In addition, when performing incremental backup, a corresponding bitmap file is also generated, which is used to calculate the location of the data blocks in the dump file. The format can be: disk id.bmp, such as 0.bmp. One incremental backup corresponds to one version of the bitmap file, which contains the data block information of the entire disk.
[0072] Specifically, during incremental backup, the block attributes of each storage block (including data blocks and empty blocks) are saved to the bitmap file corresponding to the incremental backup of the current version. The block attribute can specifically indicate whether the storage block is a data block or an empty block, and can also indicate compression information of the data block.
[0073] In addition, the attributes of the blocks (whether the blocks are data blocks or empty blocks) are marked in the bitmap file (for example, file1.bmp) of the incremental backup of the current version, and other information may also be saved.
[0074] Take the last backup as a full backup as an example, that is, the last level 2 table is the first level 2 table L2Table_1 of the full backup. The first level 2 table L2Table_1 is compared with the second level 2 table L2Table_2 item by item according to the corresponding table items. If the value of an item in the second level 2 table L2Table_2 changes, then the table item is the target table item.
[0075] Based on the target entry in the second level table L2Table_2, find the corresponding offset position and read the data of block size blockSize. The read data is the first data block. According to the offset entry subscript corresponding to the first data block in the second level table L2Table_2, write it to the corresponding position of the full backup data file (for example, file1.dump) for overwriting; before overwriting, read out the backup data block at the overwriting position and write it to the difference file of the current incremental version (such as file1.diff.backward), and append the changed information of this data block to the incremental summary file of the current version incremental backup to ensure that the backup data of all backup time points can be restored.
[0076] For example, the first secondary table L2Table_1 is compared with the second secondary table L2Table_2 item by item according to the corresponding table items. If the Kth difference appears in the Nth item of the second secondary table L2Table_2, the Nth item of the second secondary table L2Table_2 is a target table item. According to the offset corresponding to the Nth item of the second secondary table L2Table_2, the data block can be found and read, which is recorded as incBlockK, which is a first data block. In addition, the data in the Nth data block (i.e., the backup data block) of the full backup data file (e.g., file1.dump) is read out in advance and written to the Kth block of the difference file (e.g., file1.diff.backward) in the folder of the current version incremental backup. In addition, the data in the first data block incBlockK is overwritten and written to the Nth data block of the full backup data file (e.g., file1.dump), and the Nth block in the full backup data file is marked as a data block in the bitmap file (e.g., file1.bmp) in the folder of the current version incremental backup.
[0077] Optionally, the method further includes: after the incremental backup of the current version, deleting the last snapshot generated by the last backup. By deleting useless snapshots, the storage space of the host where the virtual machine is located can be saved.
[0078] It can be understood that the above steps 103 to 106 are a process of incremental backup. If incremental backup is still needed, steps 103 to 106 can be repeated.
[0079] A disk backup method provided by an embodiment of the present invention can realize full disk backup and incremental disk backup based on the characteristics of QCOW2 format files; a full backup data file is generated during full backup, and the full backup data file is iteratively updated during each incremental backup, so that the full backup data file can save all the data of the last backup. When data is subsequently restored, it can be mainly restored based on the full backup data file, especially when the latest version of data is restored, the difference file of the incremental backup is not required, the recovery process is fast, and the data recovery efficiency can be guaranteed.
[0080] Another disk backup method provided by the embodiment of the present invention is shown in Figure 2 As shown, the method includes: Step 201: When a full backup of a virtual disk file in the QCOW2 format is required, a first current snapshot of the virtual disk file is generated.
[0081] Step 202: Determine full data blocks according to the first snapshot, and write each data block into a full backup data file.
[0082] Step 203: After the full backup, if it is necessary to perform an incremental backup of the current version of the virtual disk file, generate a second snapshot of the current virtual disk file.
[0083] Step 204: Determine the first data block that has changed after the last backup according to the second snapshot and the last snapshot generated by the last backup.
[0084] Step 205: write the backup data block corresponding to the position of the first data block in the full backup data file into the differential file corresponding to the current version incremental backup, then write the first data block into the full backup data file and overwrite the backup data block.
[0085] Step 206: Write the block information of the first data block into the incremental summary file of the current version incremental backup.
[0086] The above steps 201 to 206 illustrate the process of full backup and incremental backup. Figure 1 The relevant description of the illustrated embodiment will not be repeated here.
[0087] Step 207: When performing data recovery, if the target version incremental backup to be recovered is the latest version incremental backup, determine the block information of the second data block to be recovered according to the incremental summary file of the latest version incremental backup.
[0088] Step 208: Obtain the second data block from the full backup data file according to the block information of the second data block, and restore the second data block to a corresponding position of the virtual disk file.
[0089] In this embodiment, when performing data recovery, it is necessary to determine which version of the backup to which the data needs to be recovered, that is, it is necessary to determine the target version incremental backup. If the target version incremental backup is the latest version incremental backup, then since the full backup data file stores all the data of the last backup (that is, the latest version incremental backup), the block information of the corresponding data block is extracted from the incremental summary file of the latest version incremental backup. For ease of description, the data block is referred to as the second data block. Based on the block information of these second data blocks, the data corresponding to the second data blocks can be read from the full backup data file, and then these second data blocks can be restored to the corresponding positions of the virtual disk file, and the data recovery of the latest version can be completed.
[0090] For example, if the folder of the latest version of the incremental file is "incremental version X", then the incremental summary file in "incremental version X" can be used to read the corresponding file header information and write the file header information into the file header of the virtual disk file to be restored; then, based on the incremental summary file and bitmap file in "incremental version X", determine where the second data block in the full backup data file (such as file1.dump) needs to be written in the virtual disk file, whether it needs to be decompressed, etc., and completing the header information also depends on the above-mentioned incremental summary file and bitmap file.
[0091] Optionally, after the above step 206, the method may further include steps C1 and C2.
[0092] Step C1: during data recovery, if the target version incremental backup to be recovered is not the latest version incremental backup, determine the block information of the third data block to be recovered based on all incremental summary files between the latest version incremental backup and the target version incremental backup.
[0093] Step C2: Obtain the third data blocks from the full backup data file or the difference file of other version incremental backup according to the block information of each third data block, and restore the third data blocks to the corresponding position of the virtual disk file.
[0094] In this embodiment, if the target version incremental backup to be restored is not the latest version incremental backup, an older version needs to be restored. For example, there are currently "full version", "incremental version 1", and "incremental version 2", and "incremental version 2" is the latest version; if you want to restore the disk corresponding to the "incremental version 1" time, you need to use the relevant information of all versions between the latest version incremental backup and the target version incremental backup.
[0095] Specifically, if you need to restore the disk corresponding to the "Incremental Version 1" moment, you need four files, namely the incremental summary file and bitmap file in "Incremental Version 1" and "Incremental Version 2", to determine whether the data required at the "Incremental Version 1" moment is saved in the full backup data file or in the difference file under the "Incremental Version 2" folder, and determine where the data block should be written in the disk file to be restored and whether it needs to be decompressed. In addition, completing the header information also depends on all the above-mentioned incremental summary files and bitmap files.
[0096] A disk backup method provided by an embodiment of the present invention updates the dump file corresponding to the full backup each time an incremental backup is performed. When restoring the latest version, no differential file is required, there will not be excessive calculations, and the recovery process is faster. When restoring an older version, all incremental summary files and bitmap files in all versions between the latest version and the target version are used to participate in the calculation to determine whether the required data block is in the full backup data file or in the differential file of a certain version, thereby realizing data recovery of the older version.
[0097] The above describes in detail the process of the disk backup method. The method can also be implemented by a corresponding device. The structure and function of the device are described in detail below.
[0098] Based on the same inventive concept, the embodiment of the present invention also provides a disk backup device, see Figure 3 As shown, the device comprises: The first snapshot module 301 is used to generate a current first snapshot of the virtual disk file when a full backup of the virtual disk file in the QCOW2 format is required; A full backup module 302, configured to determine a full amount of data blocks according to the first snapshot, and write each data block into a full backup data file; The second snapshot module 303 is used to generate a current second snapshot of the virtual disk file after the full backup when it is necessary to perform an incremental backup of the current version of the virtual disk file; A comparison module 304, configured to determine a first data block that has changed after the last backup according to the second snapshot and the last snapshot generated by the last backup; The incremental backup module 305 is used to write the backup data block corresponding to the position of the first data block in the full backup data file to the difference file corresponding to the current version incremental backup, and then write the first data block to the full backup data file and overwrite the backup data block; write the block information of the first data block to the incremental summary file of the current version incremental backup.
[0099] In some optional implementations, the full backup module 302 determines the full data blocks according to the first snapshot, and writes each data block into the full backup data file, including: Reading a file header of the virtual disk file, and saving the read first file header data to a full summary file; the first file header data includes a block size and a first snapshot table offset; searching snapshot information of the first snapshot according to the first snapshot table offset, and determining the first and second level tables according to the snapshot information of the first snapshot; the snapshot information of the first snapshot includes a first level table offset corresponding to the first snapshot, and the first level table offset corresponding to the first snapshot is used to locate the first level table; The data blocks corresponding to the respective table entries of the first secondary table are written into the full backup data file; the table entries of the first secondary table correspond to the offsets of the corresponding data blocks.
[0100] In some optional implementations, the comparison module 304 determines the first data block that has changed after the last backup according to the second snapshot and the last snapshot generated by the last backup, including: Reading a file header of the virtual disk file, and saving the read second file header data to an incremental summary file of the incremental backup of the current version; the second file header data includes a block size and a second snapshot table offset; searching for snapshot information of the second snapshot according to the second snapshot table offset, and determining the second secondary table according to the snapshot information of the second snapshot; the snapshot information of the second snapshot includes a primary table offset corresponding to the second snapshot, and the primary table offset corresponding to the second snapshot is used to locate the second secondary table; Compare the second secondary table with the previous secondary table corresponding to the previous snapshot generated by the previous backup item by item to determine the target table item that has changed; the target table item corresponds to the offset of the first data block; The first data block of the block size is read according to the offset of the first data block corresponding to the target table entry.
[0101] In some optional implementations, the incremental backup module 305 is further configured to: after the incremental backup of the current version, delete the last snapshot generated by the last backup.
[0102] In some optional implementations, the incremental backup module 305 is further configured to: when performing an incremental backup of the current version of the virtual disk file, save the block attributes of each storage block to a bitmap file corresponding to the incremental backup of the current version.
[0103] In some optional implementations, the device further includes a data recovery module, which is used to: When performing data recovery, if the target version incremental backup to be recovered is the latest version incremental backup, determining the block information of the second data block to be recovered according to the incremental summary file of the latest version incremental backup; The second data block is obtained from the full backup data file according to the block information of the second data block, and the second data block is restored to a corresponding position of the virtual disk file.
[0104] In some optional implementations, the device further includes a data recovery module, which is used to: When performing data recovery, if the target version incremental backup to be recovered is not the latest version incremental backup, determining the block information of the third data block to be recovered according to all incremental summary files between the latest version incremental backup and the target version incremental backup; The third data blocks are obtained from the full backup data file or a difference file of other version incremental backups according to the block information of each of the third data blocks, and the third data blocks are restored to corresponding positions of the virtual disk file.
[0105] An embodiment of the present invention further provides a computer storage medium, wherein the computer storage medium stores computer executable instructions, which include a program for executing the above-mentioned disk backup method. The computer executable instructions can execute the method in any of the above-mentioned method embodiments.
[0106] Among them, the computer storage medium can be any available medium or data storage device that can be accessed by the computer, including but not limited to magnetic storage (such as floppy disks, hard disks, magnetic tapes, magneto-optical disks (MO), etc.), optical storage (such as CD, DVD, BD, HVD, etc.), and semiconductor storage (such as ROM, EPROM, EEPROM, non-volatile memory (NANDFLASH), solid-state drive (SSD)), etc.
[0107] Figure 4 The block diagram of the structure of an electronic device of another embodiment of the present invention is shown. The electronic device 1100 may be a host server with computing capability, a personal computer PC, or a portable computer or terminal, etc. The specific embodiment of the present invention does not limit the specific implementation of the electronic device.
[0108] The electronic device 1100 includes at least one processor 1110, a communications interface 1120, a memory array 1130, and a bus 1140. The processor 1110, the communications interface 1120, and the memory array 1130 communicate with each other via the bus 1140.
[0109] The communication interface 1120 is used to communicate with network elements, where the network elements include, for example, a virtual machine management center, shared storage, etc.
[0110] The processor 1110 is used to execute programs. The processor 1110 may be a central processing unit (CPU), or an application specific integrated circuit (ASIC), or one or more integrated circuits configured to implement the embodiments of the present invention.
[0111] The memory 1130 is used for executable instructions. The memory 1130 may include a high-speed RAM memory, and may also include a non-volatile memory, such as at least one disk memory. The memory 1130 may also be a memory array. The memory 1130 may also be divided into blocks, and the blocks may be combined into virtual volumes according to certain rules. The instructions stored in the memory 1130 can be executed by the processor 1110, so that the processor 1110 can execute the disk backup method in any of the above method embodiments.
[0112] The above is only a specific embodiment of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art can easily think of changes or substitutions within the technical scope disclosed by the present invention, which should be included in the protection scope of the present invention. Therefore, the protection scope of the present invention should be based on the protection scope of the claims.
Claims
1. A disk backup method, characterized in that: include: When a full backup of a virtual disk file in QCOW2 format is required, generating a first current snapshot of the virtual disk file; Determine a full amount of data blocks according to the first snapshot, and write each data block into a full backup data file; After the full backup, if it is necessary to perform an incremental backup of the current version of the virtual disk file, generate a current second snapshot of the virtual disk file; Determine, based on the second snapshot and a previous snapshot generated by the previous backup, a first data block that has changed since the previous backup; Writing the backup data block corresponding to the position of the first data block in the full backup data file to the difference file corresponding to the incremental backup of the current version, and then writing the first data block to the full backup data file and overwriting the backup data block; The block information of the first data block is written into the incremental summary file of the current version incremental backup.
2. The method according to claim 1, characterized in that The step of determining the full amount of data blocks according to the first snapshot and writing each data block into the full backup data file includes: Reading a file header of the virtual disk file, and saving the read first file header data to a full summary file; the first file header data includes a block size and a first snapshot table offset; searching snapshot information of the first snapshot according to the first snapshot table offset, and determining the first and second level tables according to the snapshot information of the first snapshot; the snapshot information of the first snapshot includes a first level table offset corresponding to the first snapshot, and the first level table offset corresponding to the first snapshot is used to locate the first level table; The data blocks corresponding to the respective table entries of the first secondary table are written into the full backup data file; the table entries of the first secondary table correspond to the offsets of the corresponding data blocks.
3. The method according to claim 1 or 2, characterized in that: The determining, according to the second snapshot and the last snapshot generated by the last backup, the first data block that has changed after the last backup comprises: Reading a file header of the virtual disk file, and saving the read second file header data to an incremental summary file of the incremental backup of the current version; the second file header data includes a block size and a second snapshot table offset; searching for snapshot information of the second snapshot according to the second snapshot table offset, and determining the second secondary table according to the snapshot information of the second snapshot; the snapshot information of the second snapshot includes a primary table offset corresponding to the second snapshot, and the primary table offset corresponding to the second snapshot is used to locate the second secondary table; Compare the second secondary table with the previous secondary table corresponding to the previous snapshot generated by the previous backup item by item to determine the target table item that has changed; the target table item corresponds to the offset of the first data block; The first data block of the block size is read according to the offset of the first data block corresponding to the target table entry.
4. The method according to claim 1, characterized in that Also includes: After the current version is incrementally backed up, delete the last snapshot generated by the last backup.
5. The method according to claim 1, characterized in that Also includes: When the current version incremental backup is performed on the virtual disk file, the block attributes of each storage block are saved to a bitmap file corresponding to the current version incremental backup.
6. The method according to claim 1, characterized in that Also includes: When performing data recovery, if the target version incremental backup to be recovered is the latest version incremental backup, determining the block information of the second data block to be recovered according to the incremental summary file of the latest version incremental backup; The second data block is obtained from the full backup data file according to the block information of the second data block, and the second data block is restored to a corresponding position of the virtual disk file.
7. The method according to claim 1, characterized in that Also includes: When performing data recovery, if the target version incremental backup to be recovered is not the latest version incremental backup, determining the block information of the third data block to be recovered according to all incremental summary files between the latest version incremental backup and the target version incremental backup; The third data blocks are obtained from the full backup data file or a difference file of other version incremental backups according to the block information of each of the third data blocks, and the third data blocks are restored to corresponding positions of the virtual disk file.
8. A disk backup device, characterized in that: include: A first snapshot module, used for generating a current first snapshot of the virtual disk file when a full backup of the virtual disk file in the QCOW2 format is required; A full backup module, used to determine a full amount of data blocks according to the first snapshot, and write each data block into a full backup data file; A second snapshot module is used to generate a current second snapshot of the virtual disk file after the full backup when it is necessary to perform an incremental backup of the current version of the virtual disk file; a comparison module, configured to determine, based on the second snapshot and a previous snapshot generated by the previous backup, a first data block that has changed since the last backup; An incremental backup module, used for writing a backup data block corresponding to the position of the first data block in the full backup data file to a difference file corresponding to the incremental backup of the current version, and then writing the first data block to the full backup data file and overwriting the backup data block; The block information of the first data block is written into the incremental summary file of the current version incremental backup.
9. A computer storage medium, characterized in that: The computer storage medium stores computer executable instructions, and the computer executable instructions are used to execute the disk backup method according to any one of claims 1 to 7.
10. An electronic device, characterized in that: include: at least one processor; as well as, a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the disk backup method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Disk backup method and apparatus for virtual machine
CN107544870A
Virtual machine disk backup method and apparatus
CN107544871A
Block-level data backup system and method
CN113849342A
Fast synthesis backup and recovery method based on snapshots
CN114090337A
Backup method and device of virtual machine
CN117032884A
Cited By
Data backup and recovery method and device, electronic equipment and storage medium
CN121070697A