Virtual Machine Data Recovery Method, Device, Computer Equipment, Readable Storage Medium and Program Product
By creating subdisks to write data shards in parallel during virtual machine data recovery, the problem of lock restrictions of virtualization management software is solved, multi-threaded parallel recovery is realized, and the speed and efficiency of virtual machine data recovery is improved.
Patent Information
- Application Number
- CN202411428304.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-10-14
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2044-10-14
AI Technical Summary
In the prior art, the virtual machine data recovery method is limited by the locking mechanism of virtualization management software, resulting in low recovery speed and efficiency in multi-threaded parallel recovery scenarios, especially when processing large-capacity disks or multiple disks.
Generate child disks by creating snapshots of virtual machines, and open the parent disk and child disk in parallel in write state. Multiple threads use fragments to write backup data to the parent disk and child disk respectively, bypassing the lock limit and achieving multi-threaded parallel recovery.
It significantly shortens data recovery time, improves the recovery speed and efficiency of virtual machine disks, and significantly reduces recovery time in large-capacity disks or multiple disk recovery scenarios, ensuring the reliability and integrity of data recovery.
Smart Images

Figure CN119415318B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of virtual machines, and particularly to a virtual machine data recovery method, apparatus, computer device, computer-readable storage medium, and computer program product. Background Art
[0002] With the popularization of virtualization technology, more and more enterprises adopt virtualization platforms to deploy and manage their IT infrastructures, and data backup and recovery are key links to ensure business continuity and data security. To quickly and reliably recover virtual machine data, many backup software and solutions have been developed and applied.
[0003] In related technologies, virtual machine data recovery methods usually perform data operations based on the API interfaces provided by virtualization management software. However, some limitations of virtualization management software can cause bottlenecks in virtual machines in multi-threaded recovery scenarios, affecting the speed and efficiency of data recovery. Summary of the Invention
[0004] Based on this, it is necessary to provide a virtual machine data recovery method, apparatus, computer device, computer-readable storage medium, and computer program product that can improve the speed and efficiency of virtual machine data recovery for the above technical problems.
[0005] In a first aspect, this application provides a virtual machine data recovery method, including:
[0006] Determine a target virtual machine for which data recovery is to be performed;
[0007] Create a parent disk corresponding to the target virtual machine, and obtain a child disk of the target virtual machine according to the snapshot created for the target virtual machine;
[0008] Open the parent disk and the child disk in write state, and respectively and parallelly write multiple data shards corresponding to the backup data into the parent disk and the child disk to obtain the parent disk and the child disk with the data shards backed up;
[0009] Obtain the target virtual machine with data recovery completed according to the parent disk and the child disk with the data shards backed up.
[0010] In one embodiment, the opening the parent disk and the child disk in write state includes:
[0011] Open the child disk in write state, and add an associated blocking identifier to the opened child disk; the associated blocking identifier is used to indicate blocking the locking of the parent disk corresponding to the child disk when the child disk is opened in write state;
[0012] When the associated blocking identifier is added to the sub-disk, the parent disk is opened in the write state.
[0013] In one embodiment, the obtaining of the sub-disk of the target virtual machine according to the snapshot created for the target virtual machine includes:
[0014] Creating a plurality of snapshots for the target virtual machine, and obtaining a plurality of sub-disks of the target virtual machine according to the plurality of snapshots;
[0015] The parallel writing of the plurality of data shards corresponding to the backup data into the parent disk and the sub-disk respectively includes:
[0016] Writing the plurality of data shards corresponding to the backup data into the parent disk and the plurality of sub-disks respectively in parallel; the data shards written into the parent disk and the plurality of sub-disks are not repeated.
[0017] In one embodiment, the parallel writing of the plurality of data shards corresponding to the backup data into the parent disk and the plurality of sub-disks respectively includes:
[0018] Determining the number of disks corresponding to the plurality of sub-disks and the parent disk;
[0019] Invoking a plurality of idle threads corresponding to the number of disks in a pre-set thread pool, and parallelly using the plurality of idle threads to write the plurality of data shards corresponding to the backup data into the parent disk and the plurality of sub-disks respectively in parallel; the number of threads of the plurality of idle threads is equal to the number of disks.
[0020] In one embodiment, the obtaining of the target virtual machine with the data recovery completed according to the parent disk and the sub-disk storing the data shards includes:
[0021] Associating the parent disk and the sub-disk storing the data shards, and obtaining the target virtual machine with the data recovery completed according to the associated parent disk and sub-disk; the target virtual machine is used to search for target data from the complete backup data when receiving a target data query request.
[0022] In one embodiment, the associating of the parent disk and the sub-disk storing the data shards includes:
[0023] Obtaining the parent disk identifier corresponding to the parent disk storing the data shards;
[0024] Modifying the attribute information of the sub-disk storing the data shards under the parent disk identifier attribute to the parent disk identifier.
[0025] Second aspect, the present application further provides a virtual machine data recovery device, including:
[0026] A virtual machine determination module, configured to determine a target virtual machine for which data recovery is to be performed;
[0027] A snapshot creation module, configured to create a parent disk corresponding to the target virtual machine, and obtain a child disk of the target virtual machine according to a snapshot created for the target virtual machine;
[0028] A parallel writing module, configured to open the parent disk and the child disk in a write state, and respectively and parallelly write multiple data shards corresponding to backup data into the parent disk and the child disk, to obtain the parent disk and the child disk with the data shards backed up;
[0029] An arrangement module, configured to obtain the target virtual machine with data recovery completed according to the parent disk and the child disk with the data shards backed up.
[0030] Third aspect, the present application further provides a computer device, including a memory and a processor, where the memory stores a computer program, and when the processor executes the computer program, the following steps are implemented:
[0031] Determine a target virtual machine for which data recovery is to be performed;
[0032] Create a parent disk corresponding to the target virtual machine, and obtain a child disk of the target virtual machine according to a snapshot created for the target virtual machine;
[0033] Open the parent disk and the child disk in a write state, and respectively and parallelly write multiple data shards corresponding to backup data into the parent disk and the child disk, to obtain the parent disk and the child disk with the data shards backed up;
[0034] Obtain the target virtual machine with data recovery completed according to the parent disk and the child disk with the data shards backed up.
[0035] Fourth aspect, the present application further provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the following steps are implemented:
[0036] Determine a target virtual machine for which data recovery is to be performed;
[0037] Create a parent disk corresponding to the target virtual machine, and obtain a child disk of the target virtual machine according to a snapshot created for the target virtual machine;
[0038] Open the parent disk and the child disk in write state, and write multiple data shards corresponding to the backup data into the parent disk and the child disk in parallel respectively, to obtain the parent disk and the child disk with the data shards backed up;
[0039] Obtain the target virtual machine with data recovery completed according to the parent disk and the child disk with the data shards backed up.
[0040] In a fifth aspect, the present application further provides a computer program product, including a computer program, which when executed by a processor implements the following steps:
[0041] Determine a target virtual machine for which data recovery is to be performed;
[0042] Create a parent disk corresponding to the target virtual machine, and obtain a child disk of the target virtual machine according to the snapshot created for the target virtual machine;
[0043] Open the parent disk and the child disk in write state, and write multiple data shards corresponding to the backup data into the parent disk and the child disk in parallel respectively, to obtain the parent disk and the child disk with the data shards backed up;
[0044] Obtain the target virtual machine with data recovery completed according to the parent disk and the child disk with the data shards backed up.
[0045] After determining the target virtual machine for which data recovery is to be performed, the above virtual machine data recovery method, apparatus, computer device, computer-readable storage medium, and computer program product can create a parent disk corresponding to the target virtual machine, obtain a child disk of the target virtual machine according to the snapshot created for the target virtual machine, then open the parent disk and the child disk in write state, and write multiple data shards corresponding to the backup data into the parent disk and the child disk in parallel respectively, to obtain the parent disk and the child disk with the data shards backed up; obtain the target virtual machine with data recovery completed according to the parent disk and the child disk with the data shards backed up. In this embodiment, by generating a child disk of the target virtual machine using the snapshot created for the target virtual machine, more data writing channels can be provided for the original parent disk based on the newly created child disk, and the recovery operation can be dispersed to the parent disk and the child disk. Furthermore, by writing multiple data shards of the backup data into the parent disk and the child disk in parallel respectively, the locking limit of the virtual machine can be broken through, multi-threaded parallel recovery can be achieved, the data recovery time can be greatly shortened, and the data recovery speed and efficiency of the virtual machine disk can be significantly and effectively improved. Especially when dealing with large-capacity disks or multiple disks, the data recovery time of the virtual machine can be significantly reduced. Description of the Drawings
[0046] To more clearly illustrate the technical solutions in the embodiments of the present application or related technologies, the following will briefly introduce the accompanying drawings required for the description of the embodiments of the present application or related technologies. Obviously, the accompanying drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other related drawings can also be obtained based on these drawings.
[0047] Figure 1 It is a schematic flowchart of a virtual machine data recovery method in an embodiment;
[0048] Figure 2 It is a schematic flowchart of the steps for opening a parent disk and a child disk in an embodiment;
[0049] Figure 3 It is a schematic flowchart of another virtual machine data recovery method in an embodiment;
[0050] Figure 4 It is a structural block diagram of a virtual machine data recovery device in an embodiment;
[0051] Figure 5 It is an internal structure diagram of a computer device in an embodiment. Detailed implementation manners
[0052] In order to make the purpose, technical solutions and advantages of the present application more clear and understandable, the following further details the present application in combination with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0053] To enable those skilled in the art to better understand the present application, the following first introduces the virtual machine data recovery methods in related technologies.
[0054] With the rapid development of cloud computing and big data technologies and the popularization of virtualization technologies, the demand of enterprises for virtualization platforms has been increasing continuously, and the scale and complexity of virtual machine disks have also grown accordingly. This has promoted the further development of data backup and recovery technologies. Traditional backup and recovery technologies have gradually shifted from single-threaded operations to multi-threaded operations in order to improve the recovery speed while reducing the impact on the business system.
[0055] However, due to some technical limitations of virtualization management software, related technologies still cannot fully meet the requirements of efficient recovery. Specifically, when backing up virtual machine data, according to the content of the backup set, the data is written to the virtual disk sequentially. The recovery process will be carried out in a single thread, or in the case of multi-threading, multi-threading is used to perform serial operations on different parts of the disk. However, due to the existence of the locking mechanism of the virtual disk development kit of the virtualization management software, it is impossible to open the same virtual disk file in write mode multiple times at the same time, which limits the feasibility of multi-threaded parallel recovery and can only rely on single-threaded or multi-threaded serial operations, reducing the recovery efficiency. For scenarios of large-capacity virtual disks or simultaneous recovery of multiple disks, the recovery time is often long and it is difficult to meet the requirements of rapid recovery.
[0056] Based on this, the present application provides a virtual machine data recovery method, device, computer device, computer-readable storage medium and computer program product to at least solve the above technical problems.
[0057] In one embodiment, as Figure 1 shown, a virtual machine data recovery method is provided. In this embodiment, it is exemplified that the method is applied to a server. It can be understood that the method can also be applied to a terminal, and can also be applied to a system including a terminal and a server, and is implemented through the interaction between the terminal and the server. In this embodiment, the method includes the following steps S101 to S104:
[0058] S101, determine the target virtual machine for which data recovery is to be performed.
[0059] Among them, the target virtual machine is any virtual machine for which data recovery is to be performed. Data recovery to be performed can be understood as recovering specified backup data.
[0060] S102, create a parent disk corresponding to the target virtual machine, and obtain a child disk of the target virtual machine according to the snapshot created for the target virtual machine.
[0061] After determining the target virtual machine, a corresponding parent disk can be created for the target virtual machine first. After creating the parent disk, according to the virtual machine snapshot mechanism of the virtualization management software, a snapshot of the target virtual machine can be created, and a child disk of the target virtual machine can be obtained according to the snapshot created for the target virtual machine.
[0062] In some exemplary embodiments, before restoring the backup data, a new virtual machine disk file (Virtual Machine Disk, VMDK) can be created for the target virtual machine, and the parent disk corresponding to the target virtual machine can be obtained based on the newly created original virtual machine disk file. In some examples, the virtual machine disk file can contain the virtual machine operating system, applications, data, etc. Among them, the parent disk can be used to store the data before the creation of the target virtual machine snapshot, and the child disk can record the data changes after the creation of the target virtual machine snapshot. The combination of the two forms a complete virtual machine disk image. When creating the parent disk and the child disk before restoring the backup data, the parent disk and the child disk can be regarded as disks that do not store the backup data.
[0063] A snapshot of a virtual machine is a mechanism for recording the state of a virtual machine at a certain point in time. Snapshots allow users to restore to a specific moment in time and are commonly used for backup, testing, and recovery operations. In some exemplary embodiments, in a virtualized environment, the parent disk can be the main disk of the target virtual machine, storing the operating system, applications, and data of the target virtual machine, etc. Creating the parent disk is a basic step in deploying a virtual machine and can be completed through virtualization management software. Then, the "Snapshot" or "Snapshot Management" option can be found in the operation menu of the virtual machine, and "Create Snapshot" or "New Snapshot" can be clicked. In some examples, the snapshot to be created can be one or more. When creating a snapshot, the name and description of the snapshot can be entered to facilitate differentiating different snapshots in the future. After that, confirm the operation of creating the snapshot and wait for the target virtual machine to complete the snapshot creation process. When creating a snapshot, the virtualization management software can capture the current state of the target virtual machine, including memory, CPU registers, and disk data, etc., and save this information in one or more snapshot files. These files can include one or more such as.vmsd (database of snapshot information and snapshot manager information),.vmsn (current configuration and active state of the virtual machine), and *.vmdk (virtual disk file).
[0064] Each time a snapshot of the target virtual machine is created, a new sub-disk (or incremental disk) can be generated. This sub-disk stores all the changes to the virtual machine's disk data since the last snapshot. Over time, a snapshot chain is formed, where each snapshot corresponds to a sub-disk. A series of associated snapshot files form a snapshot chain through a parent-child relationship. The snapshot chain allows the state of the virtual machine to be restored from multiple points in time. The sub-disk can be understood as a pointer to the parent disk or the previous sub-disk, plus the changes that have occurred since the last snapshot. These changes are stored incrementally and thus do not occupy a large amount of storage space. In some examples, when a particular snapshot needs to be accessed, the virtualization management software loads the sub-disk corresponding to that snapshot and merges it with the parent disk or the previous sub-disk to restore the state of the virtual machine. This process is usually fast because only the necessary incremental changes need to be read and merged.
[0065] S103, open the parent disk and the sub-disk in write state, and write the multiple data shards corresponding to the backup data into the parent disk and the sub-disk in parallel respectively, to obtain the parent disk and the sub-disk with the data shards backed up.
[0066] In the related art, due to the locking mechanism of the virtualization management software, it is often impossible to open the same virtual disk file in write mode multiple times at the same time. This limitation makes it so that even if there are multiple threads, the recovery process can only serially recover data from the virtual disk file by multiple threads, and the advantages of multi-threaded parallel operations cannot be effectively utilized. This limitation results in low recovery efficiency and long recovery time when dealing with the recovery tasks of large-capacity disks or multiple disks.
[0067] In response to this, in this embodiment, by creating a virtual machine snapshot and generating a sub-disk, multiple disks that are related to each other and whose disk states can be independent of each other can be obtained, namely the parent disk and at least one sub-disk associated with the parent disk. Since the parent disk and the sub-disk can be opened in write state at the same time, the multiple data shards obtained by pre-splitting the backup data can be written into the parent disk and the sub-disk respectively, to obtain the parent disk and the sub-disk with the data shards backed up. By innovatively introducing the snapshot mechanism in this embodiment, during the virtual machine data recovery process, a sub-disk associated with the parent disk can be additionally introduced on the basis of the original parent disk, to obtain multiple associated disks that can write data at the same time, bypassing the locking limitation, ensuring the smooth progress of parallel operations, overcoming the shortcoming in the traditional technology that multi-threaded writing operations cannot be performed on the same disk file at the same time, and realizing multi-threaded parallel recovery of different data blocks to the parent disk and the sub-disk.
[0068] Exemplarily, a virtual disk development kit (such as the VDDK (Virtual Disk Development Kit) library) can be used to open the parent disk and the child disk files respectively, ensuring that both can be operated in parallel in the write state; wherein, the virtual disk development kit can be a virtual disk development kit provided by virtualization management software, which consists of a set of application programming interfaces (Application Programming Interface, API) and tools, and can be used to develop applications to interact and operate with the virtual disks in virtual machines, so as to implement operations such as reading, writing, and managing VMDK virtual disk files, and it supports one or more operations such as backup, recovery, and cloning of virtual machine disk files.
[0069] S104. Obtain a target virtual machine that has completed data recovery based on the parent disk and the child disk with data shards backed up.
[0070] After writing each data shard into the parent disk and the child disk respectively, a target virtual machine that has completed data recovery can be obtained based on the current parent disk and child disk. In some embodiments, the parent disk and the child disk can be associated; in other embodiments, after the data shards are written into the parent disk and the child disk, disk association may not be performed, and the user can directly start the virtual machine. The system can first attempt to read the target data from the child disk. When it is determined that the child disk does not have the required target data, it will attempt to read the target data from the parent disk.
[0071] In the above virtual machine data recovery method, after determining the target virtual machine to be recovered, a parent disk corresponding to the target virtual machine can be created, and based on the snapshot created for the target virtual machine, the child disk of the target virtual machine can be obtained. Then, the parent disk and the child disk are opened in the write state, and multiple data shards corresponding to the backup data are respectively written into the parent disk and the child disk in parallel to obtain the parent disk and the child disk with data shards backed up; a target virtual machine that has completed data recovery is obtained based on the parent disk and the child disk with data shards backed up. In this embodiment, by generating the child disk of the target virtual machine using the snapshot created for the target virtual machine, more data writing channels can be provided for the original parent disk based on the newly created child disk, and the recovery operation can be dispersed to the parent disk and the child disk. Furthermore, by writing multiple data shards of the backup data into the parent disk and the child disk in parallel respectively, the locking limit of the virtual machine can be broken through, and multi-threaded parallel recovery can be achieved, greatly shortening the data recovery time, significantly and effectively improving the data recovery speed and efficiency of the virtual machine disk. Especially when dealing with large-capacity disks or multiple disks, the data recovery time of the virtual machine is significantly reduced.
[0072] In one embodiment, such as Figure 2As shown, in step S103, opening the parent disk and the child disk in a write state may include the following steps:
[0073] S201, opening a sub-disk in a write state, and adding an associated blocking mark to the opened sub-disk; the associated blocking mark is used to indicate that when the sub-disk is opened in a write state, the parent disk corresponding to the sub-disk is prevented from being locked.
[0074] In the process of using traditional virtualization management software, you can open the child disk or the parent disk by opening the relevant tools. Since data will be read in the order of child disk and parent disk, when the child disk is opened, the parent disk to which it belongs will be opened by default at the same time. Some implementations will automatically generate a .lck file named after the disk after opening the disk, indicating that the file is locked and the file cannot be opened later. For example, there is a parent disk a.vmdk and a child disk a-000001.vmdk. When a-000001.vmdk is opened, the parent disk will be opened and locked automatically, so two files will be automatically generated, namely a.vmdk.lck and a-000001.vmdk.lck, which indicate that the parent disk a.vmdk and the child disk a-000001.vmdk are locked respectively.
[0075] In this regard, in this embodiment, in order to enable the child disk to be opened in a write state at the same time, the child disk can be opened first, and an associated blocking mark can be added to the opened child disk. The associated blocking mark can indicate that the parent disk corresponding to the child disk is blocked from being locked when the child disk is opened in a write state. For example, when opening a child disk, a flag "VIXDISKLIB_FLAG_OPEN_SINGLE_LINK" can be added as an associated blocking mark, indicating that the parent disk to which it belongs is not opened and the parent disk is not locked. Accordingly, when opening a-000001.vmdk, only one file a-000001.vmdk.lck will be generated, which will not affect the subsequent opening of the parent disk.
[0076] S202: When a related isolation flag is added to the child disk, the parent disk is opened in a write state.
[0077] Furthermore, the parent disk can be opened in a write state while the associated isolation flag is added to the child disk.
[0078] In this embodiment, the child disk is first opened in the write state, and then an association barrier flag is added to avoid automatic association with the parent disk, thereby avoiding the potential locking restriction on simultaneous data writing of the parent disk and the child disk.
[0079] In one embodiment, in step S102, obtaining the sub-disks of the target virtual machine according to the snapshots created for the target virtual machine may include the following steps: creating multiple snapshots for the target virtual machine, and obtaining multiple sub-disks of the target virtual machine according to the multiple snapshots.
[0080] In practical applications, multiple snapshots can be created for the target virtual machine by using the snapshots of the virtual machine. For each created snapshot, the corresponding sub-disks of the target virtual machine can be obtained according to the snapshot, and thus multiple sub-disks are obtained. The process of obtaining multiple sub-disks according to multiple snapshots can refer to the description in step S102 above, and will not be elaborated here.
[0081] Correspondingly, in step S103, writing the multiple data shards corresponding to the backup data into the parent disk and the sub-disks in parallel respectively includes: writing the multiple data shards corresponding to the backup data into the parent disk and multiple sub-disks in parallel respectively; the data shards written into the parent disk and multiple sub-disks are not repeated.
[0082] Although the shard recovery method in the related art can divide the backup data into blocks to improve the recovery efficiency, since multi-threaded recovery cannot be truly executed in parallel, and multiple threads may compete for writing when operating on different data blocks, resulting in data corruption or recovery failure, affecting the consistency of the recovered data and further reducing the reliability of the recovery.
[0083] In response to this, in this embodiment, the backup data can be divided into several parts to obtain multiple data shards, and the multiple data shards together constitute the complete backup data; then, the multiple data shards can be written into a parent disk and multiple sub-disks in parallel respectively, and the data shards written into each disk are different from those written into other disks.
[0084] In this embodiment, on the one hand, by obtaining multiple sub-disks according to multiple snapshots, more additional data writing channels can be provided for the original disk of the target virtual machine, so that more data can be written into the disk of the target virtual machine at the same time during the data recovery process; on the other hand, by dividing the backup data into multiple data shards and writing them into the parent disk and sub-disks in parallel, it can be ensured that the writing operations of different data shards in different parent disks or sub-disks will not affect each other, avoiding data corruption or recovery failure, and improving the reliability and integrity of data recovery.
[0085] In one embodiment, writing the multiple data shards corresponding to the backup data into the parent disk and the multiple sub-disks in parallel respectively may include the following steps:
[0086] Determine the number of disks corresponding to multiple sub-disks and the parent disk; call multiple idle threads corresponding to the number of disks in a pre-set thread pool, and use the multiple idle threads in parallel to write multiple data shards corresponding to the backup data into the parent disk and multiple sub-disks respectively in parallel.
[0087] Among them, the number of threads of the multiple idle threads is equal to the number of disks.
[0088] In practical applications, after obtaining multiple sub-disks, the number of multiple sub-disks and the parent disk can be counted to obtain the disk number reflecting the total disk volume information of the multiple sub-disks and the parent disk. Then, multiple idle threads corresponding to the number of disks can be called in a pre-set thread pool. For example, if the total number of disks of the sub-disks and the parent disk is m (m is a positive integer greater than or equal to 3), then m idle threads can be called from the thread pool. Furthermore, multiple idle threads can be used to write multiple data shards corresponding to the backup data into the parent disk and multiple sub-disks respectively in parallel.
[0089] In this embodiment, by calling multiple idle threads matching the number of disks in the thread pool to manage the parallel recovery task, multi-threaded parallel recovery can be achieved, ensuring that the write operations of different data shards do not affect each other, avoiding data corruption or recovery failure, and improving the reliability and integrity of data recovery; at the same time, the write operations of multiple data shards on the virtual machine disk can be processed simultaneously using multi-threaded technology, significantly increasing the amount of data written per unit time, improving the recovery speed, and helping to significantly improve the data recovery efficiency of the virtual machine in the large data volume recovery scenario.
[0090] In one embodiment, in step S104, according to the parent disk and sub-disks with data shards backed up, obtaining the target virtual machine that has completed data recovery may include the following steps:
[0091] Associate the parent disk with data shards backed up and the sub-disks, and obtain the target virtual machine that has completed data recovery according to the associated parent disk and sub-disks; the target virtual machine is used to find the target data from the complete backup data when receiving a target data query request.
[0092] In this embodiment, after the data shards are separately backed up to the parent disk and the child disks, the parent disk and each child disk in the target virtual machine can be automatically associated. In some embodiments, the association between the parent disk and the child disks can be achieved through Snapshot Consolidation. For example, after the target virtual machine is started, the user can choose to delete the snapshot. After choosing to delete the snapshot, the data merge will be automatically triggered, and the data changes in the child disks will be merged back to the parent disk to ensure the consistency of the final disk state. Snapshot Consolidation can be understood as a process of merging the data of the snapshot file (such as the virtual disk file of the child disk) back to the main disk file (such as the virtual disk file of the parent disk). After merging the snapshots, the storage space occupancy can be reduced, and the performance of the target virtual machine can be closer to the state when using a single disk file. Exemplarily, if there are multiple child disks, the multiple child disks can be merged first, and finally the merged multiple child disks can be merged into the parent disk.
[0093] Furthermore, when the target virtual machine receives a target data query request from the system or other devices, it can, according to this request, search for the target data from the complete backup data obtained after the disk association.
[0094] In this embodiment, by automatically associating the parent disk and the child disks that back up the data shards, the correctness, consistency, and integrity of the final recovery result (including the backup data recovered on the target virtual machine) can be effectively ensured.
[0095] In the related art, when dealing with complex scenarios (such as parallel recovery of multiple disks), it is often necessary for the user to manually manage the logic of multi-threading and shard recovery, and also manually merge the snapshots after recovery. This not only increases the code complexity but also may lead to errors during the recovery process, increasing the difficulty of maintenance and debugging, and further increasing the complexity of the operation.
[0096] In view of this, in one embodiment, associating the parent disk that backs up the data shards and the child disks may include the following steps:
[0097] Obtain the parent disk identifier corresponding to the parent disk that backs up the data shards; modify the attribute information of the child disk that backs up the data shards under the parent disk identifier attribute to the parent disk identifier.
[0098] In some embodiments, each time a disk is opened and modified, a new random string will be generated. This random string can be used to verify the integrity and legality of the disk chain, that is, when the data in the disk is modified next time, a new random string will be generated again. In this embodiment, the random string obtained after backing up the data shards to the parent disk can be used as the disk identifier (content ID), and thus the parent disk identifier corresponding to the parent disk can be obtained.
[0099] On the other hand, a disk may have a parent disk identification attribute (parentCID). The parent disk identification attribute can be used to indicate the parent disk associated with the current disk. For the parent disk, since the parent disk has no parent disk, the parent disk identification attribute of the parent disk can take a fixed value, such as ffffffff. In this embodiment, after obtaining the parent disk identification, the attribute information of each sub-disk with a data shard backup under the parent disk identification attribute can be automatically and uniformly modified to the obtained parent disk identification.
[0100] In this embodiment, after completing the data shard recovery for the sub-disk and the parent disk, by automatically updating the association information of the sub-disk (i.e., the attribute information of the sub-disk under the parent disk identification attribute), automated snapshot management and sub-disk description information update are achieved, ensuring that it can be correctly associated with the parent disk, guaranteeing data consistency and integrity when the virtual machine starts, being able to correctly read the complete data chain, avoiding data inconsistency problems that may occur in traditional recovery methods, and ensuring the integrity of data recovery; at the same time, it also simplifies the operation steps in the recovery process, reduces the complexity of development and maintenance, and effectively improves the existing virtual machine disk recovery technology.
[0101] To enable those skilled in the art to better understand the above steps, the following provides an exemplary illustration of the embodiments of the present application through an example, but it should be understood that the embodiments of the present application are not limited thereto.
[0102] In this example, as Figure 3 shown, the server can perform the following steps:
[0103] S301, determine the target virtual machine for which data recovery is to be performed.
[0104] S302, create a parent disk corresponding to the target virtual machine, create multiple snapshots for the target virtual machine, and obtain multiple sub-disks of the target virtual machine according to the multiple snapshots.
[0105] S303, open each sub-disk in the write state and add an association blocking identifier to each opened sub-disk. In the case where each sub-disk has an association blocking identifier, open the parent disk in the write state.
[0106] S304, determine the number of disks m corresponding to the multiple sub-disks and the parent disk, call m idle threads corresponding to the number of disks m in a pre-set thread pool, and parallelly use the m idle threads to write the multiple data shards corresponding to the backup data into the parent disk and the multiple sub-disks respectively.
[0107] S305, obtain the parent disk identification corresponding to the parent disk with a data shard backup, and modify the attribute information of each sub-disk with a data shard backup under the parent disk identification attribute to the parent disk identification.
[0108] It should be understood that although the steps in the flowcharts involved in the above-described embodiments are shown in sequence according to the indications of the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless there is a clear indication in this article, the execution of these steps has no strict order limitation, and these steps can be executed in other orders. Moreover, at least a part of the steps in the flowcharts involved in the above-described embodiments may include multiple steps or multiple stages. These steps or stages are not necessarily executed at the same moment, but can be executed at different moments. The execution order of these steps or stages is not necessarily sequential, but can be executed alternately or in turn with at least a part of other steps or steps or stages in other steps.
[0109] Based on the same inventive concept, an embodiment of the present application also provides a virtual machine data recovery device for implementing the virtual machine data recovery method described above. The implementation solutions for solving problems provided by this device are similar to the implementation solutions described in the above method. Therefore, the specific limitations in one or more embodiments of the virtual machine data recovery device provided below can refer to the limitations on the virtual machine data recovery method in the above text, and will not be repeated here.
[0110] In an exemplary embodiment, as Figure 4 shown, a virtual machine data recovery device is provided, including:
[0111] A virtual machine determination module 401, configured to determine a target virtual machine for which data recovery is to be performed;
[0112] A snapshot creation module 402, configured to create a parent disk corresponding to the target virtual machine, and obtain a child disk of the target virtual machine according to the snapshot created for the target virtual machine;
[0113] A parallel writing module 403, configured to open the parent disk and the child disk in a write state, and respectively and parallelly write multiple data shards corresponding to backup data into the parent disk and the child disk to obtain the parent disk and the child disk that have backed up the data shards;
[0114] An arrangement module 404, configured to obtain the target virtual machine for which data recovery is completed according to the parent disk and the child disk that have backed up the data shards.
[0115] In an embodiment, the parallel writing module 403 is configured to:
[0116] Open the child disk in a write state, and add an associated blocking identifier to the opened child disk; the associated blocking identifier is used to indicate that when the child disk is opened in a write state, locking of the parent disk corresponding to the child disk is blocked;
[0117] When the associated blocking identifier is added to the sub-disk, the parent disk is opened in a write state.
[0118] In one embodiment, the snapshot creation module 402 is configured to:
[0119] Create multiple snapshots for the target virtual machine, and obtain multiple sub-disks of the target virtual machine according to the multiple snapshots;
[0120] The parallel writing module 403 is configured to:
[0121] Parallelly write multiple data shards corresponding to the backup data into the parent disk and the multiple sub-disks respectively; the data shards written in the parent disk and the multiple sub-disks do not repeat.
[0122] In one embodiment, the parallel writing module 403 is configured to:
[0123] Determine the number of disks corresponding to the multiple sub-disks and the parent disk;
[0124] Call multiple idle threads corresponding to the number of disks in a pre-set thread pool, and parallelly use the multiple idle threads to parallelly write multiple data shards corresponding to the backup data into the parent disk and the multiple sub-disks respectively; the number of threads of the multiple idle threads is equal to the number of disks.
[0125] In one embodiment, the sorting module 404 is configured to:
[0126] Associate the parent disk and the sub-disks that have backed up the data shards, and obtain the target virtual machine that has completed data recovery according to the associated parent disk and sub-disks; the target virtual machine is used to find target data from the complete backup data when receiving a target data query request.
[0127] In one embodiment, the sorting module 404 is configured to:
[0128] Obtain the parent disk identifier corresponding to the parent disk that has backed up the data shards;
[0129] Modify the attribute information of the sub-disk that has backed up the data shards under the parent disk identifier attribute to the parent disk identifier.
[0130] Each module in the virtual machine data recovery device described above can be implemented in whole or in part by software, hardware, or a combination thereof. Each of the above modules can be embedded in the processor of the computer device in hardware form or be independent of it, or can be stored in the memory of the computer device in software form, so that the processor can call and execute the operations corresponding to each of the above modules.
[0131] In an exemplary embodiment, a computer device is provided. The computer device can be a server, and its internal structure diagram can be as Figure 5 shown. The computer device includes a processor, a memory, an input / output interface (Input / Output, abbreviated as I / O), and a communication interface. Among them, the processor, the memory, and the input / output interface are connected through a system bus, and the communication interface is connected to the system bus through the input / output interface. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store backup data. The input / output interface of the computer device is used to exchange information between the processor and external devices. The communication interface of the computer device is used to communicate with external terminals through a network connection. The computer program, when executed by the processor, implements a virtual machine data recovery method.
[0132] Those skilled in the art can understand that Figure 5 the structure shown in
[0133] is only a block diagram of some structures related to the solution of this application, and does not constitute a limitation on the computer device to which the solution of this application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine some components, or have a different component layout.
[0134] In an embodiment, a computer device is provided, including a memory and a processor. A computer program is stored in the memory, and when the processor executes the computer program, the steps in the above method embodiments are implemented.
[0134] In an embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by the processor, the steps in the above method embodiments are implemented.
[0135] In an embodiment, a computer program product is provided, including a computer program. When the computer program is executed by the processor, the steps in the above method embodiments are implemented.
[0136] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in this application are all information and data that have been authorized by the user or fully authorized by all parties, and the collection, use, and processing of relevant data need to comply with relevant regulations.
[0137] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, database, or other medium used in the embodiments provided in this application can include at least one of non-volatile memory and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetoresistive random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc. The databases involved in the embodiments provided in this application can include at least one of relational databases and non-relational databases. Non-relational databases can include distributed databases based on blockchain, etc., without limitation. The processors involved in the embodiments provided in this application can be general-purpose processors, central processors, graphics processors, digital signal processors, programmable logic devices, data processing logics based on quantum computing, artificial intelligence (AI) processors, etc., without limitation.
[0138] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope recorded in this application.
[0139] The above-described embodiments merely represent several implementation manners of this application. The description thereof is relatively specific and detailed, but it should not be construed as a limitation on the scope of the patent of this application. It should be noted that for those of ordinary skill in the art, without departing from the concept of this application, several deformations and improvements can still be made, and these all belong to the protection scope of this application. Therefore, the protection scope of this application shall be subject to the appended claims.
Claims
1. A virtual machine data recovery method, characterized in that, The method includes: Determine a target virtual machine for which data recovery is to be performed; Create a parent disk corresponding to the target virtual machine, and obtain a child disk of the target virtual machine according to a snapshot created for the target virtual machine; Open the parent disk and the child disk in a write state, and respectively and parallelly write multiple data shards corresponding to backup data into the parent disk and the child disk to obtain the parent disk and the child disk with the data shards backed up; Associate the parent disk and the child disk with the data shards backed up through snapshot merging, and obtain a target virtual machine with data recovery completed according to the associated parent disk and child disk; the snapshot merging is used to merge the virtual disk file of the child disk back into the virtual disk file of the parent disk; the target virtual machine is used to search for target data from the complete backup data when receiving a target data query request.
2. The method according to claim 1, characterized in that, The opening the parent disk and the child disk in a write state includes: Open the child disk in a write state, and add an association blocking identifier to the opened child disk; the association blocking identifier is used to indicate blocking the locking of the parent disk corresponding to the child disk when the child disk is opened in a write state; When the association blocking identifier is added to the child disk, open the parent disk in a write state.
3. The method according to claim 1, characterized in that, The obtaining the child disk of the target virtual machine according to a snapshot created for the target virtual machine includes: Create multiple snapshots for the target virtual machine, and obtain multiple child disks of the target virtual machine according to the multiple snapshots; The respectively and parallelly writing multiple data shards corresponding to backup data into the parent disk and the child disk includes: Respectively and parallelly write multiple data shards corresponding to backup data into the parent disk and the multiple child disks; the data shards written into the parent disk and the multiple child disks are not repeated.
4. The method according to claim 3, wherein The respectively and parallelly writing multiple data shards corresponding to backup data into the parent disk and the multiple child disks includes: Determine the number of disks corresponding to the multiple child disks and the parent disk; Call multiple idle threads corresponding to the number of disks in a pre-set thread pool, and respectively and parallelly write multiple data shards corresponding to backup data into the parent disk and the multiple child disks by using the multiple idle threads in parallel; the number of threads of the multiple idle threads is equal to the number of disks.
5. The method according to claim 1, wherein The associating the parent disk and the child disk with the data shards backed up includes: Obtain a parent disk identifier corresponding to the parent disk with the data shards backed up; Modify the attribute information of the child disk with the data shards backed up under the parent disk identifier attribute to the parent disk identifier.
6. A virtual machine data recovery device, characterized in that, The apparatus includes: A virtual machine determination module, configured to determine a target virtual machine for which data recovery is to be performed; A snapshot creation module, configured to create a parent disk corresponding to the target virtual machine, and obtain a child disk of the target virtual machine according to a snapshot created for the target virtual machine; A parallel writing module, configured to turn on the parent disk and the child disk in a write state, and respectively and parallelly write multiple data shards corresponding to backup data into the parent disk and the child disk, so as to obtain the parent disk and the child disk with the data shards backed up; An organizing module, configured to associate the parent disk and the child disk with the data shards backed up through snapshot merging, and obtain a target virtual machine with data recovery completed according to the associated parent disk and child disk; the snapshot merging is used to merge the virtual disk file of the child disk back into the virtual disk file of the parent disk; the target virtual machine is configured to search for target data from the complete backup data when receiving a target data query request.
7. A computer device, comprising a memory and a processor, the memory storing a computer program, characterized in that, When the processor executes the computer program, the steps of the method according to any one of claims 1 to 5 are implemented.
8. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 5 are implemented.
9. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 5 are implemented.
Citation Information
Patent Citations
Backup method and device of virtual machine
CN117032884A