Efficiently provide virtual machine reference points

By introducing the virtual machine reference point mechanism, the problems of waste of virtual machine snapshot resources and performance impacts are solved, efficient virtual machine status tracking and data backup are realized, and the resource utilization efficiency and application performance of virtual machines are improved.

CN112988326BActive Publication Date: 2025-07-18MICROSOFT TECHNOLOGY LICENSING LLC
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202110227269.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2014-12-17
Filing Date
2015-12-03
Publication Date
2025-07-18
Estimated Expiration
2035-12-03

AI Technical Summary

Technical Problem

The prior art has problems with wasted resource usage and performance impacts when maintaining point-in-time snapshots of virtual machines, especially when tracking changes in virtual machine storage devices, resulting in an increase in I/O throughput.

Method used

Using the virtual machine reference point (VM reference point) mechanism, by generating stable checkpoints and converting them to include only data storage identifiers, releasing resource overhead, recording only the serial number to track incremental changes, reducing the need for precise image of storage state.

Benefits of technology

It improves the resource utilization efficiency of virtual machines, reduces overhead for I/O operations, improves application performance, and supports efficient data backup and recovery.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN112988326B_ABST
    Figure CN112988326B_ABST
Patent Text Reader

Abstract

The embodiments are directed to establishing an efficient virtual machine reference point and specifying a virtual machine reference point to query incremental changes. In one scenario, a computer system accesses a stable virtual machine checkpoint that includes a portion of underlying data stored in a data storage device, where the checkpoint is associated with a particular point in time. The computer system then queries the data storage device to determine a data storage identifier that references the point in time associated with the checkpoint and stores the determined data storage identifier as the virtual machine reference point, where each subsequent change to the data storage device results in an update to the data storage identifier such that the virtual machine reference point can be used to identify incremental changes starting from the particular point in time.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Related Applications

[0002] This application is a divisional application of the invention patent application with international application number PCT / US2015 / 063565, international filing date December 3, 2015, entering the Chinese national phase on June 19, 2017, Chinese national application number 201580069480.X, and invention title "Efficiently Providing Virtual Machine Reference Points". Background Art

[0003] Computing systems have become ubiquitous, ranging from small embedded devices to mobile phones and tablet computers to PCs and backend servers. Each of these computing systems is designed to process software code. Software allows users to perform functions and thereby interact with the hardware provided by the computing system. In some cases, these computing systems allow users to create and run virtual machines. These virtual machines can provide functionality not provided by the host operating system or can collectively include different operating systems. In this way, virtual machines can be used to extend the functionality of the computing system. Virtual machines can be backed up on virtual storage devices, which themselves can be backed up to physical or virtual storage devices. The virtual machine host can also be configured to take snapshots, which represent a point-in-time image of the virtual machine. A VM snapshot or "checkpoint" includes the CPU state, memory state, storage state, and other information necessary to fully recreate or restore the virtual machine to that point in time. Summary of the Invention

[0004] Embodiments described herein are directed to establishing efficient virtual machine reference points and specifying virtual machine reference points to query incremental changes. As used herein, a virtual machine reference point allows a computer system to identify incremental changes starting from a particular point in time. For example, in one embodiment, a computer system accesses a stable virtual machine checkpoint that includes a portion of the underlying data stored in a data storage device, where the checkpoint is associated with a particular point in time. The computer system then queries the data storage device to determine the data storage identifier that references the data storage at the time point associated with the checkpoint and stores the determined data storage identifier as a virtual machine reference point or a virtual machine reference point artifact, where each subsequent change to the data storage device results in an update to the data storage identifier such that the virtual machine reference point can be used to identify incremental changes starting from a particular point in time. The virtual machine reference point artifact allows for the case where a virtual machine has two (or more) virtual disks. Each virtual disk can have a different identifier for the same point in time, and the reference point artifact allows the computer system to associate those two time points as a common point. This will be further explained below.

[0005] In another embodiment, a computer system performs a method for specifying a virtual machine reference point for querying incremental changes. The computer system establishes a stable, non-changing state within the virtual machine, where the stable state is associated with a checkpoint including corresponding state data and stored data. The computer system accesses a previously generated reference point to identify differences in the virtual machine state between the current stable state and a selected past stable time point. The computer system also copies the differences between the current stable state and the virtual machine state at the selected past stable time point. The differences can be copied as an incremental backup to a data storage device, or can be used for remote replication or disaster recovery purposes.

[0006] This Summary is provided to introduce a selection of concepts in a simplified form that are further described below in the Detailed Description. This Summary is not intended to identify key features or essential features of the claimed subject matter, nor is it intended to be used as an aid in determining the scope of the claimed subject matter.

[0007] Additional features and advantages will be set forth in the description which follows, and in part will be obvious from the description, or can be learned by practice of the teachings herein. The features and advantages of the embodiments described herein can be realized and obtained by means of the instrumentalities and combinations particularly pointed out in the appended claims. The features of the embodiments described herein will be more fully apparent from the following description and appended claims. BRIEF DESCRIPTION OF THE DRAWINGS

[0008] To further clarify the above and other features of the embodiments described herein, a more particular description will be rendered by reference to the accompanying drawings. It is to be appreciated that these drawings merely illustrate examples of the embodiments described herein and are therefore not to be considered limiting of their scope. The embodiments will be described and explained with additional feature and details by using the accompanying drawings, in which:

[0009] Figure 1 Illustrates a computer architecture in which the embodiments described herein can operate, the operations including establishing an efficient virtual machine reference point.

[0010] Figure 2 Illustrates a flowchart of an example method for establishing an efficient virtual machine reference point.

[0011] Figure 3 Illustrates a flowchart of an example method for specifying a virtual machine reference point for querying incremental changes.

[0012] Figure 4 Illustrates a computer architecture in which the embodiments can operate, the operations including specifying a virtual machine reference point for querying incremental changes. DETAILED DESCRIPTION

[0013] The embodiments described herein are directed to establishing an efficient virtual machine reference point and specifying a virtual machine reference point to query incremental changes. In one embodiment, a computer system accesses a stable virtual machine checkpoint that includes a portion of underlying data stored in a data storage device, where the checkpoint is associated with a specific point in time. The computer system then queries the data storage device to determine a data storage identifier that references the data stored at the time associated with the checkpoint, and stores the determined data storage identifier as the virtual machine reference point, where each subsequent change to the data storage device results in an update to the data storage identifier such that the virtual machine reference point can be used to identify incremental changes starting from the specific point in time.

[0014] In another embodiment, a computer system performs a method for specifying a virtual machine reference point to query incremental changes. The computer system establishes a stable, non-changing state within the virtual machine, where the stable state is associated with a checkpoint that includes corresponding state data and stored data. The computer system accesses a previously generated reference point to identify differences in the virtual machine state between the current stable state and a selected past stable point in time. The computer system then copies the differences in the virtual machine state between the current stable state and the selected past stable point in time. The differences can be copied to a data storage device as an incremental backup, or can be used for remote replication or disaster recovery purposes.

[0015] The following discussion now refers to several methods and method acts that can be performed. It should be noted that although method acts may be discussed in a certain order or illustrated in a flowchart as occurring in a particular order, a particular ordering is not necessarily required unless specifically stated, or because one act depends on another act being completed before that act is performed.

[0016] The embodiments described herein can be implemented in various types of computing systems. These computing systems are now increasingly taking on various forms. A computing system can be, for example, a handheld device such as a smart phone or a feature phone, a household appliance, a laptop computer, a wearable device, a desktop computer, a mainframe, a distributed computing system, or even a device that is not conventionally considered a computing system. In this description and in the claims, the term "computing system" is broadly defined to include any device or system (or combination thereof) that includes at least one physical and tangible processor, and a physical and tangible memory capable of having computer-executable instructions thereon that can be executed by the processor. The computing system can be distributed across a network environment and can include multiple constituent computing systems.

[0017] As in Figure 1As shown in the figure, computing system 101 typically includes at least one processing unit 102 and a memory 103. The memory 103 can be a physical system memory, which can be volatile, non-volatile, or some combination of the two. The term "memory" may also be used herein to refer to non-volatile mass storage devices (such as physical storage media). If the computing system is distributed, the processor and / or storage capabilities may also be distributed.

[0018] As used herein, the terms "executable module" or "executable component" may refer to software objects, routines, or methods that can be executed on a computing system. The different components, modules, engines, and services described herein may be implemented as objects or processes (e.g., as separate threads) that execute on a computing system.

[0019] In the following description, embodiments are described with reference to actions performed by one or more computing systems. If such actions are implemented in software, one or more processors of the associated computing system that executes the actions direct the operation of the computing system in response to having executed computer-executable instructions. For example, such computer-executable instructions may be embodied on one or more computer-readable media that form a computer program product. Examples of such operations involve the manipulation of data. The computer-executable instructions (and the data being manipulated) may be stored in the memory 103 of the computing system 101. The computing system 101 may also include a communication channel that allows the computing system 101 to communicate with other message processors via a wired or wireless network.

[0020] The embodiments described herein may include or utilize a special-purpose or general-purpose computer system including computer hardware, as discussed in more detail below, such as, for example, one or more processors and system memory. The system memory may be included within the total memory 103. The system memory may also be referred to as "main memory" and includes memory locations that are addressable by the at least one processing unit 102 via a memory bus, in which case the address locations are asserted on the memory bus itself. System memory has traditionally been volatile, but the principles described herein also apply to cases where the system memory is partially or even completely non-volatile.

[0021] Embodiments within the scope of the present invention also include physical and other computer-readable media for carrying or storing computer-executable instructions and / or data structures. Such computer-readable media can be any available media that can be accessed by a general or special-purpose computer system. A computer-readable media that stores computer-executable instructions and / or data structures is a computer storage media. A computer-readable media that carries computer-executable instructions and / or data structures is a transmission media. Thus, by way of example and not limitation, embodiments of the present invention can include at least two distinctly different types of computer-readable media: computer storage media and transmission media.

[0022] Computer storage media is physical hardware storage media that stores computer-executable instructions and / or data structures. Physical hardware storage media includes computer hardware such as RAM, ROM, EEPROM, solid state drives (“SSD”), flash memory, phase change memory (“PCM”), optical disk storage devices, magnetic disk storage devices, or any other hardware storage device(s) that can be used to store program code in the form of computer-executable instructions or data structures, and that can be accessed and executed by a general or special-purpose computer system to implement the disclosed functionality of the present invention.

[0023] Transmission media can include a network and / or a data link that can be used to carry program code in the form of computer-executable instructions or data structures, and that can be accessed by a general or special-purpose computer system. A “network” is defined as one or more data links that enable the transport of electronic data between computer systems and / or modules and / or other electronic devices. When information is transferred or provided to a computer system via a network or another communication connection (wired, wireless, or a combination of wired or wireless), the computer system can consider the connection to be a transmission media. Combinations of the above should also be included within the scope of computer-readable media.

[0024] In addition, when program code in the form of computer-executable instructions or data structures arrives at various computer system components, it can automatically transfer from the transmission media to the computer storage media (or vice versa). For example, computer-executable instructions or data structures received via a network or data link can be cached in RAM within a network interface module (e.g., “NIC”), and then ultimately transferred to the computer system RAM and / or a less volatile computer storage media at the computer system. Thus, it should be understood that computer storage media can be included in computer system components that also (or even primarily) utilize transmission media.

[0025] Computer-executable instructions include, for example, instructions and data that, when executed at one or more processors, cause a general-purpose computer system, a special-purpose computer system, or a special-purpose processing device to implement a certain function or a set of functions. Computer-executable instructions can be, for example, binary numbers, intermediate format instructions (such as assembly language or even source code).

[0026] Those skilled in the art will appreciate that the principles described herein can be practiced in a network computing environment having many types of computer system configurations, including personal computers, desktop computers, laptop computers, messaging processors, handheld devices, multiprocessor systems, microprocessor-based or programmable consumer electronics, network PCs, minicomputers, mainframe computers, mobile phones, PDAs, tablet computers, pagers, routers, switches, etc. The present invention can also be practiced in a distributed computing environment where both local and remote computing systems perform tasks, the local and remote computing systems being linked by a network (by a hardwired data link, a wireless data link, or a combination of hardwired and wireless data links). Thus, in a distributed computing environment, a computer system can include multiple constituent computer systems. In a distributed system environment, program modules can be located in both local and remote memory storage devices.

[0027] Those skilled in the art will also appreciate that the present invention can be practiced in a cloud computing environment. A cloud computing environment can be distributed (although this is not required). When it is distributed, a cloud computing environment can be distributed internationally within an organization and / or have components owned across multiple organizations. In this description and the following claims, "cloud computing" is defined as a model for enabling on-demand network access to a shared pool of configurable computing resources (e.g., networks, servers, storage devices, applications, and services). The definition of "cloud computing" is not limited to any of the many other advantages that can be obtained from such a model when properly deployed.

[0028] Still further, the system architectures described herein can include multiple independent components, each contributing to the functionality of the system as a whole. This modularity allows for increased flexibility in the face of platform scalability issues and, for this purpose, provides a variety of advantages. System complexity and growth can be more easily managed by using smaller-scale parts within a limited functional scope. Platform fault tolerance is enhanced by using these loosely coupled modules. Individual components can grow incrementally as dictated by business needs. Modular development also translates into reduced time to market for new functionality. New functionality can be added or subtracted without affecting the core system.

[0029] Figure 1FIG. illustrates a computer architecture 100 in which at least one embodiment may be implemented. The computer architecture 100 includes a computer system 101. The computer system 101 may be any type of local or distributed computer system, including a cloud computing system. The computer system 101 includes modules for performing various different functions. For example, a communication module 104 may be configured to communicate with other computing systems. The communication module 104 may include any wired or wireless communication component that may receive data from and / or transmit data to other computing systems. The communication module 104 may be configured to interact with databases, mobile computing devices (such as mobile phones or tablets), embedded or other types of computing systems.

[0030] Figure 1 The computer system 101 may further include a checkpoint generation module 105. A checkpoint 106 (or "snapshot" herein) generated by the module 105 may include various portions of corresponding checkpoint data 116 that may be stored in a data warehouse 115. The checkpoint data may include, for example, state data 117 and storage data 118. The state data 117 may include CPU state, memory state, and device state for various computer system devices. The storage data 118 may include data files, application files, operating system files, backup files, or other data associated with the checkpoint. Thus, the checkpoint data 116 associated with the checkpoint 106 includes sufficient data to perform a full recovery or backup from the data. However, the storage and state data for a given computer system or virtual machine may include a significant amount of data.

[0031] In some cases, multiple checkpoints may be stored for a single computer system or virtual machine (VM). Each checkpoint may be generated at different time periods. Then, once two or more checkpoints have been generated, they may be compared to each other to determine the differences between them. These differences may then be applied to a differential backup that only backs up the differences between the two checkpoints. The embodiments described herein introduce the concept of a "reference point", "virtual machine reference point", or "VM reference point" herein. The VM reference point allows the deletion of previous stored, memory, and device states associated with a checkpoint while still retaining the ability to create differential backups. This is accomplished by recording sequence numbers for state transitions.

[0032] Incremental backup of a virtual machine involves tracking those changes to the virtual machine's storage device that have occurred since a previously specified VM time point (or time points). Traditionally, a time point snapshot of a VM is represented by a VM checkpoint. As mentioned above, storing a full checkpoint can introduce a significant overhead on the VM's I / O throughput because copy-on-write (or similar) techniques are typically used to maintain the time point image of the virtual storage device. Maintaining a VM checkpoint solely for reference to a previous time point for querying incremental changes is wasteful in terms of resource usage and adversely affects the performance of applications running in the VM.

[0033] The VM reference points described herein do not require maintaining an exact time point image of the VM state (e.g., storage, memory, device state). A VM reference point provides a representation of a previous moment in the VM. The reference point can be represented by a unique identifier (which can be a globally unique identifier (GUID)), a sequence number, or other identifier.

[0034] In one embodiment, a VM reference point (e.g., Figure 1 110) is generated in the following manner: First, a time point image (checkpoint 106) is generated for the VM. This provides a stable copy from which to backup. Along with the creation of this image, a mechanism for tracking changes to the virtual storage device is triggered. The trigger mechanism can be based on user interaction, input from an application or other source, or can be manually triggered. Second, once the computer system or VM has been backed up, the checkpoint is converted / downgraded to a reference point (i.e., just a representation of the time point without the backup of the corresponding machine state). This releases the overhead associated with the checkpoint while allowing tracking of the last backup time point. Third, during a subsequent backup, the user can specify the reference point to query incremental changes to the VM since the specified time point.

[0035] The VM reference point 110 includes minimal metadata that enables querying of incremental changes to the virtual storage device since the time point represented by the free reference point. Example change tracking techniques involve a virtual storage subsystem for maintaining a list of changed blocks across discrete time points represented by sequence numbers. In such a system, the VM reference point for the VM will include only the sequence number of the virtual storage device corresponding to that discrete time point. The sequence number will increment each time a memory register changes. Since the sequence number corresponds to the time point for the checkpoint, the sequence number increases for each memory block written to the disk. In at least some embodiments, the sequence number is the only content stored in the VM reference point.

[0036] When converting a VM checkpoint (e.g., 106) to a VM reference point (e.g., 110), the system releases all (or substantially all) resources that have been used to maintain the time - point image of the VM checkpoint. For example, in one scenario, the system can release the differential virtual hard disk that has been used to maintain the time - point image of the checkpoint, thereby eliminating the overhead of performing I / O on the differential virtual hard disk (VHD). The VM state data of this overhead is replaced by metadata about the reference point. The reference - point metadata only contains identifiers (e.g., serial numbers) corresponding to the time - point for the virtual hard disk. This enables the subsequent use of the reference point to query for incremental changes since that time - point.

[0037] Thus, in Figure 1 once the checkpoint generation module 105 has generated a checkpoint 106 for a physical or virtual computer system, the query generation module 108 generates a query 113 to identify which data - storage identifiers 114 reference the time - point associated with the checkpoint 106. These data - storage identifiers 114 are sent to the computer system 101 and are implemented by the VM reference - point generation module 109 to generate a VM reference point 110. It should be understood here that the data warehouse 115 can be inside or outside the computer system 101, and can include a single storage device (e.g., hard disk or optical disk) or can include many storage devices, and can actually include a storage network (such as a cloud - storage network). Thus, the data transfer for the query 113 and the data - storage identifiers 114 can be internal (e.g., via a hardware bus) or can be external via a computer network.

[0038] The VM reference point 110 thus includes those data - storage identifiers that point to a specific time - point. Each subsequent change to the data - storage device results in an update to the data - storage identifier. Thus, the virtual - machine reference point 110 can be used to identify incremental changes starting from a specific time - point. These concepts will be further explained below with respect to Figure 2 and 3 the methods 200 and 300 and the embodiments illustrated in Figure 4 respectively.

[0039] Given the systems and architectures described above, the methods that can be implemented in accordance with the disclosed subject matter can be better understood with reference to the Figure 2 and 3 flowcharts. For purposes of simplicity of explanation, the methods are shown and described as a sequence of blocks. However, it should be understood and appreciated that the claimed subject matter is not limited by the order of the blocks, as some blocks can occur in a different order and / or concurrently with other blocks than those depicted and described herein. In addition, not all of the illustrated blocks may be required to implement the methods described hereinafter.

[0040] Figure 2The flowchart of method 200 for establishing an efficient virtual machine reference point is illustrated. Method 200 will now be described with frequent reference to the components and data of environment 100.

[0041] Method 200 includes accessing a stable virtual machine checkpoint, which includes one or more portions of underlying data stored in a data storage device, and the checkpoint is associated with a specific point in time (210). For example, the checkpoint access module 107 can access checkpoint 106. Checkpoint 106 can be generated based on a running computer system or virtual machine. The checkpoint can include operating system files, application files, register files, data files, or any other type of data (including data currently stored in RAM or other memory areas). Checkpoint 106 thus has different state data 117 and stored data 118 in its underlying checkpoint data 116. Checkpoint data 116 can be stored in data warehouse 115 or some other data warehouse. The data warehouse can include an optical storage device, a solid-state storage device, a magnetic storage device, or any other type of data storage hardware.

[0042] Method 200 then includes querying the data storage device to determine one or more data storage identifiers that reference the point in time associated with the checkpoint (220) and storing the determined data storage identifiers as a virtual machine reference point, where each subsequent change to the data storage device results in an update to the data storage identifiers such that the virtual machine reference point can be used to identify incremental changes starting from a specific point in time (230). For example, the query generation module 108 can generate query 113, which queries data warehouse 115 to determine which data storage identifiers reference the point in time associated with checkpoint 106. These data storage identifiers 114 are then stored as virtual machine reference point 110. The VM reference point generation module 109 can thus access an existing checkpoint and degrade or convert it from a checkpoint fully supported by state data 117 and stored data 118 to a VM reference point that only includes identifiers. In some cases, these identifiers can simply be the serial numbers of the storage devices. There can be multiple data storage identifiers for a single disk, or there can be only data storage identifiers for the disk. These concepts can be better understood with reference to Figure 4 the computing architecture 400. However, it will be understood that the example shown in Figure 4 is only one embodiment, and many different embodiments can be implemented.

[0043] Figure 4 The computing system 420 is illustrated, which allows each subsequent change to the data storage device to result in an update to the data storage identifier 409, and thereby allows the virtual machine reference point to be used to identify incremental changes starting from a specific point in time. For example, assumeFigure 4 The original data 401 is stored in different data blocks in a data warehouse (e.g., Figure 4 the data warehouse 408). The original data 401 can be stored locally on the local disk 403 or can be stored on an external data warehouse. This data is backed up (e.g., by the backup generation module 411) in the initial backup 404 and stored as the full backup 410. The full backup 410 thus includes every block of the original data 401. This is similar to creating a checkpoint 106, where all the data 116 of the checkpoint is stored in the data warehouse 115.

[0044] Portions of the original data 401 can currently be in the memory 402 and other portions can be on the disk 403. Since no changes have been made since the full backup 410, the memory and the disk are empty. When changes are made, those changes occur in the original data 401. This can be similar to any change to a data file, an operating system file, or other data on a physical or virtual machine. Memory blocks in the memory 402 that include the updated data can be marked with a "1" or other changed block data identifier 409. Again, it will be understood that although a sequence number is used as the identifier in this example, essentially any type of identifier can be used, including a user-selected name, a GUID, a bitmap, or other identifier. These identifiers are associated with a point in time. Then, using that point in time, a virtual machine reference point can be used to identify incremental changes starting at that point in time.

[0045] If additional changes are made to the original data 401 at a later point in time, the memory 402 as well as the identifiers on the disk will indicate that new changes have occurred. An identifier (e.g., "1") indicates a memory block that was changed at the first point in time, and "2" can indicate that the memory block has been changed at the second point in time. If a memory block includes the "1,2" identifier, this can indicate that the memory block was changed at both the first and second points in time. If a power failure occurs at this point, the data on the disk will most likely be preserved while any content in the memory (e.g., in RAM) will be lost.

[0046] If a backup is performed at this stage, given a power failure, the original data 401 will be backed up as a differential 406 and will include all the data in the changed data blocks at the second time point (as indicated by the identifier "2"). Thus, the differential 406 will include the data changes that have occurred since the full backup, and the complete merged backup 407 will include the differential 406 combined with the full backup 410. The reference point does not need to maintain an exact time point image of the physical or virtual machine state (e.g., storage, memory, device state). It is simply a representation of the previous moment of the physical or virtual machine. Using this VM reference point 412, data backup can be performed, data migration can be performed, data recovery can be performed, and other data-related tasks can be performed.

[0047] In some embodiments, Figure 1 the state management module 111 can be used to establish a stable non-changing state within the virtual machine at the current time. A stable non-changing state means that all appropriate buffers have been emptied, no transactions are pending, and the state is not subject to change. The stable state can be established by performing any one of the following: caching subsequent data changes within the virtual machine, implementing a temporary copy-on-write for subsequent data changes within the virtual machine, and generating a checkpoint for the virtual machine that includes one or more portions of the underlying data.

[0048] This stable state is then associated with the virtual machine reference point 110, which includes a data storage identifier 114 corresponding to a specific stable time point (such as the current time point). The computer system 101 can then access the previously generated virtual machine reference point to identify the differences in the virtual machine state between the current stable state and a selected past stable time point, and use the data identified by the stored data storage identifier and the current data storage identifier 114 to perform at least one operation. As mentioned above, these data-related tasks can include backing up data, restoring data, copying data, and providing data to a user or other specified entity (such as a third-party service or client).

[0049] The data identified by the stored data storage identifier and the current data storage identifier can be combined with the previously generated checkpoint (the initial backup 404), where the data storage identifier "1,2" is used to combine the data identified by the stored data storage identifier and the initial backup 404. In this way, differential backups can be simply provided using the data storage identifier to update the state changes.

[0050] In some embodiments, an application programming interface (API) may be provided that allows third parties to store data storage identifiers as virtual machine reference points. In this way, VM reference point functionality can be extended to third parties in a unified manner. These VM reference points may refer to full backups as well as incremental backups. Using these APIs, multiple vendors can perform data backups simultaneously. Each VM reference point (e.g., 110) includes metadata that includes a data storage identifier 114. Thus, the virtual machine reference point is lightweight and not supported by checkpoint data 116 that includes data storage, memory, or virtual machine state. In some cases, a VM reference point can be converted from a checkpoint and thus can go from having data storage, memory, or virtual machine state as a checkpoint to having only metadata that includes a data storage identifier.

[0051] In some cases, a data-supported checkpoint can be reconstructed using the changes identified between a virtual machine reference point 110 and a future point in time. For example, computer system 101 can use the VM reference point plus a change log (with pure metadata describing what has changed, rather than the data itself) to create a full checkpoint by obtaining data from a service that monitors changes. Even further, at least in some cases, if a virtual machine is migrated, the virtual machine reference point information can be transferred along with the VM. Thus, if a virtual machine moves to a different computing system, the data identified by the virtual machine reference point is recoverable at the new computing system. Accordingly, various embodiments are described in which VM reference points can be created and used to backup data.

[0052] Figure 3 A flowchart of a method 300 for specifying virtual machine reference points to query incremental changes is illustrated. Method 300 will now be described with frequent reference Figure 1 to the components and data of environment 100.

[0053] Method 300 includes establishing a stable non-changing state within a virtual machine, the stable state being associated with a checkpoint that includes corresponding state data and stored data (310). For example, Figure 1 a state management module 111 can establish a stable non-changing state within a virtual machine. The stable state can be associated with a checkpoint 106 that includes corresponding state data 117 and stored data 118. The stable state can be established by caching any subsequent data changes within the virtual machine, by implementing a temporary copy-on-write for subsequent data changes within the virtual machine such that all subsequent data changes are stored, and / or by generating a checkpoint for the virtual machine that includes underlying checkpoint data 116.

[0054] Method 300 also includes accessing one or more previously generated reference points to identify one or more differences in the virtual machine state between the current stable state and a selected past stable time point (320), and copying the differences between the virtual machine state between the current stable state and the selected past stable time point (330). Computer system 101 may be configured to access VM reference point 110 to identify differences in the VM state between the established stable state and another stable time point (identified by the data storage identifier of the VM reference point). The copy module 112 of computer system 101 may then copy the identified differences between the current stable state and the selected past stable time point. These copied differences may form an incremental backup. This incremental backup may be used for remote replication, disaster recovery, or other purposes.

[0055] In some embodiments, computer system 101 may be configured to cache any data changes that occur when determining differences in the virtual machine state. Then, once the differences in the VM state have been determined, the differences may be merged with the cached data into the live virtual machine state, which includes data backed up from the selected time point. This allows the user to select substantially any past stable time point (since the creation of the checkpoint) and determine the state at that point, and merge it with the current state to form a live virtual machine that includes data backed up from the selected time point. The selected past stable time point may thus include any available previous stable time point represented by a virtual machine reference point (which need not be the immediately preceding reference point).

[0056] The differential virtual hard drive may be configured to keep track of data changes that occur when determining differences in the virtual machine state. It should also be noted that the establishment, access, and copy steps (310, 320, and 330) of method 300 may each continue to operate during storage operations, including during the creation of checkpoints. Thus, while a live storage operation is being performed, the embodiments herein may still establish a stable state, access previously generated checkpoints, and copy differences in the VM state between the current state and a selected past stable state.

[0057] In the embodiments described above, reference is made to Figure 4, the sequence number is used as a data storage identifier. In some cases where the sequence number is used in this way, a separate sequence number can be assigned to each physical or virtual disk associated with a virtual machine. Thus, reference points can be created for multiple disks (including asynchronous disks). Elastic change tracking understands the sequence number and can simultaneously create available VM reference points for many disks. Each VM reference point has a sequence number that references the changes made to the disk. For multiple asynchronous disks (e.g., added at different times), each disk has its own sequence (reference) number, but the VM reference point tracks the sequence numbers for all disks of the VM. Any data associated with a checkpoint can be moved to a recovery storage device so that the data does not have to be located on the production server. This reduces the load on the production server and increases its ability to process data more efficiently.

[0058] Claim support: In one embodiment, a computer system is provided that includes at least one processor. The computer system performs a computer-implemented method for establishing an efficient virtual machine reference point, where the method includes the steps of: accessing a stable virtual machine checkpoint 116 that includes one or more portions of underlying data 116 stored in a data storage device 115, the checkpoint being associated with a specific point in time, querying the data storage device to determine one or more data storage identifiers 114 that reference the point in time associated with the checkpoint 106, and storing the determined data storage identifiers as a virtual machine reference point 110, where each subsequent change to the data storage device causes an update to the data storage identifier such that the virtual machine reference point can be used to identify incremental changes starting from the specific point in time.

[0059] The method further includes establishing a stable non-changing state within the virtual machine at the current time, the stable state being associated with the virtual machine reference point, the virtual machine reference point including one or more data storage identifiers corresponding to a specific stable point in time, accessing one or more previously generated virtual machine reference points to identify one or more differences in the virtual machine state between the current stable state and a selected past stable point in time, and performing at least one operation using the data identified by the stored data storage identifier and the current data storage identifier.

[0060] In some cases, the at least one operation includes one or more of the following: backing up data, restoring data, copying data, and providing data to a user or other specified entity. The data identified by the stored data storage identifier and the current data storage identifier is combined with previously generated checkpoints. The stable state is established by performing at least one of the following: caching subsequent data changes within the virtual machine, implementing a temporary copy-on-write for subsequent data changes within the virtual machine, and generating a checkpoint for the virtual machine that includes one or more portions of the underlying data.

[0061] In some cases, one or more provided application programming interfaces (APIs) allow multiple different third parties to store data storage identifiers as virtual machine reference points. The virtual machine reference points include metadata that includes the data storage identifier, such that the virtual machine reference points are lightweight and not supported by checkpoint data that includes data storage, memory, or virtual machine state. A virtual machine checkpoint that includes data storage, memory, or virtual machine state is converted into a virtual machine reference point that separately includes metadata. Additionally, one or more identified changes between the virtual machine reference point and a future time point are used to reconstruct a data-supported checkpoint.

[0062] In another embodiment, a computer system is provided that includes at least one processor. The computer system performs a computer-implemented method for specifying a virtual machine reference point to query incremental changes, where the method includes the steps of: establishing a stable non-changing state within the virtual machine, the stable state being associated with a checkpoint 106 that includes corresponding state data 117 and stored data 118, accessing one or more previously generated reference points 110 to identify one or more differences in the virtual machine state between the current stable state and a selected past stable time point, and copying the differences in the virtual machine state between the current stable state and the selected past stable time point.

[0063] In some cases, the selected past stable time point includes any available previous stable time point represented by a virtual machine reference point. Additionally, a differential virtual hard drive keeps track of data changes that occur when determining differences in the virtual machine state. The data storage identifier includes a sequence number that increments each time the memory register changes.

[0064] In yet another embodiment, a computer system is provided that includes: one or more processors for accessing a checkpoint access module 107 for a stable virtual machine checkpoint 106 that includes one or more portions of underlying data 116 stored in a data storage device 115, the checkpoint being associated with a specific time point, a query generation module 108 for generating a query 113 that queries the data storage device to determine one or more data storage identifiers 114 that reference the time point associated with the checkpoint, and a virtual machine reference point generation module 109 for storing the determined data storage identifiers as virtual machine reference points 110, where each subsequent change to the data storage device results in an update to the data storage identifier such that the virtual machine reference point can be used to identify incremental changes starting from the specific time point. The virtual machine reference point information is transferred along with the virtual machine such that if the virtual machine moves to a new computing system, the data identified by the virtual machine reference point can be restored.

[0065] Accordingly, a method, a system, and a computer program product for establishing an efficient virtual machine reference point are provided. In addition, a method, a system, and a computer program product for specifying a virtual machine reference point for querying incremental changes are provided.

[0066] The concepts and features described herein may be embodied in other specific forms without departing from their spirit or descriptive characteristics. The embodiments are to be considered in all respects only illustrative and not restrictive. The scope of protection of the present disclosure is thus represented by the appended claims rather than the foregoing description. All changes that fall within the meaning and range of equivalents of the claims are to be embraced within their scope. Although attempts have been made in the foregoing description to draw attention to those features regarded as important, it should be understood that the applicant may seek protection by means of the claims for any patentable feature or combination of features referred to and / or shown in the foregoing description and / or in the drawings whether or not emphasis has been placed thereon.

Claims

1. A computer-implemented method for establishing a virtual machine reference point at a computer system, the computer system including at least one processor, the method comprising: Maintaining a plurality of data storage identifiers that identify changed data blocks among a plurality of data blocks of a data storage device corresponding to a virtual machine; Accessing a stable virtual machine checkpoint that includes a recoverable image of the virtual machine at a point in time and stores a representation of data when at least one of the plurality of data blocks of the data storage device existed at the point in time; Converting the virtual machine checkpoint into a virtual machine reference point that includes a representation of the virtual machine at the point in time, wherein the virtual machine reference point information is transferable with the virtual machine such that if the virtual machine is moved to a different computing system, any data identified by the virtual machine reference point is recoverable, the conversion including: Querying the data storage device to determine at least one data storage identifier corresponding to the at least one data block of the virtual machine checkpoint at the point in time; Storing the determined at least one data storage identifier as part of the virtual machine reference point; and Releasing the representation of the data of the at least one data block from the virtual machine checkpoint; and After converting the virtual machine checkpoint into the virtual machine reference point, using the virtual machine reference point to identify one or more changes among the plurality of data blocks of the data storage device since the point in time.

2. The method according to claim 1, further comprising: Establishing a stable unchanging state within the virtual machine, the stable state being associated with the virtual machine reference point; Accessing one or more previously generated virtual machine reference points to identify one or more differences in the virtual machine state between the current stable state and a selected past stable point in time; And Using the data identified by the stored data storage identifiers and current data storage identifiers to perform at least one operation.

3. The method according to claim 2, wherein the at least one operation includes one or more of the following: backing up data, restoring data, copying data, or providing data to a user or other designated entity.

4. The method according to claim 2, wherein the data identified by the stored data storage identifiers and current data storage identifiers is combined with a previously generated checkpoint.

5. The method according to claim 2, wherein the stable state is established by performing at least one of the following: caching subsequent data changes within the virtual machine, implementing a temporary copy-on-write for subsequent data changes within the virtual machine, or generating a checkpoint for the virtual machine that includes one or more portions of underlying data.

6. The method according to claim 1, further comprising providing one or more application programming interfaces (APIs) that allow a plurality of different third parties to store data storage identifiers as virtual machine reference points.

7. The method according to claim 6, wherein the plurality of third parties are supported to perform data backup simultaneously by using the provided one or more APIs.

8. The method according to claim 1, wherein the data storage identifier includes at least one of a serial number or a bitmap.

9. The method according to claim 1, wherein the virtual machine reference point includes metadata, the metadata includes the at least one data storage identifier, such that the virtual machine reference point is lightweight and does not include checkpoint data of the virtual machine checkpoint of the data storage, memory, or virtual machine state.

10. The method according to claim 9, wherein the virtual machine checkpoint including the data storage, memory, and virtual machine state is converted into a virtual machine reference point including only metadata.

11. The method according to claim 1, wherein the data-supported checkpoint is reconstructed using one or more identified changes between the virtual machine reference point and a future time point.

12. A computer system, comprising: one or more processors; and one or more computer-readable media having computer-executable instructions stored thereon, the computer-executable instructions being executable by the one or more processors to cause the computer system to establish a virtual machine reference point, the computer-executable instructions including instructions executable to cause the computer system to at least perform the following: maintain a plurality of data storage identifiers that identify changed data blocks among a plurality of data blocks of a data storage device corresponding to a virtual machine; access a stable virtual machine checkpoint that includes a recoverable image of the virtual machine at a point in time and stores a representation of data when at least one data block among the plurality of data blocks of the data storage device existed at the point in time; convert the virtual machine checkpoint into a virtual machine reference point including a representation of the virtual machine at the point in time, wherein the virtual machine reference point information is transferable with the virtual machine such that if the virtual machine is moved to a different computing system, any data identified by the virtual machine reference point is recoverable, the conversion including: querying the data storage device to determine at least one data storage identifier corresponding to the at least one data block of the virtual machine checkpoint at the point in time; storing the determined at least one data storage identifier as part of the virtual machine reference point; and releasing the representation of the data of the at least one data block from the virtual machine checkpoint; and after converting the virtual machine checkpoint into the virtual machine reference point, using the virtual machine reference point to identify one or more changes among the plurality of data blocks of the data storage device since the point in time.

13. The computer system according to claim 12, wherein data identified by the stored data storage identifier and the current data storage identifier is combined with a previously generated checkpoint.

14. The computer system according to claim 12, wherein the stable state is established by performing at least one of the following: caching subsequent data changes within the virtual machine, implementing a temporary copy-on-write for subsequent data changes within the virtual machine, or generating a checkpoint for the virtual machine that includes one or more portions of underlying data.

15. The computer system according to claim 12, wherein the data storage identifier includes at least one of a sequence number or a bitmap.

16. The computer system according to claim 12, wherein the virtual machine reference point includes metadata that includes the at least one data storage identifier, such that the virtual machine reference point is lightweight and does not include the checkpoint data of the virtual machine checkpoint that includes data storage, memory, or virtual machine state.

17. The computer system according to claim 16, wherein a virtual machine checkpoint that includes data storage, memory, and virtual machine state is converted into a virtual machine reference point that includes only metadata.

18. The computer system according to claim 12, wherein a data-supported checkpoint is reconstructed using one or more identified changes between the virtual machine reference point and a future point in time.

19. The computer system according to claim 12, wherein the at least one operation includes one or more of the following: backing up data, restoring data, copying data, or providing data to a user or other designated entity.

20. A computer program product, comprising one or more hardware storage devices having computer-executable instructions stored thereon, the computer-executable instructions being executable by one or more processors to cause a computer system to establish a virtual machine reference point, the computer-executable instructions including instructions executable to cause the computer system to at least perform the following: Maintain a plurality of data storage identifiers that identify changed data blocks among a plurality of data blocks of a data storage device corresponding to a virtual machine; Access a stable virtual machine checkpoint that includes a recoverable image of the virtual machine at a point in time and stores a representation of data when at least one data block among the plurality of data blocks of the data storage device existed at the point in time; Convert the virtual machine checkpoint into a virtual machine reference point that includes a representation of the virtual machine at the point in time, wherein the virtual machine reference point information is transferable with the virtual machine such that if the virtual machine is moved to a different computing system, any data identified by the virtual machine reference point is recoverable, the conversion including: Querying the data storage device to determine at least one data storage identifier corresponding to the at least one data block of the virtual machine checkpoint at the point in time; Storing the determined at least one data storage identifier as part of the virtual machine reference point; and Releasing the representation of the data of the at least one data block from the virtual machine checkpoint; and After converting the virtual machine checkpoint into the virtual machine reference point, the virtual machine reference point is used to identify one or more changes in the plurality of data blocks of the data storage device since the time point.

Citation Information

Patent Citations

  • Storage checkpointing in a mirrored virtual machine system

    CN103562878A

  • Methods and systems for storage system generation and use of differential block lists using copy-on-write snapshots

    US20080140963A1