Hot restart of hypervisor

The hot restart of a hypervisor by initializing a second hypervisor within a service partition and synchronizing its state minimizes downtime, enabling quick transitions with minimal disruption to virtual machine services.

JP2025102848AActive Publication Date: 2025-07-08MICROSOFT TECHNOLOGY LICENSING LLC
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
JP2025051079
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2020-04-15
Filing Date
2025-03-26
Publication Date
2025-07-08
Estimated Expiration
2041-03-15

AI Technical Summary

Technical Problem

The interruption and downtime caused by the reset or restart of a hypervisor significantly affect the workload and services running on virtual machines due to the need to interrupt all virtual machines during the process.

Method used

A method for hot restarting a hypervisor by initializing a second hypervisor within a service partition with an identity mapping between guest and host physical address spaces, synchronizing the state of the first hypervisor, and de-virtualizing the second hypervisor to replace the first with minimal disruption to running guest partitions.

Benefits of technology

Reduces downtime to seconds or less, allowing seamless continuation of services without noticeable interruption, as the second hypervisor is synchronized and initialized while the first is still running, thus maintaining high availability of virtual machine services.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025102848000001_ABST
    Figure 2025102848000001_ABST
Patent Text Reader

Abstract

To provide a computing system that achieves hot restart of a hypervisor by replacing a first hypervisor in execution with a second hypervisor with minimum down time for a guest partition.SOLUTION: A computing system executes a first hypervisor, creates one or more virtual partitions, generates a service partition during a hot restart, initializes a second hypervisor, migrates and synchronizes at least part of a runtime state of the first hypervisor to the second hypervisor using a reverse hypercall, and de-virtualizes the second hypervisor from the service partition in order to replace the first hypervisor after synchronization. De-virtualization includes transferring control of hardware resources from the first hypervisor to the second hypervisor using a previously migrated and synchronized runtime state.SELECTED DRAWING: Figure 6A
Need to check novelty before this filing date? Find Prior Art

Description

Background Art

[0001] Background

[0001] Virtualization in computing often refers to abstracting physical components into logical objects. A virtual machine (VM) can virtualize hardware resources including a processor, memory, storage, and network connectivity, and present the virtualized resources to the host operating system. The process of virtualizing a VM includes at least two parts: (1) mapping virtual resources or states, such as registers, memory, or files, to real resources within the underlying physical hardware, and (2) using machine instructions and / or system calls to execute actions specified by virtual machine instructions and / or system calls, such as emulating the application binary interface (ABI) or instruction set architecture (ISA) interface of the virtual machine.

[0002]

[0002] A hypervisor is a software layer that provides an environment (i.e., a virtualized hardware partition) in which virtual machines operate. The hypervisor is located between the physical resources on a physical hardware computing system and the VMs. Without a hypervisor, an operating system communicates directly with the hardware beneath it, i.e., disk operations go directly to the disk subsystem and memory calls are fetched directly from physical memory. When multiple operating systems of multiple VMs are running simultaneously on a single machine, the hypervisor manages the interaction between each VM and the shared hardware to prevent simultaneous control of the shared hardware by multiple VMs.

[0003]

[0003] When the hypervisor is reset or restarted (e.g., due to a software upgrade), all VMs running on the hypervisor are interrupted (e.g., restarted, paused, etc.), which can significantly affect the workload and the services running on the VMs.

[0004]

[0004] The content claimed in this specification is not limited to embodiments that solve any disadvantages or embodiments that operate only in the environments as described above. Rather, this background is provided only to illustrate an exemplary technical field in which some of the embodiments described in this specification can be practiced.

Summary of the Invention

Means for Solving the Problems

[0005] Brief Summary

[0005] This summary is provided to introduce, in a simplified form, a series of concepts that are further described in the detailed description below. This summary is not intended to identify key features or essential features of the content recited in the claims, nor is it intended to be used as an aid in determining the scope of the content recited in the claims.

[0006]

[0006] The embodiments described in this specification relate to a hot restart of a hypervisor that replaces a first hypervisor with a second hypervisor with little interference to a currently running guest partition. The embodiments described in this specification are implemented in a computing system. First, the computing system executes a first hypervisor on the computing system. The first hypervisor is configured to create one or more guest partitions. A service partition is created during a soft restart of the hypervisor. A second hypervisor is initialized within the service partition. The service partition is created with an identity mapping between its guest physical address (GPA) space and host physical address (HPA) space. Optionally, the service partition may be given additional privileges beyond conventional partitions / virtual machines to facilitate its initialization. Any execution environment that can meet the boot and security requirements of the platform can create and / or initialize the service partition. For example, components within a trusted computing base (TCB) of a computing system, such as the first hypervisor, may be tasked with creating or initializing the second hypervisor. During initialization, at least a part of the state of the first hypervisor is initialized to the second hypervisor. The state of the first hypervisor includes, but is not limited to, (1) a static system state, (2) a dynamic system state, and (3) a logical software state. The static system state includes, but is not limited to, the system topology, memory map, and memory layout of the first hypervisor. The dynamic system state includes, but is not limited to, a hardware architecture state visible to the guest, including general-purpose registers, mode control registers. The dynamic system state also includes, but is not limited to, a state of a guest VM, such as a state of a virtual central processing unit (CPU), a second-level page table (also known as a nested page table), and a list of allocated hardware devices (e.g., network cards), a virtualization instruction set architecture (ISA)-specific hardware state.The logical software state includes, but is not limited to, a page frame number (PFN) database. Since the dynamic system state and the logical software state can change continuously during runtime, these states are also referred to as the runtime state. Finally, a second hypervisor is de-virtualized from the service partition to replace the first hypervisor, and the state of each VM is restored by the second hypervisor.

[0007]

[0007] In some embodiments, at least one of one or more guest partitions includes a privileged parent partition, and the parent partition operates a host operating system including an orchestrator configured to coordinate the initialization and synchronization of the second hypervisor. In some embodiments, the first hypervisor supports by enabling the orchestrator to register and complete certain requests received from the second hypervisor (e.g., hypercall intercepts, register intercepts). Since these requests (sent from the second hypervisor to the first hypervisor) are intercepted (e.g., registered and completed) by the orchestrator, such requests are also referred to as "intercepts". In some embodiments, in response to receiving an intercept, the orchestrator issues a "reverse hypercall" to the second hypervisor in the service partition according to the application binary interface (ABI) of the hypercall to transition the relevant state. A hypercall is a call from a guest partition to a hypervisor, and a "reverse hypercall" is a call from a privileged software or hardware component (e.g., the orchestrator) to the second hypervisor within the service partition. In some embodiments, the orchestrator may share a portion of the memory page with the first hypervisor and the second hypervisor for communication purposes (e.g., hypercalls, intercepts, and / or reverse hypercalls).

[0008]

[0008] Examples of embodiments are implemented within a dedicated partition (e.g., a service partition) for providing services to a hypervisor. However, the overall technique of pre-initializing data structures within a virtual machine and / or migrating / synchronizing the runtime state within a VM / partition to reduce the service downtime and provide better uptime and continuity can be used in any part of the OS of any VM.

[0009]

[0009] Further features and advantages are described in the following description, become apparent in part from the description, or may be learned by practicing the teachings herein. The features and advantages of the present invention can be realized and attained by means of the instrumentalities and combinations particularly pointed out in the appended claims. The features of the present invention will become more fully apparent from the following description and appended claims, or may be learned by practicing the invention as set forth hereinafter.

[0010] Brief Description of the Drawings

[0010] To explain the advantages and features mentioned above and other advantages and features that can be obtained, a more specific description of the content briefly described above is made by referring to specific embodiments shown in the accompanying drawings. It is understood that these drawings merely illustrate typical embodiments and should not be regarded as limiting the scope. The embodiments will be described and explained with further specificity and detail using the accompanying drawings.

Brief Description of the Drawings

[0011]

Figure 1A

[0011] An example of a computing system including a hypervisor that hosts one or more virtual machines is shown.

Figure 1B

[0011] An example of a computing system including a hypervisor that hosts one or more virtual machines is shown.

Figure 2

[0012] A schematic diagram of the isomorphism between a guest partition and a host computing system is shown.

Figure 3A

[0013] An example of an embodiment of VM state management is shown.

Figure 3B

[0013] An example of an embodiment of VM state management is shown.

Figure 4

[0014] An example of Embodiment 400 of memory virtualization using a multi-level memory page table is shown.

Figure 5A

[0015] An example of a computing system in which hot restart of a hypervisor is enabled is shown.

Figure 5B

[0015] An example of a computing system in which hot restart of a hypervisor is enabled is shown.

Figure 6A

[0016] A flowchart of an example of a method for hot restart of a hypervisor is shown.

Figure 6B

[0017] A flowchart of an example of a method for initializing a service partition is shown.

Figure 6C

[0018] A flowchart of an example of a method for synchronizing a part of the runtime state of a first hypervisor with a second hypervisor is shown.

Figure 6D

[0019] A flowchart of an example of a method for collecting the runtime state of a first hypervisor is shown.

Figure 6E

[0020] A flowchart of an example of a method for sharing the runtime state of a first hypervisor with a second hypervisor is shown.

Figure 6F

[0021] A flowchart of an example of a method for replicating the runtime state of a first hypervisor to a second hypervisor is shown.

Figure 6G

[0022] A flowchart of an example of a method for de-virtualizing a second hypervisor is shown.

Figure 7

[0023] An example of a computing system that can employ the principles described herein is shown.

Mode for Carrying Out the Invention

[0012] Detailed Description

[0024] Embodiments described herein relate to a hot restart of a hypervisor that replaces a first hypervisor with a second hypervisor with little interference to a currently running guest partition. Embodiments described herein are implemented in a computing system. First, the computing system runs a first hypervisor on the computing system. The first hypervisor is configured to create one or more virtual machines / partitions each hosting a guest operating system. During a soft restart of the hypervisor, a service partition having a second-level page table with identities mapped is created. Then, the second hypervisor is initialized within the service partition and synchronized with at least a part of the state of the first hypervisor. Any execution environment capable of meeting the boot and security requirements of the platform can create and / or initialize the second hypervisor. For example, components within a trusted computing base (TCB) such as the first hypervisor can be tasked with creating and / or initializing the service partition.

[0013]

[0025] In some embodiments, the state of the first hypervisor includes, but is not limited to, (1) a static system state, (2) a dynamic system state, and (3) a logical software state. The static system state includes, but is not limited to, the system topology, memory map, and memory layout of the first hypervisor. The dynamic system state includes the hardware architecture state visible to the guest, including but not limited to general-purpose registers, mode control registers. The dynamic system state also includes the state of the guest VM, such as the state of the virtual central processing unit (CPU), the second-level page table (also known as the nested page table), and the list of allocated hardware devices (e.g., network cards), which are virtualization instruction set architecture (ISA)-specific hardware states. The logical software state includes, but is not limited to, the page frame number (PFN) database. Since the dynamic system state and the logical software state can change continuously during runtime, these states are also called runtime states. Finally, to replace the first hypervisor, the second hypervisor is virtualization-released from the service partition, and the state of each VM is restored by the second hypervisor.

[0014]

[0026] In some embodiments, at least one of the one or more guest partitions includes a privileged parent partition, and the parent partition operates a host operating system that includes an orchestrator configured to adjust the initial setup and synchronization of a second hypervisor. In some embodiments, a first hypervisor is part of the system's trusted computing base (TCB), and the orchestrator is executed within a trusted execution environment (TEE) to maintain the reliability of the inputs generated for the second hypervisor. In some embodiments, the first hypervisor enables and supports the orchestrator to register and complete certain requests received from the second hypervisor. These requests (sent from the second hypervisor to the first hypervisor) are intercepted (e.g., registered and completed) by the orchestrator, and such requests are also referred to as "intercepts". In some embodiments, in response to receiving an intercept, the orchestrator issues a "reverse hypercall" to the second hypervisor according to the application binary interface (ABI) of the hypercall to transition the relevant state. A hypercall is typically a call from a guest partition that requests an appropriate privileged operation to the hypervisor. The switch from the guest partition context to the hypervisor context is realized using platform-specific instructions. For using platform-specific instructions / mechanisms, control is returned from the hypervisor context to the guest context upon completion of the privileged operation. A "reverse hypercall" requests an operation from a privileged component (e.g., the orchestrator) to the hypervisor within a service partition that uses the hypercall ABI. The switch from the orchestrator to the second hypervisor context is realized by permitting the execution of a virtual processor of the service partition that has registers formatted as required by the hypercall ABI.Upon completion of the reverse hypercall, the second hypervisor generates a very special intercept that causes an intentional context switch from the second hypervisor within the service partition to the original orchestrator. In some embodiments, the orchestrator may share one or more memory pages with the first hypervisor and the second hypervisor for communication purposes (e.g., hypercalls, intercepts, and / or reverse hypercalls).

[0015]

[0027] In some embodiments, initializing the second hypervisor includes the first hypervisor generating a loader block for the second hypervisor. The loader block includes a logical construct that describes a part of the static system state for initialization. The static system state includes various system invariant conditions such as system topology, memory map, and layout. The memory map includes an identity map of substantially all system memory. In some embodiments, the identity map includes at least the memory visible from the first hypervisor being replaced. For example, in some cases, both the first (old) hypervisor and the second (new) hypervisor know about all of the memory in the system. In some cases, at least one of the first hypervisor and / or the second hypervisor does not know about all of the memory, for example, when some RAM including errors has to be taken offline, or when the second (new) hypervisor has to know about newly hot-added memory during a hot restart. The orchestrator obtains one or more system invariant conditions of the computing system, including at least one or more features supported by the hardware resources of the computing system, (1) directly from the hardware or system resources if the specific resources are not virtualized by the hypervisor, and / or (2) from the first hypervisor if the access of the orchestrator to the underlying physical hardware resources is virtualized by the first hypervisor. In some embodiments, the orchestrator transfers the relevant system invariant conditions by reverse hypercalls. Then, the second hypervisor is initialized based on the one or more shared system invariant conditions. In some embodiments, initializing the second hypervisor may also include providing the second hypervisor with read-only exclusive access to certain physical resources.

[0016]

[0028] After the second hypervisor is initialized using the static system state within the service partition, the computing system migrates the runtime state of the first hypervisor to the second hypervisor and synchronizes it. The runtime state includes at least the dynamic system state and the logical software state. Synchronizing the runtime state of the first hypervisor may include the orchestrator collecting the runtime state of the first hypervisor and sharing / migrating the runtime state with the second hypervisor over a reverse hypercall. As described above, the dynamic system state includes, but is not limited to, the hardware architecture state visible to the guest including general-purpose registers, mode control registers. The dynamic system state also includes, but is not limited to, the state of the guest VM such as the state of the virtual central processing unit (CPU), the second-level page table (also known as the nested page table), and the list of allocated hardware devices (e.g., network cards), the virtualization instruction set architecture ISA-specific hardware state. The logical software state includes, but is not limited to, the page frame number (PFN) database. The second hypervisor then duplicates the shared second-level memory page table and / or PFN database and puts them to sleep until the second hypervisor is finally de-virtualized.

[0017]

[0029] In addition, during this synchronization, the guest partitions running on the first hypervisor continue to operate and issue hypercalls to the first hypervisor. Therefore, after the initial synchronization of the compute-intensive state, additional incoming hypercalls to the first hypervisor that affect the state transmitted in the past may also have to be synchronized to the second hypervisor. In some embodiments, such state changes are also transmitted using reverse hypercalls.

[0018]

[0030] When the first hypervisor and the second hypervisor are fully or substantially synchronized, the second hypervisor is de-virtualized. In some embodiments, de-virtualization includes the first hypervisor "trampolining" to the second hypervisor and relinquishing all physical hardware control to the second hypervisor. "Trampolining" may refer to a memory location that holds addresses indicating interrupt service routines, I / O routines, etc. Here, "trampolining" refers to transferring control of all physical hardware from the first hypervisor to the second hypervisor. In some embodiments, trampolining can be achieved by reusing an existing control transfer mechanism of the computing system for kernel soft reboot. In some embodiments, no other system or user software, or booting system firmware, is permitted to execute during the trampolining process. However, a programmed direct memory access (DMA) before de-virtualization can continue to execute and complete. In some embodiments, the second hypervisor can be subjected to additional initialization, verify the hardware state, re-initialize the hardware, and / or initialize new hardware not previously programmed by the first hypervisor. De-virtualization may also include freezing each of the guest partitions currently running on the first hypervisor and transmitting the final state of the first hypervisor to the second hypervisor. Thereafter, each of the guest partitions is switched onto the second hypervisor. Then, the second hypervisor unfreezes each of the guest partitions. The memory footprint of the first hypervisor can be intermittently or slowly reclaimed by the computing system that bears the overall memory load.

[0019]

[0031] The embodiments described herein are implemented within a VM environment that can simultaneously support multiple guest partitions that each execute a unique operating system and associated application programs, and thus some introductory explanations regarding virtualization and hypervisors with respect to FIGS. 1-4 are provided.

[0020]

[0032] Hypervisor-based virtualization often enables a privileged host operating system (operating within a "parent" partition) and multiple guest operating systems (operating within "child" partitions) to simultaneously share access to the hardware of a single computing system, with each operating system given the illusion of having access to the entire set of system resources. To create this illusion, in some embodiments a hypervisor in the computing system creates multiple partitions that operate as virtual hardware machines (i.e., VMs) that each execute a respective operating system and associated application programs. Each operating system controls and manages a set of virtualized hardware resources.

[0021]

[0033] FIG. 1A shows an example of a computing system 100A that includes a hypervisor 140A hosting one or more VMs 110A, 120A. The computing system 100A includes various hardware devices 150 such as one or more processors 151, one or more storage devices 152, and / or one or more peripheral devices 153. The peripheral device 153 can be configured to perform input and output (I / O) for the computing system 100A. The ellipsis 154 indicates that different or additional hardware devices can be included within the computing system 100A. Hereinafter, the physical hardware 150 that executes the hypervisor 140A is also referred to as the "host computing system" 100A (also referred to as the "host"), and the VMs 110A, 120A hosted on the hypervisor 140A are also referred to as "guest partitions" (also referred to as "guests").

[0022]

[0034] As shown in FIG. 1A, hypervisor 140A hosts one or more guest partitions (e.g., VM A110, VM B120). The ellipsis 130 indicates that any number of guest partitions may be included in computing system 100A. Hypervisor 140A allocates a portion of the physical hardware resources 150 of computing system 100A to each of guest partitions 110A, 120A, giving guest partitions 110A, 120A the illusion of owning resources 114A, 124A. These “illusory” resources are called “virtual” resources. Each of VMs 110A, 120A uses its own virtual resources 114A, 124A to execute its own operating system (OS) 113A, 123A. Operating systems 113A, 123A can allocate virtual resources 114A, 124A to their various user applications 112A, 122A. For example, guest partition 110A executes OS 113A and user application 112A. If OS 113A is a Windows® operating system, user application 112A is a Windows® application. As another example, VM 120A executes OS 123A and user application 1122A. If OS 123A is a Linux operating system, user application 122A is a Linux application.

[0023]

[0035] Each virtual hardware resource 114A, 124A may or may not have a corresponding physical hardware resource 150. When the corresponding physical hardware resource 150 is available, the hypervisor 140A determines how to provide access to the guest partitions 110A, 120A that request its use. For example, the resource 150 can be partitioned or time-shared. If the virtual hardware resources 114A, 124A do not have a matching physical hardware resource, the hypervisor 140A can typically emulate the desired hardware resource actions by a combination of software and other hardware resources physically available on the host computing system 100A.

[0024]

[0036] As shown in FIG. 1A, in some embodiments, the hypervisor 140A can be the only software executed within the highest privilege level defined by the system architecture. Such a system is called a native VM system. Conceptually, in a native VM system, the hypervisor 140A is first installed on the bare hardware, and then the guest partition VMs A 110A and B 120A are installed on the hypervisor 140A. The guest operating systems 113A, 123A and other less privileged applications 112A, 122A are executed within a privilege level lower than that of the hypervisor. This generally means that the privilege levels of the guest OSs 113A, 123A may be emulated by the hypervisor 140A.

[0025]

[0037] Alternatively, in some embodiments, a hypervisor is installed on a host platform that is already running an existing OS. Such a system is referred to as a host-based VM system. In a host-based VM system, the hypervisor utilizes functions already available on the host OS to control and manage the resources desired by each of the guest partitions. In a host-based VM system, the hypervisor can be implemented at the user level or at the privileged level similar to the host operating system. Alternatively, a part of the hypervisor is implemented at the user level and another part of the hypervisor is implemented at the privileged level.

[0026]

[0038] In some embodiments, one of the guest partitions running on the same computing system can be considered to be more privileged than the other guest partitions. FIG. 1B shows an example of such an embodiment, where the more privileged guest partition can be referred to as the parent partition 110B and the remaining partitions can be referred to as child partitions 120B. In some embodiments, the parent partition 110B includes an operating system 113B and / or a virtualization service module 115B. Each of the operating system 113B and / or the virtualization service module 115B can have direct access to the hardware device 150 by the device driver 114B. Accordingly, the operating system 113B can also be referred to as the host operating system. Further, the virtualization service module 115B can create child partitions 120B using hypercalls. Depending on the configuration, the virtualization service module 115B can expose a subset of the hardware resources to each child partition 140B by the virtualization client module 125B. The child partitions 120B generally do not have direct access to the physical processor and / or cannot process real interrupts. Instead, the child partitions 120B can use the virtualization client module 125B to obtain a virtual view of the virtual hardware resources 124B.

[0027]

[0039] In some embodiments, the parent partition 110B may also include a VM management service application 112B that enables a user (e.g., a system administrator) to view and modify the configuration of the virtualization service module 115B. For example, the hypervisor 140B can be a Microsoft® Hyper-V hypervisor, and the parent partition can run a Windows® server. The user interface of the parent partition may provide a window that displays the full user interface of the child partition 120B. Interaction with applications running on the child partition 120B can occur within the window. When the host operating system 113B is Windows®, a graphical window can be established on the desktop interface to interact with the child partition 120B on the same platform. Elements 113B, 122B, 123B, and / or 124B in FIG. 1B are similar to elements 113A, 122A, 123A, and / or 124A in FIG. 1A, and thus will not be discussed further.

[0028]

[0040] Whether a native VM system or a hosted VM system, the relationship between the hypervisor and the guest partition is generally similar to the relationship between an operating system and an application program in a conventional computing system. In a conventional computing system, the operating system generally functions within a privilege level higher than the application level, e.g., within kernel mode relative to user mode. Similarly, in a VM environment, the hypervisor also operates within a privilege mode higher than the mode of the guest partition. When the guest partition needs to perform a privileged operation such as updating a page table, the guest partition uses a hypercall to request such an operation in the same way as a system call in a conventional operation.

[0029]

[0041] Therefore, the described embodiments of the present invention are applicable to both native VM systems and host-based VM systems, and the term "hypervisor" in this specification refers to a hypervisor implemented by any type of VM system.

[0030]

[0042] To better understand how the hypervisor operates, it is also necessary to understand how the hypervisor maintains the state of each guest partition. In a computing system, the configured state of the computing system is contained within and maintained by the hardware resources of the computing system. Typically, there is a configured hierarchy of state resources that spans from registers at one end of the hierarchy to secondary storage (e.g., hard drive) at the other end of the hierarchy.

[0031]

[0043] In a VM environment, each guest partition may or may not have its own configured state information and there may or may not be sufficient physical resources within the host computing system to map each element of the guest's state to its natural level within the host's memory hierarchy. For example, the register state of a guest can be actually maintained within the main memory of the host platform as part of a register context block.

[0032]

[0044] In normal operation, the hypervisor 140A periodically switches control between guest partitions 110A, 120A. When an operation is performed on the state of a guest, the state maintained on the host computing system 100A is modified as would be done on the guest operating systems 113A, 123A. In some embodiments, the hypervisor 140A constructs an isomorphism that maps the state of the virtual guest operating systems 113A, 123A to the state of the physical host computing system 100A.

[0033]

[0045] Figure 2 shows a schematic diagram of an isomorphism 200 between a guest partition 210 and a host computing system 220. The isomorphism 200 uses functions 230, namely V(state A) = state A' and V(state B) = state B', to map guest states 211, 212 to host states 221, 222. For a series of operations within the guest partition 210 that modify the state of the guest partition 210 from state A 211 to state B 212, there is a corresponding series of operations within the host computing system 220 that modify the state of the host computing system from state A' 221 to state B' 222. The isomorphism 200 between the guest partition 210 and the host computing system 220 is managed by a hypervisor (e.g., 140A, 140B).

[0034]

[0046] In an embodiment, there are two basic ways to manage the guest state so that this VM isomorphism is realized. One way is to use a certain degree of indirection by holding the state of each guest in a fixed position within the memory hierarchy of the host computing system with the pointer managed by the hypervisor indicating the currently active guest state. When the hypervisor switches between guest partitions, the hypervisor changes the pointer to match the current guest. FIG. 3A shows an example of such an embodiment 300A of state management by indirection. Referring to FIG. 3A, the memory 320 of the host computing system managed by the hypervisor stores the register values of each of VMs A and B in register context blocks 321, 322. The register block pointer 311A of the processor 310A points to the register context block 322 of the currently active guest partition (e.g., VM B). When a different guest partition is activated, the hypervisor changes the pointer 311A stored in the processor to point to the register context block 321, 322 of the activated guest partition, loads the program counter to point to the activated VM program, and starts execution.

[0035]

[0047] Another way to manage the guest state is to always replicate the guest state information to its natural level within the memory hierarchy when activated by the hypervisor and replicate it back when a different guest is activated. FIG. 3B shows an example of such an embodiment 300B of guest state management by replication. As shown in FIG. 3B, the host memory 320 managed by the hypervisor also stores the register values of VMs A and B of each guest partition in register context blocks 321, 322. However, unlike in FIG. 3A, here the hypervisor replicates all guest register contents 322 into the register file 311B of the processor 310B when VM B is activated (after saving the registers of the previous guest back to memory 320).

[0036]

[0048] The choice between indirect reference and replication can be determined, for example, by the frequency of use and whether the guest state managed by the hypervisor is held in different types of hardware resources than on the native system. For frequently used state information such as general-purpose registers, it may be preferable to swap the state of the virtual machine to the corresponding physical resources each time the virtual machine is activated. However, as shown in FIGS. 3A and 3B, in either case, the register values 321, 322 of each VM are often kept within the main memory 320 of the host platform as part of the register context block.

[0037]

[0049] In addition to VM state management, memory management is also worth explaining. In a VM environment, each guest partition has a set of unique virtual memory tables, also called "first-level" memory page tables. Address translation in each of the first-level memory page tables transforms an address within its virtual address space to a location within guest physical memory. Here, guest physical memory does not correspond to host physical memory on the host computing system. Instead, the guest physical address (GPA) is further mapped to determine an address within the physical memory of the host hardware, also called the host physical address (HPA). This mapping from GPA to HPA is performed by another set of virtual memory tables of the host computing system, also called "second-level" or nested memory page tables. Note that the combined total size of all guests' guest physical memory can be larger than the actual physical memory on the system. In an embodiment, the hypervisor maintains a separate unique swap space for each guest, and the hypervisor manages physical memory by swapping guest physical pages in and out of its own swap space. Further, the state of all virtually or physically allocated pages, and their corresponding attributes, are stored in a list called the page frame number (PFN) list. The tracking of virtually or physically allocated pages is stored in a database called the PFN database.

[0038]

[0050] Figure 4 shows an example of an embodiment 400 of memory virtualization that uses multi-level memory page tables 440 and 450. Each entry in the first-level memory table 440 maps a location (e.g., PFN) in the PFN database 410 of virtual memory to a location (e.g., PFN) in the PFN database 420 of guest physical memory. As shown in Figure 4, a portion of the PFN database 410 tracks the virtually allocated pages of a program running on VM A, and a portion of the PFN database 420 tracks the physical memory of the guest VM of VM A. Further, in order to convert GPA to HPA, the hypervisor also maintains a second-level memory page table 450 that maps guest physical pages to host physical pages. There is also a PFN database 430 that tracks the physical memory of the host computing system.

[0039]

[0051] As shown in Figure 4, the physical page frame numbered 1500 is allocated to the guest physical page frame numbered 2500, and this guest physical page frame is allocated to the virtual memory page frame numbered 3000. Similarly, the physical page frame 2000 is allocated to the guest physical page frame numbered 6000, and this guest physical page frame is allocated to the virtual memory page frame numbered 2000. The remaining physical memory pages can be allocated to other VMs or to the hypervisor itself. These remaining physical memory pages (e.g., memory 320), including those allocated to the hypervisor itself for recording the register values of each VM, are also tracked by the PFN database 430.

[0040]

[0052] Figure 4 is merely a schematic diagram for showing a simplified concept of memory management using a multi-level page table. Additional mechanisms may be implemented to achieve the same or similar memory management purposes. For example, in some embodiments, page translation is supported by a combination of a page table and a translation lookaside buffer (TLB).

[0041]

[0053] Referring to FIGS. 1A to 4, a virtual environment and how a hypervisor manages and virtualizes various hardware resources will be described. Next, specific embodiments of the hot restart of the hypervisor will be described with respect to FIGS. 5A and 5B.

[0042]

[0054] FIG. 5A shows an example of a computing system 500A in which the hot restart of the hypervisor is enabled. The computing system 500A may include a native VM system in which the hypervisor 520A is the only software executed within the highest privilege level defined by the system architecture, as shown in FIG. 1A. Alternatively, the computing system 500A may include a host-type VM system in which the hypervisor 520A is installed on a computing system that is already executing a host OS.

[0043]

[0055] Whether the computing system 500A includes a native VM system or a host-type VM system, a new partition called the service partition 560 is generated during the hot restart of the hypervisor. In an embodiment, the service partition 560 is treated differently from other types of partitions. When the service partition 560 is created, hardware resources including at least some processor resources and memory resources are allocated to the service partition 560. In some embodiments, the allocation of hardware resources may be based on user input. In alternative embodiments, the hypervisor 520A or the component that created the service partition 560 automatically allocates a predetermined portion of the processor resources and / or memory resources to the service partition 560.

[0044]

[0056] The allocation of processor resources can specify the total processing capacity required by service partition 560, leave the allocation of available processors to workload management software, or service partition 560 or hypervisor 520A may specify that a particular processor in the system be dedicated to the use of service partition 560. Service partition 560 or hypervisor 520A may specify that service partition 560 requires a certain number of processors, but that service partition 560 is willing to share those processors with other partitions. For example, if service partition 560 requires a total of eight processing devices, service partition 560 or hypervisor 520A may specify that service partition 560 requires eight dedicated processors, or that service partition 560 requires 16 processors, but only half of the available computing capacity of each processor. The allocation of memory (including RAM and / or hard disk) may specify a chunk of a particular granularity, e.g., an amount of memory in units of 1MB.

[0045]

[0057] Next, the service partition 560 is initialized. Any component within the trusted computing base (TCB) may be tasked with creating and / or initializing the service partition 560. In some embodiments, the first hypervisor 520A is part of the TCB, and the first hypervisor 520A generates and / or initializes the service partition 560. The initialization process includes bootstrapping, which includes a series of actions, each action activating a function that enables the next action to be taken until the entire system is activated. In some embodiments, the hypervisor 520A constructs a loader block for the second hypervisor 561. Executing the initialization code enables other aspects of the service partition 560 to be initialized. As shown in FIG. 5A, unlike other normal guest partitions 540A, 550A where the conventional operating systems 541A, 551A are loaded, the hypervisor 561 is loaded within the service partition 560. For clarity, hereinafter the hypervisors 520A and 561 are referred to as the first hypervisor 520A and the second hypervisor 561. In some embodiments, initializing the second hypervisor 561 may also include providing the second hypervisor with read-only exclusive access to certain physical resources 510.

[0046]

[0058] The purpose of the hypervisor hot restart is to ultimately replace the first hypervisor 520A with the second hypervisor 561 with minimal to imperceptible interruptions to the guest virtual machines. Before the second hypervisor 561 replaces the first hypervisor 520A, the first hypervisor 520A initializes the second hypervisor 561 using available system invariant conditions (e.g., features supported by the hardware resources 510), and then synchronizes the runtime state 512 with the second hypervisor 561. In some embodiments, the runtime state 512 is stored in memory (e.g., RAM) managed by the hypervisor 520A. As described with respect to FIGS. 3A, 3B, and 4, the runtime state 512 may include hardware architecture states such as register values 321, 322, etc. of each guest partition (e.g., guest A 540A, 550A), virtualized hardware states such as the second-level memory page table 450 of each guest partition (e.g., guest A 540A, 550A), and / or software-defined states such as the PFN database 430 associated with the first hypervisor 520A.

[0047]

[0059] In some embodiments, the communication between the first hypervisor 520A and the second hypervisor 561 during initialization and synchronization is coordinated by the orchestrator 530A. The first hypervisor 520A enables and supports the orchestrator 530A to register and complete certain requests received from the second hypervisor 561. The orchestrator 530A is a software component of the host computing system 500A configured to coordinate the communication between the first hypervisor 520A and the second hypervisor 561. In some embodiments, the first hypervisor 520A is part of the system's trusted computing base (TCB), and the orchestrator 530A is executed within a trusted execution environment (TEE) to maintain the reliability of the inputs generated for the second hypervisor.

[0048]

[0060] In some embodiments, the orchestrator 530A can use reverse hypercalls to transmit the state to the second hypervisor 561, and / or the second hypervisor 561 can use intercepts to request some services (e.g., characteristics of physical resources to which the second hypervisor 561 does not have access) from the orchestrator 530A. In some embodiments, the orchestrator 530A shares a portion of the memory page 511 with the second hypervisor 561 so that data can be efficiently communicated among the orchestrator 530A, the first hypervisor 520A, and the second hypervisor 561. In some embodiments, the orchestrator 530A issues a reverse hypercall to a second hypervisor that conforms to the application binary interface (ABI) of the hypercall to transfer the relevant state.

[0049]

[0061] As discussed above, when the service partition 560 is created and loaded, the second hypervisor 561 first needs to obtain various system invariant conditions such as the characteristics of the hardware resources 510. In some embodiments, there is a tight coupling between the characteristics recognized by the first hypervisor 520A and the characteristics recognized by the second hypervisor 561.

[0050]

[0062] Alternatively, in other embodiments, there is a loose coupling between the first hypervisor and the second hypervisor. The orchestrator 530A can query the first hypervisor via a hypercall to obtain the characteristics and properties of the hardware resources related to the initial configuration of the second hypervisor 561, and / or can directly query the hardware resources. For example, the second hypervisor 561 can attempt to read CPUID or MSR values. However, if it is within the virtual machine itself and the second hypervisor 561 can only access virtualized values rather than the physical values of the underlying physical computing system, in that case, the orchestrator 530A can call a hypercall to the first hypervisor 520A to obtain the corresponding physical values of the corresponding physical resources. For example, the second hypervisor 561 may want to query the processor support for the XSAVE function and instructions. If the first hypervisor 520A does not support virtualization of XSAVE, the orchestrator 530A can query the characteristics of the underlying processor and determine that the processor supports XSAVE. The obtained query results can also be stored in the shared memory page 511 so that the second hypervisor 561 has access to the query results. In some cases, the second hypervisor 561 may not be able to obtain all the characteristics of the hardware resources 510. In some embodiments, such characteristics are left unknown during initialization and obtained later after virtualization is removed.

[0051]

[0063] In addition to obtaining system invariant conditions, the runtime state of the first hypervisor 520A may also be synchronized with the second hypervisor 561. The runtime state includes at least a dynamic system state and a logical software state. The dynamic system state includes a hardware architecture state visible to a guest including, but not limited to, general-purpose registers and control registers. The dynamic system state includes a virtualization instruction set architecture ISA-specific hardware state including, but not limited to, the state of a guest VM such as the state of a virtual central processing unit (CPU), a second-level page table (also known as a nested page table), and a list of allocated hardware devices (e.g., a network card). Synchronizing the runtime state 512 of the first hypervisor includes at least synchronizing one or more second-level memory page tables and / or one or more PFN databases to the second hypervisor 561. The second-level memory table is a memory page table that maps the GPA of each guest partition to an HPA (e.g., the second-level memory page table 450 of FIG. 4). The PFN database to be synchronized may include a PFN database (e.g., PFN database 430) that tracks the physical memory of the host computing system.

[0052]

[0064] However, after the initial synchronization of the runtime state between the two hypervisors 520A and 561 and before the de-virtualization of the second hypervisor 561, the guest partitions 540A, 550A are still running on the first hypervisor 520A, and their guest partitions 540A, 550A can still call hypercalls. When an incoming hypercall is serviced by the first hypervisor 520A, the state of the guest partition 540A or 550A that called the hypercall changes, and the previously synchronized state of the second hypervisor 561 is no longer accurate. To solve this problem, incoming hypercalls (before the de-virtualization of the second hypervisor) also need to be recorded and synchronized with the second hypervisor 561.

[0053]

[0065] In some embodiments, the orchestrator 530A is also tasked with the task of recording and synchronizing each incoming hypercall. For example, when the first hypervisor 520A receives a hypercall called by the guest partitions 540A, 540B and provides services, the orchestrator 530A records a log of the hypercall and the actions that occurred during the service of the hypercall by the first hypervisor 520A. At the same time, the orchestrator 530A feeds the hypercall to the second hypervisor 561 via a reverse hypercall. In some embodiments, the instruction point of the second hypervisor 561 is within the hypercall dispatch loop, so that when a reverse hypercall is supplied to the second hypervisor 561, the second hypervisor 561 processes it. Upon receiving the reverse hypercall, the second hypervisor 561 switches the context of the guest partitions 540A, 540B that called the hypercall and processes the reverse hypercall to reconstruct the necessary software state and / or suspended hardware state. Upon completion of the reverse hypercall, the second hypervisor 561 notifies the orchestrator 530A of the completion.

[0054]

[0066] In some cases, a series of hypercalls may be serviced in a short period of time, and the orchestrator 530A may only need to send a state related to the last or relevant operation of the series of hypercalls. In that case, the second hypervisor 561 can only replay a partial log of the actions, that is, perform a condensed replay. This process can be repeated as many times as necessary until the second hypervisor 561 is fully or at least substantially synchronized with the first hypervisor 520A. Thereafter, the second hypervisor 561 is de-virtualized from the service partition 560 to replace the first hypervisor 520A. De-virtualization includes the first hypervisor 520A trampolining to the second hypervisor and relinquishing all physical hardware control to the second hypervisor 561. In some embodiments, the trampoline can be implemented by reusing an existing control transfer mechanism of the computing system 500A for kernel soft reboot.

[0055]

[0067] In some embodiments, the second hypervisor 561 can be subjected to additional initialization, verify the hardware state, re-initialize the hardware, and / or initialize new hardware not previously programmed by the first hypervisor. In some embodiments, de-virtualization may include the first hypervisor 520A freezing all guest partitions 540A, 550A, transmitting the details of the final state to the second hypervisor 561, and migrating the guest partitions onto the second hypervisor. The second hypervisor then unfreezes each of the guest partitions. When the second hypervisor 561 starts de-virtualization, the first hypervisor 520A effectively ends. No other system software, user software, and / or system firmware is permitted to execute during de-virtualization. However, the DMA programmed before de-virtualization can continue to execute and complete.

[0056]

[0068] For example, when a guest partition (e.g., child A 540A, 550A) provides access to a physical device, the guest partition can initiate DMA using its GPA as the source or target of the DMA operation. (Programmable in an input / output memory management unit (IOMMU) by a first hypervisor) The second-level page table performs the necessary permission checks, in addition to converting the GPA to an HPA and providing the HPA to the DMA engine. As described above, the first hypervisor transmitted / synchronized the architectural guest state and the architectural virtual state to the second hypervisor. The architectural virtual state includes, among other things, the second-level page table for the CPU, IOMMU, and / or domain information of the device. Thus, when the second hypervisor virtualization is removed and the hardware is re-initialized, the second hypervisor carefully programs the hardware using the new page table it constructed during the previous synchronization phase. Since the effective address translation and permission of the new page table are identical while being two different instances, all new translation requests from the DMA engine can continue to execute without loss of fidelity and can successfully use the same page table.

[0057]

[0069] In addition, in some embodiments, the memory footprint of the first hypervisor 520A can be intermittently or slowly reclaimed by a computing system that receives the overall memory load.

[0058]

[0070] Figure 5B shows another example of a computing system 500B in which hot restart of the hypervisor is enabled. Computing system 500B corresponds to computing system 100B of FIG. 1B, and one of the guest partitions (e.g., parent partitions 110B, 530B) is given more privileges than the remaining guest partitions (e.g., child partitions 120B, 540B, 550B). In some embodiments, the parent partition 530B includes a host operating system and / or a virtualization service module 531B, either of which can be configured to create and manage the child partitions 540B and 550B and handle various system management functions and device drivers. Similar to the embodiment shown in FIG. 5A, a service partition 560 is created, a second hypervisor 561 is initialized within the service partition 560, and finally the first hypervisor 520B is replaced with the second hypervisor 561 to complete the hot restart of the hypervisor.

[0059]

[0071] In some embodiments, since the parent partition 530B has a higher level of privilege than the child partitions 540B, 550B within the computing system 500B, an orchestrator 532B can be implemented within the parent partition 530B as part of the host operating system or virtualization service module 531B of the parent partition 530B. The orchestrator 532B functions in the same manner as the orchestrator 530B to coordinate communication between the first hypervisor 520B and the second hypervisor 561. Elements 541B, 551B of FIG. 5B are the same as elements 541A, 551A of FIG. 5A and thus will not be discussed further.

[0060]

[0072] In addition, in some embodiments, a kernel soft reboot or reset of guest partitions 540A, 550A, 530B, 540B, 550B may also be associated with a hot restart of hypervisors 520A, 520B. In a kernel soft reboot, guest partitions 540A, 550A, 530B, 540B, or 550B may be recreated as new partitions and reinitialized, and the runtime state of the corresponding guest partition may be synchronized with the new partition. When all runtime states of the guest partition are synchronized with the new partition, the new partition can replace the corresponding guest partition to complete the kernel soft reboot of the corresponding guest partition. In some embodiments, only the parent partition 530B is hot restarted in association with the hot restart of hypervisor 520B. Alternatively or in addition, each of guest partitions 540A, 550A, and / or child partitions 540B, 550B is hot restarted along with the hot restart of hypervisors 520A, 520B.

[0061]

[0073] Some embodiments restart VM-related components within the parent partition 530B without restarting the host operating system within the parent partition 530B. For example, some embodiments restart the virtualization service module 531B in relation to the hot restart of the hypervisor without restarting the host operating system 113B. In this way, the hypervisor can be upgraded and restarted together with its operating system-level management components without restarting the host operating system 113B.

[0062]

[0074] The hot restart of the hypervisor described in this specification significantly reduces the interruption time given to the running guest VMs, unlike the restart of a conventional hypervisor. Unlike the restart of a normal hypervisor that can take several minutes depending on the number of hosted VMs and the amount of hardware resources being managed, the second hypervisor 561 described in this specification is initialized and synchronized while the first hypervisor 520A or 520B is still running. Thus, in each running guest partition, there is only a short (e.g., 1 second or less than a few seconds) freeze period that the user may not even notice.

[0063]

[0075] Next, the following explanations refer to some methods and acts that can be performed. The acts of the method may be discussed in a certain order or may be shown in a flowchart as being performed in a specific order, but no specific ordering is required unless otherwise specified or required because an act depends on another act that is completed before that act is performed.

[0064]

[0076] Figure 6A shows a flowchart of an example of a method 600 for hot restart of a hypervisor. The method 600 is implemented on a computing system that may correspond to the computing systems 100A, 100B, or 500A, 500B. The method 600 includes executing a first hypervisor (610). In response, the first hypervisor creates one or more guest partitions, each of which may host a guest operating system (620). The purpose of the hot restart of the hypervisor is to replace the first hypervisor with a new hypervisor. When the hot restart of the hypervisor is performed, the computing system creates a service partition (630) and initializes a second hypervisor within the service partition (640). Next, at least a part of the runtime state of the first hypervisor is synchronized with the second hypervisor (650). When the synchronization is complete or almost complete, the computing system de-virtualizes the second hypervisor to replace the first hypervisor with the second hypervisor (660).

[0065]

[0077] FIG. 6B shows a flowchart of an example of a method 640 for initializing a service partition corresponding to step 640 of FIG. 6A. The method 640 includes generating (641) a loader block for a second hypervisor by a first hypervisor. The loader block is a logical construct that describes some of the static system characteristics and resources for initialization. The method 640 also includes obtaining (642) one or more system invariant conditions and sharing (643) the one or more system invariant conditions with the second hypervisor. Then the second hypervisor is initialized (644) based on the one or more system invariant conditions. In some embodiments, the system invariant conditions include (645) one or more characteristics of the hardware resources of the computing system. In some embodiments, there is a tight coupling (646) between the first hypervisor and the second hypervisor. In some embodiments, there is a loose coupling between the first hypervisor and the second hypervisor. In that case, one or more characteristics of the hardware resources can be obtained (647) from the first hypervisor by a hypercall, or directly obtained (648) from the hardware resources by a system call.

[0066]

[0078] Figure 6C shows a flowchart of an example of method 650 for synchronizing a portion of the runtime state of a first hypervisor with a second hypervisor, corresponding to step 650 of FIG. 6A. Method 650 may be executed by orchestrators 530A, 532B of FIGS. 5A or 5B. Method 650 includes gathering (651) the runtime state of the first hypervisor. Method 650 also includes sharing (652) the runtime state with the second hypervisor, which may be done by the orchestrators 530A, 532B via reverse hypercalls and may or may not further rely on shared memory pages to assist with signaling and / or message / data passing. The shared runtime state of the first hypervisor is then replicated (653) by the second hypervisor. Before the second hypervisor is de-virtualized, the guest partitions running on the first hypervisor continue to operate and can still make hypercalls, and the gathered runtime state quickly becomes inaccurate. Thus, this process may be repeated several times until the second hypervisor and the first hypervisor are fully or substantially synchronized.

[0067]

[0079] Figure 6D shows a flowchart of an example of method 651 for gathering the runtime state of a first hypervisor, corresponding to step 651 of FIG. 6C. Method 651 includes gathering (651-A) current data related to one or more second-level memory page tables and one or more PFNs. Method 651 also includes gathering (651-B) data related to incoming hypercalls. Specifically, when a hypercall is invoked by a guest partition, the first hypervisor receives (651-C) the received hypercall and services it (651-D). The computing system (e.g., orchestrator 530A or 532B) then records (651-E) a log of the received hypercall and the actions that occurred during the servicing of the hypercall.

[0068]

[0080] Figure 6E shows a flowchart of an example of method 652 for sharing the runtime state of a first hypervisor with a second hypervisor, corresponding to step 652 of FIG. 6C. Method 652 includes feeding one or more second-level page tables and one or more PFNs to the second hypervisor (652-A). Method 652 also includes feeding an in-hypercall to the second hypervisor (652-B) and feeding a log of actions that occurred during servicing of the hypercall (by the first hypervisor) to the second hypervisor (652-C).

[0069]

[0081] Figure 6F shows a flowchart of an example of method 653 for replicating the runtime state of a first hypervisor with a second hypervisor, corresponding to step 653 of FIG. 6C. Method 653 includes receiving a reverse hypercall from an orchestrator within the second hypervisor (653-A). The second hypervisor then switches the context of the guest partition that made the hypercall (653-B) and processes the reverse hypercall to reconstruct the necessary software state and / or suspended hardware state (653-C). Upon completion of processing the reverse hypercall, the second hypervisor notifies the orchestrator (653-D).

[0070]

[0082] Figure 6G shows a flowchart of an example of method 660 for de-virtualizing a second hypervisor, corresponding to step 660 of FIG. 6A. Method 660 includes freezing each guest partition currently executing on the first hypervisor (661). Next, the final state of the first hypervisor is transmitted to the second hypervisor (662). Each of the guest partitions is then switched onto the second hypervisor (663), and the second hypervisor unfreezes each of the guest partitions (664). Finally, the first hypervisor is terminated (665).

[0071]

[0083] Note that even if the examples of the above embodiments are implemented within a dedicated partition (e.g., a service partition) for providing services to a hypervisor, the overall techniques of pre-initializing data structures and / or migrating / synchronizing runtime states to save time can be used in any part of any VM.

[0072]

[0084] Finally, since the principles described herein are implemented in the context of a computing system (e.g., the computing systems 100A, 100B of FIGS. 1A and 1B and / or 500A, 500B of FIGS. 5A or 5B), some introductory explanation of the computing system will be provided with respect to FIG. 7.

[0073]

[0085] Currently, computing systems are taking on an increasingly diverse range of forms. A computing system can be, for example, a portable device, an appliance, a laptop computer, a desktop computer, a mainframe, a distributed computing system, a data center, or even a device that has not conventionally been considered a computing system, such as a wearable (e.g., glasses). In this description and the claims, the term "computing system" is broadly defined to include any device or system (or combination thereof) that includes at least one physical and tangible processor and a physical and tangible memory on which computer-executable instructions executable by the processor can be had. The memory can take any form and can depend on the nature and form of the computing system. The computing system can be distributed across a network environment and can include multiple constituent computing systems.

[0074]

[0086] As shown in FIG. 7, in its most basic configuration, computing system 700 typically includes at least one hardware processing unit 702 and a memory 704. The processing unit 702 can include a general-purpose processor and can also include a field-programmable gate array (FPGA), an application-specific integrated circuit (ASIC), or any other dedicated circuit. The memory 704 can be a physical system memory that can be volatile, non-volatile, or some combination of the two. The term "memory" may also be used herein to refer to a non-volatile mass storage area such as a physical storage medium. If the computing system is distributed, the processing, memory, and / or storage functions may also be distributed.

[0075]

[0087] Computing system 700 also has a number of structures often referred to as "executable components" thereon. For example, the memory 704 of computing system 700 is shown as including an executable component 706. The term "executable component" is a name for a structure well known to those of ordinary skill in the computing art that can be a structure that can be software, hardware, or a combination thereof. For example, when implemented by software, the structure of an executable component can include software objects, routines, methods, etc. that can be executed on a computing system whether such executable component exists within the heap of the computing system or on a computer-readable storage medium, as would be understood by one of ordinary skill in the art.

[0076]

[0088] Those skilled in the art will recognize that when interpreted by one or more processors of a computing system (e.g., by a processor thread), the structure of an executable component exists on a computer-readable medium so as to cause the computing system to perform a function. Such a structure can be directly computer-readable by a processor (as would be the case if the executable component were binary). Alternatively, this structure can be configured to be interpretable and / or compiled (in one or multiple stages) to produce such binary that is directly interpretable by a processor. When using the term "executable component", such an understanding of examples of the structure of an executable component is well within the understanding of those skilled in the computing art.

[0077]

[0089] The term "executable component" is also well understood by those skilled in the art to include structures such as hard-coded or hard-wired logic gates that are implemented exclusively or almost exclusively in hardware, such as within a field-programmable gate array (FPGA), an application-specific integrated circuit (ASIC), or any other dedicated circuit. Thus, whether implemented by software, by hardware, or by a combination, the term "executable component" is a term for a structure well understood by those skilled in the computing art. In this description, terms such as "component", "agent", "manager", "service", "engine", "module", "virtual machine", etc. may also be used. When used in this description and in this example, these terms (whether expressed with a modifier or not) are synonymous with the term "executable component" and are thus intended to also have a structure well understood by those skilled in the computing art.

[0078]

[0090] In the foregoing description, embodiments have been described with respect to acts performed by one or more computing systems. When such acts are implemented in software, in response to executing computer-executable instructions that make up an executable component, one or more processors (of the associated computing system performing the act) direct the operation of the computing system. For example, such computer-executable instructions may be embodied by one or more computer-readable media that form a computer program product. An example of such an operation includes the manipulation of data. When such acts are implemented exclusively or almost exclusively in hardware, such as within an FPGA or ASIC, the computer-executable instructions can be hard-coded or hard-wired logic gates. The computer-executable instructions (and the data being operated on) may be stored within the memory 704 of the computing system 700. The computing system 700 may also include, for example, a communication channel 708 that enables the computing system 700 to communicate with other computing systems over a network 710.

[0079]

[0091] Not all computing systems require a user interface, but in some embodiments, the computing system 700 includes a user interface system 712 for use in interfacing with a user. The user interface system 712 can include an output mechanism 712A as well as an input mechanism 712B. The principles described herein are not limited to a particular output mechanism 712A or input mechanism 712B as they depend on the nature of the device. However, the output mechanism 712A can include, for example, speakers, displays, tactile outputs, holograms, and the like. Examples of the input mechanism 712B can include, for example, microphones, touchscreens, holograms, cameras, keyboards, mice, or other pointer inputs, any type of sensors, and the like.

[0080]

[0092] As discussed in more detail below, the embodiments described herein can include or utilize a special-purpose computing system or a general-purpose computing system that includes computer hardware such as, for example, one or more processors and system memory. The embodiments described herein also include physical and other computer-readable media for carrying or storing computer-executable instructions and / or data structures. Such computer-readable media can be any available media that can be accessed by a general-purpose computing system or a special-purpose computing system. A computer-readable medium that stores computer-executable instructions is a physical storage medium. A computer-readable medium that carries computer-executable instructions is a transmission medium. Thus, by way of example and not limitation, embodiments of the invention can include at least two distinctly different types of computer-readable media, namely storage media and transmission media.

[0081]

[0093] A computer-readable storage medium can be used to store desired program code means in the form of computer-executable instructions or data structures and can include RAM, ROM, EEPROM, CD-ROM, or other optical disk storage, magnetic disk storage, or other magnetic storage devices, or any other physical and tangible storage medium accessible by a general-purpose computing system or a special-purpose computing system.

[0082]

[0094] "Network" is defined as one or more data links that enable the carrying of electronic data between computing systems and / or modules and / or other electronic devices. When information is transferred or provided to a computing system over a network or another communication connection (either hardwired, wireless, or a combination of hardwired or wireless), the computing system properly views the connection as a transmission medium. A transmission medium can be used to carry desired program code means in the form of computer-executable instructions or data structures and can include a network and / or data link accessible by a general-purpose computing system or a special-purpose computing system. Combinations of the above are also to be included within the scope of computer-readable media.

[0083]

[0095] Further, upon reaching various computing system components, program code means in the form of computer-executable instructions or data structures can be automatically transferred from a transmission medium to a storage medium (or vice versa). For example, computer-executable instructions or data structures received over a network or data link are buffered in RAM within a network interface module (e.g., a "NIC") and then gradually transferred to the RAM of the computing system and / or a less volatile storage medium in the computing system. Thus, it should be understood that a storage medium can include computing system components that also utilize (or primarily utilize) a transmission medium.

[0084]

[0096] Computer-executable instructions, when executed in a processor for example, include instructions and data that cause a general-purpose computing system, a special-purpose computing system, or a special-purpose processing device to perform certain functions or groups of functions. Alternatively or in addition, computer-executable instructions can configure a computing system to perform certain functions or groups of functions. Computer-executable instructions can be binary or instructions that are subject to some transformation (such as compilation) before being directly executed by a processor, such as intermediate form instructions like assembly language or even source code.

[0085]

[0097] Although this content has been described in language specific to structural features and / or methodological acts, it should be understood that the content defined in the appended claims is not necessarily limited to the described features or acts. Rather, the described features and acts are disclosed as examples of forms for implementing the claims.

[0086]

[0098] Those skilled in the art will understand that the present invention can be practiced in a network computing environment having many types of computing system configurations, including personal computers, desktop computers, laptop computers, message processors, portable devices, multiprocessor systems, microprocessor-based or programmable household appliances, network PCs, minicomputers, mainframe computers, mobile phones, PDAs, pagers, routers, switches, data centers, wearables (such as glasses), and the like. The present invention can also be practiced in a distributed system environment where both local and remote computing systems linked via a network (by a hardwired data link, a wireless data link, or a combination of a hardwired data link and a wireless data link) perform tasks. In a distributed system environment, program modules can be located in both local and remote memory storage devices.

[0087]

[0099] Those skilled in the art will also understand that the present invention can be practiced in a cloud computing environment. A cloud computing environment may be distributed, but this is not essential. When distributed, the cloud computing environment may be internationally distributed within an organization and / or may have components maintained across multiple organizations. In this description and the appended claims, "cloud computing" is defined as a model that enables on-demand network access to a shared pool of configurable computing resources (e.g., networks, servers, storage, applications, and services). The definition of "cloud computing" is not limited to any of the many other advantages that can be obtained from such a model when properly introduced.

[0088]

[0100] The remaining drawings may discuss various computing systems that may correspond to the previously described computing system 700. The computing systems in the remaining drawings include various components or functional blocks that may implement the various embodiments disclosed herein. The various components or functional blocks can be implemented on a local computing system, or on a distributed computing system that includes elements in the cloud or implements aspects of cloud computing. The various components or functional blocks can be implemented as software, hardware, or a combination of software and hardware. The computing systems in the remaining drawings can include more or fewer components than those illustrated, and some components may be combined depending on the situation. Although not necessarily shown, the various components of the computing system can access and / or utilize processors and memories such as processor 702 and memory 704 when necessary to perform their various functions.

[0089]

[0101] In the processes and methods disclosed herein, the operations performed within the processes and methods may be implemented in a different order. Further, the operations outlined are merely illustrative, and some of the operations may be made optional without departing from the essence of the disclosed embodiments, combined into fewer steps and operations, supplemented with additional operations, or extended into additional operations.

[0090]

[0102] The present invention can be implemented in other specific forms without departing from the spirit or characteristics of the present invention. The described embodiments should be considered in all respects to be illustrative rather than restrictive. Accordingly, the scope of the present invention is indicated by the appended claims rather than by the above description. All modifications that fall within the meaning and scope of equivalence of the claims are embraced by the claims.

Claims

1. One or more processors and, when executed by the one or more processors, executing a first hypervisor to create one or more guest partitions, creating a service partition, initializing a second hypervisor within the service partition using a static system state of the first hypervisor, the static system state including at least an identity mapping of substantially all guest physical addresses to physical addresses of a computing system, initializing; synchronizing at least a portion of a runtime state of the first hypervisor with the second hypervisor, and unvirtuallizing the second hypervisor from the service partition to replace the first hypervisor One or more computer-readable media having thereon computer-executable instructions configured to cause the computing system to perform the foregoing, and A computing system comprising the foregoing.

2. The computing system of claim 1, wherein initializing the second hypervisor within the service partition further comprises granting read-only exclusive access to a specific portion of the one or more computer-readable media of the computing system.

3. The computing system of claim 1, wherein the runtime state includes at least one of (1) one or more second-level memory page tables that map guest physical memory of the guest partitions to host physical memory of the computing system, (2) one or more page frame number databases, or (3) a list of physical devices attached to the guest partitions and assignments of the list of physical devices.

4. At least one of the one or more guest partitions includes a privileged parent partition, the parent partition operates a host operating system or a virtualization service module, the host operating system or the virtualization service module includes an orchestrator configured to coordinate the initializing and synchronizing of the second hypervisor, the first hypervisor enables the orchestrator to intercept specific requests from the second hypervisor to the first hypervisor The computing system according to claim 1.

5. When the orchestrator intercepts a request from the second hypervisor to the first hypervisor, the orchestrator issues a reverse hypercall to the second hypervisor and shares a part of a memory page with the second hypervisor for communication purposes. The computing system according to claim 4.

6. The initial setting of the second hypervisor is generating, by the first hypervisor, a loader block for the second hypervisor, the loader block including a logical construct that describes at least a part of the static system state for initial setting, obtaining, by the orchestrator, one or more system invariant conditions of the computing system, sharing the one or more system invariant conditions with the second hypervisor, and initializing the second hypervisor based on the one or more system invariant conditions The computing system according to claim 5, including.

7. The computing system according to claim 6, wherein the one or more system invariant conditions include one or more features supported by one or more hardware resources of the computing system.

8. The computing system according to claim 7, wherein the orchestrator obtains the one or more features from at least one of (1) a hypercall from the first hypervisor or (2) directly from a hardware resource.

9. The computing system according to claim 8, wherein the orchestrator obtains at least one of the features by a binary interface that discovers or queries a hardware function.

10. The sharing of the one or more system invariant conditions with the second hypervisor is intercepting, on shared memory, a request for at least one of the system invariant conditions from the first hypervisor, and issuing a reverse hypercall to the second hypervisor to transmit the at least one invariant condition to the first hypervisor via the shared memory The computing system according to claim 8, including.

Citation Information

Patent Citations

  • Logical interval type computer system

    JP2000259434A

  • High-availability system and execution state control method

    JP2009080695A

  • Fault tolerant calculator system, switch device connected to multiple physical servers and storage device, and server synchronous control method

    JP2012014239A

  • Hypervisor replacing method and information processor

    JP2012220990A