Fast access to virtual machine memory backed by the host computer's virtual memory
By skipping the surface level of SLAT in the virtual machine environment and maintaining memory correlation directly at the lowest level of the page table, the problem of low memory access efficiency in the virtual machine environment is solved, and more efficient SLAT traversal and memory access are achieved.
Patent Information
- Application Number
- CN201980076709.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2019-05-27
- Filing Date
- 2019-11-14
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2039-11-14
AI Technical Summary
In virtual machine environments, memory access efficiency is low, mainly due to latency caused by layer 2 address table (SLAT) hierarchical traversal.
By skipping or not referring to the table levels within the SLAT, memory relevance is maintained directly at the lowest level of the page table, thereby improving the efficiency of SLAT traversal and memory access.
It realizes more efficient SLAT traversal and memory access, maintaining the memory utilization efficiency advantage of using host virtual memory to support virtual machine processes.
Smart Images

Figure CN113168379B_ABST
Abstract
Description
Background Art
[0001] Modern computing devices operate by executing computer executable instructions from high-speed volatile storage in the form of random access memory (RAM), and the execution of such computer executable instructions generally requires reading data from the RAM. Due to cost, physical size limitations, power requirements, and other similar constraints, computing devices generally include less RAM than is generally required for processes executed on such computing devices. To accommodate such constraints, virtual memory is utilized, whereby it appears that more memory is available to processes executed on the computing device than is provided by the physical memory circuitry. The relationship between virtual memory and physical memory is generally managed by one or more memory managers, which implement, maintain, and / or reference "page tables," information of which describes the relationship between one or more virtual memory addresses and the location of corresponding data, whether the corresponding data is in physical memory or on some form of storage medium. To accommodate the amount of memory associated with modern computing devices and the processes executed thereon, modern page tables generally consist of multiple levels of tables, wherein higher level tables have entries that respectively identify different lower level tables, and the lowest level tables contain entries that do not identify another table, but rather identify the memory address itself.
[0002] Among the processes that can be executed by a computing device, processes that virtualize or abstract the underlying hardware of the computing device are included. Such processes include virtual machines that can simulate a complete underlying computing device as a process executed within a virtualized computing context provided by such a virtual machine. A virtual machine monitor or similar computer executable instruction set can facilitate the provision of a virtual machine by virtualizing or abstracting the underlying hardware of a physical computing device that hosts such a virtual machine monitor. The virtual machine monitor can maintain a second layer address table (SLAT), which can also be arranged in layers in a manner similar to the above-mentioned page table. The SLAT can maintain information describing the relationship between one or more memory addresses of physical memory locations that appear to be processes executing on the virtual machine monitor and the memory locations of the actual physical memory itself, including processes executed in the context of a virtual machine whose virtualization of the underlying computing hardware is facilitated by the virtual machine monitor.
[0003] For example, when a process executing in the context of a virtual machine accesses memory, two different lookups may be performed. One lookup may be performed in the virtual machine context itself to correlate the requested virtual memory address with the physical memory address. Because this lookup is performed in the virtual machine context itself, the physical memory address identified is only the physical memory address perceived by the process executing in the virtual machine context. Such a lookup may be performed by a memory manager executing in the virtual machine context and may be generated by referencing a page table present in the virtual machine context. A second lookup may then be performed outside the virtual machine context. More specifically, the physical memory address identified by the first lookup (within the context of the virtual machine) may be correlated with an actual physical memory address. Such a second lookup may require one or more processing units of the computing device to reference the SLAT, which may correlate the perceived physical memory address with the actual physical memory address.
[0004] As noted, both the page table and the SLAT can be hierarchical arrangements of different levels of tables. Thus, whether a table lookup is performed by a memory manager referencing a page table or a virtual machine monitor referencing a SLAT, it may be necessary to determine the appropriate table entry within the highest level table, reference the lower level table identified by the table entry, determine the appropriate table entry in the lower level table, reference the lower level table identified by the table entry, and so on, until the lowest level table is reached, and then the various entries of the lowest level table identify one or more specific addresses or address ranges of the memory itself, such as another table should be identified. Each reference to a lower level table consumes processor cycles and increases the duration of the memory access.
[0005] For example, in the case of a memory access from a process executing in the context of a virtual machine, the duration of such a memory access may include a traversal of the hierarchy of page tables performed by the memory manager in the context of the virtual machine, and a traversal of the SLAT hierarchy performed by the virtual machine monitor. The additional latency introduced by the lookup performed by the virtual machine monitor (referenced to the SLAT) makes memory accesses made by a process executing in the context of a virtual machine, or indeed any process accessing memory through a virtual machine monitor, less efficient than if the process accessed memory more directly. Such inefficiencies may prevent users from gaining security benefits, as well as other benefits that come with accessing memory through a virtual machine monitor. Summary of the invention
[0006] In order to improve memory utilization efficiency of virtual machine processes executing on a host computing device, these virtual machine processes can be backed by the host's virtual memory, so that traditional virtual memory efficiencies can be applied to memory consumed by such virtual machine processes. In the case where guest physical memory of a virtual machine environment is backed by virtual memory allocated to one or more processes executing on the host computing device, in order to improve the speed of traversing levels of a second-level address table (SLAT) as part of a memory access, one or more levels of tables within the SLAT can be skipped or otherwise not referenced, resulting in more efficient SLAT traversal and more efficient memory access. Although the SLAT can be populated with memory associativity at higher levels of the table, a page table of the host computing device that supports the host computing device providing virtual memory can maintain a corresponding set of contiguous memory associativity at the lowest level of the table, so that the host computing device can page out or otherwise manipulate smaller memory blocks. If such manipulation occurs, the SLAT can be repopulated with memory associativity at the lowest level of the table. Conversely, if the host can reorganize a sufficiently large set of contiguous small pages, the SLAT can be populated with associativity again at a higher level of the table, thereby again achieving more efficient SLAT traversal and more efficient memory access. In this way, more efficient SLAT traversal can be achieved while maintaining the memory utilization efficiency advantage of using the virtual memory of the host computing device to support virtual machine processes.
[0007] This Summary is provided to introduce some concepts in a simplified form, which are further described in the Detailed Description below. This Summary is neither intended to identify key features or essential features of the claimed subject matter nor to be used to limit the scope of the claimed subject matter.
[0008] Other features and advantages will become apparent from the following detailed description, which proceeds with reference to the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0009] The following detailed description may be better understood in conjunction with the accompanying drawings, in which:
[0010] Figure 1 is a block diagram of an exemplary computing device
[0011] Figure 2 is a system diagram of an exemplary hardware and software computing system;
[0012] Figure 3 is a block diagram illustrating the terms used herein;
[0013] Figure 4 is a flow chart illustrating an exemplary method for backing guest physical memory with host virtual memory;
[0014] Figure 5 is a flow chart illustrating various exemplary actions in the life cycle of data in guest physical memory backed by host virtual memory;
[0015] Figure 6 is a block diagram illustrating an exemplary mechanism for backing guest physical memory with host virtual memory;
[0016] Figure 7 is a block diagram illustrating an exemplary mechanism for more efficiently translating between guest physical memory and host physical memory while backing guest physical memory with host virtual memory; and
[0017] Figure 8 is a flow chart illustrating an exemplary method for more efficiently translating between guest physical memory and host physical memory while backing guest physical memory with host virtual memory. DETAILED DESCRIPTION
[0018] The following description relates to improving the efficiency of memory access while providing more efficient memory utilization on a computing device hosting one or more virtual machine processes providing a virtual machine computing environment. In order to improve the memory utilization efficiency of virtual machine processes executing on a host computing device, these virtual machine processes can be supported by the virtual memory of the host, so that traditional virtual memory efficiencies can be applied to the memory consumed by such virtual machine processes. In the case where the guest physical memory of the virtual machine environment is supported by the virtual memory allocated to one or more processes executing on the host computing device, in order to improve the speed of traversing the levels of a second-level address table (SLAT) as part of a memory access, one or more levels of the table within the SLAT can be skipped or otherwise not referenced, resulting in more efficient SLAT traversal and more efficient memory access. Although the SLAT can be populated with memory associativity at higher levels of the table, the page table of the host computing device supporting the host computing device to provide virtual memory can maintain a corresponding set of contiguous memory associativity at the lowest level of the table, so that the host computing device can swap out or otherwise manipulate smaller memory blocks. If such manipulation occurs, the SLAT can be repopulated with memory associativity at the lowest level of the table. Conversely, if the host can reorganize a sufficiently large set of contiguous small pages, the SLAT can be populated with associativity again at a higher level of the table, thereby again achieving more efficient SLAT traversal and more efficient memory access. In this way, more efficient SLAT traversal can be achieved while maintaining the memory utilization efficiency advantage of using the virtual memory of the host computing device to support virtual machine processes.
[0019] Although not required, the following description will be in the general context of computer-executable instructions (such as program modules) executed by computing devices. More specifically, unless otherwise indicated, the description will refer to actions and symbolic representations of operations performed by one or more computing devices or peripherals. As such, it will be understood that such actions and operations, sometimes referred to as computer-executed, include manipulations by a processing unit of electrical signals representing data in a structured form. The manipulations transform the data or maintain it at a location in a memory, which reconfigures or otherwise changes the operation of the computing device or peripheral in a manner well known to those skilled in the art. The data structures in which the data is maintained are physical locations that have specific properties defined by the data format.
[0020] Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. In addition, those skilled in the art will appreciate that computing devices need not be limited to conventional personal computers, but include other computing configurations, including servers, handheld devices, multiprocessor systems, microprocessor-based or programmable consumer electronics, network PCs, minicomputers, mainframe computers, etc. Similarly, computing devices need not be limited to stand-alone computing devices, as these mechanisms can also be practiced in distributed computing environments where tasks are performed by remote processing devices linked through a communications network. In a distributed computing environment, program modules may be located in local and remote storage devices.
[0021] Before continuing to describe the memory allocation and access mechanism mentioned above in detail, refer to Figure 1 The illustrated exemplary computing device 100 provides a detailed description of an exemplary host computing device that provides context for the following description. The exemplary computing device 100 may include, but is not limited to, one or more central processing units (CPUs) 120, a system memory 130, and a system bus 121 that couples various system components, including the system memory, to the processing unit 120. The system bus 121 may be any of several types of bus structures, including a memory bus or memory controller, a peripheral bus, and a local bus using any of a variety of bus architectures. The computing device 100 may optionally include graphics hardware, including, but not limited to, a graphics hardware interface 160 and a display device 161, which may include a display device capable of receiving touch-based user input, such as a touch-sensitive or multi-touch display device. Depending on the particular physical implementation, one or more of the CPU 120, the system memory 130, and the other components of the computing device 100 may be physically co-located, such as on a single chip. In this case, some or all of the system bus 121 may be simply silicon paths within a single chip structure, and its Figure 1 The illustrations in the figure may be for the convenience of illustration only.
[0022] The computing device 100 also typically includes computer-readable media, which may include any available media that the computing device 100 can access, and includes volatile and non-volatile media, as well as removable and non-removable media. As an example and not limitation, computer-readable media may include computer storage media and communication media. Computer storage media include media implemented in any method or technology for storing content such as computer-readable instructions, data structures, program modules, or other data. Computer storage media include, but are not limited to: RAM, ROM, EEPROM, flash memory or other storage technology, CD-ROM, digital versatile disk (DVD) or other optical disk storage, magnetic cassettes, tapes, disk storage or other magnetic storage devices, or any other medium that can be used to store desired content and can be accessed by the computing device 100. However, computer storage media do not include communication media. Communication media typically embody computer-readable instructions, data structures, program modules, or other data in modulated data signals such as carrier waves or other transmission mechanisms, and include any content delivery media. As an example and not limitation, communication media include wired media such as wired networks or direct wired connections, and wireless media such as acoustic, RF, infrared, and other wireless media. Combinations of the any of the above should also be included within the scope of computer-readable media.
[0023] System memory 130 includes computer storage media in the form of volatile and / or nonvolatile memory, such as read-only memory (ROM) 131 and random access memory (RAM) 132. A basic input / output system 133 (BIOS), containing basic routines that help transfer content between elements within computing device 100, such as during startup, is typically stored in ROM 131. RAM 132 typically contains data and / or program modules that are immediately accessible to and / or currently being operated on by processing unit 120. By way of example and not limitation, Figure 1 Operating system 134 , other program modules 135 , and program data 136 are shown.
[0024] The computing device 100 may also include other removable / non-removable volatile / non-volatile computer storage media. For example only, Figure 1A hard disk drive 141 is shown that reads or writes from a non-removable non-volatile magnetic medium. Other removable / non-removable volatile / non-volatile computer storage media that can be used with the exemplary computing device include, but are not limited to, magnetic tape cassettes, flash memory cards, digital versatile disks, digital video tapes, solid-state RAM, solid-state ROM, and other computer storage media defined and described above. The hard disk drive 141 is typically connected to the system bus 121 via a non-volatile memory interface such as interface 140.
[0025] discussed above and in Figure 1 The drives and their associated computer storage media shown in FIG. 1 provide storage of computer readable instructions, data structures, program modules and other data for the computing device 100. Figure 1 14, for example, hard drive diagram 141 is shown as storing operating system 144, other program modules 145, and program data 146. Note that these components may be the same as or different from operating system 134, other program modules 135, and program data 136. Operating system 144, other program modules 145, and program data 146 are given different numbers to illustrate that they are at least different copies.
[0026] The computing device 100 can operate in a networked environment using logical connections to one or more remote computers. The computing device 100 is shown connected to a general network connection 151 (to a network 190) through a network interface or adapter 150, which in turn is connected to a system bus 121. In a networked environment, program modules depicted with respect to the computing device 100, or portions or peripherals thereof, may be stored in the memory of one or more other computing devices communicatively coupled to the computing device 100 through a general network connection 161. It should be understood that the network connections shown are exemplary and other ways that may be used to establish a communications link between computing devices.
[0027] Although described as a single physical device, the exemplary computing device 100 may be a virtual computing device, in which case the functionality of the aforementioned physical components, such as the CPU 120, the system memory 130, the network interface 160, and other similar components, may be provided by computer executable instructions. Such computer executable instructions may be executed on a single physical computing device, or may be distributed across multiple physical computing devices, including being distributed across multiple physical computing devices in a dynamic manner, such that the specific physical computing device hosting such computer executable instructions may dynamically change over time as needed and available. In the case where the exemplary computing device 100 is a virtualized device, the underlying physical computing device hosting such a virtualized computing device may itself include physical components similar to the aforementioned components and operating in a similar manner. In addition, virtual computing devices may be utilized in multiple layers, with one virtual computing device executing within the fabric of another virtual computing device. Therefore, as used herein, the term "computing device" refers to a physical computing device or a virtualized computing environment that includes a virtual computing device, in which computer executable instructions may be executed in a manner consistent with the manner in which a physical computing device executes them. Similarly, as used herein, a term referring to a physical component of a computing device refers to either the physical component or a virtualization that performs the same or equivalent functionality of the physical component.
[0028] Go to Figure 2 , the system 200 shown therein illustrates a computing system including computing device hardware and processes executed thereon. More specifically, Figure 2 The illustrated exemplary system 200 includes the computing device hardware itself, such as the processing unit 120 and RAM 132 described above, which are shown as part of the exemplary computing device 100 described above. In addition, the exemplary system 200 includes a virtual machine monitor, such as the exemplary virtual machine monitor 210, which may include computer-executable instructions executed by the exemplary computing device 100 to enable the computing device 100 to perform functions associated with the exemplary virtual machine monitor 210. The exemplary virtual machine monitor 210 may include or may reference a layer 2 address table (SLAT), such as the exemplary SLAT 220.
[0029] The exemplary virtual machine monitor 210 can virtualize a set of computing hardware that can be comparable to the hardware of the computing device 100, or can be different therefrom, including differences in processor type and / or capacity, differences in the amount and / or type of RAM, differences in the amount and / or type of storage media, and other similar differences. Among other things, such virtualization can implement one or more virtual machine processes, such as the exemplary virtual machine processes 260 and 270. The virtual machine monitor 210 can present the appearance of executing directly on the computing device hardware 110 to both processes executing within these virtual machine contexts and to any other processes executing on top of the virtual machine monitor. Figure 2 The illustrated exemplary system 200 illustrates an operating system, such as the exemplary operating system 240, executing on top of the exemplary virtual machine monitor 210. Furthermore, as shown in the exemplary system 200, one or more computer-executable applications, such as the virtual machine processes 260 and 270 described above, may be executed on the exemplary operating system 240.
[0030] The exemplary operating system 240 may include various components, subcomponents, or aspects thereof as described below. Figure 2 132. The exemplary memory manager 241 may be a memory manager, such as the exemplary memory manager 241. More specifically, the exemplary memory manager 241 may read code or data from one or more computer-readable storage media and copy such code or data into a memory that may be supported by the RAM 132.
[0031] As used herein, the term "virtual memory" is formed with reference to a virtual machine, but rather with reference to the concept of presenting to an application or other process executing on a computing device the appearance that a greater amount of memory than physically exists in RAM 132 is accessible. Thus, for example, the virtual memory functionality provided by the operating system 240 may allow code and / or data to be loaded into memory that is ultimately backed by physical memory, such as physical memory 132, but the amount of virtual memory is greater than the physical memory. A memory manager 241 executed as part of the operating system 240 may maintain, modify, and / or utilize a page table 250 to resolve an association between one or more virtual memory addresses and one or more physical memory addresses. More specifically, the page table 250 provides address translation between one memory addressing scheme and another, different memory addressing scheme. To accomplish this, the page table 250 may include multiple levels of tables, described in further detail below, that may associate one or more memory addresses that conform to one memory addressing scheme (i.e., virtual memory addresses) with one or more memory addresses that conform to another, different memory addressing scheme (i.e., physical memory addresses).
[0032] Figure 2 The illustrated exemplary system 200 illustrates applications executing on top of an operating system 240. In particular, such applications may be virtual machine applications that may instantiate virtual machine processes such as exemplary virtual machine processes 260 and 270. Each virtual machine process may create a virtual computing environment that may host its own operating system, applications, and other similar functionality that may be hosted by a physical computing environment such as that presented by computing device 100. Figure 2 , exemplary virtual machine process 260 is shown as creating virtual machine environment 261, and exemplary virtual machine process 270 is shown as creating virtual machine environment 271. Exemplary virtual machine environment 271 is shown as hosting an operating system in the form of exemplary operating system 281, which can be similar to exemplary operating system 240, or can be a different type of operating system. For illustrative purposes, exemplary operating system 281 executing within exemplary virtual machine environment 271 is shown as including a memory manager, namely, exemplary memory manager 282, which in turn references page tables (i.e., exemplary page tables 283) to provide virtual memory within virtual machine environment 271. Such virtual memory within virtual machine environment 271 can be utilized by processes executing on top of exemplary operating system 281, such as exemplary processes 291 and 292, which are executed in Figure 2 is also shown as being executed in an exemplary virtual machine environment 271.
[0033] In a manner similar to that implemented by memory managers 241 and 282, virtual machine monitor 210 may similarly convert memory addresses from one memory addressing scheme to another different memory addressing scheme when facilitating the presentation of a set of virtual computing hardware including memory. More specifically, virtual machine monitor 210 may convert memory addresses provided by virtual machine monitor 210 as physical memory addresses to actual physical memory addresses, such as memory addresses that identify physical locations within memory 132. Such conversions may be performed by referencing another hierarchically organized table (i.e., SLAT 220), which may also provide memory address conversions between one memory addressing scheme (in this case, addresses are viewed as physical memory addresses by processes executing on top of virtual machine monitor 210) and another different memory addressing scheme (i.e., in this case, actual physical memory addresses).
[0034] According to one aspect, the virtual memory provided by the operating system 240 can be used to support the physical memory of the virtual machine, rather than using non-paged physical memory allocations on the host computing device. This allows the memory manager 241 to manage the host physical memory associated with the guest physical memory. In particular, the memory management logic (such as the memory manager 241) already existing on the host can now be utilized to manage the physical memory of the guest virtual machine. In terms of the amount of code used to implement the virtual machine monitor (such as the exemplary virtual machine monitor 210), this can allow the use of a smaller virtual machine monitor. A smaller virtual machine monitor can be safer because there is less code that can be exploited or may have errors. In addition, this can increase the virtual machine density on the host because less host physical memory can be utilized to implement the virtual machine. For example, previously, because the virtual machine process requires non-paged memory, if the exemplary computing device 100 includes 3GB of RAM 132, it can only host one virtual machine process that creates a virtual machine with 2GB of RAM. In contrast, if the memory consumed by a virtual machine process such as the exemplary virtual machine process 270 is virtual memory provided by a memory manager 241 such as the operating system 240, the 4 GB virtual memory space can provide support for two virtual machine processes, each creating a virtual machine with 2 GB of RAM. In this way, the exemplary computing device 100 can host two virtual machine processes instead of just one, thereby increasing the density of virtual machine processes executing on the exemplary computing device 100.
[0035] Other benefits include a single memory management code base that manages all memory on the system (host and virtual machines). Therefore, making improvements, fixes and / or adjustments in one code base is beneficial to everyone. In addition, this can lead to reduced engineering costs since only one code base needs to be maintained. Another example advantage can be that virtual machines can immediately benefit from density improvements such as paging, page sharing, working set aging and pruning, fault clustering, etc. Another example advantage can be that virtual memory and physical memory consumption limits can be set on virtual machines, just like other processes, enabling administrators to control system behavior. Another example advantage can be that other features can be added to the host memory manager to provide more performance, density and functionality to virtual machines (other non-virtual machine workloads may also benefit from this).
[0036] Go to Figure 3 , system 300 illustrates a system used herein to refer to a Figure 2 200 and the corresponding memory addressing scheme. More specifically, within the virtual machine environment 271, the exemplary memory manager 282 may reference the page table 283 to associate virtual memory addresses with physical memory addresses, again in the context of the virtual machine environment 271. As described above, the term "virtual memory" does not refer to memory in a virtual machine environment, but rather refers to the amount of memory available to processes executing on a computing device that is supported by, but greater than, the physical memory capacity. Therefore, in order to nominally distinguish between virtual memory in a virtual machine environment and virtual memory in a non-virtualized or "bare metal" computing environment, the adjectives "guest" and "host" will be used, with the adjective "guest" referring to the virtual machine computing environment and the adjective "host" referring to the bare metal computing environment on which virtual machine processes that provide the virtual machine environment are executed.
[0037] Thus, and as shown in system 301, the Figure 2 The page table 283 in the virtual machine environment 271 shown in FIG. 1 associates addresses in the "guest virtual memory" 311 addressing scheme with addresses in the "guest physical memory" 312 addressing scheme. Similarly, the page table 283 in the virtual machine environment 271 shown in FIG. 1 associates addresses in the "guest virtual memory" 311 addressing scheme with addresses in the "guest physical memory" 312 addressing scheme. Figure 2 Page table 250, which is a portion of an exemplary operating system 240 shown executing in a non-virtualized computing environment, associates addresses within a "host virtual memory" 321 addressing scheme with addresses within a "host physical memory" 322 addressing scheme. Figure 3301, but according to one aspect, and as will be described in further detail below, the guest physical memory 312 of the virtual machine environment can be backed by the host virtual memory 321, thereby providing the above-mentioned benefits. Such backing of the guest physical memory 312 by the host virtual memory 321 may require modifications to the SLAT 220, which provides translation between the "guest physical memory" 312 addressing scheme and the "host physical memory" 322 addressing scheme, as shown in the exemplary system 301.
[0038] Go to Figure 3 The illustrated system 302 shows a hierarchical arrangement of tables, such as can be used to implement the exemplary page table 250, the exemplary page table 283 and / or the exemplary SLAT 220. In order to accommodate the amount of memory suitable for modern computing devices and application processes, the page table 250, 283 and / or the SLAT 220 will have multiple (usually four) hierarchical tables. However, for ease of illustration, the system 302 only shows two hierarchical tables. The higher level table 340 may include multiple table entries, such as exemplary table entries 341, 342, 343 and 344. For example, the higher level table 340 may be 4KB in size and may contain 512 discrete table entries, each of which may be eight bytes in size. Each table entry may include a pointer or other similar identifier to a lower level table such as one of the exemplary lower level tables 350 and 360. Thus, for example, the exemplary table entry 341 may include an identification of the exemplary lower level table 350, as shown by arrow 371. If the higher level table 340 is 4KB in size and each table entry is eight bytes in size, at least some of the eight bytes of data may be a pointer or other similar identifier to the lower level table 350. The exemplary table entry 342 may include an identification of the lower level table 360 in a similar manner, as indicated by arrow 382. Similarly, the remaining discrete table entries of the exemplary higher level table 340 may identify unique lower level tables. Thus, in an example where the higher level table 340 is 4KB in size and includes 512 table entries, each of which is eight bytes in size, such a higher level table 340 may include 512 identifications of 512 unique lower level tables, such as the exemplary lower level tables 350 and 360.
[0039] Each lower level table may include individual table entries in a similar manner. For example, exemplary lower level table 350 may include exemplary table entries 251, 352, 353, and 354. Similarly, exemplary lower level table 360 may include table entries 361, 362, 363, and 364. As an example, each lower level table may have a size and structure equivalent to that of higher level table 340. Thus, exemplary lower level table 350 may be 4KB in size and may include 512 table entries, such as exemplary table entries 351, 352, 353, and 354, each of which may be eight bytes in size. In the example shown in system 302, the table entries of the lower level table (e.g., exemplary table entries 351, 352, 353, and 354 of exemplary lower level table 350) may each identify a continuous memory address range. For example, exemplary table entry 351 may identify memory address range 371, as indicated by arrow 391. Memory address range 371 may include memory "pages," and may be the smallest individually manageable amount of memory. Thus, actions may be performed on a page-by-page basis, such as temporarily moving information stored in volatile memory to non-volatile storage media, or vice versa, as part of implementing an amount of virtual memory that is greater than the amount of physical memory installed. Additionally, access permissions may be established on a page-by-page basis. For example, because memory range 371 is individually identified by table entry 351, access permissions may apply to all memory addresses within memory range 371. Conversely, access permissions for memory addresses within memory range 371 may be independent of the access permissions established for memory range 372, which may be uniquely identified by a different table entry (i.e., table entry 352), as indicated by arrow 392.
[0040] The amount of memory in a single memory page, such as memory range 371, may depend on various factors, including processor design and other hardware factors, communication connections, etc. In one example, a page size of memory, such as represented by memory range 371, may be 4KB.
[0041] When a request to access a memory location identified by one or more memory addresses is received, a table entry in a higher level table may be identified that corresponds to a memory address range that includes the memory address to be accessed. For example, if the memory address to be accessed includes the memory represented by memory 371, then table entry 341 may be identified in higher level table 340. Identification of table entry 341 may in turn result in identification of lower level table 350. In lower level table 350, it may be determined that the memory address to be accessed is a portion of the memory identified by table entry 351. The information in table entry 351 may then identify memory 371 that is being attempted to be accessed.
[0042] Each such traversal of a lower level of the page table increases the latency between receiving a request to access memory and returning access to the memory. According to one aspect, a table entry of a higher level table (e.g., exemplary table entry 341 of higher level table 340) may not identify a lower level table, such as exemplary lower level table 350, but may instead identify a page of memory that may correspond to the entire memory range that has been identified by the table entry in the lower level table. Thus, for example, if each page of memory 371, 372, and 373 identified by individual table entries 351, 352, and 353 of lower level table 350 is 4KB in size, then a memory range of 2MB in size may be identified by the total number of all individual table entries of lower level table 350, since exemplary lower level table 350 may include 512 table entries, each identifying a memory range of 4KB in size. In such an example, if all 2MB of memory are considered as a single 2MB memory page, such a single large memory page can be directly identified by the exemplary table entry 341 of the higher level table 340. In this case, because the exemplary table entry 341 will directly identify a single 2MB memory page, such as the exemplary 2MB memory page 370, and will not identify the lower level table, there will be no need to reference any lower level table. For example, if the page table or SLAT contains four levels, then the use of large memory pages (such as the exemplary large memory page 370) will allow one of these levels of the table to be skipped and not referenced, thereby providing memory address translation between virtual memory and physical memory by referencing only three levels of the table instead of four levels, thereby providing more efficient memory access, that is, 25% efficiency improvement. Because a page is the smallest amount of individually manageable memory, if a 2MB sized memory page is used, the memory access permissions can be the same for all 2MB of memory in such a large memory page, for example.
[0043] As another example, if there are hierarchically arranged tables above the level of exemplary table 340, and each entry in exemplary table 340 (such as entries 341, 342, 343, and 344) can address a 2MB memory page, such as exemplary 2MB memory page 370, then 512 such table entries can address 1GB of memory as a whole. If a single 1GB memory page is used, references to table 340 and subsequent references to lower level tables (such as exemplary table 350) can be avoided. Returning to the example in which the page table or SLAT includes four levels, utilizing such a huge memory page will enable two of these levels of the table to be skipped and not referenced, thereby providing memory address translation between virtual memory and physical memory by referencing only the levels of the table instead of four levels, thereby providing memory access in about half the time. Similarly, since a page is the smallest amount of individually manageable memory, if a 1GB-sized memory page is used, the memory access permissions can be the same for all 1GB of memory in such a huge memory page, for example.
[0044] As used herein, the term "large memory page" refers to a contiguous memory range that is sized so that it encompasses all memory ranges that would otherwise be identified by a single table at the lowest level and that can be uniquely and directly identified by a single table entry of a table one level above the lowest level table. In the specific size example provided above, a "large memory page" as defined herein would be a memory page of 2MB in size. However, as previously discussed, the term "large memory page" does not refer to a specific size, but rather the amount of memory in a "large memory page" depends on the hierarchical design of the page table itself and the amount of memory referenced by each table entry at the lowest level of the page table.
[0045] In a similar manner, as used herein, the term "huge memory page" refers to a contiguous memory range that is sized so that it encompasses all memory ranges that would otherwise be identified by a single table at the next lowest level and that can be uniquely and directly identified by a single table entry of a table at one level above the next lowest level table. In the specific size example provided above, as used herein, a "huge memory page" would be a memory page of 1GB in size. Again, as shown, the term "huge memory page" does not refer to a specific size, but rather the amount of memory in a "huge memory page" depends on the hierarchical design of the page table itself and the amount of memory referenced by each table entry at the lowest level of the page table.
[0046] Back to Figure 2 , refer to Figure 2The illustrated exemplary virtualization stack 230 and other virtualization stack descriptions provide a description of a mechanism by which host virtual memory (of a host computing device) may be used to support guest physical memory (of a virtual machine). A user mode process may be executed on a host computing device to provide virtual memory for supporting guest virtual machines. One such user mode process may be created for each guest machine. Alternatively, a single user mode process may be used for multiple virtual machines, or multiple processes may be used for a single virtual machine. Alternatively, virtual memory may be implemented in other ways than using user mode processes, as described below. For purposes of illustration, virtual memory assigned to virtual machine process 270 may be used to support guest physical memory presented by virtual machine process 270 in virtual machine environment 271, and similarly, virtual memory assigned to virtual machine process 260 may be used to support guest physical memory presented by virtual machine process 260 in virtual machine environment 261. Although in Figure 2 This is described in the exemplary system 200 of , but the above-mentioned user mode processes can be separated from virtual machine processes such as exemplary virtual machine processes 260 and 270.
[0047] According to one aspect, a virtualization stack such as the exemplary virtualization stack 230 can allocate host virtual memory in the address space of a designated user mode process hosting a virtual machine such as the exemplary virtual machine process 260 or 270. The host memory manager 241 can treat this memory as any other virtual allocation, which means that it can be paged, the physical pages supporting it can be changed to meet a continuous memory allocation elsewhere on the system; the physical pages can be shared with another virtual allocation in another process (which can in turn be another virtual machine supporting allocation or any other allocation on the system). At the same time, many optimizations can be performed to allow the host memory manager to treat the virtual machine supporting the virtual machine allocation specially as needed. In addition, if the virtualization stack 230 chooses to prioritize performance over density, it can perform many operations supported by the operating system memory manager 241, such as locking pages in memory to ensure that the virtual machine does not experience paging of these portions. Similarly, large pages can be used to provide higher performance for virtual machines, which will be described in detail below.
[0048] A given virtual machine may place all of its guest physical memory addresses in guest physical memory backed by host virtual memory, or it may place some of its guest physical memory addresses in guest physical memory backed by host virtual memory and some of its guest physical memory addresses in guest physical memory backed by traditional mechanisms, such as non-paged physical memory allocations from host physical memory.
[0049] When creating a new virtual machine, the virtualization stack 230 can use a user mode process to host virtual memory allocation to support guest physical memory. This can be a newly created empty process, an existing process that hosts multiple virtual machines, or a process of each virtual machine that contains other virtual machine-related virtual allocations (e.g., virtualization stack data structures) that are invisible to the virtual machine itself. The kernel virtual address space can also be used to support the virtual machine. Once such a process is found or created, the virtualization stack 230 can perform private memory virtual allocation (or section / file mapping) in its address space, and the virtual memory corresponds to the number of guest physical memories that the virtual machine should have. Specifically, the virtual memory can be a private allocation, a file mapping, a partial mapping supported by a page file, or any other type of allocation supported by the host memory manager 241. As described below, this allocation can provide access speed advantages. On the contrary, this allocation can be an allocation of any number.
[0050] Once the virtual memory allocation is completed, it can be registered with the components that will manage the physical address space of the virtual machine (such as the exemplary virtual machine environment 271) and synchronized with the host physical memory pages, which the host memory manager 241 will select to support the virtual memory allocation. These components can be the virtual machine monitor 210 and the virtualization stack 230, which can be implemented as part of the host kernel and / or driver. The virtual machine monitor 210 can manage the conversion between the guest physical memory address range and the corresponding host physical memory address range by utilizing the SLAT 220. In particular, the virtualization stack 230 can update the SLAT 220 with the host physical memory pages that support the corresponding guest physical memory pages. When a guest virtual machine (such as the exemplary virtual machine 271) performs a certain access type to a given guest physical memory address, the virtual machine monitor 210 can enable the virtualization stack 230 to have the ability to receive interception. For example, when the guest virtual machine 271 writes a certain physical address, the virtualization stack 230 can request to receive interception.
[0051] According to one aspect, when a virtual machine such as the exemplary virtual machine 271 is first created, the SLAT 220 may not contain any valid entries corresponding to the virtual machine because there may not be host physical memory addresses allocated to support the guest physical memory addresses of such a virtual machine (although as described below, in some embodiments, the SLAT 220 may be pre-populated for certain guest physical memory addresses at or about the same time as the virtual machine 271 is created). The virtual machine monitor 210 may know the range of guest physical memory addresses that the virtual machine 27 will utilize, but does not need to have any host physical memory to support them at this time. When the virtual machine process 270 begins execution, it may begin accessing its (guest) physical memory pages. As each new physical memory address is accessed, since the corresponding SLAT entry has not yet been populated with the corresponding host physical memory address, it may generate an intercept of the appropriate type (read / write / execute). The virtual machine monitor 210 may receive the guest access intercepts and forward them to the virtualization stack 230. The virtualization stack 230 may in turn reference a data structure maintained by the virtualization stack 230 to find the host virtual memory address range that corresponds to the requested guest physical memory address range (and the host process from which the supporting virtual address space was allocated, such as the exemplary virtual machine process 270). At this point, the virtualization stack 230 may know the specific host virtual memory address that corresponds to the guest physical memory address that generated the intercept.
[0052] The virtualization stack 230 may then issue a virtual fault to the host memory manager 241 in the context of a process that hosts the virtual address range, such as the virtual machine process 270. The virtual fault may be issued with the corresponding access type (read / write / execute) of the original intercept that occurred when the virtual machine 271 accessed a physical address in the guest physical memory. The virtual fault may execute the same or similar code path as a regular page fault to make the specified virtual address valid and accessible by the host CPU. One difference is that the virtual fault code path may return a physical page number that the memory manager 241 uses to make the virtual address valid. The physical page number may be a host physical memory address that supports the host virtual address, which in turn supports the guest physical memory address that originally generated the access intercept in the virtual machine monitor 210. At this point, the virtualization stack 230 may generate or update a table entry in the SLAT 220 that corresponds to the original guest physical memory address that generated the intercept, and update the table entry with the host physical memory address and the access type (read / write / execute) used to make the virtual address in the host valid. Once this is done, the guest physical memory address can be immediately accessed using the access type of the guest virtual machine 271. For example, a parallel virtual processor in the guest virtual machine 271 can immediately access such an address without encountering an intercept. Since the SLAT contains a table entry that can associate the requested guest physical address with the corresponding host physical memory address, the original intercept process can be completed, and the original virtual processor that generated the intercept can retry its instruction and continue to access memory.
[0053] If and / or when the host memory manager 241 decides to perform any action that may or will change the host physical address support of the host virtual address that is valid through the virtual fault, it can perform a translation buffer (TLB) flush for the host virtual address. It can already perform such an action to comply with the existing contract that the host memory manager 241 may have with the hardware CPU on the host. The virtualization stack 230 can intercept such a TLB flush and invalidate the corresponding SLAT entry of any host virtual address that is flushed to support any guest physical memory address in any virtual machine. The TLB flush call can identify the virtual address range that is flushed. The virtualization stack 230 can then look up the flushed host virtual address against its data structure, which can be indexed by the host virtual address to find the guest physical range that can be supported by the given host virtual address. If any such range is found, the SLAT entries corresponding to these guest physical memory addresses can be invalidated. In addition, the host memory manager can handle virtual allocations that support virtual memory in different ways as needed or when needed to optimize TLB flushing behavior (e.g., reducing SLAT invalidation time, subsequent memory interception, etc.).
[0054] The virtualization stack 230 may carefully synchronize the updates of the SLAT 220 with the host physical memory page number returned from the virtual fault (serviced by the memory manager 241) against the TLB flush performed by the host (issued by the host memory manager 241). Doing so may avoid adding complex synchronization between the host memory manager 241 and the virtualization stack 230. The physical page number returned by the virtual fault may be out of date by the time it is returned to the virtualization stack 230. For example, the virtual address may have been invalid. By intercepting the TLB flush call from the host memory manager 241, the virtualization stack 230 may know when this contention occurs and retry the virtual fault to obtain the updated physical page number.
[0055] When the virtualization stack 230 invalidates the SLAT entry, any subsequent access by the virtual machine 271 to the guest physical memory address will again generate an intercept to the virtual machine monitor 210, which is then forwarded to the virtualization stack 230 to be resolved as described above. The same process can be repeated when the guest physical memory address is first accessed for reading and then written. The write operation will generate a separate intercept because the SLAT entry can only be made valid by the "read" access type. The intercept can be forwarded to the virtualization stack 230 as usual, and a virtual fault with "write" access rights can be issued to the host memory manager 241 to obtain the appropriate virtual address. The host memory manager 241 can update its internal state (usually in a page table entry (or "PTE")) to indicate that the host physical memory page is now dirty. This can be done before allowing the virtual machine to write to its guest physical memory address, thereby avoiding data loss and / or corruption. If and / or when the host memory manager 241 decides to trim the virtual address (which will perform a TLB flush and invalidate the corresponding SLAT entry), the host memory manager 241 can know that the page is dirty and needs to be written to a non-volatile computer-readable storage medium (such as a page file on a hard drive) before reuse. In this way, the above sequence is no different than what would happen for a regular private virtual allocation for any other process running on the host computing device 100.
[0056] As an optimization, the host memory manager 241 may choose to perform a page combining pass on all of its memories 120. This may be an operation where the host memory manager 241 finds identical pages in all processes and combines them into one read-only copy of a page shared by all processes. If and / or when any combined virtual address is written, the memory manager 241 may perform a copy-on-write operation to allow the write to proceed. Such optimizations may now work transparently across virtual machines, such as exemplary virtual machines 261 and 271, to increase the density of virtual machines executing on a given computing device by combining identical pages across virtual machines, thereby reducing memory consumption. When page combining occurs, the host memory manager 241 may update the PTEs that map the affected virtual addresses. During this update, it may perform a TLB flush because the host physical memory address may change from a unique private page to a shared page for these virtual addresses. As part of this, the virtualization stack 230 may invalidate the corresponding SLAT entry, as described above. If and / or when a guest physical memory address whose virtual address is assembled to point to a shared page is read, the virtual fault solution may return the physical page number of the shared page during the intercept process, and the SLAT 220 may be updated to point to the shared page.
[0057] If the virtual machine writes to any combined guest physical memory address, the virtual fault with write access can perform a copy-on-write operation and a new dedicated host physical memory page number can be returned and updated in SLAT 220. For example, virtualization stack 230 can instruct host memory manager 241 to perform a page combining process. In some embodiments, virtualization stack 230 can specify which portion of memory to scan for combing or which processes should be scanned. For example, virtualization stack 230 can identify processes such as exemplary virtual machine processes 260 and 270 as processes to be scanned for combining, whose host virtual memory allocation supports the guest physical memory of the corresponding virtual machine (i.e., exemplary virtual machines 261 and 271).
[0058] Even when the SLAT 220 is updated to allow write access due to an execution virtual failure, the virtual machine monitor 210 can also support triggering interception of write operations to such SLAT entries if requested by the virtualization stack 230. This is useful because the virtualization stack 230 may want to know when writes occur regardless of the fact that the host memory manager 241 can accept that these writes occur. For example, live migration of virtual machines or snapshots of virtual machines may require the virtualization stack 230 to monitor writes. Even if the state of the host memory manager has been updated accordingly for writes, such as when the PTE has been marked as dirty, the virtualization stack 230 can still be notified when a write occurs.
[0059] The host memory manager 241 may be able to maintain an accurate access history for each host virtual page that supports the guest physical memory address space, just as it does for regular host virtual pages allocated in any other process address space. For example, the "access bit" in the PTE may be updated as part of handling memory interception during a virtual fault. When the host memory manager clears the access bit on any PTE, it can already flush the TLB to avoid memory corruption. As previously described, this TLB flush may invalidate the corresponding SLAT entry, which may generate an access interception again if the virtual machine accesses its guest physical memory address again. As part of handling the interception, the virtual fault processing in the host memory manager 241 may set the access bit again, thereby maintaining the correct access history for the page. Alternatively, for performance reasons, such as avoiding access interception in the virtual machine monitor 210 as much as possible, the host memory manager 241 may use page access information collected from the SLAT entry directly from the virtual machine monitor 210 (if supported by the underlying hardware). Host memory manager 241 may cooperate with virtualization stack 230 to translate access information in SLAT 220 (organized by guest physical memory addresses) into host virtual memory addresses that support these guest physical memory addresses to know which addresses are accessed.
[0060] By having an accurate page access history, the host memory manager 241 can run intelligent aging and trimming algorithms of its usual processing process working set. This can allow the host memory manager 241 to examine the state of the entire system and, as needed or for other reasons, intelligently select addresses to be trimmed and / or pages to be swapped out to disk or other similar operations to relieve memory pressure.
[0061] In some embodiments, the host memory manager 241 can impose virtual and physical memory limits on the virtual machine, just like any other process on the system. This can help system administrators sandbox or otherwise constrain or enable virtual machines, such as exemplary virtual machines 261 and 271. The host system can use the same mechanism as native processes to achieve this purpose. For example, in some embodiments where higher performance is required, the virtual machine can have all of its guest physical memory directly supported by the host physical memory. Alternatively, some parts can be supported by the host virtual memory, while other parts are supported by the host physical memory 120. In another example, where lower performance is acceptable, the virtual machine can be supported primarily by the host virtual memory, and the host virtual memory can be limited to less than the full support of the host physical memory. For example, a virtual machine such as the exemplary virtual machine 271 can present a guest physical memory of 4GB in size, and the guest physical memory of the above-mentioned process can be supported by 4GB of host virtual memory. However, the host virtual memory can be limited to only 2GB of host physical memory. This may cause paging or other performance obstacles to the disk, but can be a way for administrators to restrict or impose other controls on virtual machine deployment based on service levels. Similarly, guaranteeing a certain amount of physical memory to a virtual machine (while still enabling the virtual machine to be backed by host virtual memory) may be supported by the host memory manager to provide a certain level of performance.
[0062] The amount of guest physical memory provided within the virtual machine environment can be changed dynamically. When guest physical memory needs to be added, another host virtual memory address range can be allocated as described above. Once the virtualization stack 230 is ready to handle access interception on the memory, the guest physical memory address range can be added to the virtual machine environment. When the guest physical memory address is removed, the portion of the host virtual address range that supports the removed guest physical memory address can be released by the host memory manager 241 (and updated accordingly in the virtualization stack 230 data structure). Alternatively or additionally, various host memory manager APIs can be called on these portions of the host virtual address range to release host physical memory pages without releasing the host virtual address space for it. Alternatively, in some embodiments, nothing may be done at all, because the host memory manager 241 can eventually trim these pages from the working set and eventually write them to disk, such as a page file, because they will no longer be accessed in practice.
[0063] According to one aspect, some or all of the guest physical memory addresses may be pre-populated into the SLAT 220 to host the host physical memory address mapping. This may reduce the number of fault handling operations performed when the virtual machine is initialized. However, as the virtual machine is running, the entries in the SLAT 220 may become invalid for various reasons, and the above-described fault handling may be used to re-associate the guest physical memory addresses with the host physical memory addresses. The SLAT entries may be pre-populated prior to starting the virtual computing environment or during runtime. Alternatively, the entire SLAT 220 or only a portion thereof may be pre-populated.
[0064] Another optimization may be to prefetch other portions of the host virtual memory into the host physical memory to support the guest physical memory that was previously swapped out, so that when a subsequent memory intercept arrives, the virtual fault can be satisfied faster because it can be satisfied without going to disk to read the data.
[0065] Go to Figure 4 , a method 400 is shown, which provides an overview of the mechanism described in detail above. Method 400 may include actions for backing guest physical memory with host virtual memory. The method includes: from a virtual machine executing on a host computing device, attempting to access guest physical memory using guest physical memory access (action 402). For example, Figure 2 The illustrated virtual machine environment 271 can access guest physical memory, which appears to be actual physical memory to processes executing within the virtual machine environment 271.
[0066] Method 400 may also include determining that a guest physical memory access references a guest physical memory address that does not have a valid entry in a data structure that associates guest physical memory addresses with host physical memory addresses (act 404). Figure 2 SLAT 220 is shown with no valid entries.
[0067] As a result, method 400 may include identifying a host virtual memory address corresponding to the guest physical memory address and identifying a host physical memory address corresponding to the host virtual memory address (act 406). Figure 2 The virtualization stack 230 shown may identify host virtual memory addresses corresponding to guest physical memory addresses and also Figure 2 The illustrated memory manager 241 may identify host physical memory addresses that correspond to host virtual memory addresses.
[0068] Method 400 may also include updating a data structure that correlates the guest physical memory address with the host physical memory address using the association of the guest physical memory address with the identified host physical memory address (act 408). Figure 2 The virtualization stack 230 shown may be derived from Figure 2 The memory manager 241 shown obtains the host physical memory address and updates the SLAT 220 (eg, Figure 2 shown).
[0069] Method 400 can be practiced by causing an interception. The interception can be forwarded to a virtualization stack on the host. This may cause the virtualization stack to identify a host virtual memory address corresponding to the guest physical memory address and issue a fault to the memory manager to obtain a host physical memory address corresponding to the host virtual memory address. The virtualization stack can then utilize the association of the guest physical memory address and the identified host physical memory address to update a data structure that associates the guest physical memory address with the host physical memory address.
[0070] The method 400 may also include determining a type of the guest physical memory access and updating a data structure associating the guest physical memory address with the host physical memory address using the determined type associated with the guest physical memory address and the identified host physical memory address. For example, if the guest physical memory access is a read, the SLAT may be updated to indicate so.
[0071] Method 400 may also include performing an action that can change the host physical memory address that supports the host virtual memory address. As a result, the method may include invalidating an entry in a data structure that associates the guest physical memory address with the host physical memory address, and the data structure associates the guest physical memory address with the host physical memory address. This may cause subsequent accesses to the guest physical memory address to generate a fault, which may be used to update the data structure that associates the guest physical memory address with the host physical memory address using the correct association of the host virtual memory that supports the guest physical memory address. For example, the action may include a page combination operation. Page combination may be used to increase the density of virtual machines on the host.
[0072] Method 400 may include initializing a guest virtual machine. As part of initializing the guest virtual machine, method 400 may include pre-populating at least a portion of a data structure associating guest physical memory addresses with host physical memory addresses using some or all of the guest virtual machine's guest physical memory address to host physical memory address mappings. Thus, for example, host physical memory may be pre-allocated for the virtual machine and the appropriate association entered into the SLAT. This will result in fewer exceptions required to initialize the guest virtual machine.
[0073] Reference now Figure 5 , an example flow 500 is shown, which illustrates various actions that may occur during a portion of the life cycle of certain data at an example guest physical memory address (hereinafter referred to as "GPA") 0x1000. As shown at 502, the SLAT entry for GPA 0x1000 is invalid, which indicates that there is no SLAT entry for this particular GPA. As shown at 504, the virtual machine attempts to perform a read at GPA 0x1000, resulting in a virtual machine (hereinafter referred to as "VM") read intercept. As shown at 506, the virtual machine monitor forwards the intercept to the virtualization stack on the host. At 508, the virtualization stack performs a virtualization lookup on the host virtual memory address (VA) corresponding to GPA 0x1000. The lookup result is VA 0x8501000. At 510, a virtual fault is generated for the read access to VA 0x8501000. The system physical address (hereinafter referred to as "SPA") returned by the virtual fault processing of the memory manager is 0x88000, which defines the address where the data at GPA 0x1000 in the system memory is physically located. Therefore, as shown in 514, the SLAT is updated to associate GPA 0x1000 with SPA 0x88000 and mark the data access as "read-only". At 516, the virtualization stack completes the read interception process, and the virtual machine monitor resumes the guest virtual machine execution.
[0074] As shown in 518, a period of time passes. At 520, the virtual machine attempts a write access at GPA 0x1000. At 522, the virtual machine monitor forwards the write access to the virtualization stack on the host. At 524, the virtualization stack performs a virtualization lookup for the host VA for GPA 0x1000. As previously described, this is at VA 0x8501000. At 526, a virtual fault occurs for the write access at VA 0x8501000. At 528, the virtual fault returns SPA 0388000 in physical memory. At 530, the SLAT entry for GPA 0x1000 is updated to indicate that the data access is "read / write". At 532, the virtualization stack completes the write intercept processing. The virtual machine monitor resumes execution of the guest virtual machine.
[0075] As shown in 534, a period of time passes. At 536, the host memory manager runs a page combination traversal to combine any pages in the host physical memory that are functionally identical. At 538, the host memory manager finds a combination candidate for VA 0x8501000 and another virtual address in another process. At 540, the host performs a TLB flush on VA 0x8501000. At 542, the virtualization stack intercepts the TLB flush. At 544, the SLAT entry for GPA 0x1000 is invalidated. At 546, a virtual machine intercept is performed for GPA 0x1000. At 548, a virtual fault occurs for a read access to VA 0x8501000. At 550, the virtual fault returns SPA0x52000, which is a shared page between the N processes in the page combination traversal of 536. At 552, the SLAT entry for GPA 0x1000 is updated to be associated with SPA 0x52000, where the access permission is set to "read only".
[0076] As shown in 554, a period of time passes. At 556, a virtual machine write intercept occurs for GPA 0x1000. At 558, a virtual fault for write access occurs on VA 0x8501000. At 560, the host memory manager performs a copy write to VA 0x8501000. At 562, the host performs a TLB flush on VA 0x850100. As shown in 564, this invalidates the SLAT entry for GPA 031000. At 566, the virtual fault returns SPA 0x11000, which is a private page after the copy write. At 568, the SLAT entry for GPA 0x1000 is updated to SPA 0x1000, where the access permission is set to "read / write". At 570, the virtualization stack completes the read intercept processing and the virtual machine monitor resumes virtual machine execution.
[0077] Thus, as described above, the virtual machine physical address space is backed by host virtual memory (typically allocated in the user address space of the host process), which is managed by the host memory manager for regular virtual memory. The virtual memory backing the physical memory of the virtual machine can be of any type supported by the host memory manager 118 (private allocation, file mapping, partial mapping backed by a page file, large page allocation, etc.). The host memory manager can perform its existing operations and apply policies and / or apply dedicated policies on the virtual memory to know that the virtual memory supports the physical address space of the virtual machine when necessary.
[0078] Go to Figure 6 , as described above, the exemplary system 600 shown therein includes Figure 4 Overview and Figure 5More specifically, guest virtual memory 311, guest physical memory 312, host virtual memory 321, and host physical memory 322 are shown as rectangular boxes, whose widths represent the range of the memory. To simplify the description, the widths of the rectangular boxes are also roughly equal, even though in actual operation, the amount of virtual memory may exceed the amount of physical memory supporting such virtual memory.
[0079] As mentioned above, when in a virtual machine computing environment (such as Figure 2When a process executing in the exemplary virtual machine computing environment 271 shown accesses a portion of virtual memory provided by an operating system executing in such a virtual machine computing environment, a page table of the operating system (such as the exemplary page table 283) may include a PTE that can associate the accessed virtual memory address with a physical memory address, at least as perceived by the process executing within the virtual machine computing environment. Thus, for example, if the virtual memory represented by region 611 is accessed, a PTE 631 that can be part of the page table 283 can associate the region 611 of the guest virtual memory 311 with a region 621 of the guest physical memory 312. As described in further detail above, the guest physical memory 312 can be backed by a portion of the host virtual memory 321. Thus, the virtualization stack 230 described above can include a data structure of entry 651 that can associate a region 621 of the guest physical memory 312 with a region 641 of the host virtual memory 321. Then, a process executing on the host computing device, including the above-described page table 250, may associate the region 641 of the host virtual memory 321 with the region 661 of the host physical memory 322. As also detailed above, the SLAT 220 may include a hierarchically arranged set of tables that may be updated with a table entry 691 that may associate the region 621 of the guest physical memory 312 with the region 661 of the host physical memory 322. More specifically, mechanisms such as those implemented by the virtualization stack 230 may detect the association between the region 621 and the region 661, and may update or generate a table entry in the table SLAT 220, such as the entry 691, as shown in action 680, to include such an association. In this manner, the guest physical memory 312 may be backed by the host virtual memory 321 while continuing to allow existing mechanisms, including the SLAT 220, to function in their conventional manner. For example, a subsequent access to region 611 of guest virtual memory 311 by a process executing in the virtual machine computing environment may require two page table lookups in a conventional manner: (1) a page table lookup performed in the virtual machine computing environment by referencing page table 283, which may include entries such as entry 631 that may associate the requested region 611 with a corresponding region 621 in guest physical memory 312, and (2) a page table lookup performed on the host computing device by referencing SLAT 220, which may include entries such as entry 691 that may associate region 621 in guest physical memory 312 with region 661 in host physical memory 322, thereby completing the path to the relevant data stored in the host's physical memory.
[0080] As described above, various page tables, such as the exemplary page table 283 in a virtual memory computing environment, the exemplary page table 250 on a host computing device, and the exemplary SLAT 220, may include hierarchically arranged layers of tables such that entries at the lowest level of the table may identify a "small page" of memory, as that term is explicitly defined herein, which may be the smallest memory address range that may be individually and carefully maintained and manipulated by a process using a corresponding page table. Figure 6 In the exemplary system 600 shown, various page table entries (PTEs) 631, 671, and 691 may be at the lowest level of the table, such that regions 611, 621, 641, and 661 may be small memory pages. As described above, in a common microprocessor architecture, such a small memory page may include 4KB of contiguous memory addresses, as explicitly defined herein.
[0081] According to one aspect, to improve the speed and efficiency of memory accesses from a virtual machine computing environment whose physical memory is backed by the virtual memory of a host computing device, SLAT 220 may maintain associations at a higher level of the table. In other words, SLAT 220 may associate "large pages" of memory or "huge pages" of memory, as these terms are clearly defined herein. In one common microprocessor architecture, a large memory page, as that term is clearly defined herein, may include 2MB of contiguous memory addresses, while a huge memory page, as that term is clearly defined herein, may include 1GB of contiguous memory addresses. Again, as previously described, the terms "small page," "large page," and "huge page" are clearly defined with reference to the level in a hierarchically arranged set of page tables, rather than being defined based on a specific amount of memory, as the amount of memory may vary between different microprocessor architectures, operating system architectures, etc.
[0082] Go to Figure 7 , shows a mechanism for creating large page entries in SLAT 220, thereby improving the speed and efficiency of accessing memory from a virtual machine computing environment whose physical memory is backed by the virtual memory of a host computing device. Figure 6 As shown, in Figure 7 In the illustrated system 700, guest virtual memory 311, guest physical memory 312, host virtual memory 321, and host physical memory 322 are illustrated as rectangular boxes, the width of which represents the range of the memory. In addition, access to region 611 may be performed as described in detail above and as follows: Figure 6250 . However, in addition to generating the lowest level table entries within SLAT 220, the mechanisms described herein may generate higher level table entries, such as exemplary table entry 711. As described above, such higher level table entries may associate large pages or huge pages of memory. Thus, when virtualization stack 230 causes table entry 711 to be entered into SLAT 220, table entry 711 may associate large page size region 761 of guest physical memory 312 with large page size region 720 of host physical memory 322. However, as will be described in detail below, from the perspective of a memory manager executing on a host computing device and utilizing page table 250, such large page size region is not a large page, and therefore, such large page size region, for example, cannot be swapped out as a single large page by a memory manager executing on the host computing device, nor is it otherwise considered as a single indivisible large page.
[0083] Conversely, in order to maintain the efficiency, density, and other advantages provided by the mechanisms detailed above that enable guest physical memory, such as host virtual memory 312, to be backed by host virtual memory, such as exemplary host virtual memory 321, the page table 250 executing on the host may not include higher level table entries similar to higher level table entry 711 in SLAT 220, but may include only lower level table entries, such as Figure 7 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91
[0084] Because virtual faults may have been generated for less than a large page size memory region, only certain entries in the lower level table entries, such as exemplary lower level table entries 671, 731, and 732, may be utilized, while the remaining lower level table entries may still be marked as free memory. For example, an application executed within a virtual machine computing environment utilizing guest physical memory 312 may initially require only a small page size memory. In this case, a single higher level table entry 711 may identify a large page size memory region used, such as large page size region 720. However, from the perspective of a memory manager of a host computing device, referencing page table 250, only a single lower level table entry (such as exemplary lower level table entry 671) may be indicated as having been utilized, and the remaining lower level table entries are indicated as available. In order to prevent other entries from being considered available, the virtualization stack 230 may also mark the remaining lower level table entries, such as exemplary lower level table entries 731, 732, etc., as used. When an application executing in the virtual machine computing environment needs to utilize additional small page size memory, the virtualization stack 230 can utilize memory locations previously indicated as used but not actually used by the application executing in the virtual machine computing environment to meet further memory needs of the application.
[0085] According to one aspect, tracking of identical sequences of lower level table entries in page table 250 (such as the above-mentioned tags) can be coordinated by virtualization stack 230, such as Figure 7 In this case, if the memory manager executing on the host makes a change to one or more of these lower-level table entries, the virtualization stack 230 can trigger appropriate changes to the corresponding entries in the SLAT 220 (such as the higher-level table entry 711).
[0086] The sequence of lower-level table entries in the page table 250 (equivalent to the higher-level table entry 711 in the SLAT 220) can associate a sequence of host virtual memory ranges (such as the exemplary ranges 641, 741, and 742) to host physical memory ranges, such as the exemplary ranges 661, 721, and 722, respectively. Thus, large page size regions of the host virtual memory 321 (such as the exemplary large page size region 751) are associated with large page size regions of the host physical memory 322 (such as the exemplary large page size region 720) through the sequence of lower-level table entries in the page table 250 (i.e., the exemplary lower-level table entries 671, 731, 732, etc.). The virtualization stack 230 can maintain a data structure, such as a hierarchically arranged table, that can coordinate host virtual memory locations to guest physical memory locations. Thus, the virtualization stack may maintain one or more data entries, such as the exemplary data entry 771, which may associate the large page size region 751 of the host virtual memory 321 with the corresponding large page size region 761 of the guest physical memory 312. Although the data entry 771 is shown as a single entry, the virtualization stack 230 may maintain the above association as a single entry, as multiple entries, such as by associating sub-portions of the region 761 with the region 751 individually, including sub-portions that are sized as small pages, or a combination thereof. To complete the cycle, the above-mentioned higher-level table entry 711 in the SLAT 220 may then associate the large page size region 761 in the guest physical memory 312 with the large page size region 720 in the host physical memory 322. Again, from the perspective of SLAT 220, having only a single higher level table entry 711, large page size regions 720 and 761 are large pages handled uniformly by the SLAT (according to the single higher level table entry 711), whereas from the perspective of the memory manager executing on the host, page table 250, and indeed virtualization stack 230, such regions are not unitary but include individually manageable small page regions identified, for example, by lower level table entries in page table 250.
[0087] The availability of contiguous small memory pages (equivalent to one or more large or huge memory pages, as these terms have been defined herein) in the guest physical memory 312 can be direct because the guest physical memory 312 is perceived as physical memory by processes executing in the corresponding virtual machine computing environment. Therefore, the following description focuses on the continuity in the host physical memory 322. In addition, because the guest physical memory 312 is perceived as physical memory, the entire memory access rights (such as "read / write / execute" rights) are still retained within the entire range of the guest physical memory 312, and therefore, it is less likely that there will be discontinuities in memory access rights between lower-level table entries in the page table 250.
[0088] Considering the page table 250 of the host computing device, if one or more large page size regions of the host physical memory 322 remain unused, then according to one aspect, when a virtual fault is triggered on a small page within a memory range partitioned into large page size regions that remain unused, the remaining small pages within the large page size regions may generate a single higher level table entry (e.g., exemplary higher level table entry 711) in the SLAT 220.
[0089] However, because from the perspective of the page table 250, the large page 720 corresponding to the higher level table entry 711 is a sequence of small pages, such as the exemplary small pages 661, 721, 722, etc., due to the sequence of lower level table entries, such as the exemplary lower level table entries 671, 731, 732, etc., the memory management mechanism using such a page table 250 can treat each small page separately, which may include swapping out such small pages to a non-volatile storage medium (such as a hard drive). In this case, according to one aspect, the higher level table entry 711 in the SLAT 220 can be replaced with an equivalent sequence of lower level table entries. More specifically, the lower level table entries generated within the SLAT 220 can be contiguous and can reference the same memory address range as the higher level table entry 711 as a whole, except that the specific lower level table entry corresponding to the page swapped out by the memory manager may be lost, or may not be created.
[0090] According to other aspects, if a large page or huge page size region of the host physical memory 322 is not available when a virtual fault is triggered as described above, an available large page or huge page size region of the host physical memory can be constructed at the time of interception by assembling an appropriate number of contiguous small pages. For example, a large page size region can be constructed from contiguous small pages or other similar enumerations of available host physical memory pages that can be obtained from a "free list". As another example, if small pages currently in use by other processes are needed to establish the continuity necessary to assemble a sufficient number of small pages into a large page size region, then these small pages can be obtained from such other processes, and the data contained therein can be transferred to other small pages or swapped out to disk. Once such a large page size region is constructed, the processing of the virtual fault can be performed in the manner detailed above, including generating a higher level table entry in the SLAT 220, such as the exemplary higher level table entry 711.
[0091] In some cases, a memory manager executing on a host computing device may implement a background or opportunistic mechanism when utilizing page table 250 by which an available large page or huge page size region is constructed from a sufficient number of contiguous small pages. To the extent that such a mechanism results in an available large page size region, for example, when a virtual fault is triggered, processing may proceed as noted above. However, to the extent that such a mechanism has not yet completed assembling an available large page size region, construction of such a large page size region may be completed at a higher priority upon interception. For example, the process constructing such a large page size region may be given a higher execution priority, or may be moved from the background to the foreground, or may be executed more aggressively.
[0092] Alternatively, if a large page size region of the host physical memory is not available when a virtual fault is triggered, a plurality of small pages may be utilized to process as described above. Subsequently, when a large page size region is available, the data in the plurality of previously used small pages (due to the large page size region being unavailable at the time) may be copied to the large page, and the previously used small pages may then be released. In this case, a lower level table entry in SLAT 220 may be replaced with a higher level table entry, such as exemplary higher level table entry 711.
[0093] Additionally, because of the skipping of table levels during SLAT lookups and the initial delay in constructing a large page size region from small pages, the parameters under which a large page size region is constructed can be varied based on a balance between faster memory accesses, either at interception or opportunistically. One such parameter can be the number of contiguous small pages sufficient to trigger the swapping or swapping out of other small pages required to complete contiguity throughout the large page size region. For example, if 512 contiguous small pages are required to construct a large page and there are contiguous ranges of 200 small pages and 311 small pages, and one of the small pages between them is currently in use by another process, such small pages can be swapped out, or swapped with another available small page, so that all 512 contiguous small pages can be constructed as a large page size region. In such an example, the fragmentation of the available small pages may be very low. In contrast, a contiguous range of only 20-30 small pages that is interrupted contiguously by other small pages currently in use by other processes can describe a small page range with very large fragmentation. While such a highly fragmented set of small pages may still be swapped, swapped out, or otherwise reorganized to generate a contiguous range of small pages equivalent to a large page size region, such work may take longer. Therefore, according to one aspect, a fragmentation threshold may be set that may describe whether work is expended to construct a large page size region from a range of small pages with higher or lower fragmentation.
[0094] According to one aspect, the construction of large page size regions can take advantage of existing utilization of small pages by processes whose memory is used to support guest physical memory of a virtual machine computing environment. As a simple example, if a process whose memory supports guest physical memory 312 has used a contiguous range of 200 small pages and 311 small pages, and only one small page between these two contiguous ranges is needed to establish a contiguous range of small pages equivalent to large pages, then once the small page is released (such as by swapping out or swapping), the only data that needs to be copied may be the data corresponding to the small page, and the rest of the data may remain in place, even though the corresponding lower level table entry in SLAT 220 may be invalidated and replaced with a single higher level table entry containing the same memory range.
[0095] In order to avoid having to reconstruct large page size regions of memory from contiguous ranges of small pages, an existing set of contiguous lower level table entries in the page table 250 that includes the equivalent memory range as a single higher level table entry 711 in the SLAT 220 as a whole may be locked so that they can be prevented from being swapped out. Thus, for example, the virtualization stack 230 may request to lock the lower level table entries 671, 731, 732, etc. so that the corresponding ranges of the host physical memory 661, 721, 722, etc. may retain the data stored therein without swapping it out to disk. According to one aspect, the decision such as the request made by the virtualization stack 230 to lock such entries may depend on various factors. For example, one of such factors may be the frequency of use of the data stored in these memory ranges, where more frequently used data causes the entries to be locked, while less frequently used data causes the entries to be subsequently unlocked. Other factors may also be referenced in deciding whether to request to lock the page table entry that supports the higher level table entry in the SLAT, such as the amount of memory fragmentation of the entire system, the need for a higher density virtual machine computing environment on the host computing device, and other similar factors.
[0096] Steering Figure 8 , wherein a flowchart 800 is shown illustrating the above mechanism as an exemplary series of steps. Initially, at step 810, a virtual fault as described above may occur. Subsequently, at step 815, it may be determined whether the virtual fault of step 810 references memory within a large page size memory region, wherein the virtual fault has not previously referenced other memory within the large page size region. In particular, the host physical memory may be divided into segments of a large page size. A first virtual fault that references memory within a large page size region may result in a positive determination being made at step 815, whereby processing may proceed to step 835.
[0097] At step 835, a single higher level table entry in the SLAT may be generated, and accordingly, at step 840, a plurality of lower level table entries in the host page table may be utilized, wherein the relevant number of the plurality of lower level table entries refers to a memory region having the same range as the generated single higher level table entry in the SLAT, i.e., a memory range corresponding to a large page size memory region containing a memory location whose access permissions triggered the virtual fault of step 810. As previously described, since the host page table includes a plurality of lower level table entries, rather than a single higher level table entry, the memory management of the host computing device may choose to swap out one or more memory regions or pages (identified by such lower level table entries in the host page table), such as by copying their contents to disk. If such a swap-out occurs, as determined at step 845, then processing may proceed to step 850, and the higher level table entry in the SLAT may be replaced with a plurality of lower level table entries corresponding to the memory regions referenced by the lower level table entries in the host page table that were not swapped out (and those level table entries corresponding to the memory regions referenced by the lower level table entries in the host page table that were actually swapped out are removed from the SLAT).
[0098] According to one aspect, if in step 845 a small page is swapped out from a large page size memory region, and the higher level table entries in the SLAT are invalidated and replaced with a corresponding sequence of lower level table entries (except for any lower level table entries that reference the swapped out small page), then optionally, processing can proceed to step 855 to continue checking for available large page size regions of free memory. As described above, for example, a memory management process executing on a host computing device can reconstruct large page size regions of available memory in a timely manner. If a large page size region of memory becomes available as determined in step 855, then processing can proceed to step 860, in which the lower level table entries in the SLAT can be replaced with a single higher level table entry in the SLAT, and, if appropriate, data can be copied from the memory region referenced by the previous lower level table entry to the memory region now referenced by the single higher level table entry (if the two memory regions are different).
[0099] Returning to step 815, if the virtual fault of step 810 is not the first fault that references memory within the previously described large page size region, then processing may proceed to step 820. At step 820, a determination may be made as to whether a large page should be created upon interception, such as in the manner detailed above. More specifically, if steps 835 and 840 were previously performed for a large page size memory region, and the SLAT now includes a single higher level table entry and the page table includes equivalent lower level table entries that reference the memory region referenced by the single higher level table entry in the SLAT, then subsequent memory accesses to any memory location in the memory region referenced by the single higher level table entry in the SLAT will not trigger a virtual fault, as detailed above. Thus, if a virtual fault is triggered at step 810, and at step 815, it is determined that the virtual fault is not first directed to a memory location within the previously depicted large page size memory region, then the virtual fault triggered at step 810 may have been triggered because steps 835 and 840 were previously executed for a memory location within the large page size region, but then, as determined at step 845, one or more small pages comprising the large page size memory region were swapped out, and therefore, step 850 was previously executed. Thus, at step 820, it may be determined whether a large page size memory region should be constructed upon interception. If such a large page size region should be constructed, processing may proceed to step 830, and an appropriate number of contiguous small pages may be assembled, as described above. Processing may then continue with step 835 as before. Conversely, if on-demand construction of a large page size region of available memory is not requested at step 830, then according to the above reference Figure 4 and 5 With the mechanism described in detail, processing can proceed to utilize the host page table and lower level table entries in the SLAT.
[0100] As a first example, a method for improving the access speed of a computer memory is described above, the method comprising: detecting a first memory access directed to a first memory range from a first process executed in a virtual machine computing environment; as a prerequisite for completing the first memory access, generating a first entry in a hierarchically arranged two-layer address translation table, the hierarchically arranged two-layer address translation table associating a host physical memory address with a guest physical memory address, the first entry being at least one level above the lowest level of the table so as to enable identification of a second memory range identified by the first entry without referencing any table in the lowest level of the table; and in response to generating the first entry in the hierarchically arranged two-layer address translation table, marking a first plurality of entries in a hierarchically arranged page table as used, the hierarchically arranged page table associating the host physical memory address with a host virtual memory address. The invention relates to a method of storing a guest physical memory address associated with a host computer system, wherein the first plurality of entries as a whole reference the same second memory range as the first entry in the hierarchically arranged layer-2 address translation table, wherein entries in the first plurality of entries are at a lowest level of the table; wherein the guest physical memory address is perceived by a process executing within the virtual machine computing environment as an address to physical memory; wherein the host physical memory address is an address to actual physical memory of a host computing device hosting the virtual machine computing environment; wherein the host virtual memory address is an address to virtual memory provided by a memory manager executing on the host computing device, the memory manager utilizing the hierarchically arranged page tables to provide the virtual memory; and wherein the guest physical memory identified by the guest physical memory address is backed by a portion of the host virtual memory identified by a portion of the host virtual memory address.
[0101] The second example is a method according to the first example, further comprising: detecting a second memory access from the first process executed in the virtual machine computing environment, the second memory access being directed to a third memory range different from the first memory range; and satisfying the second memory access by referencing a second subset of the first plurality of entries in the hierarchically arranged page table; wherein the first memory access is satisfied by referencing a first subset of the first plurality of entries in the hierarchically arranged page table, to a third memory range different from the first memory range.
[0102] The third example is a method according to the first example, further comprising: after the first memory access is completed, detecting that a first subset of the first plurality of entries in the hierarchically arranged page table already has data that was originally stored in the corresponding host physical memory address and subsequently swapped out to a non-volatile storage medium; in response to the detection, invalidating the first entry in the hierarchically arranged second-layer address translation table; and generating a second plurality of entries in the hierarchically arranged second-layer address translation table to replace the first entry in the hierarchically arranged second-layer address translation table, the second plurality of entries as a whole referencing at least some of the same second memory range as the first entry, wherein the entries in the second plurality of entries are at the lowest level of the table.
[0103] A fourth example is the method of the third example, wherein the second plurality of entries references a portion of the second memory range that was previously accessed by the first process and has not been swapped out.
[0104] A fifth example is the method according to the first example, further comprising: assembling a first plurality of consecutive small page size regions of the host physical memory into a single large page size region of the host physical memory.
[0105] A sixth example is the method according to the fifth example, wherein the assembling occurs after the detecting the first memory access and before the generating the first entry in the hierarchically arranged layer-2 address translation table.
[0106] The seventh example is a method according to the fifth example, further comprising: generating a second entry in the hierarchically arranged two-layer address translation table, the second entry in the at least one level above the lowest level of the table enabling identification of a third memory range identified by the second entry without referencing any table in the lowest level of the table, the third memory range referencing the single large page size area of the host physical memory into which the first plurality of consecutive small page size areas of the host physical memory are assembled; and in response to generating the second entry in the hierarchically arranged two-layer address translation table, marking the second plurality of entries in the hierarchically arranged page table as used, the second plurality of entries referencing the first plurality of consecutive small page size areas of the host physical memory assembled into the single large page size area of the host physical memory.
[0107] The eighth example is a method according to the seventh example, further comprising: generating a second entry in the hierarchically arranged two-layer address translation table, the second entry in the at least one level above the lowest level of the table enabling identification of a third memory range identified by the second entry without referencing any table in the lowest level of the table, the third memory range referencing the single large page size area of the host physical memory into which the first plurality of consecutive small page size areas of the host physical memory are assembled; and in response to generating the second entry in the hierarchically arranged two-layer address translation table, marking the second plurality of entries in the hierarchically arranged page table as used, the second plurality of entries referencing the first plurality of consecutive small page size areas of the host physical memory assembled into the single large page size area of the host physical memory.
[0108] The ninth example is a method according to the seventh example, further comprising: copying data from a second group of one or more small page size areas of the host physical memory to at least a portion of the first plurality of continuous small page size areas; and invalidating a second plurality of entries in the hierarchically arranged second-layer address translation table, the second plurality of entries comprising both of the following: (1) a first subset of entries referencing the second group of one or more small page size areas of the host physical memory, and (2) a second subset of entries referencing at least some of the first plurality of continuous small page size areas of the host physical memory, wherein entries in the second plurality of entries are at the lowest level of the table; wherein the generated second entries in the hierarchically arranged second-layer address translation table are utilized to replace the invalid second plurality of entries.
[0109] The tenth example is a method according to the fifth example, wherein the assembling of the first plurality of consecutive small page size areas of the host physical memory includes copying data from some of the first plurality of consecutive small page size areas to other small page size areas of the host physical memory that are different from the first plurality of consecutive small page size areas.
[0110] The eleventh example is a method according to the tenth example, wherein the copying of the data from some of the first plurality of consecutive small page size areas to the other small page size areas is performed only when fragmentation of the first plurality of consecutive small page size areas is below a fragmentation threshold.
[0111] A twelfth example is the method according to the first example, further comprising: preventing paging of the second memory range.
[0112] A thirteenth example is the method according to the twelfth example, wherein the preventing the paging of the second memory range is performed only when an access frequency of one or more portions of the second memory range is greater than an access frequency threshold.
[0113] The fourteenth example is a method according to the twelfth example, further comprising: if the access frequency of one or more parts of the second memory range is less than an access frequency threshold, eliminating the blocking of paging of the second memory range.
[0114] A fifteenth example is the method according to the first example, wherein the size of the second memory range is 2MB.
[0115] The sixteenth example is a computing device, the computing device comprising: one or more central processing units; a random access memory (RAM); one or more computer-readable media, comprising: a first set of computer-executable instructions, which when executed by the computing device causes the computing device to provide a memory manager, the memory manager referencing a hierarchically arranged page table to convert a host virtual memory address to a host physical memory address identifying a location on the RAM; a second set of computer-executable instructions, which when executed by the computing device causes the computing device to provide a virtual machine computing environment, wherein a process executed within the virtual machine computing environment perceives a guest physical memory address as an address to a physical memory, and wherein additional guest physical memory identified by the guest physical memory address is supported by a portion of the host virtual memory identified by a portion of the host virtual memory address; and a third set of computer-executable instructions, which when executed by the computing device causes the computing device to Preparation: detecting a first memory access directed to a first memory range from a first process executed in a virtual machine computing environment; as a prerequisite for completing the first memory access, generating a first entry in a hierarchically arranged two-layer address translation table, the hierarchically arranged two-layer address translation table associating a host physical memory address with a guest physical memory address, the first entry being at least one level above the lowest level of the table so as to enable identification of a second memory range identified by the first entry without referencing any table in the lowest level of the table, the second memory range being larger than the first memory range; and in response to generating the first entry in the hierarchically arranged two-layer address translation table, marking a first plurality of entries in the hierarchically arranged page table as used, the first plurality of entries as a whole referencing the same second memory range as the first entry in the hierarchically arranged two-layer address translation table, wherein entries in the first plurality of entries are at the lowest level of the table.
[0116] The seventeenth example is a computing device according to the sixteenth example, wherein the third set of computer-executable instructions includes additional computer-executable instructions that, when executed by the computing device, cause the computing device to perform the following operations: after the first memory access is completed, detect that a first subset of the first plurality of entries in the hierarchically arranged page table already has data that was originally stored in the corresponding host physical memory address and was subsequently swapped out to a non-volatile storage medium; in response to the detection, invalidate the first entry in the hierarchically arranged second-layer address translation table; and generate a second plurality of entries in the hierarchically arranged second-layer address translation table in place of the first entry in the hierarchically arranged second-layer address translation table, the second plurality of entries as a whole referencing at least some of the same second memory range as the first entry, wherein entries in the second plurality of entries are at the lowest level of the table.
[0117] The eighteenth example is a computing device according to the sixteenth example, wherein the third set of computer-executable instructions includes additional computer-executable instructions that, when executed by the computing device, cause the computing device to perform the following operations: assembling a first plurality of consecutive small page size regions of the host physical memory into a single large page size region of the host physical memory; wherein the assembling occurs after detecting the first memory access and before generating the first entry in the hierarchically arranged layer-2 address translation table.
[0118] A nineteenth example is the computing device of the sixteenth example, wherein the third set of computer-executable instructions includes further computer-executable instructions that, when executed by the computing device, cause the computing device to: prevent paging of the second memory range.
[0119] The twentieth example is one or more computer-readable storage media including computer-executable instructions, which, when executed, cause a computing device to: detect a first memory access directed to a first memory range from a first process executing in a virtual machine computing environment; generate, as a prerequisite for completing the first memory access, a first entry in a hierarchically arranged second-layer address translation table, the hierarchically arranged second-layer address translation table associating a host physical memory address with a guest physical memory address, the first entry being at least one level above a lowest level of the table so as to enable identification of a second memory range identified by the first entry without referencing any table in the lowest level of the table; and in response to generating the first entry in the hierarchically arranged second-layer address translation table, mark a first plurality of entries in a hierarchically arranged page table as used, the hierarchically arranged page table associating the host physical memory address with a guest physical memory address. The address is associated with a host virtual memory address, the first plurality of entries as a whole refer to the same second memory range as the first entry in the hierarchically arranged layer 2 address translation table, wherein entries in the first plurality of entries are at a lowest level of the table; wherein the guest physical memory address is perceived by a process executing within the virtual machine computing environment as an address to physical memory; wherein the host physical memory address is an address to actual physical memory of a host computing device hosting the virtual machine computing environment; wherein the host virtual memory address is an address to virtual memory provided by a memory manager executing on the host computing device, the memory manager utilizing the hierarchically arranged page tables to provide the virtual memory; and wherein the guest physical memory identified by the guest physical memory address is backed by a portion of the host virtual memory identified by a portion of the host virtual memory address.
[0120] From the above description it can be seen that a mechanism has been described that can accelerate memory access through SLAT.In view of the many possible variations of the subject matter described herein, we claim as the invention all such embodiments that come within the scope of the appended claims and their equivalents.
Claims
1. A method for improving the access speed of a computer memory, the method comprising: detecting, from a first process executing in a virtual machine computing environment, a first memory access directed to a first memory range; As a precondition for completing the first memory access, generating a first entry in a hierarchically arranged layer 2 address translation table that associates a host physical memory address with a guest physical memory address, the first entry being at least one level above a lowest level of the table such that a second memory range identified by the first entry can be identified without referencing any table in the lowest level of the table; as well as responsive to generating the first entry in the hierarchically arranged layer 2 address translation table, marking a first plurality of entries in a hierarchically arranged page table as used, the hierarchically arranged page table associating the host physical memory address with a host virtual memory address, the first plurality of entries as a whole referencing the same second memory range as the first entry in the hierarchically arranged layer 2 address translation table, wherein entries in the first plurality of entries are at a lowest level of the table; wherein the guest physical memory address is perceived by a process executing within the virtual machine computing environment as an address to a physical memory; wherein the host physical memory address is an address to an actual physical memory of a host computing device hosting the virtual machine computing environment; wherein the host virtual memory address is an address to a virtual memory provided by a memory manager executing on the host computing device, the memory manager utilizing the hierarchically arranged page tables to provide the virtual memory; as well as Wherein the guest physical memory identified by the guest physical memory address is backed by a portion of the host virtual memory identified by a portion of the host virtual memory address.
2. The method according to claim 1, further comprising: detecting a second memory access from the first process executing in the virtual machine computing environment, the second memory access being directed to a third memory range different from the first memory range; and satisfying the second memory access by referencing a second subset of the first plurality of entries in the hierarchically arranged page table; Wherein the first memory access is satisfied by referencing a first subset of the first plurality of entries in the hierarchically arranged page table, the first subset comprising entries in the hierarchically arranged page table that are different from the second subset.
3. The method according to claim 1, further comprising: After the first memory access is completed, detecting that a first subset of the first plurality of entries in the hierarchically arranged page table already has data, the data initially stored in corresponding host physical memory addresses and subsequently swapped out to a non-volatile storage medium; In response to the detecting, invalidating the first entry in the hierarchically arranged layer 2 address translation table; as well as Generate a second plurality of entries in the hierarchically arranged layer-2 address translation table to replace the first entry in the hierarchically arranged layer-2 address translation table, wherein the second plurality of entries as a whole reference at least some of the same second memory range as the first entry, wherein entries in the second plurality of entries are at the lowest level of the table. 4 . The method of claim 3 , wherein the second plurality of entries references portions of the second memory range that were previously accessed by the first process and that have not been swapped out.
5. The method according to claim 1, further comprising: A first plurality of contiguous small page size regions of the host physical memory are assembled into a single large page size region of the host physical memory.
6. The method of claim 5, wherein said assembling occurs after said detecting said first memory access and before said generating said first entry in said hierarchically arranged layer-2 address translation table.
7. The method according to claim 5, further comprising: generating a second entry in the hierarchically arranged layer 2 address translation table, the second entry being at least one level above the lowest level of the table so as to identify a third memory range identified by the second entry without referencing any table in the lowest level of the table, the third memory range referencing the single large page size region of the host physical memory into which the first plurality of contiguous small page size regions of the host physical memory are assembled; as well as In response to generating the second entry in the hierarchically arranged second-layer address translation table, a second plurality of entries in the hierarchically arranged page table are marked as used, the second plurality of entries referencing the first plurality of contiguous small page size regions of the host physical memory assembled into the single large page size region of the host physical memory.
8. The method according to claim 7, further comprising: copying data from a second plurality of small page size regions of the host physical memory to at least a portion of the first plurality of contiguous small page size regions, the second plurality of small page size regions being at least partially non-contiguous; as well as invalidating a second plurality of entries in the hierarchically arranged layer-2 address translation table that reference the second plurality of small page size regions of the host physical memory, wherein entries in the second plurality of entries are at a lowest level of the table; The generated second entry in the hierarchically arranged layer-2 address translation table is utilized to replace the invalidated second plurality of entries.
9. The method according to claim 7, further comprising: copying data from a second set of one or more small page size regions of the host physical memory to at least a portion of the first plurality of contiguous small page size regions; as well as invalidating a second plurality of entries in the hierarchically arranged layer 2 address translation table, the second plurality of entries comprising both: (1) a first subset of entries referencing the second set of one or more small page size regions of the host physical memory, and (2) a second subset of entries referencing at least some of the first plurality of contiguous small page size regions of the host physical memory, wherein entries in the second plurality of entries are at a lowest level of the table; The generated second entry in the hierarchically arranged layer-2 address translation table is utilized to replace the invalidated second plurality of entries.
10. The method of claim 5, wherein said assembling said first plurality of contiguous small page size regions of said host physical memory comprises: Data is copied from some of the first plurality of continuous small page size areas to other small page size areas of the host physical memory that are different from the first plurality of continuous small page size areas.
11. The method of claim 10, wherein the copying of the data from the some of the first plurality of consecutive small page size regions to the other small page size regions is performed only when fragmentation of the first plurality of consecutive small page size regions is below a fragmentation threshold.
12. The method according to claim 1, further comprising: Paging of the second memory range is prevented.
13. The method of claim 12, wherein said preventing said paging of said second memory range is performed only if an access frequency of one or more portions of said second memory range is greater than an access frequency threshold.
14. The method according to claim 12, further comprising: If the access frequency of one or more portions of the second memory range is less than an access frequency threshold, the blocking of paging of the second memory range is removed.
15. The method of claim 1, wherein the size of the second memory range is 2MB.
16. A computing device comprising: one or more central processing units; Random Access Memory (RAM); as well as One or more computer-readable media comprising: a first set of computer executable instructions that, when executed by the computing device, cause the computing device to provide a memory manager that references a hierarchically arranged page table to translate a host virtual memory address into a host physical memory address identifying a location on the RAM; a second set of computer executable instructions that, when executed by the computing device, cause the computing device to provide a virtual machine computing environment, wherein processes executing within the virtual machine computing environment perceive guest physical memory addresses as addresses to physical memory, and wherein additional guest physical memory identified by the guest physical memory address is backed by a portion of host virtual memory identified by a portion of the host virtual memory address; and A third set of computer executable instructions, when executed by the computing device, causes the computing device to: detecting, from a first process executing in a virtual machine computing environment, a first memory access directed to a first memory range; As a precondition for completing the first memory access, generating a first entry in a hierarchically arranged layer 2 address translation table that associates a host physical memory address with a guest physical memory address, the first entry being at least one level above a lowest level of the table such that a second memory range identified by the first entry can be identified without referencing any table in the lowest level of the table, the second memory range being larger than the first memory range; and In response to generating the first entry in the hierarchically arranged second-layer address translation table, a first plurality of entries in the hierarchically arranged page table are marked as used, the first plurality of entries as a whole referencing the same second memory range as the first entry in the hierarchically arranged second-layer address translation table, wherein entries in the first plurality of entries are at the lowest level of the table.
17. The computing device of claim 16, wherein the third set of computer-executable instructions further comprises additional computer-executable instructions that, when executed by the computing device, cause the computing device to: detecting, after completion of the first memory access, that a first subset of the first plurality of entries in the hierarchically arranged page table already had data that was originally stored in corresponding host physical memory addresses and subsequently swapped out to a non-volatile storage medium; In response to the detecting, invalidating the first entry in the hierarchically arranged layer 2 address translation table; as well as Replacing the first entry in the hierarchically arranged layer-2 address translation table, generating a second plurality of entries in the hierarchically arranged layer-2 address translation table, wherein the second plurality of entries as a whole reference at least some of the same second memory range as the first entry, wherein entries in the second plurality of entries are at the lowest level of the table.
18. The computing device of claim 16, wherein the third set of computer-executable instructions includes additional computer-executable instructions that, when executed by the computing device, cause the computing device to: assembling a first plurality of contiguous small page size regions of the host physical memory into a single large page size region of the host physical memory; wherein said assembling occurs after said detecting said first memory access and before said generating said first entry in said hierarchically arranged layer-2 address translation table.
19. The computing device of claim 16, wherein the third set of computer-executable instructions includes additional computer-executable instructions that, when executed by the computing device, cause the computing device to: Paging of the second memory range is prevented.
20. One or more computer-readable storage media comprising computer-executable instructions that, when executed, cause a computing device to: detecting, from a first process executing in a virtual machine computing environment, a first memory access directed to a first memory range; As a precondition for completing the first memory access, generating a first entry in a hierarchically arranged layer 2 address translation table, the hierarchically arranged layer 2 address translation table associating a host physical memory address with a guest physical memory address, the first entry being at least one level above a lowest level of the table such that a second memory range identified by the first entry can be identified without referencing any table in the lowest level of the table, the second memory range being larger than the first memory range; as well as responsive to generating the first entry in the hierarchically arranged layer 2 address translation table, marking a first plurality of entries in a hierarchically arranged page table as used, the hierarchically arranged page table associating the host physical memory address with a host virtual memory address, the first plurality of entries as a whole referencing the same second memory range as the first entry in the hierarchically arranged layer 2 address translation table, wherein entries in the first plurality of entries are at a lowest level of the table; wherein the guest physical memory address is perceived by a process executing within the virtual machine computing environment as an address to a physical memory; wherein the host physical memory address is an address to an actual physical memory of a host computing device hosting the virtual machine computing environment; wherein the host virtual memory address is an address to a virtual memory provided by a memory manager executing on the host computing device, the memory manager utilizing the hierarchically arranged page tables to provide the virtual memory; and Wherein the guest physical memory identified by the guest physical memory address is backed by a portion of the host virtual memory identified by a portion of the host virtual memory address.
Citation Information
Patent Citations
Memory mirroring and redundancy generation for high availability
CN103597451A
Virtual machines backed by host virtual memory
CN107466397A