Fast access to virtual machine memory supported by virtual memory of host computer
By skipping the SLAT level on the host computing device and maintaining continuous memory correlation using page tables, the problem of low memory access efficiency in the virtual machine environment is solved, and more efficient memory access and memory utilization of virtual machine processes is achieved.
Patent Information
- Application Number
- CN202510501453.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2019-05-27
- Filing Date
- 2019-11-14
- Publication Date
- 2025-07-25
AI Technical Summary
In virtual machine environments, memory access efficiency is low in the prior art, especially due to the increase in latency caused by multi-level traversal of layer two address tables (SLATs), which affects the efficiency and performance of memory access.
By skipping or not referring to multiple levels within the SLAT on the host computing device, maintaining continuous memory correlation at the lowest level with the page table of the host computing device, enabling more efficient SLAT traversal and memory access.
It improves the efficiency of memory access, reduces the consumption of processor cycles, improves the speed and efficiency of memory access, and maintains the memory utilization efficiency of virtual machine processes.
Smart Images

Figure CN120371735A_ABST
Abstract
Description
[0001] Related Applications
[0002] This application is a divisional application of a patent application for invention, with an international filing date of November 14, 2019, entering the Chinese national phase on May 20, 2021, with a Chinese national application number of 201980076709.0 and an invention title of "Fast Access to Virtual Machine Memory Supported by Virtual Memory of a Host Computer". Background Art
[0003] Modern computing devices operate by executing computer-executable instructions in high-speed volatile memory in the form of random access memory (RAM), and the execution of such computer-executable instructions typically requires reading data from the RAM. Due to cost, physical size limitations, power requirements, and other similar constraints, computing devices typically include less RAM than is required by the processes typically executed on such computing devices. To accommodate such constraints, virtual memory is utilized, whereby the memory available to the processes that can be executed on the computing device appears to be greater than the memory provided by the physical memory circuitry. The relationship between virtual memory and physical memory is typically managed by one or more memory managers that implement, maintain, and / or reference a "page table", the information of which describes the relationship between one or more virtual memory addresses and the location of the corresponding data, whether the corresponding data is in the physical memory or on some form of storage medium. To accommodate the memory quantities associated with modern computing devices and the processes executed thereon, modern page tables typically consist of multiple levels of tables, where the higher-level tables have entries that respectively identify different lower-level tables, and the entries contained in the lowest-level table do not identify additional tables but rather identify the memory addresses themselves.
[0004] Among the processes that can be executed by a computing device are those that virtualize or abstract the underlying hardware of the computing device. Such processes include virtual machines, which can simulate a complete underlying computing device as a process executed within a virtualized computing context provided by such a virtual machine. A hypervisor or similar set of computer-executable instructions can facilitate the provision of virtual machines by virtualizing or abstracting the underlying hardware of the physical computing device that hosts such a hypervisor. The hypervisor can maintain a second-level address table (SLAT), which can also be hierarchically arranged in a manner similar to the above-described page table. The SLAT can maintain information that describes the relationship between one or more memory addresses that appear to be the physical memory locations to the processes executed on top of the hypervisor, including the processes executed within a virtual machine context, where the virtualization of the underlying computing hardware by the virtual machine is facilitated by the hypervisor, and the memory locations of the actual physical memory itself.
[0005] For example, when a process executing in the context of a virtual machine accesses memory, two different lookups can be performed. One lookup can be performed within the virtual machine context itself to correlate the requested virtual memory address with a physical memory address. Since this lookup is performed within the virtual machine context itself, the identified physical memory address is only the physical memory address as perceived by the process executing within the virtual machine context. Such a lookup can be performed by a memory manager executing within the virtual machine context and can be generated by referencing a page table that exists within the virtual machine context. A second lookup can then be performed outside of the virtual machine context. More specifically, the physical memory address identified by the first lookup (within the context of the virtual machine) can be correlated with an actual physical memory address. Such a second lookup may require one or more processing units of the computing device to reference a SLAT, which can correlate the perceived physical memory address with the actual physical memory address.
[0006] As noted, both the page table and the SLAT can be hierarchical arrangements of different levels of tables. Thus, whether the table lookup is performed by the memory manager referencing the page table or by the hypervisor referencing the SLAT, it may be necessary to determine an appropriate table entry within the highest level table, reference the lower level table identified by that table entry, determine an appropriate table entry within that lower level table, reference the lower level table identified by that table entry, and so on until the lowest level table is reached, whereupon the entries of the lowest level table identify one or more specific addresses or address ranges of the memory itself, as should identify another table. Each reference to a lower level table consumes processor cycles and increases the duration of the memory access.
[0007] For example, in the case of a memory access from a process executing in the context of a virtual machine, the duration of such a memory access can include traversing the levels of the page table performed by the memory manager within the virtual machine context and traversing the levels of the SLAT performed by the hypervisor. The additional latency introduced by the lookup performed by the hypervisor (referencing the SLAT) makes the memory access from a process executing within the virtual machine context or indeed any process accessing memory through the hypervisor less efficient compared to a process accessing memory more directly. Such inefficiencies may prevent a user from obtaining security benefits and other benefits that come with accessing memory through the hypervisor. SUMMARY OF THE INVENTION
[0008] To improve the memory utilization efficiency of virtual machine processes executed on a host computing device, these virtual machine processes can be supported by the host's virtual memory, enabling traditional virtual memory efficiency to be applied to the memory consumed by such virtual machine processes. In the case where the guest physical memory in a virtual machine environment is supported by virtual memory allocated to one or more processes executed on a host computing device, to increase the speed of traversing the levels of a second-level address table (SLAT) as part of a memory access, one or more levels of the tables within the SLAT can be skipped or otherwise not referenced, resulting in a more efficient SLAT traversal and more efficient memory access. Although the SLAT can be populated with memory affinity at higher levels of the table, the page tables of the host computing device that supports the host computing device in providing virtual memory can maintain a corresponding set of contiguous memory affinity at the lowest level of the table, enabling the host computing device to page out or otherwise manipulate smaller memory blocks. If such manipulation occurs, the SLAT can be repopulated with memory affinity at the lowest level of the table. Conversely, if the host can reorganize a sufficiently large set of contiguous small pages, the SLAT can again be populated with affinity at higher levels of the table, again enabling a more efficient SLAT traversal and more efficient memory access. In this way, a more efficient SLAT traversal can be achieved while maintaining the memory utilization efficiency advantage of using the host computing device's virtual memory to support virtual machine processes.
[0009] This "Summary of the Invention" is provided to introduce some concepts in a simplified form that will be further described in the "Detailed Description" below. This "Summary of the Invention" is not intended to identify key features or essential features of the claimed subject matter, nor is it intended to be used to limit the scope of the claimed subject matter.
[0010] Other features and advantages will become apparent from the following detailed description when taken in conjunction with the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0011] The following detailed description can be better understood when taken in conjunction with the accompanying drawings, in which:
[0012] Figure 1 is a block diagram of an exemplary computing device
[0013] Figure 2 is a system diagram of an exemplary hardware and software computing system;
[0014] Figure 3 is a block diagram illustrating terms used herein;
[0015] Figure 4 is a flowchart illustrating an exemplary method for supporting guest physical memory with host virtual memory;
[0016] Figure 5 is a flowchart showing various exemplary actions in the life cycle of data in guest physical memory supported by host virtual memory;
[0017] Figure 6 is a block diagram showing an exemplary mechanism for supporting guest physical memory with host virtual memory;
[0018] Figure 7 is a block diagram showing an exemplary mechanism for more efficiently translating between guest physical memory and host physical memory while supporting guest physical memory with host virtual memory; and
[0019] Figure 8 is a flowchart showing an exemplary method for more efficiently translating between guest physical memory and host physical memory while supporting guest physical memory with host virtual memory. DETAILED DESCRIPTION
[0020] The following description relates to improving the efficiency of memory access while providing more efficient memory utilization on a computing device hosting one or more virtual machine processes providing a virtual machine computing environment. To improve the memory utilization efficiency of virtual machine processes executing on a host computing device, these virtual machine processes can be supported by the host's virtual memory, such that traditional virtual memory efficiency can be applied to the memory consumed by such virtual machine processes. In the case where the guest physical memory of a virtual machine environment is supported by virtual memory allocated to one or more processes executing on a host computing device, to improve the speed of traversing the levels of a second-level address table (SLAT) as part of a memory access, one or more levels of the tables within the SLAT can be skipped or otherwise not referenced, resulting in a more efficient SLAT traversal and more efficient memory access. While the SLAT can be populated with memory affinity at higher levels of the table, the page tables of the host computing device that support the host computing device in providing virtual memory can maintain a corresponding set of contiguous memory affinity at the lowest level of the table, such that the host computing device can swap out or otherwise manipulate smaller memory blocks. If such manipulation occurs, the SLAT can be repopulated with memory affinity at the lowest level of the table. Conversely, if the host can reorganize a large enough set of contiguous small pages, the SLAT can again be populated with affinity at higher levels of the table, again resulting in a more efficient SLAT traversal and more efficient memory access. In this way, a more efficient SLAT traversal can be achieved while maintaining the memory utilization efficiency benefits of supporting virtual machine processes with the virtual memory of the host computing device.
[0021] Although not required, the following description will be in the general context of computer-executable instructions, such as program modules, executed by a computing device. More specifically, unless otherwise indicated, the description will refer to the acts and symbolic representations of operations performed by one or more computing devices or peripheral devices. Thus, it will be understood that such acts and operations, sometimes referred to as computer-executed, include the manipulation of electrical signals by a processing unit on data represented in a structured form. This manipulation transforms the data or maintains its position in memory, which reconfigures or otherwise changes the operation of the computing device or peripheral device in a manner well known to those skilled in the art. The data structures in which data is maintained are physical locations that have specific properties defined by the data format.
[0022] Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. In addition, those skilled in the art will understand that a computing device need not be limited to a conventional personal computer, but includes other computing configurations, including servers, handheld devices, multiprocessor systems, microprocessor-based or programmable consumer electronics, network PCs, minicomputers, mainframe computers, etc. Similarly, a computing device need not be limited to a stand-alone computing device, as these mechanisms can also be practiced in a distributed computing environment where tasks are performed by remote processing devices linked through a communication network. In a distributed computing environment, program modules can be located in local and remote storage devices.
[0023] Before continuing with the detailed description of the memory allocation and access mechanisms mentioned above, reference Figure 1 The exemplary computing device 100 shown provides a detailed description of an exemplary host computing device that provides context for the following description. The exemplary computing device 100 may include, but is not limited to, one or more central processing units (CPUs) 120, a system memory 130, and a system bus 121 that couples various system components, including the system memory, to the processing unit 120. The system bus 121 can be any of several types of bus structures, including a memory bus or memory controller using any of various bus architectures, a peripheral bus, and a local bus. The computing device 100 may optionally include graphics hardware, including but not limited to a graphics hardware interface 160 and a display device 161, which may include a display device capable of receiving touch-based user input, such as a touch-sensitive or multi-touch display device. Depending on the specific physical implementation, one or more of the CPU 120, the system memory 130, and other components of the computing device 100 may be physically located in the same place, such as on a single chip. In such a case, some or all of the system bus 121 may be merely a silicon path within a single-chip structure, and itsFigure 1 The illustration in
[0024] The computing device 100 generally also includes a computer-readable medium, which can include any available medium accessible by the computing device 100, and includes volatile and non-volatile media, as well as removable and non-removable media. By way of example and not limitation, the computer-readable medium can include computer storage media and communication media. Computer storage media includes media for storing content such as computer-readable instructions, data structures, program modules, or other data implemented in any method or technology. Computer storage media includes, but is not limited to: RAM, ROM, EEPROM, flash memory or other storage technologies, CD-ROM, digital versatile disk (DVD) or other optical disk storage, magnetic cassettes, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other medium that can be used to store the desired content and can be accessed by the computing device 100. However, computer storage media does not include communication media. Communication media typically embodies computer-readable instructions, data structures, program modules, or other data in a modulated data signal such as a carrier wave or other transmission mechanism, and includes any content delivery medium. By way of example and not limitation, communication media includes wired media such as a wired network or direct wired connection, and wireless media such as acoustic, RF, infrared, and other wireless media. Any combination of the above should also be included within the scope of computer-readable media.
[0025] The system memory 130 includes computer storage media in the form of volatile and / or non-volatile memory, such as read-only memory (ROM) 131 and random access memory (RAM) 132. The basic input / output system 133 (BIOS), which contains basic routines that help transfer content between elements within the computing device 100 during startup, is typically stored in the ROM 131. The RAM 132 typically contains data and / or program modules that are immediately accessible to and / or currently being operated on by the processing unit 120. By way of example and not limitation, Figure 1 an operating system 134, other program modules 135, and program data 136 are shown.
[0026] The computing device 100 may also include other removable / non-removable volatile / non-volatile computer storage media. Merely by way of example, Figure 1Illustrated is a hard disk drive 141 that reads from or writes to a non-removable, non-volatile magnetic medium. Other removable / non-removable volatile / non-volatile computer storage media that may be used with the exemplary computing device include, but are not limited to, magnetic tape cartridges, flash memory cards, digital versatile disks, digital video tapes, solid state RAM, solid state ROM, and other computer storage media as defined and described above. The hard disk drive 141 is typically connected to the system bus 121 through a non-volatile memory interface such as interface 140.
[0027] The drives and their associated computer storage media discussed above and illustrated in Figure 1 provide storage of computer-readable instructions, data structures, program modules, and other data for the computing device 100. In Figure 1 , for example, the hard disk drive illustration 141 is shown storing an operating system 144, other program modules 145, and program data 146. Note that these components may be the same as or different from the operating system 134, other program modules 135, and program data 136. The operating system 144, other program modules 145, and program data 146 are given different numbers to illustrate that they are at least different copies.
[0028] The computing device 100 may operate in a networked environment using logical connections to one or more remote computers. The computing device 100 is shown connected to a general network connection 151 (to network 190) through a network interface or adapter 150, which in turn is connected to the system bus 121. In a networked environment, program modules depicted relative to the computing device 100 or portions thereof or peripherals may be stored in the memory of one or more other computing devices communicatively coupled to the computing device 100 through the general network connection 161. It should be understood that the network connections shown are exemplary and other means by which communication links may be established between computing devices.
[0029] Although described as a single physical device, the exemplary computing device 100 can be a virtual computing device, in which case the functions of the above-described physical components such as the CPU 120, system memory 130, network interface 160, and other similar components can be provided by computer-executable instructions. Such computer-executable instructions can be executed on a single physical computing device or can be distributed across multiple physical computing devices, including dynamically across multiple physical computing devices, such that the particular physical computing device hosting such computer-executable instructions can change dynamically over time according to need and availability. In the case where the exemplary computing device 100 is a virtualized device, the underlying physical computing device hosting such virtualized computing device can itself include physical components similar to and operating in a similar manner as the above-described components. Additionally, virtual computing devices can be utilized in multiple layers, where one virtual computing device executes within the construction of another virtual computing device. Thus, as used herein, the term "computing device" refers to a physical computing device or a virtualized computing environment including virtual computing devices, in which computer-executable instructions can be executed in a manner consistent with the way a physical computing device executes them. Similarly, as used herein, the term referring to the physical components of a computing device refers to the physical components or the virtualization that performs the same or equivalent functions of the physical components.
[0030] Go to Figure 2 , the system 200 shown therein illustrates a computing system including computing device hardware and processes executing thereon. More specifically, Figure 2 the exemplary system 200 shown includes the computing device hardware itself, such as the processing unit 120 and RAM 132 described above, which are shown as part of the exemplary computing device 100. Additionally, the exemplary system 200 includes a hypervisor, such as the exemplary hypervisor 210, which can include computer-executable instructions executed by the exemplary computing device 100 to enable the computing device 100 to perform functions associated with the exemplary hypervisor 210. The exemplary hypervisor 210 can include or can reference a second-level address table (SLAT), such as the exemplary SLAT 220.
[0031] The exemplary hypervisor 210 can virtualize a set of computing hardware, which can be equivalent to or different from the hardware of the computing device 100, including differences in processor type and / or capacity, the amount and / or type of RAM, the amount and / or type of storage media, and other similar differences. Among other things, such virtualization can enable one or more virtual machine processes, such as the exemplary virtual machine processes 260 and 270. The hypervisor 210 can present the appearance of executing directly on the computing device hardware 110 to both the processes executing within these virtual machine contexts and any other processes executing above the hypervisor. Figure 2 The illustrated exemplary system 200 illustrates an operating system, such as the exemplary operating system 240 executing above the exemplary hypervisor 210. Further, as shown in the exemplary system 200, one or more computer-executable applications, such as the virtual machine processes 260 and 270 described above, can execute on the exemplary operating system 240.
[0032] The exemplary operating system 240 can include various components, sub-components, or aspects thereof related to the following description. Although shown as visually distinct elements in Figure 2 this visual description does not imply a particular computational independence, such as process boundaries, independent memory silos, or other similar computational delineations. One aspect of the operating system 240 can be a memory manager, such as the exemplary memory manager 241. More specifically, the exemplary memory manager 241 can read code or data from one or more computer-readable storage media and copy such code or data into memory that can be supported by the RAM 132.
[0033] As used herein, the term "virtual memory" is formed not with reference to virtual machines, but rather with reference to the concept of presenting to an application or other process executing on a computing device the appearance of being able to access more memory than physically exists in the RAM 132. Thus, for example, the virtual memory functionality provided by the operating system 240 may allow code and / or data to be loaded into memory ultimately supported by physical memory such as the physical memory 132, but with a virtual memory quantity greater than the physical memory. A memory manager 241 that executes as part of the operating system 240 may maintain, modify, and / or utilize a page table 250 to resolve the association between one or more virtual memory addresses and one or more physical memory addresses. More specifically, the page table 250 provides an address translation between one memory addressing scheme and another different memory addressing scheme. To achieve this, the page table 250 may include multiple levels of tables, as will be described in further detail below, which may associate one or more memory addresses (i.e., virtual memory addresses) conforming to one memory addressing scheme with one or more memory addresses (i.e., physical memory addresses) conforming to another different memory addressing scheme.
[0034] Figure 2 The exemplary system 200 shown illustrates an application executing on top of the operating system 240. In particular, such an application may be a virtual machine application, which may instantiate virtual machine processes such as the exemplary virtual machine processes 260 and 270. Each virtual machine process may create a virtual computing environment that may host its own operating system, applications, and other similar functionality that may be hosted by the physical computing environment presented by, for example, the computing device 100. Within Figure 2 the exemplary system 200 shown, the exemplary virtual machine process 260 is shown as creating a virtual machine environment 261, and the exemplary virtual machine process 270 is shown as creating a virtual machine environment 271. The exemplary virtual machine environment 271 is shown as hosting an operating system in the form of the exemplary operating system 281, which may be similar to the exemplary operating system 240 or may be a different type of operating system. For illustrative purposes, the exemplary operating system 281 executing within the exemplary virtual machine environment 271 is shown as including a memory manager, namely, the exemplary memory manager 282, which in turn references a page table (i.e., the exemplary page table 283) to provide virtual memory within the virtual machine environment 271. Such virtual memory within the virtual machine environment 271 may be utilized by processes such as the exemplary processes 291 and 292 that execute on top of the exemplary operating system 281 and that are also shown as executing within the exemplary virtual machine environment 271 in Figure 2 which is also shown.
[0035] In a manner similar to that implemented by memory managers 241 and 282, when facilitating the presentation of a set of virtual computing hardware including memory, hypervisor 210 can similarly translate memory addresses from one memory addressing scheme to another different memory addressing scheme. More specifically, hypervisor 210 can translate a memory address provided by hypervisor 210, which is a physical memory address, into an actual physical memory address, such as a memory address identifying a physical location within memory 132. Such translation can be performed by referring to another hierarchically organized table (i.e., SLAT 220), which can also provide a memory address translation between one memory addressing scheme (in this case, the address is treated as a physical memory address by a process executing above hypervisor 210) and another different memory addressing scheme (in this case, the actual physical memory address).
[0036] According to one aspect, the virtual memory provided by operating system 240 can be used to support the physical memory of a virtual machine instead of using non-paged physical memory allocation on the host computing device. This allows memory manager 241 to manage the host physical memory associated with guest physical memory. In particular, the memory management logic already present on the host (such as memory manager 241) can now be utilized to manage the physical memory of guest virtual machines. In terms of the amount of code used to implement a hypervisor (such as exemplary hypervisor 210), this can allow for the use of a smaller hypervisor. A smaller hypervisor can be more secure because there is less code that can be exploited or that may have bugs. Additionally, this can increase the virtual machine density on the host because fewer host physical memory resources can be used to implement virtual machines. For example, previously, because virtual machine processes required non-paged memory, if exemplary computing device 100 included 3GB of RAM 132, it could only host one virtual machine process that created a virtual machine with 2GB of RAM. In contrast, if the memory consumed by a virtual machine process such as exemplary virtual machine process 270 is virtual memory provided by a memory manager 241 such as operating system 240, then 4GB of virtual memory space can provide support for two virtual machine processes, each creating a virtual machine with 2GB of RAM. In this way, exemplary computing device 100 can host two virtual machine processes instead of just one, thereby increasing the density of virtual machine processes executing on exemplary computing device 100.
[0037] Other benefits include a single memory management codebase for all memory on the management system (host and virtual machines). Thus, improvements, fixes, and / or tuning in one codebase benefit everyone. Additionally, since only one codebase needs to be maintained, this can result in a reduction in engineering costs. Another example advantage can be that virtual machines can immediately benefit from density improvements such as paging, page sharing, working set aging and trimming, fault clustering, etc. Another example advantage can be that virtual memory and physical memory consumption limits can be set on virtual machines, just like other processes, enabling the administrator to control system behavior. Another example advantage can be that additional features can be added to the host memory manager to provide more performance, density, and functionality to virtual machines (other non-virtual machine workloads may also benefit from this).
[0038] Go to Figure 3 , system 300 illustrates the terms used herein to refer to the multi-level page table and corresponding memory addressing scheme applicable to Figure 2 the system 200 shown. More specifically, within the virtual machine environment 271, the exemplary memory manager 282 may reference the page table 283 to again associate virtual memory addresses with physical memory addresses in the context of the virtual machine environment 271. As described above, the term "virtual memory" does not refer to the memory in the virtual machine environment, but rather to the amount of memory available to a process executing on a computing device, which is supported by the physical memory capacity but is greater than the physical memory capacity. Thus, to nominally distinguish the virtual memory in the virtual machine environment from the virtual memory in a non-virtualized or "bare metal" computing environment, the adjectives "guest" and "host" will be used, with the adjective "guest" representing the virtual machine computing environment and the adjective "host" representing the bare metal computing environment on which the virtual machine process that provides the virtual machine environment is executed.
[0039] In this way, and as shown in system 301, the addresses in the "guest virtual memory" 311 addressing scheme can be associated with the addresses in the "guest physical memory" 312 addressing scheme through the page table 283 in the virtual machine environment 271 as shown. Figure 2 Similarly, the addresses within the "host virtual memory" 321 addressing scheme can be associated with the addresses within the "host physical memory" 322 addressing scheme through the page table 250 that is part of the exemplary operating system 240 executing in a non-virtual computing environment as shown. Figure 2 Although not shown in Figure 3As shown by exemplary system 301, but according to one aspect and as will be described in further detail below, the guest physical memory 312 of the virtual machine environment can be supported by the host virtual memory 321, thereby providing the above-mentioned benefits. This support of the host virtual memory 321 for the guest physical memory 312 may require modification of the SLAT 220, which provides a translation between the "guest physical memory" 312 addressing scheme and the "host physical memory" 322 addressing scheme, as shown by exemplary system 301.
[0040] Go to Figure 3 The system 302 shown illustrates a hierarchical arrangement of tables, such as may be used to implement exemplary page table 250, exemplary page table 283, and / or exemplary SLAT 220. To accommodate the amount of memory applicable to modern computing devices and application processes, the page tables 250, 283, and / or SLAT 220 will have multiple (usually four) levels of tables. However, for purposes of illustration, system 302 shows only two levels of tables. The higher-level table 340 may include a plurality of table entries, such as exemplary table entries 341, 342, 343, and 344. For example, the higher-level table 340 may be 4KB in size and may contain 512 discrete table entries, each discrete table entry being eight bytes in size. Each table entry may include a pointer or other similar identifier to a lower-level table, such as one of exemplary lower-level tables 350 and 360. Thus, for example, exemplary table entry 341 may include the identification of exemplary lower-level table 350, as shown by arrow 371. If the higher-level table 340 is 4KB in size and each table entry is eight bytes in size, then at least some of the eight bytes of data may be a pointer or other similar identifier to the lower-level table 350. Exemplary table entry 342 may, in a similar manner, include the identification of lower-level table 360, as shown by arrow 382. Similarly, the remaining discrete table entries of exemplary higher-level table 340 may identify unique lower-level tables. Thus, in an example where the higher-level table 340 is 4KB in size and includes 512 table entries, each table entry being eight bytes in size, such a higher-level table 340 may include 512 identifications of 512 unique lower-level tables, such as exemplary lower-level tables 350 and 360.
[0041] Each lower-level table may include individual table entries in a similar manner. For example, the exemplary lower-level table 350 may include exemplary table entries 251, 352, 353, and 354. Similarly, the exemplary lower-level table 360 may include table entries 361, 362, 363, and 364. As an example, each lower-level table may have the same size and structure as the higher-level table 340. Thus, the exemplary lower-level table 350 may be 4KB in size and may include 512 table entries, such as exemplary table entries 351, 352, 353, and 354, and each table entry may be eight bytes in size. In the example shown in system 302, the table entries of the lower-level table (e.g., exemplary table entries 351, 352, 353, and 354 of the exemplary lower-level table 350) may each identify a contiguous memory address range. For example, the exemplary table entry 351 may identify the memory address range 371, as shown by arrow 391. The memory address range 371 may include a memory "page" and may be the smallest individual manageable amount of memory. Thus, actions may be performed on a page-by-page basis, such as temporarily moving information stored in volatile memory to non-volatile storage media, or vice versa, as part of implementing a virtual memory amount greater than the installed physical memory amount. Additionally, access permissions may be established on a page-by-page basis. For example, since the memory range 371 is uniquely identified by the table entry 351, the access permission may apply to all memory addresses within the memory range 371. Conversely, the access permission for the memory addresses within the memory range 371 may be independent of the access permissions established for the memory range 372, which may be uniquely identified by a different table entry (i.e., table entry 352), as shown by arrow 392.
[0042] The amount of memory in a single memory page (such as the memory range 371) may depend on various factors, including processor design and other hardware factors, as well as communication connections, etc. In one example, the page size of the memory, such as represented by the memory range 371, may be 4KB.
[0043] Upon receiving a request to access a memory location identified by one or more memory addresses, a table entry in the higher-level table may be identified that corresponds to the memory address range that includes the memory address to be accessed. For example, if the memory address to be accessed includes the memory represented by memory 371, then the table entry 341 may be identified in the higher-level table 340. The identification of the table entry 341 may in turn result in the identification of the lower-level table 350. In the lower-level table 350, it may be determined that the memory address to be accessed is part of the memory identified by the table entry 351. Then, the information in the table entry 351 may identify the memory 371 that is being attempted to be accessed.
[0044] Each such traversal of the lower levels of the page table increases the latency between receiving a request to access memory and returning the access to that memory. According to one aspect, a table entry of a higher level table (e.g., exemplary table entry 341 of higher level table 340) may not identify a lower level table, such as exemplary lower level table 350, but rather identify a page of memory that may correspond to the entire memory range that table entries in the lower level table have identified. Thus, for example, if each of the pages of memory 371, 372, and 373 identified by individual table entries 351, 352, and 353 of lower level table 350 is 4KB in size, a 2MB memory range may be identified by the total number of all individual table entries of lower level table 350, since exemplary lower level table 350 may include 512 table entries, each identifying a 4KB memory range. In such an example, if all 2MB of memory are considered a single 2MB memory page, such a single large memory page may be directly identified by exemplary table entry 341 of higher level table 340. In this case, since exemplary table entry 341 will directly identify a single 2MB memory page, such as exemplary 2MB memory page 370, and will not identify a lower level table, no reference to any lower level table will be required. For example, if a page table or SLAT includes four levels, using a large memory page (such as exemplary large memory page 370) will allow one of these levels of the table to be skipped and not referenced, thus providing memory address translation between virtual memory and physical memory by referencing only three levels of the table instead of four levels, thereby providing more efficient memory access, i.e., a 25% increase in efficiency. Since a page is the smallest amount of individually manageable memory, if 2MB sized memory pages are used, then for example, the memory access permissions may be the same for all 2MB of memory within such a large memory page.
[0045] As another example, if there is a hierarchically arranged table above the level of exemplary table 340, and each entry in exemplary table 340 (such as entries 341, 342, 343, and 344) can address a 2MB memory page, such as exemplary 2MB memory page 370, then 512 such table entries can address 1GB of memory as a whole. If a single 1GB memory page is used, the reference to table 340 and subsequent references to lower-level tables (such as exemplary table 350) can be avoided. Returning to the example where the page table or SLAT includes four levels, using such a huge memory page will enable two of these levels of the table to be skipped and not referenced, thereby providing memory address translation between virtual memory and physical memory by referring only to the levels of the table rather than all four levels, and thus providing memory access in approximately half the time. Also, since a page is the smallest amount of individually manageable memory, if a 1GB-sized memory page is used, for example, for all 1GB of memory in such a huge memory page, the memory access permissions can be the same.
[0046] As used herein, the term "large memory page" refers to the following contiguous memory range, the size of which is determined such that it encompasses all the memory ranges that would otherwise be identified by a single table at the lowest level and can be uniquely and directly identified by a single table entry of a table at one level above the lowest-level table. In the specific size example provided above, a "large memory page" as defined herein would be a 2MB-sized memory page. However, as previously mentioned, the term "large memory page" does not refer to a specific size, but rather the amount of memory in a "large memory page" depends on the hierarchical design of the page table itself and the amount of memory referenced by each table entry of the table at the lowest level of the page.
[0047] In a similar manner, as used herein, the term "huge memory page" refers to the following contiguous memory range, the size of which is determined such that it encompasses all the memory ranges that would otherwise be identified by a single table at the second-lowest level and can be uniquely and directly identified by a single table entry of a table at one level above the second-lowest-level table. In the specific size example provided above, as used herein, a "huge memory page" would be a 1GB-sized memory page. Again, as shown, the term "huge memory page" does not refer to a specific size, but rather the amount of memory in a "huge memory page" depends on the hierarchical design of the page table itself and the amount of memory referenced by each table entry of the table at the lowest level of the page.
[0048] Returning to Figure 2 , referring to such as Figure 2Descriptions of virtualization stacks such as the exemplary virtualization stack 230 provide a description of the mechanism by which the host virtual memory (of the host computing device) can be used to support the guest physical memory (of the virtual machine). A user-mode process can be executed on the host computing device to provide virtual memory for supporting guest virtual machines. One such user-mode process can be created for each guest machine. Alternatively, a single user-mode process can be used for multiple virtual machines, or multiple processes can be used for a single virtual machine. Alternatively, virtual memory can be implemented in other ways different from using user-mode processes, as described below. For illustrative purposes, the virtual memory assigned to the virtual machine process 270 can be used to support the guest physical memory presented by the virtual machine process 270 in the virtual machine environment 271, and similarly, the virtual memory assigned to the virtual machine process 260 can be used to support the guest physical memory presented by the virtual machine process 260 in the virtual machine environment 261. Although this is illustrated in the exemplary system 200 of Figure 2 , the above-described user-mode process can be separate from virtual machine processes such as the exemplary virtual machine processes 260 and 270.
[0049] According to one aspect, a virtualization stack such as the exemplary virtualization stack 230 can allocate host virtual memory in the address space of a designated user-mode process that hosts a virtual machine such as the exemplary virtual machine process 260 or 270. The host memory manager 241 can treat this memory as any other virtual allocation, which means that it can be paged, the physical pages supporting it can be changed to satisfy contiguous memory allocations elsewhere on the system; the physical pages can be shared with another virtual allocation in another process (and that virtual allocation can in turn be another virtual machine support allocation or any other allocation on the system). At the same time, many optimizations can be made to enable the host memory manager to specifically handle virtual machines that support virtual machine allocations as needed. Additionally, if the virtualization stack 230 chooses to prioritize performance over density, it can perform many operations supported by the operating system memory manager 241, such as locking pages in memory to ensure that the virtual machine does not experience paging of these portions. Similarly, large pages can be used to provide higher performance for the virtual machine, which will be described in detail below.
[0050] A given virtual machine can place all of its guest physical memory addresses in the guest physical memory supported by the host virtual memory, or it can place some of its guest physical memory addresses in the guest physical memory supported by the host virtual memory and some in the guest physical memory supported by traditional mechanisms, such as non-paged physical memory allocations from the host physical memory.
[0051] When creating a new virtual machine, the virtualization stack 230 can use a user-mode process to host virtual memory allocation to support guest physical memory. This can be a newly created empty process, an existing process hosting multiple virtual machines, or a process for each virtual machine that contains other virtual machine-related virtual allocations (e.g., virtualization stack data structures) that are not visible to the virtual machine itself. The kernel virtual address space can also be used to support the virtual machine. Once such a process is found or created, the virtualization stack 230 can make a private memory virtual allocation (or section / file mapping) in its address space, which corresponds to the amount of guest physical memory that the virtual machine should have. Specifically, the virtual memory can be a private allocation, a file mapping, a page file-supported partial mapping, or any other type of allocation supported by the host memory manager 241. As described below, this allocation can provide access speed advantages. Conversely, this allocation can be any number of allocations.
[0052] Once the virtual memory allocation is complete, it can be registered with the components (such as the exemplary virtual machine environment 271) that will manage the physical address space of the virtual machine and synchronized with the host physical memory pages that the host memory manager 241 will select to support the virtual memory allocation. These components can be the hypervisor 210 and the virtualization stack 230, which can be implemented as part of the host kernel and / or driver. The hypervisor 210 can manage the translation between the guest physical memory address range and the corresponding host physical memory address range by leveraging the SLAT 220. Specifically, the virtualization stack 230 can update the SLAT 220 with the host physical memory pages that support the corresponding guest physical memory pages. When a guest virtual machine (such as the exemplary virtual machine 271) performs a certain type of access to a given guest physical memory address, the hypervisor 210 can enable the virtualization stack 230 to have the ability to receive an intercept. For example, when the guest virtual machine 271 writes to a certain physical address, the virtualization stack 230 can request to receive an intercept.
[0053] According to one aspect, when a virtual machine such as the exemplary virtual machine 271 is first created, the SLAT 220 may not contain any valid entries corresponding to that virtual machine because no host physical memory addresses may have been allocated to support the guest physical memory addresses of such a virtual machine (although, as shown below, in some embodiments, the SLAT 220 may be pre-populated for some guest physical memory addresses at the same time or approximately the same time as the virtual machine 271 is created). The hypervisor 210 may know the range of guest physical memory addresses that the virtual machine 27 will utilize, but there need not be any host physical memory to support them at this time. When the virtual machine process 270 begins execution, it may start accessing its (guest) physical memory pages. When each new physical memory address is accessed, since the corresponding SLAT entry has not been populated with the corresponding host physical memory address, it may generate an intercept of the appropriate type (read / write / execute). The hypervisor 210 may receive the guest access intercept and forward it to the virtualization stack 230. The virtualization stack 230, in turn, may reference data structures maintained by the virtualization stack 230 to find the range of host virtual memory addresses corresponding to the requested range of guest physical memory addresses (and the host process from which the supported virtual address space was allocated, such as the exemplary virtual machine process 270). At this point, the virtualization stack 230 may know the specific host virtual memory address corresponding to the guest physical memory address that generated the intercept.
[0054] Then, the virtualization stack 230 can issue a virtual fault to the host memory manager 241 in the context of a process (such as the virtual machine process 270) that is hosting a virtual address range. The virtual fault can be issued with the corresponding access type (read / write / execute) of the original intercept that occurred when the virtual machine 271 accessed a physical address in the guest physical memory. The virtual fault can execute the same or a similar code path as a regular page fault to make the specified virtual address valid and accessible by the host CPU. One difference is that the virtual fault code path can return a physical page number that the memory manager 241 uses to make the virtual address valid. This physical page number can be the host physical memory address that supports the host virtual address, which in turn supports the guest physical memory address where the access intercept was originally generated in the hypervisor 210. At this point, the virtualization stack 230 can generate or update the table entry in the SLAT 220 corresponding to the original guest physical memory address where the intercept was generated, and update the table entry with the host physical memory address and the access type (read / write / execute) used to make the virtual address valid in the host. Once this is done, the guest physical memory address can be accessed immediately using the access type of the guest virtual machine 271. For example, the parallel virtual processors in the guest virtual machine 271 can access such an address immediately without encountering an intercept. Since the SLAT contains table entries that can associate the requested guest physical address with the corresponding host physical memory address, the original intercept handling can be completed, and the original virtual processor that generated the intercept can retry its instruction and continue accessing memory.
[0055] If and / or when the host memory manager 241 decides to perform any action that may or will change the host physical address support for a host virtual address that is made valid through a virtual fault, it can perform a translation lookaside buffer (TLB) flush for that host virtual address. It may already perform such actions to comply with existing contracts that the host memory manager 241 may have with the hardware CPUs on the host. The virtualization stack 230 can intercept such a TLB flush and invalidate the corresponding SLAT entries for any host virtual address that supports any guest physical memory address in any virtual machine being flushed. The TLB flush call can identify the range of virtual addresses being flushed. Then, the virtualization stack 230 can look up the flushed host virtual address against its data structure, which can be indexed by the host virtual address to find the guest physical ranges that can be supported by a given host virtual address. If any such ranges are found, the SLAT entries corresponding to these guest physical memory addresses can be invalidated. Additionally, the host memory manager can handle virtual allocations that support virtual memory in different ways as needed or when needed to optimize TLB flush behavior (e.g., reduce SLAT invalidation time, subsequent memory intercepts, etc.).
[0056] The virtualization stack 230 can carefully synchronize the update of the SLAT 220 with the host physical memory page number returned from a virtual fault (served by the memory manager 241) in contrast to a TLB flush performed by the host (issued by the host memory manager 241). Doing so can avoid adding complex synchronization between the host memory manager 241 and the virtualization stack 230. The physical page number returned from a virtual fault may be stale by the time it is returned to the virtualization stack 230. For example, the virtual address may have become invalid. By intercepting the TLB flush calls from the host memory manager 241, the virtualization stack 230 can know when this race occurs and retry the virtual fault to obtain an updated physical page number.
[0057] When the virtualization stack 230 invalidates a SLAT entry, any subsequent access by the virtual machine 271 to that guest physical memory address will again generate an intercept to the hypervisor 210, which is then forwarded to the virtualization stack 230 to be resolved in the manner described above. The same process can be repeated when first accessing the guest physical memory address for reading and then writing. The write operation will generate a separate intercept because a SLAT entry can only be made valid by a "read" access type. The intercept can be forwarded to the virtualization stack 230 as usual, and a virtual fault with "write" access rights can be issued to the host memory manager 241 to obtain the appropriate virtual address. The host memory manager 241 can update its internal state (usually in a page table entry (or "PTE")) to indicate that the host physical memory page is now dirty. This can be done before allowing the virtual machine to write to its guest physical memory address, thus avoiding data loss and / or corruption. If and / or when the host memory manager 241 decides to trim the virtual address (which will perform a TLB flush and invalidate the corresponding SLAT entry), the host memory manager 241 can know that the page is dirty and needs to be written to a non-volatile computer-readable storage medium (such as a page file on a hard drive) before being re-used. In this way, the above sequence is no different from what would happen for a regular private virtual allocation for any other process running on the host computing device 100.
[0058] As an optimization, the host memory manager 241 may choose to perform a page combing pass on all of its memory 120. This can be an operation where the host memory manager 241 finds the same pages across all processes and combines them into a read-only copy of the pages shared by all processes. If and / or when any combined virtual address is written, the memory manager 241 may perform a copy-on-write operation to allow the write to occur. Such an optimization can now work transparently across virtual machines (such as the exemplary virtual machines 261 and 271) to increase the density of virtual machines executing on a given computing device by combining the same pages across virtual machines, thereby reducing memory consumption. When page combing occurs, the host memory manager 241 may update the PTEs that map the affected virtual addresses. During this update, it may perform a TLB flush since the host physical memory addresses may change from unique private pages to shared pages for these virtual addresses. As part of this, as described above, the virtualization stack 230 may invalidate the corresponding SLAT entries. If and / or when the virtual address is combined to point to the guest physical memory address of a shared page and is read, the virtual fault resolution may return the physical page number of the shared page during the intercept handling, and the SLAT 220 may be updated to point to the shared page.
[0059] If a virtual machine writes to any combined guest physical memory address, the virtual fault with write access may perform a copy-on-write operation and may return and update the new private host physical memory page number in the SLAT 220. For example, the virtualization stack 230 may instruct the host memory manager 241 to perform the page combing process. In some embodiments, the virtualization stack 230 may specify which portion of the memory to scan for combing or which processes should be scanned. For example, the virtualization stack 230 may identify processes such as the exemplary virtual machine processes 260 and 270 as processes to be scanned for combination, where the host virtual memory allocation for the process supports the guest physical memory of the corresponding virtual machine (i.e., the exemplary virtual machines 261 and 271).
[0060] Even when the SLAT 220 is updated to allow write access due to the execution of a virtual fault, the hypervisor 210 may support intercepting a write operation to such a SLAT entry if requested by the virtualization stack 230. This is useful because the virtualization stack 230 may want to know when a write occurs regardless of the fact that the host memory manager 241 may accept these writes as occurring. For example, live migration of a virtual machine or a virtual machine snapshot may require the virtualization stack 230 to monitor writes. Even if the state of the host memory manager has been updated accordingly for the write, such as when the PTE has been marked as dirty, the virtualization stack 230 may still be notified when the write occurs.
[0061] The host memory manager 241 can be able to maintain an accurate access history for each host virtual page that supports the guest physical memory address space, just as it does for regular host virtual pages allocated in any other process address space. For example, the "access bit" in the PTE can be updated during a virtual fault as part of handling the memory interception. When the host memory manager clears the access bit on any PTE, it can already flush the TLB to avoid memory corruption. As previously mentioned, this TLB flush can invalidate the corresponding SLAT entry, which can in turn generate an access interception if the virtual machine accesses its guest physical memory address again. As part of handling the interception, the virtual fault handling in the host memory manager 241 can set the access bit again, thus maintaining the correct access history for the page. Alternatively, for performance reasons, such as avoiding access interceptions in the hypervisor 210 as much as possible, the host memory manager 241 can directly use the page access information collected from the SLAT entry by the hypervisor 210 (if the underlying hardware supports it). The host memory manager 241 can cooperate with the virtualization stack 230 to convert the access information in the SLAT 220 (organized by guest physical memory address) into host virtual memory addresses that support these guest physical memory addresses to know which addresses have been accessed.
[0062] With an accurate page access history, the host memory manager 241 can run its normal processing process of intelligent aging and trimming algorithms for the working set. This can allow the host memory manager 241 to examine the state of the entire system and, based on need or other reasons, intelligently select the addresses to be trimmed and / or the pages to be swapped out to disk or other similar operations to relieve memory pressure.
[0063] In some embodiments, like any other process on the system, the host memory manager 241 can impose virtual and physical memory limits on virtual machines. This can help system administrators sandbox or otherwise constrain or enable virtual machines, such as exemplary virtual machines 261 and 271. The host system can use the same mechanisms as native processes to achieve this. For example, in some embodiments where higher performance is required, a virtual machine can have all of its guest physical memory directly supported by host physical memory. Alternatively, some portions can be supported by host virtual memory while other portions are supported by host physical memory 120. In yet another example, where lower performance is acceptable, a virtual machine can be primarily supported by host virtual memory, and the host virtual memory can be restricted to less than full support of the host physical memory. For example, a virtual machine such as exemplary virtual machine 271 can present a guest physical memory size of 4GB, and the guest physical memory for this process can be supported by 4GB of host virtual memory. However, the host virtual memory can be restricted to being supported by only 2GB of host physical memory. This can result in paging to disk or other performance handicaps, but can be a way for an administrator to limit based on service levels or impose other controls on virtual machine deployment. Similarly, a certain amount of physical memory can be guaranteed to a virtual machine by the host memory manager (while still enabling the virtual machine to be supported by host virtual memory) to provide a certain level of performance.
[0064] The amount of guest physical memory provided within the virtual machine environment can be changed dynamically. When additional guest physical memory is needed, another host virtual memory address range can be allocated as described above. Once the virtualization stack 230 is ready to handle access interception on the memory, the guest physical memory address range can be added to the virtual machine environment. When a guest physical memory address is removed, the host memory manager 241 can release the portion of the host virtual address range that supported the removed guest physical memory address (and update the virtualization stack 230 data structures accordingly). Alternatively or additionally, various host memory manager APIs can be called on these portions of the host virtual address range to free host physical memory pages without freeing the host virtual address space for them. Alternatively, in some embodiments, nothing may be done at all because the host memory manager 241 can eventually trim these pages from the working set and eventually write them to disk, such as a page file, as they will no longer be accessed in practice.
[0065] According to one aspect, some or all of the guest physical memory addresses may be pre-populated into the SLAT 220 to host the host physical memory address mapping. This can reduce the number of fault handling operations performed during virtual machine initialization. However, as the virtual machine runs, due to various reasons, the entries in the SLAT 220 may become invalid, and the above-mentioned fault handling can be used to re-associate the guest physical memory address with the host physical memory address again. The SLAT entries may be pre-populated before starting the virtual computing environment or at runtime. Additionally, the entire SLAT 220 or only a portion thereof may be pre-populated.
[0066] Another optimization may be to prefetch other parts of the host virtual memory into the host physical memory to support the guest physical memory that was previously swapped out, such that when subsequent memory intercepts arrive, the virtual fault can be satisfied more quickly because it can be satisfied without having to read data from disk.
[0067] Turning to Figure 4 , method 400 is shown, which provides an overview of the mechanisms described in detail above. Method 400 may include actions for supporting guest physical memory with host virtual memory. The method includes: attempting to access guest physical memory using a guest physical memory access from a virtual machine executing on a host computing device (action 402). For example, Figure 2 the virtual machine environment 271 shown may access guest physical memory that appears as actual physical memory to a process executing within the virtual machine environment 271.
[0068] Method 400 may also include determining a guest physical memory address for which the guest physical memory access reference has no valid entry in a data structure that associates guest physical memory addresses with host physical memory addresses (action 404). For example, it may be determined that there is no valid entry in the Figure 2 SLAT 220 shown.
[0069] As a result, method 400 may include: identifying a host virtual memory address corresponding to the guest physical memory address and identifying a host physical memory address corresponding to the host virtual memory address (action 406). For example, Figure 2 the virtualization stack 230 shown may identify a host virtual memory address corresponding to the guest physical memory address, and also the Figure 2 memory manager 241 shown may identify a host physical memory address corresponding to the host virtual memory address.
[0070] Method 400 may further include updating a data structure that correlates a guest physical memory address with a host physical memory address by leveraging the correlation between the guest physical memory address and the identified host physical memory address (act 408). For example, Figure 2 the illustrated virtualization stack 230 may obtain a host physical memory address from a memory manager 241 also as Figure 2 illustrated, and update the SLAT 220 by leveraging the correlation between the guest physical memory address and the identified host physical memory address (as Figure 2 illustrated).
[0071] Method 400 may be practiced by causing an interception. The interception may be forwarded to the virtualization stack on the host. This may cause the virtualization stack to identify a host virtual memory address corresponding to the guest physical memory address and issue a fault to the memory manager to obtain a host physical memory address corresponding to the host virtual memory address. The virtualization stack may then update a data structure that correlates the guest physical memory address with the host physical memory address by leveraging the correlation between the guest physical memory address and the identified host physical memory address.
[0072] Method 400 may further include: determining the type of guest physical memory access, and updating a data structure that correlates the guest physical memory address with the host physical memory address by leveraging the determined type associated with the guest physical memory address and the identified host physical memory address. For example, if the guest physical memory access is a read, the SLAT may be updated to indicate so.
[0073] Method 400 may further include performing an action that may change a host physical memory address that supports a host virtual memory address. As a result, the method may include invalidating an entry in a data structure that correlates the guest physical memory address with the host physical memory address, the data structure that correlates the guest physical memory address with the host physical memory address. This may cause subsequent accesses to the guest physical memory address to generate a fault, which may be used to update a data structure that correlates the guest physical memory address with the host physical memory address by leveraging the correct correlation of the host virtual memory that supports the guest physical memory address. For example, the action may include a page consolidation operation. Page consolidation may be used to increase the density of virtual machines on the host.
[0074] Method 400 may include initializing a guest virtual machine. As part of initializing the guest virtual machine, method 400 may include pre-populating at least a portion of a data structure that associates guest physical memory addresses with host physical memory addresses using some or all of the guest physical memory addresses of the guest virtual machine to host physical memory address mappings. Thus, for example, host physical memory may be pre-allocated for the virtual machine, and the appropriate associations are entered into the SLAT. This will result in fewer exceptions being required to initialize the guest virtual machine.
[0075] Now referring to Figure 5 , an example flow 500 is shown that illustrates various actions that may occur during a portion of the life cycle of certain data at an example guest physical memory address (hereinafter simply referred to as "GPA") 0x1000. As shown at 502, the SLAT entry for GPA 0x1000 is invalid, indicating that there is no SLAT entry for that particular GPA. As shown at 504, the virtual machine attempts to perform a read at GPA 0x1000, resulting in a virtual machine (hereinafter simply referred to as "VM") read intercept. As shown at 506, the hypervisor forwards the intercept to the virtualization stack on the host. At 508, the virtualization stack performs a virtualization lookup for the host virtual memory address (VA) corresponding to GPA 0x1000. The lookup result is VA 0x8501000. At 510, a virtual fault is generated for the read access to VA 0x8501000. The system physical address (hereinafter simply referred to as "SPA") returned by the virtual fault handling of the memory manager is 0x88000, which defines the address in system memory where the data at GPA 0x1000 physically resides. Thus, as shown at 514, the SLAT is updated to associate GPA 0x1000 with SPA 0x88000 and the data access is marked as "read-only". At 516, the virtualization stack completes the read intercept handling, and the hypervisor resumes execution of the guest virtual machine.
[0076] As shown at 518, some time passes. At 520, the virtual machine attempts a write access at GPA 0x1000. At 522, the hypervisor forwards the write access to the virtualization stack on the host. At 524, the virtualization stack performs a virtualization lookup for the host VA of GPA 0x1000. As previously mentioned, this is at VA 0x8501000. At 526, a virtual fault occurs for the write access on VA 0x8501000. At 528, the virtual fault returns SPA 0388000 in physical memory. At 530, the SLAT entry for GPA 0x1000 is updated to indicate that the data access is "read / write". At 532, the virtualization stack completes the write intercept handling. The hypervisor resumes execution of the guest virtual machine.
[0077] As shown at 534, a period of time passes. At 536, the host memory manager runs a page combination traversal to combine any pages that are functionally identical in the host physical memory. At 538, the host memory manager finds a combination candidate for VA 0x8501000 and another virtual address in another process. At 540, the host performs a TLB flush on VA 0x8501000. At 542, the virtualization stack intercepts the TLB flush. At 544, the SLAT entry for GPA 0x1000 is invalidated. At 546, a virtual machine intercept is performed for GPA 0x1000. At 548, a virtual fault occurs for a read access to VA 0x8501000. At 550, the virtual fault returns SPA 0x52000, which is a shared page among N processes in the page combination traversal at 536. At 552, the SLAT entry for GPA 0x1000 is updated to be associated with SPA 0x52000, with the access permission set to "read-only".
[0078] As shown at 554, a period of time passes. At 556, a virtual machine write intercept occurs for GPA 0x1000. At 558, a virtual fault occurs for a write access on VA 0x8501000. At 560, the host memory manager performs a copy write on VA 0x8501000. At 562, the host performs a TLB flush on VA 0x85010. As shown at 564, this invalidates the SLAT entry for GPA 031000. At 566, the virtual fault returns SPA 0x11000, which is a private page after the copy write. At 568, the SLAT entry for GPA 0x1000 is updated to SPA 0x1000, with the access permission set to "read / write". At 570, the virtualization stack completes the read intercept processing, and the hypervisor resumes virtual machine execution.
[0079] Thus, as described above, the virtual machine physical address space is supported by host virtual memory (usually allocated in the user address space of a host process), and the host virtual memory is subject to regular virtual memory management by the host memory manager. The virtual memory that supports the physical memory of the virtual machine can be of any type supported by the host memory manager 118 (private allocation, file mapping, page file - backed partial mapping, large page allocation, etc.). The host memory manager can perform its existing operations and apply policies and / or dedicated policies on the virtual memory to know that the virtual memory supports the physical address space of the virtual machine when necessary.
[0080] Moving on to Figure 6 , as described above, where the exemplary system 600 shown includes Figure 4 Overview and Figure 5Block diagram of the detailed mechanism. More specifically, guest virtual memory 311, guest physical memory 312, host virtual memory 321, and host physical memory 322 are shown as rectangular boxes, the width of which represents the range of the memory. For simplicity of illustration, the widths of the rectangular boxes are also approximately equal, even though in actual operation the number of virtual memories may exceed the number of physical memories that support such virtual memories.
[0081] As described above, when in a virtual machine computing environment (such as Figure 2When a process executing in the exemplary virtual machine computing environment 271) accesses a portion of the virtual memory provided by the operating system executing in such a virtual machine computing environment, the page table of the operating system (such as the exemplary page table 283) may include a PTE that may associate the accessed virtual memory address with a physical memory address, at least as perceivable by the process executing within the virtual machine computing environment. Thus, for example, if the virtual memory represented by region 611 is accessed, the PTE 631, which may be part of the page table 283, may associate this region 611 of the guest virtual memory 311 with a region 621 of the guest physical memory 312. As described in further detail above, the guest physical memory 312 may be supported by a portion of the host virtual memory 321. Thus, the virtualization stack 230 described above may include a data structure of entry 651 that may associate the region 621 of the guest physical memory 312 with a region 641 of the host virtual memory 321. Then, a process executing on the host computing device, including the above page table 250, may associate the region 641 of the host virtual memory 321 with a region 661 of the host physical memory 322. Also as described in detail above, the SLAT 220 may include a set of hierarchical arrangements of tables that may be updated with a table entry 691 that may associate the region 621 of the guest physical memory 312 with a region 661 of the host physical memory 322. More specifically, mechanisms such as those implemented by the virtualization stack 230 may detect the association between region 621 and region 661 and may update or generate a table entry in the table SLAT 220, such as entry 691, as shown by action 680, to include such an association. In this way, the guest physical memory 312 may be supported by the host virtual memory 321 while continuing to allow existing mechanisms, including the SLAT 220, to operate in their traditional manner. For example, subsequent accesses by a process executing in this virtual machine computing environment to the region 611 of the guest virtual memory 311 may require two page table lookups in the traditional manner: (1) a page table lookup executed in the virtual machine computing environment by referencing the page table 283, which may include entries such as entry 631 that may associate the requested region 611 with the corresponding region 621 in the guest physical memory 312, and (2) a page table lookup executed on the host computing device by referencing the SLAT 220, which may include entries such as entry 691 that may associate the region 621 in the guest physical memory 312 with the region 661 in the host physical memory 322, thus completing the path to the relevant data stored in the host's physical memory.
[0082] As described above, various page tables, such as the exemplary page table 283 in a virtual memory computing environment, the exemplary page table 250 on a host computing device, and the exemplary SLAT 220, may include layers of a hierarchical arrangement of the table such that entries of the lowest level table may identify "small pages" of memory, as that term is specifically defined herein, which may be the smallest memory address ranges that can be individually and carefully maintained and operated on by a process using the corresponding page table. In Figure 6 the exemplary system 600 shown, various page table entries (PTEs) 631, 671, and 691 may be at the lowest level of the table such that regions 611, 621, 641, and 661 may be small memory pages. As described above, in one common microprocessor architecture, such a small memory page may include 4KB of contiguous memory addresses, as specifically defined herein.
[0083] According to one aspect, to improve the speed and efficiency of memory access from a virtual machine computing environment, the physical memory of which is supported by the virtual memory of a host computing device, the SLAT 220 may maintain associations at a higher level of the table. In other words, the SLAT 220 may associate "large pages" of memory or "huge pages" of memory, as these terms are specifically defined herein. In one common microprocessor architecture, a large memory page, as the term is specifically defined herein, may include 2MB of contiguous memory addresses, while a huge memory page, as the term is specifically defined herein, may include 1GB of contiguous memory addresses. Again, as mentioned above, the terms "small page", "large page", and "huge page" are specifically defined with reference to levels in a hierarchical set of page tables and not based on a particular amount of memory, as the amount of memory may vary between different microprocessor architectures, operating system architectures, etc.
[0084] Turning to Figure 7 , a mechanism for creating large page entries in the SLAT 220 is shown to improve the speed and efficiency of accessing memory from a virtual machine computing environment, the physical memory of which is supported by the virtual memory of a host computing device. As Figure 6 shown, in Figure 7 the system 700 shown, guest virtual memory 311, guest physical memory 312, host virtual memory 321, and host physical memory 322 are shown as rectangular boxes, the width of which represents the range of memory. Additionally, access to region 611 may be as detailed above and as Figure 6In the manner shown. However, in addition to generating the lowest-level table entries within the SLAT 220, the mechanisms described herein can generate higher-level table entries, such as the exemplary table entry 711. As described above, such higher-level table entries can associate large or huge pages of memory. Thus, when the virtualization stack 230 causes the table entry 711 to be entered into the SLAT 220, the table entry 711 can associate a large-page-sized region 761 of the guest physical memory 312 with a large-page-sized region 720 of the host physical memory 322. However, as will be described in detail below, from the perspective of a memory manager executing on the host computing device and utilizing the page table 250, such a large-page-sized region is not a large page, and thus, such a large-page-sized region, for example, cannot be swapped out as a single large page by the memory manager executing on the host computing device, nor will it otherwise be treated as a single indivisible large page.
[0085] Conversely, in order to maintain the efficiency, density, and other advantages provided by the mechanisms detailed above, which enable guest physical memory such as the host virtual memory 312 to be supported by host virtual memory such as the exemplary host virtual memory 321, the page table 250 executing on the host may not include higher-level table entries similar to the higher-level table entry 711 in the SLAT 220, but may instead include only lower-level table entries, such as Figure 7 the exemplary lower-level table entries 671, 731, and 732 shown. More specifically, a particular set of lower-level table entries in the page table 250 (such as the exemplary lower-level table entries 671, 731, 732, etc.) can identify a contiguous range of the host physical memory 322, such that the overall contiguous range of the host physical memory 322 identified by the entire equivalent sequence of lower-level table entries in the page table 250 encompasses the same memory range as a single higher-level table entry 711 in the SLAT 220. Thus, as an example, in a microprocessor architecture where the small-page size is 4 KB and the large-page size is 2 MB, 512 contiguous small-page entries in the page table 250 can collectively identify a contiguous 2 MB of the host physical memory 322.
[0086] Because virtual faults may have been generated for memory regions less than the large page size, only specific entries in the lower-level table entries can be utilized, such as the exemplary lower-level table entries 671, 731, and 732, while the remaining lower-level table entries may still be marked as free memory. For example, an application executing within a virtual machine computing environment that utilizes the guest physical memory 312 may initially only require a small page size of memory. In this case, a single higher-level table entry 711 can identify the large page size memory region being used, such as the large page size region 720. However, from the perspective of the memory manager of the host computing device, referring to the page table 250, only a single lower-level table entry (such as the exemplary lower-level table entry 671) can be indicated as being utilized, and the remaining lower-level table entries are indicated as available. To prevent the other entries from being considered available, the virtualization stack 230 can also mark the remaining lower-level table entries, such as the exemplary lower-level table entries 731, 732, etc., as being used. When an application executing in the virtual machine computing environment needs to utilize additional small page size memory, the virtualization stack 230 can utilize memory locations that were previously indicated as being used but are not actually being used by the application executing in the virtual machine computing environment to satisfy the application's further memory requirements.
[0087] According to one aspect, the tracking of the equivalent sequence of lower-level table entries in the page table 250 (such as the markings described above) can be coordinated by the virtualization stack 230, as Figure 7 shown by action 730 in. In this case, if the memory manager executing on the host makes a change to one or more of these lower-level table entries, the virtualization stack 230 can trigger an appropriate change to the corresponding entry (such as the higher-level table entry 711) in the SLAT 220.
[0088] The sequence of lower-level table entries in page table 250 (equivalent to the higher-level table entries 711 in SLAT 220) can associate a sequence of host virtual memory ranges (such as exemplary ranges 641, 741, and 742) to respectively host physical memory ranges, such as exemplary ranges 661, 721, and 722. Thus, through the sequence of lower-level table entries in page table 250 (i.e., exemplary lower-level table entries 671, 731, 732, etc.), a large page size region of host virtual memory 321 (such as exemplary large page size region 751) is related to a large page size region of host physical memory 322 (such as exemplary large page size region 720). The virtualization stack 230 can maintain data structures such as hierarchically arranged tables that can coordinate host virtual memory locations to guest physical memory locations. Thus, the virtualization stack can maintain one or more data entries, such as exemplary data entry 771, which can associate the large page size region 751 of host virtual memory 321 with the corresponding large page size region 761 of guest physical memory 312. Although data entry 771 is shown as a single entry, the virtualization stack 230 can maintain the above associations as a single entry, as multiple entries, such as by individually relating sub-parts of region 761 to region 751, including sub-parts of small page size, or combinations thereof. To complete the cycle, the above higher-level table entry 711 in SLAT 220 can then associate the large page size region 761 in guest physical memory 312 with the large page size region 720 in host physical memory 322. Again, from the perspective of SLAT 220, with only a single higher-level table entry 711, the large page size regions 720 and 761 are large pages (according to the single higher-level table entry 711) that are uniformly processed by SLAT, while from the perspective of the memory manager executing on the host, page table 250, and in fact the virtualization stack 230, such regions are not single, but include, for example, individually manageable small page regions identified by the lower-level table entries in page table 250.
[0089] The availability of contiguous small memory pages in guest physical memory 312 (equivalent to one or more large or huge memory pages as defined herein) can be direct because guest physical memory 312 is perceived as physical memory by processes executing in the corresponding virtual machine computing environment. Thus, the following description focuses on the contiguity of host physical memory 322. Additionally, since guest physical memory 312 is perceived as physical memory, the entire memory access rights (such as "read / write / execute" rights) remain across the entire range of guest physical memory 312, and thus, there is unlikely to be a discontinuity in memory access rights among the lower-level table entries in page table 250.
[0090] Considering the page table 250 of the host computing device, if one or more large page size regions of the host physical memory 322 remain unused, then according to one aspect, when a virtual fault is triggered on a small page within a memory range that is divided into large page size regions that are still unused as described above, the remaining small pages within that large page size region can generate a single higher-level table entry (e.g., exemplary higher-level table entry 711) in the SLAT 220.
[0091] However, since from the perspective of the page table 250, the large page 720 corresponding to the higher-level table entry 711 is a sequence of small pages, such as exemplary small pages 661, 721, 722, etc., due to the lower-level table entry sequence, such as exemplary lower-level table entries 671, 731, 732, etc., each small page can be processed separately using the memory management mechanism of this page table 250, which can include swapping out such small pages to a non-volatile storage medium (such as a hard disk drive). In this case, according to one aspect, the higher-level table entry 711 in the SLAT 220 can be replaced with an equivalent sequence of lower-level table entries. More specifically, the lower-level table entries generated within the SLAT 220 can be consecutive and can as a whole reference the same memory address range as the higher-level table entry 711, except that the specific lower-level table entry corresponding to the page swapped out by the memory manager may be lost or may not be created.
[0092] According to other aspects, if a large page or huge page size region of the host physical memory 322 is unavailable when a virtual fault is triggered as described above, an available large page or huge page size region of the host physical memory can be constructed at interception by assembling an appropriate number of consecutive small pages. For example, a large page size region can be constructed from consecutive small pages obtainable from a "free list" or other similar enumeration of available host physical memory pages. As another example, if small pages currently being used by other processes are needed to establish the continuity necessary to assemble a sufficient number of small pages into a large page size region, then these small pages can be obtained from such other processes, and the data contained therein can be transferred to other small pages or swapped out to disk. Once such a large page size region is constructed, the virtual fault can be processed in the detailed manner described above, including generating a higher-level table entry in the SLAT 220, such as exemplary higher-level table entry 711.
[0093] In some cases, a memory manager executing on a host computing device may implement a background or opportunistic mechanism when utilizing page table 250, through which available large page or huge page size regions can be constructed from a sufficient number of contiguous small pages. To the extent that such a mechanism results in an available large page size region, for example, when a virtual fault is triggered, processing can proceed in the manner indicated above. However, to the extent that such a mechanism has not yet completed assembling an available large page size region, the construction of such a large page size region can be completed at a higher priority upon interception. For example, the process of constructing such a large page size region can be given a higher execution priority, or can be moved from the background to the foreground, or can be executed more aggressively.
[0094] Alternatively, if a large page size region of the host physical memory is not available when a virtual fault is triggered, multiple small pages can be utilized to process as described above. Subsequently, when the large page size region becomes available, the data in the multiple previously used small pages (due to the unavailability of the large page size region at that time) can be copied to the large page, and the previously used small pages can then be released. In such a case, a lower-level table entry in the SLAT 220 can be replaced with a higher-level table entry, such as the exemplary higher-level table entry 711.
[0095] Additionally, due to the skipping of table levels during SLAT lookup and the initial latency in constructing a large page size region from small pages, the parameters under which a large page size region is constructed can be varied based on a balance for faster memory access, either upon interception or opportunistically. One such parameter can be the number of contiguous small pages sufficient to trigger the swapping or eviction of other small pages required to complete contiguity across the entire large page size region. For example, if 512 contiguous small pages are required to construct a large page and there are contiguous ranges of 200 small pages and 311 small pages, and one small page between them is currently being used by another process, then such a small page can be evicted, or swapped with another available small page, so that all 512 contiguous small pages can be constructed into a large page size region. In such an example, the fragmentation of available small pages may be very low. In contrast, a contiguous range of only 20 - 30 small pages being continuously interrupted by other small pages currently being used by other processes can describe a range of small pages with very high fragmentation. Although it is still possible to swap, evict, or otherwise reorganize such a highly fragmented set of small pages to generate a contiguous range of small pages equivalent to a large page size region, such an effort may take longer. Thus, according to one aspect, a fragmentation threshold can be set, which can describe whether it is worth the effort to construct a large page size region from a range of small pages with higher or lower fragmentation.
[0096] According to one aspect, the construction of the large page size region can utilize the existing utilization of small pages by a process whose memory is used to support the guest physical memory of the virtual machine computing environment. As a simple example, if a process whose memory supports guest physical memory 312 has used a contiguous range of 200 small pages and 311 small pages, and only needs one small page between these two contiguous ranges to establish a contiguous range of small pages equivalent to a large page, then once the small page is freed (such as by swapping out or swapping), the only data that needs to be copied can be the data corresponding to that small page, and the rest of the data can remain in place, even though the corresponding lower-level table entry in the SLAT 220 can be invalidated and replaced with a single higher-level table entry that includes the same memory range.
[0097] To avoid having to reconstruct the large page size region of the memory from the contiguous range of small pages, an existing set of contiguous lower-level table entries in the page table 250 that as a whole includes the same memory range as a single higher-level table entry 711 in the SLAT 220 can be locked so that they can be avoided from being swapped out. Thus, for example, the virtualization stack 230 can request to lock the lower-level table entries 671, 731, 732, etc., so that the corresponding ranges of the host physical memory 661, 721, 722, etc. can retain the data stored therein without swapping it out to disk. According to one aspect, the decision to lock an entry such as the request made by the virtualization stack 230 can depend on various factors. For example, one such factor can be the frequency of use of the data stored in these memory ranges, where more frequently used data causes the entry to be locked, while less frequently used data causes the entry to be unlocked subsequently. Other factors can also be similarly referred to when deciding whether to request to lock the page table entry that supports the higher-level table entry in the SLAT, such as the amount of memory fragmentation in the entire system, the demand for a higher density virtual machine computing environment on the host computing device, and other similar factors.
[0098] Turning Figure 8 , where the flowchart 800 shown depicts the above mechanism as an exemplary series of steps. Initially, at step 810, a virtual fault as described above may occur. Subsequently, at step 815, it can be determined whether the virtual fault at step 810 references memory within the large page size memory region, where the virtual fault has not previously referenced other memory within that large page size region. In particular, the host physical memory can be divided into segments of large page size. A first virtual fault that references memory within a large page size region can result in an affirmative determination at step 815, whereby the processing can proceed to step 835.
[0099] In step 835, a single higher-level table entry in the SLAT can be generated. Correspondingly, in step 840, multiple lower-level table entries in the host page table can be utilized, where the multiple lower-level table entries of the relevant quantity refer to the memory regions having the same range as the generated single higher-level table entry in the SLAT, that is, the memory range corresponding to the large page size memory region that contains the memory location where the access right triggered the virtual fault in step 810. As mentioned above, since the host page table includes multiple lower-level table entries instead of a single higher-level table entry, the memory management of the host computing device can choose to swap out one or more memory regions or pages (identified by such lower-level table entries in the host page table), such as by copying their contents to the disk. If such a swap-out occurs, as determined at step 845, the process can proceed to step 850, and the higher-level table entry in the SLAT can be replaced with multiple lower-level table entries corresponding to the memory regions referenced by the non-swapped-out lower-level table entries in the host page table (and these level table entries corresponding to the memory regions actually swapped out in the host page table are removed from the SLAT).
[0100] According to one aspect, if in step 845 a small page in the large page size memory region is swapped out and the higher-level table entry in the SLAT is invalidated and replaced with a corresponding sequence of lower-level table entries (except for any lower-level table entry that references the swapped-out small page), then optionally, the process can proceed to step 855 to continue checking for available large page size regions of free memory. As mentioned above, for example, the memory management process executed on the host computing device can timely reconstruct the large page size regions of the available memory. If the large page size region of the memory becomes available as determined at step 855, the process can proceed to step 860, in which the lower-level table entries in the SLAT can be replaced with a single higher-level table entry in the SLAT, and, if appropriate, data can be copied from the memory region referenced by the previous lower-level table entry to the memory region now referenced by the single higher-level table entry (if these two memory regions are different).
[0101] Return to step 815. If the virtual fault at step 810 is not the first fault referencing memory within the previously described large page size region, the process can proceed to step 820. At step 820, it can be determined whether a large page should be created upon interception, such as in the manner detailed above. More specifically, if steps 835 and 840 were previously executed for a large page size memory region and the SLAT now includes a single higher-level table entry while the page table includes equivalent lower-level table entries that reference the memory region referenced by the single higher-level table entry in the SLAT, subsequent memory accesses to any memory location within the memory region referenced by the single higher-level table entry in the SLAT will not trigger a virtual fault, as detailed above. Thus, if a virtual fault was triggered at step 810 and at step 815 it is determined that the virtual fault was not first directed to a memory location within the previously depicted large page size memory region, the virtual fault triggered at step 810 may have been triggered because steps 835 and 840 were previously executed for a memory location within that large page size region, but subsequently, as determined at step 845, one or more small pages including that large page size memory region were swapped out, and thus step 850 was previously executed. Accordingly, at step 820, it can be determined whether a large page size memory region should be constructed upon interception. If such a large page size region should be constructed, the process can proceed to step 830 and an appropriate number of consecutive small pages can be assembled, as described above. The process can then continue with step 835 as before. Conversely, if on-demand construction of a large page size region of available memory is not requested at step 830, the process can continue to utilize the lower-level table entries in the host page table and the SLAT, in accordance with the mechanisms described in detail above with reference to Figure 4 and 5 the mechanisms described in detail above.
[0102] As a first example, a method for increasing the access speed of a computer memory is described above. The method includes: detecting, from a first process executing in a virtual machine computing environment, a first memory access directed to a first memory range; generating, as a prerequisite for completing the first memory access, a first entry in a two-level address translation table arranged in a hierarchy, the two-level address translation table arranged in a hierarchy associating a host physical memory address with a guest physical memory address, the first entry enabling identification of a second memory range identified by the first entry without referring to any table in the lowest-level table at at least one level above the lowest-level table; and in response to generating the first entry in the two-level address translation table arranged in a hierarchy, marking a first plurality of entries in a page table arranged in a hierarchy as used, the page table arranged in a hierarchy associating the host physical memory address with a host virtual memory address, the first plurality of entries as a whole referring to the same second memory range as the first entry in the two-level address translation table arranged in a hierarchy, wherein the entries in the first plurality of entries are in the lowest-level table; wherein the guest physical memory address is perceived by a process executing within the virtual machine computing environment as an address to physical memory; wherein the host physical memory address is an address to the actual physical memory of the host computing device hosting the virtual machine computing environment; wherein the host virtual memory address is an address to virtual memory provided by a memory manager executing on the host computing device, the memory manager using the page table arranged in a hierarchy to provide the virtual memory; and wherein the guest physical memory identified by the guest physical memory address is supported by a portion of the host virtual memory identified by a portion of the host virtual memory address.
[0103] A second example is the method according to the first example, further including: detecting, from the first process executing in the virtual machine computing environment, a second memory access directed to a third memory range different from the first memory range; and satisfying the second memory access by referring to a second subset of the first plurality of entries in the page table arranged in a hierarchy; wherein the first memory access is satisfied by referring to a first subset of the first plurality of entries in the page table arranged in a hierarchy, to a third memory range different from the first memory range.
[0104] A third example is the method according to the first example, further comprising: after the first memory access is completed, detecting that a first subset of the first plurality of entries in the hierarchically arranged page table has data that was originally stored in a corresponding host physical memory address and subsequently swapped out to a non-volatile storage medium; in response to the detection, invalidating the first entry in the hierarchically arranged two-level address translation table; and generating a second plurality of entries in the hierarchically arranged two-level address translation table to replace the first entry in the hierarchically arranged two-level address translation table, the second plurality of entries as a whole referencing at least some of the same second memory range as the first entry, wherein the entries in the second plurality of entries are in the lowest-level table.
[0105] A fourth example is the method according to the third example, wherein the second plurality of entries reference a portion of the second memory range that was previously accessed by the first process and not swapped out.
[0106] A fifth example is the method according to the first example, further comprising: assembling a first plurality of consecutive small page-size regions of the host physical memory into a single large page-size region of the host physical memory.
[0107] A sixth example is the method according to the fifth example, wherein the assembling occurs after the detecting of the first memory access and before the generating of the first entry in the hierarchically arranged two-level address translation table.
[0108] A seventh example is the method according to the fifth example, further comprising: generating a second entry in the hierarchically arranged two-level address translation table, the second entry in at least one level above the lowest-level table enabling identification of a third memory range identified by the second entry without referencing any table in the lowest-level table, the third memory range referencing the single large page-size region of the host physical memory into which the first plurality of consecutive small page-size regions of the host physical memory were assembled; and in response to generating the second entry in the hierarchically arranged two-level address translation table, marking a second plurality of entries in the hierarchically arranged page table as used, the second plurality of entries referencing the first plurality of consecutive small page-size regions of the host physical memory that were assembled into the single large page-size region of the host physical memory.
[0109] The eighth example is the method according to the seventh example, further comprising: generating a second entry in the two-level address translation table of the hierarchical arrangement, the second entry enabling identification of a third memory range identified by the second entry without referring to any table in the lowest-level table in at least one of the levels above the lowest-level table, the third memory range referring to the single large-page-size region of the host physical memory assembled from the first plurality of consecutive small-page-size regions of the host physical memory; and in response to generating the second entry in the two-level address translation table of the hierarchical arrangement, marking a second plurality of entries in the hierarchical page table as used, the second plurality of entries referring to the first plurality of consecutive small-page-size regions of the host physical memory assembled into the single large-page-size region of the host physical memory.
[0110] The ninth example is the method according to the seventh example, further comprising: copying data from a second set of one or more small-page-size regions of the host physical memory to at least a portion of the first plurality of consecutive small-page-size regions; and invalidating a second plurality of entries in the two-level address translation table of the hierarchical arrangement, the second plurality of entries including both: (1) a first subset of entries referring to the second set of one or more small-page-size regions of the host physical memory, and (2) a second subset of entries referring to at least some of the first plurality of consecutive small-page-size regions of the host physical memory, wherein the entries in the second plurality of entries are in the lowest-level table; wherein the generated second entry in the two-level address translation table of the hierarchical arrangement is utilized to replace the invalidated second plurality of entries.
[0111] The tenth example is the method according to the fifth example, wherein assembling the first plurality of consecutive small-page-size regions of the host physical memory includes copying data from some of the first plurality of consecutive small-page-size regions to other small-page-size regions of the host physical memory different from the first plurality of consecutive small-page-size regions.
[0112] The eleventh example is the method according to the tenth example, wherein the copying of the data from the some of the first plurality of consecutive small-page-size regions to the other small-page-size regions is performed only when the fragmentation of the first plurality of consecutive small-page-size regions is below a fragmentation threshold.
[0113] The twelfth example is the method according to the first example, further comprising: preventing paging of the second memory range.
[0114] The thirteenth example is the method according to the twelfth example, wherein the paging of the second memory range is performed only when the access frequency of one or more parts of the second memory range is greater than an access frequency threshold.
[0115] The fourteenth example is the method according to the twelfth example, further comprising: if the access frequency of one or more parts of the second memory range is less than the access frequency threshold, removing the block on the paging of the second memory range.
[0116] The fifteenth example is the method according to the first example, wherein the size of the second memory range is 2MB.
[0117] The sixteenth example is a computing device, the computing device comprising: one or more central processing units; a random access memory (RAM); one or more computer-readable media, comprising: a first set of computer-executable instructions that, when executed by the computing device, cause the computing device to provide a memory manager that references a hierarchically arranged page table to convert a host virtual memory address to a host physical memory address identifying a location on the RAM; a second set of computer-executable instructions that, when executed by the computing device, cause the computing device to provide a virtual machine computing environment, wherein a process executing within the virtual machine computing environment perceives a guest physical memory address as an address to physical memory, and wherein additional guest physical memory identified by the guest physical memory address is supported by a portion of the host virtual memory identified by a portion of the host virtual memory address; and a third set of computer-executable instructions that, when executed by the computing device, cause the computing device to: detect a first memory access directed to a first memory range from a first process executing in the virtual machine computing environment; as a prerequisite for completing the first memory access, generate a first entry in a hierarchically arranged two-level address translation table that associates a host physical memory address with a guest physical memory address, the first entry enabling identification of a second memory range identified by the first entry without referencing any tables in the lowest-level table at at least one level above the lowest-level table, the second memory range being larger than the first memory range; and in response to generating the first entry in the hierarchically arranged two-level address translation table, mark a first plurality of entries in the hierarchically arranged page table as used, the first plurality of entries as a whole referencing the same second memory range as the first entry in the hierarchically arranged two-level address translation table, wherein the entries in the first plurality of entries are in the lowest-level table.
[0118] The seventeenth example is a computing device according to the sixteenth example, wherein the third set of computer-executable instructions includes additional computer-executable instructions that, when executed by the computing device, cause the computing device to perform the following operations: after the first memory access is completed, detect that a first subset of the first plurality of entries in the hierarchical page table has data that was originally stored in a corresponding host physical memory address and was subsequently swapped out to a non-volatile storage medium; in response to the detection, invalidate the first entry in the hierarchical two-level address translation table; and instead of the first entry in the hierarchical two-level address translation table, generate a second plurality of entries in the hierarchical two-level address translation table, the second plurality of entries as a whole referencing at least some of the same second memory range as the first entry, wherein the entries in the second plurality of entries are in the lowest-level table.
[0119] The eighteenth example is a computing device according to the sixteenth example, wherein the third set of computer-executable instructions includes additional computer-executable instructions that, when executed by the computing device, cause the computing device to perform the following operations: assemble a first plurality of consecutive small page size regions of the host physical memory into a single large page size region of the host physical memory; wherein the assembly occurs after the detection of the first memory access and before the generation of the first entry in the hierarchical two-level address translation table.
[0120] The nineteenth example is a computing device according to the sixteenth example, wherein the third set of computer-executable instructions includes additional computer-executable instructions that, when executed by the computing device, cause the computing device to perform the following operation: prevent paging of the second memory range.
[0121] A twentieth example is one or more computer-readable storage media including computer-executable instructions that, when executed, cause a computing device to: detect a first memory access directed to a first memory range from a first process executing in a virtual machine computing environment; as a prerequisite for completing the first memory access, generate a first entry in a hierarchically arranged two-level address translation table that associates host physical memory addresses with guest physical memory addresses, the first entry enabling identification of a second memory range identified by the first entry without reference to any table in the lowest-level table in at least one level above the lowest-level table; and in response to generating the first entry in the hierarchically arranged two-level address translation table, mark a first plurality of entries in a hierarchically arranged page table as used, the hierarchically arranged page table associating the host physical memory addresses with host virtual memory addresses, the first plurality of entries as a whole referencing the same second memory range as the first entry in the hierarchically arranged two-level address translation table, where the entries in the first plurality of entries are in the lowest-level table; where the guest physical memory address is perceived by a process executing within the virtual machine computing environment as an address to physical memory; where the host physical memory address is an address to actual physical memory of a host computing device hosting the virtual machine computing environment; where the host virtual memory address is an address to virtual memory provided by a memory manager executing on the host computing device, the memory manager utilizing the hierarchically arranged page table to provide the virtual memory; and where guest physical memory identified by the guest physical memory address is supported by a portion of host virtual memory identified by a portion of the host virtual memory address.
[0122] From the above description, it can be seen that mechanisms that can accelerate memory access through SLAT have been described. Considering many possible variations of the subject matter described herein, we claim all such embodiments that fall within the scope of the appended claims and their equivalents as the present invention.
Claims
1. A method for improving the access speed of a computer memory, the method comprising: Receiving a request to load first data into the memory, wherein the first data includes at least one of the following: executable code or readable data; Determining a first memory access right expected to be set for a first memory, wherein the first data is stored in the first memory; If a memory block associated with the first memory access right has not been created, or if none of the previously created memory blocks associated with the first memory access right have an available memory amount sufficient to accommodate the first data, allocating a first memory block and loading only the executable code, readable data, or a combination of executable code and readable data expected to have the first memory access right into the first memory block; If the first memory block has been created and has the available memory amount sufficient to accommodate the first data, identifying a first available memory in the first memory block; And Loading the first data into the identified first available memory; Wherein in a page table that associates a memory address in a first memory addressing scheme with a memory address in a different second memory addressing scheme, a first memory range covered by the first memory block is identified by referring to a single table entry in a table that is at least one level higher than the lowest-level table in the page table, and the first memory range is identified without referring to any table in the lowest-level table.
2. The method according to claim 1, wherein the page table is a second-level access table SLAT maintained by a hypervisor.
3. The method according to claim 2, further comprising: Providing access to the first data to an operating system process when the first data is loaded into the identified first available memory; Wherein when the data is loaded into the identified first available memory, the operating system process verifies the data, and if the data is correctly verified, the operating system process instructs the hypervisor to set the first memory access right for the identified first available memory in the SLAT.
4. The method according to claim 1, wherein the first memory access right is a non-default memory access right, and the non-default memory access right includes one of the following: (1) read-only and execute rights, (2) read-only rights, or (3) no access rights.
5. The method according to claim 1, wherein the first memory range is 2MB.
6. The method according to claim 1, wherein the first memory range covered by the first memory block is identified by referring to a single table entry in a table that is two levels higher than the lowest-level table in the page table, and the first memory range is identified without referring to any table in the lowest-level table and without referring to any table in a second lowest-level table that is one level higher than the lowest-level table.
7. The method according to claim 6, wherein the first memory range is 1GB.
8. The method according to claim 1, wherein said allocating the first memory block comprises: Establish a starting address of the first memory range, the starting address being spaced apart from an ending address of a second memory range covered by a previous slice such that an intermediate memory range between the ending address of the second memory range and the starting address of the first memory range can be identified using a large memory page or a huge memory page.
9. The method according to claim 1 further comprises: Prevent demand paging for the first memory range.
10. The method according to claim 1, further comprising: Write a known secure data pattern to remaining available memory in the first memory block.
11. The method according to claim 1, further comprising: Determine that the identified first available memory into which the first data is loaded has a second memory access permission different from the first memory access permission; If a memory block associated with the second memory access permission has not been created, or if none of all previously created memory blocks associated with the second memory access permission have an amount of available memory sufficient to accommodate the first data, allocate a second memory block and load only executable code, readable data, or a combination of executable code and readable data expected to have the second memory access permission into the second memory block; If the second memory block has been created and has an amount of available memory sufficient to accommodate the first data, identify second available memory in the second memory block; Load the first data into the identified second available memory; and Remove the first data from the identified first available memory, the identified first available memory being part of the first memory block.
12. The method according to claim 1, further comprising: Pre-allocate at least one memory block corresponding to at least some non-default sets of memory access permissions; wherein the non-default sets of memory access permissions include: (1) read-only and execute permissions, (2) read-only permissions, and (3) no access permissions.
13. One or more computer-readable storage media, including computer-executable instructions that, when executed, cause a computing device to: Receive a request to load first data into a memory, wherein the first data includes at least one of the following: executable code or readable data; Determine a first memory access permission expected to be set for a first memory, the first data being stored in the first memory; If a memory block associated with the first memory access permission has not been created, or if none of all previously created memory blocks associated with the first memory access permission have an amount of available memory sufficient to accommodate the first data, allocate a first memory block and load only executable code, readable data, or a combination of executable code and readable data expected to have the first memory access permission into the first memory block; If the first memory block has been created and has an amount of available memory sufficient to accommodate the first data, identify first available memory in the first memory block; and Load the first data into the identified first available memory; Wherein in a page table that associates memory addresses in a first memory addressing scheme with memory addresses in a different second memory addressing scheme, a first memory range covered by the first memory block is identified by referencing a single table entry in a table that is at least one level higher than the lowest level table in the page table, and the first memory range is identified without referencing any tables in the lowest level table.
14. The computer-readable storage medium according to claim 13, wherein the first memory range covered by the first memory block is identified by referencing a single table entry in a table that is two levels higher than the lowest level table in the page table, and the first memory range is identified without referencing any tables in the lowest level table and without referencing any tables in a second lowest level table that is one level higher than the lowest level table.
15. The computer-readable storage medium according to claim 13, wherein the computer-executable instructions for allocating the first memory block include computer-executable instructions for: establishing a start address of the first memory range, the start address being spaced apart from an end address of a second memory range covered by a previous slice such that an intermediate memory range between the end address of the second memory range and the start address of the first memory range can be identified using one large memory page or one huge memory page.
16. The computer-readable storage medium according to claim 13, comprising additional computer-executable instructions that, when executed, cause the computing device to: Prevent demand paging for the first memory range.
17. The computer-readable storage medium according to claim 13, comprising additional computer-executable instructions that, when executed, cause the computing device to: Write a known secure data pattern to the remaining available memory in the first memory block.
18. The computer-readable storage medium according to claim 13, comprising additional computer-executable instructions that, when executed, cause the computing device to: Determine that the identified first available memory into which the first data is loaded has a second memory access right different from the first memory access right; If a memory block associated with the second memory access right has not been created, or if none of the previously created memory blocks associated with the second memory access right have an amount of available memory sufficient to accommodate the first data, allocate a second memory block and load only executable code, readable data, or a combination of executable code and readable data that is expected to have the second memory access right into the second memory block; If the second memory block has been created and has an available memory amount sufficient to accommodate the first data, identify second available memory in the second memory block; Load the first data into the identified second available memory; And Remove the first data from the identified first available memory, the identified first available memory being part of the first memory block.
19. The computer-readable storage medium according to claim 13, comprising additional computer-executable instructions that, when executed, cause the computing device to: Pre-allocate at least one memory block corresponding to at least some non-default sets of memory access rights; Among them, the memory access permissions of the non-default set include: (1) read-only and execute rights, (2) read-only rights, and (3) no access rights.
20. A computing device, comprising: One or more processing units; Hardware memory; And One or more computer-readable media comprising computer-executable instructions that, when executed by at least some of the one or more processing units, cause the computing device to: Receive a request to load first data into memory supported by the hardware memory, where the first data includes at least one of the following: executable code or readable data; Determine a first memory access right expected to be set for a first memory, the first data being stored in the first memory; If a memory block associated with the first memory access right has not been created, or if all previously created memory blocks associated with the first memory access right do not have an available memory amount sufficient to accommodate the first data, allocate a first memory block and load only the executable code, readable data, or a combination of executable code and readable data expected to have the first memory access right into the first memory block; If the first memory block has been created and has an available memory amount sufficient to accommodate the first data, identify first available memory in the first memory block; And Load the first data into the identified first available memory; wherein in a page table that associates memory addresses in a first memory addressing scheme with memory addresses in a different second memory addressing scheme, a first memory range covered by the first memory block is identified by reference to a single table entry in a table at least one level higher than the lowest-level table in the page table, and the first memory range is identified without reference to any table in the lowest-level table.