Dynamically allocatable physically addressed metadata storage - Patents.com
Patent Information
- Application Number
- JP2024504992
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2021-07-26
- Filing Date
- 2022-06-02
- Publication Date
- 2025-06-11
- Estimated Expiration
- 2042-06-02
AI Technical Summary
Existing hardware security schemes that associate metadata with physical memory result in unacceptable memory overhead, especially in cloud deployments where compute nodes reserve large portions of physical memory for metadata storage, leading to inefficient use of this expensive resource.
Implementing a mechanism with a small, fixed reservation of physical memory for a metadata summary table that stores pointers to dynamically allocated memory for metadata, allowing for fine-grained metadata storage while minimizing processing overhead.
This approach significantly reduces memory overhead by dynamically allocating metadata storage, ensuring efficient use of physical memory and maintaining performance by separating data and metadata storage, thus optimizing resource utilization in cloud environments.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
[Background technology]
[0001] Some hardware security schemes involve associating metadata with physical memory, so that when data is read from or written to the physical memory, the associated metadata can be examined to apply certain security policies or modify the behavior of the system.
[0002] The embodiments described below are not limited to implementations that address some or all of the shortcomings of known approaches that store and use metadata in physical memory. Summary of the Invention [Problem to be solved by the invention]
[0003] The following presents a simplified summary of the disclosure in order to provide the reader with a basic understanding. This summary is not intended to identify key features or essential features of the claimed subject matter, nor is it intended to be used to limit the scope of the claimed subject matter. Its sole purpose is to present a selection of concepts disclosed herein in a simplified form as a prelude to the more detailed description that is presented later. [Means for solving the problem]
[0004] In various examples, there is a computing device including a processor, the processor having a memory management unit. The computing device also has a memory storing instructions that, when executed by the processor, cause the memory management unit to receive a memory access instruction including a virtual memory address, convert the virtual memory address into a physical memory address of a memory, and obtain information associated with the physical memory address, the information being one or more of permission information and memory type information. In response to the information indicating that the metadata is permitted to be associated with the physical memory address, a metadata summary table stored in the physical memory is checked to see if the metadata is compatible with the physical memory address.
[0005] In response to the confirmation being negative, a trap is sent to system software of the computing device to allow dynamic allocation of physical memory for storing the metadata associated with the physical memory address. In some examples, the metadata summary table is cacheable in a processor cache.
[0006] Many of the attendant features will be more readily appreciated as the same becomes better understood by reference to the following detailed description considered in conjunction with the accompanying drawings, in which:
[0007] The present description will be better understood from the following detailed description read in conjunction with the accompanying drawings. [Brief description of the drawings]
[0008] [Figure 1] FIG. 1 is a schematic diagram of a data center showing the physical memories of the compute nodes of the data center, which store data and metadata separately. [Diagram 2] FIG. 2 is a schematic diagram of a computing device's central processing unit, cache lines, virtual memory, and physical memory, where the physical memory stores both data and metadata. [Diagram 3] FIG. 3 is a schematic diagram of a computing device's central processing unit, cache lines, virtual memory, and physical memory, where the physical memory stores data and metadata separately. [Figure 4] FIG. 4 is a schematic diagram of a memory management unit and tag controller of a computing device. [Diagram 5] FIG. 5 is a flow diagram of a method performed by the memory management unit. [Figure 6] FIG. 6 is a flow diagram of a method performed by the system software. [Figure 7] FIG. 7 is a flow diagram of a method performed by a tag controller for cache line eviction. [Figure 8] FIG. 8 is a flow diagram of a method performed by the tag controller for cache fill. [Figure 9] FIG. 9 illustrates an exemplary computing-based device upon which embodiments may be implemented. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0009] In the accompanying drawings, like reference numbers are used to denote like parts.
[0010] The detailed description provided below in conjunction with the accompanying drawings is intended as a description of the examples and is not intended to represent the only manner in which the examples may be constructed or utilized. The description sets forth the functions of the examples and the sequence of operations for constructing and deploying the examples. However, the same or equivalent functions and sequences may be accomplished by different examples.
[0011] In this document, the term "page" is used to refer to a unit of memory that may be the same as or different from the unit of memory used by a memory management unit.
[0012] The examples described herein describe a processor that is a CPU. The examples also work if the processor is any processor connected to a cache coherent interconnect, such as an accelerator.
[0013] As previously mentioned, some hardware security schemes involve associating metadata with physical memory, and can check the associated metadata to apply certain security policies when data is read from or written to the physical memory. Examples of such hardware security schemes include, but are not limited to, Memory Tag Extension (MTE) and Functional Hardware Extended RISC Instructions (CHERI). Metadata associated with physical memory can also be used for purposes other than hardware security schemes.
[0014] A typical hardware security scheme that involves associating metadata with physical memory carves out a range of physical memory to store the metadata. The inventors have recognized that for cloud deployments where compute nodes in the cloud deployment carve out ranges of physical memory to store metadata associated with the physical memory for hardware security schemes or other purposes, unacceptable memory overhead is incurred. In the case of MTE, the memory overhead can be as much as 1 / 32 of the total memory, even if only a single virtual machine (VM) on a server utilizes the metadata.
[0015] In various embodiments, there is a mechanism that includes a fixed reservation of physical memory (much smaller, typically 1 / 512) to store the metadata summary table, which can store the coarse-grained metadata if necessary, and stores a pointer to dynamically allocated physical memory that stores the fine-grained metadata. Implementing such a metadata summary table is not trivial, as it must impact as little as possible to operations that do not use metadata, while keeping processing overhead low. In embodiments, there is a mechanism for deploying a metadata summary table to store the fine-grained metadata in a typical central processing unit (CPU) design, dynamic allocation of physical memory, and a set of software abstractions to manage it.
[0016] 1 is a schematic diagram of a data center 100, showing the physical memory 114 of the computing nodes 104 of the data center 100, which individually store data and metadata associated with the data. The data center 100 comprises a number of computing nodes 102, 104. Each computing node has a central processing unit (CPU) and a physical memory 114. The computing nodes 102, 104 optionally communicate with each other via a communication network. The computing nodes are used to provide cloud computing services, such as by running applications and providing functionality of the applications to client devices over a communication network such as the Internet.
[0017] In the example, there is a first plurality of computing nodes 102 and a second plurality of computing nodes 102. Each computing node 104 has only virtual machines that use metadata. In contrast, each computing node 102 has at least one virtual machine that does not use metadata.
[0018] Figure 1 shows an expanded view of the physical memory of one of the computational nodes according to the first approach (left side) and according to the second approach (right side). In both the first and second approaches, data and metadata associated with the data are stored separately in the physical memory. Separately stored means that the locations in memory for a given data and its associated metadata are not necessarily contiguous.
[0019] According to the first approach (left side), a range of physical memory 106 is reserved in physical memory 114 for storing metadata, and the remaining physical memory 108 is for storing data. The inventors have recognized that the first approach results in unacceptable memory overhead, since each computing node reserves a range of physical memory 106 for storing metadata, even if the computing node is one of the computing nodes 102 with only one virtual machine using hardware security or other schemes. According to the first approach (left side), the range of physical memory 106 reserved for storing metadata is allocated at the boot time of the computing node, and often amounts to around 3% of the physical memory. Assume that a data center has multiple tenants, only 25% of which use hardware security or other schemes. The metadata allocation is not adjusted for the (non-)use of the scheme by the tenants, but still imposes an effective memory overhead of 12-13%, since 25% of the tenants using the scheme, and 25% of the total data physical memory, account for 100% of the metadata allocation. Memory is one of the most expensive resources in a data center, so reserving 12-13% of physical memory is unacceptable.
[0020] According to the second approach (right side), and according to an embodiment of the present disclosure, at boot time, a small (not necessarily contiguous) range of physical memory 110 is reserved for the metadata summary table by making a fixed reservation. The metadata summary table is stored, possibly using 1 / 512 of the physical memory. The metadata summary table allows for indirection (by storing pointers to locations in physical memory that store the metadata) and possibly for caching (where the metadata is stored). The metadata summary table is a representation of physical memory that holds information about which areas of physical memory store metadata, which areas store data, and which areas are not yet allocated. The metadata summary table itself stores some metadata in certain circumstances, as will be explained in more detail below, and in this respect provides a kind of caching functionality for the metadata it stores. Data is stored in the remaining physical memory 112, and in addition, a portion of the remaining physical memory 112 is dynamically allocated to store metadata. A pointer to the dynamically allocated physical memory area for storing metadata is stored in the metadata summary table. In this way, the data and its associated metadata are not stored together in physical memory, i.e., not spatially adjacent. An embodiment of the present technology relates to this second approach in which a metadata summary table is used. The metadata summary table is much smaller than the extent of the physical memory 106 in the first approach, resulting in significant memory savings compared to the first approach. The metadata summary table allows for indirection by storing pointers to locations in physical memory that store the metadata, so that even if the metadata summary table is much smaller than the extent of the physical memory in the first approach, there can still be enough memory to store the metadata.
[0021] Introducing indirection using a metadata summary table can increase the computational burden compared to the first approach. However, enabling some caching functionality can alleviate the burden. In some circumstances, storing the metadata in the metadata summary table itself reduces the computational burden by eliminating the need to look up pointers to locations in physical memory where the metadata is already present in the metadata summary table. In an example, if the metadata is the same in a contiguous region of physical memory, the metadata is kept in the metadata summary table.
[0022] In the approach of FIG. 1, data and metadata are stored separately in physical memory. Other approaches include storing metadata in bands, i.e., storing metadata in physical memory along with its associated data. Doing this involves using specialist physical memory that directly supports metadata. However, memory chips that directly support metadata are expensive and are not deployed in many current data centers. The inventors also recognized that using memory chips that directly support metadata would not provide the advantage of locality properties that can be exploited to improve performance, as described below. Some operations on metadata benefit from being able to inspect the metadata separately from the data. To perform these efficiently using physical memory that directly stores metadata, the memory needs to be able to query not only rows of data and metadata, but also columns of metadata.
[0023] 2 is a schematic diagram of a computing device's central processing unit 200, cache lines 202, virtual memory 206, and physical memory 210, which stores both data and metadata. FIG. 2 is background material included to aid in understanding this disclosure.
[0024] Assume that an application (APP in FIG. 2) running on a virtual machine on a computing device uses pages 212, 214 of memory assigned to a virtual memory address space 206. The pages of virtual memory 212, 214 are mapped to pages 216, 218 of pseudo-physical / guest-physical memory 208. The pages 216, 218 of pseudo-physical / guest-physical memory 208 are then mapped to physical memory pages 220, 222. To store metadata, the physical memory pages 220, 222 are slightly expanded as indicated by the shaded rectangles. The metadata is stored in physical memory 210, and the fact that address translation occurs does not affect the ability to discover the metadata. That is, given a virtual memory address, conventional address translation methods can be used to translate the virtual memory address to a physical memory address and use the physical memory pages to retrieve the metadata stored in physical memory. When data and its associated metadata are retrieved from physical memory and put into a cache line (as indicated by the arrow from physical memory page 220 to cache line chunk 204 in FIG. 2), the data and its metadata flow together and are stored together in cache line 202. However, this type of approach does not work because there is typically not enough space in physical memory to store the metadata along with the data. Therefore, embodiments of the present disclosure store the data and metadata separately in physical memory.
[0025] FIG. 3 is included as background material to aid in understanding the present disclosure. FIG. 3 is a schematic diagram of a central processing unit 200, cache lines 202, virtual memory 206, and physical memory 210 of a computing device, where the physical memory 210 stores data and metadata separately. FIG. 3 illustrates one approach to avoid a fixed carving of physical memory for storing metadata. However, the approach of FIG. 3 has drawbacks as explained below, which are addressed in the present disclosure. By storing data and metadata separately, there is sufficient space for metadata. With reference to FIG. 2, an application runs in a virtual machine of a computing device and uses virtual memory pages in virtual memory 206. In this example, the virtual memory pages comprise a virtual memory page 300 for metadata only and virtual memory pages 302, 304 for data only. With reference to FIG. 2, each page of virtual memory maps to a page in simulated physical / guest physical memory 208. With reference to FIG. 2, each page in simulated physical / guest physical memory 208 maps to a page in physical memory 210. When a cache line fill operation occurs, data and its associated metadata are retrieved from separate pages in physical memory 210 and stored in separate locations in cache line 202, as indicated by arrows 306, 308 between physical memory 210 and cache line 202 in Figure 3. Because the data and associated metadata are not stored together in cache line 202, subsequent operations on the data and associated metadata are not straightforward. Thus, embodiments of the present technology enable data and metadata to be stored together in cache line 202 or in a CPU's cache, such as a hierarchical cache, even when the data and its associated metadata are stored separately in physical memory.
[0026] FIG. 4 is a schematic diagram of a memory management unit 400 and a tag controller 406 of a computing device according to an embodiment of the present disclosure. In various embodiments of the present disclosure, the functionality for updating and making the metadata summary table available is implemented in the memory management unit 400 and / or the tag controller 406. The term “tag” is used to refer to a type of metadata, so the “tag controller” is a function for controlling the metadata. The tag controller does not necessarily operate in response to what the CPU is currently doing. The tag controller is part of the cache hierarchy, and therefore performs cache line evictions or cache fills. The memory management unit 400 is part of the CPU 200. The CPU communicates with the physical memory 210 through a cache, which in the example of FIG. 4 is a hierarchical cache that includes a level 1 cache 402, a level 2 cache 404, possible further levels of cache (shown by the dotted lines in FIG. 4), and the tag controller 406. Additional cache layers may also exist between the tag controller 406 and the physical memory 210. The data and metadata are stored together in any layer of cache between the tag controller 406 and the CPU 200, and separately in any layer between the tag controller 406 and the main memory 210. In an embodiment of the present disclosure, the CPU can operate as if the data and associated metadata are stored together in physical memory (when in fact they are stored separately). This is achieved by using a tag controller such that there is a single address in the physical memory address space that can be used to retrieve the data and its associated metadata. By using a tag controller as described herein, after address translation, there is a single physical address that is used to retrieve the data and the metadata.
[0027] Exactly how the data and metadata is stored in physical memory is transparent to the processor and most caches.
[0028] 4, the tag controller is at the bottom of the cache hierarchy immediately after the physical memory 210. Advantages of having the tag controller immediately after the physical memory include simplicity of implementation in some systems.
[0029] The computing device of Figure 4 has one or more memory controllers, which are not shown in Figure 4 for clarity. Because the memory controllers communicate directly with the physical memory, in some embodiments the tag controller 406 is internal to one of the memory controllers. In this case, when the tag controller 406 is internal to one of the memory controllers, the tag controller 406 has a small amount of cache internal to the tag controller for storing the metadata and, optionally, the metadata summary table entries.
[0030] Advantages of having a tag controller inside the memory controller include providing a single point of integration.
[0031] In some embodiments, the tag controller immediately precedes the last level cache, which stores data and metadata separately (unlike other cache levels where data and metadata are stored together). Advantages of this approach include dynamically trading the amount of data and metadata stored in the general purpose cache depending on the workload.
[0032] FIG. 5 is a flow diagram of a method performed by a memory management unit according to an embodiment of the present disclosure. The memory management unit has the functionality of a conventional memory management unit to translate between virtual memory addresses, pseudo / guest physical memory addresses, and physical memory addresses. In addition, the memory management unit is extended as described below. The memory management unit receives a memory access instruction (500), for example, as a result of an application running in a virtual machine running on a CPU. The memory access instruction includes a virtual memory address. The memory management unit translates the virtual memory address to obtain a physical memory address (502). As part of the translation process, the memory management unit also obtains information that is permission information and / or memory type information. The translation information and permission are provided to the MMU by software, either by explicit instruction or by software that maintains a data structure, such as a page table, that the MMU can directly reference. The permission information indicates whether permission is granted to store metadata associated with the physical memory address. The memory type information indicates whether the type of memory at the physical memory address is of a type for which metadata may potentially be associated with the physical memory address.
[0033] The memory management unit uses (at operation 504) the information obtained at operation 502 to determine whether the memory access instruction allows metadata. In response to the memory access instruction not allowing metadata, the memory management unit proceeds to operation 510 and resumes normal translation operations. In this manner, processing overhead is not affected when metadata is not used, such as when a hardware security scheme or other scheme that uses metadata is not required. This is very useful as it allows a zero-cost feature for virtual machines that do not use metadata storage. For computing devices with hardware that supports metadata and virtual machines that do not enable the use of metadata, the address translation does not allow metadata.
[0034] In response to the memory access instruction permitting metadata (at operation 504), the memory management unit proceeds to operation 506. At operation 506, the memory management unit checks whether the memory access instruction is compatible with the metadata. This check includes a query of a metadata summary table. In response to the memory access instruction being compatible with the metadata, the memory management unit proceeds with normal translation operations 510. Because the memory access instruction is compatible with the metadata, the physical memory address already has metadata storage associated with it.
[0035] In response to the memory access instruction being incompatible with the metadata (at operation 506), a trap is sent to the system software (508). A trap is a message or flag that indicates to the computing device's system software that there is a fault. The fault is that the memory access instruction requires metadata storage space in physical memory that is not currently allocated. The trap allows, enables, or triggers the system software to dynamically allocate space in physical memory to store the metadata for the memory access instruction. Having the memory management unit send the trap is beneficial in that the memory management unit operates more or less synchronously with the central processing unit. In contrast, the tag controller does not synchronize with the central processing unit. Thus, the inventors have recognized that it is beneficial to use the memory management unit, rather than the tag controller, to generate the trap, because generating a trap in the tag controller would likely result in a situation where the software attempting to allocate metadata space would trigger more cache evictions, making it very difficult to guarantee success.
[0036] Note that in FIG. 5, there are two decision diamonds 504, 506, which provides an advantage as opposed to having a single decision diamond. The process works even if decision diamond 504 is omitted. In the case of a single decision diamond, it is inferred (compatibility check 506) that metadata is allowed (i.e., decision diamond 504) from the fact that storage has already been allocated for the metadata. However, in this case, the metadata summary table needs to be known to the CPU, or at least to the MMU, which is problematic. Also, in the case of a single decision diamond, there is not enough information available to be able to distinguish between the various types of metadata schemes used in the data center. Given that some virtual machines in the data center use only MTE, some only CHERI, and some CHERI and MTE, the fact that some metadata storage is available is not enough to tell whether a particular type of metadata can be stored in a given page. Note that the term "page" refers to a unit of memory and is not necessarily the same as the unit of memory called a page used by the MMU.
[0037] FIG. 5 illustrates how instructions executed on a computing device may cause a memory management unit to proceed with a translation operation in response to the permission information indicating that metadata is not permitted at a physical memory address.
[0038] FIG. 5 illustrates how, in response to a metadata summary positively determining, instructions in a computing device cause a memory management unit to proceed with a transformation operation.
[0039] FIG. 6 is a flow diagram of a method performed by system software in a computing device according to an embodiment of the present disclosure. A non-exhaustive list of examples of system software are operating systems and hypervisors. The system software receives a trap as a result of operation 508 of FIG. 5 (600). The system software identifies a faulty address associated with the trap by inspecting the information received in the trap (602). The faulty address is a physical memory address examined as a result of the memory access instruction of FIG. 5, which is found to have no associated metadata storage even though it requires an associated metadata storage. The system software identifies locations in available physical memory to be allocated for metadata storage (604). The system software invokes one or more instructions or invokes firmware of the computing device to configure the metadata storage (606). The system software may explicitly or implicitly communicate with system software in other trust domains during this process. During configuration, microcode or firmware allocates the identified locations in physical memory for metadata storage. This includes marking a new page for use as metadata storage or using a new slot in an existing page.
[0040] The microcode or firmware performs a check to determine whether the specified location can be used for metadata storage for the specified data page within a hardware security scheme, or other scheme in which metadata is used.
[0041] The system software returns (608) from the trap handler so that the memory management unit knows that the trap has been handled. The MMU then retries the memory access instruction, i.e., repeats the process of Figure 5 from operation 500. This time, when the MMU reaches operation 506, the result is positive since the metadata storage has been allocated, and the MMU can proceed to operation 510.
[0042] FIG. 6 shows the system software: Receives a trap from the memory management unit, Identifying the faulty address, which is a physical memory address; Identifying a memory location for storing metadata associated with the physical memory address; Use microcode or firmware to 13 illustrates how instructions may be provided for configuring a specified memory location for storing metadata.
[0043] FIG. 6 illustrates how the system software includes instructions to direct the memory management unit to continue the translation operation depending on the successful configuration of the memory location identified for storing the metadata.
[0044] FIG. 7 is a flow diagram of a method according to an embodiment of the present disclosure. The method of FIG. 7 is performed by the tag controller for cache line eviction, where a cache line is evicted to a lower level cache or physical memory. When a cache line is evicted to a lower level cache, the lower level cache is a cache that stores data and metadata separately. As previously described, the tag controller is part of the cache hierarchy and is responsible for evicting cache lines from and filling cache lines into that part of the cache hierarchy. The tag controller optionally has its own internal cache, referred to herein as the tag controller's internal cache. When a cache line is evicted (700), the tag controller checks its own internal cache (702) to determine whether configured metadata storage exists for the location in the physical memory or lower level cache where the cache line contents are evicted. If the results from the internal cache are inconclusive, the tag controller queries the metadata summary table (704) to determine whether configured metadata storage exists. If the tag controller has an internal cache, the tag controller optionally caches the results of the query in a metadata summary table in its internal cache.
[0045] Thus, the tag controller may determine whether metadata storage is configured for where the cache line contents are to be evicted (706). In response to metadata storage not being configured, the tag controller begins discarding metadata in the evicted cache line (708).
[0046] In response to the metadata storage being configured at decision diamond 706, the tag controller checks (710) whether the metadata in the cache line being evicted has sufficient locality. Sufficient locality means that the metadata is the same in relatively nearby memory locations. A cache line is made up of multiple contiguous chunks of data, and each chunk may have an item of metadata stored with the chunk in the cache line (because in a cache line, metadata and data are stored together). If the metadata is the same for multiple contiguous chunks, there is sufficient locality. The tag controller checks the chunk's metadata to check whether there are more than a threshold number of contiguous chunks with the same metadata. The threshold is preset and depends on the amount of space available in the summary table.
[0047] If the metadata locality in the cache line is not sufficient, the tag controller writes the metadata to physical memory (714). The data from the evicted cache line is also written to physical memory. Since data and metadata are stored separately in physical memory, the metadata is written to a different location in physical memory than the data. The tag controller optionally stores the metadata associated with the cache line in its internal cache (or needs to write the metadata to physical memory or a lower level cache). Over time, metadata accumulates in the tag controller's internal cache, and when the internal cache becomes full, the tag controller writes the metadata back to physical memory. In this way, the frequency of writes to physical memory decreases.
[0048] If there is sufficient metadata locality in the cache line, the tag controller writes the metadata in a compressed format to the metadata summary table (718). It is assumed that the metadata for each chunk in the cache line is the same. In this case, the metadata written to the metadata summary table is one instance of the metadata for a single chunk, indicating that the same metadata fits across the entire cache line. The compressed format of the metadata is the value of the metadata and the range of memory locations where the metadata value fits. The tag controller optionally uses its internal cache to accumulate the compressed format of the metadata before writing it to the metadata summary table (716). The caching and compression of the metadata serves to reduce accesses (both reads and writes) to the metadata storage in physical memory. In various examples, there is a cache of both the metadata summary table and the metadata.
[0049] The inventors have recognized that sufficient locality checking in operation 710 of FIG. 7 provides useful performance advantages. Examples of MTE and CHERI locality properties are given below. Both MTE and CHERI have the same basic requirement from the memory system: storing out-of-band data transmitted through the cache hierarchy. MTE uses 4 bits per 16-byte granule to store metadata. CHERI uses 1 bit per 16-byte granule to indicate whether the data is a valid function or not. Both forms of metadata have useful locality properties, such as: The MTE metadata is the same for all granules in an allocation, so large allocations have consecutive runs of the same metadata values. This is especially true for mapping large files, where entire pages have the same metadata values. Because most programs contain large amounts of contiguous non-pointer data, CHERI systems have long runs with invalid tag bit values. These can even cover entire pages.
[0050] FIG. 7 illustrates the use of a tag controller that is part of a cache hierarchy of a computing device, the tag controller being configured to take metadata into account during cache line evictions and / or cache fills of at least a portion of the cache hierarchy.
[0051] Boxes 700, 702, and 704 in FIG. 7 show how the tag controller is configured, as part of a cache line eviction, to determine whether metadata storage is configured for the physical memory address from which the cache line is being evicted by checking one or more of the tag controller's cache and metadata summary table.
[0052] The negative result of diamond 706 in FIG. 7 illustrates how the tag controller may be configured to discard metadata for the evicted cache line in response to finding that metadata storage is not configured.
[0053] Diamond symbol 710 in FIG. 7 illustrates how the tag controller is configured, in response to finding that metadata storage is configured, to check whether the metadata of the evicted cache line has sufficient locality, and in response to sufficient locality being found, to write the metadata of the evicted cache line to a metadata summary table, and in response to sufficient locality not being found, to write the metadata to physical memory.
[0054] FIG. 8 is a flow diagram of a method performed by a tag controller for cache fill according to an embodiment of the present disclosure. A cache to be filled is a cache that stores both metadata and data. The tag controller receives a cache fill request from a higher level cache in the cache hierarchy (800). The cache fill request includes a physical memory address from which data is desired to be obtained to fill the cache. The tag controller checks its own internal cache (if it has an internal cache available) to determine whether there is metadata storage configured for the physical memory address. If the result is inconclusive (i.e., the metadata is not found in the internal cache), the tag controller queries a metadata summary table and, optionally, caches the results in the internal cache (804), since the metadata summary table has information about the physical memory and whether metadata is configured for the unit of physical memory.
[0055] At decision diamond 806, the tag controller is in a position to determine (using the information obtained in acts 802 and 804) whether metadata storage is already configured for the physical memory address. If not, the tag controller retrieves the data from physical memory (808). The tag controller then sets the metadata associated with the data in the cache to default values (812). The cache stores both data and metadata. In this case, the data does not have associated metadata, so the tag controller fills the metadata field of each chunk in the cache with default values. The tag controller then fills the cache (810). If the cache supports independent tracking of the validity of the data and metadata, the tag controller may provide the data and metadata as soon as they are available, rather than combining them into a single message.
[0056] If, at decision diamond 806, it is determined that metadata storage is configured for the physical memory address, the tag controller proceeds to decision diamond 814. At decision block 814, the tag controller checks the metadata summary table entry for the physical memory address to see if the metadata summary table entry has sufficient information. If so, the metadata summary table entry includes the metadata. The tag controller proceeds to retrieve data from the physical memory address and add the metadata to the data (816). The tag controller then fills the cache with the data and the added metadata (818).
[0057] At decision diamond 814, if the metadata summary table entry does not already contain the metadata, the metadata summary table entry comprises a pointer to the location in physical memory where the metadata is stored. Thus, the tag controller gets the location in physical memory where the metadata is stored (820) and gets the metadata (822). The tag controller also gets the data from physical memory (because data is stored in physical memory separately from the metadata) (824), adds the metadata to the data, and fills the cache with the data and the added metadata (818).
[0058] Boxes 800, 802, and 804 in FIG. 8 illustrate how the tag controller, as part of a cache fill of a cache hierarchy, is configured to determine whether metadata storage is configured for a physical memory address from which data is to be written into the cache hierarchy by checking one or more of the tag controller's cache and metadata summary table.
[0059] The negative result of diamond 806 in FIG. 8 illustrates how the tag controller is configured, in response to finding that metadata storage is not configured, to fill the caches of the cache hierarchy with data and metadata where the metadata is set to default values.
[0060] The positive result of diamond 814 in FIG. 8 illustrates how the tag controller is configured, in response to finding that metadata storage is configured, to check the metadata summary table, and, in response to associated metadata being found in the metadata summary table, to fill the cache hierarchy with data and associated metadata.
[0061] The negative result of diamond 814 in FIG. 8 illustrates how the tag controller is configured, in response to finding that metadata storage is configured, to check the metadata summary table, and, in response to the associated metadata not being found in the metadata summary table, to obtain the location of the associated metadata in physical memory from the metadata summary table.
[0062] The methods described above with reference to Figures 7 and 8 allow for the compression of long runs of the same metadata and for example, providing a fast path when the metadata is uniform across a page. Caching such runs of metadata within the tag controller can avoid many reads to physical memory, and in fact can omit the more detailed metadata store entirely for pages where system software guarantees uniform metadata. With CHERI, opportunistic compression of runs of zeroed metadata is also useful for system software to sweep memory for functions where metadata function data does not need to be retrieved from memory for such runs.
[0063] The embodiments described herein enable dynamically allocated hierarchical metadata storage, allowing storage to be tailored per VM or per application.
[0064] In some embodiments described herein, a computing device has a policy for metadata storage.
[0065] In examples, policies for metadata storage are managed by the operating system kernel of the computing device in non-virtualized systems, the hypervisor in virtualized systems, and the hypervisor combined with privileged microcode or firmware in systems with fuzzy privilege boundaries and traditional virtualization. In embodiments in which the computing device implements confidential computing, the mechanisms for metadata storage are implemented in the privileged firmware or microcode. If there is a confidential computing system that removes the hypervisor from the trusted computing base for confidentiality or integrity reasons, the privileged firmware or microcode is configured such that no entity in the system can inspect or modify the metadata for a page, where the data cannot be modified either. In systems without support for confidential computing, other system software, such as an operating system (OS) or hypervisor, can be trusted to maintain the same guarantees.
[0066] In embodiments in which the computing device implements confidential computing, the metadata summary table is managed by privileged firmware or microcode, and its layout does not need to be architecturally exposed outside the privileged firmware or microcode. There is a simple mapping from any physical address to a corresponding entry in the metadata summary table. In systems that provide confidential computing support, the metadata summary table information used by the MMU may be stored in the same location as the reverse mapping or similar metadata for confidential computing. In this way, no additional memory accesses are required for the MMU.
[0067] The cache and memory system has coherence points below which metadata and data are stored separately and above which they are merged and flow into a single cache line. The tag controller is responsible for assembling data and metadata on loads, splitting them on evictions, and maintaining a cache of metadata values.
[0068] The embodiments described herein provide the advantage of "you don't pay for what you don't use": large memory carve-ups exist only for VMs that use architectural features that require reserved physical memory, and the carve-up is no larger than is required by the particular feature requested.
[0069] The embodiments described herein offer the advantage that the (micro)architecture is as simple as possible in order to achieve high performance.
[0070] In examples using confidential computing, the hypervisor, or microcode, or firmware is configured to ensure that metadata storage is not removed from pages while the metadata may still be present in the cache. When a metadata page is removed from physical memory, the stale data in the cache is safely discarded since the corresponding data page has also been removed. As with returning a data page to an unsecure memory region, it is the responsibility of the microcode or firmware to ensure that such data is flushed from the cache before returning the metadata page to the untrusted hypervisor.
[0071] In instances using confidential computing, as described herein, there is a software interface to the privileged firmware or microcode, which comprises function calls that operate across the privilege boundary.
[0072] The privileged firmware or microcode exposes a small number of functions that manage the metadata summary table and the metadata it points to. The metadata summary table can be thought of as a contract between the hardware and firmware, read by both the tag controller and the MMU, and maintained by the firmware. The tag controller may write to the metadata summary table and store metadata as the top level in a hierarchical tag storage design. The system software at the top of the firmware has a consistent interface to the firmware that does not expose the implementation details of the metadata summary table.
[0073] Each metadata page is logically treated as an array of metadata stores, one metadata store per data page. The number of metadata stores in a single metadata page varies depending on the features supported. If a VM has the globally disabled metadata usage feature, the MMU can skip metadata storage checks entirely. This means that the hypervisor can either provide metadata storage as each page is assigned to the VM, or lazily dynamically allocate detailed metadata storage for a page after the VM maps it as a metadata page.
[0074] The hypervisor may also over-provision memory if necessary. For example, a cloud provider may add 3.125% of the memory allocated to a regular VM with MTE enabled to the pool available for other VMs, to be reused if metadata storage for those pages is desired. On a physical machine with 1TiB of memory, if all VMs have MTE enabled but use MTE for only half of the pages, this leaves 16GiB of memory that can be temporarily used for other VMs as long as it is available for reuse by the hypervisor when it detects a failure.
[0075] In various embodiments, a computing device includes a privilege boundary, and instructions executed on the more trusted side of the boundary expose a number of functions across the privilege boundary to manage one or more of the metadata summary table and the metadata to which the metadata summary table points.
[0076] In an example, one of the functions receives as arguments a physical page address and a set of metadata features, returns true if the page can be used for metadata storage, where the page is the same or a different unit of memory as the unit of memory used by the memory management unit, or creates a mapping between a data page and a metadata storage slot, or decouples a data page from its associated metadata page such that confidential computing guarantees of the host architecture are met.
[0077] In an example, one of the functions handles transitioning a page from metadata storage usage to data storage usage, such that if the page transitions across a privilege boundary, the page is not used for metadata storage.
[0078] Querying the size of metadata A particular implementation may need to reserve different amounts of storage for different micro-architectural features, which may exceed the amount of architectural state. In an example, some VMs in a data center use MTE and some VMs use CHERI. In this case, it is useful to have the ability to query the firmware or microcode to find out the amount of space required for metadata storage (as it differs depending on whether MTE or CHERI is being used). In an example, this function receives as an argument an indication of the type of metadata (e.g., MTE or CHERI) to be stored for a given page.
[0079] Creating a Metadata Page In an example, the privileged firmware or microcode exposes a function to transfer a page from a data page to a metadata page. For confidential computing, the function for creating the metadata page is configured to keep the data page at an appropriate privilege level so that confidential computing is not compromised. The function receives a physical page address as an argument and returns true or false depending on whether the physical page was successfully transferred to the page for storing the metadata.
[0080] Using this feature, physical memory is removed from the pool of generally available physical memory before it can be used to store metadata.
[0081] For confidential computing, the firmware checks whether the page passed as an argument to this function is owned by the privilege level of the caller.
[0082] Metadata storage allocation for the page In an example, the privileged firmware or microcode exposes a function to allocate metadata storage for a page.
[0083] Each data page that has associated metadata has an associated metadata storage slot.
[0084] A metadata storage slot is a unit of physical memory within a page for storing metadata. A function is used to create that mapping. In the example, the function takes as an argument the physical address of a data page. Another argument is the physical address of the page allocated for storing the metadata. Another argument is the index or address of an available slot in the metadata page.
[0085] This function performs sufficient checks to enforce the confidential computing guarantees of the host architecture. If metadata storage was successfully allocated for the page, the function returns true; otherwise, the function returns false. If the metadata or data page is not owned by the caller, if the metadata page is not a metadata page, or if the data page already has metadata allocated, the function returns failure.
[0086] If there are no errors, the metadata summary table entry for the data page is updated to point to the specified slot in the given metadata page, and the state of the metadata summary table associated with the metadata page is updated so that the metadata page stores at least enough metadata to prevent it from being released while still referenced by the data page.
[0087] Removing metadata storage from a data page In an example, the privileged firmware or microcode exposes the ability to remove metadata storage from a data page.
[0088] Before a metadata page is reused for normal use, any data pages that use it to store metadata are detached using a function that removes the metadata storage from the data page. The caller of this function removes the data page from the translation table and invalidates any cache lines containing data from the data page. This ensures that no cache lines containing metadata exist above the tag controller.
[0089] The function to remove metadata storage from a data page takes the physical address of the data page as an argument and returns a value indicating success or failure. The function returns failure if the data page is not owned by the caller, is not a data page, or is a data page with no assigned metadata.
[0090] If the function call is successful, the metadata summary table entry for the data page is reset to indicate that there is no metadata, the recorded state associated with the metadata page is updated to reflect the fact that this reference is gone, and the firmware invalidates cache lines that reference this data page.
[0091] Querying the metadata storage for a data page In an example, the privileged firmware or microcode exposes the ability to query metadata storage for a data page.
[0092] This function takes the physical address of a data page as an argument and returns a value indicating success or failure.
[0093] If the data page is not owned by the caller, is not a data page, or does not have metadata associated with it, the function returns failure.
[0094] Reusing Metadata Pages In an example, the privileged firmware or microcode exposes functionality for reclaiming metadata pages.
[0095] Dynamic metadata storage is intended to allow system software to move memory between data and metadata usage and vice versa. This function handles the transition of pages from metadata storage usage to data storage usage.
[0096] This function takes the physical address of a metadata page as an argument and returns true if the page has been deleted, false otherwise.
[0097] This facility performs sufficient validation so that the confidential computing guarantees of the host architecture can be enforced.
[0098] FIG. 9 illustrates various components of an exemplary computing-based device 900 embodied as any form of computing and / or electronic device, such as a mobile phone, a data center computing node, a desktop personal computer, a wearable computer, etc., in which embodiments of dynamically allocable metadata storage are implemented in some examples.
[0099] The computing-based device 900 includes one or more processors 910, which may be a microprocessor, controller, or any other suitable type of processor, for processing computer-executable instructions that control the operation of the device to dynamically allocate metadata storage. The processor has a memory management unit 914. In some examples, for example when a system-on-chip architecture is used, the processor 910 includes one or more fixed function blocks (also referred to as accelerators) that implement portions of any of the methods of Figures 5-8 (rather than software or firmware). The cache hierarchy 916 includes a tag controller (not shown in Figure 9). System software 904 is provided in the computing-based device to enable application software 906 to execute in the device.
[0100] The computer-executable instructions are provided using any computer-readable medium accessible by the computing-based device 900. Computer-readable media includes, for example, computer storage media, such as memory 902, and communication media. Computer storage media, such as memory 902, include volatile and nonvolatile, removable and non-removable media implemented in any method or technology for storage of information, such as computer-readable instructions, data structures, program modules, and the like. Computer storage media includes, but is not limited to, random access memory (RAM), dynamic random access memory (DRAM), read-only memory (ROM), erasable programmable read-only memory (EPROM), electronically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassettes, magnetic tapes, magnetic disk storage or other magnetic storage devices, or any other non-transmission medium used to store information for access by a computing device. In contrast, communication media embodies computer readable instructions, data structures, program modules, etc., in a modulated data signal such as a carrier wave, or other transport mechanism. As defined herein, computer storage media does not include communication media. Thus, computer storage media should not be interpreted as a propagating signal itself. While computer storage media (memory 902) is illustrated within computing-based device 900, it will be understood that in some instances the storage is distributed or remotely located and accessed over a network or other communications link (e.g., using communications interface 912).
[0101] The term "computer" or "computing-based device" is used herein to refer to any device with processing capabilities to execute instructions. Those skilled in the art will recognize that such processing capabilities may be incorporated in many different devices, and thus the terms "computer" and "computing-based device" each include personal computers (PCs), servers, mobile phones (including smartphones), tablet computers, set-top boxes, media players, gaming consoles, personal digital assistants, wearable computers, and many other devices.
[0102] The methods described herein are in some examples performed by software in machine-readable form in a tangible storage medium, such as in the form of a computer program comprising computer program code means adapted to perform all operations of one or more of the methods described herein when the program is executed on a computer, the computer program may be embodied in a computer readable medium. The software is suitable for execution on a parallel or serial processor such that the method operations may be performed in any suitable order, or simultaneously.
[0103] Those skilled in the art will recognize that storage devices utilized to store program instructions are, optionally, distributed across a network. For example, a remote computer can store an example of the process described as software. A local or terminal computer can access the remote computer and download some or all of the software to execute the program. Alternatively, the local computer may download parts of the software as needed, or execute some software instructions at the local terminal and some software instructions at the remote computer (or computer network). Those skilled in the art will also recognize that all or some of the software instructions can be executed by dedicated circuitry, such as a digital signal processor (DSP) or programmable logic array, by utilizing conventional techniques known to those skilled in the art.
[0104] As will be apparent to one skilled in the art, any ranges or device values given herein may be expanded or modified without losing the desired effect.
[0105] Although the subject matter has been described in language specific to structural features and / or methodological acts, it is to be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or acts described above. Rather, the specific features and acts described above are disclosed as example forms of implementing the claims.
[0106] It will be understood that the benefits and advantages described above may relate to one embodiment or to several embodiments. The embodiments are not limited to those that solve some or all of the problems described or those that have any or all of the benefits and advantages described. Further, it will be understood that references to "an" or "an" item refer to one or more of those items.
[0107] The operations of the methods described herein may be performed in any suitable order, or simultaneously as appropriate. Additionally, individual blocks may be deleted from any method without departing from the scope of the subject matter described herein. Aspects of any of the above-described examples may be combined with aspects of any of the other described examples to form further examples without losing the desired effect.
[0108] The term "comprising" is used herein to mean that, although specified method blocks or elements are included, such blocks or elements do not comprise an exclusive list and the method or apparatus may include additional blocks or elements.
[0109] The term "subset" is used herein to refer to a proper subset of a set, where the subset does not comprise all elements of the set (i.e., at least one element of the set is missing from the subset).
[0110] It will be understood that the above description is given by way of example only, and that various modifications may be made by those skilled in the art. The above specification, examples, and data provide a complete description of the structure and use of the exemplary embodiments. Although various embodiments have been described above with a degree of particularity, or with reference to one or more individual embodiments, those skilled in the art will be able to make many modifications to the disclosed embodiments without departing from the scope of the present specification.
Claims
1. A computing device, comprising: a processor having a memory management unit; a memory for storing instructions; wherein the instructions, when executed by the processor, cause the memory management unit to: receive a memory access instruction including a virtual memory address; convert the virtual memory address to a physical memory address of the memory and obtain information associated with the physical memory address, the information including permission information and / or memory type information, the permission information indicating whether permission is given to store metadata associated with data at the physical memory address, and the memory type information indicating whether the metadata is associated with the physical memory address; determine whether the obtained information indicates that it is permitted to associate the metadata with the physical memory address; based on the determination that it is permitted to associate the metadata with the physical memory address, execute a query on a metadata summary table stored in physical memory, the metadata summary table enabling indirection by storing pointers to locations within the physical memory that store metadata; based on executing the query on the metadata summary table that stores the pointer, determine whether the metadata is compatible with the physical memory address based on the physical memory address configured to store the metadata; when it is determined that the metadata is incompatible, send a trap to the system software of the computing device, the trap triggering dynamic allocation of the physical memory for storing metadata associated with the physical memory address; use a tag controller, the tag controller being part of the cache hierarchy of the computing device, the tag controller being configured to perform cache line eviction and / or cache fill of at least a part of the cache hierarchy based on the metadata. As part of the cache line eviction, the tag controller is configured to determine that the metadata storage is configured for the physical memory address from which the cache line is evicted by checking the cache of the tag controller and / or the metadata summary table. The tag controller is configured to check whether the metadata of the cache line to be evicted is the same within a threshold number of consecutive chunks, and if the metadata is the same within a threshold number of consecutive chunks, write the metadata of the cache line to be evicted to the metadata summary table, and if the metadata is not the same within a threshold number of consecutive chunks, write the metadata to the physical memory. A computing device. **Claim 2** If the obtained information indicates that the metadata is not permitted at the physical memory address, the instruction causes the memory management unit to proceed with the conversion operation, or if the metadata is determined to be compatible, the instruction causes the memory management unit to proceed with the conversion operation. The computing device according to claim 1. **Claim 3** The memory stores system software for execution in the processor, and the system software receives the trap from the memory management unit, identifies a faulty address that is the physical memory address where it has been determined that there is no associated metadata storage despite the need for associated metadata storage for the memory access instruction, identifies a memory location for storing the metadata associated with the physical memory address, and includes instructions for configuring the identified memory location to store the metadata using microcode or firmware. The computing device according to claim 1. **Claim 4** The system software includes instructions for instructing the memory management unit to continue the conversion operation if the identified memory location for storing the metadata is successfully configured. The computing device according to claim 3. **Claim 5** The tag controller is either part of the memory controller of the computing device, or is the last level cache of the cache hierarchy and is placed immediately before the last level cache that stores data and metadata separately, or is between the last level cache of the cache hierarchy and the memory controller of the computing device. The computing device according to claim 1.
6. A computing device according to claim 1, comprising a cache hierarchy, wherein the physical memory stores the data and the associated metadata separately, and the cache hierarchy stores the data and the associated metadata together.
7. The tag controller is configured to discard the metadata of the cache line to be evicted when it discovers that the metadata storage is not configured. The computing device according to claim 1.
8. The tag controller is configured to determine whether metadata storage is configured for a physical memory address from which data is written to the cache hierarchy at the physical memory address by checking the cache of the tag controller and / or the metadata summary table as part of the cache fill of the cache hierarchy. The computing device according to claim 1.
9. The tag controller is configured to fill the cache of the cache hierarchy with the data and metadata set to default values when it discovers that the metadata storage is not configured. The computing device according to claim 8.
10. The tag controller is configured to check the metadata summary table when it discovers that the metadata storage is configured, and if the relevant metadata is found in the metadata summary table, to fill the cache hierarchy with the data and the relevant metadata. The computing device according to claim 8.
11. When the tag controller discovers that the metadata storage is configured, it checks the metadata summary table. If the related metadata is not found in the metadata summary table, it is configured to obtain, using a pointer, the location in the physical memory where the related metadata is stored, and the pointer is stored in the metadata summary table. The computing device according to claim 8.
12. The computing device according to claim 1, comprising a privilege boundary, and instructions executed on the more reliable side of the privilege boundary expose multiple functions across the privilege boundary to manage the metadata summary table and / or the metadata pointed to by the metadata summary table.
13. One of the multiple functions receives, as arguments, a physical page address and a set of metadata characteristics, and returns true if the page can be used for metadata storage and false otherwise, where the page is the same or a different unit of memory as the unit of memory used by the memory management unit, or creates a mapping between a data page and a metadata storage slot so that the confidential computing guarantee of the host architecture is met, or detaches a data page from an associated metadata page, The computing device according to claim 12.
14. One of the multiple functions handles the transition of a page from metadata storage use to data storage use, and the page is transitioned across the privilege boundary when the page is not used for metadata storage. The computing device according to claim 12.
15. A method comprising: using a memory management unit of a processor, receiving a memory access instruction including a virtual memory address, converting the virtual memory address to a physical memory address of physical memory, obtaining permission information and memory type information associated with the physical memory address, the permission information indicating whether permission is given to store metadata associated with data at the physical memory address, and the memory type information indicating whether the metadata is associated with the physical memory address, Determine whether the obtained permission information and memory type information indicate that it is permitted for the metadata to be associated with the physical memory address, By executing a query on the metadata summary table stored in the physical memory, determine whether the metadata is compatible with the physical memory address based on the physical memory address configured to store the metadata. The metadata summary table enables indirection by storing a pointer to the location within the physical memory where the metadata is stored. If it is determined that the metadata is incompatible, send a trap to the system software of the computing device. The trap triggers the dynamic allocation of the physical memory for storing the metadata associated with the physical memory address. Using a tag controller, the tag controller is part of the cache hierarchy of the computing device. The tag controller is configured to perform cache line eviction and / or cache fill for at least a part of the cache hierarchy based on the metadata. As part of cache line eviction, the tag controller is configured to determine, by checking the cache of the tag controller and / or the metadata summary table, whether the metadata storage is configured for the physical memory address from which the cache line is evicted. The tag controller checks whether the metadata of the cache line to be evicted is the same within a threshold number of consecutive chunks. If the metadata is the same within the threshold number of consecutive chunks, write the metadata of the cache line to be evicted to the metadata summary table. If the metadata is not the same within the threshold number of consecutive chunks, configure to write the metadata to the physical memory.