System memory reduction by oversubscribing to physical memory shared by computation entities supported by the system.
The logical pool memory system addresses memory inefficiencies in multi-tenant systems by oversubscribing to shared memory, enhancing utilization and reducing physical memory needs.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- MICROSOFT TECHNOLOGY LICENSING LLC
- Filing Date
- 2022-05-26
- Publication Date
- 2026-05-15
AI Technical Summary
Existing systems face inefficiencies in memory utilization, particularly in multi-tenant computing environments where allocated memory resources are often underutilized, leading to a need for additional physical memory allocation.
Implementing a logical pool memory system that allows oversubscription by combining local physical memory with logical pooled memory, where the logical pool memory is overcommitted, enabling efficient allocation and utilization of shared memory resources across multiple compute entities.
This approach reduces the overall physical memory requirements by allowing more compute entities to be supported without additional physical memory allocation, optimizing memory usage and reducing costs.
Smart Images

Figure 0007858994000004 
Figure 0007858994000005 
Figure 0007858994000006
Abstract
Description
Background Art
[0001] Multiple users or tenants may share a system that includes a computing system and a communication system. The computing system may include a public cloud, a private cloud, or a hybrid cloud having both a public portion and a private portion. The public cloud includes a global network of servers that perform various functions including storage and management of data, execution of applications, and delivery of content or services such as video streaming, email provisioning, office productivity software provisioning, or social media handling. While the public cloud provides services to the public over the Internet, a business may use a private cloud or a hybrid cloud. Both the private cloud and the hybrid cloud also include a network of servers housed in a data center.
[0002] Multiple tenants may use computing, storage, and networking resources associated with servers in the cloud. As an example, computing entities associated with different tenants may be allocated a certain amount of computing resources and memory resources. In many such situations, the resources allocated to various computing entities, including memory resources, may not be fully utilized. Similar underutilization of memory resources may occur in other systems such as a communication system including a base station.
Summary of the Invention
[0003] As an example, the present disclosure relates to a method for allocating a portion of memory associated with a system to compute entities, wherein the portion of memory includes a combination of a portion of a first type of first physical memory and a portion of logical pool memory used with a plurality of compute entities associated with the system. As used herein, the term “computed entity” includes, but is not limited to, any executable code (in the form of hardware, firmware, software, or any combination thereof) that implements a functionality, virtual machine, application, service, microservice, container, unikernel for serverless computing, or a portion of the foregoing. As used herein, the term “logical pool memory” refers to memory including overcommitted physical memory shared by a plurality of compute entities, which may correspond to a single host or a plurality of hosts. The logical pool memory may be mapped to a first type of second physical memory, and the amount of logical pool memory indicated as available for allocation to a plurality of compute entities may be greater than the amount of first type of second physical memory. The method may further include indicating to a logical pool memory controller associated with the logical pool memory that all pages associated with the logical pool memory to be initially allocated to any of the plurality of compute entities are known pattern pages. The method may further include the logical pool memory controller tracking both the status of whether a page of logical pool memory allocated to any of the multiple compute entities is a known pattern page, and the relationship between the physical memory address and the logical memory address associated with any allocated logical pool memory to any of the multiple compute entities. The method may further include, in response to a write operation initiated by a compute entity, the logical pool memory controller allowing the writing of data to any available space in a second physical memory of a first type, but only up to the extent of the physical memory corresponding to the portion of logical pool memory previously allocated to the compute entity.
[0004] In other examples, the disclosure relates to a system having memory used with a plurality of compute entities, wherein the portion of memory includes a combination of a portion of a first type of first physical memory and a portion of logical pool memory associated with the system. The logical pool memory may be mapped to a second type of physical memory, and the amount of logical pool memory indicated as available for allocation to the plurality of compute entities may be greater than the amount of the second type of physical memory. The system may further include a logical pool memory controller coupled to the logical pool memory, configured to (1) track both the state of whether a page of logical pool memory allocated to any of the plurality of compute entities is a known pattern page and the relationship between a physical memory address and a logical memory address associated with any allocated logical pool memory to any of the plurality of compute entities, and (2) in response to a write operation initiated by a compute entity to write any data other than a known pattern, enable the write operation to write data to any available space in the second type of physical memory, but only up to the extent of the physical memory corresponding to the portion of logical pool memory previously allocated to the compute entity.
[0005] In further other examples, the disclosure relates to a system comprising a plurality of host servers configured to run one or more of a plurality of compute entities. The system may further include memory, the memory portion comprising a combination of a portion of first physical memory of a first type and a portion of logical pool memory shared among the plurality of host servers. The logical pool memory may be mapped to second physical memory of a first type, and the amount of logical pool memory directed as available for allocation to the plurality of compute entities may be greater than the amount of second physical memory of a first type. The system may further include a logical pool memory controller, which is coupled to the logical pool memory associated with the system and is configured to (1) track both the state of whether a page of logical pool memory allocated to any of a plurality of compute entities is a known pattern page and the relationship between the physical memory address and the logical memory address associated with any allocated logical pool memory to any of the plurality of compute entities, and (2) in response to a write operation initiated by a compute entity, which is performed by a processor associated with any of the plurality of host servers to write any data other than known patterns, the write operation allows the data to be written to any available space in a second physical memory of a first type only up to the range of physical memory corresponding to the portion of logical pool memory previously allocated to the compute entity.
[0006] This summary is provided to introduce some of the concepts described below in a simplified form in the more detailed explanation. This summary is not intended to identify any important or essential features of the claims, nor is it intended to be used to limit the scope of the claims.
[0007] This disclosure is illustrative and not limited by the accompanying drawings. In the drawings, the same reference numerals indicate similar elements. Elements in the drawings are shown for brevity and clarity and are not necessarily to actual size. [Brief explanation of the drawing]
[0008] [Figure 1] This is a block diagram of a system environment including a host server coupled with a logical pool memory system, following one example. [Figure 2] The oversubscription process is illustrated by following an example. [Figure 3] This is a block diagram of a system including a host server with a logical pool memory system configured to allow oversubscription, a machine learning (ML) system, and a VM scheduler, following one example. [Figure 4] This is a block diagram of a logical pool memory system configured to allow memory oversubscription, following one example. [Figure 5] An example of a mapping table using a Translation Lookaside Buffer (TLB) is shown below. [Figure 6] The block diagram shows an example system including multiple host servers coupled with a logical pool memory system configured to allow memory oversubscription. [Figure 7] A block diagram is shown of an example system that implements at least part of the method for allowing oversubscription of pool memory allocated to compute entities. [Figure 8] This flowchart shows an example of how to oversubscribe to the physical memory used with computation entities. [Figure 9] A flowchart illustrating other examples of methods for oversubscribing to physical memory used with computation entities is shown. [Modes for carrying out the invention]
[0009] The examples described herein relate to reducing memory within a system by oversubscribing to physical memory shared among computing entities supported by the system. A specific example relates to oversubscribing to physical memory used with virtual machines in a multitenant computing system. A multitenant computing system may be a public cloud, a private cloud, or a hybrid cloud. A public cloud includes a global network of servers that perform a variety of functions, including data storage and management, application execution, and delivery of content or services such as streaming video, email, office productivity software, or social media. Servers and other components may be located in data centers around the world. While a public cloud provides services to the public over the internet, businesses may use a private cloud or a hybrid cloud. Both private and hybrid clouds also include a network of servers housed in data centers. Computing entities may run using the computing and memory resources of data centers. As used herein, the term “computing entity” includes, but is not limited to, any executable code (in the form of hardware, firmware, software, or any combination thereof) that implements a functionality, virtual machine, application, service, microservice, container, unikernel for serverless computing, or a portion thereof. Alternatively, the computing entities may run on hardware associated with edge computing devices, on-premises servers, or other types of systems, including communication systems such as base stations (5G or 6G base stations).
[0010] In accordance with the examples of this disclosure, a compute entity is allocated a combination of local physical memory and logical pooled memory under the assumption that its memory usage is typically less than or equal to the memory allocated to the compute entity. This allows compute entities to oversubscribe the installed memory in the system, so that more compute entities can be supported without the system needing to allocate additional physical memory. For example, in a multi-tenant computing or communications system, each tenant is allocated a portion of the total system memory. It is observed that not all tenants use their entire allocation of memory simultaneously. Thus, for example, a host server in a data center may be allocated logical pooled memory exposed by a pooled memory system. Each VM may be allocated a combination of local physical memory on the host server and a portion of the logical pooled memory available to the host server. The logical pooled memory may be a combination of known pattern pages and physical pooled memory that are not backed by any physical memory. In short, as used herein, the term “logical pooled memory” refers to memory containing overcommitted physical memory shared by multiple compute entities, which may correspond to a single host or multiple hosts. In this example, "overcommitted" means that the logical pool memory has less physical memory than the total amount of memory indicated as available to the compute entities. As used herein, the term "known pattern page" refers to any page containing only 0s, only 1s, or any other known pattern of values or symbols. A VM uses local physical memory first, and once that is used, the VM can access the portion of the logical pool memory. In a multitenant computing system, if all VMs are allocated memory in this manner, they are unlikely to exhaust the entire allocation of physical pool memory supporting the logical pool memory.This is because the actual physical memory behind the logical pool memory controller may be less than the logical pool memory exposed to the host server. This, in turn, can allow for a reduction in the amount of physical memory (e.g., DRAM) that needs to be allocated as part of a multi-tenant computing or communications system.
[0011] In the above implementation of logical pool memory, some portion of the logical pool memory remains unused, but it is difficult to predict which portion of the logical pool memory will remain unused at a given time. Therefore, a mechanism is needed to flexibly map pooled physical memory to the logical pool memory exposed to compute entities, so that the logical pool memory space in use is supported by the physical pool memory, without mapping the unused physical pool memory. In some cases, the mapping function may be performed using a page mapping table. Other mechanisms may also be used. Furthermore, in some cases, memory usage may be further optimized by compressing the data into physical pool memory before storing it. If data compression is used, the mapping function may be further modified to track the progress of the compressed memory pages.
[0012] Figure 1 is a block diagram of a system environment 100, including host servers (e.g., host servers 110, 120, and 130) coupled to a logical pool memory system 180, according to one example. Each host server may be configured to run several compute entities. In this example, host server 110 may be running compute entities 112 (e.g., CE-0, CE-1, and CE-2), host server 120 may be running compute entities 122 (e.g., CE-3, CE-4, and CE-5), and host server 130 may be running compute entities 132 (e.g., CE-6, CE-7, and CE-7). The logical pool memory system 180 may include a logical pool memory 190, which may include several memory modules. Although not shown in Figure 1, the logical pool memory system 180 may also include a logical pool memory controller (described later). As will be described in more detail later, the logical pool memory 190 may include an address space larger than that of the physical memory configured to support the logical address space. In this example, the logical pool memory 190 may include memory modules PM-1, PM-2, and PM-M, where M is an integer greater than 1. As used herein, the term “memory module” includes any memory based on memory technology that can operate as cacheable memory. Examples of such memory modules include, but are not limited to, dual inline memory modules (DIMMs) or single inline memory modules (SIMMs). As used herein, the term “cacheable memory” includes any type of memory in which data retrieved by a read operation associated with the memory can be copied to a memory cache so that the same data can be retrieved from the memory cache the next time it is accessed. The memory may be dynamic random access memory (DRAM), flash memory, static random access memory (SRAM), phase-change memory, magnetic random access memory, or any other type of memory technology that enables the memory to operate as cacheable memory.
[0013] Continuing with Figure 1, each of the host servers 110, 120, and 130 can be linked to memory modules contained in the logical pool memory 190. As an example, host server 110 is shown as being linked to memory module PM-1 via link 142, to memory module PM-2 via link 144, and to memory module PM-M via link 146. Host server 120 is shown as being linked to memory module PM-1 via link 152, to memory module PM-2 via link 154, and to memory module PM-M via link 156. As another example, host server 130 is shown as being linked to memory module PM-1 via link 162, to memory module PM-2 via link 164, and to memory module PM-M via link 166.
[0014] Referring still to Figure 1, in this example, the fabric manager 170 can manage the allocation of memory ranges to each host server. A processor (e.g., a CPU) that provides services to compute entities can issue load or store instructions. Load or store instructions can result in a read or write to local memory (e.g., DRAM associated with the CPU of a host server configured to run compute entities), or a read / write to a memory module associated with logical pool memory 190. In this example, load or store instructions resulting in a read / write to a memory module associated with logical pool memory 190 can be translated by the fabric manager 170 into a transaction that is completed using one of the links (e.g., links 142, 144, or 146). In one example, the fabric manager 170 may be implemented using a fabric manager based on the Compute Express Link standard.
[0015] Continuing with reference to Figure 1, compute entities 112, 122, and 132 may provide compute, storage, and networking resources using a set of clusters that may be included as part of a data center. As used in this disclosure, the term “data center” may include, but is not limited to, some or all of a data center owned by a cloud service provider, some or all of the data owned and operated by a cloud service provider, some or all of the data centers owned by a cloud service provider operated by a customer of the service provider, any other combination of data centers, a single data center, or several clusters within a particular data center. For example, each cluster may include several identical servers. Thus, a cluster may include servers with a certain number of CPU cores and a certain amount of memory. Alternatively, compute entities 112, 122, and 132 may run on hardware associated with edge computing devices, on-premises servers, or other types of systems, including communication systems such as base stations (e.g., 5G or 6G base stations). Figure 1 shows the system environment 100 as having a certain number of components, including compute entities and memory components arranged in a particular manner, but the system environment 100 may include additional or fewer components arranged in a different manner.
[0016] Figure 2 shows an example of an oversubscription process flow. This example represents an oversubscription process that uses a virtual machine (VM) as an exemplary compute entity that runs using hardware associated with a host server. Furthermore, this example assumes that the host server is located in a data center containing a cluster of host servers. This example further assumes that the cluster of host servers is part of a multi-tenant cloud. However, the oversubscription process may be used in other settings, including when the compute entity runs on hardware associated with edge computing devices, on-premises servers, or other types of systems, including communication systems such as base stations (e.g., 5G or 6G base stations). In response to a user request to initialize a VM, the VM scheduler may schedule the VM for execution using hardware resources associated with the host server. As part of this process, the VM may be allocated specific compute resources (e.g., CPU) and memory (e.g., DRAM associated with the CPU). Memory associated with the host server may be better utilized by using oversubscription alone or in combination with pooling. For example, memory can be oversubscribed in most cases due to the expectation that a VM will only use the memory that is normally used by such a VM or that is expected to be used by such a VM. In one example, the average of a VM's peak usage may be allocated to the VM as local physical memory. Thus, a certain amount of memory may be allocated to a VM, while being aware that such memory may not always be available to such a VM, for example, during peak usage or other high-memory-usage scenarios. Pooling can help with oversubscription, and oversubscription can help with pooling when the two techniques are used together, because memory from the pool can be used by VMs to which oversubscribed memory has been allocated.
[0017] Referring still to Figure 2, an exemplary oversubscription process is shown for a VM's request for 32GB of memory (210). In this example, in response to this request, the hypervisor (or other coordinator for allocating resources to the VM) may allocate 19GB of physical memory and 13GB of memory mapped to logical pool memory. In this example, the allocation of 19GB of memory may be based on the expected memory usage by the requesting VM. In this example, the pooled memory may have 256GB of logical pool memory along with 128GB of physical pool memory. Thus, the 13GB of memory mapped to the pool may or may not be available, as the 256GB of logical pool memory is shared with other host servers supporting additional VMs. A VM allocated 19GB of physical memory and 13GB of memory mapped to pooled memory can run without making any requests to the pooled memory, as in most cases the memory actually used by the VM is less than 19GB (220). Alternatively, in rare cases, a VM's memory usage may exceed 19GB (230). In such cases, additional memory may be allocated from physical pool memory 242 within logical pool memory 240. As part of the oversubscription process, a high-water mark may be associated with the use of physical pool memory 242, thereby initiating certain mitigation processes when the VM's use of this memory approaches the high-water mark. For example, upon reaching the high-water mark, the logical pool memory controller may begin compressing pages before storing them in physical pool memory 242. In another example, physical pool memory 242 may be backed up by additional physical memory that may have higher latency and is therefore likely to be cheaper to allocate than the memory associated with physical pool memory 242. In such a situation, upon reaching the high-water mark, pages may be written to the slower memory.
[0018] Software memory may be used to enable a compute entity to distinguish between physical and logical pool memory, ensuring that it uses the portion of physical pool memory (e.g., 19GB) allocated to it before using the portion of logical pool memory (e.g., 13GB) allocated to it. For example, to ensure better performance on a host server supporting Non-Uniform Memory Access (NUMA) nodes, pooled memory may be exposed to a VM as a compute-less virtual NUMA node. This allows the VM to distinguish between local physical memory and logical pool memory, resulting in higher latency. Thus, the VM may initially continue using local physical memory, and only after that is exhausted can it begin using the pooled memory exposed to the VM as a compute-less virtual NUMA node. Other mechanisms may also be used to enable the VM to distinguish between local physical memory and logical pool memory. Figure 2 illustrates an oversubscription process with a total of 32GB and 256GB of logical pool memory requested by the VM, although the VM may request additional or smaller amounts of memory (e.g., somewhere between 4GB and 128GB). Similarly, the size of the logical pool memory may be smaller or larger depending on the number of compute entities serviced by the computing system. Furthermore, Figure 2 shows the amount of local physical memory as 19GB and the amount of logical pool memory as 13GB, but these amounts are... Computational entities It may vary between these two points. Furthermore, the oversubscription process does not require memory pooling between host servers (or other systems); it only requires multiple compute entities that share logical pool memory.
[0019] Figure 3 is a block diagram of a system 300 that includes a host server 310 with resources 330 configured to allow oversubscription, a machine learning (ML) system 304, and a VM scheduler 302, according to one example. The host server 310 may also include a hypervisor 320, compute resources (e.g., CPU-0 332 and CPU-1 338), local memory resources (e.g., local memory-0 340 and local memory-1 342), a fabric manager 350, and a logical pool memory system 360. The VM scheduler 302 may be configured to receive requests for VMs from any tenant associated with a multi-tenant cloud and to schedule the placement of VMs (e.g., VM312, VM314, VM316, and VM318 on the host server 310). In this example, each of the resources 330 may be implemented as a hardware component, such as using one or more SoCs, controllers, processing cores, memory modules, ASICs, FPGAs, CPLDs, PLSs, silicon IP, etc. Furthermore, in this example, the functionalities associated with the VM scheduler 302, the ML system 304, and the hypervisor 320, respectively, may be performed by instructions stored in a non-temporary medium when executed by the processor. Figure 3 describes system 300 with respect to the VM, but other computing entities may be used by system 300 in a similar manner. Furthermore, the VM scheduler 302 may be configured to handle requests from computing entities included in other types of systems, such as edge computing devices, on-premises servers, or communication systems, such as base stations (e.g., 5G or 6G base stations).
[0020] The ML system 304 can be configured to provide resource predictions, including predicted memory usage, to the VM scheduler 302. Requests for placing sets of VMs can reach the VM scheduler 302. In this example, the VM scheduler 302 can then send a query to the ML system 304 to request the predicted memory usage. The ML system 304 can provide memory usage predictions. Using these predictions, the VM scheduler 302 can allocate memory resources to the VMs. Alternatively, the VM scheduler 302 may be provided with memory usage predictions or other memory usage information so that it does not need to query the ML system 304 every time a VM needs to be installed or scheduled for execution. As previously explained in connection with FIG. 2, the allocated memory can be a combination of a mapping to physical memory (e.g., local DRAM) and logical pooled memory.
[0021] Continuing to refer to Figure 3, the ML system 304 may include several components having instructions configured to perform various functions related to the ML system 304. The ML system 304 may include both an offline training component and an online prediction component. The offline training component may be involved in training various machine learning models, validating the models, and publishing the validated models. The online prediction component may generate predictions relating to various aspects, including predictions about memory usage by the VM and arbitrary behavioral patterns related to the VM. Before training the ML model, features may be selected that enable the ML model to predict metrics based on inputs. The training phase may include the use of backward propagation or other techniques that enable the ML to learn the relationship between specific input parameters and specific predictions based on those parameters. For example, a neural network model trained using stochastic gradient descent may be used to predict memory usage for computational entities such as VMs. Other types of ML models, including Bayesian models, may be used.
[0022] Generally, a supervised learning algorithm that can be trained based on input data can be implemented, and once it is trained, predictions or prescriptions can be made based on the training. Learning and inference techniques such as linear regression, support vector machines (SVMs) configured for regression, random forests configured for regression, gradient boosting trees configured for regression, and neural networks can be used. Linear regression may include modeling the past relationship between independent variables and a dependent output variable. A neural network may include artificial neurons used to generate an input layer, one or more hidden layers, and an output layer. Each layer may be encoded as a matrix or vector of weights represented in the form of coefficients or constants obtained by offline training of the neural network. The neural network may be implemented as a regression neural network (RNN), a long short-term memory (LSTM) neural network, or a gated recurrent unit (GRU). All information required by a model based on supervised learning can be converted into a vector representation corresponding to any of these techniques.
[0023] Taking LSTM as an example, an LSTM network may have a succession of repeating RNN layers or other types of layers. Each layer of the LSTM network may consume an input at a given time step, e.g., the state of the layer from the previous time step, and generate a new set of outputs or states. When using LSTM, a single chunk of content may be encoded into a single vector or multiple vectors. As an example, a word or combination of words (e.g., phrase, sentence, or paragraph) may be encoded as a single vector. Each chunk may be encoded into an individual layer (e.g., a specific time step) of the LSTM network. The LSTM layer may be described using a set of equations such as:
Number
[0024] In this example, within each LSTM layer, the input and hidden states may be processed, as needed, by a combination of vector operations (e.g., dot product, inner product, or vector addition) or nonlinear operations. Instructions corresponding to the machine learning system may be encoded as hardware corresponding to the A / I processor. In this case, some or all of the functionality related to the ML system 304 may be hardcoded or supplied separately as part of the A / I processor. For example, the A / I processor may be implemented using an FPGA with the necessary functionality.
[0025] In other examples, the memory allocated to a VM may be dynamically changed or determined. For example, behavioral changes may correspond to the workload handled by the VM over time, which can be used to dynamically change memory allocation. An example of behavioral changes may be related to a VM serving a web search workload or other similar workload. Such a VM may have a diurnal pattern in which the workload requires more resources during the day but fewer resources at night. In one example, the Fast Fourier Transform (FFT) algorithm can be used to detect periodicity in the behavior associated with the VM. While the FFT can be used to detect periodicity on multiple time scales, the VM's memory allocation may only be dynamically changed if the workload handled by the VM has a periodicity that matches a particular pattern, such as a diurnal pattern. Periodicity analysis can generate ground truth labels that can also be used to train ML models to predict which VMs are likely to exhibit a diurnal pattern or another pattern. In other examples, other features (e.g., cloud subscription ID, the user who created the VM, VM type, VM size, guest operating system) may also be used to statically or dynamically change the range of physical and pooled memory allocated to the VM.
[0026] In yet another example, instead of predicting memory usage or determining behavior / usage patterns, a VM may exclusively allocate local physical memory in response to a request to schedule it. At a later point, when the local physical memory associated with a host server (or collection of host servers) begins to reach a certain utilization level, the scheduled VM may be allocated a combination of local physical memory and mapped logical pool memory. In this way, memory can be dynamically allocated to the VM currently running.
[0027] Still referring to Figure 3, hypervisor 320 may manage virtual machines including VMs 312, 314, 316, and 318. CPU-0 332 may contain one or more processing cores, and CPU-1 338 may contain one or more processing cores. Instructions corresponding to hypervisor 320 may provide management functions provided by hypervisor 320 when executed by any of the processing cores. Hypervisor 320 may ensure that any threads associated with any given VM are scheduled only for the logical cores of that group. CPU-0 332 may be coupled to local memory-0 340, and CPU-1 338 may be coupled to another local memory-1 342. Local memory-0 340 and local memory-1 342 each may contain one or more memory modules. Examples of such memory modules include, but are not limited to, dual inline memory modules (DIMMs) or single inline memory modules (SIMMs). Any type of memory that is “cacheable memory” may be used to implement a memory module. As used herein, the term “cacheable memory” includes any type of memory in which data retrieved by a read operation associated with the memory can be copied to a memory cache so that the same data can be retrieved from the memory cache the next time it is accessed. The memory may be dynamic random access memory (DRAM), flash memory, static random access memory (SRAM), phase-change memory, magnetic random access memory, or any other type of memory technology that enables the memory to operate as cacheable memory.
[0028] Continuing to refer to Figure 3, in this example, CPU-0 332 may be coupled to Fabric Manager 350 via Fabric Interface-0 334, and CPU-1 338 may be coupled to Fabric Manager 350 via another Fabric Interface-1 336. Fabric Manager 350 may enable any VM running on processing cores associated with CPU-0 332 and CPU-1 338 to access the logical pool memory system 360. In one example, any CPU associated with Host Server 310 may issue a load or store instruction to the logical address space set up by Hypervisor 320. A load or store instruction may result in a read or write to local memory (e.g., either local memory-0 340 or local memory-1 342, depending on the CPU issuing the load or store instruction), or a read / write transaction associated with the logical pool memory system 360. In one example, if local memory is DRAM, then one of the relevant Double Data Rate (DDR) protocols associated with accessing DRAM may be used by the CPU to read from / write to local memory. In this example, load or store instructions resulting in read / write operations to memory modules associated with the logical pool memory system 360 may be translated by the fabric manager 350 (or other similar subsystem) into transactions that are completed using one of the links connecting the memory and VM associated with the logical pool memory system 360. Read / write transactions may be executed using the fabric manager 350 and protocols associated with the logical pool memory system 360. If the memory modules associated with the logical pool memory system 360 are implemented as DRAM modules, they may be accessed by the controller associated with the logical pool memory system 360 using one of the relevant DDR protocols.
[0029] In one example, the fabric manager 350 may be implemented using a fabric manager based on the Compute Express Link (CXL) standard. In this example, the memory modules associated with the logical pool memory system 360 may be configured as Type 3 CXL devices. The logical address space exposed to the virtual machine by the hypervisor 320 may be at least a subset of the address range exposed by the controller associated with the CXL bus / link. Thus, as part of this example, transactions associated with the CXL.io protocol, a PCIe-based non-coherent I / O protocol, may be used to configure memory devices and the link between the memory modules in the logical pool memory system 360 and the CPU. The CXL.io protocol may also be used by the CPU associated with the host server 310 for device discovery, enumeration, error reporting, and management. Alternatively, any other I / O protocol that supports such configuration transactions may be used. Memory access to the memory modules may be handled via transactions associated with the CXL.mem protocol, a memory access protocol that supports memory transactions. For example, load and store instructions associated with any of the CPUs of the host server 310 may be processed via the CXL.mem protocol. Alternatively, any other protocol that enables the translation of CPU load / store instructions into read / write transactions associated with memory modules contained in the logical pool memory system 360 may be used. Figure 3 shows a host server 310 with certain components arranged in a particular way, but the host server 310 may include additional or fewer components arranged in a different way. For example, Figure 3 shows local memory coupled to the CPU, but non-local memory may be coupled to the CPU. Alternatively, the CPU may only access memory associated with the logical pool memory system 360.Furthermore, although Figure 3 shows VMs hosted by a single host server, VMs 312, 314, 316, and 318 may be hosted using a group of host servers (for example, a cluster of host servers in a data center). The logical pool memory system 360 may be shared among the host servers or may be dedicated to each host server.
[0030] Figure 4 is a block diagram of a logical pool memory system 400 configured to allow memory oversubscription according to one example. In one example, the logical pool memory system 360 of Figure 3 may be implemented as the logical pool memory system 400. In this example, the logical pool memory system 400 may include a logical pool memory 410, a logical pool memory controller 450, and a pool memory control structure 480. The logical pool memory 410 may include a physical pool memory 412 and a known pattern page 414. The physical pool memory 412 may include a memory module with physical memory (e.g., DRAM). The physical pool memory 412 may correspond to a logical address space. The known pattern page 414 may represent a logical address space not supported by any physical memory. The logical pool memory controller 450 may manage access to the logical pool memory 410 by interacting with the pool memory data structure 480. The logical pool memory controller 450 may be implemented as any combination of hardware, firmware, or software instructions. Instructions corresponding to the logical pool memory controller 450 may be stored in memory associated with the logical pool memory system 400. Such instructions, when executed by a processor associated with the logical pool memory system 400, may provide at least some of the functions associated with the logical pool memory controller 450. Other functions may be provided by control logic associated with the logical pool memory controller 450, including finite state machines and other logic.
[0031] In this example, the pool memory control structure 480 may include a mapping table 482 and a physical memory free page list 484. Both the mapping table 482 and the physical memory free page list 484 may be stored in the physical pool memory 412 or other memory (e.g., a cache associated with the logical pool memory controller 450). The mapping table 482 may maintain mappings between the physical address space and logical address space associated with the physical pool memory 412. The mapping table 482 may also track the state of a page regarding whether or not it is a known pattern page. A known pattern page does not need to have any allocated physical space as part of the physical pool memory 412. In this example, physical memory addresses associated with a known pattern page are considered invalid in the mapping table 482 (identified as Not Applicable (N / A)). If additional space is freed up in the physical pool memory 412 using compression, this can also be tracked using the mapping table 482. Table 1 below shows an example mapping table 482. [Table 1]
[0032] As shown above in Table 1, the mapping table 482 may be indexed by page number. An indicator of a known pattern page (e.g., a flag for a known pattern page) may be associated with each page. A logical high value (e.g., 1) may indicate to the logical pool memory controller 450 that no physical memory space is allocated in the physical pool memory 412 corresponding to the page number. A logical low value (e.g., 0) may indicate to the logical pool memory controller 450 that physical memory space is allocated in the physical pool memory 412 corresponding to the page number. The logical pool memory controller 450 may change the indicator of a known pattern page based on any status change related to the page number. At the time of VM scheduling or creation, the hypervisor (or host operating system) may map a portion of the memory for the VM to the physical memory associated with the host server and the remainder to the logical pool memory 410. At this point, the hypervisor (or operating system) may command the logical pool memory controller 450 to set the indicator of a known pattern page to a logical high value (e.g., 1) for pages associated with the logical pool memory allocated to the VM. The VM may also write known patterns to the logical pool memory allocated to it when it is initialized. The logical pool memory controller 450 may detect a write operation that writes known patterns as an indication that physical memory does not need to be allocated to the VM at this point. Furthermore, the logical pool memory controller 450 may ensure that the indicators for known pattern pages for the corresponding pages are accurately reflected in the mapping table 482.
[0033] Initially, the logical pool memory controller 450 may map all of the logical pool memory for the VM to known pattern pages 414. A newly scheduled VM may also work with the hypervisor to ensure that any pages corresponding to the logical pool memory 410 allocated to the VM are identified as known pattern pages in the mapping table 482. Later, when the VM has exhausted all of the local physical memory allocated to it and begins writing unknown patterns (e.g., non-zero) to the logical memory address space, the logical pool memory controller 450 may begin allocating pages from within the logical address space (exposed by the logical pool memory controller 450) corresponding to the physical pool memory 412 to the VM. The physical memory free page list 484 may contain a list of pages corresponding to free physical pool memory 412 that are not mapped by the mapping table 482. Table 2 below shows an example physical memory free page list 484. This example lists pages corresponding to freely allocable physical pool memory 412. [Table 2]
[0034] Although address translation is described as single-level address translation, it may also be multi-level address translation, and therefore the use of additional mapping tables may be required. In any case, the logical pool memory controller 450 may be configured to maintain mapping tables and ensure proper tracking of known pattern pages, enabling pooling and oversubscription of memory allocated to VMs, as described above. Furthermore, although mapping table 482 is described in relation to pages, other groupings of memory, including memory blocks, may also be used with mapping table 482. Also, the functionality of mapping table 482 may be implemented using registers associated with the processor and controller or other similar hardware structures. Furthermore, although mapping table 482 is described as being stored in physical pool memory 412, at least a portion of the mapping table may be stored in a cache associated with the logical pool memory controller 450 to accelerate the address translation process.
[0035] Figure 5 shows an example of a mapping table using a Translation Lookaside Buffer (TLB). As previously described in relation to Figure 4, the mapping table 482 may be used by the logical pool memory controller 450 to translate logical addresses received from the hypervisor (or other source) to physical addresses associated with the physical pool memory 412. In this example, the logical pool memory controller 520 may include a cache implemented as a Translation Lookaside Buffer (TLB) 530. The TLB 530 may be implemented as an associative memory (CAM), a fully associative cache, a two-way set associative cache, a direct-mapped cache, or any other suitable type of cache. The TLB 530 may be implemented to achieve both the translation from logical pool addresses to physical pool memory addresses and the status check of known pattern pages. For each address translation request, the logical pool memory controller 520 may first check the TLB 530 to determine whether the address translation has already been cached. A TLB HIT signal may be generated if it is known that address translation is cached in TML530. Furthermore, the physical address corresponding to the page number that generated the TLB HIT signal may be supplied to the physical pool memory 560 to enable access to the data or instruction stored at that physical address. Alternatively, if address translation is not cached, a TLB MISS signal may be generated. This may then result in a table walk across the mapping table 562 unless the flags for known pattern pages are set to logical true. As shown in Figure 5, a physical memory free page list 564 may also be stored in the physical pool memory 560.
[0036] Figure 6 shows a block diagram of system 600, which includes several host servers coupled to a logical pool memory system 670 configured to allow memory oversubscription. System 600 may include host servers 610, 620, and 630 coupled to the logical pool memory system 670. As previously described, each host server may host multiple compute entities (e.g., VMs), which may be purchased or otherwise paid for by different tenants of system 600. Each server may include local memory (e.g., memory modules described in relation to host server 310 in Figure 3) and may also have access to logical pool memory, which may be a portion of the logical pool memory contained in the logical pool memory system 670. As previously described, the fabric manager (implemented based on CXL) may expose a subset of the logical pool memory associated with the logical pool memory system 670 to hosts 610, 620, and 630, respectively. The logical pool memory system 670 may be implemented as previously described.
[0037] Each compute entity scheduled or otherwise made available to tenants via host server 610 may be allocated a portion of local memory 612 and a portion of pooled memory 614. As previously described, the allocation to each compute entity may be based on oversubscription to logical pool memory managed by logical pool memory system 670. In this example, each compute entity hosted by host server 610 may be connected to logical pool memory system 670 via link 616 (or a set of links). Each compute entity scheduled or otherwise made available to tenants via host server 620 may be allocated a portion of local memory 622 and a portion of pooled memory 624. As previously described, the allocation to each compute entity hosted by host server 620 may be based on oversubscription to logical pool memory managed by logical pool memory system 670. In this example, each compute entity hosted by host server 620 may be connected to logical pool memory system 670 via link 626 (or a set of links). Each compute entity scheduled or otherwise made available to tenants via host server 630 may be allocated a portion of local memory 632 and a portion of pooled memory 634. As previously described, allocation to each compute entity hosted by host server 630 may be based on oversubscription to logical pool memory managed by logical pool memory system 670. In this example, each compute entity hosted by host server 630 may be coupled to logical pool memory system 670 via link 636 (or a set of links). For example, load and store instructions associated with any of the CPUs of host servers 610, 620, and 630 may be handled by the CXL.mem protocol.Alternatively, any other protocol that enables the translation of CPU load / store instructions into read / write transactions associated with memory modules contained in the logical pool memory system 670 may be used. Figure 6 shows system 600 as containing certain components arranged in a particular way, but system 600 may contain additional or fewer components arranged in a different way.
[0038] Figure 7 shows a block diagram of system 700 that implements at least part of a method for allowing oversubscription of pooled memory allocated to computation entities. System 700 may include a processor 702, an I / O component 704, memory 706, a presentation component 708, a sensor 710, a database 712, a networking interface 714, and an I / O port 716, which may be interconnected via a bus 720. The processor 702 may execute instructions stored in memory 706. The I / O component 704 may include components such as a keyboard, mouse, speech recognition processor, or touchscreen. The memory 706 may be any combination of non-volatile storage and volatile storage (e.g., flash memory, DRAM, SRAM, or other types of memory). The presentation component 708 may include a display, a holographic device, or other presentation device. The display may be any type of display, such as an LCD, LED, or other type of display. Sensor 710 may include telemetry or other types of sensors configured to detect and / or receive information (e.g., collected data). Sensor 710 may include telemetry or other types of sensors configured to detect and / or receive information (e.g., memory usage by various computing entities running on host servers within a data center). Sensor 710 may include sensors configured to detect conditions related to the CPU, memory, or other storage components, FPGAs, motherboards, board management controllers, etc. Sensor 710 may also include sensors configured to detect conditions related to racks, chassis, fans, power supply units (PSUs), etc. Sensor 710 may include sensors configured to detect conditions related to network interface controllers (NICs), top-of-rack (TOR) switches, middle-of-rack (MOR) switches, routers, power distribution units (PDUs), rack-level uninterruptible power supply (UPS) systems, etc.
[0039] Still referring to Figure 7, the database 712 may be used to store any collected or recorded data necessary for performing the methods described herein. The database 712 may be implemented as a collection of distributed databases or as a single database. The network interface 714 may include communication interfaces such as Ethernet®, cellular radio, Bluetooth® radio, UWB radio, or other types of wireless or wired communication interfaces. The I / O port 716 may include Ethernet ports, fiber optic ports, wireless ports, or other communication or diagnostic ports. Although Figure 7 shows the system 700 as including a certain number of components arranged and combined in a particular way, it may include fewer or additional components arranged and combined in a different way. Furthermore, the functionality associated with the system 700 may be distributed as needed.
[0040] Figure 8 shows a flowchart 800 of an example of a method for oversubscribing physical memory used with compute entities. In one example, the steps associated with this method may be performed by various components of the system described earlier. Step 810 may include allocating a portion of memory associated with the system to compute entities, the portion of memory having a combination of a portion of first type first physical memory and a portion of logical pool memory associated with the system. The amount of logical pool memory that is mapped to second type first physical memory and directed to be available for allocation to multiple compute entities is greater than the amount of second type first physical memory. As described earlier with respect to Figure 2, a virtual machine, which is a type of compute entity, may be allocated 19 GB of physical memory (e.g., a memory module containing DRAM configured as local memory for the CPU) and 13 GB of logical pool memory that may be a portion of logical pool memory 240 in Figure 2, which may also be a memory module containing DRAM in Figure 2. As previously stated, the term “cacheable memory” includes any type of memory in which data retrieved by a memory-related read operation can be copied to a memory cache so that it can be retrieved from the memory cache the next time the same data is accessed. The memory may be dynamic random access memory (DRAM), flash memory, static random access memory (SRAM), phase-change memory, magnetic random access memory, or any other type of memory technology that enables the memory to operate as cacheable memory. In short, the first physical memory of the first type and the second physical memory of the first type must be of a cacheable type, but they may be of different types in other respects (e.g., DRAM versus phase-change memory, or SRAM versus DRAM).
[0041] Step 820 may include indicating to the logical pool memory controller associated with the logical pool memory that all pages associated with the logical pool memory that are initially allocated to any of the multiple compute entities are known pattern pages. For example, a hypervisor (e.g., hypervisor 320 in Figure 3) may indicate to the logical pool memory controller (e.g., logical pool memory controller 450 in Figure 4) that the status of the initially allocated pages is a known pattern page (e.g., known pattern page 414).
[0042] Step 830 may include the logical pool memory controller tracking both the status of a page of logical pool memory allocated to one of several compute entities—whether it is a known pattern page—and the relationship between the physical memory address and the logical memory address associated with any allocated logical pool memory to one of several compute entities. In this example, as part of this step, as previously described in relation to Figures 4 and 5, the logical pool memory controller (e.g., logical pool memory controller 450 in Figure 4) may track the status of a page (whether it is a known pattern page or not) and the relationship between the physical address and page number in physical memory (e.g., physical pool memory 412 in Figure 4) using a mapping table (e.g., mapping table 482 in Figure 4). As previously described, to accelerate access to logical pool memory, at least a portion of the mapping table may be implemented as a translation lookaside buffer. Furthermore, the mapping table (or similar structure) may be implemented using registers.
[0043] Step 840 may include, in response to a write operation initiated by a compute entity, allowing the logical pool memory controller to write data to any available space in the second physical memory of the first type only up to the range of physical memory corresponding to the portion of logical pool memory previously allocated to the compute entity. As previously described, when a compute entity (e.g., a VM) runs out of all the local physical memory allocated to it and begins writing an unknown pattern to the logical memory address space, the hypervisor may begin allocating pages to the compute entity from within the logical address space (exposed by the logical pool memory controller 450) corresponding to the physical memory (e.g., physical pool memory 412 in Figure 4). The logical pool memory controller may access the free page list (e.g., physical memory free page list 484 in Figure 4) to access free pages that are not mapped by the mapping table. Furthermore, as previously described, a compute entity may not be allocated more physical memory managed by the logical pool memory controller than previously allocated to the compute entity (e.g., 13GB in the example described with respect to Figure 2). This method allows computation entities to support more computing entities without requiring the system to allocate additional physical memory (e.g., DRAM), thus enabling oversubscription of installed memory within the system by computing entities.
[0044] Figure 9 shows a flowchart 900 of an example of another method for oversubscribing to physical memory used with compute entities. In one example, the steps associated with this method may be performed by various components of the system described earlier. Step 910 may include allocating a portion of memory associated with the system to compute entities, the portion of memory having a combination of a portion of first physical memory of a first type and a portion of logical pool memory associated with the system. The logical pool memory may be mapped to a second physical pool memory of a first type, and the amount of logical pool memory indicated as available for allocation to multiple compute entities may be greater than the amount of second physical pool memory of a first type. As described earlier with respect to Figure 2, a virtual machine, which is a type of compute entity, may be allocated 19 GB of physical memory (e.g., a memory module containing DRAM configured as local memory for the CPU) and 13 GB of logical pool memory, which may be a portion of logical pool memory 240 in Figure 2, which may also be a memory module containing DRAM in the physical pool memory 242 in Figure 2. For example, the logical pool memory may be the logical pool memory included as part of the logical pool memory system 670 in Figure 6, and the physical pool memory may correspond to the physical memory (e.g., DRAM) included as part of the logical pool memory system 670 in Figure 6.
[0045] Step 920 may include translating any load or store instruction directed to logical pool memory into a memory transaction for completion via the respective links between each of the multiple compute entities and each of the physical memory devices included as a portion of the second physical pool memory of the first type. In one example, a fabric manager implemented using a fabric manager based on the Compute Express Link (CXL) specification (e.g., fabric manager 350 in Figure 3 or fabric manager 650 in Figure 6), as previously described, may translate load and store instructions into transactions related to the CXL specification. The logical address space exposed to the compute entities may be at least a subset of the address range exposed by the controller associated with the CXL bus / link.
[0046] Step 930 may include indicating to the logical pool memory controller associated with the logical pool memory that all pages associated with the logical pool memory that are initially allocated to any of the multiple compute entities are known pattern pages. For example, a hypervisor (e.g., hypervisor 320 in Figure 3, or a hypervisor associated with any of the host servers 610, 620, and 630 in Figure 6) may indicate to the logical pool memory controller (e.g., logical pool memory controller 450 in Figure 4) that the status of the initially allocated pages is a known pattern page (e.g., known pattern page 414). A combination of hypervisors may also cooperate to manage memory allocation to compute entities supported by a computing system or communication system having multiple hypervisors managing shared logical pool memory.
[0047] Step 940 may include the logical pool memory controller tracking both the status of a page of logical pool memory allocated to one of several compute entities—whether it is a known pattern page—and the relationship between the physical memory address and the logical memory address associated with any allocated logical pool memory to one of several compute entities. In this example, as part of this step, as previously described in relation to Figures 4-6, the logical pool memory controller (e.g., logical pool memory controller 450 in Figure 4) may track the status of a page (whether it is a known pattern page or not) and the relationship between the physical address and page number in physical memory (e.g., physical pool memory 412 in Figure 4) using a mapping table (e.g., mapping table 482 in Figure 4). As previously described, to accelerate access to logical pool memory, at least a portion of the mapping table may be implemented as a translation lookaside buffer. Furthermore, the mapping table (or similar structure) may be implemented using registers.
[0048] Step 950 may include, in response to a write operation initiated by a compute entity to write arbitrary data other than known patterns, allowing the logical pool memory controller to write data to any available space in the second physical pool memory of the first type, but only up to the extent of the physical memory corresponding to the portion of logical pool memory previously allocated to the compute entity. As previously described, when a compute entity (e.g., a VM) has exhausted all of the local physical memory allocated to it and has begun writing an unknown pattern into the logical memory address space, the hypervisor may begin allocating pages to the compute entity from within the logical address space (exposed by the logical pool memory controller 450) corresponding to the physical memory (e.g., physical pool memory 412 in Figure 4, or the physical pool memory associated with the logical pool memory system 670 in Figure 6). The logical pool memory controller may access the free page list (e.g., physical memory free page list 484 in Figure 4) to access free pages that are not mapped by the mapping table. Furthermore, as explained earlier, a compute entity may not be allocated more physical memory managed by the logical pool memory controller than previously allocated to it (for example, 13GB in the example shown in Figure 2). This also allows compute entities to support more compute entities without requiring the system to allocate additional physical memory (e.g., DRAM), thus enabling oversubscription of installed memory within the system by compute entities.
[0049] In summary, the disclosure relates to a method for allocating a portion of memory associated with a system to compute entities, wherein the portion of memory comprises a combination of a portion of a first type of first physical memory and a portion of logical pool memory used with a plurality of compute entities associated with the system. The logical pool memory may be mapped to a second type of physical memory, and the amount of logical pool memory indicated as available for allocation to a plurality of compute entities may be greater than the amount of the second type of physical memory. The method may further include indicating to a logical pool memory controller associated with the logical pool memory that all pages associated with the logical pool memory to be first allocated to any of the plurality of compute entities are known pattern pages. The method may further include the logical pool memory controller tracking both the status of whether pages of logical pool memory allocated to any of the plurality of compute entities are known pattern pages and the relationship between physical memory addresses and logical memory addresses associated with any allocated logical pool memory to any of the plurality of compute entities. The method may further include, in response to a write operation initiated by a compute entity, allowing the logical pool memory controller to write data to any available space in the second physical memory of the first type only up to the extent of the physical memory corresponding to the portion of logical pool memory previously allocated to the compute entity.
[0050] In this way, the amount of logical pool memory associated with the system may be equal to the amount of the first type of second physical memory, combined with the amount of memory corresponding to known pattern pages that are not supported by any type of physical memory. Tracking both the state of whether a page of logical pool memory allocated to any of several compute entities is a known pattern page, and the relationship between physical memory addresses and logical memory addresses, may involve maintaining a mapping table. In one example, the mapping table may be implemented using a translation lookaside buffer.
[0051] The method may further include allocating a portion of the first type of first physical memory based on the expected use of the first type of first physical memory by the compute entity. The method may further include dynamically changing the amount of the portion of the first type of first physical memory allocated to the compute entity based on the usage pattern associated with the use of the first type of first physical memory allocated to the compute entity.
[0052] The method may further include exposing the portion of logical pool memory allocated to a compute entity via a software mechanism so that the compute entity can distinguish between the portion of logical pool memory allocated to the compute entity and the portion of first type first physical memory allocated to the compute entity. The logical pool memory may be coupled to a processor running one of several compute entities to which at least a portion of the logical pool memory associated with the system is allocated, via each link managed by the fabric manager.
[0053] In other examples, the disclosure relates to a system having memory used with a plurality of compute entities, wherein the portion of memory has a combination of a portion of a first type of first physical memory and a portion of logical pool memory associated with the system. The logical pool memory may be mapped to a second type of second physical memory, and the amount of logical pool memory indicated as available for allocation to the plurality of compute entities may be greater than the amount of the first type of second physical memory. The system may further include a logical pool memory controller coupled to the logical pool memory, configured to (1) track both the state of whether a page of logical pool memory allocated to any of the plurality of compute entities is a known pattern page and the relationship between a physical memory address and a logical memory address associated with any allocated logical pool memory to any of the plurality of compute entities, and (2) in response to a write operation initiated by a compute entity to write any data other than a known pattern, enable the write operation to write data to any available space in the second type of second physical memory, but only up to the extent of the physical memory corresponding to the portion of logical pool memory previously allocated to the compute entity.
[0054] As part of the system, the amount of logical pool memory associated with the system may be equal to the amount of a first type of second physical memory, combined with the amount of memory corresponding to known pattern pages that are not supported by any type of physical memory. The logical pool memory controller may be further configured to maintain a mapping table to track both the state of whether a page of logical pool memory allocated to any of several compute entities is a known pattern page, and the relationship between the physical memory address and the logical memory address associated with any allocated logical pool memory to any of the several compute entities.
[0055] As part of the system, for example, a scheduler may be configured to allocate a portion of a first type of first physical memory based on the expected use of that first type of first physical memory by compute entities. The scheduler may also be configured to dynamically change the amount of a portion of a first type of first physical memory allocated to compute entities based on usage patterns associated with the use of the first type of first physical memory allocated to compute entities. A logical pool memory controller may be further configured to indicate to each of the compute entities that all pages associated with the logical pool memory initially allocated to any of the compute entities are known pattern pages.
[0056] In further other examples, the disclosure relates to a system comprising multiple host servers configured to run one or more of a plurality of compute entities. The system may further include memory, the memory portion having a combination of a portion of first physical memory of a first type and a portion of logical pool memory shared among the multiple host servers. The amount of logical pool memory mapped to second physical memory of a first type and directed as available for allocation to the plurality of compute entities may be greater than the amount of second physical memory of a first type. The system may further include a logical pool memory controller configured to be coupled to a logical pool memory associated with the system and to track both the status of whether a page of logical pool memory allocated to any of several compute entities is a known pattern page, and the relationship between a physical memory address and a logical memory address associated with any allocated logical pool memory to any of several compute entities, and (2) in response to a write operation initiated by a compute entity, which is performed by a processor associated with any of several host servers to write any data other than known patterns, the write operation allows the data to be written to any available space in a second physical memory of a first type, but only up to the extent of the physical memory corresponding to the portion of logical pool memory previously allocated to the compute entity.
[0057] As part of the system, the amount of logical pool memory associated with the system may be equal to the amount of a first type of second physical memory, combined with the amount of memory corresponding to known pattern pages that are not supported by any type of physical memory. The logical pool memory controller may be further configured to maintain a mapping table to track both the state of whether a page of logical pool memory allocated to any of several compute entities is a known pattern page, and the relationship between the physical memory address and the logical memory address associated with any allocated logical pool memory to any of the several compute entities.
[0058] As part of the system, for example, a scheduler may be configured to allocate a portion of a first type of first physical memory based on the expected use of that first type of first physical memory by compute entities. The scheduler may also be configured to dynamically change the amount of a portion of a first type of first physical memory allocated to compute entities based on usage patterns associated with the use of the first type of first physical memory allocated to compute entities. A logical pool memory controller may be further configured to indicate to each of the compute entities that all pages associated with the logical pool memory initially allocated to any of the compute entities are known pattern pages.
[0059] It should be understood that the methods, modules, and components described herein are merely illustrative. Alternatively, or additionally, the functions described herein can be performed, at least partially, by one or more hardware logic components. For example, exemplary types of hardware logic components that can be used include, without limitation, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standards (ASSPs), system-on-chip systems (SOCs), composite programmable logic devices (CPLDs), and so on. In an abstract but clear sense, any arrangement of components to achieve the same function is effectively “associated” in such a way that the desired function is achieved. Thus, any two components combined herein to achieve a particular function can be considered “associated” with each other, regardless of architecture or intermediate components, in such a way that the desired function is achieved. Similarly, any two components thus associated can be considered “operably connected” or “joined” with each other in such a way that the desired function is achieved. The fact that a component is described herein as being coupled to another component does not necessarily mean that the component is a separate component. For example, component A described as being coupled to another component B may be a subcomponent of component B, component B may be a subcomponent of component A, or components A and B may be coupled subcomponents of another component C.
[0060] Some of the functions related to the examples described herein may also include instructions stored on non-temporary media. As used herein, the term “non-temporary media” refers to any medium that stores data and / or instructions that cause a machine to operate in a particular way. Examples of non-temporary media include non-volatile media and / or volatile media. Non-volatile media include, for example, hard disks, solid-state drives, magnetic disks or tapes, optical disks or tapes, flash memory, EPROM, NVRAM, PRAM, or other such media, or networked versions thereof. Volatile media include, for example, dynamic memory such as DRAM, SRAM, cache, or other such media. Non-temporary media are different from, but can be used together with, transmission media. Transmission media are used to transfer data and / or instructions to and from a machine. Examples of transmission media include coaxial cables, fiber optic cables, copper wires, and wireless media such as radio waves.
[0061] Furthermore, those skilled in the art will recognize that the boundaries between the functions of the above operations are merely illustrative. The functions of multiple operations may be combined into a single operation, and / or the functions of a single operation may be distributed to additional operations. Moreover, alternative embodiments may include multiple instances of a particular operation, and the order of operations may be modified in various other embodiments.
[0062] While this disclosure provides specific examples, various modifications and changes may be made without exceeding the scope of this disclosure as set forth in the claims below. Accordingly, the specification and drawings should be considered illustrative rather than restrictive, and all such modifications are intended to be within the scope of this disclosure. No advantage, effect, or solution to a problem described herein with respect to the specific examples is intended to be construed as an essential, necessary, or required feature or element of any or all of the claims.
[0063] Furthermore, as used herein, the term “one” (a or an) is defined as one or more. Also, the use of introductory phrases such as “at least one” or “one or more” in the claims should not be interpreted as implying that the introduction of other claim elements by the indefinite article (a or an) limits any particular claim containing such introduced claim elements to inventions containing at least one such element, even if the same claim contains the introductory phrase “one or more” or “at least one” along with an indefinite article (e.g., a or an). The same applies to the use of the definite article.
[0064] Unless otherwise stated, terms such as “First” and “Second” are used to arbitrarily distinguish between the elements they describe. Therefore, these terms are not necessarily intended to indicate their temporal or other priority.
Claims
1. The allocation of a portion of memory associated with a system to compute entities, wherein the portion of memory includes a combination of a portion of a first type of first physical memory and a portion of logical pool memory used with a plurality of compute entities associated with the system, the logical pool memory being mapped to a second type of physical memory, and the amount of the logical pool memory indicated as available for allocation to the plurality of compute entities is greater than the amount of the second type of physical memory. To indicate to the logical pool memory controller associated with the logical pool memory that all pages associated with the logical pool memory that are initially allocated to any of the multiple compute entities are known pattern pages, The logical pool memory controller tracks both the status of whether a page of the logical pool memory assigned to any of the plurality of compute entities is a known pattern page, and the relationship between the physical memory address and the logical memory address associated with any assigned logical pool memory to any of the plurality of compute entities. In response to a write operation initiated by the compute entity, the logical pool memory controller permits writing data to any available space in the second physical memory of the first type, but only up to the extent of the physical memory corresponding to the portion of the logical pool memory previously allocated to the compute entity. A method of having.
2. The amount of the logical pool memory associated with the system is equal to the amount of the second physical memory of the first type, combined with the amount of memory corresponding to known pattern pages not supported by any type of physical memory. The method according to claim 1.
3. Tracking both the state of whether the page of the logical pool memory assigned to any of the plurality of compute entities is a known pattern page, and the relationship between the physical memory address and the logical memory address, includes maintaining a mapping table. The method according to claim 1.
4. The aforementioned mapping table is performed using a translation lookaside buffer. The method according to claim 3.
5. The method further comprises allocating the portion of the first physical memory of the first type based on the expected use of the first physical memory of the first type by the computing entity, The method according to claim 1.
6. The method further comprises dynamically changing the amount of the portion of the first physical memory of the first type allocated to the computing entity based on a usage pattern associated with the use of the first physical memory of the first type allocated to the computing entity. The method according to claim 1.
7. The computing entity further includes exposing the portion of the logical pool memory allocated to it via a software mechanism, so as to enable the computing entity to distinguish between the portion of the logical pool memory allocated to it and the portion of the first type of first physical memory allocated to it. The method according to claim 1.
8. The logical pool memory is coupled via each link managed by the fabric manager to a processor that runs one of the plurality of compute entities to which at least a portion of the logical pool memory associated with the system is allocated. The method according to claim 1.
9. It is a system, The memory portion includes a combination of a portion of a first type first physical memory and a portion of logical pool memory associated with the system, wherein the amount of the logical pool memory, which is mapped to the first type second physical memory and indicated as available for allocation to multiple compute entities, is greater than the amount of the first type second physical memory, and the memory, A logical pool memory controller is coupled to the logical pool memory associated with the system and is configured to (1) track both the state of whether a page of the logical pool memory assigned to any of the plurality of compute entities is a known pattern page and the relationship between a physical memory address and a logical memory address associated with any assigned logical pool memory to any of the plurality of compute entities, and (2) in response to a write operation initiated by a compute entity to write any data other than a known pattern, enable the write operation to write the data to any available space in the second physical memory of the first type, but only up to the extent of the physical memory corresponding to the portion of the logical pool memory previously assigned to the compute entity. A system that has
10. The amount of the logical pool memory associated with the system is equal to the amount of the second physical memory of the first type, combined with the amount of memory corresponding to known pattern pages not supported by any type of physical memory. The system according to claim 9.
11. The logical pool memory controller is further configured to maintain a mapping table to track both the state of whether the page of the logical pool memory assigned to any of the plurality of compute entities is the known pattern page, and the relationship between the physical memory address and the logical memory address associated with any assigned logical pool memory to any of the plurality of compute entities. The system according to claim 9.
12. The system further includes a scheduler configured to allocate the portion of the first physical memory of the first type to the computing entity based on the expected use of the first physical memory of the first type by the computing entity. The system according to claim 9.
13. The scheduler is configured to dynamically change the amount of the portion of the first physical memory of the first type allocated to the compute entity based on the usage pattern associated with the use of the first physical memory of the first type allocated to the compute entity. The system according to claim 12.
14. The logical pool memory controller is further configured to indicate to each of the plurality of compute entities that all pages associated with the logical pool memory that are initially allocated to any of the plurality of compute entities are known pattern pages. The system according to claim 9.
15. It is a system, Multiple host servers configured to run one or more of multiple compute entities, The memory portion includes a combination of a portion of first type first physical memory and a portion of logical pool memory shared among the multiple host servers, wherein the amount of logical pool memory mapped to the second physical memory of the first type and indicated as available for allocation to the multiple compute entities is greater than the amount of the second physical memory of the first type, and the memory, A logical pool memory controller, coupled to the logical pool memory associated with the system, is configured to (1) track both the state of whether a page of the logical pool memory allocated to any of the plurality of compute entities is a known pattern page and the relationship between a physical memory address and a logical memory address associated with any allocated logical pool memory to any of the plurality of compute entities, and (2) in response to a write operation initiated by a compute entity, performed by a processor associated with any of the plurality of host servers to write any data other than known patterns, the write operation allows the data to be written to any available space in the second physical memory of the first type, but only up to the extent of the physical memory corresponding to the portion of the logical pool memory previously allocated to the compute entity. A system that has