Dynamic Page Allocation in Memory
By introducing a memory-side cache monitoring unit into the operating system page allocator, the page allocation strategy is dynamically adjusted, and the problem of memory address conflict in the memory-side cache is solved, improving memory usage efficiency and application performance.
Patent Information
- Application Number
- CN201810982166.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2017-09-27
- Filing Date
- 2018-08-27
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2038-08-27
AI Technical Summary
In the prior art, direct mapping of memory-side caches results in increased memory address conflicts, resulting in inefficient memory usage, and the operating system page allocator cannot dynamically adjust the policy to reduce conflicts.
By introducing a memory-side cache monitoring unit into the operating system page allocator, feedback is provided to dynamically adjust the page allocation strategy, reducing memory address collapse and conflict.
Improves the efficiency of memory-side cache usage, reduces memory address conflicts, and improves application performance.
Smart Images

Figure CN109558338B_ABST
Abstract
Description
Background Art
[0001] Memory devices are typically provided as internal semiconductor integrated circuits in a computer or other electronic device. There are many different types of memory, including volatile memory such as dynamic random access memory (DRAM) and synchronous dynamic random access memory (SDRAM), and non-volatile memory (NVM) such as flash memory.
[0002] Flash memory devices typically use single-transistor memory elements, which allow for high storage density, high reliability, and low power consumption. The data state of each cell is determined by programming a charge storage node (e.g., a floating gate or charge trapping) to change the threshold voltage of the cell. Other NVM such as phase change memory (PCM) uses other physical phenomena such as physical material change or polarization to determine the data state of each cell. Common uses of flash memory and other solid-state memory include personal computers, personal digital assistants (PDAs), digital cameras, digital media players, digital recorders, games, appliances, vehicles, wireless devices, cellular phones, and removable portable memory modules, among others. The uses of such memory continue to expand. Brief Description of the Drawings
[0003] The features and advantages of embodiments of the present invention will become apparent from the following detailed description in conjunction with the accompanying drawings, which illustrate the inventive features by way of example; and, wherein:
[0004] Figure 1 Illustrates a memory address access pattern mapped to a location in a memory-side cache according to an example embodiment;
[0005] Figure 2 Illustrates an operating system (OS) page free list according to an example embodiment;
[0006] Figure 3 Illustrates a computing device including a cache monitoring unit according to an example embodiment, the cache monitoring unit providing feedback to an operating system (OS) page allocator to enable the OS page allocator to dynamically adjust a page allocation policy.
[0007] Figure 4 Illustrates a system operable to allocate physical pages of memory according to an example embodiment;
[0008] Figure 5 Illustrates a memory device operable to allocate physical pages of memory according to an example embodiment;
[0009] Figure 6 Is a flowchart illustrating an operation for allocating physical pages of memory according to an example embodiment; and
[0010] Figure 7 FIG. 2 shows a computing system including a data storage device according to an example embodiment.
[0011] Reference will now be made to the illustrated exemplary embodiments, and specific language will be used herein to describe the same. However, it should be understood that no limitation of the scope of the invention is thereby intended. DETAILED DESCRIPTION
[0012] Before describing embodiments of the disclosed invention, it is to be understood that the invention is not limited to the specific structures, process steps, or materials disclosed herein but extends to their equivalents, as would be recognized by one of ordinary skill in the relevant art. It should also be understood that the terminology used herein is for the purpose of describing particular examples or embodiments only and is not limiting. Like reference numerals in different figures indicate like elements. The numbers provided in the flowcharts and processes are for clarity in illustrating steps and operations and do not necessarily indicate a particular order or sequence.
[0013] References throughout this specification to "an example" mean that a particular feature, structure, or characteristic described in connection with the example is included in at least one embodiment of the invention. Thus, the phrases "in an example" or "an embodiment" appearing throughout this specification are not necessarily all referring to the same embodiment.
[0014] As used herein, for convenience, multiple items, structural elements, constituent elements, and / or materials may be presented in a common list. However, these lists should be interpreted as if each member of the list were individually identified as a separate and unique member. Thus, unless otherwise indicated, no individual member of such a list should be construed as being in fact equivalent to any other member of the same list solely based on their presentation in a common group. Additionally, various examples and embodiments herein may be referred to in connection with alternatives to various components thereof. It should be understood that these embodiments, examples, and alternatives are not to be construed as being in fact equivalents of one another but are to be considered as separate and autonomous representations under the present disclosure.
[0015] Furthermore, the described features, structures, or characteristics may be combined in any suitable manner in one or more embodiments. In the following description, numerous specific details are provided, such as examples of layouts, distances, network examples, etc., to provide a thorough understanding of the embodiments of the invention. However, one of ordinary skill in the relevant art will recognize that the technology may be practiced without one or more of the specific details or with other methods, components, layouts, etc. In other instances, well-known structures, materials, or operations may not be shown or described in detail to avoid obscuring aspects of the present disclosure.
[0016] In the present disclosure, terms such as "comprising", "comprising of", "containing", and "having" can have the meanings given to them in the United States Patent Law and can mean "including", "including of", etc., and are generally interpreted as open-ended terms. The term "consisting of" or "consisting essentially of" is a closed term and includes only the components, structures, steps, etc. specifically listed in conjunction with these terms, and those that comply with the United States Patent Law. "Consisting essentially of" or "consisting essentially of" has the meanings generally given to them in the United States Patent Law. In particular, these terms are generally closed terms, but allow the inclusion of other items, materials, components, steps, or elements that do not materially affect the basic and novel features or functions of the item used in conjunction with them. For example, trace elements present in a composition are allowed under the language of "consisting essentially of" if they do not affect the properties or characteristics of the composition, even if not explicitly listed in the list following these terms. When open-ended terms such as "comprising" or "including" are used in this specification, it should be understood that the language "consisting essentially of" and the language "consisting of" should also be directly supported as if explicitly stated, and vice versa.
[0017] The terms "first", "second", "third", "fourth", etc. in the specification and claims are used to distinguish similar elements and are not necessarily used to describe a particular order or temporal sequence. It should be understood that any terms used in this way are interchangeable under appropriate circumstances, such that the embodiments described herein can, for example, be operated in an order different from that shown or otherwise described herein. Similarly, if a method is described herein as including a series of steps, the order of these steps shown here is not necessarily the only order in which these steps can be performed, and certain of the recited steps may be omitted and / or certain other steps not described herein may be added to the method.
[0018] As used herein, comparative terms such as "increased", "decreased", "better", "worse", "higher", "lower", "enhanced", etc. refer to the properties of a device, component, or activity that are significantly different from other devices, components, or activities in the surrounding or adjacent area, in a single device or multiple similar devices, in a group or category, or in multiple groups or categories, or compared to the state of the art. For example, a data region with an "increased" risk of corruption can refer to a region in a memory device that is more likely to have a write error than other regions in the same memory device. Many factors can contribute to such an increased risk, including location, manufacturing process, the number of program pulses applied to the region, etc.
[0019] As used herein, the term "substantially" refers to the complete or nearly complete degree or level of an action, characteristic, property, state, structure, item, or result. For example, an object that is "substantially" enclosed means that the object is either completely enclosed or nearly completely enclosed. In some cases, the exact allowable degree of deviation from absolute completeness may depend on the particular circumstances. However, generally speaking, if the degree of approximation would have the same overall result as obtaining complete absolute and total completion. When used in a negative sense, the use of "substantially" also applies to refer to the complete or nearly complete lack of an action, characteristic, property, state, structure, item, or result. For example, a composition that is "substantially free of" particles either has no particles at all or has very few particles, and this effect is the same as the effect of having no particles at all. In other words, a composition that is "substantially free of" a component or element may actually still contain such substances, as long as there is no measurable effect.
[0020] As used herein, the term "about" is used to provide flexibility to a numerical range endpoint by providing a value that can be "slightly higher" or "slightly lower" than the given endpoint. However, it should be understood that even when the term "about" is used in conjunction with a specific numerical value in this specification, support is provided for the exact numerical value that is different from the term "about".
[0021] Numerical quantities and data can be expressed or presented herein in a range format. It should be understood that such range formats are used merely for convenience and brevity and should be interpreted flexibly as including not only the explicitly recited numerical values that are the limits of the range, but also all the individual numerical values or sub-ranges that are included within that range, as if each numerical value and sub-range were explicitly recited. By way of illustration, the numerical range of "about 1 to about 5" should be interpreted as including not only the explicitly recited values of about 1 to about 5, but also the individual values and sub-ranges within the indicated range. Thus, included within this numerical range are the individual values such as 2, 3, and 4, and the sub-ranges such as 1-3, 2-4, and 3-5, etc., as well as individually 1, 1.5, 2, 2.3, 3, 3.8, 4, 4.6, 5, and 5.1.
[0022] The same principle applies to ranges that recite only one numerical value as the minimum or maximum. In addition, this interpretation should be applied regardless of the breadth of the range or the characteristics being described.
[0023] An initial overview of the technical embodiments is provided below, and then specific technical embodiments are described in further detail later. This preliminary overview is intended to help the reader understand the technology more quickly, but is not intended to identify key or essential technical features, nor is it intended to limit the scope of the claimed subject matter. Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure pertains.
[0024] Memory devices can utilize non-volatile memory (NVM), which is a storage medium that does not require power to maintain the state of data stored by the medium. NVM has traditionally been used for memory storage or long-term persistent storage tasks, but new and evolving memory technologies allow NVM to be used for roles beyond traditional memory storage. An example of such a role is using NVM as main memory or system memory. Non-volatile system memory (NVMsys) can combine the data reliability of traditional storage with ultra-low latency and high-bandwidth performance, offering many advantages compared to traditional volatile memory, such as high density, large capacity, low power consumption, and reduced manufacturing complexity, to name a few. For example, byte-addressable in-place write NVM (e.g., three-dimensional (3D) cross-point memory) can operate as a byte-addressable memory, similar to dynamic random access memory (DRAM), or as a block-addressable memory, similar to NAND flash. In other words, such NVM can operate as system memory or as persistent storage memory (NVMstor). In some cases where NVM is used as system memory, when power to NVMsys is interrupted, the stored data may be discarded or otherwise rendered unreadable. NVMsys also improves data management flexibility by providing non-volatile, low-latency memory that can be closer to the processor in a computing device. In some examples, NVMsys can reside on a DRAM bus, enabling NVMsys to provide ultra-fast DRAM-like access to data. NVMsys can also be used in computing environments that frequently access large and complex data sets, as well as in environments that are sensitive to downtime due to power failures or system crashes.
[0025] Non-limiting examples of NVM can include planar or three-dimensional (3D) NAND flash, including single-threshold or multi-threshold NAND flash, NOR flash, single-level or multi-level phase change memory (PCM), such as chalcogenide glass PCM, planar or 3D PCM, cross-point array memory, including 3D cross-point memory, non-volatile dual in-line memory module (NVDIMM)-based memory, such as flash-based (NVDIMM-F) memory, flash / DRAM-based (NVDIMM-N) memory, persistent memory-based (NVDIMM-P) memory, 3D cross-point-based NVDIMM memory, resistive RAM (ReRAM), including metal-oxide or oxygen-vacancy-based ReRAM, such as HfO2-, Hf / HfO x -, Ti / HfO2-, TiO x - and TaO-based xReRAMs, filament-based ReRAMs such as Ag / GeS2-, ZrTe / Al2O3-, and Ag-based ReRAMs, programmable metallization cell (PMC) memories such as conductive-bridging RAM (CBRAM), silicon-oxide-nitride-oxide-silicon (SONOS) memories, ferroelectric RAM (FeRAM), ferroelectric transistor RAM (Fe-TRAM), antiferroelectric memories, polymer memories (e.g., ferroelectric polymer memories), magnetoresistive RAM (MRAM), in-situ write non-volatile MRAM (NVMRAM), spin-transfer torque (STT) memories, spin-orbit torque (SOT) memories, nanowire memories, electrically erasable programmable read-only memory (EEPROM), nanotube RAM (NRAM), other memristor- and thyristor-based memories, spin-electronic magnetic junction-based memories, magnetic tunnel junction (MTJ)-based memories, domain wall (DW)-based memories, etc., including combinations thereof. The term "memory device" may refer to the die itself and / or a packaged memory product. The NVM may be byte- or block-addressable. In some examples, the NVM may conform to one or more standards promulgated by the Joint Electron Device Engineering Council (JEDEC), such as JESD21-C, JESD218, JESD219, JESD220-1, JESD223B, JESD223-1, or other suitable standards (the JEDEC standards cited herein are available from www.jedec.org). In a particular example, the NVM may be 3D cross-point memory.
[0026] In one example, with the advent of NVMsys, memory-side caches have become a beneficial type of cache in memory systems. Memory-side caches can be different from central processing unit (CPU)-side caches in that memory-side caches are caches of data in system memory, or in other words, are caches of data that are referenced to a memory layer below the memory-side cache rather than the CPU-side cache, and CPU-side caches are caches referenced to the CPU for the CPU. Memory-side caches can provide fast data access and can thus be directly mapped to enable faster lookup and retrieval of data stored therein. Additionally, memory-side caches can span tens to hundreds of gigabits (GB) in capacity, which is typically capable of being 1 / 8 th or 1 / 16 thCaching is performed on the left and right. The memory - side cache can include any type of volatile memory that can be used as a cache. Volatile memory can include any type of volatile memory and is not considered restrictive. Non - restrictive examples of volatile memory can include random access memory (RAM), such as static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), etc., including combinations thereof. SDRAM memory can include any of its variants, such as single - data - rate SDRAM (SDR DRAM), double - data - rate (DDR) SDRAM, including DDR, DDR2, DDR3, DDR4, DDR5, etc., collectively referred to as DDRx, and low - power DDR (LPDDR) SDRAM, including LPDDR, LPDDR2, LPDDR3, LPDDR4, etc., collectively referred to as LPDDRx. In some examples, DRAM complies with standards promulgated by JEDEC, such as JESD79F for DDR SDRAM, JESD79 - 2F for DDR2 SDRAM, JESD79 - 3F for DDR3 SDRAM, JESD79 - 4A for DDR4 SDRAM, JESD209B for LPDDR SDRAM, JESD209 - 2F for LPDDR2 SDRAM, JESD209 - 3C for LPDDR3 SDRAM, and JESD209 - 4A for LPDDR4 SDRAM (these standards are available at www.jedec.org; DDR5 SDRAM is upcoming). These standards (and similar standards) can be referred to as DDR - based or LPDDR - based standards, and communication interfaces that implement such standards can be referred to as DDR - based or LPDDR - based interfaces. In one specific example, the volatile memory can be DRAM. In another specific example, the volatile memory can be DDRx SDRAM. In yet another specific aspect, the volatile memory can be LPDDRx SDRAM. Additionally, in some cases, the volatile memory can be functionally coupled to a memory card, such as a dual - inline memory module (DIMM), which is coupled into the system to be used as a memory - side cache.
[0027] In one example, the increased capacity of the memory - side cache is a new variation of the memory hierarchy in a memory device. The capacity of the memory - side cache is large enough such that an entire data set can fit within it. However, due to the direct - mapped nature of the memory - side cache, a workload may utilize only a small fraction of the available capacity in the memory - side cache. For example, based on the specific access pattern of the workload, only a small fraction of the available capacity can be utilized due to conflicts in the memory - side cache. Accessing memory locations in the memory - side cache according to a specific memory access pattern can result in folding onto a small group of memory locations in the direct - mapped memory - side cache. In other words, several physical memory addresses can map to the same memory location in the memory - side cache, which can lead to inefficient use and an increase in the number of conflicts in the memory - side cache.
[0028] Figure 1 An example of a traditional memory - address access pattern that maps to memory locations in the memory - side cache is shown. As depicted, the memory - address access pattern can fold onto a small group of memory locations in the memory - side cache, which can be a direct - mapped memory - side cache. In this example, the memory - address access pattern maps to only two locations in the memory - side cache out of all available locations, resulting in inefficient use of the available cache space and an increase in the number of conflicts (i.e., at the two locations).
[0029] In one configuration, an operating - system (OS) page allocator can allocate physical pages in memory for a workload, and the memory - address access pattern of the workload can be based on the physical pages that the OS page allocator allocates to the workload in memory. Physical pages in memory can be associated with memory addresses, and the memory - address access pattern can be based on the physical pages in memory. As an example, physical pages X1, Y1, and Z1 can have addresses 100, 200, and 300 respectively, and these addresses can be accessed according to the memory - address access pattern.
[0030] In one example, while access to a 4 kilobyte (K) span (i.e., the lower 12 bits within a page) can be based on the workload itself, 4K is an upper limit controlled by the workload. The higher order bits of the physical address (e.g., 13 and higher) can be determined by the OS page allocator. For example, a workload that is sequentially streaming in virtual address space can be given discontinuous random physical pages, which can result in a non-streaming physical address access pattern in memory (outside of 4K blocks). On the other hand, a workload with some arbitrary access pattern may inadvertently be forced by the OS page allocator to collide on a reduced set of locations in a direct mapped cache, which is beneficial for non-critical applications in a shared environment but detrimental for critical applications.
[0031] Figure 2 An example of an operating system (OS) page free list is shown. As shown, the OS page free list can typically be folded to a power of 2, which can lead to an increased number of collisions on a small number of entries in the direct mapped cache. This problem can be seen in multi-channel dynamic random access memory (MCDRAM) caches and two-level memory (2LM) caches, which can result in poor application-level performance. Over time, the OS page free list can shrink or become unordered or scrambled, and the pages that are distributed can become less efficient. Over time, the pages that are distributed are more likely to collide with existing pages that are currently distributed, and in previous solutions, there has been no mechanism for determining the most efficient page to distribute from the OS page free list to avoid collisions.
[0032] In one configuration, the OS page allocator can apply page allocation patterns and policies, which can be used to determine the effectiveness and usage of the memory side cache. The application of the page allocation patterns and policies can be a control mechanism independent of the application. However, in previous solutions, the OS page allocator was separated from the static and dynamic behavior of the memory side cache. In previous solutions, the OS page allocator did not have a mechanism to adjust the free list / page allocation policy based on feedback or specifications. Additionally, in previous solutions, when folding to a small set of memory locations occurred in a direct mapped memory side cache, the solution was to statically configure a random page allocation policy, but this solution was ineffective in reducing the folding to the small set of memory locations.
[0033] In the present technology, the OS page allocator in the system can receive system-level or per-process guidance from an application, or feedback from a memory-side cache monitoring unit associated with the memory-side cache, and the guidance and / or feedback can enable the OS page allocator to dynamically adjust the page allocation mode and policy. For example, the OS page allocator can receive hints and feedback from the memory-side cache monitoring unit regarding memory address folding, memory-side cache usage, memory filling in the system, the size and interleaving scheme of the memory-side cache, the memory access pattern of the application, etc. Based on the hints and feedback received from the memory-side cache monitoring unit, the OS page allocator can appropriately adjust its page allocation mode and policy. In other words, the OS page allocator can implement a dynamic scheme for modifying the page allocation mode and policy based on the hints and feedback received from the memory-side cache monitoring unit, which can reduce the likelihood of a small set of memory locations being folded into the directly mapped memory-side cache. Therefore, the memory-side cache monitoring unit can coordinate the use of the memory-side cache with physical page allocation, which can improve application performance as well as the use and effectiveness of the memory-side cache.
[0034] Figure 3 An exemplary computing device 300 is shown. The computing device 300 can include a network interface controller (NIC) that enables communication between nodes in the computing device 300. The computing device 300 can include one or more cores or processors, such as Core 0, Core I, and Core 2. One or more cores or processors in the computing device 300 can be used to execute an application (e.g., App B). The computing device 300 can include a cache agent (CA) and a last-level cache (LLC). The cache agent can process memory data requests from the cores in the computing device 300. The computing device 300 can include NVM and DDRx memory. The DDRx memory can include a cache 330, such as a memory-side cache. The cache 330 can use direct mapping to store data to provide fast access to the cache data. However, as previously mentioned, direct mapping can lead to a small set of memory locations being folded into the directly mapped cache 330, which can result in only a small portion of the available capacity in the cache 330 being utilized due to memory address conflicts in the cache 330.
[0035] In one configuration, computing device 300 may include a cache monitoring unit 310 communicatively coupled to a cache 330 in the computing device 300. For example, the cache monitoring unit 310 may be a memory - side cache monitoring unit that monitors the cache 330 (e.g., a memory - side cache). The cache monitoring unit 310 may also be referred to as an intelligent eviction allocation module, which may be a hardware unit in the computing device 300. The cache monitoring unit 310 can be used to monitor the usage and activity in the cache 330, and based on this monitoring, the cache monitoring unit 310 can identify or receive feedback from the cache 330. The cache monitoring unit 310 can forward the feedback to an operating system (OS) page allocator 320 in the computing device 300. The OS page allocator 320 can dynamically adjust the page allocation mode or policy of the system (e.g., the cache 330 and / or NVM) based on the feedback received from the cache monitoring unit 310. In other words, the OS page allocator 320 can adjust the page allocation policy based on the feedback received from the cache monitoring unit 310, where the page allocation policy defines the physical pages allocated by the OS page allocator.
[0036] In one example, the OS page allocator 320 can dynamically change the page allocation mode to obtain an improved memory address access pattern with a reduced number of address conflicts. Adjusting the page allocation policy can prevent folding into a reduced set of locations in the cache 330. Physical pages can have addresses, and the page allocation policy can determine which physical pages (or addresses) to use. As an example, physical pages X1, Y1, Z1 may have addresses 100, 200, 300, and these addresses can fold in a cache with 100 addresses (e.g., 100 mod 100 equals 200 mod 100, which equals 300 mod 100). In this example, feedback regarding memory address folding can be provided to the OS page allocator 320, and then the OS page allocator 320 can allocate a new set of physical pages. For example, the OS page allocator 320 can allocate pages X2, Y2, and Z2 with addresses 120, 140, and 160. For this new set of physical pages, these addresses do not fold because 120 mod 100 is not equal to 140 mod 100, etc. Thus, the OS page allocator 320 can allocate a different set of physical pages, which can target a different set of physical addresses in the memory (e.g., the cache 330 and / or NVM), resulting in reduced folding and reduced memory address conflicts in the memory. As a result, an increased portion of the cache 330 can be successfully utilized to store data.
[0037] In one example, the OS page allocator 320 can adjust the page allocation mode or policy to modify the memory address access mode for the cache 330, which can reduce a reduced set of memory addresses that are folded into the memory of the computing device 300 (e.g., cache 330 and / or NVM). In other words, adjusting the page allocation mode or policy can modify the physical pages allocated by the OS page allocator 320, and the adjustment to the page allocation mode or policy can be used to modify the memory address access mode. The physical pages can be associated with memory addresses, and the memory address access mode can be based on the physical pages allocated by the OS page allocator 320. Additionally, based on the reduced folding in memory, the number of memory address conflicts in the memory of the computing device 300 can be reduced.
[0038] In one example, the cache monitoring unit 310 can identify feedback from the cache 330 that includes memory address folding information. The memory address folding information can identify a memory address range in the cache 330 and / or NVM associated with an increased amount of memory address folding and / or certain address ranges in the cache 330 and / or NVM associated with an increased number of memory address conflicts. The cache monitoring unit 310 can provide the address folding information (including the problematic memory address ranges) to the OS page allocator 320. The OS page allocator 320 can adjust the page allocation policy to allocate physical pages that avoid the memory address ranges associated with the increased amount of memory address folding and the increased number of memory address conflicts. For example, the OS page allocator 320 can allocate or distribute pages from different regions in the cache 330 and / or NVM. In other words, the OS page allocator 320 can avoid the more problematic (i.e., having increased folding and conflicts) regions in the cache 330 and / or NVM.
[0039] In another example, the feedback identified at the cache monitoring unit 310 may include cache miss information. Based on the cache miss information, the cache monitoring unit 310 may identify certain address bits that result in an increased number of memory address folds and memory address conflicts in the cache 330. In other words, the cache monitoring unit 310 may identify problem address bits (e.g., 3 bits in the conflicting memory addresses of a memory address range) that result in an increased number of memory address folds and memory address conflicts. The cache monitoring unit 310 may incorporate this information into the feedback and then provide the feedback to the OS page allocator 320. Based on this feedback, the OS page allocator 320 may identify the specific address bits that result in an increased number of memory address folds and memory address conflicts. The OS page allocator 320 may shuffle, hash, or randomize these address bits within the address field and then allocate or distribute physical pages in the cache 330 and / or NVM based on these different address bits. In other words, by modifying the address bits, the OS page allocator 320 may allocate or distribute physical pages from different regions in the cache 330 and / or NVM, thus avoiding the more problematic (i.e., having increased folds and conflicts) regions in the cache 330 and / or NVM.
[0040] In one example, the basic input / output system (BIOS) may be modified to interleave memory addresses such that a combination of higher-order bits and lower-order bits may be used to determine the cache-to-remote memory mapping distribution, and these bits may be exposed to the OS page allocator 320 that allocates physical address pages.
[0041] In one example, the cache monitoring unit 310 may observe that passing feedback to the OS page allocator 310 does not improve cache behavior in terms of memory address folds and conflicts. In such a case, the cache monitoring unit 310 may determine to modify the attributes of the cache 330, such as the associativity of the cache 330, in order to improve cache behavior.
[0042] In one example, the feedback identified at the cache monitoring unit 310 can include an indication of the size of the cache 330, the interleaving scheme of the cache 330, the memory address access pattern of the application that utilizes the data stored in the cache 330, the application identifier (ID) or process ID associated with increased priority, the usage information of the cache 330, and / or real-time telemetry information. Additionally, the cache monitoring unit 310 can identify system model specific registers (MSRs) that can specify a particular type of feedback to be provided to the OS page allocator 320. In other words, the system MSRs can provide a level of system configurability for the feedback provided to the OS page allocator 320. The cache monitoring unit 310 can provide feedback (e.g., the size of the cache 330, interleaving scheme, memory address access pattern of the application, application ID, process ID, usage information, telemetry data, etc.) to the OS page allocator 320, and the OS page allocator 320 can adjust the page allocation policy based on this feedback, where the page allocation policy defines the physical pages allocated by the OS page allocator 320.
[0043] In one example, the OS page allocator 320 can have a number of resource pools available in memory (e.g., cache 330 and / or NVM) and a page allocation policy based on these resource pools. Based on the input from the cache monitoring unit 310, the OS page allocator 320 can adjust the page allocation policy to one of, for example, a random page allocation policy, a large page allocation policy, a range page allocation policy, etc. Based on the random page allocation policy, the OS page allocator 320 can allocate physical pages in a semi-random manner to maximize expansion. Based on the large page allocation policy, the OS page allocator 320 can utilize XYZ bit folding to allocate physical pages of increased size (e.g., in the range of megabytes (MB) or gigabytes (GB) instead of 4 kilobytes (KB)). Based on the range page allocation policy, the OS page allocator 320 can allocate physical pages with physical addresses within a specific range. Alternatively, the OS page allocator 320 can adjust the page allocation policy according to different mechanisms.
[0044] Figure 4System 400 that is operable to allocate physical pages of a memory is shown. System 400 may include a memory - side cache 410. System 400 may include a memory - side cache monitoring unit 420 coupled to the memory - side cache 410. System 400 may include an operating system (OS) page allocator 430. The OS page allocator 430 may receive feedback from the memory - side cache monitoring unit 420. The OS page allocator 430 may adjust a page allocation policy based on the feedback received from the memory - side cache monitoring unit 420, where the page allocation policy defines the physical pages allocated by the OS page allocator 430.
[0045] Figure 5 Memory device 500 that is operable to allocate physical pages of a memory 510 is shown. Memory device 500 may include a memory 510. Memory device 500 may include a memory - side cache 520. Memory device 500 may include a memory - side cache monitoring unit 530 coupled to the memory - side cache 520. Memory device 500 may include an operating system (OS) page allocator 540. The OS page allocator 540 may receive feedback from the memory - side cache monitoring unit 530. The OS page allocator 540 may adjust a page allocation policy based on the feedback received from the memory - side cache monitoring unit 530, where the page allocation policy defines the physical pages allocated by the OS page allocator 540.
[0046] Another example provides a method 600 for allocating physical pages of a memory, as Figure 6 shown in the flowchart. The method may be executed as instructions on a machine, where the instructions are included on at least one computer - readable medium or one non - transitory machine - readable storage medium. The method may include an operation of receiving feedback at an operating system (OS) page allocator in a memory device from a memory - side cache monitoring unit in the memory device, where the memory - side cache monitoring unit is coupled to a memory - side cache in the memory device, as shown in block 610. The method may include an operation of adjusting a page allocation policy at the OS page allocator based on the feedback received from the memory - side cache monitoring unit, where the page allocation policy defines the physical pages allocated by the OS page allocator, as shown in block 620.
[0047] Figure 7Illustrated is a general computing system or device 700 that can be used in the present technology. The computing system 700 can include a processor 702 that communicates with a memory 704. The memory 704 can include any device, combination of devices, circuitry, etc. capable of storing, accessing, organizing, and / or retrieving data. Non-limiting examples include SAN (Storage Area Network), cloud storage network, volatile or non-volatile RAM, phase change memory, optical media, hard drive type media, etc., including combinations thereof.
[0048] The computing system or device 700 further includes a local communication interface 706 for connections between the various components of the system. For example, the local communication interface 706 can be a local data bus and / or any associated address or control bus as needed.
[0049] The computing system or device 700 can also include an I / O (input / output) interface 708 for controlling the I / O functions of the system and for I / O connections to devices external to the computing system 700. A network interface 710 can also be included for network connections. The network interface 710 can control network communications within and outside the system. The network interface can include a wired interface, wireless interface, Bluetooth interface, optical interface, etc., including suitable combinations thereof. Additionally, the computing system 700 can also include a user interface 712, a display device 714, and various other components beneficial to such a system.
[0050] The processor 702 can be a single or multiple processors, and the memory 704 can be a single or multiple memories. The local communication interface 706 can serve as a path to facilitate communication between any single processor, multiple processors, single memory, multiple memories, various interfaces, etc. in any useful combination.
[0051] Various technologies or certain aspects or portions thereof may take the form of program code (i.e., instructions) embodied in a tangible medium, such as a floppy disk, CD-ROM, hard disk drive, non-transitory computer-readable storage medium, or any other machine-readable storage medium, wherein when the program code is loaded into and executed by a machine such as a computer, the machine becomes an apparatus for practicing the various technologies. A circuit may include hardware, firmware, program code, executable code, computer instructions, and / or software. A non-transitory computer-readable storage medium may be a computer-readable storage medium that does not include a signal. In the case of executing program code on a programmable computer, a computing device may include a processor, a processor-readable storage medium (including volatile and NVM and / or storage elements), at least one input device, and at least one output device. The volatile and NVM and / or storage elements may be RAM, EPROM, flash drive, optical disk drive, magnetic hard disk drive, solid state drive, or other media for storing electronic data. Nodes and wireless devices may also include transceiver modules, counter modules, processing modules, and / or clock modules or timer modules. One or more programs that implement or utilize the various technologies described herein may use application programming interfaces (APIs), reusable controls, etc. These programs may be implemented in a high-level procedural or object-oriented programming language to communicate with a computer system. However, if desired, the programs may be implemented in assembly language or machine language. In any case, the language may be a compiled or interpreted language and combined with a hardware implementation. Exemplary systems or devices may include, but are not limited to, laptop computers, tablet computers, desktop computers, smart phones, computer terminals, and servers, storage databases, and other electronic devices that utilize circuits and programmable memory, such as household appliances, smart TVs, digital video disc (DVD) players, heating, ventilation, and air conditioning (HVAC) controllers, light switches, etc.
[0052] Example
[0053] The following examples relate to specific embodiments of the invention and point out specific features, elements, or steps that may be used in implementing these embodiments or otherwise combined.
[0054] In one example, a device for allocating physical pages of memory is provided. The device may include an operating system (OS) page allocator and a communication interface for coupling the OS page allocator to a memory-side cache and a memory-side cache monitoring unit coupled to the memory-side cache. The OS page allocator is operable to receive feedback from the memory-side cache monitoring unit and adjust a page allocation policy based on the feedback received from the memory-side cache monitoring unit, the page allocation policy defining the physical pages allocated by the OS page allocator.
[0055] In one example of the device, the OS page allocator is operable to adjust the page allocation policy to allocate physical pages as follows, the physical pages avoiding memory address ranges associated with an increasing number of memory address folds, wherein the memory address ranges are identified in the memory address fold information included in the feedback received from the memory side cache monitoring unit.
[0056] In one example of the device, the OS page allocator is operable to adjust the page allocation policy to modify the memory address access pattern for the memory side cache, thereby reducing a reduced set of memory addresses folded into the memory side cache and the resulting memory address conflicts in the memory side cache.
[0057] In one example of the device, the OS page allocator is operable to adjust the page allocation policy to modify the memory address access pattern for the memory side cache, wherein the physical pages are associated with memory addresses, and the memory address access pattern is based on the physical pages allocated by the OS page allocator.
[0058] In one example of the device, the feedback received from the memory side cache monitoring unit includes an indication of the size of the memory side cache and an interleaving scheme for the memory side cache.
[0059] In one example of the device, the memory side cache monitoring unit is configured to determine to modify one or more attributes of the memory side cache when one or more of the amount of memory address folds in the memory side cache or the number of memory address conflicts in the memory side cache exceeds a defined threshold, wherein the one or more attributes of the memory side cache include the associativity attribute of the memory side cache.
[0060] In one example of the device, the feedback received from the memory side cache monitoring unit includes the memory address access pattern of an application that uses the data stored in the memory side cache.
[0061] In one example of the device, the feedback received from the memory side cache monitoring unit includes an application identifier (ID) or a process ID associated with an increased priority.
[0062] In one example of the device, the feedback received from the memory side cache monitoring unit includes one or more of the following: usage information for the memory side cache or real-time telemetry information.
[0063] In one example of the device, the OS page allocator is operable to adjust the page allocation policy to one of the following: a random page allocation policy, a large page allocation policy, or a range page allocation policy.
[0064] In one example of the device, a non-volatile memory (NVM) is communicatively coupled to the memory side cache.
[0065] In one example, a memory system is provided that is operable to allocate physical pages of memory. The memory system may include NVM. The memory system may include a memory side cache. The memory system may include a memory side cache monitoring unit coupled to the memory side cache. The memory system may include an operating system (OS) page allocator. The OS page allocator may receive feedback from the memory side cache monitoring unit. The OS page allocator may adjust the page allocation policy based on the feedback received from the memory side cache monitoring unit, where the page allocation policy defines the physical pages allocated by the OS page allocator.
[0066] In one example of the memory system, the OS page allocator is operable to adjust the page allocation policy to allocate physical pages that avoid a memory address range associated with an increased amount of memory address folding, where the memory address range is identified in the memory address folding information included in the feedback received from the memory side cache monitoring unit.
[0067] In one example of the memory system, the OS page allocator is operable to adjust the page allocation policy to modify the memory address access pattern for the memory side cache, thereby reducing a reduced set of memory addresses folded into the memory side cache and the resulting memory address conflicts in the memory side cache.
[0068] In one example of the memory system, the OS page allocator is operable to adjust the page allocation policy to modify the memory address access pattern for the memory side cache, where the physical pages are associated with memory addresses and the memory address access pattern is based on the physical pages allocated by the OS page allocator.
[0069] In one example of the memory system, the OS page allocator is operable to adjust the page allocation policy to define physical pages allocated in one or more of the following: NVM or the memory side cache.
[0070] In one example of the memory system, the feedback received from the memory side cache monitoring unit includes an indication of the size of the memory side cache and an interleaving scheme for the memory side cache.
[0071] In one example of a memory system, feedback received from a memory - side cache monitoring unit includes a memory - address access pattern of an application that utilizes data stored in the memory - side cache.
[0072] In one example of a memory system, feedback received from a memory - side cache monitoring unit includes an application identifier (ID) or a process ID associated with an increased priority.
[0073] In one example of a memory system, feedback received from a memory - side cache monitoring unit includes one or more of the following: usage information for the memory - side cache or real - time telemetry information.
[0074] In one example of a memory system, an OS page allocator is operable to adjust a page - allocation policy to one of the following: a random page - allocation policy, a large - page - allocation policy, or a range page - allocation policy.
[0075] In one example of a memory system, the memory - side cache uses direct mapping to store data in the memory - side cache.
[0076] In one example of a memory system, the NVM and the memory - side cache are located on the same dual - in - line memory module (DIMM).
[0077] In one example, a method for allocating physical pages of memory is provided. The method can include receiving, at an operating system (OS) page allocator, feedback from a memory - side cache monitoring unit in a memory device, where the memory - side cache monitoring unit is coupled to a memory - side cache in the memory device. The method can include adjusting, at the OS page allocator, a page - allocation policy based on the feedback received from the memory - side cache monitoring unit, where the page - allocation policy defines physical pages allocated by the OS page allocator.
[0078] In one example of a method for allocating physical pages of memory, the method can further include adjusting the page - allocation policy to allocate physical pages that avoid a memory - address range associated with an increased amount of memory - address folding, where the memory - address range is identified in memory - address - folding information included in the feedback received from the memory - side cache monitoring unit.
[0079] In one example of a method for allocating physical pages of memory, the method can further include adjusting the page - allocation policy to modify a memory - address access pattern for the memory - side cache, thereby reducing a reduced set of memory addresses that fold into the memory - side cache and resulting memory - address conflicts in the memory - side cache.
[0080] In one example of a method for allocating physical pages of memory, the method may further include adjusting a page allocation policy to modify a memory address access pattern for a memory - side cache, where the physical pages are associated with memory addresses, and the memory address access pattern is based on the physical pages allocated by an OS page allocator.
[0081] In one example of a method for allocating physical pages of memory, the method may also include adjusting the page allocation policy to one of the following: a random page allocation policy, a large - page allocation policy, or a range page allocation policy.
[0082] In one example of a method for allocating physical pages of memory, feedback received from a memory - side cache monitoring unit includes one or more of the following: an indication of the size of the memory - side cache; an interleaving scheme for the memory - side cache; a memory address access pattern of an application that uses data stored in the memory - side cache; an application identifier (ID) or process ID associated with increased priority; usage information of the memory - side cache; or real - time telemetry information.
[0083] In one example, a device operable to allocate physical pages of memory is provided. The device may include a memory - side cache. The device may include a memory - side cache monitoring unit coupled to the memory - side cache. The device may include an operating system (OS) page allocator. The OS page allocator may receive feedback from the memory - side cache monitoring unit. The OS page allocator may adjust the page allocation policy based on the feedback received from the memory - side cache monitoring unit, where the page allocation policy defines the physical pages allocated by the OS page allocator.
[0084] In one example of the device, the OS page allocator is operable to adjust the page allocation policy to allocate physical pages that avoid a memory address range associated with an increased amount of memory address folding, where the memory address range is identified in memory address folding information included in the feedback received from the memory - side cache monitoring unit.
[0085] In one example of the device, the OS page allocator is operable to adjust the page allocation policy to modify a memory address access pattern for the memory - side cache, thereby reducing a reduced set of memory addresses folded into the memory - side cache and resulting memory address conflicts in the memory - side cache.
[0086] In one example of the device, the OS page allocator is operable to adjust the page allocation policy to modify the memory address access pattern for the memory - side cache, where physical pages are associated with memory addresses and the memory address access pattern is based on the physical pages allocated by the OS page allocator.
[0087] In one example of the device, the feedback received from the memory - side cache monitoring unit includes an indication of the size of the memory - side cache and an interleaving scheme for the memory - side cache.
[0088] In one example of the device, the memory - side cache monitoring unit is configured to determine to modify one or more attributes of the memory - side cache when one or more of the amount of memory - address folding in the memory - side cache or the number of memory - address conflicts in the memory - side cache exceeds a defined threshold, where the one or more attributes of the memory - side cache include the associativity attribute of the memory - side cache.
[0089] In one example of the device, the feedback received from the memory - side cache monitoring unit includes the memory - address access pattern of an application that utilizes the data stored in the memory - side cache.
[0090] In one example of the device, the feedback received from the memory - side cache monitoring unit includes an application identifier (ID) or a process ID associated with an increased priority.
[0091] In one example of the device, the feedback received from the memory - side cache monitoring unit includes one or more of the following: usage information for the memory - side cache or real - time telemetry information.
[0092] In one example of the device, the OS page allocator is operable to adjust the page allocation policy to one of the following: a random page allocation policy, a large - page allocation policy, or a range page allocation policy.
[0093] In one example of the device, a non - volatile memory (NVM) is communicatively coupled to the memory - side cache.
[0094] While the foregoing examples illustrate the principles of the embodiments of the present invention in one or more particular applications, it will be apparent to those of ordinary skill in the art that several modifications may be made to the form, use, and details of the implementations without exercising inventive faculty and without departing from the principles and concepts of the disclosure.
Claims
1. A device for allocating physical pages of memory, comprising: An operating system (OS) page allocator; And A communication interface for coupling the OS page allocator to a direct mapped memory side cache and to a memory side cache monitoring unit coupled to the direct mapped memory side cache; Wherein the OS page allocator is operable to; Receive feedback from the memory side cache monitoring unit, wherein the feedback includes memory address folding information for identifying a memory address range in the direct mapped memory side cache associated with an increased amount of memory address folding, wherein the memory address folding occurs when multiple physical memory addresses accessed based on a memory address access pattern for a workload are mapped to the same memory location in the direct mapped memory side cache; and Based on the feedback received from the memory side cache monitoring unit, adjust a page allocation policy that defines physical pages allocated by the OS page allocator for the workload to avoid the memory address range associated with the increased amount of memory address folding.
2. The device according to claim 1, wherein The OS page allocator is operable to adjust the page allocation policy to modify a memory address access pattern for the direct mapped memory side cache, thereby reducing a reduced set of memory addresses folded into the direct mapped memory side cache and resulting memory address conflicts in the direct mapped memory side cache.
3. The device according to claim 1, wherein The OS page allocator is operable to adjust the page allocation policy to modify a memory address access pattern for the direct mapped memory side cache, wherein the physical pages are associated with memory addresses and the memory address access pattern is based on the physical pages allocated by the OS page allocator.
4. The device according to claim 1, wherein, The memory side cache monitoring unit is configured to determine to modify one or more attributes of the direct mapped memory side cache when one or more of a memory address folding amount in the direct mapped memory side cache or a number of memory address conflicts in the direct mapped memory side cache exceeds a defined threshold, wherein the one or more attributes of the direct mapped memory side cache include an associativity attribute of the direct mapped memory side cache.
5. The device according to claim 1, wherein, The feedback received from the memory side cache monitoring unit includes an indication of the size of the direct mapped memory side cache and an interleaving scheme for the direct mapped memory side cache.
6. The device according to claim 1, wherein, The feedback received from the memory side cache monitoring unit includes a memory address access pattern of an application that uses data stored in the direct mapped memory side cache.
7. The device according to claim 1, wherein, The feedback received from the memory side cache monitoring unit includes an application identifier (ID) or process ID associated with an increased priority.
8. The device according to claim 1, wherein The feedback received from the memory - side cache monitoring unit includes one or more of the following: usage information of the direct - mapped memory - side cache or real - time telemetry information.
9. The device according to claim 1, wherein The OS page allocator is operable to adjust the page - allocation policy to one of the following: a random page - allocation policy, a large - page - allocation policy, or a range page - allocation policy.
10. The apparatus of claim 1, further comprising a non - volatile memory (NVM) communicatively coupled to the direct - mapped memory - side cache.
11. A memory system operable to allocate physical pages of memory, comprising: A non - volatile memory (NVM); A direct - mapped memory - side cache communicatively coupled to the NVM; A memory - side cache monitoring unit communicatively coupled to the direct - mapped memory - side cache; and An operating - system (OS) page allocator operable to: Receive feedback from the memory - side cache monitoring unit, wherein the feedback includes memory - address folding information for identifying a memory - address range in the direct - mapped memory - side cache associated with an increased amount of memory - address folding, wherein the memory - address folding occurs when multiple physical memory addresses accessed based on a memory - address access pattern for a workload are mapped to the same memory location in the direct - mapped memory - side cache; and Based on the feedback received from the memory - side cache monitoring unit, adjust a page - allocation policy that defines physical pages allocated by the OS page allocator for the workload to avoid the memory - address range associated with the increased amount of memory - address folding.
12. The system according to claim 11, wherein, The OS page allocator is operable to adjust the page - allocation policy to modify a memory - address access pattern for the direct - mapped memory - side cache, thereby reducing a reduced set of memory addresses folded into the direct - mapped memory - side cache and the resulting memory - address conflicts in the direct - mapped memory - side cache.
13. The system according to claim 11, wherein, The OS page allocator is operable to adjust the page - allocation policy to modify a memory - address access pattern for the direct - mapped memory - side cache, wherein the physical pages are associated with memory addresses and the memory - address access pattern is based on the physical pages allocated by the OS page allocator.
14. The system according to claim 11, wherein, The OS page allocator is operable to adjust the page - allocation policy to define physical pages allocated in one or more of the following: the NVM or the direct - mapped memory - side cache.
15. The system according to claim 11, wherein, The feedback received from the memory - side cache monitoring unit includes a memory - address access pattern of an application that utilizes data stored in the direct - mapped memory - side cache.
16. The system according to claim 11, wherein The OS page allocator is operable to adjust the page - allocation policy to one of the following: a random page - allocation policy, a large - page - allocation policy, or a range page - allocation policy.
17. The system according to claim 11, wherein The NVM and the direct mapped memory side cache are located in the same dual in-line memory module (DIMM) in the memory device.
18. A method for allocating physical pages of memory, the method comprising: Receiving, at an operating system (OS) page allocator in a memory device, feedback from a memory side cache monitoring unit in the memory device, wherein the memory side cache monitoring unit is coupled to a direct mapped memory side cache in the memory device, wherein the feedback includes memory address folding information for identifying a memory address range in the direct mapped memory side cache associated with an increased amount of memory address folding, wherein the memory address folding occurs when multiple physical memory addresses accessed based on a memory address access pattern for a workload are mapped to the same memory location in the direct mapped memory side cache; and At the OS page allocator, adjusting a page allocation strategy that defines physical pages allocated by the OS page allocator for a workload based on the feedback received from the memory side cache monitoring unit to avoid the memory address range associated with the increased amount of memory address folding.
19. The method according to claim 18, further comprising: Adjusting the page allocation strategy to modify a memory address access pattern for the direct mapped memory side cache, thereby reducing a reduced set of memory addresses folded into the direct mapped memory side cache and resulting memory address conflicts in the direct mapped memory side cache.
20. The method according to claim 18, further comprising: Adjusting the page allocation strategy to modify a memory address access pattern for the direct mapped memory side cache, wherein the physical pages are associated with memory addresses and the memory address access pattern is based on physical pages allocated by the OS page allocator.
21. The method according to claim 18, further comprising: Adjusting the page allocation strategy to one of: a random page allocation strategy, a large page allocation strategy, or a range page allocation strategy.
22. The method according to claim 18, wherein, The feedback received from the memory side cache monitoring unit includes one or more of: An indication of the size of the direct mapped memory side cache; An interleaving scheme for the direct mapped memory side cache; A memory address access pattern of an application that uses data stored in the direct mapped memory side cache; An application identifier (ID) or process ID associated with an increased priority; Usage information of the direct mapped memory side cache; or Real-time telemetry information.
23. A computer program product comprising instructions that, when executed by a processor, cause the processor to perform the method according to any one of claims 18-22.
24. A computer-readable storage medium having stored thereon instructions that, when executed by a processor, cause the processor to perform the method according to any one of claims 18-22.
Citation Information
Patent Citations
Cache memory, working method thereof and processor
CN106776365A