A dynamic random access memory cache is provided as a second type of memory for each application process.
By dynamically allocating DRAM as a cache for slower, larger-capacity memory types, the invention addresses inefficiencies in processor-memory balance, enhancing system performance through optimized access times and capacities for individual application processes.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2022-10-24
- Publication Date
- 2026-04-03
AI Technical Summary
Modern computing systems face challenges in efficiently balancing processor performance with memory performance, as memory mechanisms often have slow access times and varying capacities, leading to inefficiencies in data processing and storage.
Implementing a region of DRAM as a cache for slower, larger-capacity memory types like non-volatile memory (NVM) flash or phase-change material (PCM) by dynamically allocating private areas of DRAM as caches for each application process, adjusting cache sizes based on performance requirements.
This approach enhances memory efficiency by optimizing access times and capacities for individual application processes, improving overall system performance and reducing latency.
Smart Images

Figure 0007840401000001 
Figure 0007840401000002 
Figure 0007840401000003
Abstract
Description
Technical Field
[0001] The present invention generally relates to computing systems, and more particularly to various embodiments for providing dynamic random access memory (DRAM) as a cache of a second type of memory for each application process that uses one or more computing processors.
Background Art
[0002] In modern society, computer systems have become commonplace. Computer systems can be found in workplaces, homes, schools, etc. A computer system may include a data storage system or a disk storage system for processing and storing data. In recent years, both software technology and hardware technology have made remarkable progress. New technologies have added more functions and provided higher convenience in the use of these computer systems. Currently, the amount of information to be processed has increased significantly. Therefore, processing, storing, searching for, or a combination of various information has become an important issue to be solved.
Summary of the Invention
[0003] Various embodiments are illustrated for a processor to provide a first type of memory as a second type of memory to a computing system. An application process running within the computing system may be identified. An area of the first type of memory may be provided as a cache of the second type of memory for the application process.
[0004] To facilitate understanding of the advantages of the present invention, a more specific description of the invention, as briefly outlined above, will be made by reference to specific embodiments illustrated in the accompanying drawings. With the understanding that these drawings depict only typical embodiments of the invention and are therefore not intended to limit its scope, the invention will be described and explained more specifically and in detail through the use of the accompanying drawings. [Brief explanation of the drawing]
[0005] [Figure 1] A block diagram showing an exemplary cloud computing node according to one embodiment of the present invention. [Figure 2] This is an additional block diagram illustrating an exemplary cloud computing environment according to one embodiment of the present invention. [Figure 3] This is an additional block diagram showing an abstraction model layer according to one embodiment of the present invention. [Figure 4] An additional block diagram illustrating exemplary functional relationships between various aspects of the present invention. [Figure 5] This is an additional block diagram illustrating an operation in which an aspect of the present invention can be realized by providing DRAM as a second type of memory cache per application process. [Figure 6] This flowchart illustrates an exemplary method for providing a region of a first type of memory as a cache of a second type of memory for each application process, according to an aspect of the present invention. [Figure 7] This flowchart illustrates an exemplary method for providing a region of DRAM in a computing system for use as a cache of a second type of memory for each application process, in which aspects of the present invention can be realized. [Modes for carrying out the invention]
[0006] A computing environment may include main storage (sometimes called "main memory") and secondary storage. Main storage is considered to be faster access storage compared to secondary storage, such as storage-class memory. Furthermore, the addressing of main storage is considered simpler than that of storage-class memory. Storage-class memory, which is an external storage space outside of classic main storage, provides faster access and persistence than direct-access storage. Storage-class memory may be implemented as a group of solid devices connected to the computing system via multiple input / output (I / O) adapters, which are used to map the technology of the I / O devices to the memory bus of the central processing unit.
[0007] Furthermore, memory devices are used in a variety of applications, including computer systems. Computer systems and other electronic devices that include microprocessors or similar devices commonly include system memory, which is typically implemented using DRAM.
[0008] Furthermore, processor performance and memory performance are improving at different rates. As processor performance and memory capacity increase, so do virtual machines and virtual containers. Depending on the memory mechanism, some have slow access times, and therefore the memory capacity of computing systems is often balanced between fast-access volatile memory and slow-access persistent storage. Some existing systems use storage class or persistent memory in combination with DRAM. Some systems with storage class memory and DRAM of various speeds use memory buffers that accommodate the different speeds of the DRAM to overcome the slowness of the included DRAM. Some systems use local and remote DRAM that are communicated to different processors, sometimes called "memory inception."
[0009] Since this type of system is implemented in this way, it is desirable to use a region of the first type of memory (e.g., DRAM) as a cache for the second type of memory for each application process. The second type of memory may be DRAM that is slower, has a larger capacity, is cheaper, or a combination of these. Alternatively, the second type of memory may be other types of memory, such as non-volatile memory (NVM) flash, phase-change material (PCM), or other emerging memory. In one embodiment, emerging memory has a larger capacity but is generally slower than DRAM (in terms of latency and bandwidth). Therefore, in this invention, a region of DRAM is used as a cache for the second type of memory for each application process.
[0010] In some implementations, a computing system may identify a first type of memory and a second type of memory. The second type of memory has a larger storage capacity than the first type of memory, but is slower. Multiple application processes running within the computing system may be identified. A private area of the first type of memory may be provided by the processor to each of the application processes as a cache of the second type of memory. One of the private areas of the first type of memory may be dynamically made available or unavailable by the processor to each of the multiple application processes when each of the multiple application processes becomes active or inactive. The size of one of the private areas in the first type of memory may be changed for each of the multiple application processes according to one or more performance requirements for each of the multiple application processes.
[0011] In other embodiments, the present invention provides a region of a first type of memory (e.g., DRAM) that can be configured as a cache for an SCM (Storage Class Memory; the second type of memory assuming that the system DRAM and the second type of memory are on separate memory channels (e.g., Open Coherent Accelerator Processor interface "OCAPI") or on memory buffers accompanying both the DRAM and the SCM or on both). In one aspect, the DRAM cache of the SCM may operate at a similar cost and capacity to the second type of memory (e.g., a 4 terabyte "TB" DRAM cache backed up by a 32 TB SCM (with latency and bandwidth close to that of 32 TB of DRAM on average)).
[0012] This disclosure includes a detailed description of cloud computing, but the implementations of the teachings described herein are not limited to cloud computing environments. Rather, embodiments of the present invention can be implemented in any other type of computer environment that is currently known or may be developed in the future.
[0013] Cloud computing is a service delivery model that enables convenient, on-demand network access to a shared pool of configurable computing resources (e.g., networks, network bandwidth, servers, processing, memory, storage, applications, virtual machines, and services) that can be rapidly provisioned and released with minimal administrative effort or interaction with service providers. This cloud model may include at least five characteristics, at least three service models, and at least four implementation models.
[0014] The characteristics are as follows:
[0015] On-demand self-service: Cloud consumers can unilaterally prepare computing power, such as server time and network storage, automatically as needed, without requiring human interaction with service providers.
[0016] Broad network access: Computing power is available over the network and accessible through standard mechanisms. This facilitates use by heterogeneous thin or thick client platforms (e.g., mobile phones, laptops, personal digital assistants (PDAs)).
[0017] Resource pooling: A provider's computing resources are pooled and delivered to multiple consumers using a multi-tenant model. Various physical and virtual resources are dynamically allocated and reallocated as needed. Generally, consumers have a sense of location independence because they do not manage or know the exact location of the resources provided. However, consumers may be able to identify the location at a higher level of abstraction (e.g., country, state, data center).
[0018] Rapid Elasticity: Computing power can be prepared quickly and flexibly, allowing it to scale out automatically and immediately, and to be quickly released and scale in immediately. To consumers, the computing power available for preparation often appears unlimited and can be purchased in any quantity at any time.
[0019] Measured Services: Cloud systems leverage metric capabilities at a certain level of abstraction, appropriate for the type of service (e.g., storage, processing, bandwidth, active user accounts), to automatically control and optimize resource usage. Resource usage can be monitored, controlled, and reported, providing transparency to both service providers and consumers.
[0020] The service model is as follows.
[0021] Software as a Service (SaaS): The function provided to consumers is that they can use the provider's application operating on the cloud infrastructure. The application can be accessed from various client devices via a client interface such as a web browser (e.g., webmail). Consumers do not manage or control the underlying cloud infrastructure, including the network, server, operating system, storage, and even individual application functions. However, this does not apply to limited settings of user-specific application configurations.
[0022] Platform as a Service (PaaS): The function provided to consumers is to deploy the applications created or obtained by consumers on the cloud infrastructure using the programming languages and tools supported by the provider. Consumers do not manage or control the underlying cloud infrastructure including the network, server, operating system, and storage, but can control the deployed applications and, in some cases, also control the configuration of the hosting environment.
[0023] Infrastructure as a Service (IaaS): The function provided to consumers is to prepare processors, storage, networks, and other basic computing resources that allow consumers to deploy and run any software that may include an operating system and applications. Consumers do not manage or control the underlying cloud infrastructure, but can control the operating system, storage, and deployed applications, and in some cases, can partially control some network components (e.g., host firewall).
[0024] The deployment model is as follows.
[0025] Private cloud: This cloud infrastructure is operated exclusively for a specific organization. This cloud infrastructure can be managed by the organization or a third party and can exist on-premises or off-premises.
[0026] Community cloud: This cloud infrastructure is shared by multiple organizations and supports a specific community with common concerns (e.g., mission, security requirements, policies, and compliance). This cloud infrastructure can be managed by the organization or a third party and can exist on-premises or off-premises.
[0027] Public cloud: This cloud infrastructure is provided to an unspecified number of people or large industry groups and is owned by an organization that sells cloud services.
[0028] Hybrid cloud: This cloud infrastructure is a combination of two or more cloud models (private, community, or public). Each model retains its own unique entity but is bound by standard or individual technologies to achieve data and application portability (e.g., cloud bursting for load distribution between clouds).
[0029] The cloud computing environment is a service-oriented type environment that emphasizes statelessness, low coupling, modularity, and semantic interoperability. At the core of cloud computing is an infrastructure that includes a network of interconnected nodes.
[0030] Figure 1 schematically shows an example of a cloud computing node. Note that cloud computing node 10 is merely an example of a suitable cloud computing node and does not imply any limitation on the use or functionality of the embodiments of this disclosure described herein. In any case, cloud computing node 10 can be implemented, perform any of the functions described above, or both.
[0031] Cloud computing nodes 10 contain computer systems / servers 12, which can operate with a number of other general-purpose or dedicated computing system environments or configurations. Examples of well-known computing systems, environments, configurations, or combinations suitable for use with computer systems / servers 12 include personal computer systems, server computer systems, thin clients, thick clients, handheld or laptop devices, multiprocessor systems, microprocessor-based systems, set-top boxes, programmable consumer electronics, network PCs, minicomputer systems, mainframe computer systems, and distributed cloud computing environments that include any of the above systems or devices.
[0032] The computer system / server 12 can be described in general terms in relation to computer system executable instructions, such as program modules executed by the computer system. Generally, program modules can include routines, programs, objects, components, logic, data structures, etc., that perform a specific task or implement a specific data type. The computer system / server 12 can be implemented in a distributed cloud computing environment where tasks are executed by remote processing units linked via a communication network. In a distributed cloud computing environment, program modules can be stored in both local and remote computer system storage media, including memory storage devices.
[0033] As shown in Figure 1, the computer system / server 12 in the cloud computing node 10 is shown as a general-purpose computer device. Examples of components of the computer system / server 12 include one or more processors or processing units 16, system memory 28, and a bus 18 that connects various system components, including the system memory 28, to the processor 16.
[0034] Bus 18 represents one or more of several types of bus structures, including memory buses or memory controllers using various bus architectures, peripheral buses, accelerated graphics ports (AGP), and processor or local buses. Examples of such architectures include the Industry Standard Architecture (ISA) bus, Microchannel Architecture (MCA) bus, Expansion ISA (EISA) bus, Video Electronics Standards Association (VESA) local bus, and Peripheral Component Interconnect (PCI) bus.
[0035] The computer system / server 12 generally includes various computer system-readable media. Such media may be any available media accessible by the computer system / server 12 and may include both volatile and non-volatile media, as well as both removable and non-removable media.
[0036] The system memory 28 may include computer system-readable media as volatile memory, such as RAM 30, cache memory 32, or both. The computer system / server 12 may further include other removable / non-removable computer system-readable media and volatile / non-volatile computer system-readable media. As an example, the storage system 34 may be provided for reading and writing to a non-removable non-volatile magnetic medium (not shown; commonly referred to as a “hard drive”). Also, although not shown, a magnetic disk drive for reading and writing to removable non-volatile magnetic disks (e.g., floppy disks) and an optical disk drive for reading and writing to removable non-volatile optical disks (such as CD-ROMs, DVD-ROMs, or other optical media) may be provided. In these examples, each may be connected to the bus 18 by one or more data medium interfaces. As further illustrated and described below, the system memory 28 may include at least one program product having a set (e.g., at least one) of program modules configured to perform the functions of embodiments of the present invention.
[0037] As an example, a program / utility 40 having a set (at least one) of program modules 42 can be stored in memory 28, as well as an operating system, one or more application programs, other program modules, and program data. Each of the operating system, one or more application programs, other program modules, and program data, or several combinations thereof, may include an implementation of a network environment. The program modules 42 generally perform functions or methods, or both, of the embodiments of the present invention described herein.
[0038] Furthermore, the computer system / server 12 can communicate with one or more external devices 14 such as a keyboard, pointing device, or display 24, one or more devices that enable interaction between the user and the computer system / server 12, or any device that enables communication between the computer system / server 12 and one or more other computer devices (e.g., a network card or modem), or a combination thereof. Such communication can be performed via the input / output interface 22. In addition, the computer system / server 12 can communicate with one or more networks (such as a local area network (LAN), a general-purpose wide area network (WAN), or a public network (e.g., the Internet), or a combination thereof) via the network adapter 20. As shown in the figure, the network adapter 20 can communicate with other components of the computer system 12 via the bus 18. Although not shown in the figure, other hardware components, software components, or both can be used in conjunction with the computer system 12. Examples of these include microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data archive storage systems.
[0039] Referring to Figure 2, an exemplary cloud computing environment 50 is shown. As illustrated, the cloud computing environment 50 includes one or more cloud computing nodes 10. Local computer devices used by cloud consumers (e.g., personal digital assistants (PDAs) or mobile phones 54A, desktop computers 54B, laptop computers 54C, or automotive computer systems 54N, or a combination thereof) can communicate with these nodes. The nodes 10 can communicate with each other. The nodes 10 can be grouped physically or virtually (not shown) in one or more networks, such as the private, community, public, or hybrid clouds or a combination thereof. This allows the cloud computing environment 50 to provide infrastructure, platforms, or software as a service, or a combination thereof, without requiring cloud consumers to maintain resources on their local computer devices. Note that the types of computer devices 54A-N shown in Figure 2 are illustrative only, and it should be understood that the computing nodes 10 and the cloud computing environment 50 can communicate with any type of electronic device via any type of network or network addressable connection (e.g., using a web browser) or both.
[0040] Referring to Figure 3, a set of functional abstraction layers provided by the cloud computing environment 50 (Figure 2) is shown. It should be understood that the components, layers, and functions shown in Figure 3 are illustrative only, and embodiments of the present invention are not limited to these. As illustrated, the following layers and corresponding functions are provided.
[0041] Device Layer 55 includes physical or virtual devices, or both, that incorporate or stand-alone electronic devices, sensors, actuators, and other objects to perform various tasks in the cloud computing environment 50. Each device in Device Layer 55 incorporates networking capabilities to other functional abstraction layers, such that information obtained from the device is provided thereto, or information from other abstraction layers is provided to the device, or both. In one embodiment, various devices including Device Layer 55 may incorporate a network of entities collectively known as the “Internet of Things” (IoT). Such a network of entities enables the communication, collection, and transmission of data for a wide variety of purposes, as will be understood by those skilled in the art.
[0042] The illustrated device layer 55 includes, as shown, a sensor 52, an actuator 53, a “learning” thermostat 56 integrating processing, sensor, and networking electronics, a camera 57, a controllable household outlet / socket 58, and a controllable electrical switch 59. Other possible devices include, but are not limited to, a variety of additional sensor devices, networking devices, electronic devices (such as remote control devices), additional actuator devices, so-called “smart” home appliances such as refrigerators and washing machines / dryers, and a wide variety of other possible interconnected objects.
[0043] The hardware and software layer 60 includes hardware components and software components. Examples of hardware components include a mainframe 61, a reduced instruction set computer (RISC) architecture-based server 62, server 63, blade server 64, storage 65, and a network and network components 66. In some embodiments, the software components include network application server software 67 and database software 68.
[0044] The virtualization layer 70 provides an abstraction layer. From this layer, for example, the following virtual entities can be provided: virtual servers 71, virtual storage 72, virtual networks 73 including virtual private networks, virtual applications and operating systems 74, and virtual clients 75.
[0045] As an example, the management layer 80 can provide the following functions: Resource preparation 81 enables the dynamic procurement of computing resources and other resources used to perform tasks within the cloud computing environment. Metering and pricing 82 enables cost tracking as resources are used within the cloud computing environment and billing or invoicing for the consumption of these resources. As an example, these resources may include licenses for application software. Security enables not only protection of data and other resources, but also identification and verification of cloud consumers and tasks. The user portal 83 provides consumers and system administrators with access to the cloud computing environment. Service level management 84 enables the allocation and management of cloud computing resources to ensure that requested service levels are met. Service Level Agreement (SLA) planning and execution 85 enables the pre-arrangement and procurement of cloud computing resources that are expected to be needed in the future in accordance with the SLA.
[0046] The workload layer 90 provides examples of the capabilities available to the cloud computing environment. Examples of workloads and capabilities that can be provided from this layer include mapping and navigation 91, software development and lifecycle management 92, virtual classroom education delivery 93, data analysis processing 94, transaction processing 95, and workloads and capabilities 96 for providing a region of the first type of memory as a cache of a second type of memory per application process in the computing environment. Furthermore, the workloads and capabilities 96 for providing a region of the first type of memory as a cache of a second type of memory per application process in the computing environment may include operations such as data analysis (including data collection and processing from various environmental sensors) or analytical operations or both. Those skilled in the art will understand that the workloads and capabilities 96 for providing a region of the first type of memory as a cache of a second type of memory per application process in the computing environment can also work in conjunction with other parts of various abstraction layers, such as those in hardware and software 60, virtualization 70, management 80, and other workloads 90 (e.g., data analysis processing 94) to achieve various objectives of the illustrated embodiments of the present invention.
[0047] As previously stated, the present invention provides a novel solution for providing a region of a first type of memory as a cache of a second type of memory per application process in a computing environment. In operation, existing DRAM may be accessed and utilized. The region of the first type of memory may be configured as a cache per software process. A set of address ranges stored in the cache controller defines a region of local first type of memory (e.g., local DRAM or "fast memory" or "low latency" memory) as a per-process cache. In additional implementations, another set of address ranges stored in the cache controller defines a region in the second type of memory (e.g., slow memory or high latency memory) as a second type region per process.
[0048] In operation, when an application process (e.g., a software process) becomes active, the operating system ("OS") makes the application process executable / running on the processor core. Furthermore, the OS may load each address range of the process into the cache controller. Thus, the cache size of the defined regions in the first type of memory region and the second type of memory region is unique to each software process and may be adjusted according to the application's performance based on the cache region (e.g., increased or decreased based on performance, the benefits derived from the cache, or both).
[0049] Thus, the caches in the first and second types of memory for each individual process remain active for the duration of its activity (e.g., while the application process is active). When an active application process becomes inactive, the addresses of each cache region in the first type of memory are cleared from the cache controller. However, if the capacity of the first type of memory is sufficient, the cache contents of the process may remain in the first type of memory. When two or more application processes are active, each cache region is private to that process and cannot be replaced by the contents of the region of another active process.
[0050] In some implementations, a computing system may identify a first type of memory and a second type of memory. The second type of memory has a larger storage capacity than the first type of memory, but is slower. Multiple application processes running within the computing system may be identified. A private area of the first type of memory may be provided by the processor to each of the application processes as a cache of the second type of memory. One of the private areas in the first type of memory may be dynamically made available or unavailable by the processor to each of the multiple application processes when each of the multiple application processes becomes active or inactive. The size of one of the private areas in the first type of memory may be changed for each of the multiple application processes according to one or more performance requirements for each of the multiple application processes.
[0051] Next, turning to Figure 4, we see a block diagram illustrating exemplary functional components of System 400 for providing a region of first type memory as a cache of second type memory per application process in a computing environment, according to various mechanisms of the illustrated embodiment. In one embodiment, one or more of the components, modules, services, applications, or functions or combinations thereof described in Figures 1-3 may be used in Figure 4. As can be seen from this, many of the functional blocks may be considered "modules" or "components" of a function in the same descriptive sense as described earlier in Figures 1-3.
[0052] As those skilled in the art will understand, the depiction of the various functional units within System 400 is for illustrative purposes only, as the functional units may be located inside or outside the computer system / server 12 in Figure 1, or elsewhere within or between distributed computing components, or both.
[0053] In one embodiment, system 400 may provide virtualized computing services (i.e., virtualized computing, virtualized storage, virtualized networking, etc.). More specifically, system 400 may provide virtualized computing, virtualized storage, virtualized networking, and other virtualization services that run on a hardware board.
[0054] An orchestration service 410 (e.g., a dynamic scheduling agent) is shown, incorporating a processing unit 420 ("processor") for performing various calculations, data processing, and other functions according to various aspects of the present invention. In one embodiment, the processor 420 and memory 430 may be internal to or external to the orchestration service 410, or both, and may be internal to or external to the computing system / server 12. The orchestration service 410 may be included in the computer system / server 12, or external to the computer system / server 12, or both, as shown in Figure 1. The processing unit 420 may communicate with memory 430. The orchestration service 410 may include a configuration component 440, a cache controller component 450, and a range component 460.
[0055] In some implementations, the orchestration service 410 may use the configuration component 440, the cache controller component 450, the range component 460, and the storage management component 470 to identify that an application process is running on the computing system and to provide a region of the first type of memory as a cache for a second type of memory for the application process.
[0056] In other implementations, the orchestration service 410 may use a configuration component 440, a cache controller component 450, a range component 460, and a storage management component 470 to identify a first type of memory and a second type of memory in the computing system. The orchestration service 410 may also use the configuration component 440, the cache controller component 450, the range component 460, and the storage management component 470 to identify a plurality of application processes running in the computing system, and the processor may provide each of the plurality of application processes with a private area of the first type of memory as a cache of the second type of memory, and the processor may dynamically make one of the private areas in the first type of memory available or unavailable to each of the plurality of application processes when each of the plurality of application processes becomes active or inactive, and may dynamically change the size of one of the private areas in the first type of memory for each of the plurality of application processes according to one or more performance requirements for each of the plurality of application processes.
[0057] Component 440 may define a first set of address ranges stored in a range of registers in the cache controller to serve as a private area of a first type of memory per application process. Component 440 may define a second set of address ranges stored in a range of registers in the cache controller to serve as a second private area of a second type of memory for application processes.
[0058] Component 440 is a cache controller. componentIn conjunction with 450 and range component 460, access to the private area of the first type of memory may be provided to application processes whose private area is configured as a cache of the second type of memory. Component 440 is a cache controller component In relation to 450 and range component 460, the second private area may provide access to a second private area in the second type of memory for the application process, which is an additional cache of the second type of memory for the application process.
[0059] Cache controller component The 450 can coordinate access by the cache controller to the application process between the private area of the first type of memory and the second private area within the second type of memory.
[0060] Cache controller component 450 may optionally adopt component 440 and adjust the size of the private area of the first type of memory based on the performance of the application process using the private area, increasing or decreasing the size.
[0061] The cache controller component 450 may identify that an application process is active if it uses a first type of memory and a second type of memory. The cache controller component 450 may load and store a first set of address ranges stored in the cache controller to serve as a private area of the first type of memory for the active application process, and alternative address ranges for each inactive application process are removed from the cache controller.
[0062] In one embodiment, providing a private area of a first type of memory as a cache of a second type of memory for each application process includes identifying a cache line having a cache address in the first type of memory (e.g., local DRAM). The cache line is compressed within the first type of memory, generating a compressed cache line and an open memory space within the cache line. A cache tag is generated in the open memory space, and a validation value is generated in the open memory space for the compressed cache line. A cache hit is determined for the cache line based on the cache address, cache tag, and validation value.
[0063] In some implementations, a cache line has a cache address in a first type of memory. The DRAM cache or local DRAM may be a first type of memory associated with or accessible by one or more processors in the computing system. In some embodiments, the local DRAM is a cache of a set of non-uniform memory access (NUMA) DRAM, storage class memory (SCM), or variable-speed DRAM. The NUMA DRAM, SCM, variable-speed DRAM, backing storage, and second storage may be a second type of memory associated with or accessible by one or more processors in the computing system. In some embodiments, the DRAM cache is a cache of 1 terabyte or more. The backing storage, second storage, or other memory components may be 64 terabytes or more. In some examples, the local address space of one or more processors in the computer system is divided into a first region and a second region. The first type of memory (e.g., the DRAM cache) may be mapped to the first region. The second type of memory (e.g., the SCM or backing storage) may be mapped to the second region. Cache lines may be compressed within a first type of memory. Compressing a cache line generates a compressed cache line.
[0064] In one embodiment, the compression may be light compression. Light compression may compress the cache line address information and cache line data together in local DRAM. In some embodiments, cache line compression creates an open memory space in local DRAM. Light compression of a cache line may create an open memory space as a 1-byte open space. In such a case, the cache line may initially be a 128-byte cache line and lightly compressed to become a 127-byte compressed cache line. In some embodiments, the open memory space is contiguous with the compressed cache line in local DRAM. The compressed or encoded cache line may pass transparently through processor nesting or other memory architectures to simplify the implementation of the disclosure. In some embodiments, the cache line is a set of cache lines. The set of cache lines may be stored in a first type of memory. In such a case, the compressibility of each cache line in the set of cache lines may be determined in order to compress the set of cache lines. The compression component may determine that a first subset of the cache lines in the set of cache lines is compressible. In such a case, the compression component 120 compresses each line of a first subset of cache lines in the local DRAM (e.g., a first region of the processor's local address space). Compressing each line of the first subset of cache lines may generate a set of compressed cache lines and a set of open memory spaces.
[0065] In some implementations, a second subset of cache lines in a set of cache lines is not compressible. In such cases, the instruction for the second subset of cache lines is passed to the cache controller component 450, which can store the second subset of cache lines in backing memory, which is one of the following: remote DRAM, SCM, or memory inception. Cache tags may be generated in the open memory space of the compressed cache lines. The tag may be a 1-byte tag, or a tag configured to be stored in the 1-byte space freed by the compression of the cache lines. While a specified size or amount of space is mentioned, it should be understood that the tag may be any appropriate size sufficient to be stored in the open memory space. In some examples, the tag is a log2(SCMsize / DRAMsize) bit cache tag. In such cases, the base of the logarithm is 2. For example, with 1 terabyte of DRAM and 64 terabytes of SCM, log2(64) is 6 bits, so the tag will be 6 bits. In embodiments where a compressed cache line is one of a set of compressed cache lines, a separate cache tag may be generated for each compressed cache line in the set. In some embodiments, metadata may be generated and stored next to the cache tag in the same open space created by the compression. The metadata may include the physical address of the cache line in remote memory. The metadata may also include security keys, substitution information, prefetch status, combinations thereof, or other information that enables retrieval of data related to the cache line.
[0066] The verification value is generated in the open memory space of the compressed cache line. The verification value may be 1 bit. In some embodiments, the verification value indicates whether the cache line is held in the DRAM cache or in the SCM or backing memory. In embodiments where the verification value is generated for a set of compressed cache lines, the verification value may be generated for each compressed cache line in the set of compressed cache lines. In some embodiments, the verification value generated for each compressed cache line in a set of compressed cache lines (e.g., a first subset of cache lines) may be a first verification value. The first verification value may indicate that the compressed cache line and cache tag are present in the local DRAM. In such cases, the verification value may be generated as 1, a constant value, or any other appropriate value.
[0067] In embodiments where a second subset of cache lines is determined to be incompressible and therefore stored in backing memory, a second verification value may be generated for each cache line in the second subset of cache lines. In some examples, while the second subset of incompressible cache lines is stored in a second region of the local address space, a verification value indicating that a cache line is invalid may be stored in the first region to indicate that each cache line is stored in the second region. The second verification value may be stored in local DRAM. The second verification value may be generated as zero, a constant value different from the first verification value, or any other appropriate value. A verification value of zero indicates that the local DRAM does not contain a cache line and that the line is stored in backing memory.
[0068] In some embodiments, if a second subset of cache lines is determined to be incompressible and stored in backing memory (e.g., a second region of the processor's local address space), a fixed tag bit pattern or fixed data pattern may be generated in the first type of memory. A direct-mapped cache may be generated. The cache size may be increased up to the range of the capacity of the first type of memory. In some examples, the cache size may be increased up to the range of the capacity of the first type of memory to overcome a lack of associativity in the first type of memory cache.
[0069] In some embodiments, the cache may be known by an address transmitter and cache controller / decoder south of L2. The cache controller component 450 can identify, access, retrieve cache lines and respond to cache requests. Cache requests may be cache reads or cache writes. A cache read operation may determine the existence of a cache or cache line. Determining the existence of a cache or cache line may determine whether a cache hit or cache miss has occurred for the cache line. A cache write may compress a cache line and store the tag and verification value in the free space created by compressing the cache line. If the sector cannot be compressed, the cache write operation may generate a zero verification value in the free or open space and write the cache line to the next level of memory instead of the DRAM cache.
[0070] In some embodiments, the cache controller component 450 determines a cache hit for a cache line. A cache hit may be determined based on a cache address, a cache tag, and a validation value. Once a cache hit is determined, the cache controller component 450 may retrieve the cache line from the first type of memory.
[0071] In some embodiments, the cache controller component 450 determines a cache miss for a cache line. In such cases, the cache line may be stored in backing memory based on the fact that the cache line is not compressible. The cache controller component 450 may determine a cache miss based on a second verification value set to zero in local DRAM, indicating that the cache line is stored in backing memory and is not in a first type of memory.
[0072] In some embodiments, the cache controller component 450 determines a cache miss for a cache line based on a fixed tag bit pattern or a fixed data pattern. If a pattern exists, the cache contents are considered invalid, causing the cache controller component 450 to return a cache miss. When a cache miss occurs, the cache controller component 450 uses memory inception, passing instructions through one or more processors to retrieve the cache line from non-local DRAM, SCM, backing memory, or other memory components or modules other than local DRAM. In some embodiments, the cache controller component 450 accesses the SCM, backing memory, slow DRAM, or other memory components or modules to retrieve the cache line affected by the cache miss in the first type of memory. In such cases, the cache controller component 450 directly accesses the memory component or module through one or more processors communicatively coupled to the first type of memory.
[0073] The cache controller component 450 may perform the operations described above without a directory. Similarly, the operation of system 400 may be performed without an on-chip or off-chip array. The cache controller component 450 may be a simple, hardware-efficient cache controller. The cache controller component 450 may use a portion of the local DRAM as a cache for the SCM and inception cluster memory.
[0074] In some embodiments, if the remote memory, SCM, or backing memory is discontinuous or page-based, the remote address of the data may be mapped to a cache address in the local DRAM cache. In such cases, the DRAM cache controller may access the DRAM cache, but other functions may be prevented from accessing the DRAM cache. For example, if a dataset is mapped from remote memory to a portion of the local DRAM address CSSSSSS, then 0SSSSSS may be SCM backing storage mapped from the remote address to a thread of a page table entry (PTE), and 1000000 is in the DRAM cache. The information in the DRAM cache may not be mapped to a thread of the PTE remote address. The processor's cache coherence logic may operate with the address in the 0SSSSSS format while the address 1000000 is not cached. Reads and pushes to 1000000 may be directed to the remote memory. In some examples, the DRAM cache may have cache lines installed from functions or components from L3 cache writes. Both the DRAM cache and backing storage may be updated with cache lines. In return, the DRAM cache controller works with one or more components of the cache system to compress the cache line and push it to an address (e.g., 1000000) with a cache tag (e.g., Tag=SSSSSS). The raw cache line may be pushed to an address (e.g., 0SSSSSS) on page 10 / 30 of the backing storage of P202100205US01. If the line originally came from the DRAM cache, the cache line may be compressed and pushed to an address with a cache tag that has a bit that is not set by an L3 cache write. If the line came from the DRAM cache, this bit may be designed in L2 and set to 1.If the line does not originate from the DRAM cache, the bit may be reset to zero. In the event of conflicting cache writes, the last cache write may prevail or overwrite the previous cache write.
[0075] In some examples, the cache controller component 450 (which may function as a lookup component) responds to a cache read request by attempting to read the cache line. If the cache controller component 450 detects a cache miss, it may issue a read to the backing storage, 0SSSSSS. If readable and no intervention is needed, the cache controller component 450 may read the cache line from the DRAM cache by substituting 1000000 for the remote address. To reduce latency, the cache controller component 450 may initiate a cache read from the DRAM cache earlier. If an intervention or cache miss occurs while attempting to read from the DRAM cache, the data may be discarded. When reading from the DRAM cache, the cache controller component 450 may check the tag of the returned cache line. If the tag is a valid cache tag and the validation value is deemed valid, the cache controller component 450 determines a cache hit and decompresses the cache line. The cache line is then transferred to the processor's load-store unit (LSU) and placed in the L2 cache. If the cache tag or validation value is not considered valid, the cache controller component 450 determines a cache miss. The cache controller component 450 may read a cache line from the SCM or backing memory.
[0076] For further explanation, Figure 5 is an additional block diagram illustrating the operation of providing dynamic random access memory ("DRAM") as a second type of memory cache per application process in which embodiments of the present invention can be realized. In one embodiment, one or more of the components, modules, services, applications, or functions, or combinations thereof, described in Figures 1-4 may be used in Figure 5. Repeated descriptions of similar elements, components, modules, services, applications, or functions, or combinations thereof, employed in other embodiments described herein are omitted for brevity.
[0077] As illustrated, the processor 510 (e.g., processor device / chip) may include a cache controller 512. The processor 510 communicates with a first type of memory 520 (e.g., DRAM) and a second type of memory 530. The second type of memory 530 may be slow DRAM, SCM, NVM flash, PCM, and other emerging memory. In one embodiment, the second type of memory has a larger capacity but is slower (in terms of latency and bandwidth) than the first type of memory 520 (e.g., DRAM). The majority of the data is stored in the second type of memory 530. In one embodiment, cache hits go to the first type of memory 520 and cache misses go to the second type of memory 530.
[0078] The processor 510 may also include a cache controller 512 having one or more range registers, such as range registers 514A and 514B. The first type of memory 520 may include one or more regions within the first type of memory 520 that constitute a cache of a second type of memory for application processes, such as region 522 for application process A and region 524 for process B. That is, region 522 is a first set of address ranges stored in a range of registers in the cache controller 512, such as range register 514A (e.g., range register 1), so as to function as region 522 of the first type of memory 520 for application process A. Region 524 is another set of address ranges stored in a range of registers in the cache controller 512, such as range register 514A (e.g., range register 1), so as to function as region 524 of the first type of memory 520 for application process B. Thus, each “range” of the first type of memory 520 configured as a cache of a second type of memory is per application process.
[0079] The cache controller 512 may also include one or more alternative sets of address ranges stored in a range of registers, such as range register 514B (e.g., range register 2), which may function as memory regions of a second type of memory 530 for application processes, such as application process A and application process B (e.g., region 532 for application process A and region 536 for application process B).
[0080] In terms of operation, consider the following: Assume that application process A becomes active. When application process A becomes active, region 522 is defined and configured for application process A. Region 522 is a first set of address ranges stored in a range of registers in the cache controller 512, for example, range register 514A (e.g., range register 1), so that it functions as region 522 of a first type of memory 520 for the application process. The location (e.g., the set of address ranges) of region 522 of the first type of memory 520 is stored in the cache controller 512.
[0081] The cache controller 512 may also allocate a second type of memory 530 for an application process, such as application process A. The allocation may include storing a location (e.g., a set of address ranges) of the second type of memory 530 within the cache controller 512, such as in a range register 514B (e.g., range register 2), to function as a memory region (e.g., region 532 for application process A).
[0082] During execution, application process A can perform load and store data operations, which primarily involve loading and storing data in area 522 of the first type of memory 520. However, if a cache miss occurs for area 522 of the first type of memory 520, the cache controller 512 uses a range register 514B (e.g., range register 2) to act as a memory area (e.g., area 532 of application process A) to indicate that the data is stored in the second type of memory 530. The cache controller 512 allows application process A to load or store data from an area allocated for application process A, such as memory area 532.
[0083] It should be noted that while application process A is active, application B is inactive. Therefore, the cache controller 512 may remove the address range of memory range 524 of the inactive application process B from the cache controller 512, but the address range of memory range 536 is stored in the second type of memory 530. When application process B becomes active, the address range of memory range 524 is reloaded and stored in the cache controller 512.
[0084] Next, turning to Figure 6, we see a method 600 for a computing environment in which a portion of a first type of memory functions as a cache for a second type of memory, in which various aspects of the illustrated embodiment may be implemented. Functionality 600 may be implemented as an instruction executed on a machine, the instruction contained on at least one computer-readable medium or a non-transient machine-readable storage medium. Functionality 600 can begin in block 602.
[0085] As in block 604, the application process is identified as running on the computing system. For example, the application process starts execution on a processor core. Such activity may incentivize the processor core to send a signal to the cache controller indicating that the application process is active. As in block 606, a region of the first type of memory may be provided as a cache of the second type of memory for the application process. The region may be configured to function as a cache similar to the cache of the second type of memory. Functionality 600 may end in block 610.
[0086] In one embodiment, in combination with or as part of at least one block in Figure 6, or both, the operation of method 600 may include each of the following: The operation of functionality 600 may define a first set of address ranges stored in a range of registers in the cache controller to function as a private area of a first type of memory per application process. The operation of functionality 600 may define a second set of address ranges stored in a range of registers in the cache controller to function as a second private area of a second type of memory for an application process.
[0087] The operation of functionality 600 may provide access to a private area of a first type of memory for an application process, wherein the private area is configured as a cache for a second type of memory. The operation of functionality 600 may also provide access to a second private area in the second type of memory for an application process, wherein the second private area is an additional cache for the second type of memory for the application process.
[0088] The operation of Function 600 may involve the cache controller coordinating access to a private area in a first type of memory and a second private area in a second type of memory for application processes. The operation of Function 600 may also involve adjusting the size of the private area in the first type of memory based on the performance of the application processes using the private area, increasing or decreasing its size. The operation of Function 600 may identify that an application process is active if an active application process uses the first type of memory and the second type of memory. The operation of Function 600 may load and store a first set of address ranges stored in the cache controller to serve as a private area in the first type of memory for active application processes, where alternative address ranges for each inactive application process are removed from the cache controller and stored in the first type of memory.
[0089] The operation of Function 600 may partition a cache line into a set of quad-word sectors to provide a region of the first type of memory as a cache of the second type of memory for each application process (for example, a private region in the first type of memory acts / functions as a cache of the second type of memory). In some embodiments, the cache line is a 128-byte cache line. The operation of Function 600 may treat the cache line as four 32-byte quad-word-sized sectors and partition the cache line to reflect the quad-word-sized sectors. In some embodiments, the operation of Function 600 partitions the cache line into octoword sectors using a processor bus unit for data transfer. The octoword sectors may include critical octowords.
[0090] The operation of functionality 600 allows a set of quadword sectors or octoword sectors to be sequentially or contiguously placed in the first type of memory to provide a region of the first type of memory as a cache of the second type of memory for each application process. In some embodiments, the cache lines, quadword sectors, or octoword sectors may be mapped from contiguous or non-contiguous locations in the second type of memory (e.g., remote memory, SCM, or backing memory) to the cache of the first type of memory. Addresses in the second type of memory may be mapped to cache addresses using any appropriate operation. For example, addresses in the second type of memory may be mapped to cache addresses as represented by CacheAddr=RemoteAddr modulo CacheSize.
[0091] In these examples, simple algebraic operations may be used to map addresses in a similar manner to directly mapped cache operations. If the cache lines, quad words, or octowords are discontinuous or page-based, a mapping table may be used to translate linear or contiguous cache addresses to actual cache addresses.
[0092] The operation of function 600 may compress each quad-word sector in a set of quad-word sectors in order to provide a region of the first type of memory as a cache of the second type of memory for each application process. Compressing a set of quad-word sectors produces a compressed set of quad-word sectors. In some embodiments, the operation of function 600 for providing a region of the first type of memory as a cache of the second type of memory for each application process may compress each quad-word sector in the cache line lightly or trivially to make open space or free space available. For example, each quad-word sector may be compressed sufficiently to create 6 bits of open space. The operation of function 600 may compress each quad-word sector as described above in order to provide a region of the first type of memory as a cache of the second type of memory for each application process.
[0093] The operation of Function 600 may similarly compress each octoword sector, where the cache line is partitioned into octowords, in order to provide a region of the first type of memory as a cache of the second type of memory for each application process. The operation of Function 600 may generate a cache tag for each compressed quadword in the set of compressed quadword sectors in order to provide a region of the first type of memory as a cache of the second type of memory for each application process. The cache tag may be a K-bit cache address tag. The operation of Function 600 may generate a validation value for each compressed quadword sector in the set of compressed quadword sectors in order to provide a region of the first type of memory as a cache of the second type of memory for each application process. If a validation value is generated, the cache tag may be a K-bit cache address tag with valid bits added. For example, the cache tag is generated as a 6-bit tag with valid bits.
[0094] The operation of Function 600 may similarly generate a cache tag and a verification value for each octoword, where the cache line is divided into octoword sectors, in order to provide a region of the first type of memory as a cache of the second type of memory for each application process. The operation of Function 600 may similarly or similarly generate a cache tag and a verification value for each compressed quadword in a manner similar to or similar to the above, in order to provide a region of the first type of memory as a cache of the second type of memory for each application process. The operation of Function 600 may perform a cache write operation in order to provide a region of the first type of memory as a cache of the second type of memory for each application process. The cache write operation may compress each sector and write the tag and verification value to the free or open space created by compressing each quadword sector. If one or more sectors cannot be compressed, the cache write operation writes a verification value of zero to the free or open space and writes the sectors of the cache line to the next level of memory, such as an SCM, instead of the DRAM cache.
[0095] The operation of functionality 600 may respond to cache requests based on a compressed set of quadword sectors in order to provide a region of first type memory as a cache of second type memory for each application process. In some embodiments, the compressed set of quadwords includes critical quadwords. The operation of functionality 600 may respond to cache requests by determining a cache hit or cache miss based on critical quadwords in the compressed set of quadword sectors in order to provide a region of first type memory as a cache of second type memory for each application process. A cache hit may occur when an address and tag match for a quadword, such as a critical quadword. Furthermore, a cache hit may occur when an address and tag match and the state is determined to be valid based on a validation value. A cache miss occurs when an address and tag do not match, or when the quadword is determined to be invalid. If a cache miss occurs, the cache line is read from the next level of memory, such as an SCM.
[0096] The operation of functionality 600 may also respond to cache requests based on critical octowords in order to provide a region of the first type of memory as a cache of the second type of memory for each application process. In some examples, octoword sectors of the cache line may be arbitrarily recorded in the DRAM cache.
[0097] Functionality 600 may first perform a cache read operation on critical quad words in order to provide a region of the first type of memory as a cache of the second type of memory for each application process. Critical quad words are read from the cache of the first type of memory. If the verification value (e.g., the valid bit) is equal to 1 or otherwise determined to be valid and the tag matches, the critical quad words are expanded and delivered to the cache controller.
[0098] The operation of functionality 600 may involve expanding important quad words to provide a region of the first type of memory as a cache of the second type of memory for each application process. The remaining quad words in the cache line may then be expanded, either subsequently or sequentially, and each expanded quad word may be retrieved.
[0099] The operation of Function 600 may be performed in a manner similar to or similar to that described above in order to provide a region of the first type of memory as a cache of the second type of memory for each application process. The operation of Function 600 may identify, create, allocate, or generate an address in the first region from an address in the second region for each quad-word sector in order to provide a region of the first type of memory as a cache of the second type of memory for each application process. In response to a cache hit, a cache line may be read from the first region by expanding the quad-word sectors of the set of quad-word sectors associated with the cache hit, and the quad-word sectors may be delivered to one or more processors. In some examples, the address is generated from an address in the second region to the first region. In response to a cache miss, a cache line is read from the second region. The first region is the processor's local address space mapped to a DRAM cache, and the second region is the processor's local address space mapped to a second memory, such as an SCM or backing memory.
[0100] Next, turning to Figure 7, we see a method 700 for a region of a first type of memory as a cache for a second type of memory in a computing environment, and various aspects of the illustrated embodiment can be implemented. Functionality 700 may be implemented as a method executed as an instruction on a machine, where the instruction resides on at least one computer-readable medium or one non-transient machine-readable storage medium. Functionality 700 may begin in block 702.
[0101] An application process may be identified as active in the computing system, as in block 704. A region of the first type of memory may be provided as a cache of the second type of memory for the application process, as in block 706. Access to a private region of the first type of memory may be provided for the application process, as in block 708. Access to a second private region in the second type of memory may be provided for the application process, as in block 710, where the second private region is an additional cache of the second type of memory for the application process. Access to the private region of the first type of memory and the second private region in the second type of memory may be coordinated for the application process by a cache controller, as in block 712. Functionality 700 may end in block 714.
[0102] The present invention may be a system, a method, a computer program product, or a combination thereof. The computer program product may include a computer-readable storage medium storing computer-readable program instructions for causing a processor to execute an aspect of the present invention.
[0103] A computer-readable storage medium can be a tangible device capable of holding and storing instructions used by an instruction execution device. A computer-readable storage medium may, but is not limited to, an electronic storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or a suitable combination thereof. More specific examples of computer-readable storage media include portable computer diskettes, hard disks, RAM, ROM, EPROM (or flash memory), SRAM, CD-ROM, DVD, memory stick, floppy disk, punch cards, or grooved raised structures, and mechanically encoded devices on which instructions are recorded, and suitable combinations thereof. The computer-readable storage medium as used herein should not be interpreted as transient signals themselves, such as radio waves or other freely propagating electromagnetic waves, electromagnetic waves propagating through waveguides or other transmission media (e.g., light pulses passing through optical fiber cables), or electrical signals transmitted through wires.
[0104] The computer-readable program instructions described herein can be downloaded from a computer-readable storage medium to each computing / processing device, or to an external computer or external storage device via a network (e.g., the Internet, a local area network, a wide area network, or a wireless network, or a combination thereof). The network consists of copper transmission cables, optical transmission fibers, wireless transmissions, routers, firewalls, switches, gateway computers, or edge servers, or a combination thereof. The network adapter card or network interface of each computing / processing device receives computer-readable program instructions from the network and transfers the computer-readable program instructions for storage on the computer-readable storage medium within each computing / processing device.
[0105] The computer-readable program instructions for performing the operation of the present invention may be assembler instructions, instruction set architecture (ISA) instructions, machine instructions, machine-dependent instructions, microcode, firmware instructions, state setting data, configuration data for integrated circuits, or source code or object code written in any combination of one or more programming languages, including object-oriented programming languages such as Smalltalk and C++ and procedural programming languages such as the C programming language or similar programming languages. The computer-readable program instructions are executable as a standalone software package, either entirely on the user's computer or partially on the user's computer. Alternatively, they may be executable partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the latter scenario, the remote computer may be connected to the user's computer via any type of network, including a local area network (LAN) or a wide area network (WAN), or to an external computer (for example, via the Internet using an Internet service provider). In some embodiments, for example, an electronic circuit including a programmable logic circuit, a field-programmable gate array (FPGA), or a programmable logic array (PLA) can execute computer-readable program instructions by personalizing them using state information of computer-readable program instructions in order to perform aspects of the present invention.
[0106] Aspects of the present invention are described herein with reference to flowcharts or block diagrams, or both, of methods, apparatus (systems), and computer program products according to embodiments of the present invention. It will be understood that each block in a flowchart or block diagram, or both, and any combination of blocks in a flowchart or block diagram, or both, can be implemented by computer-readable program instructions.
[0107] These computer-readable program instructions can be provided to a computer processor or other programmable data processing device to generate a machine, such that instructions executed via the processor of the computer or other programmable data processing device generate means for implementing functions / operations specified in one or more blocks of a flowchart or block diagram or both. These computer-readable program instructions can also be stored in a computer-readable storage medium that can be connected to a computer, a programmable data processing device, or other device or combination of devices that function in a particular way, such that the computer-readable storage medium on which the instructions are stored constitutes one of the outputs containing instructions that implement the modes of functions / operations specified in one or more blocks of a flowchart or block diagram or both.
[0108] Computer-readable program instructions, like instructions that perform a function / action specified in one or more blocks of a flowchart or block diagram or both on a computer, other programmable device, or other device, can also be loaded into a computer, other programmable data processing device, or other device and perform a series of operational steps on the computer, other programmable device, or other device to produce a computer-implemented process.
[0109] The flowcharts and block diagrams in the figures illustrate the configuration, function, and operation of executable implementations of systems, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in a flowchart or block diagram may represent a module, segment, or part of an instruction, which constitutes one or more executable instructions for implementing a specified logical function. In some alternative embodiments, the functions shown in the blocks may differ from the order shown in the figures. For example, two blocks shown consecutively may actually be executed substantially simultaneously, or the blocks may be executed in reverse order depending on the functions they relate to. It should also be noted that each block in a block diagram or flowchart diagram, or both, and any combination of blocks in a block diagram or flowchart diagram, or both, can be implemented by a special-purpose hardware-based system that performs a specified function or operation, or a combination of special-purpose hardware and computer instructions.
Claims
1. A processor-based method in a computing environment, Identifying a first type of memory and a second type of memory in a computing system, wherein the second type of memory has a larger storage capacity than the first type of memory, but is slower than the first type of memory. Identifying multiple application processes running on the aforementioned computing system, The processor provides the private area of the first type of memory as a cache of the second type of memory for each of the plurality of application processes, The processor dynamically makes one of the private areas in the first type of memory available or unavailable for one of the plurality of application processes to become active or inactive. Dynamically changing the size of one of the private areas in the first type of memory for one of the multiple application processes based on the performance requirements for that one of the multiple application processes, For each of the aforementioned multiple application processes, a second set of address ranges stored in a range of registers within the cache controller is defined to function as a second private area of the second type of memory. including, method.
2. For each of the plurality of application processes, further includes defining a first set of address ranges stored in a range of registers in the cache controller to function as a private area of the first type of memory. The method according to claim 1.
3. A processor-based method in a computing environment, Identifying a first type of memory and a second type of memory in a computing system, wherein the second type of memory has a larger storage capacity than the first type of memory, but is slower than the first type of memory. Identifying multiple application processes running on the aforementioned computing system, The processor provides the private area of the first type of memory as a cache of the second type of memory for each of the plurality of application processes, The processor dynamically makes one of the private areas in the first type of memory available or unavailable for one of the plurality of application processes to become active or inactive. Dynamically changing the size of one of the private areas in the first type of memory for one of the multiple application processes based on the performance requirements for that one of the multiple application processes, To provide each of the plurality of application processes with access to the private area of the first type of memory, wherein the private area is configured as the cache of the second type of memory, To provide each of the plurality of application processes with access to a second private area in the second type of memory, wherein the second private area is an additional cache in the second type of memory for each of the plurality of application processes, including, method.
4. The cache controller further includes coordinating access for each of the plurality of application processes to the private area of the first type of memory and the second private area in the second type of memory, The method according to claim 3.
5. The method further includes adjusting the size of the private area of the first type of memory based on the performance of each of the multiple application processes that use the private area, wherein the size increases or decreases. The method according to claim 1.
6. A processor-based method in a computing environment, Identifying a first type of memory and a second type of memory in a computing system, wherein the second type of memory has a larger storage capacity than the first type of memory, but is slower than the first type of memory. Identifying multiple application processes running on the aforementioned computing system, The processor provides the private area of the first type of memory as a cache of the second type of memory for each of the plurality of application processes, The processor dynamically makes one of the private areas in the first type of memory available or unavailable for one of the plurality of application processes to become active or inactive. Dynamically changing the size of one of the private areas in the first type of memory for one of the multiple application processes based on the performance requirements for that one of the multiple application processes, Identifying that one or more of the plurality of application processes is active, and that one or more of the plurality of application processes uses the first type of memory and the second type of memory, Loading and storing a first set of address ranges stored in a cache controller so as to serve as a private area of the first type of memory for application processes, wherein alternative address ranges for each inactive application process are removed from the cache controller; including, method.
7. A system for improving computing efficiency in a computing environment, The system includes one or more computers having executable instructions, and when an executable instruction is executed, it has the following properties: Identifying a first type of memory and a second type of memory in a computing system, wherein the second type of memory has a larger storage capacity than the first type of memory, but is slower than the first type of memory. Identifying multiple application processes running on the aforementioned computing system, The private area of the first type of memory is provided as a cache of the second type of memory for each of the multiple application processes, When one of the multiple application processes becomes active or inactive, one of the private areas in the first type of memory is made available or unavailable to each of the multiple application processes. Dynamically changing the size of one of the private areas in the first type of memory for each of the multiple application processes, according to one or more performance requirements for each of the multiple application processes. and, For each of the aforementioned multiple application processes, a second set of address ranges stored in a range of registers within the cache controller is defined to function as a second private area of the second type of memory. A system that executes an action.
8. The executable instruction, when executed, causes the system to define a first set of address ranges stored in a range of registers in the cache controller, which will function as a private area of the first type of memory for each of the plurality of application processes. The system according to claim 7.
9. A system for improving computing efficiency in a computing environment, The system includes one or more computers having executable instructions, and when an executable instruction is executed, it has the following properties: Identifying a first type of memory and a second type of memory in a computing system, wherein the second type of memory has a larger storage capacity than the first type of memory, but is slower than the first type of memory. Identifying multiple application processes running on the aforementioned computing system, The private area of the first type of memory is provided as a cache of the second type of memory for each of the multiple application processes, When one of the multiple application processes becomes active or inactive, one of the private areas in the first type of memory is made available or unavailable to each of the multiple application processes. Dynamically changing the size of one of the private areas in the first type of memory for each of the multiple application processes, according to one or more performance requirements for each of the multiple application processes. and, To provide each of the plurality of application processes with access to the private area of the first type of memory, wherein the private area is configured as the cache of the second type of memory, To provide each of the plurality of application processes with access to a second private area in the second type of memory, wherein the second private area is an additional cache in the second type of memory for each of the plurality of application processes, and to allow each of the plurality of application processes to perform the following: system.
10. A system for improving computing efficiency in a computing environment, The system includes one or more computers having executable instructions, and when an executable instruction is executed, it has the following properties: Identifying a first type of memory and a second type of memory in a computing system, wherein the second type of memory has a larger storage capacity than the first type of memory, but is slower than the first type of memory. Identifying multiple application processes running on the aforementioned computing system, The private area of the first type of memory is provided as a cache of the second type of memory for each of the multiple application processes, When one of the multiple application processes becomes active or inactive, one of the private areas in the first type of memory is made available or unavailable to each of the multiple application processes. Dynamically changing the size of one of the private areas in the first type of memory for each of the multiple application processes, according to one or more performance requirements for each of the multiple application processes. and, The cache controller causes each of the multiple application processes to coordinate access to the private area of the first type of memory and the second private area in the second type of memory. system.
11. When the executable instruction is executed, it causes the system to adjust the size of the private area of the first type of memory based on the performance of each of the multiple application processes that use the private area, and the size increases or decreases. The system according to claim 7.
12. A system for improving computing efficiency in a computing environment, The system includes one or more computers having executable instructions, and when an executable instruction is executed, it has the following properties: Identifying a first type of memory and a second type of memory in a computing system, wherein the second type of memory has a larger storage capacity than the first type of memory, but is slower than the first type of memory. Identifying multiple application processes running on the aforementioned computing system, The private area of the first type of memory is provided as a cache of the second type of memory for each of the multiple application processes, When one of the multiple application processes becomes active or inactive, one of the private areas in the first type of memory is made available or unavailable to each of the multiple application processes. Dynamically changing the size of one of the private areas in the first type of memory for each of the multiple application processes, according to one or more performance requirements for each of the multiple application processes. and, Identifying that one or more of the plurality of application processes is active, and that one or more of the plurality of application processes uses the first type of memory and the second type of memory, Loading and storing a first set of address ranges stored in a cache controller so that they serve as a private area of the first type of memory for application processes, wherein alternative address ranges for each inactive application process are removed from the cache controller, system.
13. A computer program for improving computing efficiency in a computing environment, wherein a computer Identifying a first type of memory and a second type of memory in a computing system, wherein the second type of memory has a larger storage capacity than the first type of memory, but is slower than the first type of memory. Identifying multiple application processes running on the aforementioned computing system, The private area of the first type of memory is provided as a cache of the second type of memory for each of the multiple application processes, When one of the multiple application processes becomes active or inactive, one of the private areas in the first type of memory is made available or unavailable to each of the multiple application processes. In accordance with one or more performance requirements for each of the multiple application processes, the size of one of the private areas in the first type of memory for each of the multiple application processes is dynamically changed. For each of the aforementioned multiple application processes, a second set of address ranges stored in a range of registers within the cache controller is defined to function as a second private area of the second type of memory. A computer program that executes something.
14. For each of the plurality of application processes, further perform the task of defining a first set of address ranges to be stored in a range of registers in the cache controller so as to function as a private area of the first type of memory. The computer program according to claim 13.
15. A computer program for improving computing efficiency in a computing environment, wherein a computer Identifying a first type of memory and a second type of memory in a computing system, wherein the second type of memory has a larger storage capacity than the first type of memory, but is slower than the first type of memory. Identifying multiple application processes running on the aforementioned computing system, The private area of the first type of memory is provided as a cache of the second type of memory for each of the multiple application processes, When one of the multiple application processes becomes active or inactive, one of the private areas in the first type of memory is made available or unavailable to each of the multiple application processes. In accordance with one or more performance requirements for each of the multiple application processes, the size of one of the private areas in the first type of memory for each of the multiple application processes is dynamically changed. To provide each of the plurality of application processes with access to the private area of the first type of memory, wherein the private area is configured as the cache of the second type of memory, To provide each of the plurality of application processes with access to a second private area in the second type of memory, wherein the second private area is an additional cache in the second type of memory for each of the plurality of application processes, The cache controller coordinates access for each of the multiple application processes to the private area of the first type of memory and the second private area in the second type of memory, A computer program that executes something.
16. The program further includes a program instruction to adjust the size of the private area of the first type of memory based on the performance of each of the multiple application processes that use the private area, wherein the size increases or decreases. The computer program according to claim 13.
17. A computer program for improving computing efficiency in a computing environment, wherein a computer Identifying a first type of memory and a second type of memory in a computing system, wherein the second type of memory has a larger storage capacity than the first type of memory, but is slower than the first type of memory. Identifying multiple application processes running on the aforementioned computing system, The private area of the first type of memory is provided as a cache of the second type of memory for each of the multiple application processes, When one of the multiple application processes becomes active or inactive, one of the private areas in the first type of memory is made available or unavailable to each of the multiple application processes. In accordance with one or more performance requirements for each of the multiple application processes, the size of one of the private areas in the first type of memory for each of the multiple application processes is dynamically changed. Identifying that one or more of the plurality of application processes is active, and that one or more of the plurality of application processes uses the first type of memory and the second type of memory, Loading and storing a first set of address ranges stored in a cache controller so as to serve as a private area of the first type of memory for application processes, wherein alternative address ranges for each inactive application process are removed from the cache controller; A computer program that executes something.
Citation Information
Patent Citations
Method for using cache memory
JP2005071046A
Method and apparatus for dynamically resizing cache partitions based on task execution phases
JP2009528610A
Cache memory control program, processor including cache memory, and cache memory control method
JP2015036873A