TECHNOLOGIES FOR REGION-BASED CACHING MANAGEMENT

Region-based cache management in NFV systems addresses unpredictable cache latency by allocating cache memory based on physical regions and probabilistic eviction, enhancing data handling efficiency and reducing interference between VMs and VNFs.

DE112017001654B4Active Publication Date: 2025-07-31INTEL CORP
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
DE112017001654
Authority / Receiving Office
DE · DE
Patent Type
Patents
Current Assignee / Owner
Priority Date
2016-03-31
Filing Date
2017-03-01
Publication Date
2025-07-31
Estimated Expiration
2037-03-01

AI Technical Summary

Technical Problem

In network function virtualization (NFV) deployments, cache latency varies unpredictably due to multiple virtual machines (VMs) sharing the same cache, making it difficult to minimize the impact of one VM on others, especially in data streams between virtualized network functions (VNFs).

Method used

Implement region-based cache management that allocates cache memory based on physical memory areas, dividing it into regions and assigning indicators and configuration parameters to manage cache handling differently for each region, using probabilistic methods to determine cache line eviction.

Benefits of technology

This approach reduces cache latency and improves data handling efficiency by prioritizing cache usage for frequently accessed data, minimizing interference between VMs and VNFs, thus optimizing network traffic processing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

A network computing device (106) for region-based cache management, the network computing device (106) comprising: a processor (202) having a cache memory (206); a main memory (210), different from the cache memory (206), coupled to the processor (202); and one or more data storage devices (212) having stored therein a plurality of instructions that, when executed by the processor (202), cause the network computing device (106) to: select a cache line for eviction from a plurality of cache lines of the cache memory (206); determine whether the cache line selected for eviction is located in a cache block of the cache memory (206) that is currently associated with a corresponding memory region of the main memory (210), the cache block comprising one or more of the plurality of cache lines;in response to a determination that the cache block is associated with a corresponding memory region, retrieve a bias value associated with the corresponding memory region, the bias value corresponding to a fractional probability; generate a bias comparison value for the corresponding memory region; compare the bias value of the corresponding memory region and the bias comparison value generated for the corresponding memory region; and evict the cache line in response to a determination that a result of the comparison indicates evicting the cache line.
Need to check novelty before this filing date? Find Prior Art

Description

BACKGROUND

[0001] Network operators and communications service providers typically rely on various network visualization technologies to manage complex large-scale data centers, which may include a variety of network computing devices (e.g., servers, switches, routers, etc.) to process network traffic through a data center. For example, network operators and service provider networks may rely on network function virtualization (NFV) deployments to deploy network services (e.g., firewall services, network address translation (NAT) services, deep packet inspection (DPI) services, evolved packet core (EPC) services, mobility management entity (MME) services, packet data network gateway (PGW) services, serving gateway (SGW) services, billing services, transmission control protocol (TCP) optimization services, etc.To provide scalability to meet network traffic processing requirements and reduce operational costs, virtual network functions (VNFs) are typically deployed to handle certain network function operations. Such operations typically run on one or more virtual machines (VMs) in a virtualized environment on top of the hardware network infrastructure. Data flows occurring between such VNFs (i.e., inter-VNF flows) are commonly optimized by inter-VM shared memory (IVSHMEM), which relies on cache memory to provide critical speed advantages. However, when consolidating multiple VMs, latency can vary unpredictably, which is common in NFV deployments.Accordingly, developing interface applications typically requires careful design to ensure that critical accesses are cache-supported. However, in implementations where multiple VNFs are deployed, and each VNF relies on one or more VMs sharing the same cache, minimizing the impact of one VM on the others can be difficult to achieve.

[0002] US 2008 / 0 022 048 A1 shows a device for avoiding the sharing of cache lines in virtual machines, which can be implemented in a system with a host and several guest operating systems.

[0003] US 2008 / 0 256 303 A1 shows an apparatus for processing data with a cache memory having a plurality of cache lines, each of which is operable to store a cache line of data values.

[0004] US 2009 / 0 204 764 A1 discloses a method and apparatus for cache pooling. Threads are prioritized based on their importance and tasks. The most important threads are allocated to main memory locations so that they are subject to limited or no cache contention.

[0005] US 5 915 262 A shows a computer system with a processor, a main memory and a cache memory used to mark different memory areas in order to define and select cache properties of transfers between the processor and the memory via the cache.

[0006] WO 2015 / 047 348 A1 shows cache operations for a memory-side cache, such as a byte-addressable non-volatile memory, comprising a first operation, wherein the first operation comprises deleting cache entries from the cache memory according to a replacement policy that is designed to prefer cache entries with clean cache lines over deleting cache entries with faulty cache lines.

[0007] The stated object is achieved according to the invention by the features of patent claim 1. Further embodiments of the invention are presented in the subclaims. BRIEF DESCRIPTION OF THE DRAWINGS

[0008] The concepts described herein are illustrated in the accompanying figures by way of example only and not by way of limitation. For simplicity and clarity of illustration, elements illustrated in the figures are not necessarily drawn to scale. Where considered appropriate, reference numerals are repeated among the figures to indicate corresponding or analogous elements. Fig. 1 is a simplified block diagram of at least one embodiment of a system for region-based cache management; Fig. 2 is a simplified block diagram of at least one embodiment of the network computing device of the system of Fig. 1; Fig. 3 is a simplified block diagram of at least one embodiment of an environment implemented by the network computing device of Fig. 2 can be created; Fig. 4 is a simplified block diagram of another embodiment of an environment used by the network computing device of Fig. 2 can be created; Fig. 5 is a simplified flow diagram of at least one embodiment of a method for laying out region-based cache blocks performed by the network computing device of Fig. 2 can be executed; and the Fig. 6 and Fig. 7 is a simplified flow diagram of at least one embodiment of a method for performing region-based cache line eviction performed by the network computing device of Fig. 2 can be executed. DETAILED DESCRIPTION OF THE DRAWINGS

[0009] While the concepts of the present disclosure are susceptible to numerous modifications and alternative forms, specific embodiments thereof have been shown by way of example in the drawings and will be described in detail herein. It should be understood, however, that there is no intention to limit the concepts of the present disclosure to the particular forms disclosed; on the contrary, the intention is to cover all modifications, equivalents, and alternatives consistent with the present disclosure and the appended claims.

[0010] References in the specification to "one embodiment," "an embodiment," "an illustrative embodiment," etc., indicate that the described embodiment may include a particular feature, structure, or characteristic, but each embodiment may or may not necessarily include that particular feature, structure, or characteristic. Furthermore, such language does not necessarily refer to the same embodiment. Furthermore, where a particular feature, structure, or characteristic is described in connection with one embodiment, it is also presumed that it is within the skill of one skilled in the art to effect such feature, structure, or characteristic in connection with other embodiments, whether explicitly described or not.It should also be recognized that items included in a list in the form of "at least one of A, B, and C" may mean (A); (B); (C); (A and B); (A and C); (B and C); or (A, B, and C). Similarly, items listed in the form of "at least one of A, B, or C" may mean (A); (B); (C); (A and B); (A and C); (B and C); or (A, B, and C).

[0011] The disclosed embodiments may, in some cases, be implemented in hardware, firmware, software, or any combination thereof. The disclosed embodiments may also be implemented as instructions carried by or stored on one or more transient or non-transient machine-readable (e.g., computer-readable) storage media (e.g., memory, storage, etc.) that can be read and executed by one or more processors. A machine-readable storage medium may be embodied as any storage device, mechanism, or other physical structure for storing or transmitting information in a machine-readable form (e.g., volatile or non-volatile memory, disk, or other media device).

[0012] In the drawings, some structural or methodological features may be shown in specific arrangements and / or sequences. However, it should be recognized that such specific arrangements and / or sequences may not be required. Rather, in some embodiments, such features may be arranged in a different manner and / or in a different order than shown in the illustrative figures. Additionally, the inclusion of a structural or methodological feature in a particular figure is not intended to imply that such feature is required in all embodiments, and in some embodiments, it may not be included or may be combined with other features.

[0013] Now with reference to Fig. 1, in one illustrative embodiment, a system 100 for region-based cache management includes a source endpoint node 102 communicatively coupled to a destination endpoint node 108 via a network computing device 106 of a network 104. The network computing device 106 is configured to perform various operations (e.g., services) on the network traffic (i.e., network packets, messages, etc.) received at the network computing device 106. Such network traffic may be received from the source endpoint node 102, the destination endpoint node 108, or another network computing device 106. Accordingly, while only a single network computing device 106 is shown in the network 104 of the illustrative system 100, it should be appreciated that the network 104 may include more than one network computing device 106, which may be coupled and configured in various architectures.

[0014] Illustratively, the network computing device 106 is configured to process a network packet to determine whether a network packet should be dropped or forwarded. To process the network packet, any number of services, or network functions, may be executed on the network packet, such as firewall services, network address translation (NAT) services, deep packet inspection (DPI) services, evolved packet core (EPC) services, mobility management entity (MME) services, packet data network gateway (PGW) services, serving gateway (SGW) services, billing services, transmission control protocol (TCP) optimization services, etc. In some embodiments, the network functions may be operated using one or more virtual machines (VMs) (e.g., in a service function chain), native applications, or applications running in containers (e.g., Docker).For example, the network device 106 may be configured to operate the services across a number of virtual network functions (VNFs) using one or more VMs. Therefore, it should be appreciated that in such embodiments, the network computing device 106 is configured to map virtual representations of physical components of the network computing device 106 to virtualized components of the various VMs and / or applications, so that an inter-VNF / application data stream can be managed between VMs (i.e., between VMs in a service function chain).

[0015] Unlike current technologies focused on cache control and processor core allocation (e.g., Cache Allocation Technology, CAT) that prioritize cache for VMs and / or processes, and similarly unlike current technologies focused on input / output of a data stream to and from the processor cache (e.g., Data Direct I / O (DDIO)) that prioritize cache based on devices (e.g., Peripheral Component Interconnect Express (PCI-e) network interface controllers (NICs), solid state drives (SSDs, etc.), the region-based cache management employed by the network computing device 106 prioritizes and personalizes both caching and cache coherence based on physical memory regions.

[0016] In use, the network computing device 106 adapts an existing cache strategy for a defined memory buffer corresponding to a particular portion of physical memory to favor the use of a region-based buffer per application to maintain frequently used data in a particular cache region associated with the corresponding memory buffer. To do this, the network computing device 106 is configured to allocate a portion of physical memory and connect or otherwise associate a block of cache memory with the allocated portion of physical memory or a memory buffer associated with it. The network computing device 106 then divides the allocated physical memory into multiple regions. Therefore, the block of the corresponding cache memory is also divided into corresponding regions (i.e.Cache regions), such as may be performed by associating each cache region with the respective shared portion of the allocated physical memory. The network computing device 106 further indicates to the respective hardware those portions of the shared portions of physical memory that may benefit from differentiated cache handling, such as the region-based cache management handling described herein. Accordingly, only data from an associated portion of physical memory, and its associated memory buffer, may be stored in a particular block of cache memory, or cache region.

[0017] For example, in an embodiment where network computing device 106 is to evict a cache line to load new data from memory, but network computing device 106 has determined that no memory is allocated to a memory address associated with one of the region-mapped cache regions, network computing device 106 may use only an unmapped portion of cache (i.e., shared cache) to store the data therein. In other words, if network computing device 106 determines that data to be loaded from memory is not associated with one of the region-based mapped portions of cache, network computing device 106 may load only the data from shared cache (i.e., a non-region-based mapped portion of cache) into a cache line.

[0018] Alternatively, if the network computing device 106 determines that data to be loaded from memory is associated with one of the region-based mapped portions of cache, the network computing device 106 may choose between region-mapped cache regions and shared cache. If the network computing device 106 selects to load the data into shared cache, the network computing device 106 is configured to use the non-directed eviction strategy if the shared cache cache line is to be evicted or otherwise moved. Alternatively, if the network computing device 106 selects to load the data into a region-based mapped portion of cache, the network computing device 106 is configured to use the directed eviction strategy if the shared cache cache line is to be evicted or otherwise moved.

[0019] To identify the regions of physical memory that may benefit from differentiated cache treatment, the network computing device 106 assigns a region indicator and one or more configuration parameters to each corresponding cache block region, which may be based on a cache allocation of one or more contiguous physical memory areas. The region indicator may be embodied as any type of data structure that can be used to identify a corresponding region of allocated memory. In one illustrative example, the network computing device 106 divides the allocated portion of physical memory into a first group, which is used for sharing data between VM instances in a VNF or between VNFs (e.g.,a shared region), and a second group used for transient data communications between VM producer and consumer instances in a VNF or between VNFs (e.g., a relay region).

[0020] Similarly, the network computing device 106 shares the cache blocks associated with the corresponding physical memory regions. In such an embodiment, the network computing device 106 may use an architecturally exposed mechanism to indicate which cache block region corresponds to which type of data (i.e., which type of data corresponds to which region). For example, the first group may be assigned a shared region indicator, and the second group may be assigned a relay region indicator. Accordingly, the cache block region corresponding to the shared region (i.e., the first group) may be assigned the shared region indicator, and similarly, the cache block region corresponding to the relay region (i.e., the second group) may be assigned the relay region indicator.

[0021] To continue the illustrative example, the network computing device 106 may additionally assign one or more configuration parameters to each cache block region (i.e., the cache lines associated with each cache block region). The configuration parameters may include any data that defines or otherwise characterizes a cache block region, such as the region indicators, slice addresses, slice offsets (e.g., a head offset, a tail offset, etc.), as well as any other properties associated with the cache block region. For example, the configuration parameters may include a head indicator and a tail indicator (e.g., stored as offsets in corresponding MSRs (Model-Specific Registers)) in embodiments where a particular region (e.g.,the shared region and / or the relay region) is treated by hardware and software of the network computing device 106 as if it were a ring buffer.

[0022] In some embodiments, the configuration parameters (e.g., the head and tail indicators) may be stored as offsets in corresponding MSRs. The cache characteristics additionally include a bias value (i.e., a value between zero and one) corresponding to a fractional probability that can be used to determine whether a particular cache line selected for eviction should be evicted. Accordingly, the bias value corresponding to a region of the cache block can be compared to a probabilistic outcome to determine whether the cache line should be evicted (i.e., evict the cache line or select another cache line for eviction). It should be recognized that the probabilistic outcome can be determined using any known probabilistically determinable technique, such as a directed coin toss simulation or another type of selection randomization technique.

[0023] The network computing device 106 may be embodied as any type of network traffic processing device capable of performing the functions described herein, such as, among others, a server (e.g., standalone, rack-mounted, blade, etc.), a network device (e.g., physical or virtual), a switch (e.g., rack-mounted, standalone, fully managed, partially managed, full-duplex and / or half-duplex communication mode capable, etc.), a router, an Internet device, a distributed computing system, a processor-based system, and / or a multiprocessor system. As described in Fig. 2, the illustrative network computing device 106 includes a processor 202, an input / output (I / O) subsystem 208, a main memory 210, a data storage device 212, and communication circuitry 214. Of course, the network computing device 106 may include other or additional components, such as those commonly found in other embodiments of a computing device. Additionally, in some embodiments, one or more of the illustrative components may be integrated with or otherwise form part of another component. For example, in some embodiments, the main memory 210, or portions thereof, may be integrated with the processor 202. Further, in some embodiments, one or more of the illustrative components may be omitted from the network computing device 106.

[0024] The processor 202 may be embodied as any type of processor capable of performing the functions described herein, such as, but not limited to, a single physical multi-core processor chip or package. The illustrative processor 202 includes a number of processor cores 204, each embodied as an independent logical execution unit capable of executing programmed instructions. It should be appreciated that, in some embodiments of the network computing device 106 (e.g., supercomputers), the network computing device 106 may include thousands of processor cores 204. The processor 202 may be connected to a physical connector, or socket, on a motherboard (not shown) of the network computing device 106 that is configured to receive a single physical processor package (i.e., a multi-core physical integrated circuit).

[0025] The illustrative processor 202 additionally includes a cache 206, which may be embodied as any type of cache that the processor 202 can access faster than the main memory 210, such as an on-die cache or on-processor cache. In other embodiments, the cache 206 may be an off-die cache but may be located on the same system-on-a-chip (SoC) as the processor 202. It should be appreciated that, in some embodiments, the cache 206 may have a multi-level architecture. In other words, in such multi-level architecture embodiments, the cache 206 may be embodied as, for example, an L1, L2, or L3 cache. It should further be appreciated that, in some embodiments, the network computing device 106 may include more than one processor 202.In other words, in such embodiments, the network computing device 106 may include more than one physical processor package, each of which may be connected to a motherboard (not shown) of the network computing device 106 via an individual socket, each of which may be communicatively coupled to one or more independent hardware memory locations.

[0026] Main memory 210 is communicatively coupled to processor 202 via I / O subsystem 208, which may be embodied as circuitry and / or components to facilitate input / output operations with processor 202, main memory 210, and other components of network computing device 106. For example, I / O subsystem 208 may be embodied as, or otherwise include, memory control hubs, input / output control hubs, firmware devices, communication links (i.e., point-to-point links, bus links, wires, cables, light guides, traces on printed circuit boards, etc.), and / or other components and subsystems to facilitate input / output operations.In some embodiments, the I / O subsystem 208 may form part of a system on a chip (SoC) and may be integrated along with the processor 202, the main memory 210, and other components of the network computing device 106 on a single integrated circuit chip.

[0027] The data storage device 212 may be embodied as any type of device or devices designed for the short- or long-term storage of data, such as, for example, memory devices and circuits, memory cards, hard disk drives, solid-state drives, or other data storage devices. It should be appreciated that the data storage device 212 and / or the main memory 210 (e.g., the computer-readable storage media) may store various data as described herein, including operating systems, applications, programs, libraries, drivers, instructions, etc., capable of being executed by a processor (e.g., the processor 202) of the network computing device 106.

[0028] The communication circuitry 214 may be embodied as any communication circuit, device, or collection thereof capable of facilitating communications between the network computing device 106 and other computing devices (e.g., the source endpoint node 102, the destination endpoint node 108, another network computing device 106, etc.) over a network (e.g., the network 104). The communication circuitry 214 may be configured to support any one or more communication technologies (e.g., wireless or wired communication technologies) and protocols associated therewith (e.g., Ethernet, Bluetooth ® , Wi-Fi ® , WiMAX, LTE, 5G, etc.) to effect such communication.

[0029] The illustrative communication circuitry 214 includes a network interface controller (NIC) 216. The NIC 216 may be embodied as one or more expansion cards, daughter cards, network interface cards, controller chips, chipsets, or other devices that may be used by the network computing device 106. For example, in some embodiments, the NIC 216 may be integrated with the processor 202, embodied as an expansion card coupled to the I / O subsystem 208 via an expansion bus (e.g., PCI Express), be part of an SoC that includes one or more processors, or be included in a multi-chip package that also includes one or more processors.Additionally or alternatively, in some embodiments, functionality of the NIC 216 may be integrated into one or more components of the network computing device 106 at the card level, socket level, chip level, and / or other levels.

[0030] Referring again to Fig. 1, the source endpoint node 102 and / or the destination endpoint node 108 may be embodied as any type of computing or computer device capable of performing the functions described herein, including, but not limited to, a portable computing device (e.g., smartphone, tablet, laptop, notebook, wearable, etc.) including mobile hardware (e.g., processor, memory, storage, wireless communication circuitry, etc.) and software (e.g., an operating system) to support mobile architecture and portability, a computer, a server (e.g., standalone, rack-mounted, blade, etc.), a network device (e.g., physical or virtual), an Internet device, a distributed computing system, a processor-based system, and / or a multiprocessor system. Accordingly, the source endpoint node 102 and / or the destination endpoint node 108 may include components similar to those described in Fig. 2 of the illustrative network computing device, such as a processor, an I / O subsystem, memory, a data storage device, and / or communication circuitry, which are not shown for clarity of description. Therefore, the descriptions of similar components will not be repeated herein, with the understanding that the description of the corresponding components given above with respect to the network computing device 106 equally applies to the corresponding components of the source endpoint node 102 and / or the destination endpoint node 108.

[0031] The network 104 may be implemented as any type of wired or wireless communications network, including a wireless local area network (WLAN), a wireless personal area network (WPAN), a cellular network (e.g., Global System for Mobile Communications (GSM), Long-Term Evolution (LTE), etc.), a telephony network, a digital subscriber line (DSL) network, a cable network, a local area network (LAN), a wide area network (WAN), a worldwide network (e.g., the Internet), or any combination thereof. It should be appreciated that, in such embodiments, the network 104 may serve as a centralized network and, in some embodiments, may be communicatively coupled to another network (e.g., the Internet).Accordingly, the network 104 may include a variety of other computing devices (e.g., virtual and physical routers, switches, network hubs, servers, storage devices, data processing devices, etc.) as needed to facilitate communication between the source endpoint node 102 and the destination endpoint node 108, as well as between the network computing devices 106, which are not shown for clarity of description.

[0032] It should further be appreciated that, in some embodiments, the network 104 may include additional computing devices, such as a network controller (not shown), configured to provide one or more policies (e.g., network policies, cache eviction policies, security policies, etc.) or instructions to the network computing device 106. In such embodiments, the network controller may be a separate computing device communicatively coupled to the network computing device 106 configured to operate in a software-defined network (i.e., an SDN controller) and / or a network function virtualization (NFV) environment (i.e., an NFV manager and network orchestrator (MANO)).

[0033] Now referring to Fig. 3, in one illustrative embodiment, the network computing device 106 establishes an environment 300 during operation. The illustrative environment 300 includes a communications management module 310, a network function management module 320, a main memory management module 330, and a cache management module 340. The various modules of the environment 300 may be embodied as hardware, software, firmware, or a combination thereof. Therefore, in some embodiments, one or more of the modules of the environment 300 may be embodied as circuitry or an assembly of electrical devices (e.g., a communications management circuit 310, a network function management circuit 320, a main memory management circuit 330, a cache management circuit 340, etc.).

[0034] It should be appreciated that in such embodiments, one or more of the communication management circuitry 310, the network function management circuitry 320, the main memory management circuitry 330, and the cache management circuitry 340 may form part of one or more of the processor 202, the I / O subsystem 208, the communication circuitry 214, and / or other components of the network computing device 106. Additionally, in some embodiments, one or more of the illustrative modules may form part of another module, and / or one or more of the illustrative modules may be independent of one another.Furthermore, in some embodiments, one or more of the modules of environment 300 may be embodied as virtualized hardware components or emulated architecture that may be created and maintained by processor 202 or other components of network computing device 106.

[0035] In the illustrative environment 300, the network computing device 106 further includes main memory allocation data 302, cache eviction strategy data 304, and cache characteristic data 306, each of which may be stored in the main memory 210 and / or the data storage device 212 of the network computing device 106. Furthermore, the main memory allocation data 302, the cache eviction strategy data 304, and / or the cache characteristic data 306 may be accessed by the various modules and / or submodules of the network computing device 106. Additionally, it should be appreciated that any data stored or otherwise represented in the main memory allocation data 302, the cache eviction strategy data 304, and / or the cache characteristic data 306 may not be mutually exclusive relative to one another in some embodiments.

[0036] For example, in some implementations, data stored in main memory allocation data 302 may also be stored as part of cache eviction strategy data 304 and / or vice versa. Therefore, although the data utilized by network computing device 106 is described herein as being particularly discrete, in other embodiments, such data may be combined or aggregated and / or otherwise form part of a single or multiple data sets, including duplicate copies. It should further be appreciated that network computing device 106 may include additional and / or alternative components, subcomponents, modules, submodules, and / or devices commonly found in a computing device that, for clarity of description, are not included in Fig. 3 are shown.

[0037] The communication management module 310, which may be embodied as hardware, firmware, software, virtualized hardware, emulated architecture, and / or a combination thereof as discussed above, is configured to facilitate inbound and outbound wired and / or wireless network communications (e.g., network traffic, network packets, network flows, etc.) to and from the network computing device 106. To do so, the communication management module 310 is configured to receive and process network packets from other computing devices (e.g., the source endpoint node 102, the destination endpoint node 108, and / or other computing devices communicatively coupled to the network computing device 106, such as another network computing device 106). Additionally, the communication management module 310 is configured to prepare and forward network packets to other computing devices (e.g.,the source endpoint node 102, the destination endpoint node 108, and / or other computing devices communicatively coupled to the network computing device 106, such as another network computing device 106. Accordingly, in some embodiments, at least a portion of the functionality of the communications management module 310 may be performed by the communications circuitry 214 of the network computing device 106, or more specifically, by the NIC 216 of the communications circuitry 214.

[0038] The network function management module 320, which may be embodied as hardware, firmware, software, virtualized hardware, emulated architecture, and / or a combination thereof as discussed above, is configured to manage the physical and virtual functions of the NIC and associated VMs and / or applications of the network computing device 106. Accordingly, in such embodiments where the virtual functions are associated with VMs (e.g., running in VNFs), the network function management module 320 is additionally configured to manage the associated virtual functions (see, e.g., the VMs 402 and virtual functions 410 of Fig. 4). To do this, the network function management module 320 is configured to deploy (i.e., boot up, perform instantiation, etc.) and close (i.e., shut down, remove from the network, etc.) the VMs based on the various service functions to be performed on the network packets (e.g., based on service functions of a service function chain corresponding to the network packet stream).

[0039] Accordingly, the network function management module 320 is further configured to manage each of the virtual function drivers associated with the corresponding VMs of each VNF, as well as to manage communications between them. In other words, the network function management module 320 is configured to direct the flow of data to the appropriate network functions and between the appropriate network functions. For example, the network function management module 320 is configured to determine an intended destination (e.g., a VM) to which the data should be directed (i.e., based on an access request) and direct the data to an interface of the intended destination (i.e., a virtual function of the VM).

[0040] The main memory management module 330, which may be implemented as hardware, firmware, software, virtualized hardware, emulated architecture, and / or a combination thereof as discussed above, is configured to manage the allocation and designation of main memory (e.g., the main memory 210 of Fig. 2) the network computing device 106 (see, e.g., the method 500 for laying out region-based cache blocks of Fig. 5). To do this, the illustrative main memory management module 330 includes a main memory allocation module 332 and a main memory access management module 334.

[0041] It should be appreciated that each of the main memory allocation module 332, the main memory access management module 334, and the relay region management module 336 of the main memory management module 330 may be separately implemented as hardware, firmware, software, virtualized hardware, an emulated architecture, and / or a combination thereof. For example, the main memory allocation module 332 may be implemented as a hardware component, while the main memory access management module 334 may be implemented as a virtualized hardware component, or as another combination of hardware, firmware, software, virtualized hardware, an emulated architecture, and / or a combination thereof.

[0042] The memory allocation module 332 is configured to allocate a range of physical memory addresses for region-based cache management. In some embodiments, the memory allocation module 332 may be configured to allocate large, contiguous physical memory address ranges. In one illustrative example, the memory allocation module 332 may allocate a 1 GB huge page for inter-VM shared memory (IVSHMEM). It should be appreciated that the memory allocation module 332 may be configured to allocate more than one large, contiguous physical memory address range at any given time. The memory allocation module 332 is further configured to divide the allocated physical memory range into multiple regions, which may be based, for example, on a number of regions for which particular types of data are to be stored.

[0043] The main memory access management module 334 is configured to manage accesses (e.g., read, write, etc.) to a main memory. In other words, the main memory access management module 334 is configured to manage the flow of data to and from the main memory of the network computing device 106, such as the main memory 210 of Fig. 2. Accordingly, in some embodiments, the main memory access management module 334 may be configured to function as a memory control interface or may be otherwise implemented. In some embodiments, data related to the allocation of the main memory (e.g., memory addresses, allocation information, access logs, etc.) may be stored in the main memory allocation data 302.

[0044] The cache management module 340, which may be implemented as hardware, firmware, software, virtualized hardware, emulated architecture, and / or a combination thereof as discussed above, is configured to manage the allocation and designation of cache memory of the network computing device, such as the cache memory 206 of Fig. 2. To do so, the illustrative cache management module 340 includes a cache block allocation module 342, a cache access management module 344, a cache feature management module 346, and a cache eviction management module 348.

[0045] It should be appreciated that each of the cache block allocation module 342, the cache access management module 344, the cache feature management module 346, and the cache eviction management module 348 of the cache memory management module 340 may be separately implemented as hardware, firmware, software, virtualized hardware, an emulated architecture, and / or a combination thereof. For example, the cache block allocation module 342 may be implemented as a hardware component, while the cache access management module 344 may be implemented as a virtualized hardware component, or as another combination of hardware, firmware, software, virtualized hardware, an emulated architecture, and / or a combination thereof.

[0046] The cache block allocation module 342 is configured to manage the allocation of cache memory into cache blocks. The cache block allocation module 342 is additionally configured to manage the allocation of the cache blocks to portions of physical memory (e.g., the main memory 210 of Fig. 2) the network computing device 106. For example, as previously described, the network computing device 106 is configured to divide allocated portions of main memory based on a number of designated regions. Continuing the example, the cache blocks associated with the allocated main memory for these regions are similarly designated, such as may be indicated by a region indicator maintained by the cache feature management module 346. As such, the cache block allocation module 342 is configured to map the divided portions of the cache block to the regions of the allocated portion of physical memory.

[0047] The main memory access management module 344 is configured to manage accesses (e.g., read, write, etc.) to a cache memory. In other words, the cache access management module 344 is configured to manage the flow of data to and from the cache memory of the network computing device 106, such as the cache memory 206 of Fig. 2. Accordingly, the cache access management module 344 may be configured to interface with the cache eviction management module 348 upon determining that a cache line needs to be made available to store alternative data.

[0048] The cache feature management module 346 is configured to manage the features of each cache block, such as region indicators and any associated configuration parameters. In some embodiments, the region indicators and / or configuration parameters may be stored in an architecturally exposed mechanism, such as one or more model-specific registers (MSRs). For example, in an embodiment where the cache block has been divided into three regions, including a non-region-specific portion, a shared region (e.g., for sharing data between devices), and a relay region (e.g., for memory-based communication channels that buffer data in transit from producers to consumers), the cache feature management module 346 is configured to maintain corresponding indicators for each divided portion of the cache block.Continuing the previous example, the cache feature management module 346 may be configured to assign a shared region indicator to the shared region and a relay region indicator to the relay region.

[0049] The cache feature management module 346 is further configured to manage a bias value associated with each region-based portion of physical memory and the respective portion of the cache block mapped to the corresponding region-based portion of physical memory. The bias value imposes a bias to prevent cache lines associated with the region corresponding to that bias value from being evicted in relative preference to other cache lines (e.g., from other regions). For example, in such embodiments, the cache feature management module 346 may be configured to designate a portion of the cache block with a shared region bias value (i.e., with a default value of one), where the portion of the cache block has been mapped to the allocated portion of physical memory designated as the shared region.In some embodiments, the cache features may be stored in the cache feature data 306.

[0050] The cache eviction management module 348 is configured to manage eviction of data from cache lines of cache memory of the network computing device 106 (e.g., the cache memory 206). To do so, the cache eviction management module 348 is configured to enforce cache eviction strategies (e.g., cache replacement algorithms / strategies). It should be recognized that there may be more than one cache eviction strategy. For example, the cache eviction strategies may include a non-directed standard eviction strategy (e.g., least recently used (LRU), most recently used (MRU), 2-way associative, direct mapped cache, etc.) and the region-based cache eviction strategy as described herein (see, e.g., the method 600 for performing a region-based cache line eviction of the Fig. 6 and Fig. 7). In some embodiments, the cache eviction strategies and / or information related thereto may be stored in the cache eviction strategy data 304.

[0051] Accordingly, the cache eviction management module 348 is configured to determine which part of a cache block a cache line selected for eviction corresponds to. In other words, the cache eviction management module 348 is configured to determine which region, if any, the cache line corresponds to. The cache eviction management module 348 is further configured to retrieve a bias value of the region upon determining whether the cache line corresponds to a region specified in the region-based cache management (e.g., a shared region, a relay region, etc.). As previously described, the bias value may be used to determine whether cache lines for the corresponding region should be evicted upon a selection.To do this, the cache eviction management module 348 is configured to generate a bias comparison value, such as by using a directed coin toss simulator, to determine whether or not the cache line selected for eviction should be evicted (i.e., based on a fractional probability).

[0052] The cache eviction management module 348 is further configured to dynamically adjust the bias value, such as based on cache hit / miss rates, latency, bandwidth, flows, workloads, etc. It should be appreciated that, under certain conditions, the cache eviction management module 348 may additionally be configured to ignore the bias value. In other words, the cache eviction management module 348 may be configured to enforce the default non-directional eviction strategy even upon determining that the cache line corresponds to a particular region with a corresponding bias value.For example, such conditions may include any condition under which the region-based cache eviction strategy would likely result in undesirable inefficiency, such as, among others, cache miss rates that exceed a miss rate threshold, repeated cache line selection for eviction that results in a cache line being selected for eviction from the same region as previous attempts in which eviction was rejected based on the bias value, the cache line corresponds to a particular type of workload, the data corresponds to a particular type of workload, etc.

[0053] Now referring to Fig. 4, in another illustrative embodiment, the network computing device 106 creates an environment 400 during operation. The illustrative environment 400 includes multiple VMs 402 executing on the network computing device 106, each of which is communicatively coupled to one of a plurality of virtual functions 410 of the NIC 216. In use, the NIC 216 is divided into sets of independent virtual functions 410, where each independent virtual function 410 has its own configuration (e.g., PCI configuration space, Media Access Control (MAC) address, settings, etc.) that may be exclusively assigned to VMs (e.g., the VMs 402) or used by native applications. Further, each of the virtual functions 410 also shares one or more physical resources on the NIC 216, such as an external network port, memory, etc., with the physical functions of the NIC 216.

[0054] The illustrative VMs 402 include a first VM, labeled VM(1) 404, a second VM, labeled VM(2) 406, and a third VM, labeled VM(N) 408 (i.e., the "Nth" compute node of the VMs 402, where "N" is a positive integer and denotes one or more additional VMs 402). The illustrative virtual functions 410 include a first virtual function, labeled VF(1) 412, a second virtual function, labeled VF(2) 414, and a third virtual function, labeled VF(N) 416 (i.e., the "Nth" compute node of the virtual functions 410, where "N" is a positive integer and denotes one or more additional virtual functions 410). Each of the virtual functions 410 is managed by the NIC 216 and traffic transmitted between them is managed by the network function management module 320 described in detail above. Fig. 3. It should be appreciated that one or more of the virtual functions 410 are configured to exchange communications via a shared memory (not shown). Accordingly, the network function management module 320 is further coupled to the main memory management module 330 and the cache memory management module 340 of Fig. 3, also described in detail above, and a virtual machine monitor (VMM) 418. The VMM 418 is responsible for controlling and handling privileged instruction execution. Accordingly, in some embodiments, the network function management module 320 may form part of the VMM 418, or vice versa.

[0055] As previously described, the network computing device 106 may rely on one or more VNFs to perform virtualized network functions utilizing one or more VMs. In some embodiments, such VNFs may be dynamically chained together to form a service function chain (SFC) in a process referred to as service chaining. In an SFC, each service function is performed by one or more VMs specifically created to perform a particular service function of the SFC. Which service functions are included in an SFC may be tailored to a characteristic associated with a network packet (e.g., payload type, network packet overhead). For example, an administrator of a data center including the network computing device 106 may create an SFC consisting of multiple security service functions (e.g.,a virtualized firewall function, a virtualized intrusion detection function, etc.), each of which may be configured to process network packets received from network computing device 106 in a particular order. While the illustrative environment 400 is depicted as utilizing virtual network functions, the functions described herein may be employed to cover communication exchanges between applications, VMs, VNFs, etc., using shared memory.

[0056] In some embodiments, while processing a network packet, one of the service functions may have resulting data (i.e., resulting from the processing of the network packet) to be passed to the next service function in the SFC. Accordingly, current cache prioritization technologies, such as DDIO (Data Direct I / O) prioritization, which applies to simple passing of network packet data between network functions (e.g., ongoing in VMs, VNFs, native applications, applications running in containers), do not apply to transfers of such dynamically changing data. It should be recognized that the network functions that share data (e.g., via shared memory buffers) do not share data long-term. In other words, once a sending network function receives a network packet (e.g.,via IVSHMEM or other buffers in shared memory) to another network function, the sending network function does not refer to the network packet's memory area except to reuse it once data in it has been consumed by the receiving network function.

[0057] It should also be acknowledged that other existing cache prioritization technologies, such as CAT, typically do not cover communications across IVSHMEM or shared-memory ring buffers. Furthermore, data in IVSHMEM or shared-memory buffers typically has a high risk of eviction due to cache pressure from processor / memory-intensive work within network functions. For example, CAT addresses problems of efficient cache sharing between applications but does not support buffer sharing between network functions. In other words, unlike CAT, region-based cache management can be effective for reserving cache capacity for memory-region-bound inter-VM sharing.

[0058] Now referring to Fig. 5, the network computing device 106, in use, may perform a method 500 for laying out region-based cache blocks of the cache memory 206 of the network computing device 106. The method 500 begins at block 502, in which the network computing device 106 determines whether region-based cache management is supported. For example, in some embodiments, the network computing device 106 may rely on CPU identification (CPUID) opcode instructions to determine whether region-based cache management is supported, as well as any region-based cache management capabilities that are supported (such as cache size, allocation granularity, alignment, supported memory size, etc.). If region-based cache management is supported, the method 500 proceeds to block 504; otherwise, the method 500 loops back to block 502 to determine whether region-based cache management is supported.In block 504, the network computing device 106 allocates a range of physical memory addresses of the main memory 110 to be associated with a region of cache memory (e.g., via calls by the operating system or the VMM 418 of the network computing device 106). For example, in block 506, the network computing device 106 allocates a sequential range of physical memory addresses of the main memory 110 associated with the region in the cache.

[0059] In block 508, the network computing device 106 associates a block of cache with the allocated range of physical memory addresses (i.e., the allocated portion of main memory 110). It should be appreciated that the size of the cache allocated to the range of physical memory addresses may be the same as or smaller than that of the associated physical memory region (i.e., a 1:N allocation ratio, where N represents an integer value greater than 1). Accordingly, effects similar to those observed in cache locking paradigms may be achieved if the dynamic footprint of the region is small enough, as may be appropriate in SFCs. In block 510, the network computing device 106 divides the allocated portion of main memory 110 into multiple regions.In one illustrative example, network computing device 106 divides the allocated portion of main memory 110 into a shared region and a relay region, both described above. Accordingly, in such an embodiment, an operating system and / or VMM may enhance the caching of shared structures by moving data into the shared region. Additionally, in such embodiments, storing shared memory (e.g., IVSHMEM) in a cache (e.g., in the shared region and / or the relay region) may minimize latency associated with data exchanges. It should be appreciated that the total size of allocated physical memory, the number of regions, etc., may be designed depending on the capabilities of network computing device 106 (i.e., the supported capabilities regarding region-based cache management).

[0060] In block 512, the network computing device 106 maps a corresponding region of the cache block allocated in block 508 to each region of the main memory 110 allocated in block 504. In some embodiments, the cache block may be mapped or otherwise partitioned at implementation-specific granularity (e.g., based on cache ways, cache lines, etc.). Alternatively, in other embodiments, a probabilistic bias may be provided, which may be implemented by a hardware-based soft scheduling or randomized technique. In one illustrative example, hardware of the network computing device 106 may allocate cache lines by using a round robin (e.g., via bitmasks to select cache ways to be associated with each memory region) or hashing function to distribute physical accesses among the reserved capacity allocation in the cache memory (e.g., rather than a static distribution).

[0061] In block 514, the network computing device 106 assigns one or more cache characteristics to each region of the allocated cache. For example, in block 516, the network computing device 106 assigns a region indicator to each region of the allocated cache. As previously described, the region indicator may be any type of data structure that can be used to identify the region of allocated memory. For example, the network computing device 106 may assign a shared region indicator to the shared region of the allocated cache block and a relay region indicator to the relay region of the allocated cache block.

[0062] In block 518, the network computing device 106 assigns one or more configuration parameters to each region, such as physical memory address information corresponding to the cache block, physical starting and ending memory addresses of each region, and / or any other data that establishes or otherwise directs cache behavior for each region. As previously described, the configuration parameters may include head and tail indicators, such as may be used in regions configured as circular buffers. In such embodiments, the head indicator effectively points to the oldest data item produced and queued for consumption but has yet to be consumed, and the tail indicator effectively points to the most recent data item produced and queued in the corresponding region.Accordingly, the area of ​​memory between the head indicator and the tail indicator consists of data that has been determined to be still active, while the remaining data (e.g., the area of ​​memory between the tail indicator and the head indicator) can be considered stale (e.g., data still remaining from old, previously processed and dropped / forwarded packets).

[0063] In block 520, the network computing device 106 assigns a bias value (i.e., a fractional probability) to each region. As previously described, the bias value corresponds to a fractional probability having a value between zero and one, which can be used to determine whether a cache line selected for eviction should be evicted, as described below in the method 600 of the Fig. 6 and Fig. 7. In some embodiments, at block 522, the cache features (e.g., the region indicators, configuration parameters, bias values, etc.) may be assigned or otherwise stored in corresponding MSRs.

[0064] Now referring to the Fig. 6 and Fig. 7, the network computing device 106, in use, may perform a method 600 for performing a region-based cache line eviction of a cache line in a cache block of the cache memory 206 of the network computing device 106. The method 600 begins in block 602, in which the network computing device 106 determines whether a cache line has been selected for eviction. If so, the method 600 proceeds to block 604; otherwise, the method 600 loops back to block 602 to again determine whether a cache line has been selected for eviction. In block 604, the network computing device 106 determines whether the cache line selected for eviction is located in a portion of the cache memory that corresponds to a shared region of physical memory (e.g., the main memory 110 of Fig. 1). To do this, as previously described, the network computing device 106 checks for a region indicator (e.g., in an MSR) to determine whether the cache line selected for eviction is located in the shared region.

[0065] If the network computing device 106 determines that the cache line selected for eviction does not correspond to the shared region, the method 600 branches to block 618 of Fig. 7, described below; otherwise, the method 600 branches to block 606. In block 606, the network computing device 106 retrieves a bias value associated with the shared region (i.e., a shared region bias value) with which the cache line is associated. As previously described, the bias value corresponds to a fractional probability, having a value between zero and one, that can be used to determine whether a cache line of a corresponding region selected for eviction should be evicted, as described below.

[0066] In block 608, the network computing device 106 generates a bias comparison value for the shared region (i.e., a shared region bias comparison value). To do so, in block 610, the network computing device 106 uses a directed coin toss simulation to return a probabilistic result (e.g., a directed coin toss simulation result). Of course, it should be appreciated that, in other embodiments, the network computing device 106 may utilize other methodologies to generate or otherwise determine the bias comparison value (e.g., via a random or pseudorandom value generation method).

[0067] In block 612, the network computing device 106 determines whether the cache line should be evicted. To do so, the network computing device 106 compares the shared region bias comparison value with the shared region bias value. In other words, a result of the directed coin toss simulation is compared with the probabilistic bias value to determine whether the cache line should be evicted. It should be appreciated that, in some embodiments, the network computing device 106 may be configured to adjust the shared region bias value based on a result of block 612 (i.e., the result of whether or not the cache line should be evicted).

[0068] If the network computing device 106 determines that the cache line should not be evicted, the method 600 branches to block 614, in which the network computing device 106 selects another cache line to be evicted (e.g., by using the cache line eviction strategy) before the method 600 returns to block 604 to determine whether the newly selected cache line is in the shared region for eviction. Otherwise, if the network computing device 106 determines that the cache line should be evicted, the method 600 branches to block 616, in which the network computing device 106 evicts the cache line before the method 600 returns to block 602 to wait until another cache line is selected for eviction.

[0069] As previously described, if the network computing device 106 determines that the cache line selected for eviction is not located in the shared region in block 605, the method 600 branches to block 618. In block 618, the network computing device 106 determines whether the cache line selected for eviction is located in a portion of the cache memory that corresponds to a relay region of physical memory (e.g., the main memory 110 of Fig.1). To do so, similar to block 604, the network computing device 106 checks for a region indicator (e.g., in an MSR) to determine if the cache line selected for eviction is located in the relay region. If not, the method 600 branches to block 620, in which the cache line is evicted according to the standard non-directed eviction strategy before returning to block 602 to determine if another cache line has been selected for eviction. Otherwise, if the network computing device 106 determines that the cache line is located in the relay region, the method 600 branches to block 622.

[0070] In block 622, the network computing device 106 determines whether the cache line is active or not stale. To do this, in the illustrative embodiment, the network computing device 106 determines whether the cache line is located in an active region of a circular buffer. In other words, the network computing device 106 determines whether the cache line is located in the range of memory between a head indicator and a tail indicator of the circular buffer that contains data that has been determined to be still active. If the network computing device 106 determines that the cache line is not active in block 622 (i.e.If stale data is located in a cache line that is in the range of memory between the tail indicator and the head indicator), such as data associated with old network packets that were previously processed and dropped / forwarded and still exists, the method 600 branches to block 624. In block 624, the network computing device 106 flushes the cache line. In some embodiments, in block 626, the network computing device 106 invalidates the cache line and flushes the cache line without writing the data back to physical memory.

[0071] If the network computing device 106 determines that the cache line is active at block 622, the method 600 branches to block 626, in which the network computing device 106 retrieves a bias value corresponding to the relay region (i.e., a relay region bias value) with which the cache line is associated. As previously described, the bias value corresponds to a fractional probability having a value between zero and one. At block 628, the network computing device 106 generates a bias comparison value for the relay region (i.e., a relay region bias comparison value). To do so, at block 630, the network computing device 106 uses a directed coin toss simulation to return a probabilistic result (e.g., a directed coin toss simulation result).

[0072] In block 632, the network computing device 106 determines whether the cache line should be evicted. To do so, the network computing device 106 compares the relay region bias comparison value with the relay region bias value. In other words, a result of the directed coin toss simulation is compared with the probabilistic bias value to determine whether the cache line should be evicted. It should be appreciated that, in some embodiments, the network computing device 106 may be configured to dynamically adjust the relay region bias value based on a result of block 632 (i.e., the result of whether or not the cache line should be evicted).

[0073] If the network computing device 106 determines that the cache line should not be evicted, the method 600 branches to block 634, in which the network computing device 106 selects another cache line to be evicted (e.g., by using the cache line eviction strategy) before the method 600 returns to block 604 to determine whether the newly selected cache line is in the relay region for eviction. Otherwise, if the network computing device 106 determines that the cache line should be evicted, the method 600 branches to block 636, in which the network computing device 106 evicts the cache line before the method 600 returns to block 602 to wait until another cache line is selected for eviction.

[0074] It should be appreciated that at least a portion of one or both of methods 500 and 600 may be performed by the NIC 216 of the network computing device 106. It should further be appreciated that, in some embodiments, one or both of the methods 500 and 600 may be embodied as various instructions stored on a computer-readable medium that may be executed by a processor (e.g., processor 202, etc.), the NIC 216, and / or other components of the network computing device 106 to cause the network computing device 106 to perform the methods 500 and 600.The computer-readable medium may be implemented as any type of medium capable of being read by the network computing device 106, including, but not limited to, the main memory 210, the data storage device 212, secure storage (not shown) of the NIC 216, other memory or data storage devices of the network computing device 106, portable media readable by a peripheral device of the network computing device 106, and / or other media. EXAMPLES

[0075] Illustrative examples of the technologies disclosed herein are provided below. An embodiment of the technologies may include one or more of the examples described below and any combination thereof.

[0076] Example 1 includes a network computing device for region-based cache management, the network computing device comprising a processor having a cache memory; a main memory, different from the cache memory, coupled to the processor; and one or more data storage devices having stored therein a plurality of instructions that, when executed by the processor, cause the network computing device to select a cache line for eviction from among a plurality of cache lines of the cache memory; determine whether the cache line selected for eviction is located in a cache block of the cache memory currently associated with a corresponding memory region of the main memory, the cache block comprising one or more of the plurality of cache lines;in response to a determination that the cache block is associated with a corresponding memory region, retrieve a bias value associated with the corresponding memory region, the bias value corresponding to a fractional probability; generate a bias comparison value for the corresponding memory region; compare the bias value of the corresponding memory region and the bias comparison value generated for the corresponding memory region; and, in response to a determination that a result of the comparison indicates to evict the cache line, evict the cache line.

[0077] Example 2 includes the subject matter of Example 1, and wherein selecting the cache line for eviction comprises selecting the cache line based on a non-directed cache eviction strategy.

[0078] Example 3 includes the subject matter of any of Examples 1 and 2, and wherein generating the bias comparison value comprises generating the bias comparison value as a function of a directed coin toss simulation.

[0079] Example 4 includes the subject matter of any of Examples 1-3, and wherein determining whether the cache line selected for eviction is located in the cache block associated with the corresponding memory region comprises reading a model-specific register of the corresponding memory region that includes an indication of the corresponding memory region with which the cache block is associated.

[0080] Example 5 includes the subject matter of any of Examples 1-4, and wherein determining whether the cache line selected for eviction is located in the cache block associated with the corresponding memory region comprises determining whether the corresponding memory region is a shared memory region allocated to store data to be shared by at least one of a virtualized component or a physical component of the network computing device.

[0081] Example 6 includes the subject matter of any of Examples 1-5, and wherein determining whether the cache line selected for eviction is located in the cache block associated with the corresponding memory region comprises determining whether the corresponding memory region is a relay memory region allocated to store transient data communications.

[0082] Example 7 includes the subject matter of any of Examples 1-6, and wherein the plurality of instructions further cause the network computing device to allocate a block of main memory, the block of main memory comprising a range of memory addresses of the main memory; associate the cache block with the allocated block of main memory; divide the allocated block of main memory into a plurality of memory regions; allocate a corresponding portion of the cache block to each of the plurality of memory regions; and assign cache characteristics to each of the plurality of memory regions, the cache characteristics including a region indicator and the bias value.

[0083] Example 8 includes the subject matter of any of Examples 1-7, and wherein allocating the block of main memory comprises allocating a sequential range of memory addresses of the main memory.

[0084] Example 9 includes the subject matter of any of Examples 1-8, and wherein dividing the allocated block of main memory into the plurality of memory regions comprises dividing the allocated block of main memory into a shared region and a relay region, wherein the shared region is allocated to store data to be shared by at least one of a virtualized component or a physical component of the network computing device, and wherein the relay region is allocated to store transient data communications.

[0085] Example 10 includes the subject matter of any of Examples 1-9, and wherein assigning the region indicator to the shared region comprises storing a shared region indicator in a model-specific register.

[0086] Example 11 includes the subject matter of any of Examples 1-10, and wherein assigning the region indicator to the relay region comprises storing an indicator for a relay region in a model-specific register.

[0087] Example 12 includes the subject matter of any of Examples 1-11, and wherein the corresponding portion of the cache associated with the shared region comprises a circular buffer, wherein the cache characteristics additionally include one or more configuration parameters, and wherein the configuration parameters include a shared region head indicator and a shared region tail indicator.

[0088] Example 13 includes the subject matter of any of Examples 1-12, and wherein the plurality of instructions further cause the network computing device to adjust the bias value in response to a result of comparing the bias value of the corresponding memory region and the bias comparison value generated for the corresponding memory region.

[0089] Example 14 includes a network computing device for region-based cache management, the network computing device comprising a processor having a cache memory; and cache management circuitry for selecting a cache line of the cache memory for eviction, wherein the cache line is one of a plurality of cache lines of the cache memory; determining whether the cache line selected for eviction is located in a cache block of the cache memory currently associated with a corresponding memory region of the main memory of the network computing device, the cache block comprising one or more of the plurality of cache lines; retrieving, in response to a determination that the cache block is associated with a corresponding memory region, a bias value associated with the corresponding memory region, the bias value corresponding to a fractional probability;Generating a bias comparison value for the corresponding memory region; comparing the bias value of the corresponding memory region and the bias comparison value generated for the corresponding memory region; and evicting the cache line in response to a determination that a result of the comparison indicates evicting the cache line.

[0090] Example 15 includes the subject matter of Example 14, and wherein selecting the cache line for eviction comprises selecting the cache line based on a non-directed cache eviction strategy.

[0091] Example 16 includes the subject matter of any of Examples 14 and 15, and wherein generating the bias comparison value comprises generating the bias comparison value as a function of a directed coin toss simulation.

[0092] Example 17 includes the subject matter of any of Examples 14-16, and wherein determining whether the cache line selected for eviction is located in the cache block associated with the corresponding memory region comprises reading a model-specific register of the corresponding memory region that includes an indication of the corresponding memory region with which the cache block is associated.

[0093] Example 18 includes the subject matter of any of Examples 14-17, and wherein determining whether the cache line selected for eviction is located in the cache block associated with the corresponding memory region comprises determining whether the corresponding memory region is a shared memory region allocated to store data to be shared by at least one of a virtualized component or a physical component of the network computing device.

[0094] Example 19 includes the subject matter of any of Examples 14-18, and wherein determining whether the cache line selected for eviction is located in the cache block associated with the corresponding memory region comprises determining whether the corresponding memory region is a relay memory region allocated to store transient data communications.

[0095] Example 20 includes the subject matter of any of Examples 14-19, and further including memory management circuitry for allocating a block of main memory, the block of main memory comprising a range of memory addresses of the main memory; connecting the cache block to the allocated block of main memory; dividing the allocated block of main memory into a plurality of memory regions; allocating a corresponding portion of the cache block to each of the plurality of memory regions; and assigning cache characteristics to each of the plurality of memory regions, the cache characteristics including a region indicator and the bias value.

[0096] Example 21 includes the subject matter of any of Examples 14-20, and wherein allocating the block of main memory comprises allocating a sequential range of memory addresses of the main memory.

[0097] Example 22 includes the subject matter of any of Examples 14-21, and wherein dividing the allocated block of main memory into the plurality of memory regions comprises dividing the allocated block of main memory into a shared region and a relay region, wherein the shared region is allocated to store data to be shared by at least one of a virtualized component or a physical component of the network computing device, and wherein the relay region is allocated to store transient data communications.

[0098] Example 23 includes the subject matter of any of Examples 14-22, and wherein assigning the region indicator to the shared region comprises storing a shared region indicator in a model-specific register.

[0099] Example 24 includes the subject matter of any of Examples 14-23, and wherein assigning the region indicator to the relay region comprises storing an indicator for a relay region in a model-specific register,

[0100] Example 25 includes the subject matter of any of Examples 14-24, and wherein the corresponding portion of the cache associated with the shared region comprises a circular buffer, wherein the cache characteristics additionally include one or more configuration parameters, and wherein the configuration parameters include a shared region head indicator and a shared region tail indicator.

[0101] Example 26 includes the subject matter of any of Examples 14-25, and wherein the cache memory management circuitry is further to adjust the bias value in response to a result of comparing the bias value of the corresponding memory region and the bias comparison value generated for the corresponding memory region.

[0102] Example 27 illustrates a method for region-based cache management, the method comprising: selecting, by a network computing device, a cache line for eviction from among the plurality of cache lines, wherein the cache line comprises one of a plurality of cache lines of a cache memory residing on a processor of the network computing device; determining, by the network computing device, whether the cache line selected for eviction is located in a cache block of the cache memory currently associated with a corresponding memory region of a main memory of the network computing device externally coupled to the processor, wherein the cache block comprises one or more of the plurality of cache lines;Retrieving, by the network computing device and in response to a determination that the cache block is associated with a corresponding memory region, a bias value associated with the corresponding memory region, wherein the bias value corresponds to a fractional probability; generating, by the network computing device, a bias comparison value for the corresponding memory region; comparing, by the network computing device, the bias value of the corresponding memory region and the bias comparison value generated for the corresponding memory region; and evicting, by the network computing device, the cache line in response to a determination that a result of the comparison indicates to evict the cache line.

[0103] Example 28 includes the subject matter of Example 27, and wherein selecting the cache line for eviction comprises selecting the cache line based on a non-directed cache eviction strategy.

[0104] Example 29 includes the subject matter of any of Examples 27 and 28, and wherein generating the bias comparison value comprises generating the bias comparison value as a function of a directed coin toss simulation.

[0105] Example 30 includes the subject matter of any of Examples 27-29, and wherein determining whether the cache line selected for eviction is located in the cache block associated with the corresponding memory region comprises reading a model-specific register of the corresponding memory region that includes an indication of the corresponding memory region with which the cache block is associated.

[0106] Example 31 includes the subject matter of any of Examples 27-30, and wherein determining whether the cache line selected for eviction is located in the cache block associated with the corresponding memory region comprises determining whether the corresponding memory region is a shared memory region allocated to store data to be shared by at least one of a virtualized component or a physical component of the network computing device.

[0107] Example 32 includes the subject matter of any of Examples 27-31, and wherein determining whether the cache line selected for eviction is located in the cache block associated with the corresponding memory region comprises determining whether the corresponding memory region is a relay memory region allocated to store transient data communications.

[0108] Example 33 includes the subject matter of any of Examples 27-32, and further including allocating, by the network computing device, a block of main memory, wherein the block of main memory comprises a range of memory addresses of the main memory; connecting, by the network computing device, the cache block to the allocated block of main memory; dividing, by the network computing device, the allocated block of main memory into a plurality of memory regions; allocating, by the network computing device, a corresponding portion of the cache block to each of the plurality of memory regions; and assigning, by the network computing device, cache features to each of the plurality of memory regions, wherein the cache features include a region indicator and the bias value.

[0109] Example 34 includes the subject matter of any of Examples 27-33, and wherein allocating the block of main memory comprises allocating a sequential range of memory addresses of the main memory.

[0110] Example 35 includes the subject matter of any of Examples 27-34, and wherein dividing the allocated block of main memory into the plurality of memory regions comprises dividing the allocated block of main memory into a shared region and a relay region, wherein the shared region is allocated to store data to be shared by at least one of a virtualized component or a physical component of the network computing device, and wherein the relay region is allocated to store transient data communications.

[0111] Example 36 includes the subject matter of any of Examples 27-35, and wherein assigning the region indicator to the shared region comprises storing a shared region indicator in a model-specific register.

[0112] Example 37 includes the subject matter of any of Examples 27-36, and wherein assigning the region indicator to the relay region comprises storing an indicator for a relay region in a model-specific register.

[0113] Example 38 includes the subject matter of any of Examples 27-37, and wherein the corresponding portion of the cache memory associated with the shared region comprises a circular buffer, wherein the cache characteristics additionally include one or more configuration parameters, and wherein the configuration parameters include a shared region head indicator and a shared region tail indicator.

[0114] Example 39 includes the subject matter of any of Examples 27-38, and further including adjusting, by the network computing device, the bias value in response to a result of comparing the bias value of the corresponding memory region and the bias comparison value generated for the corresponding memory region.

[0115] Example 40 includes a computing device comprising a processor; and a memory having stored therein a plurality of instructions that, when executed by the processor, cause the network computing device to perform the method of any of Examples 27-39.

[0116] Example 41 includes one or more machine-readable storage media comprising a plurality of instructions stored thereon that, in response to being executed, cause a network computing device to perform the method of any of Examples 27-39.

[0117] Example 42 includes a network computing device for region-based cache management, the network computing device comprising means for selecting a cache line for eviction from the plurality of cache lines, wherein the cache line comprises one of a plurality of cache lines of a cache memory located on a processor of the network computing device; means for determining whether the cache line selected for eviction is located in a cache block of the cache memory currently associated with a corresponding memory region of a main memory of the network computing device externally coupled to the processor, the cache block comprising one or more of the plurality of cache lines; means for retrieving, in response to a determination that the cache block is associated with a corresponding memory region, a bias value associated with the corresponding memory region, wherein the bias value corresponds to a fractional probability;Means for generating a bias comparison value for the corresponding memory region; means for comparing the bias value of the corresponding memory region and the bias comparison value generated for the corresponding memory region; and means for evicting the cache line in response to a determination that a result of the comparison indicates evicting the cache line.

[0118] Example 43 includes the subject matter of Example 42, and wherein the means for selecting the cache line for eviction comprises means for selecting the cache line based on a non-directed cache eviction strategy.

[0119] Example 44 includes the subject matter of any of Examples 42 and 43, and wherein the means for generating the bias comparison value comprises means for generating the bias comparison value as a function of a directed coin toss simulation.

[0120] Example 45 includes the subject matter of any of Examples 42-44, and wherein the means for determining whether the cache line selected for eviction is located in the cache block associated with the corresponding memory region comprises means for reading a model-specific register of the corresponding memory region that includes an indication of the corresponding memory region with which the cache block is associated.

[0121] Example 46 includes the subject matter of any of Examples 42-45, and wherein the means for determining whether the cache line selected for eviction is located in the cache block associated with the corresponding memory region comprises means for determining whether the corresponding memory region is a shared memory region allocated to store data to be shared by at least one of a virtualized component or a physical component of the network computing device.

[0122] Example 47 includes the subject matter of any of Examples 42-46, and wherein the means for determining whether the cache line selected for eviction is located in the cache block associated with the corresponding memory region comprises means for determining whether the corresponding memory region is a relay memory region allocated to store temporary data communications.

[0123] Example 48 includes the subject matter of any of Examples 42-47, and further including main memory management circuitry for (i) allocating a block of main memory, the block of main memory comprising a range of memory addresses of the main memory, and (ii) connecting the cache block to the allocated block of main memory; means for dividing the allocated block of main memory into a plurality of memory regions; means for allocating a corresponding portion of the cache block to each of the plurality of memory regions; and means for assigning cache characteristics to each of the plurality of memory regions, the cache characteristics including a region indicator and the bias value.

[0124] Example 49 includes the subject matter of any of Examples 42-48, and wherein allocating the block of main memory comprises allocating a sequential range of memory addresses of the main memory.

[0125] Example 50 includes the subject matter of any of Examples 42-49, and wherein the means for dividing the allocated block of main memory into the plurality of memory regions comprises means for dividing the allocated block of main memory into a shared region and a relay region, wherein the shared region is allocated to store data to be shared by at least one of a virtualized component or a physical component of the network computing device, and wherein the relay region is allocated to store transient data communications.

[0126] Example 51 includes the subject matter of any of Examples 42-50, and wherein the means for assigning the region indicator to the shared region comprises means for storing a shared region indicator in a model-specific register.

[0127] Example 52 includes the subject matter of any of Examples 42-51, and wherein the means for assigning the region indicator to the relay region comprises means for storing an indicator for a relay region in a model-specific register.

[0128] Example 53 includes the subject matter of any of Examples 42-52, and wherein the corresponding portion of the cache associated with the shared region comprises a circular buffer, wherein the cache characteristics additionally include one or more configuration parameters, and wherein the configuration parameters include a shared region head indicator and a shared region tail indicator.

[0129] Example 54 includes the subject matter of any of Examples 42-53, and further including means for adjusting the bias value in response to a result of comparing the bias value of the corresponding memory region and the bias comparison value generated for the corresponding memory region.

Claims

[1] A network computing device (106) for region-based cache management, the network computing device (106) comprising: a processor (202) having a cache memory (206); a main memory (210), different from the cache memory (206), coupled to the processor (202); and one or more data storage devices (212) having stored therein a plurality of instructions that, when executed by the processor (202), cause the network computing device (106) to: selecting a cache line for eviction from a plurality of cache lines of the cache memory (206); determining whether the cache line selected for eviction is located in a cache block of the cache memory (206) that is currently associated with a corresponding memory region of the main memory (210), the cache block comprising one or more of the plurality of cache lines; in response to a determination that the cache block is associated with a corresponding memory region, retrieve a bias value associated with the corresponding memory region, the bias value corresponding to a fractional probability; to generate a bias comparison value for the corresponding memory region; compare the bias value of the corresponding memory region and the bias comparison value generated for the corresponding memory region; and to evict the cache line in response to a determination that a result of the comparison indicates to evict the cache line. [2] The network computing device (106) of claim 1, wherein selecting the cache line for eviction comprises selecting the cache line based on a non-directed cache eviction strategy. [3] The network computing device (106) of claim 1, wherein generating the bias comparison value comprises generating the bias comparison value as a function of a directed coin toss simulation. [4] The network computing device (106) of claim 1, wherein determining whether the cache line selected for eviction is located in the cache block associated with the corresponding memory region comprises reading a model-specific register of the corresponding memory region that includes an indication of the corresponding memory region with which the cache block is associated. [5] The network computing device (106) of claim 1, wherein the plurality of instructions further cause the network computing device (106) to: allocate a block of main memory (210), the block of main memory (210) comprising a series of memory addresses of the main memory (210); connecting the cache block to the allocated block of main memory (210); divide the allocated block of main memory (210) into a plurality of memory regions; allocate a corresponding portion of the cache block to each of the plurality of memory regions; and assign cache characteristics to each of the plurality of memory regions, wherein the cache characteristics include a region indicator and the bias value. [6] The network computing device (106) of claim 5, wherein allocating the block of main memory (210) comprises allocating a sequential range of memory addresses of the main memory (210). [7] The network computing device (106) of claim 5, wherein dividing the allocated block of main memory (210) into the plurality of memory regions comprises dividing the allocated block of main memory (210) into a shared region and a relay region, wherein the shared region is allocated to store data to be shared by at least one of a virtualized component and a physical component of the network computing device (106), and wherein the relay region is allocated to store transient data communications. [8] The network computing device (106) of claim 7, wherein assigning the region indicator to the shared region comprises storing a shared region indicator in a model-specific register, and wherein assigning the region indicator to the relay region comprises storing a relay region indicator in a model-specific register. [9] The network computing device (106) of claim 7, wherein the corresponding portion of the cache memory (206) associated with the shared region comprises a circular buffer, wherein the cache characteristics additionally include one or more configuration parameters, and wherein the configuration parameters include a shared region head indicator and a shared region tail indicator. [10] The network computing device (106) of claim 1, wherein the plurality of instructions further cause the network computing device (106) to adjust the bias value in response to a result of comparing the bias value of the corresponding memory region and the bias comparison value generated for the corresponding memory region. [11] A method (600) for region-based cache management, the method (600) comprising: Selecting (602), by a network computing device (106), a cache line for eviction from the plurality of cache lines, the cache line comprising one of a plurality of cache lines of a cache memory (206) located on a processor (202) of the network computing device (106); Determining, by the network computing device (106), whether the cache line selected for eviction is located in a cache block of the cache memory (206) that is currently associated with a corresponding memory region of a main memory (210) of the network computing device (106) coupled to the processor (202), the cache block comprising one or more of the plurality of cache lines; retrieving (606), by the network computing device (106) and in response to a determination that the cache block is associated with a corresponding memory region, a bias value associated with the corresponding memory region, the bias value corresponding to a fractional probability; generating (608), by the network computing device (106), a bias comparison value for the corresponding memory region; Comparing, by the network computing device (106), the bias value of the corresponding memory region and the bias comparison value generated for the corresponding memory region; and Flushing (612), by the network computing device (106), the cache line in response to a determination that a result of the comparison indicates to flush the cache line. [12] The method (600) of claim 11, wherein selecting the cache line for eviction (620) comprises selecting the cache line based on a non-directed cache eviction strategy. [13] The method (600) of claim 11, wherein generating (610) the bias comparison value comprises generating the bias comparison value as a function of a directed coin toss simulation. [14] The method (600) of claim 11, wherein determining whether the cache line selected for eviction is located in the cache block associated with the corresponding memory region comprises reading a model-specific register of the corresponding memory region that includes an indication of the corresponding memory region with which the cache block is associated. [15] The method (600) of claim 11, further comprising: Allocating, by the network computing device (106), a block of main memory (210), the block of main memory (210) comprising a series of memory addresses of the main memory; Connecting, by the network computing device (106), the cache block to the allocated block of main memory (210); Dividing, by the network computing device (106), the allocated block of main memory into a plurality of memory regions; Assigning, by the network computing device (106), a corresponding portion of the cache block to each of the plurality of memory regions; and Assigning, by the network computing device (106), cache characteristics to each of the plurality of memory regions, the cache characteristics including a region indicator and the bias value. [16] The method (600) of claim 15, wherein allocating the block of main memory (210) comprises allocating a sequential range of memory addresses of the main memory (210). [17] The method (600) of claim 15, wherein dividing the allocated block of main memory (210) into the plurality of memory regions comprises dividing the allocated block of main memory (210) into a shared region and a relay region, wherein the shared region is allocated to store data to be shared by at least one virtualized component or a physical component of the network computing device (106), and wherein the relay region is allocated to store transient data communications. [18] The method (600) of claim 17, wherein assigning the region indicator to the shared region comprises storing a shared region indicator in a model-specific register, and wherein assigning the region indicator to the relay region comprises storing a relay region indicator in a model-specific register. [19] The method (600) of claim 17, wherein the corresponding portion of the cache memory associated with the shared region comprises a circular buffer, wherein the cache characteristics additionally include one or more configuration parameters, and wherein the configuration parameters include a shared region head indicator and a shared region tail indicator. [20] The method (600) of claim 11, further comprising adjusting, by the network computing device (106), the bias value in response to a result of comparing the bias value of the corresponding memory region and the bias comparison value generated for the corresponding memory region. [21] A network computing device (106) comprising: a processor (202); and a memory having stored therein a plurality of instructions that, when executed by the processor, cause the network computing device (106) to perform the method (600) of any of claims 11-20. [22] One or more machine-readable storage media comprising a plurality of instructions stored thereon which, in response to their execution, cause a network computing device (106) to perform the method according to any one of claims 11-20.

Citation Information

Patent Citations

  • Avoiding cache line sharing in virtual machines

    US20080022048A1

  • Cache memory

    US20080256303A1

  • Cache Pooling for Computing Systems

    US20090204764A1

  • Cache system and method using tagged cache lines for matching cache strategy to I / O application

    US5915262A

  • Cache operations for memory management

    WO2015047348A1