Priority-Based Cache Line Eviction Algorithm for Flexible Cache Allocation Technology

Flex-CAT addresses the challenge of performance fluctuations for high-priority jobs in multitenant environments by implementing a dynamic priority-based eviction algorithm within the Flexible Cache Allocation Technology, ensuring efficient cache partitioning and maintaining performance determinism.

JP7682613B2Active Publication Date: 2025-05-26INTEL CORP
View PDF 8 Cites 0 Cited by

Patent Information

Application Number
JP2020150869
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2019-11-26
Filing Date
2020-09-08
Publication Date
2025-05-26
Estimated Expiration
2040-09-08

AI Technical Summary

Technical Problem

In multitenant cloud environments, high-priority (HP) jobs face performance fluctuations and quality of service (QoS) degradation due to co-located low-priority (LP) jobs, which existing cache management techniques struggle to mitigate effectively.

Method used

The Flexible Cache Allocation Technology (Flex-CAT) implements a priority-based eviction algorithm that dynamically manages cache partitions by determining the optimal way number for each cache set based on utilization, thereby suppressing HP LLC evictions by LP jobs.

Benefits of technology

Flex-CAT ensures efficient cache partitioning, prioritizing HP jobs by minimizing their evictions and maintaining performance determinism, thus effectively meeting service level agreements (SLAs) in multitenant environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007682613000003
    Figure 0007682613000003
  • Figure 0007682613000004
    Figure 0007682613000004
  • Figure 0007682613000005
    Figure 0007682613000005
Patent Text Reader

Abstract

To provide a cache line eviction algorithm.SOLUTION: A computing system 200 includes a last level cache (LLC) 204 and a cache control circuit (CCC) 201. The LLC has multiple ways, each allocated to one of multiple priorities, each specifying minimum and maximum numbers of ways to occupy. The CCC stores an incoming cache line (CL) having a requestor priority to an invalid CL if the invalid CL exists. Otherwise, when the requestor priority is a lowest priority and has an occupancy of one or more, or when the occupancy is at the maximum, the CCC evicts a least recently used (LRU) CL of the requestor priority. Otherwise, when the occupancy is between the minimum and the maximum, the CCC evicts an LRU CL of the requestor priority or a lower priority.SELECTED DRAWING: Figure 2
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The technical field of the present invention generally relates to computer processor architecture, and more particularly, to a priority-based eviction algorithm of an improved flexible cache allocation technique (Flex-CAT) for cache partitioning.

Background Art

[0002] Multitenancy is recognized as a solution that realizes high system utilization and cost savings through space sharing. Multitenancy can be realized in a cloud environment by virtualizing virtual machines (VMs) that execute user applications on each core. New computing paradigms such as Function as a Service (FaaS) utilize container virtualization to execute a large number of individual lightweight functions within containers. In a typical multitenant environment, high-priority (HP) jobs coexist with low-priority (LP) jobs on the same computing resources such as multi-core processors or cores. HP jobs are latency-sensitive jobs, while LP jobs often have loose deadlines. Some of the HP jobs require performance determinism in addition to low latency. Users submitting jobs conclude a service quality (QoS) service level agreement (SLA) with a cloud service provider (CSP) and comply with them regarding the guarantee of latency or performance determinism accordingly. The CSP needs to meet the SLA by suppressing performance fluctuations or the degradation of QoS of HP jobs caused by other co-located LP jobs.

Brief Description of the Drawings

[0003] The present invention is shown by way of example and not limitation in the figures of the accompanying drawings, in which like reference numerals indicate like elements.

[0004]

Figure 1

[0005]

Figure 2

[0006]

Figure 3

[0007]

Figure 4

[0008]

Figure 5

[0009]

Figure 6

[0010] FIG. 7A and FIG. 7B are block diagrams showing a general-purpose vector-oriented instruction format and its instruction template according to some embodiments of the present invention.

[0011]

Figure 7A

[0012]

Figure 7B

[0013]

Figure 8A

[0014]

Figure 8B

[0015]

Figure 8C

[0016]

Figure 8D

[0017]

Figure 9

[0018]

Figure 10A

[0019]

Figure 10B

[0020] Figures 11A and 11B are block diagrams of a more specific and exemplary in-order core architecture, where the core is one of several logical blocks (including other cores of the same type and / or different types) within the chip.

[0021]

Figure 11A

[0022]

Figure 11B

[0023]

Figure 12

[0024] Figures 13 to 16 are block diagrams of an exemplary computer architecture.

[0025]

Figure 13

[0026]

Figure 14

[0027]

Figure 15

[0028]

Figure 16

[0029]

Figure 17

[0030] In the following description, numerous specific details are set forth. However, it will be understood that some embodiments may be practiced without these specific details. In other instances, well-known circuits, structures, and techniques have not been shown in detail so as not to obscure the understanding of this description.

[0031] Expressions such as "one embodiment", "an embodiment", "an exemplary embodiment", etc. in this specification indicate that the features, structures, or characteristics described in the described embodiments may be included, but not necessarily all embodiments include these features, structures, or characteristics. Further, such language does not necessarily refer to the same embodiment. Additionally, when a feature, structure, or characteristic is described for one embodiment, it is to be understood by those skilled in the art that, if explicitly stated, these features, structures, or characteristics can also affect other embodiments.

[0032] As described above, a cloud service provider (CSP) needs to meet a service level agreement (SLA) by suppressing performance fluctuations or a degradation in the quality of service (QoS) of high-performance (HP) jobs caused by other low-performance (LP) jobs located in the same position. Specifically, the disclosed embodiments will describe a Flexible Cache Allocation Technology (Flex-CAT). This is a structural solution that suppresses HP last-level cache (LLC) evictions by LP jobs. The Flex-CAT approach dynamically determines the optimal way number for each set according to the utilization of cache lines (CLs) within each cache set based on various priorities. Its boundaries (minimum and maximum way numbers) are specified in a model-specific register (MSR) that provides hints when selecting eviction targets during LLC cache filling.

[0033] Flex-CAT has the advantage of providing an easily configurable and yet flexible interface for specifying cache partitions. Since Flex-CAT manages partitions at a detailed granularity, it supports a dynamic cache partitioning scheme that makes priority-based LLC eviction decisions based on real-time data. Flex-CAT contributes to meeting the requirements of cloud service providers for structural features that satisfy QoS guarantees, such as performance determinacy and separation of high-performance and low-performance jobs. In this specification, low-performance jobs may sometimes also be referred to as "noisy neighbors".

[0034] Alternative, inferior approaches have tried to resolve the imbalance between HP and LP jobs sharing resources by assigning individual sets of ways to cores and restricting HP and LP jobs to specific cores. However, such approaches have several problems. For example, there is no concept of priority in such mechanisms. Some of these approaches separate HP jobs from LP jobs by assigning all cache sets and the set of ways unique to HP cores, but such dedicated resources cannot be utilized by LP workloads when not being used by HP workloads. Further, according to some of these approaches, certain limited cache sets (e.g., the x set) become more saturated than others (e.g., if N is the total number of cache sets, the N - x set). The equal way assignment to HP jobs due to the saturation of these x sets can lead to overprovisioning across the entire N - x set and underutilization of these N - x. Further, allocating fewer ways to a core than the maximum way results in reduced associativity, more conflict misses, and a performance degradation. In the case of a method using static assignment of fixed cache ways to cores, there is no room for flexibility during cache eviction and filling.

[0035] On the other hand, the disclosed embodiments provide a flexible interface for dynamically specifying priorities and cache partitions. Priorities are enumerated in ascending order. Flexible cache partitions can be specified for the maximum and minimum number of ways per priority. Unlike some other approaches, Flex-CAT does not require software to precisely specify the cache ways assigned to each partition.

[0036] Such dynamic priority and cache partitioning specifications are supported by a class-of-service (CLOS) register. The register holds the following values for each CLOS. · CLOS priority P: Pn bits · Maximum number of ways occupied by priority P: mxwn bits · Minimum number of ways occupied by priority P: mnwn bits

[0037] For example, if the maximum number of priorities is 4, then Pn = log(4) = 2. If the maximum number of ways is 16, then mxwn = mnwn = log(16) = 4.

Table 1

[0038] According to the embodiments disclosed in this specification, the requester is the owner of the CL to be filled in the LLC. Let the priority of the requester be PF. In the system, PL is the lowest priority and PH is the highest priority. Let loc be the final storage location for the CL of the requester determined by Flex-CAT. The occupancy number O[PF] of the requester is the number of CLs occupied by the requester within the indexed cache set.

[0039] Flex-CAT is a new eviction algorithm that performs priority-based cache partitioning at the cache set granularity. As long as the occupancy number of the requester does not exceed the maximum way allocation (mxw), the LP CL is preferentially evicted. When the occupancy number of the requester reaches the maximum allocation, Flex-CAT gives priority to self-eviction over other priority evictions and tries to stay within the partition boundary. In the limited situation where the target cannot be found in these first two steps, Flex-CAT selects the HP CL for eviction to create a place for subsequent cache filling. Such a basic concept of Flex-CAT is shown in Figure 4.

[0040] The detailed algorithm is shown in the flowcharts of Figures 4 to 6 and will be described below. When filling the LLC, the disclosed embodiments determine the cache set index for the subsequent cache lines of the requester by a conventional hashing algorithm.

[0041] After indexing to the appropriate cache set, Flex-CAT first searches for invalid LLC entries within the indexed cache set. If the cache set is full and no invalid storage location is found, Flex-CAT determines the target CL that needs to be evicted from the LLC. This ensures that Flex-CAT is enabled only for saturated cache sets and that no unnecessary workload is imposed in the absence of contention. Flex-CAT scans the entire cache set and determines the index of the LRU CL, its elapsed time, and the occupancy count for each priority of the system.

[0042] If the claimant's occupancy count is less than the minimum allocation (O[PF]<PF[mnw]), Flex-CAT preferentially evicts the LP LRU CL to increase its occupancy count. If the claimant's occupancy count reaches the minimum allocation and is less than the maximum allocation (PF[mnw]≦O[PF]<PF[mxw]), Flex-CAT searches for the LRU target among the priorities and further adds the LRU CL to the candidate list. When the claimant's occupancy count reaches the maximum allocation, Flex-CAT ignores the LP LRU candidates and selects the claimant's LRU (LRUF) as the target CL, ensuring that the claimant's occupancy count never exceeds the upper limit (PF[mxw]). If no target is found in the previous step (if all lines belong to owners with higher priorities), Flex-CAT performs an HP eviction.

[0043] Figures 5 and 6 and the flowchart described below show the steps that Flex-CAT goes through after indexing to the appropriate cache set.

[0044] FIG. 1 is a block diagram showing processing components that execute instructions, according to some embodiments. As shown, storage 101 stores instructions (s) 103 to be executed. As will be further described below, in some embodiments, system 100 (also referred to herein as a "computing system") is a SIMD processor that simultaneously processes multiple elements of a packed data vector, including matrices.

[0045] During operation, instructions (s) 103 are fetched from storage 101 by fetch circuit 105. The instructions are decoded by decode circuit 109. Decode circuit 109 decodes the fetched instruction 107 into one or more operations. In some embodiments, this decoding includes generating a plurality of micro-operations to be executed by an execution circuit (such as execution circuit 117). Decode circuit 109 further decodes instruction suffixes and prefixes (when used).

[0046] In some embodiments, register naming, register allocation, and / or scheduling circuit 113 provides functions for one or more of the following: 1) renaming logical operand values to physical operand values (e.g., a register alias table in some embodiments), 2) allocating status bits and flags to decoded instructions, 3) scheduling decoded instruction 111 in the instruction pool for execution on execution circuit 117 (e.g., using a reservation station in some embodiments).

[0047] Register (register file) and / or memory 115 stores data as operands of instruction 111 to be executed by execution circuit 117. Exemplary register types are further described and shown below with reference to at least FIG. 9 and include write mask registers, packed data registers, general purpose registers, and floating point registers.

[0048] In some embodiments, the write-back circuit 119 commits the instruction execution result. The execution circuit 117 and the system 100 are further illustrated and described with reference to FIGS. 2-4, FIGS. 10A, 10B, FIGS. 11A, 11B.

[0049] FIG. 2 is a block diagram illustrating a system including a multi-core processor that executes virtual machines according to some embodiments. As shown, the computing system 200 includes a multi-core processor 202 including cores 0 206A, 1 206B, … N 206N that share a last-level cache LLC 204. Along with this, the resources of the processor 202 can function as part of a computing platform of a cloud service provider (CSP) to provide network services to one or more clients. For example, as shown, cores 0, 1, to N support VMs 0 210A, 1 210B, … N 210N. VM 0 210A supports a VNF application 0 212A (virtual network function application) and a guest OS 0 214A. Similarly, VM 1 210B supports a VNF application 1 212B and a guest OS 1 214B. Similarly, VM N 210N supports a VNF application 212N and a guest OS N 214N. During operation, the virtual machines are launched and managed using a hypervisor / VMM 208. Further, an operating system 214 can be called to manage the system.

[0050] In some embodiments, the cache control circuit 201 cooperates with the hypervisor / VMM 208 to implement the cache partitioning scheme described herein.

[0051] In some embodiments, the cache monitoring circuit 203 maintains statistics and heuristics related to cache access requests, such as the rate of low-priority cache fill requests that cause an eviction of a high-priority cache line. The cache monitoring circuit 203 can alternatively be incorporated into the processor 202, as indicated by being enclosed in dashed lines, and can be any one. In some examples, the computing system 200 is a stand-alone computing platform, but in another example, it is coupled to another computing platform via a network (not shown).

[0052] In some embodiments, the computing system 200 is a node within a data center and supports VMs that individually execute one or more VNF applications. Such applications include, for example, cloud service providers, database network services, website hosting services, routing network services, email services, firewall services, domain name services (DNS), caching services, network address translation (NAT) services, or virus scan network services. The VMs 210A - 210N in the computing system 200 can be managed or controlled by a hypervisor or virtual machine manager (VMM), such as the hypervisor / VMM 208. In other embodiments, the computing system 200 can be configured as a more conventional server having various computing resources described above housed within the same physical enclosure, chassis, or container.

[0053] According to some embodiments, a virtual machine is a software computer that, like a physical computer, executes an operating system and applications. Some virtual machines are configured by a set of configuration files and assisted by the host's physical resources. Also, a hypervisor or VMM is computer software, firmware, or hardware that creates and manages virtual machines. A computer on which a hypervisor runs one or more virtual machines is called a host machine, and each virtual machine is called a guest machine. The hypervisor or VMM presents a guest operating system with a virtual operating platform and manages the execution of the guest operating system. Multiple instances of various operating systems can share virtualized hardware resources. For example, Linux®, Windows®, and macOS® instances can all operate on a single multi-core physical processor.

[0054] In some examples, as shown in FIG. 2, at least a portion of the computing resources for computing system 200 can include processing elements such as CPU / cores 206A, 206B,... 206N having a shared last level cache (LLC) 204.

[0055] In some examples, LLC 204 is external to processor 202. According to some examples, the shared LLC 204 can be a relatively fast access memory type in order to function as a shared LLC for CPUs / Cores 206A - 206N to minimize access latency. The relatively fast access memory types included in shared LLC 204 can include, but are not limited to, volatile or non - volatile memory. Types of volatile memory can include, but are not limited to, static random access memory (SRAM), dynamic random access memory (DRAM), thyristor RAM (TRAM), or zero - capacitor RAM (ZRAM). Types of non - volatile memory can include, but are not limited to, byte or block - addressable non - volatile memory types having a three - dimensional (3D) cross - point memory structure that includes a chalcogenide phase - change material (e.g., chalcogenide glass) (hereinafter referred to as "3D cross - point memory"). Types of non - volatile memory can further include other types of byte or block - addressable non - volatile memory. Examples of this include, but are not limited to, multi - threshold NAND flash memory, NOR flash memory, single or multi - phase change memory (PCM), resistive memory, nanowire memory, ferroelectric transistor random access memory (FeTRAM), magnetic random access memory (MRAM) incorporating memristor technology, spin - transfer torque MRAM (STT - MRAM), or any combination of the above.

[0056] FIG. 3 shows an exemplary LLC cache partitioning scheme according to some embodiments. As shown, scheme 300 is an example of a scheme that can be used as the shared LLC 204 of computing system 200 as shown in FIG. 2. Here, LLC 304 is an 8 - way set where each cache way 302 includes 8 cache lines. Associativeis shown as a cache. Some of the cache lines of LLC304 are allocated to the low-priority application 306, some are allocated to the high-priority application 310, and some are invalid 308. For the sake of brevity in the description, the shared LLC354 is an 8-way set where each way contains one cache line Associative is shown as a cache. Some of the cache lines of LLC354 are allocated to the low-priority application 356, some are allocated to the high-priority application 360, and some are invalid 358. What is illustrated in FIG. 3 does not limit the disclosed embodiments to a specific configuration. Another approach may include more or fewer sets, more or fewer ways within a set, more or fewer cache lines within each way, etc. For example, LLC204 is an N-way set where each set contains M cache lines Associative can be a cache. Here, N and M are positive integers greater than or equal to 1.

[0057] During operation, as will be described further below, LLC204 is dynamically re-partitioned as needed by the applications sharing the LLC. As an advantage of the disclosed embodiments, the low-priority applications seek to minimize the eviction of cache lines allocated to the high-priority applications.

[0058] FIG. 4 is a block diagram showing cache line eviction according to some embodiments of the Flex-CAT algorithm. According to the disclosed embodiments, Flex-CAT is an eviction algorithm that performs priority-based cache partitioning at the cache set granularity. As shown, scheme 400 shows applications with priorities falling within the range of priority 402. Arcs 404, 406, 408 are subsequent cache lines from the requesting core, showing cache line eviction to make room for requests to fill the cache line. Some evictions, such as eviction 408, make room for higher priority assignments by evicting lower priority assignments. Some evictions are self-evictions, shown by arc 406 (e.g., the priority that has already assigned the maximum number of ways self-evicts to make room for subsequent CL). When the occupancy of the requester reaches the maximum assignment, Flex-CAT prioritizes self-eviction over other priority evictions to stay within the partition boundary. In limited situations where the target cannot be found in these first two steps, Flex-CAT selects the HP CL for eviction to make room for subsequent cache fills. Flex-CAT aims to maximize the eviction of LP CLs such as eviction 408 and minimize the eviction of HP CLs such as eviction 404.

[0059] FIG. 5 is a block diagram showing processing executed by a cache control circuit in response to a cache fill request according to some embodiments. For example, flow 500 can be executed by the cache control circuit (CCC) 201 of FIG. 2. As shown, flow 500 starts at 501, and the CCC is assumed to receive a subsequent cache fill request having a requester priority from the requester. For example, the requester can be one of the cores 206A, 206B,... 206N of FIG. 2, and a subsequent cache line can be obtained from a memory such as the memory 115 of FIG. 1. The requester priority can reflect the priority assigned to an application operating within the requester core. In operation 502, the CCC is assumed to determine whether there is even one invalid cache line (CL) in the LLC. If present, in operation 504, the CCC is assumed to write the subsequent CL to the storage location of the invalid cache line, and the flow ends when the storage location is found at 505. However, if it is shown that there is no invalid CL in operation 502, in 506 the CCC is assumed to determine, for each priority in the system, the priority (P), the index of the LRU CL having the priority P (LRUp), the elapsed time of LRUp, and the number of ways occupied by the priority. In operation 508, the CCC is assumed to determine whether the requester priority (P R ) is the lowest priority. If it is the lowest priority, in operation 510 the CCC is assumed to determine whether the occupancy of the requester priority (O[P R ) is equal to 0. If equal, the flow proceeds to operation 524. Here, a higher priority CL (P R ) is evicted to make room for the subsequent CL (P H ).

[0060] Operation 524 indicates a situation where a high-priority cache line is evicted to make room for lower-priority cache lines. This is a situation that the disclosed embodiments seek to avoid as much as possible, except in the cases described above. In some embodiments, a cache monitoring circuit, such as cache monitoring circuit 203 of FIG. 2, maintains a heuristic tracking response to cache fill requests that includes instances of operation 524. In some embodiments, the CCC monitors the heuristics and dynamically adjusts the minimum and maximum ways assigned to each priority as needed. This adjusts the aggressiveness of higher-priority applications with respect to the eviction of lower-priority applications that could ultimately lead to operation 524. In some embodiments, the cache monitoring circuit 203 causes the LLC to be re-partitioned. Or in some embodiments, if the maintained heuristics, such as the occurrence of operation 524, exceed a predetermined threshold, the cache monitoring circuit causes the way boundaries associated with the various priorities to be adjusted. For example, the number of ways assigned to a high-priority application can be reduced to reduce the repeated, aggressive, and complete eviction of lower-priority applications that could lead to repeated execution of operation 524. If the CCC determines in operation 510 that the requester occupancy count is not 0, the CCC shall evict the least recently used CL of the requester priority in operation 514.

[0061] Returning to operation 508, if the CCC determines that the requester priority is not the lowest priority, the CCC shall determine in operation 512 whether the occupancy count of the requester priority is the maximum. If it is the maximum, the CCC shall evict the least recently used CL of the requester priority in operation 514. The flow ends when a storage location is found in operation 515.

[0062] When returning to operation 512, if CCC determines that the occupancy count of the requester priority is not the maximum, then in operation 516, CCC determines whether the occupancy count of the requester priority is less than the maximum and is greater than or equal to the minimum for the requester priority. If so, the flow proceeds to operation 518; otherwise, the flow proceeds to operation 520. In operation 518, CCC is assumed to evict the LRU CL having a priority equal to the requester priority (P R ) or a lower priority (P L ), and when a storage location is found in 519, the flow ends. In operation 520, CCC attempts to evict the LRU CL with a priority lower than the requester priority. If such a line exists, the storage location at 522 is determined to be non-NULL, and the flow proceeds to 523, where the flow ends when a storage location is found. If there is no such CL, the storage location at 522 becomes equal to NULL, and the flow proceeds to operation 524 to evict the LRU CL having a priority higher than the requester priority. Then, the flow ends when a storage location is found in 525.

[0063] FIG. 6 is a flow diagram showing a method executed by a cache control circuit (CCC) that processes cache fill requests, according to some embodiments. For example, flow 600 can be executed by cache control circuit (CCC) 201 of FIG. 2. As shown, in operation 605, the CCC is assumed to receive a request to store a subsequent cache line (CL) having a certain requester priority among a plurality of priorities in the last level cache (LLC). In operation 610, if there is an invalid cache line (CL) in the LLC, the subsequent cache line (CL) is stored in the invalid CL. In operation 615, if the requester priority is the lowest among the plurality of priorities, has an occupancy number of 1 or more, or the occupancy number is the largest for the requester priority, the subsequent CL is stored instead of the least recently used (LRU) CL of the requester priority. In operation 620, if the occupancy number is between the maximum and the minimum for the requester priority, the subsequent CL is stored instead of the LRU CL of the requester priority or a lower priority. In operation 625, if the occupancy number is less than the minimum and there is a CL with a lower priority, the subsequent CL is stored instead of the LRU CL with a lower priority. In operation 630, if there is no invalid CL, or a CL with the requester priority or a lower priority, the subsequent CL is stored instead of the LRU CL with a higher priority. [Instruction Set]

[0064] The instruction set may include the format of one or more instructions. A given instruction format may determine, among other things, the operation to be performed (e.g., opcode) and various fields (e.g., number of bits, bit positions) and / or other data fields (if any) (e.g., mask) for specifying the operand(s) on which the operation is to be performed. Some instruction formats are further classified by the definition of instruction templates (or sub-formats). For example, an instruction template of a particular instruction format may be defined to have different subsets of the fields of the instruction format (the fields included are usually in the same order, but at least some have different bit positions because they have fewer fields included), and / or may be defined to have particular fields that are interpreted differently. Thus, each instruction of the ISA is represented using a particular instruction format (or, if defined, a particular one of the instruction templates of that instruction format) and includes fields for specifying the operation and operands. For example, an exemplary ADD instruction has an instruction format that includes a particular opcode, an opcode field for specifying that opcode, and operand fields for selecting the operands (source1 / destination / source2). When this ADD instruction appears in an instruction stream, it will have particular contents in the operand fields that select the particular operands. A series of SIMD extensions, referred to as Advanced Vector Extensions (AVX) (AVX1 and AVX2) and using the Vector Extension (VEX) coding scheme, have been released and / or published (see, e.g., the Intel® 64 and IA-32 Architectures Software Developers Manual in September 2014 and the Intel® Advanced Vector Extensions Programming Reference in October 2014). [Exemplary Instruction Format]

[0065] The embodiments of the instructions described in this specification may be embodied in different formats. Further, exemplary systems, architectures, and pipelines are detailed below. Multiple embodiments of the instructions may be executed on such multiple systems, multiple architectures, and multiple pipelines, but are not limited to these details. [Instruction Format for General-Purpose Vectors]

[0066] The vector-oriented instruction format is an instruction format suitable for vector instructions (for example, there are specific fields unique to vector operations). Embodiments that support both vector operations and scalar operations via the vector-oriented instruction format are described, but alternatively, there are also embodiments that use only vector operations with the vector-oriented instruction format.

[0067] FIGS. 7A and 7B are block diagrams showing a general-purpose vector-oriented instruction format and its instruction templates according to some embodiments of the present invention. FIG. 7A is a block diagram showing a general-purpose vector-oriented instruction format and its class A instruction template according to some embodiments of the present invention, while FIG. 7B is a block diagram showing a general-purpose vector-oriented instruction format and its class B instruction template according to some embodiments of the present invention. Specifically, the general-purpose vector-oriented instruction format 700 defined for the class A and class B instruction templates includes, for both classes, an instruction template for non-memory access 705 and an instruction template for memory access 720. The term "general-purpose" in the context of the vector-oriented instruction format relates to an instruction format that is not associated with any specific instruction set.

[0068] Embodiments of the present invention are described, where the vector-oriented instruction format supports the following. That is, a 64-byte vector operand length (or size) having a 32-bit (4-byte) or 64-bit (8-byte) data element width (or size) (thus, a 64-byte vector is composed of 16 elements of double-word size or, alternatively, 8 elements of quad-word size), a 64-byte vector operand length (or size) having a 16-bit (2-byte) or 8-bit (1-byte) data element width (or size), a 32-byte vector operand length (or size) having a 32-bit (4-byte), 64-bit (8-byte), 16-bit (2-byte), or 8-bit (1-byte) data element width (or size), and a 16-byte vector operand length (or size) having a 32-bit (4-byte), 64-bit (8-byte), 16-bit (2-byte), or 8-bit (1-byte) data element width (or size). Alternative embodiments may support larger vector operand sizes, smaller vector operand sizes, and / or different vector operand sizes (e.g., 256-byte vector operands) having larger data element widths, smaller data element widths, or different data element widths (e.g., 128-bit (16-byte) data element width).

[0069] The class A instruction templates in FIG. 7A include: 1) a non-memory access full-round control type operation 710 instruction template and a non-memory access data conversion type operation 715 instruction template shown within the non-memory access 705 instruction template, and 2) a memory access temporary 725 instruction template and a memory access non-temporary 730 instruction template shown within the memory access 720 instruction template. The class B instruction templates in FIG. 7B include: 1) a non-memory access write mask control partial round control type operation 712 instruction template and a non-memory access write mask control VSIZE type operation 717 instruction template shown within the non-memory access 705 instruction template, and 2) a memory access write mask control 727 instruction template shown within the memory access 720 instruction template.

[0070] The general-purpose vector instruction format 700 includes the fields shown below in the order shown in FIGS. 7A and 7B.

[0071] Format field 740 - The specific value (instruction format identifier value) within this field uniquely identifies the vector instruction format, and thus the occurrence of instructions within the vector instruction format in the instruction stream. Thus, this field is optional in the sense that it is not necessary for an instruction set that has only the general-purpose vector instruction format.

[0072] Base operation field 742: Its content identifies different base operations.

[0073] Register index field 744: Depending on its content, it directly or through address generation specifies the locations of source and destination operands whether they are in registers or in memory. These include sufficient bits to select N registers from P×Q (e.g., 32×512, 16×128, 32×1024, 64×1024) register files. In one embodiment, N can be up to 3 sources and 1 destination register, although alternative embodiments may support more or fewer source and destination registers (e.g., may support up to 2 sources, in which case one of these sources also functions as a destination, may support up to 3 sources, in which case one of these sources also functions as a destination, may support up to 2 sources and 1 destination).

[0074] Modifier field 746: Depending on its content, it distinguishes the appearance of an instruction that specifies a memory access in the general-purpose vector instruction format from an instruction that does not specify a memory access. That is, it distinguishes between the non-memory access 705 instruction template and the memory access 720 instruction template. The memory access operation performs a read and / or write to the memory hierarchy (optionally using values in registers to specify source and / or destination addresses), while the non-memory access operation does not (e.g., source and destination are registers). Also, in one embodiment, this field selects from 3 separate ways to perform the memory address calculation, although alternative embodiments may support more, fewer, or different ways to perform the memory address calculation.

[0075] Expansion operation field 750: Depending on its content, in addition to the base operation, it distinguishes which of various different operations is to be executed. This field is specific to the context. In some embodiments, this field is divided into a class field 768, an alpha field 752, and a beta field 754. The expansion operation field 750 enables a common operation group to be performed with a single instruction instead of two, three, or four instructions.

[0076] Scale field 760: Depending on its content, it enables scaling of the content of the index field for memory address generation (e.g., for address generation using 2 scale* index + base).

[0077] Displacement field 762A: Its content is used as part of memory address generation (e.g., for address generation using 2 scale* index + base + displacement).

[0078] Displacement coefficient field 762B (note that the displacement field 762A being juxtaposed directly above the displacement coefficient field 762B indicates that one or the other is to be used): Its content is used as part of address generation, and it specifies a displacement coefficient that is to be scaled by the size of the memory access (N). Here, N is the number of bytes in the memory access (e.g., 2 scale*(For address generation using index + base + scaled displacement). The redundant lower bits are ignored, and thus the content of the displacement coefficient field is multiplied by the total size (N) of the memory operand to generate the final displacement used in the calculation of the effective address. The value of N is determined by the processor hardware at runtime based on the full opcode field 774 (described later in this specification) and the data manipulation field 754C. The displacement field 762A and the displacement coefficient field 762B are optional in the sense that they are not used in the instruction template for non-memory access 705 and / or different embodiments can implement only one of these two or may not implement either at all.

[0079] Data element width field 764: Its content distinguishes which of a plurality of data element widths is used (in some embodiments for all instructions, in other embodiments for only some of the instructions). This field is optional in the sense that it is not needed when only one data element width is supported and / or the data element width is supported using some aspect of the opcode.

[0080] Write mask field 770: Depending on its content, for each data element position, it controls whether the data element position in the destination vector operand reflects the results of the base operation and the extension operation. Class A instruction templates support a merge write mask, while class B instruction templates support both a merge write mask and a zeroing write mask. In the case of merge, the vector mask enables any set of elements in the destination to be protected from being updated during the execution of any operation (specified by the base operation and the extension operation). In another embodiment, when the corresponding mask bit has a value of 0, the old value of each element of the destination is retained. In contrast, in the case of zeroing, the vector mask enables any set of elements in the destination to be zeroed during the execution of any operation (specified by the base operation and the extension operation). In one embodiment, when the corresponding mask bit has a value of 0, the elements of the destination are set to 0. A subset of this functionality is the ability to control the vector length of the operation being performed (i.e., the span of elements that are changed from the first to the last), although the elements that are changed do not need to be contiguous. Thus, the write mask field 770 enables partial vector operations including load, store, arithmetic, logical, etc. Embodiments of the present invention describe that the content of the write mask field 770 selects which of a plurality of write mask registers should be used (thus, indirectly identifying the masking to be performed based on the content of the write mask field 770), although alternative embodiments alternatively or additionally enable the content of the write mask field 770 to directly specify the masking to be performed.

[0081] Immediate value field 772: Depending on its content, it enables the specification of an immediate value. This field is optional in the sense that it does not exist in implementations of general-purpose vector formats that do not support immediate values and does not exist in a plurality of instructions that do not use immediate values.

[0082] Class field 768: Based on its content, it distinguishes between instructions of different classes. Referring to FIGS. 7A and 7B, the content of this field selects between class A instructions and class B instructions. In FIGS. 7A and 7B, the rounded rectangles are used to indicate that a specific value exists in the field (e.g., for class field 768 in FIGS. 7A and 7B, class A 768A and class B 768B respectively). [Instruction template for class A]

[0083] In the case of the class A non - memory - access 705 instruction template, the alpha field 752 is interpreted as the RS field 752A, and based on its content, it distinguishes which of different extended operation types is to be executed (e.g., round 752A.1 and data conversion 752A.2 are specified for the non - memory - access round - type operation 710 and non - memory - access data - conversion - type operation 715 instruction templates respectively), while the beta field 754 distinguishes which of the specified type of operations is to be executed. The non - memory - access 705 instruction template does not have a scale field 760, a displacement field 762A, and a displacement coefficient field 762B. [Non - memory - access instruction template - full - round control - type operation]

[0084] In the non-memory access full-round control type operation 710 instruction template, the beta field 754 is interpreted as a round control field 754A, and its content provides static rounding. In the embodiments described in the present invention, the round control field 754A includes a Suppress All floating-point Exceptions (SAE) field 756 and a round operation control field 758. On the other hand, alternative embodiments may encode both of these concepts in the same field, or alternative embodiments may have only one or the other of these concepts / fields (for example, may have only the round operation control field 758).

[0085] SAE field 756: Its content distinguishes whether to disable exception event reporting. When the content of the SAE field 756 indicates that suppression is enabled, a specific instruction does not report any type of floating-point exception flag and does not generate a floating-point exception handler.

[0086] Round operation control field 758: Its content distinguishes which of a group of round operations (for example, rounding up, rounding down, rounding to zero, and rounding to an approximate value) is executed. Therefore, the round operation control field 758 enables the change of the round mode on an instruction unit basis. In some embodiments where the processor includes a control register for specifying the round mode, the content of the round operation control field 750 overwrites the register value. [Non-memory access instruction template - data conversion type operation]

[0087] In the non-memory access data conversion type operation 715 instruction template, the beta field 754 is interpreted as a data conversion field 754B, and the content of the data conversion field 754B distinguishes which of a plurality of data conversions (for example, no data conversion, swizzle, broadcast) is executed.

[0088] In the case of the memory access 720 instruction template of class A, the alpha field 752 is interpreted as the revision hint field 752B, and its content distinguishes which of the revision hints should be used (in FIG. 7A, transient 752B.1 and non-transient 752B.2 are specified for the memory access transient 725 instruction template and the memory access non-transient 730 instruction template, respectively), while the beta field 754 is interpreted as the data operation field 754C, and its content distinguishes which of a plurality of data operation operations (also known as primitives) are to be executed (e.g., no operation, broadcast, source up-conversion, and destination down-conversion). The memory access 720 instruction template includes a scale field 760 and optionally includes a displacement field 762A or a displacement coefficient field 762B.

[0089] Vector memory instructions, with conversion support, perform vector loads from memory and vector stores to memory. As with normal vector instructions, vector memory instructions transfer data from / to memory for an entire data element, and the elements actually transferred are described by the content of a vector mask selected as a write mask. [Memory Access Instruction Template: Transient]

[0090] Transient data is data that is likely to be reused soon enough to benefit from caching. However, this is a hint, and different processors can implement it in different ways, including completely ignoring the hint. [Memory Access Instruction Template: Non-Transient]

[0091] Non-temporary data is data that is not likely to be reused soon enough to gain the benefits of caching in the first-level cache, and this should be given the priority of eviction. However, this is a hint, and different processors can implement it in different ways, including completely ignoring the hint. [Instruction template for Class B]

[0092] In the case of the instruction template for Class B, the alpha field 752 is interpreted as a write mask control (Z) field 752C, and depending on the content of the write mask control (Z) field 752C, it distinguishes whether the write masking controlled by the write mask field 770 should be merging or zeroing.

[0093] In the case of the non-memory access 705 instruction template for Class B, a part of the beta field 754 is interpreted as an RL field 757A, and depending on its content, it distinguishes which of different extended operation types is to be executed (for example, round 757A.1 and vector length (VSIZE) 757A.2 are specified for the non-memory access write mask control partial round control type operation 712 instruction template and the non-memory access write mask control VSIZE type operation 717 instruction template respectively). On the other hand, the remainder of the beta field 754 distinguishes which of the specified type of operations is to be executed. The non-memory access 705 instruction template does not have a scale field 760, a displacement field 762A, and a displacement coefficient field 762B.

[0094] In the non-memory access write mask control partial round control type operation 712 instruction template, the remainder of the beta field 754 is interpreted as a round operation field 759A, and exception event reporting is disabled (certain instructions do not report any type of floating-point exception flag and do not generate a floating-point exception handler).

[0095] Rounding operation control field 759A: Similar to the rounding operation control field 758, depending on its content, it differentiates which one of a group of rounding operations is to be executed (e.g., rounding up, rounding down, rounding to zero, and rounding to an approximate value). Therefore, the rounding operation control field 759A enables the change of the rounding mode in instruction units. In some embodiments, if the processor includes a control register for specifying the rounding mode, the content of the rounding operation control field 750 overwrites the register value.

[0096] In the non-memory access write mask control VSIZE type operation 717 instruction template, the remainder of the beta field 754 is interpreted as the vector length field 759B, and depending on the content of the vector length field 759B, it differentiates which one of multiple data vector lengths (e.g., 128, 256, or 512 bytes) is to be executed.

[0097] In the case of the class B memory access 720 instruction template, a part of the beta field 754 is interpreted as the broadcast field 757B, and depending on its content, it differentiates whether the operation of the broadcast type data operation is to be executed, while the remainder of the beta field 754 is interpreted as the vector length field 759B. The memory access 720 instruction template includes a scale field 760 and optionally includes a displacement field 762A or a displacement coefficient field 762B.

[0098] Regarding the general-purpose vector instruction format 700, the full opcode field 774 is shown to include the format field 740, the base operation field 742, and the data element width field 764. Although one embodiment where the full opcode field 774 includes all of these fields is shown, in embodiments that do not support all of these fields, the full opcode field 774 includes fewer fields than all of these fields. The full opcode field 774 provides an operation code (opcode).

[0099] The extended operation field 750, the data element width field 764, and the write mask field 770 enable these features to be specified on a per-instruction basis in the general-purpose vector instruction format.

[0100] The combination of the write mask field and the data element width field forms a typed instruction in that it enables masks to be applied based on different data element widths.

[0101] The various instruction templates found within Class A and Class B are beneficial in different situations. In some embodiments of the present invention, different multiple processors or different cores within a processor may support only Class A, only Class B, or both classes. For example, a high-performance general-purpose out-of-order core for general-purpose computing may support only Class B, a core mainly for graphics and / or scientific (throughput) computing may support only Class A, and a core for both may support both (of course, a core having some combination of templates and instructions of both classes, but not all templates and instructions of both classes, belongs to the scope of the present invention). Also, a single processor may include multiple cores, all of which support the same class, or different ones of them support different classes. For example, in a processor having separate graphics and general-purpose cores, one of the graphics cores mainly for graphics and / or scientific computing may support only Class A, while one or more of the general-purpose cores may be high-performance general-purpose cores using out-of-order execution and register renaming for general-purpose computing that support only Class B. Another processor without separate graphics cores may include one or more general-purpose in-order or out-of-order cores that support both Class A and Class B. Of course, in different embodiments of the present invention, functions belonging to one class may be implemented in the other class. Multiple programs written in a high-level language are in various different executable forms, including 1) a form having only instructions of the class supported by the target processor for execution, or 2) a form having alternative multiple routines described using different combinations of multiple instructions of all classes and having control flow code for selecting one of the multiple routines to execute based on the multiple instructions supported by the processor currently executing the code (e.g., compiled just-in-time or statically compiled). [Exemplary Specific Vector-Oriented Instruction Format]

[0102] Figure 8A is a block diagram showing an exemplary specific vector - oriented instruction format according to some embodiments of the present invention. Figure 8A shows a specific vector - oriented instruction format 800, which is specific in that it specifies the position, size, interpretation, and order of fields, as well as the values for some of these fields. The specific vector - oriented instruction format 800 may be used to extend the x86 instruction set, and thus, some of the fields are similar or identical to the fields used in the existing x86 instruction set and its extensions (e.g., AVX). This format maintains consistency with the prefix - encoding field, real - opcode byte field, MOD R / M field, SIB field, displacement field, and immediate - value field of an existing x86 instruction set with some extensions. The fields from Figure 7A or Figure 7B to which the fields from Figure 8A are mapped are illustrated.

[0103] Although embodiments of the present invention are described with respect to the specific vector - oriented instruction format 800 in the context of the general - purpose vector - oriented instruction format 700 for illustrative purposes, it should be understood that the present invention is not limited to the specific vector - oriented instruction format 800 except as claimed. For example, although the specific vector - oriented instruction format 800 is shown to have fields of a particular size, the general - purpose vector - oriented instruction format 700 contemplates various possible sizes for the various fields. As a particular example, the data - element width field 764 is shown as a 1 - bit field in the specific vector - oriented instruction format 800, but this does not limit the present invention (i.e., the general - purpose vector - oriented instruction format 700 contemplates data - element width fields 764 of other sizes).

[0104] The specific vector - oriented instruction format 800 includes the following fields listed below in the order shown in Figure 8A.

[0105] The EVEX prefix (bytes 0-3) 802. This is encoded in 4-byte form.

[0106] Format field 740 (EVEX byte 0, bits [7:0]). The first byte (EVEX byte 0) is format field 740, and format field 740 contains 0x62 (a unique value used in some embodiments to distinguish vector-oriented instruction formats).

[0107] The second through fourth bytes (EVEX bytes 1-3) contain a plurality of bit fields that provide specific functionality.

[0108] The REX field 805 (EVEX byte 1, bits [7-5]): Composed of the EVEX.R bit field (EVEX byte 1, bit [7] - R), the EVEX.X bit field (EVEX byte 1, bit [6] - X), and the EVEX.B bit field (EVEX byte 1, bit [5] - B). The EVEX.R bit field, EVEX.X bit field, and EVEX.B bit field provide the same functionality as the corresponding VEX bit fields, and they are encoded using one's complement form, i.e., ZMM0 is encoded as 1111B and ZMM15 is encoded as 0000B. Other fields of the instruction encode the lower 3 bits (rrr, xxx, and bbb) of the register index, so by adding EVEX.R, EVEX.X, and EVEX.B, Rrrr, Xxxx, and Bbbb can be formed.

[0109] REX'810A: This is the first part of the REX' field 810 and is the EVEX.R' bit field (EVEX byte 1, bit [4] - R') used to encode either the upper 16 or lower 16 of the extended 32 - register set. In some embodiments, along with other things shown below, this bit is stored in bit - inverted format and distinguished from the BOUND instruction (in the well - known x86 32 - bit mode). The real opcode byte of the BOUND instruction is 62, but within the MOD R / M field (described later), the MOD field does not accept the value 11. Alternative embodiments of the present invention do not store this bit and other bits described later in bit - inverted format. The value 1 is used to encode the lower 16 registers. In other words, combining EVEX.R', EVEX.R, and other RRRs of other fields forms R'Rrrr.

[0110] Opcode map field 815 (EVEX byte 1, bits [3:0] - mmmm): Its content encodes the indicated leading opcode byte (0F, 0F38, or 0F3).

[0111] Data element width field 764 (EVEX byte 2, bit [7] - W) is denoted as EVEX.W. EVEX.W is used to define the granularity (size) of the data type (32 - bit data element or 64 - bit data element).

[0112] EVEX.vvvv field 820 (EVEX byte 2, bits [6:3] - vvvv). The role of EVEX.vvvv may include the following. 1) EVEX.vvvv encodes the first source register operand in the specified inverted (one's complement) form, and EVEX.vvvv is valid for instructions with two or more source operands. 2) EVEX.vvvv encodes the destination register operand in the specified one's complement form for a particular vector shift. Or 3) EVEX.vvvv does not encode any operand, and the field should be reserved and contain 1111b. Therefore, the EVEX.vvvv field 820 encodes the four lower bits of the first source register specifier stored in inverted form (one's complement). Depending on the instruction, additional different EVEX bit fields are used to extend the specifier size to 32 registers.

[0113] EVEX.U 768 class field (EVEX byte 2, bit [2] - U): When EVEX.U = 0, it indicates class A or EVEX.U0. When EVEX.U = 1, it indicates class B or EVEX.U1.

[0114] Prefix Encoding Field 825 (EVEX byte 2, bits [1:0] - pp): Provides additional bits for the base operation field. In addition to providing support for legacy SSE instructions in the EVEX prefix format, this also has the advantage of compacting SIMD prefixes (the EVEX prefix requires only 2 bits instead of 1 byte to represent the SIMD prefix). In one embodiment, in both the legacy format and the EVEX prefix format, to support legacy SSE instructions that use SIMD prefixes (66H, F2H, F3H), these legacy SIMD prefixes are encoded in the SIMD prefix encoding field. These legacy SIMD prefixes are expanded at runtime to the legacy SIMD prefixes before being provided to the PLA of the decoder (thus, the PLA can execute both the legacy format and the EVEX format of these legacy instructions without modification). While newer instructions can use the content of the EVEX prefix encoding field directly as an opcode extension, certain embodiments expand in a similar way for consistency, allowing different means specified by these legacy SIMD prefixes. Alternative embodiments can redesign the PLA to support 2-bit SIMD prefix encoding and thus do not require expansion.

[0115] Alpha Field 752 (EVEX byte 3, bit [7] - EH. Also known as EVEX.EH, EVEX.rs, EVEX.RL, EVEX.Write Mask Control, and EVEX.N. Also denoted as α): As described above, this field is context specific.

[0116] Beta Field 754 (EVEX byte 3, bits [6:4] - SSS. Also known as EVEX.S 2-0 , EVEX.r 2-0 , EVEX.rr1, EVEX.LL0, EVEX.LLB. Also denoted as βββ): As described above, this field is context specific.

[0117] REX'810B. This is the remainder of the REX' field 810 and is the EVEX.V' bit field (EVEX byte 3, bit [3] V') that can be used to encode either the upper 16 or lower 16 of the extended 32 - register set. This bit is stored in bit - inverted format. The value 1 is used to encode the lower 16 registers. In other words, by combining EVEX.V' and EVEX.vvvv, V'VVVV is formed.

[0118] Write - mask field 770 (EVEX byte 3, bits [2:0] - kkk): As described above, its content specifies the index of the register in the write - mask register. In some embodiments, a particular value EVEX.kkk = 000 has a special operation that implies that no write - mask is used for a particular instruction (this may be implemented in various ways, including the use of a write - mask that is hard - wired to all 1s or the use of hardware that bypasses masking hardware).

[0119] The real - opcode field 830 (byte 4) is also known as the opcode byte. Part of the opcode is specified in this field.

[0120] The MOD R / M field 840 (byte 5) includes a MOD field 842, a Reg field 844, and an R / M field 846. As described above, the content of the MOD field 842 distinguishes between memory access operations and non-memory access operations. The role of the Reg field 844 can be summarized into two situations: either encoding either a destination register operand or a source register operand, or being treated as an opcode extension and not being used to encode instruction operands. The role of the R / M field 846 may include encoding an instruction operand that references a memory address, or encoding either a destination register operand or a source register operand.

[0121] The Scale, Index, Base (SIB) byte 850 (byte 6) includes an SS852 regarding scale. As described previously, the scale field 760 is used for memory address generation. SIB.xxx854 and SIB.bbb856. The content of these fields has been described with respect to register indices Xxxx and Bbbb.

[0122] Displacement field 762A (bytes 7 - 10): When the MOD field 842 includes 10, bytes 7 - 10 are the displacement field 762A, which functions the same as the legacy 32-bit displacement (disp32) and functions at the byte granularity.

[0123] Displacement Coefficient Field 762B (Byte 7): When the MOD Field 842 contains 01, Byte 7 is the Displacement Coefficient Field 762B. The position of this field is the same as that of the 8-bit displacement (disp8) of the legacy x86 instruction set that functions at the byte granularity. Since disp8 is sign-extended, disp8 can only specify addresses between -128 and 127 byte offsets. For a 64-byte cache line, disp8 uses 8 bits that can only be set to four actually useful values, -128, -64, 0, and 64. Usually, since a wider range is needed, disp32 is used, but disp32 requires 4 bytes. In contrast to disp8 and disp32, the Displacement Coefficient Field 762B is a reinterpretation of disp8. When using the Displacement Coefficient Field 762B, the actual displacement is determined by the content of the Displacement Coefficient Field multiplied by the size of the memory operand access (N). This type of displacement is called disp8*N. This reduces the average instruction length (although a single byte is used for the displacement, it comes with a very large range). Such compressed displacement is based on the premise that the effective displacement is a multiple of the memory access granularity, and thus the redundant lower bits of the address offset do not need to be encoded. In other words, the Displacement Coefficient Field 762B replaces the 8-bit displacement of the legacy x86 instruction set. Therefore, the Displacement Coefficient Field 762B is encoded in the same way as the 8-bit displacement of the x86 instruction set (therefore, there are no changes to the ModRM / SIB encoding rules), with the only exception being that disp8 is overloaded to disp8*N. That is, there are no changes to the encoding rules or the encoding length, and the only change is in the hardware's interpretation of the displacement value (the hardware needs to scale the displacement by the size of the memory operand to obtain the byte-by-byte address offset). The Immediate Field 772 performs the operation as described above. [Full Opcode Field]

[0124] FIG. 8B is a block diagram showing fields that constitute the full opcode field 774 of the specific vector-oriented instruction format 800 according to some embodiments. Specifically, the full opcode field 774 includes a format field 740, a base operation field 742, and a data element width (W) field 764. The base operation field 742 includes a prefix encoding field 825, an opcode map field 815, and a real opcode field 830. [Register Index Field]

[0125] FIG. 8C is a block diagram showing fields that constitute the register index field 744 of the specific vector-oriented instruction format 800 according to some embodiments. Specifically, the register index field 744 includes a REX field 805, a REX' field 810, a MODR / M.reg field 844, a MODR / M.r / m field 846, a VVVV field 820, an xxx field 854, and a bbb field 856. [Extended Operation Field]

[0126] FIG. 8D is a block diagram showing fields that constitute an extended operation field 750 of an instruction format 800 for a specific vector according to an embodiment of the present invention. When the class (U) field 768 contains 0, it represents EVEX.U0 (class A 768A). When the class (U) field 768 contains 1, it represents EVEX.U1 (class B 768B). When U = 0 and the MOD field 842 contains 11 (indicating a non-memory access operation), the alpha field 752 (EVEX byte 3, bit [7]-EH) is interpreted as the rs field 752A. When the rs field 752A contains 1 (round 752A.1), the beta field 754 (EVEX byte 3, bits [6:4] SSS) is interpreted as the round control field 754A. The round control field 754A includes a 1-bit SAE field 756 and a 2-bit round operation field 758. When the rs field 752A contains 0 (data conversion 752A.2), the beta field 754 (EVEX byte 3, bits [6:4] SSS) is interpreted as a 3-bit data conversion field 754B. When U = 0 and the MOD field 842 contains 00, 01, or 10 (indicating a memory access operation), the alpha field 752 (EVEX byte 3, bit [7]-EH) is interpreted as the evacuation hint (EH) field 752B, and the beta field 754 (EVEX byte 3, bits [6:4] SSS) is interpreted as a 3-bit data operation field 754C.

[0127] When U = 1, the alpha field 752 (EVEX byte 3, bit [7]-EH) is interpreted as the write mask control (Z) field 752C. When U = 1 and the MOD field 842 contains 11 (indicating a non-memory access operation), a part of the beta field 754 (EVEX byte 3, bit [4] S 0 ) is interpreted as the RL field 757A. When the RL field 757A contains 1 (round 757A.1), the remainder of the beta field 754 (EVEX byte 3, bits [6-5] S 2-1) is interpreted as the round operation field 759A, while when the RL field 757A contains 0 (VSIZE 757.A2), the remainder of the beta field 754 (EVEX byte 3, bits [6-5]S 2-1 ) is the vector length field 759B (EVEX byte 3, bits [6-5]L 1-0 ) and is interpreted as such. When U = 1 and the MOD field 842 contains 00, 01, or 10 (which means a memory access operation), the beta field 754 (EVEX byte 3, bits [6:4]SSS) is the vector length field 759B (EVEX byte 3, bits [6-5]L 1-0 ) and the broadcast field 757B (EVEX byte 3, bit [4]B) and is interpreted as such. [Exemplary Register Architecture]

[0128] Figure 9 is a block diagram of a register architecture 900 according to some embodiments. The illustrated embodiments include 32 vector registers 910 with a 512-bit width. These registers are referred to as zmm0 through zmm31. The lower 256 bits of the lower 16 zmm registers overlap with the registers ymm0~15. The lower 128 bits of the lower 16 zmm registers (the lower 128 bits of the ymm registers) overlap with the registers xmm0~xmm15. The specific vector-oriented instruction format 800 operates on these overlapping register files as shown in the following table.

Table 2

[0129] In other words, the vector length field 759B selects between a maximum length and one or more other shorter lengths, each such shorter length being half the length of the preceding length, and an instruction template without a vector length field 759B performs an operation on the maximum vector length. Further, in one embodiment, the class B instruction templates of the specific vector-oriented instruction format 800 perform operations on packed single-precision / double-precision floating-point data or scalar single-precision / double-precision floating-point data and packed integer data or scalar integer data. A scalar operation is an operation that is executed at the lowest data element position within a zmm / ymm / xmm register, and the upper data element positions are left in the same state as before the instruction or zeroed, depending on the embodiment.

[0130] In the illustrated embodiment, the write mask register 915 has eight write mask registers (k0 through k7), each 64 bits in size. In an alternative embodiment, the write mask register 915 is 16 bits in size. As noted above, in some embodiments, the vector mask register k0 is not available for use as a write mask. When an encoding that normally indicates k0 is used for a write mask, it selects a hardwired write mask of 0xffff, effectively disabling write masking for that instruction.

[0131] In the illustrated embodiment, the general-purpose register 925 has 16 64-bit general-purpose registers that are used with the existing x86 addressing modes for addressing memory operands. These registers are referred to by the names RAX, RBX, RCX, RDX, RBP, RSI, RDI, RSP, and R8 through R15.

[0132] In the illustrated embodiment, an alias of the scalar floating-point stack register file (x87 stack) 945 is shown as the MMX packed integer flat register file 950, but the x87 stack is an eight-element stack used to perform scalar floating-point operations on 32 / 64 / 80-bit floating-point data using the x87 instruction set extension. The MMX registers are used to perform operations on 64-bit packed integer data, but are also used to hold operands for some operations executed between the MMX registers and the XMM registers.

[0133] In alternative embodiments, wider or narrower registers may be used. Additionally, in alternative embodiments, more, fewer, or different register files and registers may also be used. [Exemplary Core Architecture, Processor, and Computer Architecture]

[0134] Processor cores may be implemented in different processors in different ways for different purposes. For example, the implementation of such cores may include the following. 1) General-purpose in-order cores for general-purpose computing. 2) High-performance general-purpose out-of-order cores for general-purpose computing. 3) Dedicated cores mainly for graphic and / or scientific (throughput) computing. The implementation of different processors may include the following. 1) A CPU including one or more general-purpose in-order cores for general-purpose computing and / or one or more general-purpose out-of-order cores for general-purpose computing, and 2) A coprocessor including one or more dedicated cores mainly for graphics and / or science (throughput). Such different processors result in different computer system architectures, which may include the following. 1) A coprocessor on a chip separate from the CPU. 2) A coprocessor on a different die within the same package as the CPU. 3) A coprocessor on the same die as the CPU (in this case, such a coprocessor may also be called dedicated logic such as integrated graphics and / or scientific (throughput) logic, or a dedicated core). 4) A system-on-chip that may include the CPU described on the same die (which may also be called an application core(s) or application processor(s)), the coprocessor described above, and additional functions. An exemplary core architecture is then described, followed by descriptions of exemplary processors and computer architectures. [Exemplary Core Architecture] [Block Diagrams of In-Order Cores and Out-of-Order Cores]

[0135] FIG. 10A is a block diagram showing both an exemplary in-order pipeline and exemplary register renaming, out-of-order issue / execution pipeline according to some embodiments of the present invention. FIG. 10B is a block diagram showing both an exemplary embodiment of an in-order architecture core and exemplary register renaming, out-of-order issue / execution architecture core included in a processor according to some embodiments of the present invention. The boxes shown in solid lines in FIGS. 10A and 10B illustrate the in-order pipeline and in-order core. On the other hand, any addition of boxes shown in dashed lines illustrates register renaming, out-of-order issue / execution pipeline and core. Since the in-order aspect is a subset of the out-of-order aspect, the out-of-order aspect will be described.

[0136] In FIG. 10A, a processor pipeline 1000 includes a fetch stage 1002, a length decode stage 1004, a decode stage 1006, an allocation stage 1008, a re-naming stage 1010, a scheduling (also known as dispatch or issue) stage 1012, a register read / memory read stage 1014, an execution stage 1016, a write-back / memory write stage 1018, an exception handling stage 1022, and a commit stage 1024.

[0137] FIG. 10B shows a processor core 1090 including a front-end unit 1030 coupled to an execution engine unit 1050, both being coupled to a memory unit 1070. The core 1090 may be a reduced instruction set computing (RISC) core, a complex instruction set computing (CISC) core, a very long instruction word (VLIW) core, or a hybrid or alternative core type. As yet another option, the core 1090 may be a dedicated core such as, for example, a network or communication core, a compression engine, a coprocessor core, a general-purpose computing graphics processing unit (GPGPU) core, a graphics core, etc.

[0138] The front - end unit 1030 includes a branch prediction unit 1032 coupled to an instruction cache unit 1034. The instruction cache unit 1034 is coupled to an instruction translation look - aside buffer (TLB) 1036. The instruction translation look - aside buffer (TLB) 1036 is coupled to an instruction fetch unit 1038. The instruction fetch unit 1038 is coupled to a decode unit 1040. The decode unit 1040 (i.e., decoder) may decode an instruction and may also generate, as outputs, one or more micro - operations, micro - code entry points, micro - instructions, other instructions, or other control signals that are decoded from, or that reflect, or that are derived from the original instruction. The decode unit 1040 may be implemented using a variety of different mechanisms. Examples of suitable mechanisms include, but are not limited to, look - up tables, hardware implementations, programmable logic arrays (PLAs), micro - code read - only memories (ROMs), etc. In one embodiment, the core 1090 includes a micro - code ROM, or other medium storing micro - code for specific macro - instructions (e.g., in the decode unit 1040 or otherwise within the front - end unit 1030). The decode unit 1040 is coupled to a rename / allocator unit 1052 within the execution engine unit 1050.

[0139] The execution engine unit 1050 includes a rename / allocator unit 1052 coupled to a retirement unit 1054 and a set of one or more scheduler units 1056. The scheduler units 1056 represent any number of different schedulers, including a plurality of reservation stations, a central instruction window, etc. The scheduler unit(s) 1056 is coupled to the physical register file unit(s) 1058. Each of the physical register file units 1058 represents one or more physical register files, and each of them stores one or more different data types. Such data types include scalar integers, scalar floating points, packed integers, packed floating points, vector integers, vector floating points, status (e.g., instruction pointer which is the address of the next instruction to be executed), etc. In one embodiment, the physical register file unit 1058 includes a plurality of vector register units, write mask register units, and scalar register units. These register units may provide the vector registers, vector mask registers, and general-purpose registers of the architecture. The physical register file unit 1058 overlaps with the retirement unit 1054 to show various ways in which register renaming and out-of-order execution can be implemented (e.g., using reorder buffers and retirement register files, using future files, history buffers, and retirement register files (multiple available), using multiple register maps and register pools, etc.). The retirement unit 1054 and the physical register file unit(s) 1058 are coupled to the execution cluster(s) 1060. The execution cluster 1060 includes a set of one or more execution units 1062 and a set of one or more memory access units 1064. The execution units 1062 may perform various operations (e.g., shift, add, subtract, multiply) on various types of data (e.g., scalar floating point, packed integer, packed floating point, vector integer, vector floating point).Some embodiments may include a plurality of execution units dedicated to a particular function or set of functions, while other embodiments may include only one execution unit or a plurality of execution units all of which execute all functions. The scheduler unit 1056, physical register file unit 1058, and execution cluster 1060 are shown as plural because a particular embodiment forms separate pipelines (e.g., a scalar integer pipeline, a scalar floating point / packed integer / packed floating point / vector integer / vector floating point pipeline and / or a memory access pipeline. Each of these has its own scheduler unit, physical register file unit, and / or execution cluster. In the case of a separate memory access pipeline, a particular embodiment is implemented in which only the execution cluster of this pipeline has a memory access unit(s) 1064) for a particular type of data / operation. It should also be understood that if separate pipelines are used, one or more of these pipelines may be out-of-order issue / execution and the rest may be in-order.

[0140] A set of multiple memory access units 1064 is coupled to a memory unit 1070, which includes a data TLB unit 1072 coupled to a data cache unit 1074 coupled to a level 2 (L2) cache unit 1076. In one exemplary embodiment, the memory access unit 1064 may include a load unit, a store address unit, and a store data unit, each of which is coupled to the data TLB unit 1072 of the memory unit 1070. The instruction cache unit 1034 is further coupled to the level 2 (L2) cache unit 1076 of the memory unit 1070. The L2 cache unit 1076 is coupled to one or more other levels of cache and ultimately to main memory.

[0141] As an example, an exemplary register naming, out-of-order issue / execution core architecture may implement pipeline 1000 as follows: 1) Instruction fetch unit 1038 executes fetch stage 1002 and length decode stage 1004, 2) Decode unit 1040 executes decode stage 1006, 3) Rename / allocator unit 1052 executes allocation stage 1008 and renaming stage 1010, 4) Scheduler unit 1056 executes schedule stage 1012, 5) Physical register file unit 1058 and memory unit 1070 execute register read / memory read stage 1014, execution cluster 1060 executes execution stage 1016, 6) Memory unit 1070 and physical register file unit 1058 execute write-back / memory write stage 1018, 7) Various units may participate in exception handling stage 1022, and 8) Retirement unit 1054 and physical register file unit 1058 execute commit stage 1024.

[0142] Core 1090 may support one or more instruction sets (e.g., the x86 instruction set (with some extensions with newer versions added), the MIPS instruction set of MIPS Technologies in Sunnyvale, California, the ARM instruction set of ARM Holdings in Sunnyvale, California (with any additional extensions such as NEON added)) including the instructions described herein. In one embodiment, core 1090 includes logic to support packed data instruction set extensions (e.g., AVX1, AVX2), thereby enabling a plurality of operations used by many multimedia applications to be executed using packed data.

[0143] The core can support multithreading (execution of two or more parallel sets of operations or threads), and can support it in various ways including time slice multithreading, simultaneous multithreading (where each of the threads for which a physical core is simultaneously multithreaded is provided with a single physical core as a logical core), or combinations thereof (such as time slice fetch and decode and subsequent simultaneous multithreading in, for example, Intel® Hyper-Threading Technology).

[0144] Although register renaming is described in the context of out-of-order execution, it should be understood that register renaming may be used in an in-order architecture. The illustrated embodiments of the processor may also include a separate instruction cache unit 1034 and data cache unit 1074, as well as a shared L2 cache unit 1076, but multiple alternative embodiments may have a single internal cache for both instructions and data, such as a level 1 (L1) internal cache or multiple levels of internal caches. In some embodiments, the system may include a combination of an internal cache and an external cache that is external to the core and / or the processor. Alternatively, all of the caches may be external to the core and / or the processor. [Specific Exemplary In-Order Core Architectures]

[0145] FIGS. 11A and 11B show block diagrams of more specific exemplary in-order core architectures, where the core would be one of several logical blocks (including other cores of the same type and / or different types) within the chip. The logical blocks communicate with some fixed function logic, memory I / O interfaces, and other required I / O logic through a high bandwidth interconnect network (such as a ring network), depending on the application.

[0146] FIG. 11A is a block diagram of a single processor core with a connection to an on-die interconnect network 1102 and with a local subset 1104 of a level 2 (L2) cache. In one embodiment, instruction decoder 1100 supports the x86 instruction set using a packed data instruction set extension. L1 cache 1106 enables low-latency access to cache memory both within scalar and vector units. In one embodiment, (for simplicity of design) scalar unit 1108 and vector unit 1110 use separate register sets (a plurality of scalar registers 1112 and a plurality of vector registers 1114, respectively), and data transferred between them is written to and then reread from the memory of level 1 (L1) cache 1106, although alternative embodiments of the present invention may use different approaches (e.g., using a single register set or including a communication path that enables data transfer between two register files without writing and rereading).

[0147] The local subset 1104 of the L2 cache is part of a global L2 cache that is divided into separate local subsets, one per processor core. Each processor core has a path for direct access to its own local subset 1104 of the L2 cache. Data read by a processor core is stored in that L2 cache subset 1104 and is quickly accessible in parallel with other processor cores accessing their own local L2 cache subsets. Data written by a processor core is stored in its own L2 cache subset 1104 and is flushed from other subsets if necessary. The ring network guarantees coherence for shared data. The ring network is bi-directional and enables agents such as processor cores, L2 caches, and other logical blocks to communicate with each other within the chip. Each ring data path is 1012 bits wide for each direction.

[0148] Figure 11B is an enlarged view of a part of the processor core of Figure 11A according to some embodiments of the present invention. Figure 11B includes not only more details regarding the vector unit 1110 and the vector register 1114, but also the L1 data cache 1106A which is part of the L1 cache 1106. Specifically, the vector unit 1110 is a 16-wide vector processing unit (VPU) (see 16-wide ALU 1128), and executes one or more of integer instructions, single-precision floating instructions, and double-precision floating instructions. The VPU supports register input swizzling using the swizzle unit 1120, numerical conversion using the numerical conversion units 1122A and 1122B, and replication using the replication unit 1124 for memory input. The write mask register 1126 enables prediction of the resulting vector writes.

[0149] Figure 12 is a block diagram of a processor 1200 according to some embodiments of the present invention, which may have two or more cores, may have an integrated memory controller, and may have integrated graphics. The boxes shown by solid lines in Figure 12 illustrate the processor 1200 and include a single core 1202A, a system agent unit 1210, and a set of one or more bus controller units 1216. The boxes shown by dashed lines illustrate any additional another processor 1200 and include a plurality of cores 1202A to 1202N, a set of one or more integrated memory controller units 1214 within the system agent unit 1210, and dedicated logic 1208.

[0150] Therefore, different implementations of the processor 1200 may include: 1) a CPU that integrates the dedicated logic 1208 as integrated graphics and / or scientific (throughput) logic (which may include one or more cores), and the cores 1202A - 1202N as one or more general-purpose cores (e.g., general-purpose in-order cores, general-purpose out-of-order cores, or a combination of the two); 2) a coprocessor that uses the cores 1202A - 1202N as a number of dedicated cores mainly for graphics and / or scientific (throughput) purposes; and 3) a coprocessor that uses the cores 1202A - 1202N as a number of general-purpose in-order cores. Thus, the processor 1200 may be, for example, a general-purpose processor, coprocessor, or dedicated processor such as a network processor or communication processor, compression engine, graphics processor, GPGPU (general-purpose graphics processing unit), high-throughput many integrated core (MIC) coprocessor (including 30 or more cores), embedded processor, etc. The processor may be implemented on one or more chips. The processor 1200 may be, for example, part of and / or implemented on one or more substrates using any of a plurality of process technologies such as BiCMOS, CMOS, or NMOS.

[0151] The memory hierarchy includes one or more levels of cache within the cores, a set of one or more shared cache units 1206, and an external memory (not shown) coupled to a set of integrated memory controller units 1214. The set of shared cache units 1206 may include one or more intermediate-level caches, such as level 2 (L2), level 3 (L3), level 4 (L4), or other levels of cache, a last-level cache (LLC), and / or combinations thereof. In one embodiment, a ring-based interconnect unit 1212 interconnects the integrated graphics logic 1208 (the integrated graphics logic 1208 is an example of dedicated logic and is also referred to herein as dedicated logic), the set of shared cache units 1206, and the system agent unit 1210 / integrated memory controller units 1214, while alternative embodiments may use any number of well-known techniques for interconnecting such units. In one embodiment, coherence is maintained between one or more cache units 1206 and cores 1202A - 1202N.

[0152] In some embodiments, one or more of cores 1202A - 1202N are capable of multi-threading. The system agent unit 1210 includes these components that coordinate and operate cores 1202A - 1202N. The system agent unit 1210 may include, for example, a power control unit (PCU) and a display unit. The PCU may be or include the logic and a plurality of components required to adjust the power states of cores 1202A - 1202N and integrated graphics logic 1208. The display unit is for driving one or more externally connected displays.

[0153] Cores 1202A to 1202N may be of the same or different types with respect to the architecture instruction set. That is, two or more of Cores 1202A to 1202N may be capable of executing the same instruction set, while others may be capable of executing only a subset of that instruction set or a different instruction set. [Exemplary Computer Architecture]

[0154] FIGS. 13 to 16 are block diagrams of an exemplary computer architecture. Other system designs and configurations known in the art for laptops, desktops, handheld PCs, personal digital assistants, engineering workstations, servers, network devices, network hubs, switches, embedded processors, digital signal processors (DSPs), graphic devices, video game devices, set-top boxes, microcontrollers, cell phones, portable media players, handheld devices, and other various electronic devices are also suitable. Generally, a wide variety of systems or electronic devices that can incorporate the processors and / or other execution logics disclosed herein are generally suitable.

[0155] Referring now to FIG. 13, a block diagram of a system 1300 in accordance with one embodiment of the present invention is shown. System 1300 may include one or more processors 1310, 1315 coupled to a controller hub 1320. In one embodiment, controller hub 1320 includes a graphics memory controller hub (GMCH) 1390 and an input / output hub (IOH) 1350 (which may be on a separate chip), where GMCH 1390 includes a memory controller and a graphics controller to which memory 1340 and a coprocessor 1345 are coupled, and IOH 1350 couples input / output (I / O) devices 1360 to GMCH 1390. Alternatively, one or both of the memory controller and the graphics controller may be integrated within the processor (as described herein), and memory 1340 and coprocessor 1345 are directly coupled to controller hub 1320 within a single chip having processors 1310 and IOH 1350.

[0156] The optional nature of additional processor 1315 is indicated by dashed lines in FIG. 13. Each processor 1310, 1315 may include one or more of the plurality of processing cores described herein and may be any version of processor 1200.

[0157] Memory 1340 may be, for example, dynamic random access memory (DRAM), phase change memory (PCM), or a combination of the two. For at least one embodiment, controller hub 1320 communicates with processors 1310, 1315 via a multi-drop bus such as a front side bus (FSB), a point-to-point interface such as QuickPath Interconnect (QPI), or a similar connection 1395.

[0158] In one embodiment, the coprocessor 1345 is a dedicated processor such as, for example, a high throughput MIC processor, a network or communication processor, a compression engine, a graphics processor, a GPGPU, an embedded processor, etc. In one embodiment, the controller hub 1320 may include an integrated graphics accelerator.

[0159] There can be various differences between the physical resources 1310, 1315 with respect to criteria of various advantages including those architectural, microarchitectural, thermal, and power consumption characteristics.

[0160] In one embodiment, the processor 1310 executes a plurality of instructions that control general types of data processing operations. Coprocessor instructions may be incorporated within these instructions. The processor 1310 recognizes these coprocessor instructions as being of the type to be executed by the attached coprocessor 1345. Accordingly, the processor 1310 issues these coprocessor instructions (or a plurality of control signals representing the coprocessor instructions) to the coprocessor 1345 over a coprocessor bus or other interconnect. The coprocessor(s) 1345 receives and executes the received coprocessor instructions.

[0161] Referring now to FIG. 14, there is shown a block diagram of a first more specific exemplary system 1400 in accordance with an embodiment of the present invention. As shown in FIG. 14, the multiprocessor system 1400 is a point-to-point interconnect system and includes a first processor 1470 and a second processor 1480 coupled via a point-to-point interconnect 1450. Each of processors 1470 and 1480 may be some version of processor 1200. In some embodiments of the present invention, processors 1470 and 1480 are processors 1310 and 1315 respectively, while coprocessor 1438 is coprocessor 1345. In other embodiments, processors 1470 and 1480 are processor 1310 and coprocessor 1345 respectively.

[0162] Processors 1470 and 1480 are shown as including integrated memory controller (IMC) units 1472 and 1482 respectively. Processor 1470 also includes point-to-point (P-P) interfaces 1476 and 1478 as part of its bus controller unit, and similarly, second processor 1480 includes P-P interfaces 1486 and 1488. Processors 1470, 1480 may exchange information via point-to-point (P-P) interconnect 1450 using P-P interface circuits 1478, 1488. As illustrated in FIG. 14, IMCs 1472 and 1482 couple the processors to respective memories, namely memory 1432 and memory 1434. Memories 1432 and 1434 may be part of main memories locally attached to their respective processors.

[0163] Processors 1470 and 1480 may each exchange information with chipset 1490 via respective P-P interfaces 1452 and 1454, using point-to-point interface circuits 1476, 1494, 1486, and 1498. Chipset 1490 may optionally exchange information with coprocessor 1438 via high-performance interface 1492. In one embodiment, coprocessor 1438 is a dedicated processor such as, for example, a high-throughput MIC processor, a network processor or communication processor, a compression engine, a graphics processor, a GPGPU, or an embedded processor.

[0164] A shared cache (not shown) is included either in a processor or external to both processors already connected to a processor via a P-P interconnect, such that when a processor is put into a low power mode, local cache information of either or both processors may be stored in the shared cache.

[0165] Chipset 1490 may be coupled to first bus 1416 via interface 1496. In one embodiment, first bus 1416 can be a bus such as a Peripheral Component Interconnect (PCI) bus, a PCI Express bus, or another third-generation I / O interconnect bus, although the scope of the present invention is not so limited.

[0166] As shown in FIG. 14, various I / O devices 1414 may be coupled to a first bus 1416, along with a bus bridge 1418 that couples the first bus 1416 to a second bus 1420. In one embodiment, one or more additional processors 1415, such as a coprocessor, a high throughput MIC processor, a GPGPU, an accelerator (e.g., a graphics accelerator or a digital signal processing (DSP) unit, etc.), a field programmable gate array, or any other processor, etc., are coupled to the first bus 1416. In one embodiment, the second bus 1420 may be a Low Pin Count (LPC) bus. In one embodiment, various devices may be coupled to the second bus 1420, including, for example, a keyboard and / or a mouse 1422, a communication device 1427, and a storage unit 1428, such as a disk drive or other mass storage device that may include a plurality of instructions / codes and data 1430. Further, audio I / O 1424 may be coupled to the second bus 1420. Note that other architectures are possible. For example, instead of the point-to-point architecture of FIG. 14, the system may implement a multi-drop bus or other such architecture.

[0167] Referring now to FIG. 15, a block diagram of a second more specific exemplary system 1500 is shown in accordance with an embodiment of the present invention. Identical elements in FIGS. 14 and 15 have multiple identical reference numerals, and the specific aspects of FIG. 14 are omitted from FIG. 15 so as not to obscure the other aspects of FIG. 15.

[0168] FIG. 15 shows that the processors 1470, 1480 may each include integrated memory as well as I / O control logic (``CL'') 1572 and 1582. Thus, CL 1572, 1582 includes an integrated memory controller unit and includes I / O control logic. FIG. 15 shows that not only the memories 1432, 1434 are coupled to CL 1572, 1582, but also a plurality of I / O devices 1514 are coupled to CL 1572, 1582. The legacy I / O device 1515 is coupled to the chipset 1490.

[0169] Referring now to FIG. 16, a block diagram of a SoC 1600 is shown in accordance with an embodiment of the present invention. A plurality of similar elements of FIG. 12 have the same reference numerals. Also, the dashed boxes are optional features on a more advanced SoC. In FIG. 16, the interconnect unit 1602 is coupled to an application processor 1610, a system agent unit 1210, a bus controller unit 1216, an integrated memory controller unit 1214, a set of one or more coprocessors 1620, a static random access memory (SRAM) unit 1630, a direct memory access (DMA) unit 1632, and a display unit 1640 for connecting to one or more external displays. The application processor 1610 includes a set of one or more cores 1202A - 1202N including cache units 1204A - 1204N and a shared cache unit 1206. The set of coprocessors 1620 may include integrated graphics logic, an image processor, an audio processor, and a video processor. In one embodiment, the coprocessor 1620 includes dedicated processors such as, for example, a network or communication processor, a compression engine, a GPGPU, a high throughput MIC processor, an embedded processor, etc.

[0170] Embodiments of the mechanisms disclosed herein may be implemented in hardware, software, firmware, or a combination of such implementation means. Embodiments of the present invention may be implemented as a computer program or program code executed on a programmable system, which programmable system comprises at least one processor, a storage device system (e.g., volatile and non-volatile memories and / or storage elements), at least one input device, and at least one output device.

[0171] Program code such as code 1430 illustrated in FIG. 14 may be applied to input a plurality of instructions for executing a plurality of functions described herein to generate output information. The output information may be applied in a known manner to one or more output devices. For this purpose, the processing system includes any system comprising a processor such as, for example, a digital signal processor (DSP), a microcontroller, an application specific integrated circuit (ASIC), or a microprocessor.

[0172] The program code may be implemented in a high-level procedural or object-oriented programming language for communicating with the processing system. Also, if necessary, the program code may be implemented in assembly language or machine language. In fact, the mechanisms described herein are not limited to any particular programming language within its scope. In any case, the language may be a compiled language or an interpreted language.

[0173] One or more aspects of at least one embodiment may be implemented by typical instructions stored on a machine-readable medium, which instructions represent various logics within a processor and, when read by a machine, cause the machine to create the logic for executing the techniques described herein. Such a plurality of descriptive representations, known as a plurality of "IP cores", may be stored on a tangible machine-readable medium, supplied to various customers or manufacturing facilities, and loaded into a plurality of manufacturing machines that actually create the logic or processor.

[0174] Such a machine-readable storage medium may include a non-transitory tangible configuration of an article manufactured or formed by a machine or device, and examples of such include storage media such as hard disks, floppy disks, optical disks, compact disc read-only memory (CD-ROM), compact disc rewritable (CD-RW), and any other type of disk such as magneto-optical disks, read-only memory (ROM), random access memory (RAM) such as dynamic random access memory (DRAM) and static random access memory (SRAM), semiconductor devices such as erasable programmable read-only memory (EPROM), flash memory, electrically erasable programmable read-only memory (EEPROM), phase change memory (PCM), magnetic or optical cards, or any other type of medium suitable for storing electronic instructions, but are not limited thereto.

[0175] Accordingly, embodiments of the present invention also include non-transitory tangible machine-readable media such as a hardware description language (HDL) that include instructions defining the structural, circuit, device, processor, and / or system features described herein, or include design data. Such embodiments may also be referred to as program products. [Emulation (including binary translation, code morphing, etc.)]

[0176] In some cases, an instruction converter may be used to convert instructions from a source instruction set to a target instruction set. For example, the instruction converter may interpret, morph, emulate, or otherwise convert the instructions into one or more other instructions to be processed by the core (e.g., using static binary translation, dynamic binary translation including dynamic compilation). The instruction converter may be implemented in software, hardware, firmware, or a combination thereof. The instruction converter may be present within the processor, outside the processor, or partially within or partially outside the processor.

[0177] FIG. 17 is a block diagram contrasting the use of a software instruction converter to convert binary instructions within a source instruction set to binary instructions within a target instruction set, according to some embodiments of the present invention. In the illustrated embodiment, the instruction converter is a software instruction converter, although alternatively, the instruction converter may be implemented in software, firmware, hardware, or various combinations thereof. FIG. 17 shows a program in a high-level language 1702 that can be compiled using an x86 compiler 1704 to generate x86 binary code 1706 that can be natively executed by a processor 1716 using at least one x86 instruction set core. The processor 1716 using at least one x86 instruction set core represents any processor that can perform substantially the same functions as an Intel® processor using at least one x86 instruction set core, and it does so as follows. That is, to obtain substantially the same results as an Intel® processor using at least one x86 instruction set core, (1) most of the instruction set of the Intel® x86 instruction set core, or (2) an object code version of an application or other software intended for execution on an Intel® processor using at least one x86 instruction set core, are executed in a compatible state or otherwise processed. The x86 compiler 1704 represents a compiler operable to generate x86 binary code 1706 (e.g., object code) executable on a processor 1716 using at least one x86 instruction set core, with or without further linking.Similarly, FIG. 17 shows a program in a high-level language 1702 that can be compiled using a compiler 1708 for an alternative instruction set to generate an alternative instruction set binary code 1710 that can be natively executed by a processor 1714 that does not use at least one x86 instruction set core (e.g., a processor using multiple cores that execute the MIPS instruction set of MIPS Technology of Sunnyvale, Calif., and / or the ARM instruction set of ARM Holdings of Sunnyvale, Calif.). An instruction converter 1712 is used to convert the x86 binary code 1706 into code that can be natively executed by the processor 1714 that does not use an x86 instruction set core. This converted code may not be the same as the binary code 1710 of the alternative instruction set. This is because it is difficult to create an instruction converter that can do this. However, the converted code implements common operations and is composed of instructions of the alternative instruction set. Thus, the instruction converter 1712 represents software, firmware, hardware, or a combination thereof that enables a processor or other electronic device that does not have an x86 instruction set processor or core to execute the x86 binary code 1706 through emulation, simulation, or any other process. [Further Example]

[0178] Example 1 includes an exemplary system comprising a processor having one or more cores, a last-level cache (LLC), and a cache control circuit (CCC). The LLC has a plurality of ways each assigned to one of a plurality of priorities, with each priority associated with a class of service (CLOS) register specifying a minimum and maximum number of ways occupied. The CCC, when there is an invalid cache line (CL) in the LLC, stores a subsequent cache line having the priority of the requester among the plurality of priorities in this invalid CL, or, if the priority of the requester is the lowest among the plurality of priorities and has an occupancy number of one or more or the occupancy number is the maximum for the priority of the requester, stores the subsequent CL in place of the least recently used (LRU) CL of the priority of the requester, or, if the occupancy number is between the minimum and maximum for the priority of the requester, stores the subsequent CL in place of the LRU CL of the priority of the requester or a lower priority, or, if there is an LRC CL with an occupancy number lower than the minimum and having a lower priority, stores the subsequent CL in place of this LRU CL, or, if there is no eviction candidate having the priority of the requester or a lower priority, stores the subsequent CL in place of the LRU CL of a higher priority than the priority of the requester.

[0179] Example 2 includes the content of the exemplary system of Example 1, the LLC includes a plurality of sets of ways, the plurality of ways are part of the plurality of sets, and the CCC determines in which of the plurality of sets the subsequent CL is included based on a hashing algorithm executed on the logical address of the subsequent CL before determining where to store the subsequent CL.

[0180] Example 3 includes the content of the exemplary system of Example 1 and further includes a cache monitoring circuit that maintains heuristics regarding LLC cache eviction. To create space for filling subsequent CLs with lower priorities, when a percentage of cache lines with higher priorities higher than a threshold are evicted, the CLOS register for the higher priority is updated to increase the minimum and maximum ways occupied.

[0181] Example 4 includes the content of the exemplary system of Example 1, where the plurality of ways each include N CLs, and N is a positive integer greater than or equal to 1.

[0182] Example 5 includes the content of the exemplary system of Example 1, where there are LRU CLs with lower priority, and when storing subsequent CLs instead, if there are other CLs in the way containing the LRU CL, the CCC flushes those other CLs.

[0183] Example 6 includes the content of the exemplary system of Example 1, where one or more cores each implement a virtual machine, and the CCC includes a hypervisor.

[0184] Example 7 includes the content of the exemplary system of Example 1, where the processor is one of a plurality of processors within a data center of a cloud service provider.

[0185] Example 8 includes an exemplary method executed by a cache control circuit (CCC) in a system comprising a processor having one or more cores and a last level cache (LLC) having a plurality of ways each assigned to one of a plurality of priorities, each priority being associated with a class of service (CLOS) register that specifies a minimum and maximum number of ways to occupy. Receiving a request to store a subsequent cache line (CL) having a requesting priority among the plurality of priorities in the LLC; if there is an invalid CL in the LLC, storing the subsequent CL in the invalid CL; or, if the requesting priority is the lowest among the plurality of priorities and has an occupancy number that is one or more or the occupancy number is the maximum for the requesting priority, storing the subsequent CL in place of the least recently used (LRU) CL of the requesting priority; or, if the occupancy number is between the minimum and maximum for the requesting priority, storing the subsequent CL in place of the LRU CL of the requesting priority or a lower priority; or, if the occupancy number is lower than the minimum and there is an LRC CL having a lower priority, storing the subsequent CL in place of the LRU CL; or, if there is no eviction candidate having the requesting priority or a lower priority, storing the subsequent CL in place of the LRU CL of a priority higher than the requesting priority.

[0186] Example 9 includes the content of the exemplary method of Example 8, the LLC includes a plurality of sets of ways, the plurality of ways are part of the plurality of sets, and the CCC determines which of the plurality of sets the subsequent CL is included in based on a hashing algorithm executed on the logical address of the subsequent CL before determining where to store the subsequent CL.

[0187] Example 10 includes the content of the exemplary method of Example 8 and uses an LLC cache monitoring circuit to maintain heuristics regarding LLC cache eviction and create space to fill subsequent CLs with lower priority. When a higher percentage of cache lines with high priority are evicted than the threshold, the CLOS register for high priority is updated to increase the minimum and maximum ways occupied.

[0188] Example 11 includes the content of the exemplary method of Example 8, where multiple ways each contain N CLs, and N is a positive integer greater than or equal to 1.

[0189] Example 12 includes the content of the exemplary method of Example 8. When there is an LRU CL with lower priority and storing a subsequent CL instead, if there are other CLs in the way containing the LRU CL, the CCC flushes the other CLs.

[0190] Example 13 includes the content of the exemplary method of Example 8, where one or more cores each implement a virtual machine and the CCC includes a hypervisor.

[0191] Example 14 includes the content of the exemplary method of Example 8, where the processor is one of a plurality of processors in a data center of a cloud service provider.

[0192] Example 15 includes a non - transitory computer - readable medium containing instructions to which a cache control circuit (CCC) in a system responds, the system comprising a processor having one or more cores and a last - level cache (LLC) having a plurality of ways each assigned to one of a plurality of priorities, where each priority is associated with a class - of - service (CLOS) register that specifies a minimum and maximum number of ways to be occupied. The response is to store a subsequent cache line (CL) having a certain requester priority among the plurality of priorities in the invalid CL if there is an invalid CL in the LLC; or, if the requester priority is the lowest among the plurality of priorities and has an occupancy number of one or more or the occupancy number is the maximum for the requester priority, store the subsequent CL in place of the least - recently - used (LRU) CL of the requester priority; or, if the occupancy number is between the minimum and maximum for the requester priority, store the subsequent CL in place of the LRU CL of the requester priority or a lower priority; or, if the occupancy number is lower than the minimum and there is an LRC CL with a lower priority, store the subsequent CL in place of this LRU CL; or, if there is no eviction candidate having the requester priority or a lower priority, store the subsequent CL in place of the LRU CL of a priority higher than the requester priority.

[0193] Example 16 includes the content of the non - transitory computer - readable medium of the exemplary method of Example 15, where the LLC includes a plurality of sets of ways, the plurality of ways being part of the plurality of sets, and the CCC, in further response to the instructions, determines which of the plurality of sets the subsequent CL is included in based on a hashing algorithm executed on the logical address of the subsequent CL before determining where to store the subsequent CL.

[0194] Example 17 includes the content of the non - transitory computer - readable medium of the exemplary method of Example 15, where the processor is one of a plurality of processors in a data center of a cloud service provider.

[0195] Example 18 includes the content of the exemplary non-transitory computer-readable medium of the exemplary method of Example 15, where the plurality of ways each include N CLs, and N is a positive integer greater than or equal to 1.

[0196] Example 19 includes the content of the exemplary non-transitory computer-readable medium of the exemplary method of Example 15, where there is an LRU CL with a lower priority, and when storing a subsequent CL in its place, the CCC flushes other CLs if there are other CLs in the way that includes the LRU CL.

[0197] Example 20 includes the content of the exemplary non-transitory computer-readable medium of the exemplary method of Example 15, where one or more cores each implement a virtual machine, and the CCC includes a hypervisor.

[0198] [Other conceivable items] (Item 1) A last-level cache (LLC) having a plurality of ways each assigned to one of a plurality of priorities, wherein each priority is associated with a class-of-service (CLOS) register specifying the minimum and maximum number of ways occupied, and a cache control circuit (CCC). When there is an invalid cache line (CL) in the LLC, the CCC stores a subsequent cache line (CL) having a requester priority, which is one of the plurality of priorities, in the invalid CL. If the requester priority is the lowest of the plurality of priorities, has an occupancy number of one or more, or the occupancy number is the maximum for the requester priority, the subsequent CL is stored instead of the least recently used (LRU) CL of the requester priority. If the occupancy number is between the minimum and the maximum for the requester priority, the subsequent CL is stored instead of the LRU CL of the requester priority or a lower priority. If the occupancy number is lower than the minimum and there is a CL having a lower priority, the subsequent CL is stored instead of the LRU CL having the lower priority. If there is no invalid CL or a CL having the requester priority or a lower priority, the subsequent CL is stored instead of the LRU CL of a higher priority. A system comprising the CCC. (Item 2) The LLC includes a plurality of sets of ways, and the plurality of ways are part of the plurality of sets. The CCC determines which of the plurality of sets the subsequent CL is included in based on a hashing algorithm executed on the logical address of the subsequent CL before determining where to store the subsequent CL. The system according to Item 1. (Item 3) Further comprising a cache monitoring circuit that maintains heuristics regarding LLC cache eviction. To create space for filling subsequent CLs with lower priorities, when a percentage of cache lines with higher priorities that is higher than a threshold is evicted, the CLOS register for the higher priority is updated to increase the minimum and maximum ways occupied. The system according to Item 1. (Item 4) The system according to Item 1, wherein each of the plurality of ways includes N CLs, and N is a positive integer greater than or equal to 1. (Item 5) The system according to Item 1, wherein when there is the LRU CL having a lower priority and the subsequent CL is stored in place thereof, if there are other CLs in the way including the LRU CL, the CCC flushes the other CLs. (Item 6) The system according to Item 1, further comprising a processor having one or more cores that incorporate the LLC and the CCC and implement virtual machines respectively, and the CCC includes a hypervisor. (Item 7) The system according to Item 6, wherein the processor is one of a plurality of processors in a data center of a cloud service provider. (Item 8) A method executed by a cache control circuit (CCC) in a system comprising a last-level cache (LLC) having a plurality of ways each assigned to one of a plurality of priorities, wherein each priority is associated with a class of service (CLOS) register specifying a minimum and maximum number of ways occupied, the method comprising: receiving a request to store a subsequent cache line (CL) having a requesting priority among the plurality of priorities in the LLC; storing the subsequent CL in the invalid CL if there is an invalid CL in the LLC; storing the subsequent CL instead of the least recently used (LRU) CL of the requesting priority if the requesting priority is the lowest among the plurality of priorities, has an occupancy number of one or more, or the occupancy number is the maximum for the requesting priority; storing the subsequent CL instead of the LRU CL of the requesting priority or a lower priority if the occupancy number is between the minimum and the maximum for the requesting priority; storing the subsequent CL instead of the LRU CL of the lower priority if the occupancy number is lower than the minimum and there is a CL having a lower priority; and storing the subsequent CL instead of the LRU CL of a higher priority if there is no invalid CL or a CL having the requesting priority or a lower priority. (Item 9) The method according to item 8, wherein the LLC includes a plurality of sets of ways, the plurality of ways are part of the plurality of sets, and the CCC determines which of the plurality of sets the subsequent CL is included in based on a hashing algorithm executed on the logical address of the subsequent CL before determining where to store the subsequent CL. (Item 10) The method according to item 8, wherein, using the LLC cache monitoring circuit, in order to maintain heuristics regarding LLC cache eviction and create space to fill subsequent CLs with lower priority, when a cache line with a higher priority and a ratio higher than the threshold is evicted, the CLOS register for the higher priority is updated to increase the minimum and maximum ways of occupancy. (Item 11) The method according to item 8, wherein each of the plurality of ways includes N CLs, and N is a positive integer greater than or equal to 1. (Item 12) The method according to item 8, wherein when storing the subsequent CL instead of the LRU CL with the lower priority, if there are other CLs in the way including the LRU CL, the CCC flushes the other CLs. (Item 13) The method according to item 8, wherein the system further includes a processor having one or more cores that implement virtual machines and incorporating the LLC and the CCC, and the CCC includes a hypervisor. (Item 14) The method according to item 13, wherein the processor is one of a plurality of processors in a data center of a cloud service provider. (Item 15) A non - transitory computer - readable medium including instructions for a cache control circuit (CCC) in a system comprising a last - level cache (LLC) having a plurality of ways each assigned to one of a plurality of priorities, where each priority is associated with a class - of - service (CLOS) register specifying a minimum and maximum number of ways occupied, and the response includes receiving a request to store a subsequent cache line (CL) having a certain requester priority among the plurality of priorities in the LLC; storing the subsequent CL in an invalid CL if an invalid CL exists within the LLC; storing the subsequent CL in place of the least recently used (LRU) CL of the requester priority if the requester priority is the lowest among the plurality of priorities and has an occupancy count of one or more, or if the occupancy count is the maximum for the requester priority; storing the subsequent CL in place of the LRU CL of the requester priority or a lower priority if the occupancy count is between the minimum and the maximum for the requester priority; storing the subsequent CL in place of the LRU CL of a lower priority if the occupancy count is lower than the minimum and a CL with a lower priority exists; and storing the subsequent CL in place of the LRU CL of a higher priority if no invalid CL or CL with the requester priority or a lower priority exists. (Item 16) The LLC includes a plurality of sets of ways, and the plurality of ways are part of the plurality of sets. The CCC further responds to the instructions by determining, based on a hashing algorithm executed on the logical address of the subsequent CL, which of the plurality of sets the subsequent CL is included in before determining where to store the subsequent CL. The non - transitory computer - readable medium according to Item 15. (Item 17) The non-transitory computer-readable medium according to item 15, wherein the system further includes a processor incorporating the LLC and the CCC, and the processor is one of a plurality of processors in a data center of a cloud service provider. (Item 18) The non-transitory computer-readable medium according to item 15, wherein the plurality of ways each include N CLs, and N is a positive integer greater than or equal to 1. (Item 19) The non-transitory computer-readable medium according to item 15, wherein when there is the LRU CL having the lower priority and the subsequent CL is stored instead, the CCC flushes the other CLs if there are other CLs in the way including the LRU CL. (Item 20) The non-transitory computer-readable medium according to item 15, wherein the system further includes a processor incorporating the LLC and the CCC and having one or more cores implementing virtual machines respectively, and the CCC includes a hypervisor.

Claims

1. A last-level cache (LLC) having a plurality of ways, each assigned to one of a plurality of priorities, the LLC associated with a class of service register (CLOS register) that specifies the minimum and maximum number of ways occupied by each priority, and a cache control circuit (CCC), when there is an invalid cache line (invalid CL) in the LLC, storing a subsequent cache line (subsequent CL) having a requester priority that is one of the plurality of priorities in the invalid CL, when the requester priority is the lowest of the plurality of priorities and the number of ways occupied by the requester priority is 1 or more, or when the number of ways occupied for the requester priority is the maximum, storing the subsequent CL in place of the least recently used (LRU) CL of the requester priority, when the number of ways occupied for the requester priority is between the minimum and the maximum, storing the subsequent CL in place of the LRU CL of the requester priority or a lower priority, when the number of ways occupied for the requester priority is lower than the minimum and there is a CL having a lower priority, storing the subsequent CL in place of the LRU CL having the lower priority, the CCC that stores the subsequent CL in place of the LRU CL of a higher priority when there is no invalid CL or CL having the requester priority or a lower priority, A system comprising.

2. The LLC includes a plurality of sets of ways, the plurality of ways being part of the plurality of sets, and the CCC determines which of the plurality of sets the subsequent CL is included in based on a hashing algorithm executed on the logical address of the subsequent CL before determining where to store the subsequent CL. The system according to claim 1.

3. Further comprising a cache monitoring circuit that maintains heuristics regarding LLC cache eviction, when a higher percentage of cache lines having a high priority are evicted to create space for filling subsequent CLs having a lower priority, the CLOS register for the high priority is updated to increase the minimum and maximum ways occupied. The system according to claim 1.

4. The system according to claim 1, wherein each of the plurality of ways includes N CLs, and N is a positive integer greater than or equal to 1.

5. When storing the subsequent CL instead of the LRU CL having the lower priority, if there are other CLs in the way including the LRU CL, the CCC flushes the other CLs. The system according to claim 1.

6. The system according to any one of claims 1 to 5, further comprising a processor having one or more cores that incorporate the LLC and the CCC and implement virtual machines respectively, wherein the CCC includes a hypervisor.

7. The system according to claim 6, wherein the processor is one of a plurality of processors in a data center of a cloud service provider.

8. A method executed by a cache control circuit (CCC) in a system including a last level cache (LLC) having a plurality of ways respectively assigned to one of a plurality of priorities, wherein each priority is associated with a class of service register (CLOS register) that specifies the minimum and maximum number of ways occupied, the method comprising: Receiving a request to store a subsequent cache line (subsequent CL) having a certain requester priority among the plurality of priorities in the LLC; Storing the subsequent CL in the invalid CL if there is an invalid CL in the LLC; When the requester priority is the lowest among the plurality of priorities and the number of ways occupied by the requester priority is 1 or more, or when the number of ways occupied for the requester priority is the maximum, storing the subsequent CL instead of the least recently used (LRU) CL of the requester priority; When the number of ways occupied for the requester priority is between the minimum and the maximum, storing the subsequent CL instead of the LRU CL of the requester priority or a lower priority; When the number of ways occupied for the requester priority is lower than the minimum and there is a CL having the lower priority, storing the subsequent CL instead of the LRU CL having the lower priority; If there is no invalid CL or a CL having the priority of the requester or lower than that, storing the subsequent CL instead of the LRU CL with a higher priority; A method including this.

9. The LLC includes a plurality of sets of ways, and the plurality of ways are part of the plurality of sets. Before determining where to store the subsequent CL, the CCC determines which of the plurality of sets the subsequent CL is included in based on a hashing algorithm executed on the logical address of the subsequent CL. The method according to claim 8.

10. Using an LLC cache monitoring circuit to maintain heuristics regarding LLC cache eviction, When a cache line with a high priority and a ratio higher than a threshold is evicted to create space for filling a subsequent CL with a lower priority, updating the CLOS register for the high priority to increase the minimum and maximum ways occupied. The method according to claim 8.

11. Each of the plurality of ways includes N CLs, and N is a positive integer greater than or equal to 1. The method according to claim 8.

12. When the LRU CL with the lower priority exists and the subsequent CL is stored instead, if there are other CLs in the way including the LRU CL, the CCC flushes the other CLs. The method according to claim 8.

13. The system further includes a processor having one or more cores implementing virtual machines, each incorporating the LLC and the CCC, and the CCC includes a hypervisor. The method according to any one of claims 8 to 12.

14. The processor is one of a plurality of processors in a data center of a cloud service provider. The method according to claim 13.

15. A computer program including instructions executed by a cache control circuit (CCC) in a system including a last level cache (LLC) having a plurality of ways each assigned to one of a plurality of priorities, wherein each priority is associated with a class of service (CLOS) register specifying a minimum and a maximum number of ways occupied. The execution includes Receiving a request to store a subsequent cache line (subsequent CL) having a certain requester priority among the plurality of priorities in the LLC; When there is an invalid CL in the LLC, storing the subsequent CL in the invalid CL; When the requester priority is the lowest among the plurality of priorities, the number of ways occupied by the requester priority is 1 or more, or the number of ways occupied for the requester priority is the maximum, storing the subsequent CL instead of the least recently used (LRU) CL of the requester priority; When the number of ways occupied for the requester priority is between the minimum and the maximum, storing the subsequent CL instead of the LRU CL of the requester priority or a lower priority; When the number of ways occupied for the requester priority is lower than the minimum and there is a CL having a lower priority, storing the subsequent CL instead of the LRU CL having the lower priority; When there is no invalid CL or a CL having the requester priority or a lower priority, storing the subsequent CL instead of the LRU CL having a higher priority; A computer program implemented by the above.

16. The LLC includes a plurality of sets of ways, the plurality of ways are part of the plurality of sets, and the CCC further responds to the instruction to determine which of the plurality of sets the subsequent CL is included in based on a hashing algorithm executed on the logical address of the subsequent CL before determining where to store the subsequent CL. The computer program according to claim 15.

17. The system further includes a processor incorporating the LLC and the CCC, and the processor is one of a plurality of processors in a data center of a cloud service provider. The computer program according to claim 15.

18. Each of the plurality of ways includes N CLs, and N is a positive integer of 1 or more. The computer program according to claim 15.

19. When there is the LRU CL having the lower priority than the foregoing, and storing the subsequent CL instead thereof, when there are other CLs in the way including the LRU CL, the CCC causes the other CLs to be flushed. The computer program according to claim 15.

20. The system further includes a processor having one or more cores that incorporate the LLC and the CCC and implement virtual machines, respectively, and the CCC includes a hypervisor. The computer program according to any one of claims 15 to 19.

21. A non-transitory computer-readable medium storing the computer program according to any one of claims 15 to 20.

Citation Information

Patent Citations

  • Multiprocess processor

    JP1997101916A

  • Cache memory having sector function

    JP2009163450A

  • Cache memory

    JP2011018196A

  • Providing quality of service (QoS) for cache architectures using priority information

    US20080040554A1

  • Dynamic quality of service (QoS) for a shared cache

    US20080235457A1