Core-Aware Cache System and Method for Multi-Core Processors
By adopting core-aware non-inclusive non-exclusive cache technology in multi-core processors and using core sharing agents for cache management, the problems of waste of cache capacity and high complexity in the existing technology are solved, and an efficient cache system is realized.
Patent Information
- Application Number
- CN202180004850.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-01-20
- Publication Date
- 2025-06-20
- Estimated Expiration
- 2041-01-20
AI Technical Summary
While improving computing performance, the cache strategy of existing multi-core processors faces the problems of wasted cache capacity and high complexity.
Using kernel-aware non-inclusive non-exclusive (NINE) caching technology, a non-inclusive non-exclusive cache strategy is implemented by introducing a core sharing agent in a multi-core processor, and data and instructions are cached based on physical page numbers and core identifiers or core valid bit vectors.
It effectively reduces cache misses, improves the efficiency of the cache system, and simplifies cache consistency management, providing a large effective cache capacity.
Smart Images

Figure CN115119520B_ABST
Abstract
Description
Background Art
[0001] Multi-core processors and memory caches are some common technologies in computing devices. The multiple computing cores of a multi-core are configured to run multiple applications, multiple routines within an application, multiple instances of a given routine, etc. to enhance computing performance. Memory caches are used to temporarily store data and / or instructions commonly used by the cores of a computing device to further enhance computing performance. Caches can be organized into multiple levels, can be configured to cache data, instructions, or both, and can be specific to (private, allocated, exclusive, etc.) individual computing cores or shared among multiple computing cores. Caches can be internal to a multi-core processor, external to a multi-core processor, or some cache layers can be integrated while other cache layers are external to the multi-core processor.
[0002] Please refer to Figure 1 , which shows an exemplary processor according to the prior art. Processor 100 may include, but is not limited to, multiple cores 105 - 115, multi-level caches 120 - 150, and one or more memory controllers 155 and interconnect interface HT160. The multi-level caches 120 - 150 may include one or more levels of caches 120 - 145 dedicated to each of the multiple cores 105 - 115, and one or more levels of caches 150 shared among the multiple cores 105 - 115. For example, processor 100 may include multiple level-1 (L1) caches 120 - 130 and multiple level-2 (L2) caches 134 - 145. Each level-1 (L1) cache 120 - 130 and each level-2 (L2) cache 135 - 145 may be configured to cache data and / or instructions for each of the multiple cores 105 - 115. The multi-level caches 120 - 150 may also include one or more levels of caches 150 shared by the multiple cores 105 - 115. For example, processor 100 may include one or more level-3 (L3) caches 150, which are configured to cache data and / or instructions for the multiple cores 105 - 115.
[0003] One or more interconnect interfaces may include one or more memory controllers 155, which may be configured to handle memory access requests. One or more memory controllers 155 may be coupled between one or more external memories 165 - 170 and one or more levels of caches 120 - 150. For example, the processor 100 may include a memory controller 155 coupled between one or more dynamic random access memories (DRAMs) 165 - 170 and multiple levels of caches 120 - 150. The memory controller 155 may be configured to read data from the DRAMs 165 - 170 into one or more of the multiple levels of caches 120 - 150, and write data into one or more of the multiple levels of caches 120 - 150. One or more interconnect interfaces 155 - 160 may also include an interconnect interface 160 to interconnect the processor 100 to one or more input / output devices 175, other processors, etc. For example, one or more interconnect interfaces 160 may include, but are not limited to, bidirectional serial and / or parallel communication interfaces, such as, but not limited to, a HyperTransport (HT) interface coupled between one or more input / output devices 175, one or more memory controllers 155, and one or more shared level-three (L3) caches 150.
[0004] A given cache layer may be inclusive, exclusive, or non-inclusive non-exclusive (NINE) of the next higher cache layer. As used herein, the terms lower and higher cache levels are used to refer to the relationship between cache layers. In an inclusive cache policy, data blocks and / or instructions in a higher-level cache also exist in a lower-level cache. In other words, the lower-level cache is in communication with the higher-level cache. In an exclusive cache policy, data blocks and / or instruction blocks in a lower-level cache do not exist in a higher-level cache. In other words, the lower-level cache is mutually exclusive with the higher-level cache. If the content of a lower-level cache is neither strictly inclusive nor mutually exclusive with a higher-level cache, the lower-level cache is considered non-inclusive non-exclusive. Now referring to Figure 2 , an inclusive cache method according to the prior art is shown. Reference will be made to Figure 1Describe the inclusive cache method using the secondary (L2) cache and the shared tertiary (L3) cache. The method includes: receiving a current memory access request from a given one of multiple cores at 205. At 210, it can be determined whether the data and / or instructions of the given physical page number (PPN) of the memory access request are cached in a given higher-level cache. For example, it can be determined whether the data and / or instructions are cached in a given secondary (L2) cache 140. At 215, if the data and / or instructions of the given physical page number are found in the given higher-level cache (e.g., cache hit), the data and / or instructions can be retrieved from the given higher-level cache according to the corresponding cache policy and placed in a given further higher-level cache or returned to a given one of the multiple cores. For example, the data and / or instructions can be retrieved from the given secondary (L2) cache 140 and placed in a given primary (L1) cache 125 and / or returned to the given core 110. At 220, if the data and / or instructions of the given physical page number are not found in the given higher-level cache (e.g., cache miss), it can be determined whether the data and / or instructions of the given physical page number of the memory access request are cached in a given lower-level cache. For example, if there is a miss in the given secondary (L2) cache 140, it can be determined whether the data and / or instructions of the given physical page number for the memory access request received from the given core 110 are cached in the shared tertiary (L3) cache 150. At 225, if the data and / or instructions of the given physical page number are found in the given lower-level cache, the data and / or instructions can be retrieved from the given lower-level cache and placed in the given higher-level cache. For example, the data and / or instructions can be retrieved from the shared tertiary (L3) cache 150 and placed in the given secondary (L2) cache 140. At 230, if the data and / or instructions of the given physical page number are not found in the given lower-level cache, the data and / or instructions of the given physical page number of the memory access request can be obtained from a further lower-level cache or from memory and placed in the given lower-level cache and the given higher-level cache. For example, if the data and / or instructions are not found in the shared tertiary (L3) cache 150, the data and / or instructions can be obtained from a further lower-level cache (if available) or from the memory 165 - 170. The obtained data and / or instructions can be placed in the shared tertiary (L3) cache 150 and the given secondary (L2) cache 140. At 235, if other data and / or instructions are evicted from the given lower-level cache, other data and / or instructions in the given higher-level cache are also invalidated / evicted.For example, if other data and / or instructions are evicted from the shared level-3 (L3) cache 150 to make room for fetched data and / or instructions for a given physical page number of a memory access request, the corresponding other data and / or instructions cached in a given level-2 (L2) cache 140 are also invalidated or evicted. The inclusive cache approach conveniently filters out unnecessary coherence snooping traffic. However, the inclusive cache approach also wastes effective cache capacity.
[0005] Now refer to Figure 3 , which illustrates an exclusive cache approach according to conventional techniques. Reference will be made to Figure 1The exclusive cache method is described in terms of a secondary (L2) cache and a shared tertiary (L3) cache. At 305, the method includes receiving a current memory access request from a given one of a plurality of cores. At 310, it can be determined whether the data and / or instructions for a given physical page number for the memory access request are cached in a given higher-level cache. For example, it can be determined whether the data and / or instructions are cached in a given secondary (L2) cache 140. At 315, if the data and / or instructions for a given physical page number are found in the given higher-level cache (e.g., cache hit), the data and / or instructions can be retrieved from the given higher-level cache and placed in a given further higher-level cache or returned to a given one of the plurality of cores according to the corresponding cache policy. For example, the data and / or instructions can be retrieved from the given secondary (L2) cache 140 and placed in a given primary (L1) cache 125 and / or returned to the given core 110. At 320, if the data and / or instructions for a given physical page number are not found in the given higher-level cache (e.g., cache miss), it can be determined whether the data and / or instructions for the given physical page number for the memory access request are cached in a given lower-level cache. For example, if there is a cache miss in the given secondary (L2) cache 140, it can be determined whether the data and / or instructions for the given physical page number received from the given core 110 are cached in the shared tertiary (L3) cache 150. At 325, if the data and / or instructions for a given physical page number are found in the given lower-level cache, the data and / or instructions can be moved from the given lower-level cache to the given higher-level cache. For example, the data and / or instructions can be removed from the shared tertiary (L3) cache 150 and placed in the given secondary (L2) cache 140. At 330, if other data and / or instructions are evicted from the given higher-level cache, the other data and / or instructions can be placed in the given lower-level cache. For example, if other data and / or instructions are evicted from the given secondary (L2) cache 140 to make room for the moved data and / or instructions for the given physical page number of the memory access request, the corresponding other data and / or instructions can be moved to the shared tertiary (L3) cache 150. At 335, if the data and / or instructions for a given physical page number are not found in the given lower-level cache, the data and / or instructions for the given physical page number for the memory access request can be retrieved from an even lower-level cache or from memory and placed in the given higher-level cache. For example, if the data and / or instructions are not found in the shared tertiary (L3) cache 150, the data and / or instructions can be fetched from the next lower-level cache (if applicable) or from memory 165-170.The retrieved data and / or instructions can be placed in a given level-two (L2) cache 140. At 340, likewise, if other data and / or instructions are evicted from a given higher-level cache, the other data and / or instructions can be placed in a given lower-level cache. For example, if other data and / or instructions are evicted from a given level-two (L2) cache 140 to make room for the data and / or instructions of a given physical page number for a retrieved memory access request, the corresponding other data and / or instructions can be moved to a shared level-three (L3) cache 150. The exclusive cache method advantageously provides a large effective cache capacity. However, to maintain exclusivity and cache coherence, the exclusive cache method is characterized by a relatively high complexity.
[0006] Reference is now made to Figure 4 , which illustrates a non-inclusive non-exclusive cache method according to the prior art. Reference will be made to Figure 1The non-inclusive non-exclusive cache method is described in terms of a secondary (L2) cache and a shared tertiary (L3) cache. The method may include: at 405, receiving a current memory access request from a given one of a plurality of cores. At 410, it may be determined whether the data and / or instructions of a given physical page number of the memory access request are cached in a given higher-level cache. For example, it may be determined whether the data and / or instructions are cached in a given secondary (L2) cache 140. At 415, if the data and / or instructions are in a given higher-level cache (e.g., cache hit), the data and / or instructions may be retrieved from the given higher-level cache and placed in a given further higher-level cache or returned to the given one of the plurality of cores according to the corresponding cache policy. For example, the data and / or instructions may be retrieved from a given secondary (L2) cache 140 and placed in a given primary (L1) cache 125 and / or returned to a given core 110. At 420, if the data and / or instructions of a given physical page number are not found in a given higher-level cache (e.g., cache miss), it may be determined whether the data and / or instructions of the given physical page number of the memory access request are cached in a given lower-level cache. For example, if there is a cache miss in a given secondary (L2) cache 140, it may be determined whether the data and / or instructions are cached in a shared tertiary (L3) cache 150. At 425, if the data and / or instructions of the given physical page number are in a given lower-level cache, the data and / or instructions may be retrieved from the given lower-level cache and placed in a given higher-level cache. For example, the data and / or instructions may be fetched from a shared tertiary (L3) cache 150 and placed in a given secondary (L2) cache 140. At 430, if the data and / or instructions of a given physical page number are not found in a given lower-level cache, the data and / or instructions for the given physical page number of the memory access request may be retrieved from a further lower-level cache or from memory and placed in the given lower-level cache and the given higher-level cache. For example, if the data and / or instructions are not found in a shared tertiary (L3) cache 150, the data and / or instructions may be fetched from the next lower-level cache (if applicable) or from memory 165-170. The fetched data and / or instructions may be placed in a shared tertiary (L3) cache 150 and a given secondary (L2) cache 140. In the non-inclusive non-exclusive cache method, there is no return invalid and / or purge. Compared to the exclusive cache policy, the non-inclusive non-exclusive cache method is closer to the inclusive cache policy because it retains the fetched data and / or instructions in the lower-level cache.Non-inclusive non-exclusive cache methods can be implemented relatively simply, but provide limited improvement in terms of effective cache capacity. Another characteristic of non-inclusive non-exclusive cache methods is complex cache coherence.
[0007] Although inclusive, exclusive, and non-inclusive non-exclusive cache methods provide various trade-offs, there is still a need for improved cache systems and methods. Summary of the Invention
[0008] The present technology can be best understood by reference to the following description and drawings, which are used to illustrate embodiments of the present technology that relate to core-aware non-inclusive non-exclusive (NINE) cache technology.
[0009] In one embodiment, a non-inclusive non-exclusive cache method can include receiving a memory access request from one or more of a plurality of cores. Caching data and / or instructions relative to a shared cache level and a core-specific cache level based on a physical page number (PPN) and a group of core identifiers for previously accessing each physical page number.
[0010] In another embodiment, a non-inclusive non-exclusive cache method can include receiving a memory access request from one or more of a plurality of cores. Performing core-aware non-inclusive non-exclusive caching of data and / or instructions relative to a shared cache level and a core-specific cache level based on a physical page number and a group of core valid bit vectors for each physical page number previously accessed by each of the plurality of cores.
[0011] In another embodiment, a computing system can include a plurality of computing cores, one or more cache levels specific to each of the plurality of computing cores, one or more cache levels shared by the plurality of computing cores, and a core sharing agent. The core sharing agent can be configured to perform non-inclusive non-exclusive caching of data and / or instructions in the shared cache level relative to the core-specific cache level based on the core sharing behavior of the shared cache layer.
[0012] This summary of the invention introduces a selection of concepts that are further described below in the detailed description. This summary of the invention is not intended to identify key features or essential features of the claimed subject matter, nor is it intended to be used to limit the scope of the claimed subject matter. Brief Description of the Drawings
[0013] Embodiments of the present application are illustrated in the drawings by way of example and not limitation, and wherein like reference numerals refer to like elements, and wherein:
[0014] Figure 1 An exemplary processor according to conventional technology is shown.
[0015] Figure 2 Shows an inclusive cache method according to the prior art.
[0016] Figure 3 Shows an exclusive cache method according to the prior art.
[0017] Figure 4 Shows a non-inclusive non-exclusive (NINE) cache method according to the prior art.
[0018] Figure 5 Shows an exemplary processor according to some aspects of the present application.
[0019] Figures 6A-6B Shows a core-aware non-inclusive non-exclusive cache method according to some aspects of the present application.
[0020] Figure 7 Shows a core-aware cache data array according to some aspects of the present application.
[0021] Figures 8A-8B , a core-aware non-inclusive non-exclusive cache method according to some aspects of the present application.
[0022] Figure 9 Shows a core-aware cache data array according to some aspects of the present application. Detailed Description
[0023] Reference will now be made in detail to embodiments of the present application, examples of which are illustrated in the accompanying drawings. While the present application will be described in conjunction with these embodiments, it should be understood that they are not intended to limit the present application to these embodiments. On the contrary, the present invention is intended to cover alternatives, modifications, and equivalents that may be included within the scope of the present invention as defined by the appended claims. In addition, in the following detailed description of the present application, numerous specific details are set forth in order to provide a thorough understanding of the present application. However, it should be understood that the present application may be practiced without these specific details. In other instances, well-known methods, procedures, components, and circuits have not been described in detail so as not to unnecessarily obscure aspects of the present application.
[0024] Some embodiments of the present application below are presented in the form of routines, modules, logic blocks, and other symbolic representations of operations on data within one or more electronic devices. The description and representation are means by which those skilled in the art can most effectively convey the substance of their work to others skilled in the art. A routine, module, logic block, and / or the like are herein and generally considered to be a self-consistent sequence of processes or instructions that result in a desired outcome. These processes include physical operations on physical quantities. Typically, though not necessarily, these physical operations take the form of electrical or magnetic signals capable of being stored, transmitted, compared, and otherwise manipulated in an electronic device. For convenience, and with reference to common usage, in reference to embodiments of the present technology, these signals are referred to as data, bits, values, elements, symbols, characters, terms, numbers, strings, and the like.
[0025] However, it should be remembered that these terms are to be interpreted as referring to physical operations and quantities and are merely convenient labels and will be further interpreted in light of terms commonly used in the art. Unless clearly stated otherwise from the following discussion, it should be understood that, by the discussion of the present application, the use of terms such as "receiving" refers to actions and processes of an electronic computing device such as processing and transforming data. Data is represented as physical (e.g., electrical) quantities within the logic circuits, registers, memories, and / or the like of an electronic device and is transformed into other data similarly represented as physical quantities within the electronic device.
[0026] In the present application, antonymic conjunctions are intended to include conjunctions. The use of the definite or indefinite article is not intended to denote number. In particular, a reference to "the" object or "a" object is also intended to denote one of a possible plurality of such objects. The use of terms such as "comprising," "comprises," "including," "includes," etc. specifies the presence of the stated element but does not preclude the presence or inclusion of one or more other elements and / or groups thereof. It should also be understood that although the terms first, second, etc. may be used herein to describe various elements, such elements should not be limited by these terms. These terms are used herein to distinguish one element from another. For example, without departing from the scope of the embodiments, a first element may be referred to as a second element and, similarly, a second element may be referred to as a first element. It should also be understood that when an element is referred to as "coupled" to another element, it may be directly or indirectly connected to the other element, or there may be intermediate elements. In contrast, when an element is referred to as "directly connected" to another element, there are no intermediate elements. It should also be understood that the term "and / or" includes any and all combinations of one or more related elements. It should also be understood that the wording and terminology used herein are for the purpose of description and should not be regarded as limiting.
[0027] Now refer to Figure 5, showing an exemplary processor in accordance with some aspects of the present application. Processor 500 may include, but is not limited to, multiple cores 505 - 515, multi - level caches 520 - 550, and one or more interconnect interfaces 555 - 560. The multi - level caches 520 - 550 may include one or more levels of caches 520 - 545 dedicated to each of the multiple cores 505 - 515, and one or more levels of caches 550 shared among the multiple cores 505 - 515. For example, processor 500 may include multiple level - one (L1) caches 520 - 530 and multiple level - two (L2) caches 535 - 545. Each level - one (L1) cache 520 - 530 and each level - two (L2) cache 535 - 545 may be configured to cache data and / or instructions for the respective cores among the multiple cores 505 - 515. The multi - level caches 520 - 550 may also include one or more levels of caches 550 shared by the multiple cores 505 - 515. For example, processor 500 may include one or more level - three (L3) caches 550, which are configured to cache data and / or instructions for the multiple cores 505 - 515.
[0028] One or more interconnects may include one or more memory controllers 555, which are configured to handle memory access requests. One or more memory controllers 555 may be coupled between one or more external memories 565 - 570 and the one or more levels of caches 520 - 550. For example, processor 500 may include a memory controller 555 coupled between one or more dynamic random access memories (DRAMs) 565 - 570 and the multi - level caches 520 - 550. The memory controller 555 may be configured to read data from the DRAMs 565 - 570 into one or more of the multiple levels of caches 520 - 550, and write data from one or more of the multiple levels of caches 520 - 550 into the DRAMs 565 - 570.
[0029] One or more interconnect interfaces 555 - 560 may also include an interconnect interface 560 to interconnect processor 500 to one or more input / output devices 575, other processors, etc. For example, one or more interconnect interfaces 560 may include, but are not limited to, bi - directional serial and / or parallel communication interfaces, such as, but not limited to, a HyperTransport (HT) interface coupled between one or more input / output devices 575, one or more memory controllers 555, and one or more shared level - three (L3) caches 550.
[0030] The processor 500 may also include a Core Sharing Agent (CSA) 580. In one implementation, the Core Sharing Agent 580 may be integrated into a cache at a given level or may be a discrete subsystem of the processor 500. The Core Sharing Agent 580 may be configured to implement a core-aware non-inclusive non-exclusive (NINE) cache policy. The core-aware non-inclusive non-exclusive cache policy and operation of the Core Sharing Agent 580 will be further explained with reference to Figures 6A-6B , FIGS. 7, 8A - 8B, and 9.
[0031] Now referring to Figures 6A-6B , which illustrates a core-aware non-inclusive non-exclusive (NINE) caching method in accordance with some aspects of the application. The method may include, at 605, receiving a current memory access request from a given one of a plurality of cores. At 610, it may be determined whether data and / or instructions for a given physical page number (PPN) of the current memory access request are cached in a given higher-level cache specific to each given (private, allocated, exclusive, etc.) core. For example, it may be determined whether the data and / or instructions are cached in a given level two (L2) cache 540 of a given core 510. At 615, if the data and / or instructions for the given physical page number are in the given higher-level cache (e.g., cache hit), then the data and / or instructions may be retrieved from the given higher-level cache and placed in a given further higher-level cache or returned to the given one of the plurality of cores according to the corresponding cache policy. For example, the data and / or instructions may be retrieved from the given level two (L2) cache 540 and placed in a given level one (L1) cache 525 and / or returned to the given core 510.
[0032] At 620, if the data and / or instructions for a given physical page number are not found in a given higher-level cache (e.g., cache miss), it can be determined whether the data and / or instructions for the given physical page number are cached in a given lower-level shared cache. For example, if there is a cache miss at a given level 2 (L2) cache 540, it can be determined whether the data and / or instructions are cached in a shared level 3 (L3) cache 550. At 625, if the data and / or instructions for a given physical page number are not found in a given lower-level cache, the data and / or instructions for the given physical page number of the memory access request can be fetched from a further lower-level cache or from memory and placed in the given lower-level cache and the given higher-level cache. For example, if the data and / or instructions are not found in the shared level 3 (L3) cache 550, the data and / or instructions can be fetched from the next lower-level cache (if applicable) or from memory 165 - 170, and the fetched data and / or instructions can be placed in the shared level 3 (L3) cache 550 and the given level 2 (L2) cache 540. At 630, the given physical page number and identifier of the core maintaining the current memory access request are maintained as part of the information regarding a previous memory access request. For example, the core sharing agent 580 can be configured to add the given physical page number and core number of the current memory access request to a data array 710 that includes the physical page numbers and core numbers of other memory access requests, as Figure 7 shown. In one implementation, the data array 710 can include one or more sets of physical page numbers and corresponding identifiers, such as the core number of the core that last accessed the physical page number for a previous memory access request. The core sharing agent 580 can thus act as a fully / group-associative cache, where if it is group-associative, the physical page numbers in the table are used as tag bits and index bits, and the core numbers are stored in the data array of the cache.
[0033] At 635, if data and / or instructions for the given physical page number are found in a given lower-level cache, the data and / or instructions can be fetched from the given lower-level cache and placed in a given higher-level cache. For example, data and / or instructions can be fetched from a shared level-three (L3) cache 550 and placed in a given level-two (L2) cache 540. Additionally, at 640, it can be determined whether the given core of the current access request is the same as one of the cores maintained in information regarding a previous memory access request for the given physical page number. For example, the core sharing agent 580 can be configured to determine whether the physical page number of the current memory access request matches a physical page number in a data array. If a matching physical page number exists in the data array 710, it can be determined whether the core number of the current memory access request matches the core number associated with the matching physical page number in the data array 710. At 645, if the given core of the current memory access is not the same as any of the cores maintained in information regarding a previous memory access request for the given physical page number, the cache line for the given physical page number that was fetched can be maintained in a lower-level shared cache. Additionally, at 650, if the given core of the current memory access is different from the core in the information regarding a previous access request for the given physical page number, information regarding the given core of the current memory access request can be maintained together with information regarding other cores that have accessed the given physical page number. At 655, if the given core of the current memory access is the same as one of the cores maintained in information regarding a previous memory access to the given physical page number, the data and / or instructions for the given physical page number that were fetched can be removed from the lower-level shared cache.
[0034] In a core-sharing-aware non-inclusive non-exclusive cache method, a core number identifier can identify 128 cores with one byte. Thus, compared to the following cache method based on a core valid bit vector, the core-sharing-aware non-inclusive non-exclusive cache method that utilizes a core number identifier provides relatively coarse-grained cache control.
[0035] Now refer to Figures 8A-8B, which shows a core - shared aware non - inclusive non - exclusive cache method according to some aspects of the present application. The method may include, at 805, receiving a current memory access request from a given one of a plurality of cores. At 810, determining whether data and / or instructions for a given physical page number of the memory access request are cached in a given higher - level cache specific to (private, allocated, exclusive, etc.) each given core. For example, it may be determined whether data and / or instructions are cached in a given level - 2 (L2) cache 540. At 815, if the data and / or instructions for the given physical page number are found in the given higher - level cache (e.g., cache hit), the data and / or instructions may be retrieved from the given higher - level cache and placed in a given further - higher - level cache or returned to the given one of the plurality of cores according to the corresponding cache policy. For example, the data and / or instructions may be retrieved from the given level - 2 (L2) cache 540 and placed in a given level - 1 (L1) cache 525 and / or returned to the given core 510.
[0036] At 820, if the data and / or instructions for the given physical page number are not found in the given higher - level cache (e.g., cache miss), it may be determined whether the data and / or instructions for the given physical page number of the memory access request are cached in a given lower - level shared cache. For example, if a cache miss occurs at the given level - 2 (L2) cache 540, it may be determined whether data and / or instructions are cached in a shared level - 3 (L3) cache 550. At 825, if the data and / or instructions for the given physical page number are not found in the given lower - level cache, the data and / or instructions for the given physical page number of the memory access request may be fetched from a further - lower - level cache or from memory and placed in the given lower - level cache and the given higher - level cache. For example, if the data and / or instructions are not found in the shared level - 3 (L3) cache 550, the data and / or instructions may be fetched from the next - lower - level cache (if applicable) or from memory 65 - 570. The fetched data and / or instructions may be placed in the shared level - 3 (L3) cache 550 and the given level - 2 (L2) cache 540. At 830, maintain information about the given physical page number and the given core of the current memory access request as part of the information about a previous memory access request. For example, as Figure 9As shown, the core sharing agent 580 can be configured to add in the data array 910 the bits of a given physical page number and the core valid bit vector corresponding to the corresponding core for the current memory access request. In one implementation, the data array 910 can include one or more sets of physical page numbers and corresponding core valid bit vectors, where the core valid bit vector includes one bit for each of the multiple computing cores of the processor.
[0037] At 835, if the data and / or instructions of the given physical page number are found in a given lower-level cache, the data and / or instructions can be fetched from the given lower-level cache and placed in a given higher-level cache. For example, the data and / or instructions can be fetched from the shared level-3 (L3) cache 550 and placed in a given level-2 (L2) cache 540. Additionally, at 840, it can be determined that one or more other cores among the multiple cores have previously accessed the given physical page number of the memory access request. For example, the core sharing agent 580 can be configured to determine whether one or more bits of the corresponding core valid bit vector in the data array 910 for the given physical page number of the current memory access request are in a given state, where the given state is used to indicate that one or more other cores have previously accessed the given physical page number. At 845, if one or more bits in the corresponding core valid bit vector in the data array 910 indicate that one or more other cores have accessed the given physical page number, the fetched cache line of the given physical page number can be maintained in the lower-level shared cache. Additionally, at 850, if one or more other cores have previously accessed the given physical page number, the information about the given core for the memory access request is maintained together with the information about the other cores that have accessed the given physical page number. At 855, if one or more other cores have not accessed the given physical page number, the fetched data and / or instructions of the given physical page number can be removed from the lower-level shared cache. In one implementation, the core valid bit vector in the core sharing agent data array 910 can be reset so that the data in the instructions of the corresponding physical page number is not continuously maintained in the lower-level shared cache.
[0038] The core sharing-aware non-inclusive non-exclusive caching method using the core valid bit vector can conveniently implement fine-grained cache control. The core valid bit vector can conveniently record the core access history over a period of time. Therefore, when multiple cores have accessed the corresponding physical page number, the fetched cache line can be maintained in the lower-level shared cache based on the corresponding valid core bits. However, compared with the core number identifier, the core sharing-aware non-inclusive non-exclusive caching method using the core valid bit vector has a higher storage overhead because one byte of the core valid bit vector can only represent eight computing cores.
[0039] Some aspects of the present application advantageously provide a non-inclusive non-exclusive cache policy based on core sharing behavior. The non-inclusive non-exclusive cache policy according to some aspects of the application advantageously achieves a relatively large effective capacity similar to that of an exclusive cache policy. In the case of inter-core data sharing, the non-inclusive non-exclusive cache policy according to some aspects of the present application effectively reduces cache misses.
[0040] For purposes of illustration and description, the foregoing description of specific embodiments of the present application has been presented. They are not intended to be exhaustive or to limit the application to the precise forms disclosed, and obviously many modifications and variations are possible in light of the above teachings. The selected and described embodiments are best explained the principles of the present application and its practical application, so that others skilled in the art can best utilize the present application and various embodiments with various modifications suitable for the intended specific purposes. The scope of the present application is intended to be defined by the appended claims and their equivalents.
Claims
1. A non-inclusive non-exclusive (NINE) caching method, comprising: Receive a memory access request from one or more of multiple cores; and Perform core-aware non-inclusive non-exclusive caching of data and / or instructions between a shared cache level and a core-specific cache level based on a physical page number (PPN) and a set of core identifiers for previous accesses to respective physical page numbers, including: Determine whether data and / or instructions for a given physical page number of a current memory access request received from a given core of a plurality of cores of a processor are cached in a lower-level shared cache; When data and / or instructions for a given physical page number of a current memory access request are not cached in a lower-level shared cache, maintain the given physical page number and identifier of the core of the current memory access request as part of information regarding previous memory access requests.
2. The non-inclusive non-exclusive caching method according to claim 1, further comprising: When data and / or instructions for a given physical page number of a current memory access request are not cached in a lower-level shared cache, fetch the data and / or instructions for the given physical page number of the current memory access request from a further lower-level cache or memory and place them in the lower-level cache and a given higher-level cache; When data and / or instructions for a given physical page number of a current memory access request are cached in a lower-level shared cache, fetch the data and / or instructions for the given physical page number of the current memory access request from the given lower-level cache and place them in the given higher-level cache; When data and / or instructions for a given physical page number of a current memory access request are cached in a lower-level shared cache, determine whether the given core of the current memory access is the same as the core in the information maintained regarding the previous memory access request to the given physical page number; When the given core of the current memory access is different from the core in the information maintained regarding the previous memory access request to the given physical page number, maintain the fetched data and / or instructions for the given physical page number in the lower-level shared cache; When the given core of the current memory access is different from the core in the information maintained regarding the previous memory access request to the given physical page number, maintain information about the given core of the current memory access request together with information about other cores that have accessed the given physical page number; and When the given core of the current memory access is the same as the core maintained in the information regarding the previous memory access request to the given physical page number, remove the fetched data and / or instructions for the given physical page number from the lower-level shared cache.
3. The non-inclusive non-exclusive caching method according to claim 1, wherein When data and / or instructions for a given physical page number of a current memory access request are not cached in a lower-level shared cache, maintain the given physical page number and given core of the current memory access request as part of information regarding previous memory access requests, including: Add the given physical page number and a corresponding core valid bit vector to a data array, where one bit of the core valid bit vector corresponding to the given core is set to a given state.
4. The non-inclusive non-exclusive caching method according to claim 3, wherein When one or more other cores among multiple cores have previously accessed the given physical page number of the current memory access request, maintain information about the given core of the current memory access request together with information about other cores that have accessed the given physical page number, including: Set one bit of the core valid bit vector corresponding to the given core to a given state in the core valid bit vector corresponding to the physical page number of the current memory access request.
5. The non-inclusive non-exclusive caching method according to claim 2, further comprising: Determine whether the data and / or instructions of the given physical page number of the current memory access request are cached in a given higher-level cache specific to each given core; and Fetch the data and / or instructions of the given physical page number of the current memory access request from the given higher-level cache and place them in a given further higher-level cache or return them to a given one of the multiple cores according to the corresponding cache policy.
6. The non-inclusive non-exclusive caching method according to claim 1, wherein the lower-level shared cache includes the lowest-level cache of the processor.
7. The non-inclusive non-exclusive caching method according to claim 6, wherein a given high-level cache is dedicated to a given one of a plurality of computing cores.
8. A non-inclusive non-exclusive caching method, comprising: Receive memory access requests from one or more of the multiple cores; and Perform core-aware non-inclusive non-exclusive caching of data and / or instructions between a shared cache level and a core-specific cache level based on the physical page number and a set of core valid bit vectors for each physical page number previously accessed by each of the multiple cores, including: Determine whether the data and / or instructions of a given physical page number (PPN) of a current memory access request received from a given core of the multiple cores of a processor are cached in a lower-level shared cache; When the data and / or instructions of the given physical page number of the current memory access request are not cached in the lower-level shared cache, maintain information about the given physical page number and the given core of the current memory access request as part of the information about previous memory access requests.
9. The non-inclusive non-exclusive (NINE) caching method according to claim 8, further comprising: When the data and / or instructions of the given physical page number of the current memory access request are not cached in the lower-level shared cache, fetch the data and / or instructions of the given physical page number of the current memory access request from a further lower-level cache or memory and place the data and / or instructions in the lower-level cache and a given higher-level cache; When the data and / or instructions of the given physical page number of the current memory access request are cached in the lower-level shared cache, fetch the data and / or instructions of the given physical page number of the current memory access request from the given lower-level cache and place them in a given higher-level cache; When the data and / or instructions of the given physical page number of the current memory access request are cached in the lower-level shared cache, determine whether one or more other cores among the multiple cores have previously accessed the given physical page number of the current memory access request; When one or more other cores among the multiple cores have previously accessed the given physical page number of the current memory access request, maintain the fetched data and / or instructions for the given physical page number in the lower-level shared cache; When one or more other cores among the multiple cores have previously accessed the given physical page number of the current memory access, maintain information about the given core of the current memory access request together with information about the other cores that have accessed the given physical page number; When one or more other cores among the multiple cores have not previously accessed the given physical page number of the current memory access, remove the data and / or instructions of the obtained given physical page number from the lower-level shared cache.
10. The non-inclusive non-exclusive cache method according to claim 8, wherein, When the data and / or instructions of the given physical page number of the current memory access request are not cached in the lower-level shared cache, maintain information about the given physical page number of the current memory access request as part of the information about the previous memory access request, including: Add the given physical page number and the corresponding core valid bit vector to the data array, where one bit of the core valid bit vector corresponding to the given core is set to a given state.
11. The non-inclusive non-exclusive cache method according to claim 10, wherein, When one or more other cores among the multiple cores have previously accessed the given physical page number of the current memory access, maintain information about the given core of the current memory access request together with information about the other cores that have accessed the given physical page number, including: Set one bit of the core valid bit vector corresponding to the given core to the given state in the core valid bit vector corresponding to the physical page number of the current memory access request.
12. The non-inclusive non-exclusive cache method according to claim 9, further comprising: Determine whether the data and / or instructions of the given physical page number of the current memory access request are cached in a given higher-level cache specific to each given core; and Obtain the data and / or instructions of the given physical page number of the current memory access request from the given higher-level cache and place them in a given further higher-level cache according to the corresponding cache policy or return them to a given one of the multiple cores.
13. The non-inclusive non-exclusive cache method according to claim 8, wherein the lower-level shared cache includes the lowest-level cache of the processor.
14. The non-inclusive non-exclusive cache method according to claim 13, wherein a given higher-level cache is dedicated to a given one of the multiple computing cores.
15. A processor, comprising: Multiple computing cores; One or more cache levels specific to each of the multiple computing cores; One or more cache levels shared by the multiple computing cores; and A core sharing agent configured to perform non-inclusive non-exclusive (NINE) caching of data and / or instructions in the shared cache layer relative to the core-specific cache layer based on the core sharing behavior of the shared cache layer, wherein the core sharing agent is configured to perform core-aware non-inclusive non-exclusive caching of data and / or instructions in the shared cache layer relative to the core-specific cache layer based on a core number identifier, or the core sharing agent is configured to perform core-aware non-inclusive non-exclusive caching of data and / or instructions in the shared cache layer relative to the core-specific cache layer based on a core valid bit vector.
16. The processor according to claim 15, wherein the core sharing agent is configured to: Determine whether the data and / or instructions of a given physical page number (PPN) of a current memory access request received from a given core of the multiple cores of the processor are cached in the lower-level shared cache; When the data and / or instructions of the given physical page number of the current memory access request are not cached in a lower-level shared cache, obtain the data and / or instructions of the given physical page number of the current memory access request from a further lower-level cache or memory, and place the data and / or instructions in the lower-level cache and a given higher-level cache; When the data and / or instructions of the given physical page number of the current memory access request are not cached in a lower-level shared cache, maintain information about the given physical page number and the given core of the current memory access request as part of the information about previous memory access requests; When the data and / or instructions of the given physical page number of the current memory access request are cached in a lower-level shared cache, obtain the data and / or instructions of the given physical page number of the current memory access request from the given lower-level cache, and place them in a given higher-level cache; When the data and / or instructions of the given physical page number of the current memory access request are cached in a lower-level shared cache, determine whether one or more other cores among the multiple cores have previously accessed the given physical page number of the current memory access request; When one or more other cores among the multiple cores have previously accessed the given physical page number of the current memory access request, maintain the obtained data and / or instructions for the given physical page number in the lower-level shared cache; When one or more other cores among the multiple cores have previously accessed the given physical page number of the current memory access, maintain the information about the given core of the current memory access request together with the information about the other cores that have accessed the given physical page number; When one or more other cores among the multiple cores have not previously accessed the given physical page number of the current memory access, remove the obtained data and / or instructions of the given physical page number from the lower-level shared cache.
17. The processor according to claim 16, characterized in that The lower-level shared cache includes the lowest-level cache of the processor.
18. The processor according to claim 17, wherein The given higher-level cache is dedicated to a given one of the multiple computing cores.
19. The processor according to claim 16, wherein The memory includes one or more dynamic random access memories (DRAM).
Citation Information
Patent Citations
Virtual address cache memory, processor and multiprocessor
US20110231593A1