Multiple copy scoping bits for cache memory
Multiple copy scoping bits in SMP systems enhance cache coherence by defining sharing scopes, reducing query times and operational latency, and optimizing cache management in SMP systems.
Patent Information
- Application Number
- JP2024504193
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2021-08-20
- Filing Date
- 2022-08-15
- Publication Date
- 2025-10-22
- Estimated Expiration
- 2042-08-15
AI Technical Summary
Maintaining cache coherency in symmetric multiprocessing (SMP) systems is time-consuming due to extended delays from system-wide queries to determine cache line copies, leading to performance inefficiencies.
Implementing multiple copy scoping bits to define the scope of cache line sharing within the system, reducing the need for system-level queries by utilizing chip, module, drawer, and system scopes to manage cache coherence.
Reduces query time and operational latency, conserves bus bandwidth, and minimizes controller resource consumption by optimizing cache state changes and maintaining efficient cache management.
Smart Images

Figure 0007758846000002 
Figure 0007758846000003 
Figure 0007758846000004
Abstract
Description
[Technical Field]
[0001] The present invention relates generally to data processing, and more particularly to multiple copy scoping bits for cache memories in symmetric multiprocessing computers. [Background technology]
[0002] Modern high-performance computer systems are typically implemented as multinode symmetric multiprocessing ("SMP") computers with many compute nodes. SMP is a multiprocessor computer hardware architecture in which two or more, typically many, identical processors are connected to a single shared main memory and controlled by a single operating system. Today, most multiprocessor systems use SMP architecture. In the case of multicore processors, the SMP architecture applies to the cores, treating them as separate processors. Processors may be interconnected using buses, crossbar switches, mesh networks, and the like. Each compute node typically contains several processors, each of which may have at least some local memory, at least some of which is accelerated by cache memory. Cache memory can be local to each processor, local to the compute node shared across two or more processors, or shared across the entire node. All of these architectures require maintaining cache coherence between the separate caches.
[0003] Maintaining cache coherency can become a time-consuming process as more caches are interconnected in an SMP architecture. Querying multiple caches throughout the system to determine where copies of a cache line reside can result in extended delays while waiting for the query to propagate through the system and for a response to the query, thereby reducing system performance. Summary of the Invention
[0004] According to one or more embodiments of the present invention, a computer-implemented method includes accessing a multicopy scope directory state of a cache memory indicating a sharing scope of a cache line in the cache memory system, and determining a sharing scope of a cache line in the cache memory system based on the multicopy scope directory state, where the multicopy scope directory state enumerates multiple scopes in the cache memory system. The sharing scope is used to reduce the number of queries to one or more cache memories having a sharing scope greater than the sharing scope identified in the sharing scope. The multicopy scope directory state of the cache memory is updated based on detecting a change in the sharing scope of a cache line in the cache memory system. Advantages may include faster updates in the cache memory system.
[0005] According to additional or alternative embodiments of the present invention, a scope of a share having a smaller scope may be a subset of a scope of a share having a larger scope within a cache memory system.
[0006] According to additional or alternative embodiments of the present invention, the scope of sharing may include a chip scope, a module scope, a drawer scope as the smallest scope, and a system scope as the largest scope. Advantages may include establishing scopes associated with a hierarchy of system components.
[0007] According to additional or alternative embodiments of the present invention, the operations may include, based on determining that the requesting cache has requested a shared copy of the cache line, setting a snoop scope to a minimum scope, sending a query to one or more remote cache memories within the current snoop scope that have not previously been queried, determining whether an intervention master copy of the cache line was found in the remote cache memories of the current snoop scope or encountered the cache line in a forward state, installing the shared copy of the cache line in the requesting cache, and updating a state of a directory entry holding a master copy of the cache line based on whether the intervention master copy of the cache line was found in the remote cache memories of the current snoop scope or encountered the cache line in a forward state. Advantages may include requesting shared copies of the cache line through increasingly larger scopes to reduce request and response times.
[0008] According to additional or alternative embodiments of the present invention, the operations can include fetching a cache line from a cache memory and, after not finding the cache line within an initial snoop scope, updating a request cache directory and at least one remote sourcing directory for the cache line based on finding one or more copies of the cache line in a shared non-master state. Advantages can include maintaining directory state for more efficient cache management.
[0009] According to additional or alternative embodiments of the present invention, the operations can include determining whether the requesting cache has an intervening master shared copy of the cache line based on the requesting cache requesting an exclusive copy of the cache line, requesting invalidation of the cache line in all cache memories within a minimum snoop scope that includes the directory scope, and updating the requesting cache directory for the cache line in the requesting core to an exclusive state. Advantages can include state management of shared cache lines using scope information.
[0010] According to additional or alternative embodiments of the present invention, the operations can include setting a current snoop scope to a minimum snoop scope based on determining that the requesting cache does not have an intervening master shared copy of the cache line, sending a query to one or more remote caches within the current snoop scope that have not previously been queried, updating a remote cache directory to an invalid state based on a shared non-intervening master copy of the cache line within the current snoop scope, updating a requesting cache directory for the cache line to an exclusive state based on determining that an intervening master copy of the cache line is not found in a remote cache within the current snoop scope and the maximum snoop scope has been reached, and setting the current snoop scope to a next larger snoop scope based on determining that an intervening master copy of the cache line is not found in a remote cache within the current snoop scope and the maximum snoop scope has not been reached. Advantages can include tiered queries based on scope for cache state management.
[0011] According to additional or alternative embodiments of the present invention, the operations can include updating a remote cache directory for the cache line to an invalid state and installing an exclusive copy of the cache line in an exclusive state based on finding a remote cache with an exclusive copy of the cache line within the current snoop scope and determining that no intervening master copy of the cache line is found within the remote cache within the current snoop scope, and updating a remote cache directory for the cache line to an invalid state and updating a request cache directory for the cache line to an exclusive state based on finding a remote cache with an intervening master shared copy of the cache line within the current snoop scope, where the directory scope in the remote cache is less than or equal to the current snoop scope. Advantages can include managing exclusive copies of a cache line in multiple caches.
[0012] According to additional or alternative embodiments of the present invention, the operations can include updating a remote cache directory for the cache line to an invalid state, and, based on the directory scope in the remote cache being greater than the current snoop scope, sending a query to the remote cache within a minimum snoop scope greater than the directory scope in the remote cache to which the query has not been sent, updating the remote cache directory for the cache line to an invalid state, and updating a request cache directory for the cache line to an exclusive state. Advantages can include managing cache directories in multiple scopes.
[0013] According to additional or alternative embodiments of the present invention, the operations may include using the intervening master copy of the cache line before the entire system state is resolved based on detecting that the primary scope does not include an intervening master copy of the cache line, the secondary scope includes an intervening master copy of the cache line, and a larger scope broadcast has been initiated. Benefits may include earlier use of the intervening master copy of the cache line.
[0014] Other embodiments of the present invention embody features of the above-described methods in computer systems and computer program products.
[0015] Additional technical features and advantages are realized through the techniques of the present invention. Embodiments and aspects of the invention are described in detail herein and are considered a part of the claimed subject matter. For a better understanding, reference is made to the detailed description and drawings.
[0016] The particulars of the exclusive rights set forth herein are particularly pointed out and distinctly claimed in the claims at the conclusion of this specification. The foregoing and other features and advantages of embodiments of the present invention will become apparent from the following detailed description taken in conjunction with the accompanying drawings. [Brief explanation of the drawings]
[0017] [Figure 1] 1 is a diagram of a distributed symmetric multiprocessing (SMP) system in accordance with one or more embodiments of the present invention. [Figure 2] 1 is a block diagram of a distributed symmetric multiprocessing (SMP) system utilizing multiple copy scoping bits for cache memory in accordance with one or more embodiments of the present invention. [Figure 3] FIG. 1 is a block diagram of multiple scopes of cache memory in accordance with one or more embodiments of the present invention. [Figure 4]1 is a flow diagram of a method for using multiple copy scoping bits for a cache memory in accordance with one or more embodiments of the present invention. [Figure 5] 1 is a flow diagram of a method for requesting a shared copy of a cache line in accordance with one or more embodiments of the present invention. [Figure 6A] 1 is a flow diagram of a method for requesting an exclusive copy of a cache line in accordance with one or more embodiments of the present invention. [Figure 6B] 1 is a flow diagram of a method for requesting an exclusive copy of a cache line in accordance with one or more embodiments of the present invention. [Figure 6C] 1 is a flow diagram of a method for requesting an exclusive copy of a cache line in accordance with one or more embodiments of the present invention. [Figure 7] FIG. 1 is a block diagram of a computer system in accordance with one or more embodiments of the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0018] The diagrams depicted herein are exemplary. Many variations to the diagrams or operations described herein are possible without departing from the spirit of the invention. For example, actions can be performed in a different order, or actions can be added, deleted, or modified. Also, the term "coupled" and variations thereof indicate that there is a communication path between two elements and do not imply a direct connection between elements without an intervening element / connection between them. All of these variations are considered to be part of this specification.
[0019] One or more embodiments of the present invention provide systems and methods for using multiple copy scoping bits for cache memory in an SMP environment. The multiple copy scoping bits can define the scope of sharing of cache lines across lateral caches, for example, in an SMP environment. As further described herein, the multiple copy scoping bits can reduce latency and improve performance of cache state changes for cache lines.
[0020] The traditional approach to cache line management uses the Modify-Exclusive-Shared-Invalidate (MESI) protocol to update a cache line from a shared state to an exclusive state by applying several rules. The Modified state can indicate that the cache line exists only in the associated cache, has been modified, and therefore differs from the value stored in main memory. The Exclusive state can indicate that the cache line exists only in the associated cache, is unmodified, and therefore the value matches main memory. The Shared state can indicate that the cache line may be stored in one or more other caches, is unmodified, and therefore the value matches main memory. The Invalid state can indicate that the cache line is invalid; for example, its contents should not be used and can be overwritten. Other protocol variants include the MESIF protocol, which adds forwarding state, the MOESI protocol, which adds owned state, and other such variants. When caches are distributed across multiple system components, system-level queries can be used to resolve the coherent state of the system. To switch a cache line from shared to exclusive state, all copies of the cache line in the system are typically invalidated. While this approach can work to maintain cache coherency, it can result in performance inefficiencies in larger systems. For example, a system-wide query typically results in inefficient bus usage and interface bandwidth. A system-wide query can also result in consuming controller resources and increased operation latency.
[0021] Embodiments can reduce the impact of system-level queries, for example, when switching a cache line from a shared state to an exclusive state. By augmenting scope tracking with multi-bit shared state, the scope of sharing can be defined within the system. The enhanced scope information can be used to establish a system coherence protocol that reduces system-level queries to a reduced domain scope. For example, caches at the same level, such as a level 3 cache, may be distributed throughout the system with chip-scope, module-scope, and drawer-scope. In this example, chip-scoped cache can refer to chip-level cache distributed among multiple processing cores on the same chip. Module-scoped cache can group caches where multiple chips are connected to the same module, such as two chips on the same card or board. Drawer-scoped cache can group caches where multiple modules are connected together into a physical or logical grouping, such as a drawer on a card or board, where the card or board within the drawer is considered a module. System-scoped cache can group all caches in a system together, such as multiple drawers. Embodiments recognize that latency and access times can vary within a system depending on how close cache lines are in scope. For example, accessing a cache line within the same chip can be faster than accessing a cache line within a different chip on the same module. Similarly, accessing a cache line within the same module can be faster than accessing a cache line within a different module. Access and response times can increase across drawers.
[0022] Query time can be reduced by maintaining a multiple copy scope directory bit that indicates the proximity of shared copies of a cache line based on where the cache line was last fetched from. For example, if the scope bit indicates that all copies contain the local chip, the requester need not query beyond the chip scope, avoiding the delay associated with querying the entire system. Similarly, if the scope bit indicates that all copies contain the module, the requester need not query beyond the module scope. If the scope bit indicates that all copies contain the drawer, the requester need not query beyond the drawer scope. As cache lines move through the system, the multiple copy scope state can be continually updated to reflect the furthest scope that has been seen and cleaned up as recollection occurs.
[0023] When the multiple copy scope bits indicate that all copies are contained within the local chip, there is no need to venture out into the system. Therefore, fabric broadcasts can be suppressed. When the multiple copy scope bits indicate that not all copies are contained within the local chip, a drawer-wide broadcast can occur. When the scope bits indicate that all copies are contained within the local module, the requester can leverage this information by not having to wait for responses from all remote chips. When not all copies are contained within the local module, the requester waits for the drawer response. When the drawer response indicates that not all copies are contained within the local drawer, a system-wide query can be made to search the next level, for example, multiple drawers. The requester can build information from one scoping level to another to determine how far into the system to venture to find and invalidate all copies of the cache line.
[0024] Depending on the state of the multi-copy bit, the request type, and the architectural design, a drawer-wide broadcast can be proactively performed in the event of a chip-scoped failure. If the bit indicates that all copies are contained within the local module, the requester can leverage this information by not having to wait for responses from all remote chips when a proactive drawer query is made. In scenarios where not all copies are contained within the local module, the requester must wait for the drawer's response. Scoping information is aggregated as the request traverses the system as part of the normal consistency flow.
[0025] In embodiments, the system can maintain multi-copy scope directory state, including state for directing chip containment. The fabric can reduce the need to venture into the system because other copies are not outside the chip scope. Operational latency can be reduced because requesters do not need to wait for responses from remote chips. Additionally, bandwidth can be saved and controller busy time can be reduced because remote caches outside the local chip do not need to be queried.
[0026] If not all copies are contained within the local chip, a drawer-wide broadcast can occur. By querying a sibling chip and determining that the multiple copy scope bits indicate that all copies are contained within the local module, the requester can take advantage of this information by not having to wait for responses from all other remote chips. Controller and bus utilization may not be significantly reduced in such cases, but the latency of the operation may be reduced. If the scoping bits from the drawer's response indicate that not all copies are contained within the local drawer, a system-wide broadcast can occur to search the next scoping level, such as a multi-drawer scope. Some bus configurations can allow sibling chip queries to occur simultaneously with local chip queries. If all copies of a cache line are contained within the local module, no further system outreach is necessary. Therefore, drawer-wide broadcasts can be suppressed.
[0027] FIG. 1 depicts a distributed symmetric multiprocessing (SMP) system 100 (hereinafter "system 100") according to one or more embodiments. System 100 can include four processing units, or "drawers." Each drawer 102, 104, 106, 108, in this example, includes eight (8) microprocessor (CP) chips (CP0-CP7). Each CP chip can include eight (8) cores. Each core within a CP chip includes a private L1 cache (including both an instruction cache and a data cache). These private L1 caches are backed by a semi-private L2 cache. The semi-private L2 caches can interact to provide an on-chip virtual level 3 (L3) cache. Each drawer 102, 104, 106, 108 can include up to eight CP chips with a fully connected topology providing a virtual level 4 (L4) cache. The virtual L3 and virtual L4 caches can be realized to act as a unified shared victim cache through a set of chip caching techniques that cluster independent physical L2 caches within the chip and within the drawer.
[0028] A virtual L3 / L4 cache can be implemented by defining a group / cluster of L2 caches within a CP chip, a group of CP chips, or a drawer for evicting cache lines from peer caches, or a combination thereof. That is, cache lines can be evicted from a first L2 to a peer L2 within a defined group / cluster of L2 caches according to a specified replacement policy. Embodiments can utilize virtual or physical caches using multiple copy scoping bits.
[0029] FIG. 2 depicts a block diagram of a distributed symmetric multiprocessing (SMP) system 200 with a cache memory system 201 in accordance with one or more embodiments of the present invention. System 200 includes four drawers 240-0, 240-1, 240-2, and 240-3. Each drawer includes the same components as those described for drawer 240-0. Drawer 240-0 includes eight CP chips 202-0 through 202-7. Each CP chip 202 includes eight cores 204, each with a private level 1 (L1) cache 206. Each core 204 also utilizes a semi-private level 2 (L2) cache 208. The cache memory system 201 may include cache memory distributed through a hierarchy of levels L1, L2, L3, L4, etc., and across physical locations such as drawers 240-0, 240-1, 240-2, 240-3.
[0030] Peer L2 caches (sometimes called "lateral caches") can be divided into cache clusters 214. Each drawer 240 can include one or more cache clusters 214 utilized for cache lines. While the illustrative example shows one configuration of cache clusters 214, in one or more embodiments, cache clusters 214 can include any number of L2 caches in any type of configuration, including across drawers with L2 caches in groups / clusters. Caches can be clustered within the same drawer 240-0 or within other drawers 240-1, 240-2, 240-3. Cache lines can be fetched or updated by one or more processing cores 204. Cache management policies can be implemented using a cache controller 212 to manage the cache contents and state between cache clusters 214 relative to main memory 220. The replacement policy for the lateral cache 208 on the CP chip 202 can be preferential. The CP chip 202 can have more than one defined cache cluster 214, with eight on the CP chip 202 in this example.
[0031] In an embodiment, whether implemented as a virtual or physical cache, multiple copy scoping bits may be used to manage shared cache lines in one or more cache levels that are non-private, such as in L2 cache 208 or in an L3 or L4 cache within cache cluster 214. Cache controller 212 may manage the use of multiple copy scoping bits throughout system 200.
[0032] FIG. 3 depicts a block diagram of multiple scopes 300 of cache memory, according to one or more embodiments. FIG. 3 illustrates an example of cache scopes for the L3 caches of multiple cores (core 0 through core 63) within a drawer scope 302. In the example of FIG. 3, a module scope 304 is smaller than the drawer scope 302 and includes multiple chip scopes 306. The L3 caches of cores 0, 1, 2, 3, 4, 5, 6, and 7 may all be within the same chip scope 306, while the L3 caches of cores 8, 9, 10, 11, 12, 13, 14, and 15 may be within a different chip scope 306 but within the same module scope 304 as the L3 caches of cores 0 through 7. The L3 caches of cores 16 through 63 may be within a different chip scope 306 and module scope 304 than cores 0 through 15. For example, chip scope 306 including the L3 caches of cores 0 through 7 may correspond to CP chip 202-0 in Figure 2, and the L3 caches of cores 8 through 15 may correspond to CP chip 202-1 in Figure 2, with CP chips 202-0 and 202-1 being within the same module scope 304. Cache controller 212 in Figure 2 may set and check cache directory bits for managing each cache.
[0033] In the example of FIG. 3, the L3 cache of core 0 can be a requesting cache, requesting a shared copy of a cache line. A copy of the cache line may be shared, for example, within the L3 caches of core 6, core 13, and core 33. Relative to core 0, the L3 cache of core 6 is a cache memory within the same chip scope 306, the L3 cache of core 13 is a remote cache memory within the same module scope 304, and the L3 cache of core 33 is a remote cache memory within the same drawer scope 302. If the only shared copy of the cache line in core 0's L3 cache was within core 6, response time would be improved because a system-wide or drawer-wide cache check would be avoided. For example, if the maximum scope of a shared cache line for core 13's L3 cache was module scope 304, response time could be improved by checking only within module scope 304 rather than performing a system-wide or drawer-wide check. If drawer scope 302 is the broadest scope, response time may be improved by avoiding system-wide checks, such as those for drawers 240-1, 240-2, and 240-3 in FIG.
[0034] In the example of FIG. 3, the L3 cache of core 0 has multicopy scope directory state 310, which indicates the scope of sharing for cache line 311 in the L3 cache of core 0. Each cache line may contain multiple associated cache entries containing copies of data from different memory sources, such as main memory 220 in FIG. 2. The directory for each cache may include state information defining the status and scope of sharing for each cache line. Similarly, the L3 cache of core 6 has multicopy scope directory state 312, which indicates the scope of sharing for cache line 313 in the L3 cache of core 6. The L3 cache of core 13 has multicopy scope directory state 314, which indicates the scope of sharing for cache line 315 in the L3 cache of core 13. The L3 cache of core 33 has multicopy scope directory state 316, which indicates the scope of sharing for cache line 317 in the L3 cache of core 33. Examples of cache states that can be encoded in the bits of the multicopy scope directory states 310, 312, 314, 316 for each associated cache line 311, 313, 315, 317 stored in the L3 cache are illustrated in Table 1, where the multicopy scope directory states 310, 312, 314, 316 can enumerate multiple scopes with a cache memory system, such as the L3 cache of cache memory system 201 of FIG. 2.
[0035] [Table 1]
[0036] In Table 1, the exclusive state means intervening master and is treated as the minimum directory scope. Shared scopes with smaller scopes are subsets of shared scopes with larger scopes in a cache memory system, for example, forming a scope hierarchy. In this example, chip scope is the minimum shared scope and system scope is the maximum shared scope. For each line in the cache, there is a directory entry in the cache directory. For each cache included in the system, there is a cache directory. Some systems may have a cacheless directory for tracking. Embodiments can use directory entries and can operate in systems with cacheless directories.
[0037] The use of chip, module, drawer, and system scopes is an exemplary embodiment. These can represent any hierarchical grouping of caches into multiple scopes.
[0038] For purposes of explanation, a snoop scope is a scope that can be snooped. For example, it may be snoopable only in chip, drawer, and system scopes. A multicopy scope (or directory scope) is a scope that can be represented by a directory. It may be a snoop scope or a subset of a snoop scope. For example, module scope can be a subset of drawer scope, which is also a snoop scope.
[0039] The intervening master is the most recently cached copy of a cache line in the system, or when using the MESIF protocol, when encountering a cache line in the forwarding state. This forms the highest point of coherency in the system for a given cached line, when one or more copies exist. Therefore, a non-intervening master line is considered to have an out-of-date directory state.
[0040] Cache line requests may come from a variety of sources. For example, a cache line can enter a first cache from memory with tip scope (e.g., dedicated to a cache tip scope using simple directory state). Additionally, the cache line can be fetched by another cache in a scope that remains in tip scope or is shared by two caches. A cache line can be fetched by another cache in tip scope with exclusive intent, with the other cache invalidated, and the requesting cache can install an exclusive cache line with tip scope.
[0041] As another example, a cache line can be fetched by another cache in a module that goes to module scope, where module scope can be indicated for both the source and destination caches. A cache line can be fetched with exclusive intent by yet another cache in a module and returned to exclusive tip scope where other copies are invalidated.
[0042] Benefits may include limiting fetch snoop scope within the chip, with chip directory lookups returning chip-scope state. In other types of systems, this may result in a max-scope broadcast, delaying the response to the requester until all remote copies are invalidated. Embodiments not only reduce the latency experienced by the request processor, but also eliminate the max-scope broadcast and all associated packets, improving system resource availability for other requests in the system. Furthermore, instead of traffic reduction being observed at the entire system level, traffic reduction can be observed for a partial subset of the system.
[0043] FIG. 4 depicts a flow diagram of a method 400 for using multiple copy scoping bits for a cache memory, in accordance with one or more embodiments of the present invention. At least a portion of method 400 may be performed, for example, by processor 701 shown in FIG. 7. Furthermore, method 400 may be performed in systems 100, 200 of FIGS. 1 and 2. For example, method 400 may be implemented by cache controller 212 of FIG. 2. Method 400 includes accessing a multi-copy scope directory state of a cache memory, as shown in block 402, which indicates the scope of sharing of cache lines within the cache memory system. Scope sharing may be checked at various levels, as depicted in the example of multiple scopes 300 of FIG. 3, such as cache lines 311, 313, 315, and 317 of FIG. 3. At block 404, the method 400 includes determining a sharing scope of a cache line in the cache memory system based on a multicopy scope directory state, where the multicopy scope directory state enumerates multiple scopes in the cache memory system. At block 406, the method 400 includes using the sharing scope to reduce the number of queries to one or more cache memories having a larger sharing scope than the sharing scope identified in the sharing scope. At block 408, the method 400 includes updating the multicopy scope directory state of the cache memory based on detecting a change in the sharing scope of the cache line in the cache memory system.
[0044] Additional processes and / or steps may be further included in method 400. It should be understood that the processes depicted in Figure 4 represent examples, and that other processes may be added, or existing processes may be removed, modified, or rearranged, without departing from the scope of the present disclosure.
[0045] 5 depicts a flow diagram of a method 500 for requesting a shared copy of a cache line in accordance with one or more embodiments of the present invention. Method 500 may be performed, for example, by cache controller 212 of FIG. 2.
[0046] When a requesting cache requests a shared copy of a line, such as the cache line in the L3 cache of core 0 in Figure 3, the current snoop scope is set to the minimum snoop scope in block 502. A query is sent to any remote caches in the current snoop scope that have not previously been queried in block 504. In block 506, a check is performed to determine whether an intervening master copy of the cache line was found in a remote cache of the current snoop scope or whether a cache line in a forwarding state was encountered.
[0047] If an intervening master copy is found in a remote cache within the current snoop scope or a cache line in the forwarding state is encountered in block 506, a shared copy of the cache line is installed in the request cache in block 508. The directory entry for the request copy of the cache line can be updated to be the intervening master copy, and the directory entry for the remote cache copy of the cache line can be updated as a shared non-master copy. In block 510, the state of the directory entry holding the master copy of the cache line can be updated to reflect a directory scope that is the larger of the directory scope found in the remote cache directory, or a minimum directory scope that includes both the request cache and the remote cache.
[0048] If an intervening master copy is not found within the current snoop scope at block 506, a check is performed at block 512 to determine whether this is the highest snoop scope. If so, the cache line is fetched from memory at block 514. Optionally, a response from cache memory in the system in response to the fetch may be awaited. If any cache is found in a shared non-master state, then at block 516 the request cache directory for the cache line may be updated to an intervening master shared with the requesting core and the lowest scope directory state that includes all found copies in a shared non-master state. Otherwise, the request cache directory may be updated to an exclusive state for the cache line. Thus, if the initial snoop scope did not find the cache line, and the cache line is found in the next snoop scope, the directory state is updated accordingly. The directory update may include updating at least one remote sourcing directory for the cache line.
[0049] If in block 512 this is not the highest snoop scope, then in block 518 the current snoop scope is set to the next largest snoop scope and flow returns to block 504 .
[0050] Figures 6A, 6B, and 6C collectively depict a flow diagram of a method 600 for requesting an exclusive copy of a cache line, according to one or more embodiments. Method 600 may be performed, for example, by cache controller 212 of Figure 2. While depicted as a combined process flow, embodiments of the invention may remove, subdivide, combine, or expand portions of method 600.
[0051] When a requesting cache requests an exclusive copy of the line, a check may be performed in block 602 to determine whether the requesting cache has an intervening master shared copy of the cache line. For example, the L3 cache of core 0 in FIG. 3 may be the requesting cache. If the requesting cache has an intervening master shared copy of the cache line, then invalidation of the cache line in all caches within a minimum snoop scope, including the directory scope, may be requested in block 604. A wait may be performed until all caches within the directory scope indicate that the cache line has been invalidated. In block 606, the directory for the cache line in the requesting core is updated to an exclusive state.
[0052] If the requesting cache does not have an intervening master shared copy of the cache line at block 602, then at block 608, the current snoop scope is set to the minimum snoop scope. At block 610, a query is sent to remote caches within the current snoop scope that have not previously been queried. For each remote cache in which a shared non-intervening master copy of the cache line is found within the current snoop scope, at block 612, the remote cache's directory may be set to an invalid state. If at block 614, an intervening master copy of the cache line is found in a remote cache within the current snoop scope, method 600 may perform one or more of blocks 616-622.
[0053] In block 616, the remote cache directory for the cache line is updated to an invalid state, and an exclusive copy of the cache line is installed in the request cache directory in an exclusive state based on finding a remote cache within the current snoop scope that has an exclusive copy of the cache line.
[0054] In block 618, the remote cache directory for the cache line may be updated to an invalid state, and the request cache directory for the cache line may be updated to an exclusive state based on finding a remote cache with an intervening master shared copy of the cache line within the current snoop scope, where the directory scope in the remote cache is less than or equal to the current snoop scope. For example, waiting until all caches within the scope indicated by the directory scope in the remote cache have been queried may be performed. Another possibility is waiting until all caches within the smallest directory scope that includes both the remote cache and the request cache have been queried.
[0055] In block 620, the remote cache directory for the cache line can be updated to an invalid state, and based on the directory scope in the remote cache being greater than the current snoop scope, the query can be sent to a remote cache within a minimum snoop scope greater than the directory scope in any remote cache that has not previously been queried.
[0056] In block 622, for each remote cache in which a shared non-interfering master copy of the cache line is found within a minimum snoop scope greater than the directory scope in the remote cache, the remote cache directory for the cache line is updated to an invalid state and the request cache directory for the cache line is updated to an exclusive state. A wait until all caches within the directory scope in the remote cache have been queried may be implemented.
[0057] If an intervening master copy is not found in a remote cache within the current snoop scope at block 614, then it is determined whether the maximum snoop scope has been reached at block 624. If the maximum snoop scope has been reached, then the request cache directory for the cache line is updated to an exclusive state at block 626. If the maximum snoop scope has not been reached, then the current snoop scope is set to the next larger snoop scope at block 628, and process flow returns to block 610. An intervening master copy of the cache line is available before the entire system state is resolved based on detecting that the primary scope does not contain an intervening master copy of the cache line, the secondary scope does contain an intervening master copy of the cache line, and a larger scope broadcast has been initiated.
[0058] Turning now to FIG. 7 , a computer system 700 according to one embodiment is generally illustrated. The computer system 700 can be an electronic computer framework comprising and / or employing any number of computing devices and networks, and combinations thereof, utilizing various communication technologies, as described herein. The computer system 700 can be easily scalable, extensible, and modular, allowing modifications to different services or reconfiguration of some features independently of others. The computer system 700 can be, for example, a server, desktop computer, laptop computer, tablet computer, or smartphone. In some examples, the computer system 700 can be a cloud computing node. The computer system 700 may be described in the general context of computer system-executable instructions, such as program modules, executed by a computer system. Generally, program modules may include routines, programs, objects, components, logic, data structures, etc. that perform particular tasks or implement particular abstract data types. Computer system 700 may also be practiced in a distributed cloud computing environment where tasks are performed by remote processing devices that are linked through a communications network. In a distributed cloud computing environment, program modules may be located in both local and remote computer system storage media, including memory storage devices.
[0059] As shown in FIG. 7, computer system 700 includes one or more central processing units (CPUs) 701a, 701b, 701c, etc. (collectively or generally referred to as processors 701). Processor 701 can be a single-core processor, a multi-core processor, a computing cluster, or any number of other configurations. Processor 701, also referred to as a processing circuit, is coupled to system memory 703 and various other components via a system bus 702. System memory 703 can include read-only memory (ROM) 704 and random access memory (RAM) 705. ROM 704, coupled to system bus 702, may include a basic input / output system (BIOS), which controls certain basic functions of computer system 700. RAM is read-write memory coupled to system bus 702 for use by processor 701. System memory 703 provides temporary memory space for the execution of instructions during operation. System memory 703 may include random access memory (RAM), read-only memory, flash memory, or any other suitable memory system.
[0060] Computer system 700 includes an input / output (I / O) adapter 706 and a communications adapter 707 coupled to a system bus 702. I / O adapter 706 may be a small computer system interface (SCSI) adapter that communicates with a hard disk 708 and / or any other similar components. I / O adapter 706 and hard disk 708 are collectively referred to herein as mass storage 710.
[0061] Software 711 for execution by computer system 700 may be stored on mass storage 710. Mass storage 710 is an example of a tangible storage medium readable by processor 701, and software 711 is stored as instructions for execution by processor 701 to operate computer system 700, as described later herein with reference to various figures. Examples of computer program products and the execution of such instructions are discussed in more detail herein. Communications adapter 707 interconnects system bus 702 with network 712, which may be an external network, enabling computer system 700 to communicate with other such systems. In one embodiment, system memory 703 and a portion of mass storage 710 collectively store an operating system for coordinating the functions of the various components shown in FIG. 7, which may be any suitable operating system, such as IBM Corporation's z / OS or AIX operating systems. z / OS and AIX are trademarks of IBM Corporation.
[0062] Additional input / output devices are shown connected to system bus 702 via display adapter 715 and interface adapter 716. In one embodiment, adapters 706, 707, 715, and 716 may be connected to one or more I / O buses connected to system bus 702 through intermediate bus bridges (not shown). A display 719 (e.g., a screen or display monitor) is connected to system bus 702 by display adapter 715, which may include a graphics controller and a video controller to improve performance of graphics-intensive applications. A keyboard 721, mouse 722, speaker 723, etc., may be interconnected to system bus 702 via interface adapter 716, which may include, for example, a super I / O chip that combines multiple device adapters into a single integrated circuit. Suitable I / O buses for connecting peripheral devices such as hard disk controllers, network adapters, and graphics adapters typically include common protocols such as Peripheral Component Interconnect (PCI). Thus, as configured in Figure 7, computer system 700 includes processing capabilities in the form of processor 701, storage capabilities including system memory 703 and mass storage 710, input means such as keyboard 721 and mouse 722, and output capabilities including speakers 723 and display 719.
[0063] In some embodiments, communications adapter 707 may transmit data using any suitable interface or protocol, such as an Internet Small Computer System Interface, among others. Network 712 may be a cellular network, a wireless network, a wide area network (WAN), a local area network (LAN), or the Internet, among others. External computing devices may connect to computer system 700 through network 712. In some examples, the external computing device may be an external web server or a cloud computing node.
[0064] It should be understood that the block diagram of Figure 7 is not intended to indicate that computer system 700 should include all of the components shown in Figure 7. Rather, computer system 700 may include any suitable fewer or additional components (e.g., additional memory components, embedded controllers, modules, additional network interfaces, etc.) not illustrated in Figure 7. Furthermore, the embodiments described herein with respect to computer system 700 may be implemented with any suitable logic, and logic as referred to herein may include any suitable hardware (e.g., a processor, embedded controller, or application specific integrated circuit, among others), software (e.g., an application, among others), firmware, or any suitable combination of hardware, software, and firmware, in various embodiments.
[0065] The technical effect may include improved computerized system performance. For example, by knowing the scope of sharing within a system through the use of multiple copy scoping bits, the system can dynamically determine how far a requestor must go to maintain cache coherency. This can reduce the resolution of event latency for shared-to-exclusive line conversions, thus improving the overall performance of these operations in the system. By avoiding a system-wide query and waiting for a system-wide response, bus usage and interface bandwidth can be improved, and controller busy time can be reduced when invalidating multiple copies of a cache line.
[0066] Various embodiments of the present invention are described herein with reference to the associated drawings. Alternate embodiments of the present invention may be devised without departing from the scope of the present invention. In the following description and drawings, various connections and relationships (e.g., above, below, adjacent, etc.) between elements are described. These connections and / or relationships may be direct or indirect unless otherwise specified, and the present invention is not intended to be limited in this respect. Thus, connections of entities may refer to direct or indirect connections, and relationships between entities may be direct or indirect relationships. Moreover, the various tasks and process steps described herein may be combined into a more comprehensive procedure or process having additional steps or functions not specifically described herein.
[0067] One or more of the methods described herein may be performed using any technology or combination of technologies, such as discrete logic circuits having logic gates for performing logic functions based on data signals, application specific integrated circuits (ASICs) having appropriate combinatorial logic gates, programmable gate arrays (PGAs), field programmable gate arrays (FPGAs), or the like, each of which is well known in the art.
[0068] For the sake of brevity, conventional technology relevant to the implementation and use of aspects of the present invention may or may not be described in detail herein. In particular, various aspects of computing systems and specific computer programs for implementing various technical features described herein are well known. Thus, for the sake of brevity, many conventional implementation details are only briefly mentioned herein or are omitted entirely without providing details of well-known systems and / or processes.
[0069] In some embodiments, various functions or acts may be performed at a given location, or in conjunction with the operation of one or more devices or systems, or both. In some embodiments, some of a given function or act may be performed at a first device or location, and remaining functions or acts may be performed at one or more additional devices or locations.
[0070] The terminology used herein is for the purpose of describing particular embodiments only and is not intended to be limiting. As used herein, the singular forms "a," "an," and "the" are intended to include the plural forms as well, unless the context clearly dictates otherwise. It will be further understood that the terms "comprises" and / or "comprising," when used herein, specify the presence of stated features, integers, steps, operations, elements, or components, or combinations thereof, but do not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components, or groups thereof, or combinations thereof.
[0071] Corresponding structures, materials, acts, and equivalents of all means or step and functional elements in the following claims are intended to include any structure, material, or act for performing a function in combination with other claim elements as specifically claimed. This disclosure has been presented for purposes of illustration and description, but is not intended to be exhaustive or limited to the precise form disclosed. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the disclosure. The embodiments were chosen and described in order to best explain the principles and practical application of the disclosure and to enable others skilled in the art to understand the disclosure of various embodiments with various modifications as suited to the particular uses envisioned.
[0072] The diagrams depicted herein are exemplary. There may be many variations to the diagrams or steps (or operations) described without departing from the spirit of the present disclosure. For example, actions may be performed in a different order, or actions may be added, deleted, or modified. Also, the term "coupled" refers to having a signal path between two elements, and does not imply a direct connection between elements with no intervening elements / connections between them. All of these variations are considered to be part of the present disclosure.
[0073] The following definitions and abbreviations will be used for interpreting the claims and this specification. As used herein, the terms "comprises," "comprising," "includes," "including," "has," "having," "contains," or "containing," or any other variation thereof, are intended to cover a non-exclusive inclusion. For example, a structure, mixture, process, method, article, or device that includes a list of elements is not necessarily limited to only those elements, but can include other elements not expressly listed or inherent in such structure, mixture, process, method, article, or device.
[0074] Additionally, the term "exemplary" is used herein to mean "serving as an example, instance, or illustration." Any embodiment or design described herein as "exemplary" is not necessarily to be construed as preferred or advantageous over other embodiments or designs. The terms "at least one" and "one or more" are understood to include any integer greater than or equal to one, i.e., 1, 2, 3, 4, etc. The term "plurality" is understood to include any integer greater than or equal to two, i.e., 2, 3, 4, 5, etc. The term "connected" can include both an indirect "connected" and a direct "connected."
[0075] The terms "about," "substantially," "approximately," and variations thereof are intended to include the degree of error associated with measurement of a particular quantity based on equipment available at the time of filing this application. For example, "about" can include a range of ±8% or 5%, or 2% of a given value.
[0076] The present invention may be a system, method, or computer program product, or combination thereof, at any possible level of technical detail of integration. The computer program product may include a computer-readable storage medium (or media) having computer-readable program instructions for causing a processor to carry out aspects of the present invention.
[0077] A computer-readable storage medium can be a tangible device capable of retaining and storing instructions for use by an instruction-execution device. A computer-readable storage medium may be, for example, but is not limited to, an electronic storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the foregoing. A non-exhaustive list of more specific examples of computer-readable storage media includes portable computer diskettes, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), static random access memory (SRAM), portable compact disk read-only memory (CD-ROM), digital versatile disk (DVD), memory sticks, floppy disks, mechanically encoded devices such as punch cards or ridge-in-groove structures having instructions recorded thereon, and any suitable combination of the foregoing. Computer-readable storage media as used herein should not be construed as being signals that are transitory in nature, such as radio waves or other freely propagating electromagnetic waves, electromagnetic waves propagating through a waveguide or other transmission medium (e.g., light pulses passing through fiber optic cable), or electrical signals transmitted through wires.
[0078] The computer-readable program instructions described herein can be downloaded from a computer-readable storage medium to each computing / processing device or to an external computer or external storage device over a network, such as the Internet, a local area network, a wide area network, or a wireless network, or a combination thereof. The network may include copper transmission cables, optical fiber transmissions, wireless transmissions, routers, firewalls, switches, gateway computers, or edge servers, or a combination thereof. A network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards the computer-readable program instructions for storage on a computer-readable storage medium within the respective computing / processing device.
[0079] Computer-readable program instructions for carrying out the operations of the present invention may be source or object code written in any combination of one or more programming languages, including assembler instructions, instruction set architecture (ISA) instructions, machine language instructions, machine-dependent instructions, microcode, firmware instructions, state setting data, configuration data for integrated circuits, or object code written in one or more programming languages, including object-oriented programming languages such as Smalltalk®, C++, or the like, and procedural programming languages such as the "C" programming language or similar programming languages. The computer-readable program instructions may execute entirely on the user's computer, partially on the user's computer, as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the latter scenario, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or the connection may be to an external computer (e.g., through the Internet using an Internet Service Provider). In some embodiments, electronic circuitry, including programmable logic circuits, field programmable gate arrays (FPGAs), or programmable logic arrays (PLAs), may execute computer readable program instructions by utilizing state information of the computer readable program instructions to individualize the electronic circuitry to implement aspects of the present invention.
[0080] Aspects of the present invention are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer-readable program instructions.
[0081] These computer-readable program instructions may be provided to a processor of a general-purpose computer, special-purpose computer, or other programmable data processing apparatus to produce a machine, the instructions of which execute on the processor of the computer or other programmable data processing apparatus to produce means for performing the functions / acts specified in one or more blocks of the flowcharts and / or block diagrams. These computer-readable program instructions may also be stored on a computer-readable storage medium such that the computer-readable storage medium comprises an article of manufacture containing instructions for performing aspects of the functions / acts specified in one or more blocks of the flowcharts and / or block diagrams, and can direct a computer, programmable data processing apparatus, or other device, or combination thereof, to function in a particular manner.
[0082] The computer-readable program instructions may also be loaded into a computer, other programmable data processing apparatus, or other device to perform a series of operational steps on the computer, other programmable apparatus, or other device to produce a computer-executed process, the instructions executing on the computer, other programmable apparatus, or other device to perform the functions / acts specified in one or more blocks of the flowcharts and / or block diagrams.
[0083] The flowcharts and block diagrams in the figures illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in the flowcharts or block diagrams may represent a module, segment, or portion of instructions, including one or more executable instructions for implementing specified logical functions. In some alternative implementations, the functions noted in the blocks may occur out of the order noted in the figures. For example, two blocks shown in succession may, in fact, be executed substantially concurrently, or the blocks may sometimes be executed in the reverse order, depending on the functionality involved. It should also be noted that each block in the block diagrams and / or flowchart diagrams, and combinations of blocks in the block diagrams and / or flowchart diagrams, can be implemented by a dedicated hardware-based system that performs the specified functions or acts or executes a combination of dedicated hardware and computer instructions.
[0084] The description of various embodiments of the present invention has been presented for purposes of illustration and is not intended to be exhaustive or limited to the disclosed embodiments. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described embodiments. The terminology used herein has been chosen to best explain the principles of the embodiments, their practical applications, or technical improvements over technologies found in the marketplace, or to enable others skilled in the art to understand the embodiments described herein.
Claims
1. 1. A computer-implemented method comprising: accessing a multicopy scope directory state of the cache memory indicating the scope of sharing of cache lines within the cache memory system; determining a scope of sharing of the cache line within the cache memory system based on the multicopy scope directory state, the multicopy scope directory state enumerating multiple scopes within the cache memory system; using the shared scope to reduce the number of queries to one or more cache memories having a scope greater than the shared scope identified in the shared scope; updating the multicopy scope directory state of the cache memory based on detecting a change in the sharing scope of the cache line within the cache memory system; 20. A computer-implemented method comprising:
2. 2. The computer-implemented method of claim 1, wherein the shared scope having a smaller scope is a subset of the shared scope having the larger scope in the cache memory system.
3. 3. The computer-implemented method of claim 2, wherein the scope of sharing includes a chip scope, a module scope, and a drawer scope as minimum scopes, and a system scope as maximum scope.
4. setting a snoop scope to a minimum scope based on determining that a requesting cache has requested a shared copy of the cache line; sending a query to one or more remote cache memories within the current snoop scope that have not previously been queried; determining whether an intervening master copy of the cache line was found in a remote cache memory of the current snoop scope or encountered the cache line in a forwarding state; installing a shared copy of the cache line in the request cache, and updating a state of a directory entry holding a master copy of the cache line based on the intervening master copy of the cache line being found in a remote cache memory of the current snoop scope or encountering the cache line in the forwarding state; The computer-implemented method of claim 1 further comprising:
5. fetching the cache line from the cache memory and updating a request cache directory based on finding one or more copies of the cache line in a shared non-master state after not finding the cache line within an initial snoop scope; updating at least one remote sourcing directory for said cache line; The computer-implemented method of claim 4 further comprising:
6. determining whether the requesting cache has an intervening master shared copy of the cache line based on the requesting cache requesting an exclusive copy of the cache line; requesting invalidation of the cache line in all cache memories within a minimum snoop scope that includes a directory scope; updating a request cache directory for said cache line in a request core to an exclusive state; The computer-implemented method of claim 1 further comprising:
7. setting a current snoop scope to the minimum snoop scope based on determining that the requesting cache does not have the intervening master's shared copy of the cache line; sending a query to one or more remote caches within the current snoop scope that have not previously been queried; updating a remote cache directory to an invalid state based on a shared non-intervention master copy of the cache line within the current snoop scope; updating the request cache directory for the cache line to the exclusive state based on determining that an intervening master copy of the cache line is not found in a remote cache within the current snoop scope and that a maximum snoop scope has been reached; setting the current snoop scope to the next larger snoop scope based on determining that the intervening master copy of the cache line is not found in the remote cache within the current snoop scope and the highest snoop scope has not been reached; The computer-implemented method of claim 6 further comprising:
8. updating the remote cache directory for the cache line to the invalid state and, based on finding the remote cache with the exclusive copy of the cache line within the current snoop scope and determining that the intervening master copy of the cache line is not found within the remote cache within the current snoop scope, installing the exclusive copy of the cache line in the exclusive state; updating the remote cache directory for the cache line to the invalid state, and updating the request cache directory for the cache line to the exclusive state based on finding the remote cache having the intervening master shared copy of the cache line within the current snoop scope, wherein the directory scope in the remote cache is less than or equal to the current snoop scope; The computer-implemented method of claim 7 further comprising:
9. updating the remote cache directory for the cache line to the invalid state, and based on the directory scope in the remote cache being greater than the current snoop scope, sending the query to the remote cache within the minimum snoop scope that is greater than the directory scope in the remote cache to which the query has not previously been sent; updating the remote cache directory for the cache line to the invalid state; updating the request cache directory for the cache line to the exclusive state; The computer-implemented method of claim 8 further comprising:
10. using the intervening master copy of the cache line before a full system state is resolved based on detecting that a primary scope does not include the intervening master copy of the cache line, a secondary scope does include the intervening master copy of the cache line, and a larger scope broadcast has been initiated; The computer-implemented method of claim 7 further comprising:
11. 1. A system comprising: a plurality of processors; a cache memory system; a cache controller, accessing a multicopy scope directory state of a cache memory indicating the scope of sharing of cache lines within said cache memory system; determining a scope of sharing of the cache line within the cache memory system based on the multicopy scope directory state, the multicopy scope directory state enumerating multiple scopes within the cache memory system; using the shared scope to reduce the number of queries to one or more cache memories having a scope greater than the shared scope identified in the shared scope; and updating the multicopy scope directory state of the cache memory based on detecting a change in the sharing scope of the cache line within the cache memory system; a cache controller configured to perform operations including: A system comprising:
12. 12. The system of claim 11, wherein the scope of the shares having a smaller scope is a subset of the scope of the shares having the larger scope in the cache memory system.
13. 13. The system of claim 12, wherein the scope of sharing includes a chip scope, a module scope, and a drawer scope as minimum scopes, and a system scope as maximum scope.
14. The cache controller: setting a snoop scope to a minimum scope based on determining that a requesting cache has requested a shared copy of the cache line; sending a query to one or more remote cache memories within the current snoop scope that have not previously been queried; determining whether an intervening master copy of the cache line was found in a remote cache memory of the current snoop scope or encountered the cache line in a forwarding state; installing a shared copy of the cache line in the request cache, and updating a state of a directory entry holding a master copy of the cache line based on the intervening master copy of the cache line being found in a remote cache memory of the current snoop scope or encountering the cache line in the forwarding state; The system of claim 11 configured to perform operations including:
15. The cache controller: fetching the cache line from the cache memory and updating a request cache directory based on finding one or more copies of the cache line in a shared non-master state after not finding the cache line within an initial snoop scope; updating at least one remote sourcing directory for said cache line; The system of claim 14 configured to perform operations including:
16. The cache controller: determining whether the requesting cache has an intervening master shared copy of the cache line based on the requesting cache requesting an exclusive copy of the cache line; requesting invalidation of the cache line in all cache memories within a minimum snoop scope that includes a directory scope; updating a request cache directory for said cache line in a request core to an exclusive state; The system of claim 11 configured to perform operations including:
17. The cache controller: setting a current snoop scope to the minimum snoop scope based on determining that the requesting cache does not have the intervening master's shared copy of the cache line; sending a query to one or more remote caches within the current snoop scope that have not previously been queried; updating a remote cache directory to an invalid state based on a shared non-intervention master copy of the cache line within the current snoop scope; updating the request cache directory for the cache line to the exclusive state based on determining that an intervening master copy of the cache line is not found in a remote cache within the current snoop scope and that a maximum snoop scope has been reached; setting the current snoop scope to the next larger snoop scope based on determining that the intervening master copy of the cache line is not found in the remote cache within the current snoop scope and the highest snoop scope has not been reached; 17. The system of claim 16, configured to perform operations including:
18. The cache controller: updating the remote cache directory for the cache line to the invalid state and, based on finding the remote cache with the exclusive copy of the cache line within the current snoop scope and determining that the intervening master copy of the cache line is not found within the remote cache within the current snoop scope, installing the exclusive copy of the cache line in the exclusive state; updating the remote cache directory for the cache line to the invalid state, and updating the request cache directory for the cache line to the exclusive state based on finding the remote cache having the intervening master shared copy of the cache line within the current snoop scope, wherein the directory scope in the remote cache is less than or equal to the current snoop scope; 20. The system of claim 17 configured to perform operations including:
19. The cache controller: updating the remote cache directory for the cache line to the invalid state, and based on the directory scope in the remote cache being greater than the current snoop scope, sending the query to the remote cache within the minimum snoop scope that is greater than the directory scope in the remote cache to which the query has not previously been sent; updating the remote cache directory for the cache line to the invalid state; updating the request cache directory for the cache line to the exclusive state; 20. The system of claim 18 configured to perform operations including:
20. A computer program causing one or more processors to perform a computer-implemented method according to any one of claims 1 to 10.
Citation Information
Patent Citations
Parallel computer
JP2007207151A
Methods, devices, and computer program products for making memory requests (efficient storage of metabits in system memory)
JP2015503130A