Cache snooping mode that extends coherence protection for a specific request

By employing a referee mode in cache memory systems to protect cache lines during flush or clean operations, the system ensures timely and efficient cache coherency by preventing competing accesses, addressing delays in multiprocessor systems.

JP7710447B2Active Publication Date: 2025-07-18INTERNATIONAL BUSINESS MACHINE CORPORATION
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
JP2022535608
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2019-12-17
Filing Date
2020-12-14
Publication Date
2025-07-18
Estimated Expiration
2040-12-14

AI Technical Summary

Technical Problem

In multiprocessor computer systems, cache coherency protocols face delays due to the generation of new cache line copies during flush or clean operations, leading to repeated retries and incomplete invalidation of cache lines, which can be exacerbated by competing access requests.

Method used

Implementing a cache memory system with a designated coherence member that provides protection for target cache lines through a referee mode, ensuring coherence ownership is maintained until a second flush/clean operation is completed, preventing competing accesses and eliminating delays.

Benefits of technology

This approach reduces the latency and completion time of flush or clean operations by preventing other caches from obtaining coherence ownership of the target cache line, thereby ensuring timely and efficient cache coherency maintenance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007710447000004
    Figure 0007710447000004
  • Figure 0007710447000005
    Figure 0007710447000005
  • Figure 0007710447000006
    Figure 0007710447000006
Patent Text Reader

Abstract

The cache memory includes a data array, a directory of the contents of the data array specifying coherence state information, and snooping logic that processes operations snooped from the system fabric by referencing the data array and the directory. In response to snooping a request on the system fabric for a first flush / clean memory access operation specifying the target address, the snooping logic determines whether the cache memory has coherence ownership of the target address. Based on a determination that the cache memory has coherence ownership of the target address, the snooping logic initiates a referee mode after servicing the request. In the referee mode, the snooping logic protects the memory block identified by the target address against conflicting memory access requests by multiple processor cores until the end of a second flush / clean memory access operation specifying the target address.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention generally relates to data processing, and more particularly to a cache snooping mode that extends coherence protection for flash / clean memory access requests that update system memory.

Background Art

[0002] Conventional multiprocessor (MP) computer systems (such as server computer systems) generally have multiple processing units all connected to a system interconnect that typically includes one or more address, data, and control buses. The system interconnect represents the lowest level of shared memory in the multiprocessor computer system and is coupled with system memory that is generally accessible for read and write access by all processing units. To reduce access latency to instructions and data present in the system memory, each processing unit is typically further supported by its own multi-level cache hierarchy, and the lower level cache hierarchy can be shared by one or more processor cores.

[0003] Cache memory is generally used to temporarily buffer memory blocks that the processor is likely to access, reducing access latency by loading the necessary data and instructions from system memory and speeding up processing. In some MP systems, the cache hierarchy includes at least two levels. The level 1 (L1) or upper-level cache is typically a private cache associated with a specific processor core and is not accessible to other cores in the MP system. Usually, in response to a memory access instruction such as a load instruction or a store instruction, the processor core first accesses the directory of the upper-level cache. If the requested memory block is not found in the upper-level cache, the processor core then accesses the lower-level cache (e.g., level 2 (L2) or level 3 (L3) cache) to find the requested memory block. The lowest-level cache (e.g., L2 or L3) can be shared by multiple processor cores.

[0004] Multiple processor cores may request write access to the same data cache line, and since the modified cache line is not immediately synchronized with the system memory, the cache hierarchy of a multiprocessor computer system typically implements a cache coherency protocol to ensure at least lowest-level coherence among the various "views" of the processor cores regarding the contents of the system memory. In particular, in cache coherency, it is minimally necessary to prevent a hardware thread from accessing the old copy of a memory block after accessing a copy of the memory block and then accessing an updated copy of the memory block.

[0005] Some MP systems support flush and clean operations, which copy the modified cache line associated with the address targeted by the flush or clean operation from the unique cache hierarchy that includes cache lines in the coherence state indicating write permission (which may be referred to herein as the "HPC" state), and return it to system memory if it exists. In the case of a clean operation, the target cache line also transitions to the unmodified HPC coherence state. In the case of a flush operation, the target cache line in the HPC state transitions to the invalid coherence state (regardless of whether there are modifications and if it exists). Additionally, in a flush operation, it is necessary to invalidate any one or more other copies of the target cache line in the non-HPC state in all cache hierarchies of the MP system. This invalidation may not be completed even when the cache having the target cache line in the HPC state (if it exists) has completed its processing.

[0006] In an MP system that maintains coherence by means of a snooping-based coherence protocol, generally, on the system interconnect of the MP system, a flush or clean operation is broadcast, and a cache having a target cache line in the HPC state receives a Retry coherence response until it has completed processing the flush or clean operation. For this reason, a coherence participant that initiates a flush or clean operation may be required to reissue the flush or clean operation multiple times before a cache having a target cache line in the HPC state (if any) completes processing the flush or clean operation. If the cache that had a target cache line in the HPC state has completed processing the flush or clean operation and a new copy of the target cache line in the HPC state (modified HPC state in the case of a clean operation, modified or unmodified HPC state in the case of a flush operation) has not been generated, in subsequent issuances of the clean operation, a coherence response indicating success is received, and in subsequent issuances of the flush operation, a coherence response indicating success (if there is no caching copy of the line) or a coherence response that transfers to the initiating coherence participant the authority to invalidate any other non-HPC caching copy of the target cache line is received. In any of these cases regarding the flush operation, the flush operation is considered "successful" in the sense of complete or future completion of the flush operation when one or more other non-HPC copies of the target cache line are invalidated by the initiating coherence participant (e.g., by issuing a kill operation). However, if, prior to subsequent issuances of the clean or flush operation, another coherence participant generates a new copy of the target cache line in the relevant HPC state (i.e., modified HPC state in the case of a clean operation, modified or unmodified HPC state in the case of a flush operation), subsequent reissuances of the flush or clean operation are retried, and since processing of the new copy of the target cache line in the HPC state is required, successful completion of the flush or clean operation is delayed.This delay can be exacerbated by the continued generation of new HPC copies of the cache lines targeted for the flush or clean operation. SUMMARY OF THE INVENTION

[0007] In at least one embodiment, the target cache line for the flush or clean operation is protected from competing accesses by other coherence members through a designated coherence member that provides protection for the target cache line.

[0008] In at least one embodiment, the cache memory comprises a data array, a directory of the contents of the data array that specifies coherence state information, and snooping logic that processes operations snooped from the system fabric by reference to the data array and the directory. The snooping logic determines whether the cache memory has coherence ownership of the target address in response to snooping, on the system fabric, a request for a first flush / clean memory access operation that specifies the target address. Based on the determination that the cache memory has coherence ownership of the target address, the snooping logic initiates referee mode after corresponding to the request. In referee mode, the snooping logic protects the memory block identified by the target address against collision memory access requests by a plurality of processor cores until a second flush / clean memory access operation that specifies the same target address is snooped, approved, and successfully completed. As a result, other caches are prevented from obtaining coherence ownership of the target cache line between the first and second flush / clean memory access operations, eliminating a potential cause of delay in the completion of the requested flush or clean.

[0009] The embodiments of the present invention will be described below with reference to the accompanying drawings, which are merely examples.

Brief Description of the Drawings

[0010]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Modes for Carrying Out the Invention

[0011] Reference is now made to the drawings, in which like reference numerals refer to like corresponding parts throughout. In particular, referring to FIG. 1, this figure is a high-level block diagram showing an exemplary data processing system 100 according to one embodiment. In the illustrated embodiment, the data processing system 100 is a cache-coherent multiprocessor (MP) data processing system including a plurality of processing nodes 102 for processing data and instructions. The processing nodes 102 are coupled to a system interconnect 110 that conveys address, data, and control information. The system interconnect 110 may be implemented, for example, as a bus-style interconnect, a switch-style interconnect, or a hybrid interconnect.

[0012] In the illustrated embodiment, each processing node 102 is implemented as a multi-chip module (MCM) preferably including four processing units 104 each realized as an integrated circuit. The processing units 104 within each processing node 102 are coupled to communicate with each other and with the system interconnect 110 by a local interconnect 114 that may be realized, similar to the system interconnect 110, by, for example, one or more buses or switches or both. The system interconnect 110 and the local interconnect 114 together form a system fabric.

[0013] As will be described in more detail below with reference to FIG. 2, each of the processing units 104 includes a memory controller 106 coupled to the local interconnect 114 to provide an interface to each system memory 108. Data and instructions present in the system memory 108 can generally be accessed, cached, and modified by a processor core in any processing unit 104 of any processing node 102 within the data processing system 100. Thus, in the distributed shared memory system of the data processing system 100, the system memory 108 constitutes the lowest level of memory storage. In an alternative embodiment, one or more memory controllers 106 (and system memories 108) can be coupled to the system interconnect 110 instead of the local interconnect 114.

[0014] Those skilled in the art will of course understand that the MP data processing system 100 of FIG. 1 can include many additional components not shown, such as interconnect bridges, non-volatile storage, network connection ports, or peripherals. Since these additional components are not necessary for an understanding of the described embodiments, they are neither shown in FIG. 1 nor discussed in detail below. However, it is understood that the improvements described herein are applicable to data processing systems of various architectures and are not limited to the general data processing system architecture shown in FIG. 1.

[0015] Referring now to FIG. 2, this figure is a more detailed block diagram of an exemplary processing unit 104 according to one embodiment. In the illustrated embodiment, each processing unit 104 is an integrated circuit comprising a plurality of processor cores 200 that process instructions and data. Each processor core 200 comprises one or more execution units that execute instructions, such as an LSU 202 that executes memory access instructions that request access to or generate requests for access to memory blocks. In at least some embodiments, each processor core 200 is capable of simultaneously and independently executing a plurality of hardware execution threads.

[0016] The operation of each processor core 200 is supported by a multi-level memory hierarchy having a shared system memory 108 accessed via an integrated memory controller 106 at the lowest level. At its upper levels, the memory hierarchy includes one or more cache memory levels. In the illustrated embodiment, these include a store-through level 1 (L1) cache 226 dedicated to each processor core 200 and a store-in level 2 (L2) cache 230 for each processor core 200. In some embodiments, each L2 cache 230 is implemented by a plurality of L2 cache slices that each handle memory access requests for a respective set of physical memory addresses, in order to efficiently handle multiple simultaneous memory access requests to cacheable addresses.

[0017] Although the illustrated cache hierarchy includes only two cache levels, those skilled in the art will appreciate that alternative embodiments may include additional levels of on-chip or off-chip in-line or lookaside caches (e.g., L3, L4, etc.), which may fully contain, partially contain, or not contain the contents of the upper cache levels.

[0018] Continuing to refer to FIG. 2, each processing unit 104 further includes an integrated distributed fabric controller 216 that participates in controlling the operation flow on the local interconnect 114 and the system interconnect 110, and response logic 218 that is utilized in a selected cache coherence protocol to determine a coherence response to a memory access request.

[0019] During operation, when a hardware thread being executed by the processor core 200 includes a memory access instruction that requests execution of a specified memory access operation, the LSU 202 executes this memory access instruction to determine the target physical address to be accessed. If the requested memory access cannot be fully executed by referring to the L1 cache 226 of the executing processor core 200, the processor core 200 generates a memory access request including at least, for example, the request type and the target physical address, and issues it to the associated L2 cache 230 for processing.

[0020] Referring now to FIG. 3, which is a more detailed block diagram of an exemplary embodiment of an L2 cache 230 according to one embodiment. As shown in FIG. 3, the L2 cache 230 includes a cache array 302 and a directory 308 of the contents of the cache array 302. Assuming that the cache array 302 and the directory 308 are set associative as in the prior art, a given index bit within the system memory (physical) address is utilized to map a memory location in the system memory 108 to a particular match class within the cache array 302. A particular memory block stored within a cache line of the cache array 302 is recorded in a cache directory 308 that includes one directory entry per cache line. Although not shown explicitly in FIG. 3, those skilled in the art will appreciate that each directory entry in the cache directory 308 includes various fields (e.g., a tag field that identifies the physical address of the memory block held in the corresponding cache line of the cache array 302, a state field that indicates the coherence state of the cache line, an LRU (Least Recently Used) field that indicates the replacement order of the cache line with respect to other cache lines in the same match class, and an inclusion field that indicates whether the memory block is held in the associated L1 cache 226).

[0021] The L2 cache 230 includes a plurality (e.g., 16) of Read-Claim (RC) machines 312a to 312n that independently and simultaneously process memory access requests received from the associated processor core 200. Also, to handle remote memory access requests from processor cores 200 other than the associated processor core 200, the L2 cache 230 includes a plurality of snooping (SN) machines 311a to 311m. Each SN machine 311 can independently and simultaneously handle the "snooped" remote memory access requests from the local interconnect 114. Naturally, for the RC machine 312 to respond to a memory access request, it may be necessary to replace or invalidate a memory block in the cache array 302. Therefore, the L2 cache 230 includes a CO (castout) machine 310 that manages the removal and write-back of memory blocks from the cache array 302.

[0022] The L2 cache 230 further includes an arbiter 305 that controls multiplexers M1, M2 to order the processing of local memory access requests received from the associated processor core 200 and remote requests snooped on the local interconnect 114. The memory access requests are transferred to dispatch logic 306 according to an arbitration policy implemented by the arbiter 305, and the dispatch logic processes the memory access requests regarding the directory 308 and the cache array 302 over a given number of cycles.

[0023] Also, the L2 cache 230 includes an RC queue (RCQ) 320 and a CPI (castout push intervention) queue 318 that respectively buffer data inserted into the cache array 302 and data removed from the cache array 302. The RCQ 320 includes a number of buffer entries that individually correspond to a particular one of the RC machines 312 such that each dispatched RC machine 312 reads data only from a designated buffer entry. Similarly, the CPI queue 318 includes a number of buffer entries that individually correspond to one of the CO machines 310 and the snooper 311 such that each dispatched CO machine 310 and each snooper 311 reads data only from a designated CPI buffer entry.

[0024] Also, each RC machine 312 is assigned one of a plurality of RC data (RCDAT) buffers 322 that buffer a memory block read from the cache array 302, a memory block received from the local interconnect 114 via the reload bus 323, or both. The RCDAT buffer 322 assigned to each RC machine 312 is preferably configured with connections and functions corresponding to memory access requests that the associated RC machine 312 can handle. Associated with the RCDAT buffer 322 is a store data multiplexer M4 that selects data bytes from an input and buffers them in the RCDAT buffer 322 in response to a selection signal (not shown) generated by the arbiter 305.

[0025] During operation, a processor store request including a request type (ttype), a target physical address, and store data is received from the associated processor core 200 in the store queue (STQ) 304. The store data is transmitted from the STQ 304 to the store data multiplexer M4 via the data path 324, and the store type and the target address are passed to the multiplexer M1. Also, the multiplexer M1 receives, as an input, a processor load request from the processor core 200 and a directory write request from the RC machine 312. In response to a selection signal (not shown) generated by the arbiter 305, the multiplexer M1 selects one of its input requests and transfers it to the multiplexer M2. The multiplexer M2 further receives, as an input, a remote request received from the local interconnect 114 via the remote request path 326. The arbiter 305 schedules local and remote memory access requests to be processed and generates a series of selection signals 328 based on this scheduling. In response to the selection signal 328 generated by the arbiter 305, the multiplexer M2 selects, as the next memory access request to be processed, the local request received from the multiplexer M1 or the remote request snooped from the local interconnect 114.

[0026] Referring now to FIG. 4, this figure is a spatiotemporal diagram of exemplary operations on the system fabric of the data processing system 100 of FIG. 1. It is understood that many of such operations are executed on the system fabric at any given time, and depending on the operation scenario, a plurality of these concurrent operations may specify colliding target addresses.

[0027] The operation starts at a request stage 450 where a master 400 (e.g., the RC machine 312 of the L2 cache 230) issues a request 402 on the system fabric. The request 402 preferably includes at least a request type indicating the type of the desired access and a resource identifier (e.g., a physical address) indicating the resource to which the request accesses. The requests are preferably those shown in Table I below.

[0028]

Table 1

[0029] The request 402 is received by snooping units 404a - 404n (e.g., SN machines 311a - 311m of the L2 cache 230 and snooping units (not shown) in the memory controller 106) distributed in the data processing system 100. In the case of a read type request, the SN machine 311 in the same L2 cache 230 as the master 400 of the request 402 does not perform snooping of the request 402 (i.e., generally there is no self - snooping). This is because the request 402 is transmitted on the system fabric only when the processing unit 104 cannot internally handle the read type request 402. However, in the case of other types of requests 402, such as flash / clean requests (e.g., DCBF, DCBST, AMO requests described in Table I), the SN machine 311 in the same L2 cache 230 as the master 400 of the request 402 performs self - snooping of the request 402.

[0030] The operation continues in the partial response stage 455. In the partial response stage 455, each of the snooping units 404 that receive and process the requests 402 provides a partial response ("Presp") 406 that represents at least the response to the request 402 of that snooping unit 404. The snooping unit 404 within the integrated memory controller 106 determines the partial response 406 to provide based on, for example, whether the snooping unit 404 is involved in the request address and whether there are resources currently available for responding to the request. The snooping unit 404 of the L2 cache 230 may determine its partial response 406 based on, for example, the availability of its L2 cache directory 308, the availability of the snooping logic instance 311 within the snooping unit 404 that handles the request, and the coherence state (if present) associated with the request address in the L2 cache directory 308.

[0031] The operation continues at the combined response stage 460. At the combined response stage 460, partial responses 406 of the snooper 404 are logically combined, either sequentially or simultaneously, by one or more instances of the response logic 218 to determine the overall system combined response ("Cresp") 410 to the request 402. In a preferred embodiment assumed hereinafter in this specification, the instance of the response logic 218 involved in generating the combined response 410 is located in the processing unit 400 including the master 400 that issued the request 402. The response logic 218 provides the combined response 410 to the master 400 and the snooper 404 via the system fabric, indicating the overall system response (e.g., Success, Retry, etc.) to the request 402. Cresp 410, when indicating success of the request 402, may indicate, for example, the data source (if necessary) for the requested memory block, the coherence state (if necessary) in which the requested memory block is cached by the master 400, and whether an "invalidate" operation to invalidate the cached copy of the requested memory block in one or more L2 caches 230 is necessary (if necessary).

[0032] In response to receiving the combined response 410, one or more of the master 400 and the snooper 404 typically perform one or more operations corresponding to the request 402. These operations may include supplying data to the master 400, invalidating or updating the coherence state of the data cached in one or more L2 caches 230, performing a castout operation, writing data back to the system memory 108, and the like. If required by the request 402, the requested or target memory block may be transmitted to one of the master 400 or the snooper 404 before or after the generation of the combined response 410 by the response logic 218, or the requested or target memory block may be transmitted from one of the master 400 or the snooper 404.

[0033] In the following description, regarding whether snooper 404 is the highest coherence point (HPC), the lowest coherence point (LPC), or neither with respect to the request address specified by request 402, the partial response 406 of snooper 404 to request 402 and the operations performed by snooper 404 in response to request 402 or its combined response 410 or both will be described. In this specification, the LPC is defined as a memory device or I / O device that functions as the final repository of a memory block. If there is no caching member having a copy of the memory block, the LPC has only an image of the memory block. If there is no HPC caching member for the memory block, the LPC has only the authority to permit or deny change requests for the memory block. Also, the LPC provides data for read requests or change requests for the memory block when the LPC data is up-to-date and there is no caching member capable of providing the data. If the caching member has a newer copy of the data and cannot provide it for the request, the request is retried because the LPC does not provide status data. For a representative request in the embodiment of the data processing system 100 shown in FIGS. 1-3, the LPC becomes the memory controller 106 for the system memory 108 having the reference memory block.

[0034] In this specification, the true image of a memory block (which may or may not match the corresponding memory block in the LPC) is cached, and the HPC is defined as a unique identification device that has the authority to permit or reject a change request for the memory block. Also, for the sake of explanation, the HPC (even if its copy matches the main memory behind the LPC) responds to any read request or change request for the memory block by providing a copy of the memory block to the requesting side (the cache-to-cache transfer is faster than the LPC-cache transfer). Therefore, in the case of a typical request in an embodiment of the data processing system, if the HPC exists, it becomes the L2 cache 230. Other indicators can also be used to specify the HPC of the memory block, but in a preferred embodiment, the selected cache coherence state in the L2 cache directory 308 of the L2 cache 230 is used to specify the HPC of the memory block (if it exists). In a preferred embodiment, the coherence state in the coherence protocol provides (1) an indicator of whether the cache is the HPC of the memory block, (2) whether the cached copy is unique (i.e., the only cached copy in the entire system), (3) for the phase of operation, whether the cache can provide a copy of the memory block to the master of the memory block request and at what timing, and (4) whether the cached image of the memory block matches the corresponding memory block in the LPC (system memory). These four attributes can be expressed, for example, as an exemplary variation of the well-known MESI (Modified, Exclusive, Shared, Invalid) protocol and are summarized in Table II below. For other information regarding the coherence protocol, see, for example, U.S. Patent No. 7,389,388, which is incorporated herein by reference.

[0035]

Table 2

[0036] What should be noted in Table II above is T, Te, S L each of which is a "shared" coherence state in that another cache memory may simultaneously have a copy of the cache line held in one of these states by the cache memory. State T or Te identifies the HPC cache memory that has a cache line associated with one of states M and Me and supplies a query-limited copy of the associated cache line to another cache memory. A cache memory having a cache line in coherence state T or Te has, as HPC, the authority to change the cache line or to give such authority to another cache memory. A cache memory having a cache line in state Tx (e.g., T or Te) functions as the last resort cache data source for a query-limited copy of the cache line only when a cache memory having a cache line in state S L cannot be used as the (pre-Cresp) data source such that a cache memory having a cache line in state S functions as the (post-Cresp) data source for a query-limited copy of the cache line by supplying the query-limited copy to another cache memory.

[0037] State S L is formed in a cache memory in response to the cache memory receiving a query-limited copy of a cache line from a cache memory in coherence state T. State S L is not an HPC coherence state, but a cache memory having a cache line in state S L can supply a query-limited copy of the cache line to another cache memory and can execute this prior to receiving Cresp. In response to supplying a query-limited copy of a cache line to another cache memory (assuming state S L ), the cache memory supplying the query-limited copy of the cache line updates the coherence state of the cache line from S L to S. Thus, coherence state S LWith this implementation, since many query-limited copies of cache lines with high-frequency queries are generated throughout the multiprocessor data processing system, it is convenient that the latency of query-limited access to these cache lines is reduced.

[0038] Referring to FIG. 4 again, the HPC of the memory block (if present) referenced in request 402, or the LPC of the memory block if no HPC exists, preferably assumes the protection of the transfer of the coherence ownership of the memory block in response to request 402 as needed. In the exemplary scenario shown in FIG. 4, snoopers 404n in the HPC (or LPC if no HPC) of the memory block specified by the request address of request 402 protect the transfer of the coherence ownership of the requested memory block to master 400 in protection window 412a that extends from the time snooper 404n determines its partial response 406 until snooper 304n receives the combined response 410, and in subsequent window extension 412b that extends a programmable time beyond the receipt of the combined response 410 by snooper 404n. During protection window 412a and window extension 412b, snooper 404n protects the transfer of ownership by providing partial responses 406 (e.g., Retry partial responses) to other requests specifying the same request address that prevent the acquisition of ownership by other masters until the ownership is successfully transferred to master 400. Master 400 similarly starts protection window 413 to protect the coherence ownership of the memory block requested in request 402 after receiving the combined response 410.

[0039] Since all snooping agents 404 have limited resources to handle the above-described CPU and I / O requests, multiple different levels of Presp and corresponding Cresp are possible. For example, a snooping agent within memory controller 106 that is involved with a requested memory block may respond with a partial response indicating that it can function as an LPC for the request if it has a queue available for handling requests. On the other hand, if the snooping agent does not have a queue available for handling requests, it may respond with a partial response indicating that while it is the LPC for the memory block, it cannot currently respond to the request. Similarly, snooping agent 311 in L2 cache 230 may need an available instance of snooping logic and access to L2 cache directory 406 to handle a request. If access to one (or both) of these resources is not available, the response will be a partial response (and corresponding Cresp) indicating that the request cannot be serviced due to the lack of required resources.

[0040] As described above, in a system implementing a snooping-based coherence protocol, flush operations (e.g., DCBF and AMO) and clean operations (e.g., DCBST) may face concerns about ordering in the vulnerability window between a flush / clean operation that preferably terminates in the cache hierarchy including the target cache line in the HPC state (if present) and a master of a flush / clean request that initiates a successful final flush / clean request. As will be described in detail below with reference to FIGS. 5-8, these concerns about ordering can be addressed by a coherence member (i.e., HPC) having the coherence ownership of the target cache line extending the protection window of the target cache line.

[0041] Referring now to FIG. 5, this figure is a high-level logic flowchart of an exemplary process in which master 400 (e.g., RC machine 312) of processing unit 104 performs a flash or clean-type memory access operation according to one embodiment. As described above, any number of masters may each perform their respective flash / clean memory access operations simultaneously on target addresses that may collide. As a result, in data processing system 100, multiple instances of the process shown in FIG. 5 may be executed overlapping in time.

[0042] After starting at block 500, the process of FIG. 5 proceeds to block 502, which indicates that master 400 issues a memory access operation request 402 on the system fabric of data processing system 100. In at least some embodiments, master 400, such as RC machine 312 of L2 cache 230, issues request 402 in response to receiving a memory access request from a related processor core 200 based on the execution of a corresponding instruction by LSU 202. In the above embodiment, request 402 initiates a memory access operation belonging to one of a plurality of classes or types of operations (collectively referred to herein as flash / clean (FC) operations). These FC operations include the operations DCBF, DCBST, and AMO referenced in Table I, all of which are storage modification operations that need to write any modified caching copies of the target memory block back to the associated system memory 108.

[0043] As is apparent from the foregoing description with respect to FIG. 4, the FC request 402 of the master 400 is received on the system fabric by the L2 caches 230 and the memory controller 106 distributed within the data processing system 100. In response to the reception of the FC request 402, these various snoops 404 generate respective partial responses 406 and transmit them to the relevant instances of the response logic 218. In an exemplary embodiment, the L2 cache 230 responds to the snooping of the FC request 402 with one of three Presp's: (1) heavyweight Retry, (2) lightweight Retry, or (3) Null. The heavyweight Retry Presp is provided by an L2 cache 230 that cannot access the coherence state of the target address of the FC request 402 in the directory 308 at the present time. Also, the heavyweight Retry Presp is provided by an L2 cache 230 that, although the coherence state in the directory 308 designates the target address as an HPC, cannot respond to the FC request 402 at this time or is busy processing a request for the target cache line at the present time.

[0044] The lightweight Retry Presp is provided by an L2 cache 230 where the directory 308 is accessible and indicates either state S L and S, and (1) another collision request for the target address is being processed at the present time, (2) the SN machine 311 cannot be dispatched and the SN machine 311 is not active for the target address at the present time, or (3) the SN machine 311 has been dispatched and the FC request 402 has been processed.

[0045] Of course, the L2 cache 230 will provide a heavy Retry Presp or a lightweight Retry Presp based on the coherence state associated with the target address when the request was first snooped, during the entire period in which the active SN machine 311 that processes requests for a given target address exists. For the specific operations performed by the L2 cache 230 that has the target address of the FC request 402 in the dirty HPC coherence state (e.g., M or T), it will be described in detail below with reference to FIG. 6. For the operations performed by the L2 cache 230 that has the target address of the flash request in the shared coherence state (e.g., Te, S L , or S), it will be described in detail below with reference to FIG. 8.

[0046] In response to the snooping of the FC request 402, the memory controller 106 that is not involved in the target address (i.e., not an LPC) does not provide a Presp (or provides a Null Presp). The memory controller 106, which is the LPC memory controller for the target address of the FC request 402, provides a Retry_LPC Presp if it cannot respond to the FC request 402 due to resource constraints or because it has already responded to another memory access request specifying the same address. The LPC memory controller 106 protects the target address of the FC request 402 by providing an LPC_Retry Presp to any subsequent FC operation (or other operation) until it receives a successful Cresp for the initially snooped FC request 402 when it is capable of responding to the FC request 402.

[0047] Referring again to FIG. 5, this process proceeds from block 502 to blocks 504, 505, and 506, which indicate that master 400 of request 402 for the FC operation issued at block 502 waits for the reception of the corresponding Cresp410 from response logic 218. In at least one embodiment, response logic 218 can generate Cresp410 for the FC operation based on the received Presp of snooper 404, as shown in Table III below. In at least some embodiments, the DCBST request ignores caches having target cache lines in either state S L and S, and thus does not receive the lightweight Retry Presp (as shown in rows 1, 2, 5, and 6 of Table III).

[0048] [Table 3]

[0049] In response to the reception of Cresp410, the master 400 determines whether Cresp410 indicates Retry as shown in the first three lines of Table III (block 504). If it indicates Retry, the process returns to block 502, which indicates that the master 400 reissues the request 402 for the FC operation on the system fabric. If the master 400 determines in block 504 that Cresp410 of the request 402 is not Retry, in the embodiment of Table III, the coherence result is either HPC_Success (positive determination in block 505 (the fourth line of Table III)), Success_with_cleanup (Success_CU) (positive determination in block 506 (the fifth and sixth lines of Table III)), or Success (negative determination in both blocks 505 and 506 (the seventh line of Table III)). If Cresp410 is Success (as indicated by the negative determination in both blocks 505 and 506), there is no caching copy of the target cache line in the data processing system 100. As a result, the flush or clean operation is successfully completed, and the process of FIG. 5 ends at 520.

[0050] However, when Cresp410 indicates Success_CU, one or more caches may have a copy of the target cache line other than the dirty HPC state, there was no cache having a copy of the target cache line in the HPC state, or the HPC_Ack Presp in Cresp indicating the completion of the write-back of the target cache line to the system memory 108 and the transfer of the coherence ownership of the target cache line to the master 400 responded to the request 402. Therefore, the master 400 starts protecting the target address by opening the protection window 413 and receiving the weight Retry Presp for any colliding snooping requests (block 508). Also, when the FC operation is a flush or AMO operation, the master 400 issues a sweep command (e.g., the BK or BK_Flush command in Table II) on the system fabric to invalidate one or more of the arbitrary shared caching copies of the target memory block (block 510). The process of FIG. 5 proceeds to block 512 following block 510, which indicates a determination of whether the Cresp410 of the sweep command issued in block 510 indicates Success. If it does not indicate Success, the process returns to block 510, which represents the master 400 reissuing the sweep command. When a Cresp410 indicating Success is received for the sweep command issued in block 510, the master 400 ends the protection of the target address of the FC operation request by closing the protection window 413 (block 514). Thereafter, the process of FIG. 5 ends at block 520.

[0051] When the Cresp410 of the FC operation is HPC_Success, it indicates that the processing required for the past dirty HPC cache for a given class of FC operations (e.g., the write-back of the target cache line to the system memory 108 and the update of its directory) has been completed, and there is no cache having a copy of the target cache line in the shared coherence state, the FC operation does not interact with the shared copy of the target cache line, or both. Since there is no caching copy of the target cache line in the data processing system 100 or, if it exists, it is not affected by the FC operation, the master 400 does not issue a sweep command, and the process of FIG. 5 ends at block 520. Note that the Cresp of HPC_Success is not received for the first request of the FC operation and is only received after at least one Retry, as detailed below with reference to blocks 602 and 630 - 632 of FIG. 6.

[0052] Referring now to FIG. 6, this figure is a high-level logic flowchart of an exemplary process in which an HPC cache handles a snooping request for an FC memory access operation according to one embodiment. For a deeper understanding, the flowchart of FIG. 6 will be described in conjunction with the timing diagram 700 of FIG. 7.

[0053] The process shown in FIG. 6 starts at block 600 and then proceeds to block 602, which indicates that the L2 cache memory 230 having a cache line targeted by an FC memory access operation in a dirty HPC coherence state (e.g., coherence state M or T) snoops a request for the FC memory access operation from the local interconnect 114. The FC memory access operation may be, for example, a DCBF, DCBST, or AMO operation as described above, or any other storage modification FC operation that requires writing back modified caching data to the system memory 108. In response to snooping the request for the FC memory access operation, the L2 cache memory 230 allocates the SN machine 311 for handling the FC memory access operation, and sets the SN machine 311 to a busy state so that the SN machine 311 starts protecting the target physical address specified by the request. FIG. 7 shows, at reference numeral 702, the transition of the SN machine 311 allocated for handling the request for the FC memory access operation from an idle state to a busy state. Also, since the modified data from the target cache line has not been written back to the system memory 108, the L2 cache memory 230 provides a weight Retry Presp for the request for the FC memory access operation.

[0054] The SN machine 311 performs normal processing on snooping requests during the busy state (block 604). This normal processing includes writing back the modified data of the target cache line to the associated system memory 108 via the system fabric, and invalidating the target cache line in the local directory 308 as required by the FC request 402 when the update of the system memory 108 is complete. Further, as shown in block 604, while the SN machine 311 is in the busy state, the L2 cache 230 provides a weighted Retry partial response to any snooping request that requests access to the target cache line, as shown by reference numeral 704 in FIG. 7. The weighted Retry partial response causes the associated instance of the response logic 218 to configure a Retry Cresp410 that causes the master of the collision request to reissue that request.

[0055] In some embodiments, the SN machine 311 assigned to the request of the FC memory access operation automatically sets a REF (referee) mode indicator (as shown by block 608 in FIG. 6 and reference numeral 706 in FIG. 7) based on the snooping of the request of the FC memory access operation. In other embodiments represented by the operation block 606 in FIG. 6, the SN machine 311 conditionally sets the REF mode indicator in block 608 only when a collision request for a non-FC operation is snooped by the L2 cache 230 while the SN machine 311 is busy with the work for the request of the FC memory access operation. After block 608 or when the operation block 606 is performed, following the determination that the colliding non-FC request has not been snooped, block 610 determines whether the processing of the snooping FC request by the SN machine 311 is complete. If not, the process in FIG. 6 returns to block 604, which has been described. However, as shown by reference numeral 708 in FIG. 7, if it is determined that the processing of the snooping FC request by the SN machine 311 is complete, the process in FIG. 6 proceeds from block 610 to block 612.

[0056] After the processing of the FC operation is completed, block 612 indicates that the SN machine 311 determines whether the REF mode indicator is set. If it is not set, the SN machine 311 returns to the idle state as shown in block 614 of FIG. 6 and reference numeral 710 of FIG. 7. After block 614, the process of FIG. 6 ends at block 616. Returning to block 612, if the SN machine 311 determines that the REF mode indicator is set, even after the processing of the FC operation request is completed, it does not return to the idle state. Instead, as shown in block 620 of FIG. 6 and reference numeral 712 of FIG. 7, it starts the REF (referee) mode and extends the protection of the target cache line. Also, in one embodiment, in conjunction with the start of the REF mode, the SN machine 311 starts the REF mode timer.

[0057] After block 620, the processing of the SN machine 311 enters a processing loop that monitors the expiration of the REF mode timer (block 640) and the snooping (block 630) of requests for FC operations of the same class (e.g., clean, flush, AMO, etc.). During this processing loop, the SN machine 311 targets the same cache line as the previously snooped FC request as shown in block 622 of FIG. 6 and reference numeral 714 of FIG. 7, while monitoring on the system fabric for any collision operations that are not FC operations of the same class. In response to the detection of any such collision request, the SN machine 311 provides a weighted Retry Presp for the colliding non-FC request as shown in block 624 of FIG. 6. The weighted Retry Presp causes the relevant instance of the response logic 218 to issue a Retry Cresp as described above with reference to Table III and block 504 of FIG. 5. Then, this process proceeds to block 640, which will be described later.

[0058] Also, during this processing loop, as shown in block 630 of FIG. 6 and reference numeral 716 of FIG. 7, the SN machine 311 monitors on the system fabric for any colliding FC operations that target the same cache line as the previously snooped FC request and are of the same class of FC operations (e.g., clean, flush, AMO, etc.). The colliding FC operations may be issued by the same master 400 that issued the previous FC request or may be issued by a different master. In response to the detection of any such colliding FC request, the SN machine 311 provides an HPC_Ack partial response to the colliding FC request, as shown in block 632 of FIG. 6. This HPC_Ack partial response will result in the generation of an HPC_Success Cresp (see, for example, Table III). In response to the receipt of the HPC_Success Cresp (block 634), the process of FIG. 6 proceeds to block 650, which will be described later.

[0059] Block 640 indicates that the SN machine 311 determines whether a REF mode timeout has occurred by reference to a REF mode timer. In various embodiments, the timeout may occur at a static predetermined timer value or, alternatively, may occur at a dynamic value determined based on, for example, the number of colliding FC or non-FC operations or both received by the SN machine 311 during the REF mode. If the timeout has not occurred, the process of FIG. 6 returns to block 622, which has been described. In response to the determination of the REF mode timeout at block 640, the process proceeds to block 650, which, as shown by reference numeral 718 of FIG. 7, indicates that the SN machine 311 ends the REF mode. The SN machine 311 thus finishes protecting the target cache line and returns to the idle state (block 614 of FIG. 6 and reference numeral 720 of FIG. 7). Thereafter, the process of FIG. 6 ends at block 616.

[0060] It will be apparent to those skilled in the art that FIG. 6 discloses a technique in which the HPC snooper temporarily initiates a REF mode that extends the coherence protection of the memory block targeted by the request of the FC memory access operation during the period between the end of the flash / clean operation executed by the HPC snooper and the end of a subsequent FC memory access operation (e.g., Cresp) that specifies the same target address in the same class. By extending the protection of the target cache line, no other HPCs are formed for the target cache line until the success of the FC memory access operation is guaranteed.

[0061] Referring now to FIG. 8, this figure is a high-level logic flowchart of an exemplary process in which a cache having a non-HPC shared copy of the target cache of a snooped flash-type request (e.g., DCBF or AMO) handles the request according to one embodiment. Note that a cache having a shared copy of the target cache of a clean request (e.g., DCBST) ignores the clean request.

[0062] The process of FIG. 8 starts at block 800 and then proceeds to block 802, which indicates that the L2 cache 230 having the target cache line of the flash-type request snoops the first request of the flash-type operation (e.g., DCBF or AMO) or a related sweep command (e.g., BK or BK_FLUSH). In response to the snooping of the first request or sweep command, the L2 cache 230 provides a lightweight Retry response. Also in response to the first request, the L2 cache 230 assigns the SN machine 311 to handle this first request. The assigned SN machine 311 transitions from the idle state to the busy state and starts protecting the target cache line.

[0063] In block 804, the SN machine 311 assigned to handle the requests snooped in block 802 executes normal processing for the first request or the sweep command by invalidating the target cache line in the local directory 308. Further, as shown in block 804, while the SN machine 311 is busy, the L2 cache 230 provides a lightweight Retry partial response for any snooping request that requests access to the target cache line. In block 806, it is determined whether the processing of the snooping request by the SN machine 311 has been completed. If not, the process of FIG. 8 returns to block 804, which has already been described. However, if it is determined that the processing of the snooping request by the SN machine 311 has been completed, the SN machine 311 returns to the idle state (block 808), and the process of FIG. 8 ends at block 810.

[0064] Referring now to FIG. 9, this figure is a block diagram of an exemplary design flow 900 used, for example, in semiconductor IC logic design, simulation, testing, layout, and manufacturing. Design flow 900 includes a process, machine, or mechanism, or a combination thereof, that processes the above-described design structures or devices shown in FIGS. 1 - 3, or both, to generate a logically or functionally equivalent representation of these design structures or devices. The design structures that are processed or generated, or both, by design flow 900 may include data, instructions, or both, that, when encoded on a machine-readable transmission or storage medium and executed or processed on a data processing system, generate a logically, structurally, mechanically, or functionally equivalent representation of a hardware component, circuit, device, or system. Machines may include, but are not limited to, any machine used in an IC design process, such as the design, manufacture, or simulation of a circuit, component, device, or system. For example, machines may include a lithography apparatus, a machine or device, or both, that generates a mask (e.g., an electron beam lithography apparatus), a computer or device that simulates a design structure, any device used in a manufacturing or testing process, or any machine that programs a functionally equivalent representation of a design structure onto any medium (e.g., a machine that programs a programmable gate array).

[0065] Design flow 900 may vary depending on the type of representation of the design object. For example, the design flow 900 for constructing an application-specific integrated circuit (ASIC) may be different from the design flow 900 for designing standard components, or may be different from the design flow 900 for instantiating a design in a programmable array (e.g., a programmable gate array (PGA) or a field-programmable gate array (FPGA) provided by Altera(R) Inc. or Xilinx(R) Inc.).

[0066] FIG. 9 shows such a plurality of design structures, preferably including an input design structure 920 that is processed by a design process 900. The design structure 920 may be a logical simulation design structure that generates a logically equivalent functional representation of a hardware device through generation and processing by the design process 900. Additionally or alternatively, the design structure 920 may include data, program instructions, or both that generate a functional representation of the physical structure of a hardware device when processed by the design process 900. Regardless of whether it represents functional or structural design features or both, the design structure 920 may be generated using electronic computer-aided design (ECAD) as performed by a core developer / designer. When encoded on a machine-readable data transmission, gate array, or storage medium, the design structure 920 can be accessed and processed by one or more hardware or software modules or both within the design process 900 to perform a simulation or functional representation of an electronic component, circuit, electronic or logical module, device, apparatus, or system as shown in FIGS. 1-3. For this reason, the design structure 920 may include source code, a compiled structure, and a computer-executable code structure that is readable by a human, machine, or both and that performs a functional simulation or representation of circuit or other level of hardware logic design when processed by a design or simulation data processing system. Such data structures may include data structures such as low-level hardware description language (HDL) design languages such as Verilog and VHDL or high-level design languages such as C or C++ or both, or HDL design entities that show compliance or conformance or both to such languages.

[0067] The design process 900 preferably employs and incorporates one or both of hardware or software modules that perform synthesis, transformation, or other processing of components, circuits, devices, or logical structure design / simulation functional equivalents shown in FIGS. 1-3 to generate a netlist 980 that may include design structures such as design structure 920. The netlist 980 may include a compiled or processed data structure representing a list of, for example, wires, discrete components, logic gates, control circuits, I / O devices, models, etc., that describe other elements and connections to circuits in an integrated circuit design. The netlist 980 may be synthesized using an iterative process in which the netlist 980 is resynthesized one or more times depending on the design specifications and parameters of the device. Similar to other design structure types described herein, the netlist 980 may be recorded on a machine-readable storage medium or programmed into a programmable gate array. This medium may be a non-volatile storage medium such as a magnetic or optical disk drive, programmable gate array, compact flash, or other flash memory. Additionally or alternatively, the medium may be a system or cache memory, or buffer space.

[0068] The design process 900 may include hardware and software modules that process various input data structure types, including the netlist 980. Such data structure types may exist, for example, within the library elements 930 and may include a set of common elements, circuits, and devices, such as models, layouts, and symbolic representations for a given manufacturing technology (e.g., different technology nodes, 32nm, 45nm, 90nm, etc.). The data structure types may further include the design specifications 940, the characteristic data 950, the verification data 960, the design rules 970, and a test data file 985 that may include input test patterns, output test results, and other test information. The design process 900 may further include standard mechanical design processes, such as process simulations for operations such as stress analysis, thermal analysis, mechanical event simulation, casting, molding, and die pressing. One skilled in mechanical design would be able to recognize the possible range of mechanical design tools and applications used in the design process 900 without departing from the scope and spirit of the present invention. Also, the design process 900 may include modules that perform standard circuit design processes, such as timing analysis, verification, design rule checking, placement, and routing operations.

[0069] The design process 900 processes the design structure 920, optionally with any additional mechanical design or data, by incorporating and employing logic and physical design tools such as HDL compilers and simulation model building tools, to generate a second design structure 990 that is integral with part or all of the illustrated support data structure. The design structure 990 exists on a storage medium or programmable gate array in a data format (e.g., IGES, DXF, Parasolid XT, JT, DRG, or any other suitable format storing or representing such mechanical design structures) used for the exchange of data for mechanical devices and structures. Similar to the design structure 920, the design structure 990 exists on a transmission or data storage medium and preferably includes one or more files, data structures, or other computer-encoded data or instructions that, when processed by an ECAD system, generate one or more logical or functionally equivalent forms of one or more of the embodiments of the invention shown in FIGS. 1-3. In one embodiment, the design structure 990 may include a compiled executable HDL simulation model that functionally emulates the devices shown in FIGS. 1-3.

[0070] In addition, the design structure 990 can adopt a data format or a symbol data format or both used for the exchange of integrated circuit layout data (for example, information stored in a format such as GDSII (GDS2), GL1, OASIS, a map file, or any other suitable format for storing such design data structures). The design structure 990 may include, for example, symbol data, a map file, a test data file, a design content file, manufacturing data, layout parameters, wires, metal levels, vias, shapes, data of wiring through a manufacturing line, and any other data such as information required by a manufacturer or other designer / developer to generate the devices or structures shown and described above in FIGS. 1 to 3. Thereafter, the design structure 990 may proceed to stage 995, where, for example, tape-out of the design structure 990, disclosure for manufacturing, disclosure to a mask house, sending to another design house, return to a customer, etc. are performed.

[0071] As described above, in at least one embodiment, the cache memory includes a data array, a directory of the content of the data array that specifies coherence state information, and snooping logic that processes operations snooped from the system fabric by reference to the data array and the directory. The snooping logic determines whether the cache memory has coherence ownership of the target address in response to snooping on the system fabric a request for a first flash / clean memory access operation that specifies the target address. Based on the determination that the cache memory has coherence ownership of the target address, the snooping logic starts a referee mode after corresponding to the request. In the referee mode, the snooping logic protects the memory block identified by the target address against collision memory access requests by a plurality of processor cores until the end of a second flash / clean memory access operation that specifies the target address.

[0072] Although various embodiments have been illustrated and described in detail, those skilled in the art will be able to make various changes to the form and details without departing from the spirit and scope of the appended claims, and it should be understood that all of these alternative embodiments are included in the appended claims.

[0073] The flowcharts and block diagrams in the figures illustrate the architecture, functionality, and operation of possible embodiments of systems, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in the flowchart or block diagram may represent a module, segment, or portion of instructions that include one or more executable instructions for implementing a particular logical function. In some alternative embodiments, the functions described in the blocks may occur out of the order depicted in the drawings. For example, two blocks shown in succession may in fact be executed substantially simultaneously, depending on the functions involved, or may sometimes be executed in the reverse order. Also, each block in the block diagram or flowchart diagram or both, and combinations of blocks in the block diagram or flowchart diagram or both, can be implemented by a dedicated hardware-based system that performs a particular function or operation, or by a combination of dedicated hardware and computer instructions.

[0074] Although multiple aspects of a computer system that executes program code that directs the functions of the present invention have been described, it is to be understood that the present invention may alternatively be implemented as a program product that includes a computer-readable storage device storing program code that, when processed by a processor of a data processing system, causes the data processing system to perform the above functions. Examples of computer-readable storage devices include volatile or non-volatile memories, optical or magnetic disks, etc., but exclude non-statutory subject matter such as propagation signals themselves, transmission media themselves, and energy forms themselves.

[0075] As an example, a program product may include data or instructions or both that, when executed or processed on a data processing system, generate a logical, structural, or functional equivalent representation (including a simulation model) of the hardware components, circuits, devices, or systems disclosed herein. Such data or instructions or both may include data structures such as HDL design entities that conform to or are compliant with or both a low-level hardware description language (HDL) design language such as Verilog and VHDL or a high-level design language such as C or C++. Further, the data or instructions or both may adopt a data format or a symbolic data format or both used for the exchange of integrated circuit layout data (for example, information stored in a suitable format such as GDSII (GDS2), GL1, OASIS, a map file, or any other suitable format for storing such design data structures).

Claims

1. A cache memory of a related processor core among a plurality of processor cores in a multiprocessor data processing system, wherein the multiprocessor data processing system includes a system fabric, and a memory controller of the cache memory and a system memory are communicatively coupled to receive operations on the system fabric, the cache memory including the system fabric, a data array, a directory of the contents of the data array, the directory including coherence state information, snooping logic for processing operations snooped from the system fabric by reference to the data array and the directory, setting a referee mode indicator in response to snooping, on the system fabric, a request for a first flash / clean memory access operation of one of the plurality of processor cores that specifies a target address, after the setting, determining whether the referee mode indicator is set, starting a referee mode based on a determination that the referee mode indicator is set, in the referee mode, protecting a memory block identified by the target address against collision memory access requests by the plurality of processor cores until the end of a second flash / clean memory access operation that specifies the target address, the snooping logic performing the above, a cache memory comprising the above.

2. The request for the flash or clean operation is a first request, The cache memory according to claim 1, wherein the snooping logic is configured to start the referee mode based on snooping of a second request that collides prior to completion of processing of the first request after snooping of the first request.

3. The cache memory according to claim 1, wherein the snooping logic is configured to protect the memory block against the collision memory access request by issuing a Retry coherence response to the collision memory access request.

4. The cache memory according to claim 1, wherein the snooping logic is configured to provide a first coherence response to a collision request of a flash / clean memory access operation of the same class as the first flash / clean memory access operation in the referee mode, and to provide a different second coherence response to other types of collision requests.

5. The cache memory according to claim 1, wherein the snooping logic detects a timeout condition in the referee mode and ends the referee mode in response to the detection of the timeout condition.

6. The cache memory according to claim 1, and at least one associated processor core coupled to the cache memory, a processing unit comprising the same.

7. A system fabric, a plurality of processing units according to claim 6, all coupled to the system fabric, a data processing system comprising the same.

8. A method of data processing in a multi-processor data processing system, including a cache memory of an associated processor core among a plurality of processor cores in the multi-processor data processing system, the multi-processor data processing system being a system fabric, and a memory controller of the cache memory and a system memory being communicatively coupled to receive operations on the system fabric, the method comprising: snooping, on the system fabric, a request for a first flash / clean memory access operation of one of the plurality of processor cores specifying a target address; setting a referee mode indicator in response to the snooping of the request; determining whether the referee mode indicator is set after the setting; Based on the determination that the referee mode indicator is set, starting the referee mode; In the referee mode, the cache memory protects the memory block identified by the target address against collision memory access requests from the plurality of processor cores until the end of the second flush / clean memory access operation that specifies the target address; A method comprising. **Claim 9** The request for the flash or clean operation is a first request, The method according to claim 8, wherein starting the referee mode includes starting the referee mode based on snooping of a second request that collides prior to completion of processing of the first request by snooping logic after snooping of the first request. **Claim 10** The method according to claim 8, wherein the cache memory protecting includes protecting the memory block against the collision memory access request by issuing a Retry coherence response to the collision memory access request. **Claim 11** The method according to claim 8, further comprising the cache memory providing a first coherence response to collision requests for flash / clean memory access operations of the same class as the first flash / clean memory access operation and providing a different second coherence response to other types of collision requests in the referee mode. **Claim 12** The method according to claim 8, further comprising the cache memory detecting a timeout condition in the referee mode and ending the referee mode in response to the detection of the timeout condition A method further comprising. **Claim 13** A design structure configured to be executed on a computer that simulates the logic of an integrated circuit, The design structure includes A processor core, A cache memory, A data array, A directory of the contents of the data array, including coherence state information, The directory, Snooping logic that processes operations snooped from the system fabric by reference to the data array and the directory, and in response to snooping on the system fabric a request for a first flash / clean memory access operation of one of a plurality of processor cores that specify a target address, setting a referee mode indicator; after said setting, determining whether the referee mode indicator is set; starting a referee mode based on a determination that the referee mode indicator is set; in the referee mode, protecting a memory block identified by the target address against collision memory access requests by the plurality of processor cores until the end of a second flash / clean memory access operation that specifies the target address; said snooping logic that performs; said cache memory that includes; a simulation model of the processing unit in which functions of the processing unit are simulated, including; a design structure that, when executed by a computer that simulates the logic of the integrated circuit, simulates functions of the processing unit including the processor core and the cache memory.

14. said request for the flash or clean operation is a first request, The design structure according to claim 13, wherein the snooping logic is configured to start the referee mode based on snooping of a second request that collides prior to completion of processing of the first request after snooping of the first request.

15. The design structure according to claim 13, wherein the snooping logic is configured to protect the memory block against the collision memory access request by issuing a Retry coherence response to the collision memory access request.

16. The snooping logic is configured to provide a first coherence response to a collision request for a flash / clean memory access operation of the same class as the first flash / clean memory access operation in the referee mode and to provide a different second coherence response to other types of collision requests, according to the design structure of claim 13.

17. The snooping logic is configured to detect a timeout condition in the referee mode and to end the referee mode in response to the detection of the timeout condition, according to the design structure of claim 13.

Citation Information

Patent Citations

  • Avoidance of Deadlock in a Processor-Based System Using a Retry Bus Coherence Protocol and an In-Order Response Non-Retry Bus Coherence Protocol

    JP2018533133A

  • Data processing system and method for resolving a conflict between requests to modify a shared cache line

    US20020129211A1

  • Forward progress mechanism for stores in the presence of load contention in a system favoring loads

    US20130205087A1

  • Early freeing of a snoop machine of a data processing system prior to completion of snoop processing for an interconnect operation

    US20170293559A1

  • Hardware assisted cache flushing mechanism

    US20180143903A1