Centralized distribution of multicast requests in data processing system

By introducing central request proxy and communication protocol, the problem that multicast requests are not accepted and repeatedly processed by all snoops in the multiprocessor system is solved, and efficient distribution and processing of multicast requests is realized, ping-pong locks are avoided, and system efficiency is improved.

CN120500682APending Publication Date: 2025-08-15INTERNATIONAL BUSINESS MACHINE CORPORATION
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202380087315.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2022-12-21
Filing Date
2023-12-07
Publication Date
2025-08-15

AI Technical Summary

Technical Problem

In a multiprocessor computer system, the ping-pong lock of the multicast request causes the multicast request not to be accepted and processed by all snoops, and may be processed repeatedly, affecting the system efficiency.

Method used

Introduce a central request proxy and associated communication protocols to ensure that multicast requests are assigned to a specific state machine by the central request proxy, and that a consistent response is issued only after all snoops are successfully received, avoiding retry and duplicate processing.

Benefits of technology

Ensure that multicast requests are accepted and processed by all relevant snoops at one time, avoiding ping-pong locks, and improving the processing efficiency and consistency of the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120500682A_ABST
    Figure CN120500682A_ABST
Patent Text Reader

Abstract

A data processing system includes a master device, a central request agent, and a plurality of snoopers communicatively coupled to a system fabric for communicating requests to be retried. A master device issues a multicast request on a system fabric intended for a plurality of snoopers. The central request agent receives a multicast request on the system fabric, allocates the multicast request to a particular state machine among a plurality of state machines in the central request agent, and provides a consistent response to the master device indicating a successful completion of the multicast request. The central request agent repeatedly issues the multicast request on the system fabric in association with a machine identifier identifying the particular state machine until the coherence response indicates that the multicast request is successfully received by all of the plurality of snoopers.
Need to check novelty before this filing date? Find Prior Art

Description

Background Art

[0001] The present invention relates generally to data processing, and in particular to the distribution of multicast requests (eg, translation entry invalidation requests) in a multi-threaded data processing system.

[0002] A conventional multiprocessor (MP) computer system includes multiple processing units (each of which may include one or more processor cores and their various cache memories), input / output (I / O) devices, and data storage, which may include system memory (which may be volatile or non-volatile) and non-volatile mass storage. To provide sufficient addresses for memory-mapped I / O operations and for data and instructions utilized by operating systems and application software, MP computer systems typically reference an effective address space that includes a much larger number of effective addresses than the number of physical storage locations in memory-mapped I / O devices and system memory. Therefore, to perform memory-mapped I / O or access system memory, a processor core within a computer system utilizing effective addressing is required to translate an effective address into an actual address assigned to a particular I / O device or physical storage location within system memory.

[0003] In POWER TM In a RISC architecture, the effective address space is divided into multiple uniformly sized memory pages, where each page has a corresponding associated address descriptor, called a page table entry (PTE). The PTE corresponding to a particular memory page contains the effective base address of the memory page and the associated real base address of the page frame, enabling the processor core to translate any effective address within the memory page into a real address in system memory. PTEs created in system memory by the operating system and / or hypervisor software are collected in a page frame table.

[0004] To speed up the translation of effective addresses to real addresses during the processing of memory-mapped I / O and memory access instructions (hereinafter collectively referred to as "memory reference instructions"), conventional processor cores typically employ a cache known as a Translation Lookaside Buffer (TLB) and other translation structures to buffer recently accessed PTEs within the processor core. Of course, as data is moved into and out of physical storage locations in system memory (e.g., in response to invoking a new process or context switch), entries in the TLB must be updated to reflect the presence of the new data, and TLB entries associated with data removed from system memory (e.g., paged out to non-volatile mass storage) must be invalidated. In many conventional processors (e.g., the POWER 1000 processors available from IBM Corporation), the TLB entries must be updated to reflect the presence of the new data. TM series processors), invalidation of TLB entries is the responsibility of software and is accomplished by executing explicit TLB invalidate-entry instructions (e.g., POWER TMThis is implemented using the TLBIE in the instruction set architecture (ISA).

[0005] In an MP computer system, invalidation of a PTE cached in the TLB of one processor core is complicated by the fact that each other processor core has its own corresponding TLB (which may also cache a copy of the target PTE). In order to maintain a consistent view of system memory across all processor cores, invalidation of a PTE in one processor core requires invalidation of the same PTE (if any) in the TLBs of all other processor cores. In many conventional MP computer systems, invalidation of PTEs in all processor cores in the system is achieved by executing a TLB invalidate entry instruction in the initiating processor core and broadcasting a TLB invalidate entry request from the initiating processor core to each other processor core in the system. In the instruction sequence of the initiating processor core, the TLB invalidate entry instruction (or multiple instructions if multiple PTEs are to be invalidated) may be followed by one or more synchronization instructions that ensure that the TLB entry invalidation has been executed by all processor cores.

[0006] The present disclosure recognizes that, if not handled appropriately, the broadcast of a TLB invalidate entry request from an initiating processor core to every other processor core in the system can result in a livelock in a system that allows recipients to retry interconnect operations. Therefore, it would be useful and desirable to provide an improved technique for managing the distribution of multicast requests (e.g., translation entry invalidate requests) in a data processing system that allows retry of interconnect operations. Summary of the Invention

[0007] As briefly mentioned above, data processing systems that allow snoopers to retry multicast requests on a system interconnect experience what is known as a "ping-pong" livelock. Ping-pong livelock occurs when a first subset of snoopers accepts and potentially begins processing a multicast request and a second subset of snoopers issues a retry response to the multicast request indicating that the request cannot currently be processed. In response to the retry response, the master that initiated the multicast request will reissue the multicast request on the system interconnect, potentially causing one or more snoopers in the second subset to accept the request and one or more snoopers in the first subset to issue a retry response. Furthermore, when the request is reissued by the master, one or more snoopers in the first subset may additionally accept the request again and restart processing of the request.

[0008] This situation illustrates two potential problems with multicast requests in a data processing system where interconnect requests may be retried. First, there is no guarantee that a multicast request will be accepted and processed by all snoopers necessary to complete the multicast request. Second, if retried and reissued, the multicast request may be processed repeatedly by one or more snoopers.

[0009] To avoid ping-pong livelock for multicast requests, the present disclosure employs a central request broker and associated communication protocols that ensure that a given multicast request will eventually be accepted by each relevant snooper, that a relevant snooper will only process a multicast request once (even if retried), and that the issuance of one or more other multicast requests of the same type will not prevent a given multicast request from being accepted by all relevant snoopers. The central request broker has a given number of state machines for tracking multicast requests, and the snoopers have a corresponding number of snoop machines.

[0010] Multicast requests issued by various master devices distributed within a data processing system on the system interconnect compete for acceptance by a central request agent's state machine. Once a multicast request is accepted by a given state machine in the central request agent, the central request agent forwards the multicast request to the corresponding snooper among the various snoopers distributed within the data processing system. Because the request is forwarded by the central request agent, one or more snoopers can provide a retry-consistent response to the multicast request while one or more other snoopers accept and process the multicast request. If a snooper within a snooper is still busy processing a previous multicast request assigned to the same snooper, the snooper will provide a retry response. Ultimately, the multicast request forwarded by the central request agent will be snooped and processed by all relevant snoopers, as indicated to the central request agent by a successful combined response. It should be noted that if the multicast request is the same as a multicast request that a snooper is currently processing or has just completed processing, the snooper does not issue a retry-consistent response to the snooped multicast request.

[0011] According to one embodiment, a data processing system includes a master device, a central request agent, and a plurality of snoopers communicatively coupled to a system structure for transmitting requests to be retried. The master device issues a multicast request on the system structure intended for the plurality of snoopers. The central request agent receives the multicast request on the system structure, assigns the multicast request to a specific state machine among a plurality of state machines in the central request agent, and provides a consistency response to the master device indicating successful completion of the multicast request. The central request agent repeatedly issues the multicast request on the system structure in association with a machine identifier identifying the specific state machine until a consistency response indicates that the multicast request was successfully received by all of the plurality of snoopers.

[0012] In at least one embodiment, the multicast request comprises a translation entry invalidation request.

[0013] In at least one embodiment, each of the plurality of snoopers includes a plurality of snoopers corresponding in number to the plurality of state machines in the central request agent, and each of the plurality of snoopers distributes the multicast request received from the central request unit to a specific snooper among the plurality of snoopers corresponding to the specific state machine.

[0014] In at least one embodiment, prior to repeating the multicast request, the central requesting agent modifies the multicast request to indicate that the multicast request is forwarded from the central requesting agent.

[0015] In at least one embodiment, in response to the central requesting agent repeating the multicast request, one or more of the plurality of snoopers provide a null coherence response to indicate successful receipt of the multicast request previously issued by the central requesting agent.

[0016] In at least one embodiment, the multicast request includes a plurality of operational tenures on the system architecture. In such an embodiment, the central request agent receives at least a first operational tenure and a corresponding second operational tenure, and assigns the first operational tenure and the second operational tenure to the specific state machine. Based on both the first operational tenure and the second operational tenure being assigned to the specific state machine, the central request agent marks the multicast request as available for distribution to the plurality of snoopers.

[0017] In at least one embodiment, the central request agent issues the first operational tenure of the multicast request on the system fabric with a time period indication, and based on a mismatching time period indication, one of the plurality of snoopers discards a previously snooped multicast request.

[0018] The disclosed embodiments may be implemented as a method, an integrated circuit, a data processing system, and / or a design structure. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] Figure 1 is a high-level block diagram of an exemplary data processing system according to one embodiment;

[0020] Figure 2 is a more detailed block diagram of an exemplary processing unit according to one embodiment;

[0021] Figure 3 is a detailed block diagram of a processor core and lower level cache memory according to one embodiment;

[0022] Figure 4A is a first exemplary conversion entry invalidation instruction sequence according to one embodiment;

[0023] Figure 4B is a second exemplary conversion entry invalidation instruction sequence according to one embodiment;

[0024] Figure 5 is a high-level logic flow diagram of an exemplary method for a processor core of a multi-processor data processing system to process a conversion entry invalidation instruction according to one embodiment;

[0025] Figure 6 is a high-level logic flow diagram of an exemplary method for sidecar logic of a processing unit to handle a translation entry invalidation request according to one embodiment;

[0026] Figure 7 is a high-level logic flow diagram of an exemplary method for a snooper of a processing unit to handle conversion entry invalidation requests and conversion synchronization requests according to one embodiment;

[0027] Figure 8 is a high-level logic flow diagram of an exemplary method for an arbiter of a processing unit to handle a translation entry invalidation request according to one embodiment;

[0028] Figure 9 is a high-level logic flow diagram of an exemplary method for a translation sequencer of a processor core to handle a translation entry invalidation request according to one embodiment;

[0029] Figure 10 is a high-level logic flow diagram of an exemplary method for a storage queue of a processing unit to handle a conversion invalidation completion request according to one embodiment;

[0030] Figure 11 is a high-level logic flow diagram of an exemplary method for a processor core to process a conversion synchronization instruction according to one embodiment;

[0031] Figure 12 is a high-level logic flow diagram of an exemplary method for sidecar logic of a processing unit to handle a conversion synchronization request according to one embodiment;

[0032] Figure 13 is a high-level logic flow diagram of an exemplary method for a processing core to handle a page table synchronization instruction according to one embodiment;

[0033] Figure 14 is a high-level logic flow diagram of an exemplary method for a processing unit to handle a page table synchronization request according to one embodiment;

[0034] Figure 15 is a high-level logic flow diagram of an exemplary method for snooper logic of a processing unit to handle translation entry invalidation requests, translation invalidation completion requests, and page table synchronization requests according to one embodiment;

[0035] Figure 16According to one embodiment Figure 1 A space-time diagram of interconnected operations on a system architecture of a data processing system;

[0036] Figure 17 shows an exemplary format of a tag portion of an interconnect operation according to one embodiment;

[0037] Figure 18 is a high-level block diagram of an exemplary embodiment of a central request broker according to one embodiment;

[0038] Figure 19 is a high-level logic flow chart of an exemplary method for an initiating processing unit or central request agent to issue a translation entry invalidation request according to the first embodiment;

[0039] Figure 20 is a high-level logic flow diagram of an exemplary method for a central request agent to receive a translation entry invalidation request from an originating processing unit according to a first embodiment;

[0040] Figure 21 is a high-level logic flow diagram of an exemplary method for selecting a translation entry invalidation request to be issued by a central request agent according to the first embodiment;

[0041] Figure 22 is a high-level logic flow diagram of an exemplary method for a central request agent to issue a translation entry invalidation request according to the first embodiment;

[0042] Figure 23 is a high-level logic flow diagram of an exemplary method for a processing unit to receive a translation entry invalidation request from a central request agent according to the first embodiment;

[0043] Figure 24 is a high-level logic flow chart of an exemplary method for an initiating processing unit or central request agent to issue a translation entry invalidation request according to a second embodiment;

[0044] Figure 25 is a high-level logic flow diagram of an exemplary method for a central request agent to receive a translation entry invalidation request from an originating processing unit according to a second embodiment;

[0045] Figure 26 is a high-level logic flow diagram of an exemplary method for selecting a translation entry invalidation request to be issued by a central request agent according to a second embodiment;

[0046] Figure 27 is a high-level logic flow diagram of an exemplary method for a central request agent to issue a translation entry invalidation request according to a second embodiment;

[0047] Figures 28-29together form a high-level logic flow chart of an exemplary method for a processing unit to receive a translation entry invalidation request for processing according to the second embodiment; and

[0048] Figure 30 is a data flow diagram showing the design process. DETAILED DESCRIPTION

[0049] Referring now to the drawings, wherein like reference numerals designate like and corresponding parts throughout, and particularly to Figure 1 , a high-level block diagram depicting an exemplary data processing system 100 according to one embodiment is shown. In the depicted embodiment, data processing system 100 is a cache-coherent symmetric multiprocessor (SMP) data processing system that includes a plurality of processing nodes 102 for processing data and instructions. Processing nodes 102 are coupled to a system interconnect 110 for communicating address, data, and control information. System interconnect 110 can be implemented as, for example, a bus interconnect, a switch interconnect, or a hybrid interconnect.

[0050] In the depicted embodiment, each processing node 102 is implemented as a multi-chip module (MCM) containing multiple (e.g., four) processing units 104a-104d, each preferably implemented as a respective integrated circuit. The processing units 104 within each processing node 102 are coupled to communicate with each other and with the system interconnect 110 via a local interconnect 114. The local interconnect 114 is similar to the system interconnect 110 and may be implemented, for example, using one or more buses and / or switches. Together, the system interconnect 110 and the local interconnect 114 form a system fabric.

[0051] As the following reference Figure 2 As described in more detail, processing units 104 each include a memory controller 106 coupled to a local interconnect 114 to provide an interface to a corresponding system memory 108. Data and instructions residing in system memory 108 can generally be accessed, cached, and modified by a processor core in any processing unit 104 of any processing node 102 within data processing system 100. System memory 108 thus forms the lowest level of memory storage in the distributed shared memory system of data processing system 100. In alternative embodiments, one or more memory controllers 106 (and system memory 108) may be coupled to system interconnect 110 instead of local interconnect 114.

[0052] Figure 1 The data processing system 100 is also shown to include at least one central request agent 120 coupled to the system architecture. Figure 18As discussed further below, central request agent 120 may be used to distribute certain types of multicast requests across the system fabric while avoiding livelock. In some examples, data processing system 100 may include a single central request agent 120 that may be conveniently coupled to system interconnect 110 .

[0053] Those skilled in the art will understand that Figure 1 The SMP data processing system 100 may include many additional components not shown, such as interconnect bridges, non-volatile storage devices, ports for connecting to networks or attached devices, etc. Because these additional components are not necessary for understanding the described embodiments, they are not described in detail. Figure 1 However, it should also be understood that the enhancements described herein are applicable to data processing systems of various architectures and are in no way limited to Figure 1 The general data processing system architecture shown in .

[0054] Now refer to Figure 2 , depicts a more detailed block diagram of an exemplary processing unit 104 according to one embodiment. In the depicted embodiment, each processing unit 104 is an integrated circuit that includes one or more processor cores 200 for processing instructions and data. In a preferred embodiment, each processor core 200 supports simultaneous multithreading (SMT) and is therefore capable of independently and concurrently executing multiple hardware execution threads.

[0055] The operation of each processor core 200 is supported by a multi-level memory hierarchy having, at its lowest level, a shared system memory 108 accessed via an integrated memory controller 106. As shown, the shared system memory 108 stores a page frame table 220 containing a plurality of page table entries (PTEs) 222 for performing effective to real address translations to enable access to memory locations in the system memory 108. At its higher levels, the multi-level memory hierarchy includes one or more levels of cache memory, which, in the illustrative embodiment, includes a direct-flow, store-through level 1 (L1) cache 302 within and dedicated to each processor core 200 (see FIG. Figure 3 ), and a corresponding level 2 (L2) cache 230 for each processor core 200. Although the cache hierarchy shown includes only two levels of cache, those skilled in the art will appreciate that alternative embodiments may include additional levels (L3, L4, etc.) of on-chip or off-chip, private or shared, inline or lookaside cache that may fully, partially, or exclude the contents of higher level caches.

[0056] Each processing unit 104 also includes an integrated and distributed fabric controller 216, which is responsible for controlling the flow of operations on the system fabric, including the local interconnect 114 and the system interconnect 110, and for implementing the coherent communications required to implement the selected cache coherence protocol. The processing unit 104 also includes an integrated I / O (input / output) controller 214 that supports the attachment of one or more I / O devices (not shown).

[0057] Now refer to Figure 3 , shows a more detailed block diagram of an example embodiment of a processor core 200 and its accompanying L2 cache 230 according to one embodiment.

[0058] In the illustrated embodiment, the processor core 200 includes one or more execution units 300 that execute instructions from multiple simultaneous hardware execution threads. These instructions may include, for example, arithmetic instructions, logical instructions, and memory reference instructions, as well as conversion entry invalidation instructions (hereinafter referred to as POWER TM ISA mnemonics refer to TLBIE (Translation Lookaside Buffer Invalidate Entry) and associated synchronization instructions. Execution unit 300 can generally execute the instructions of a hardware thread in any order, as long as the data dependencies and explicit ordering required by the synchronization instructions are followed.

[0059] The processor core 200 also includes a memory management unit (MMU) 308, which is responsible for translating target effective addresses determined by the execution of memory-referenced instructions in the execution unit 300 into real addresses. The MMU 308 performs effective-to-real address translation by referencing one or more translation structures 310 (e.g., a translation lookaside buffer (TLB), a block address table (BAT), segment lookaside buffers (SLBs), etc.). The number and type of these translation structures vary depending on the implementation and architecture. If present, the TLB reduces the latency associated with effective-to-real address translation by caching PTEs 222 retrieved from the page frame table 220. A translation sequencer 312 associated with the translation structure 310 handles invalidation of effective-to-real translation entries stored within the translation structure 310 and manages such invalidation with respect to running memory-referenced instructions in the processor core 200.

[0060] Processor core 200 also includes various memory facilities shared by the multiple hardware threads supported by processor core 200. The memory facilities shared by the multiple hardware threads include an L1 store queue (L1 STQ) 304, which temporarily buffers store and synchronization requests generated by the execution unit(s) 300 executing corresponding store and synchronization instructions. Because L1 cache 302 is a flow-through store cache, meaning that coherency is fully determined at lower levels of the cache hierarchy (e.g., at L2 cache 230), requests flow through L1 STQ 304 and are then passed to L2 cache 230 via bus 318 for processing. The memory facilities of processor core 200 shared by the multiple hardware threads also include a load miss queue (LMQ) 306, which temporarily buffers load requests that miss in L1 cache 302. Because such load requests have not yet been satisfied, if the address translation entry used to obtain the target real address of the load request is invalidated before the load request is satisfied, the load request experiences a memory page miss. Therefore, if a PTE or other translation entry is to be invalidated, any load requests in LMQ 306 that depend on that translation entry must be drained from LMQ 306 and satisfied before the effective address translated by the dependent translation entry can be reallocated.

[0061] Still refer to Figure 3 , the L2 cache 230 includes a cache array 332 and an L2 directory 334 for the contents of the cache array 332. Assuming that the cache array 332 and the L2 directory 334 are set associative as conventionally, memory locations in the system memory 108 are mapped to specific congruence classes within the cache array 332 by utilizing predetermined index bits within the system memory (real) address. The specific memory blocks stored within the cache lines of the cache array 332 are recorded in the L2 directory 334, which contains one directory entry for each cache line. Although Figure 3 Not explicitly depicted in, but understood by those skilled in the art, each directory entry in the cache directory 334 includes various fields, such as a tag field identifying the real address of the memory block held in the corresponding cache line of the cache array 332, a status field indicating the coherency state of the cache line, an LRU (least recently used) field indicating the replacement order of the cache line relative to other cache lines in the same congruence class, and an inclusive bit indicating whether the memory block is held in the associated L1 cache 302.

[0062] The L2 cache 230 also includes an L2 STQ 320, which receives store modification requests and synchronization requests from the L1 STQ 304 via an interface 321 and buffers such requests. It should be noted that the L2 STQ 320 is a unified store queue that buffers requests for all hardware threads of the attached processor core 200. Therefore, store requests, TLBIE requests, and associated synchronization requests for all threads flow through the L2 STQ 320. Although in most embodiments the L2 STQ 320 includes multiple entries, the L2 STQ 320 is required to operate in a deadlock-free manner regardless of depth (i.e., even if implemented as a single-entry queue). To this end, the L2 STQ 320 is coupled to associated sidecar logic 322 via an interface 321. The sidecar logic 322 includes one request buffer entry (referred to herein as a "sidecar") 324 for each hardware thread supported by the attached processor core 200. Thus, the number of sidecars 324 is independent of the number of entries in the L2 STQ 320. As further described herein, the use of the sidecar 324 allows potentially deadlocked requests to be removed from the L2 STQ 320 so that no deadlock occurs during translation entry invalidation.

[0063] L2 cache 230 also includes dispatch / response logic 336, which receives local load and store requests initiated by attached processor core 200 via buses 327 and 328, respectively, and remote requests snooped over local interconnect 114 via bus 329. These requests (including local and remote load requests, store requests, TLBIE requests, and associated synchronization requests) are processed by dispatch / response logic 336 and then dispatched to the appropriate state machine for servicing.

[0064] In the illustrated embodiment, the state machines implemented within the L2 cache 230 to service requests include multiple read request (RC) machines 342 that independently and concurrently service load (LD) and store (ST) requests received from the attached processor cores 200. To service remote memory access requests originating from processor cores 200 other than the attached processor cores 200, the L2 cache 230 also includes multiple snoop (SN) machines 344. Each snoop machine 344 can independently and concurrently process remote memory access requests snooped from the local interconnect 114. As will be appreciated, the servicing of memory access requests by the RC machines 342 may require the replacement or invalidation of memory blocks within the cache array 332 (and the L1 cache 302). Accordingly, the L2 cache 230 also includes a CO (eviction) machine 340 that manages the removal and write-back of memory blocks from the cache array 332.

[0065] In the depicted embodiment, the L2 cache 230 also includes a plurality of translation snoop (TSN) engines 346 that are used to service TLBIE requests and associated synchronization requests. It should be understood that in some embodiments, the TSN engines 346 may be implemented in another subunit of the processing unit 104, such as a non-cacheable unit (NCU) (not shown) that handles non-cacheable memory access operations. In at least one embodiment, the same number of TSN engines 346 are implemented at each L2 cache 230 to simplify the implementation of a consensus protocol (as further discussed herein) that coordinates the processing of multiple concurrent TLBIE requests within the data processing system 100.

[0066] Each TSN engine 346 is coupled to an arbiter 348, which selects requests to be processed by the TSN engine 346 for transmission to the conversion sequencer 312 in the processor core 200 via a bus 350. In at least some embodiments, the bus 350 is implemented as a unified bus that not only transmits requests from the TSN engine 346 but also returns data from the L2 cache 230 to the processor core 200, as well as performs other operations. It should be noted that the conversion sequencer 312 must accept requests from the arbiter 348 in a non-blocking manner to avoid deadlock.

[0067] Now refer to Figure 16 , depicting Figure 1 FIG1 is a time-space diagram of an exemplary interconnect operation on the interconnect structure of data processing system 100. When a master device 1600 (e.g., a read request (RC) engine 342 of L2 cache 230, a master within I / O controller 214, a central request agent 120, a master within a non-cacheable unit (NCU) (not shown), etc.) issues a request 1602 on the interconnect structure, the interconnect operation begins. Figure 17 In the example given, request 1602 includes an address field 1702 for specifying the actual address of the target to be accessed, a transaction type (ttype) field 1710 for specifying the type of access to be performed, a tag field (tag) 1720 for identifying the master device 1600 that initiated request 1602, and miscellaneous fields 1740 for specifying additional request information. In the depicted example, tag field 1720 includes a topographic identifier (ID) 1722 that uniquely identifies the physical location of the master device within data processing system 100, a machine ID 1724 that identifies which state machine (e.g., CO 340, RC 342, sidecar 324, etc.) is the master device, and a forwarding (F) field 1726 that is set to indicate that the request is being forwarded from central request agent 120 and is reset otherwise. In some embodiments, tag field 1720 may also include an era (E) field 1728 and an operation type (O) field 1730, as discussed further below.

[0068] Common types of requests 1602 include the requests set forth in Table I below.

[0069] Table I

[0070]

[0071]

[0072] like Figure 16 As shown, a request 1602 is received by snoopers 1604a-1604n distributed throughout the data processing system 100 (e.g., the snooper 344 of the L2 cache 230, the central request agent 120, and the snooper of the memory controller 106). Each snooper 1604 that receives and processes the request 1602 may each provide a respective partial response (Presp) 1606 representing at least the response of the snooper 1604 to the request 1602. The snooper within the memory controller 106 may determine which partial response 1606 to provide based on, for example, whether the snooper is responsible for the request address and whether it has resources available to service the request. The L2 cache 230 may determine its partial response 1606 based on, for example, the availability of the snooper 344 to process the request, the availability of its L2 directory 334, and the coherence state associated with the target real address in the L2 directory 334.

[0073] The partial responses 1606 of the snoopers 1604a-1604n are logically combined in stages or all at once by one or more instances of response logic 1622 to determine a system-level combined response (Cresp) 1610 to the request 1602. Although not explicitly shown, the Cresp 1610 preferably includes a tag field 1720 corresponding to the tag field of the original request 1602 to enable the master 1600 to match its request 1602 with the combined response 1610. In one embodiment assumed below, the instance of response logic 1622 responsible for generating the Cresp 1610 is located in the processing unit 104 of the master 1600 that issued the request 1602. The response logic 1622 provides the Cresp 1610 to the master 1600 and the snoopers 1604 via the interconnect structure to indicate a system-level consistent response (e.g., success, failure, retry, etc.) to the request 1602. If Cresp 1610 indicates that request 1602 is successful, Cresp 1610 can indicate, for example, the data source of the target storage block of request 1602, the coherence state of the requested storage block as it will be cached by master device 1600 (or other cache), and whether a "purge" operation is required to invalidate the requested storage block in one or more caches.

[0074] In response to receiving Cresp 1610, one or more of the master 1600 and the snooper 1604 may perform one or more additional actions to service the request 1602. These additional actions may include providing data to the master 1600, invalidating or otherwise updating the coherence state of data cached in one or more L2 caches 230, performing eviction operations, writing data back to the system memory 108, etc. If required by the request 1602, the requested memory block or target memory block may be transferred to or from the master 1600 before or after the response logic 1622 generates Cresp 1610.

[0075] Now refer to Figure 4A , shows a first exemplary translation entry invalidation instruction sequence 400 that can be executed by a processor core 200 of a data processing system 100 according to one embodiment. The purpose of the instruction sequence 400 is to: (a) disable a translation entry (e.g., PTE 222) in the page frame table 220 so that the translation entry is not reloaded by any MMU 308 of the data processing system 100, (b) invalidate any copies of the translation entry (or other translation entries that translate the same effective address as the translation entry) cached by any processor core 200 in the data processing system 100, and (c) drain all outstanding memory access requests that depend on the old translation entry before the effective address is reallocated. If the translation is updated before a store request that depends on the old translation entry is drained, the store request may corrupt the memory page identified by the old translation entry. Similarly, if a load request that depends on the old translation entry and misses the L1 cache 302 is not satisfied before the translation is reallocated, the load request will read data from a different memory page than expected and thus observe data that was not expected to be visible to the load request.

[0076] Instruction sequence 400 (which may be preceded and followed by any number of instructions) begins with one or more store (ST) instructions 402. Each store instruction 402, when executed, causes a store request to be generated that, when propagated to the associated system memory 108, marks the target PTE 222 in the page frame table 220 as invalid. Once a store request has marked a PTE 222 as invalid in the page frame table 220, the MMU 308 will no longer load invalidated translations from the page frame table 220.

[0077] One or more store instructions 402 in the instruction sequence 400 are followed by a heavy-weight synchronization (i.e., HWSYNC) instruction 404, which is a barrier that ensures that a subsequent TLBIE instruction 406 is not reordered by the processor core 200 such that the TLBIE instruction 406 executes before any store instruction 402. Thus, the HWSYNC instruction 404 ensures that if the processor core 200 reloads the PTE 222 from the page frame table 220 after the TLBIE instruction 406 invalidates the cached copy of the PTE 222, the processor core 200 is guaranteed to have observed the invalidation caused by the store instruction 402 and, therefore, will not use the target PTE 222 or reload the target PTE 222 into the translation structure 310 until the effective address translated by the target PTE 222 is deallocated and set valid.

[0078] Following the HWSYNC instruction 404 in the instruction sequence 400 is at least one TLBIE instruction 406 that, when executed, generates a corresponding TLBIE request that invalidates any translation entries that translate the target effective address of the TLBIE request in all translation structures 310 throughout the data processing system 100. In the instruction sequence 400, the one or more TLBIE instructions 406 are followed by a translation synchronization (i.e., TSYNC) instruction 408 that ensures that the TLBIE request generated by executing the TLBIE instruction 406 has completed invalidating all translations of the target effective address in all translation structures 310 throughout the data processing system 100 and that all previous memory access requests that depended on the now-invalidated translations have been flushed before the execution thread proceeds to subsequent instructions.

[0079] Instruction sequence 400 ends with a second HWSYNC instruction 410, which implements a barrier that prevents any memory-reference instructions following HWSYNC instruction 410 in program order from executing until TSYNC instruction 406 completes its processing. In this manner, any newer memory-reference instructions requiring a translation of the target effective address requested by the TLBIE will receive the new translation rather than the old translation invalidated by the TLBIE request. It should be noted that HWSYNC instruction 410 does not have any functionality directly related to invalidation of the target PTE 222 in the page frame table, invalidation of translation entries in the translation structure 310, or the ejection of memory-reference instructions that rely on old translations.

[0080] To facilitate understanding of the invention disclosed herein, reference is made to Figure 5-10 The progression of the TLBIE instruction 406 and the resulting TLBIE request from inception to completion is described. Figure 11 and12 Also depicted is the progression of the TSYNC instruction 408 and its corresponding TSYNC request, which ensure that the invalidation requested by the TLBIE request has been completed on all snooping processor cores 200 .

[0081] First reference Figure 5 , illustrates a high-level logic flow diagram of an exemplary method for processing a translation entry invalidate (e.g., TLBIE) instruction by an initiating processor core 200 of a multi-processor data processing system 100 according to one embodiment. The illustrated processes represent processing performed within a single hardware thread, meaning that multiple of these processes can be executed simultaneously (i.e., in parallel) on a single processor core 200, and further, multiple of these processes can be executed simultaneously on various different processing cores 200 throughout the data processing system 100. As a result, multiple different address translation entries buffered in various processor cores 200 of the data processing system 100 can be invalidated concurrently by different initiating hardware threads.

[0082] The illustrated process begins at block 500 and then proceeds to block 501, which illustrates execution of a TLBIE instruction 406 in an instruction sequence 400 by an execution unit 300 of a processor core 200. Execution of the TLBIE instruction 406 determines a target effective address for which all translation entries buffered in one or more translation structures 310 throughout the data processing system 100 are to be invalidated. In response to execution of the TLBIE instruction 406, the processor core 200 halts dispatch of any additional instructions in the initiating hardware thread because Figure 3 In the exemplary embodiment, the sidecar logic 322 includes only a single sidecar 324 per thread, which means that at most one TLBIE request per thread can be active at a time. In other embodiments with multiple sidecars 324 per thread, multiple concurrently active TLBIE requests per thread can be supported.

[0083] At block 504, a TLBIE request corresponding to the TLBIE instruction 406 is generated and issued to the L1STQ 304. The TLBIE request may include, for example, a transaction type indicating the type of request (i.e., a TLBIE), an effective address for which cached translations are to be invalidated, and an indication of the initiating processor core 200 and hardware thread that issued the TLBIE request. Processing of the request in the L1STQ 304 continues, and the TLBIE request is ultimately moved from the L1 STQ 304 to the L2STQ 320 via bus 318, as shown at block 506. Processing then proceeds to block 508, which illustrates that the initiating processor core 200 continues to refrain from dispatching instructions within the initiating hardware thread until it receives a TLBCMPLT_ACK signal from the memory subsystem via bus 325, indicating that processing of the TLBIE request by the initiating processor core 200 is complete. (Generation of the TLBCMPLT_ACK signal is discussed below with reference to Figure 10 ) Note also that because instruction dispatch within the initiating thread is suspended, there is no contention for the TSYNC request corresponding to the TSYNC instruction 408 to the initiating thread's sidecar 324, since only one of the two types of requests can exist in the L2STQ 320 and sidecar logic 322 at a time for any given thread.

[0084] In response to determining at block 508 that the TLBCMPLT_ACK signal has been received, the process proceeds from block 508 to block 510, which illustrates the processor core 200 resuming instruction dispatch in the initiating thread; thus, releasing the thread at block 510 allows processing of the TSYNC instruction 408 (which is the next instruction in the instruction sequence 400) as described below with reference to Figure 11 It began as described, and thereafter, Figure 5 The process ends at block 512.

[0085] Now refer to Figure 6 , depicts a high-level logic flow diagram of an example method by which the sidecar logic 322 of the L2 cache 230 handles a translation entry invalidation (eg, TLBIE) request of a hardware thread of an attached processor core 200 , according to one embodiment. Figure 6 The process is performed on a per-thread basis.

[0086] Figure 6The process begins at block 600 and then proceeds to block 602, which illustrates the sidecar logic 322 determining whether the TLBIE request for the hardware thread of the attached processor core 200 has been loaded into the L2 STQ 320. If not, the process repeats at block 602. However, in response to determining that the TLBIE for the hardware thread of the attached processor core 200 has been loaded into the L2 STQ 320, the sidecar logic 322 removes the TLBIE request from the L2 STQ 320 and moves the TLBIE request to the sidecar 324 corresponding to the initiating thread via the interface 321 (block 604). Removing the TLBIE request from the L2 STQ 320 ensures that no deadlock occurs due to the L2 STQ 320 being unable to receive incoming requests from the associated processor core 200, and enables such requests to flow through the L2 STQ 320.

[0087] At block 606, the sidecar 324 participates in a consensus protocol via the interface 326 and the local interconnect 114 to ensure that one (and only one) TSN engine 346 in each L2 cache 230 receives its TLBIE request. In some embodiments, the consensus protocol may be conventional. Figure 18-29 In other embodiments described, the consensus protocol used by the sidecar 324 to distribute TLBIE requests utilizes the central request agent 120 to manage the distribution of TLBIE requests to the L2 cache 230. After block 606, the process returns to block 602, which has already been described.

[0088] Now refer to Figure 7 , shows a high-level logic flow diagram of an exemplary method for processing TLBIE requests and TSYNC requests by the TSN engine 346 according to one embodiment. The illustrated process is performed independently and simultaneously for each TSN engine 346.

[0089] The process begins at block 700 and proceeds to blocks 702 and 720. Block 702 and subsequent block 704 illustrate that, in response to receiving notification of a TLBIE request via the consensus protocol, the TSN engine 346 buffers the TLBIE request and assumes the TLBIE_active state. Figure 6 The block 606 of the TLBIE request is broadcasted to the L2 cache 230 of the initiating processor core 200 and the L2 caches 230 of all other processor cores 200 of the data processing system 100 through the system fabric 110, 114. The TLBIE request is received by the L2 cache 230 via the bus 329, processed by the dispatch / response logic 336, and then assigned to the TSN engine 346. The TSN engine 346, having assumed the TLBIE_active state, notifies the associated arbiter 348 that the TLBIE request is ready to be processed, as described below with reference to Figure 8 802 is further described.

[0090] Block 706 illustrates that the TSN engine 346 remains in the TLBIE_active state until processing of the TLBIE request by the associated processor core 200 (i.e., invalidating the associated translation entries in the one or more translation structures 310 and draining the associated memory reference request from the processor core 200) is complete, as indicated by receipt of a TLBCMPLT_ACK signal via the bus 330. In response to receiving the TLBCMPLT_ACK signal, the TLBIE_active state is reset and the TSN engine 346 is released for reallocation (block 708). Thereafter, Figure 7 The process returns from block 708 to block 702, which has already been described.

[0091] Referring now to blocks 720-724, the TSN engine 346 determines at block 720 whether it is in the TLBIE_active state established at block 704. If not, the process repeats at block 720. However, if the TSN engine 346 is in the TLBIE_active state established at block 704, the TSN engine 346 monitors to determine whether a TSYNC request has been detected for the initiating hardware thread of its TLBIE request (block 722). If no TSYNC request is detected, the process continues to repeat at blocks 720-722. However, in response to detecting a TSYNC request for the initiating hardware thread of its TLBIE request while the TSN engine 346 is in the TLBIE_active state, the TSN engine 346 provides a retry consistency response via the system structure 110, 114, as shown in block 724. As described below with reference to Figure 12 As discussed in block 1208 of FIGURE 12, a retry coherency response by any TSN engine 346 processing a TLBIE request for the initiating hardware thread forces the TSYNC request to be reissued by the source L2 cache 230 and prevents the initiating hardware thread from proceeding to the HWSYNC instruction 410 until the TSYNC request completes without a retry coherency response. The TSYNC request completes without a retry coherency response when all processor cores 200 except the initiating processor core 200 have completed their processing of the TLBIE request. (The initiating processor core 200 does not issue a TSYNC request until it has completed processing the TLBIE request because instruction dispatch is suspended for processing the TLBIE request, as described above with reference to FIGURE 12.) Figure 5 (as discussed in block 508 of .)

[0092] Now refer to Figure 8, shows a high-level logic flow diagram of an example method for processing a TLBIE request by the arbiter 348 of the L2 cache 230 according to one embodiment. The process begins at block 800 and then proceeds to block 802, which illustrates the arbiter 348 determining whether any of its TSN machines 346 are in the TLBIE_active state. If not, then Figure 8 The process repeats at block 802. However, in response to determining that one or more of its TSN machines 346 are in the TLBIE_active state, the arbiter 348 selects a TSN machine 346 in the TLBIE_active state whose TLBIE request has not been previously forwarded, and sends its TLBIE request to the translation sequencer 312 of the attached processor core 200 via the bus 350 (block 804). To avoid deadlock, the translation sequencer 312 is configured to accept TLBIE requests within a fixed time, rather than arbitrarily delaying acceptance of TLBIE requests.

[0093] From block 804, the process proceeds to block 806, which depicts the arbiter 348 awaiting receipt of a TLBCMPLT_ACK message indicating that the attached processor core 200 has invalidated one or more relevant translation entries in the one or more translation structures 310 in response to the TLBIE request and has drained any associated memory reference requests that may have had their target addresses translated by the invalidated translation entries. Thus, at block 806, the arbiter 348 is awaiting a TLBCMPLT_ACK message, similar to both the initiating thread (block 508) and the TSN engine 346 in each L2 cache 230 (block 706). In response to receiving the TLBCMPLT_ACK message at block 806, the process returns to block 802, which has already been described. It should be noted that by the time the process returns to block 802, the previously selected TSN engine 346 will no longer be in the TLBIE_active state for the TLBIE request that has been processed, as the TLBIE_active state will have been reset, as indicated by blocks 706 and 708, before the process returns to block 802.

[0094] Figure 8 The process (particularly blocks 802 and 806) ensures that the processor core 200 processes only one TLBIE request at a time. The serial processing of TLBIE requests by the processor core 200 eliminates the need to tag the TLBCMPLT_ACK message to associate the TLBCMPLT_ACK message with the TLBIE request and simplifies the instruction tagging mechanism, as described below with reference to Figure 9 As discussed, however, those skilled in the art will recognize that in other embodiments, processor core 200 may be configured to service multiple TLBIE requests simultaneously, which has some additional complexity.

[0095] Now refer to Figure 9 , shows a high-level logic flow diagram of an exemplary method for initiating or snooping the translation sequencer 312 of the processor core 200 to handle a TLBIE request according to one embodiment. Figure 9 The process shown begins at block 900 and proceeds to block 902, which illustrates the translation sequencer 312 waiting to receive a TLBIE request forwarded by the arbiter 348, as described above with reference to FIG. Figure 8 In response to receiving the TLBIE request, the translation sequencer 312 invalidates one or more translation entries (e.g., PTEs or other translation entries) in the one or more translation structures 310 that translate the target effective address of the TLBIE request (block 904). Additionally, at block 906, the translation sequencer 312 marks all memory reference requests to be drained from the processor core 200.

[0096] In a less precise embodiment, at block 906, the translation sequencer 312 marks all memory reference requests for all hardware threads in the processor core 200 that have had their target addresses translated, provided that any such memory reference request may have had its target address translated by one or more translation entries invalidated by the TLBIE request received at block 902. Thus, in this embodiment, the marked memory reference requests would include all store requests in the L1 STQ 304 and all load requests in the LMQ 306. This embodiment advantageously eliminates the need to implement comparators for all entries of the L1 STQ 304 and LMQ 306, but may result in higher latency due to long drain times.

[0097] A more refined embodiment implements comparators for all entries of the L1 STQ 304 and LMQ 306. In this embodiment, each comparator compares a subset of the effective address bits specified by the TLBIE request (and not translated by the MMU 308) with the corresponding real address bits of the target real address specified in the associated entry of the L1 STQ 304 or LMQ 306. Only memory reference requests for which a comparator detects a match are flagged by the translation sequencer 312. Thus, this more refined embodiment reduces the number of flagged memory access requests at the expense of additional comparators.

[0098] In some implementations of the less precise and more precise tagging embodiments, the tags applied by the translation sequencer 312 are applied only to requests within the processor core 200 and persist only until the tagged requests are evicted from the processor core 200. In such implementations, the L2 cache 230 can revert to the pessimistic assumption that all store requests in progress in the L2 cache 230 may have had their addresses translated by the translation entry invalidated by the TLBIE request and force all such store requests to be evicted before processing the store request with the new translation of the target effective address of the TLBIE request. In other implementations, the more precise tagging applied by the translation sequencer 312 can also be extended to store requests in progress in the L2 cache 230.

[0099] Figure 9 The process proceeds from block 906 to block 908, which illustrates the translation sequencer 312 waiting for the requests marked at block 906 to be drained from the processor core 200. Specifically, the translation sequencer 312 waits until all load requests marked at block 906 have had their request data returned to the processor core 200 and all store requests marked at block 906 have been issued to the L2 STQ 320. In response to all marked requests being drained from the processor core 200, the translation sequencer 312 inserts a TLBCMPLT request into the L2 STQ 320 to indicate that the translation sequencer 312 is complete in servicing the TLBIE request (block 910). Thereafter, Figure 9 The process ends at box 912.

[0100] Now refer to Figure 10 , depicts a high-level logical flow diagram of an example method by which the L2 STQ 320 processes a TLB CMPLT request, according to one embodiment. Figure 10 The process begins at block 1000 and proceeds to block 1002, which illustrates the L2 STQ 320 receiving and queuing in one of its entries a TLBCMPLT request issued by its associated processor core 200, as described above with reference to FIG. Figure 9As shown in block 910 of FIG. 1004 , after receiving the TLBCMPLT request, the L2 STQ 320 waits until all older store requests from all hardware threads are drained from the L2 STQ 320. Once all older store requests have been drained from the L2 STQ 320, the process proceeds from block 1004 to block 1006, which illustrates that the L2 STQ 320 sends a TLBCMPLT_ACK signal via the bus 330 to the TSN engine 346 that issued the TLBIE request and to the arbiter 348, which are waiting for confirmation of the completion of processing the TLBIE request, as described above with reference to blocks 706 and 806.

[0101] At block 1008, the L2 STQ 320 determines whether the subordinate processor core 200 is the initiating processor core of the TLBIE request (whose completion is signaled by the TLBCMPLT request), for example, by examining the thread identification information in the TLBCMPLT request. If not (meaning that the process is being executed at the L2 cache 230 associated with the snooping processing core 200), then processing of the TLIBIE request is complete, and the L2 STQ 320 removes the TLBCMPLT request from the L2 STQ 320 (block 1014). Thereafter, the process ends at block 1016.

[0102] On the other hand, if the L2 cache 320 determines at block 1008 that its subordinate processor core 200 is the originating processor core 200 of the TLBIE request buffered in the sidecar logic 322, then the process proceeds from block 1008 to block 1009, which illustrates the L2STQ 320 issuing a TLBCMPLT_ACK signal to the sidecar logic 322 via bus 330. In response to receiving the TLBCMPLT_ACK signal, the sidecar logic 322 issues a TLBCMPLT_ACK signal to the subordinate processor core 200 via bus 325. As described above with reference to Figure 5 As described in block 508 of FIG. 5 , the receipt of the TLBCMPLT_ACK signal releases the initiating thread of the processor core 200 to resume the new instruction (ie, the TSYNC instruction 408, whose behavior is referenced Figure 11 The relevant sidecar 324 then removes the completed TLBIE request (block 1012), and the process moves to blocks 1014 and 1016, which have already been described.

[0103] Now refer to Figure 11 , illustrates a high-level logic flow diagram of an exemplary method for processor core 200 to process a transition synchronization (eg, TSYNC) instruction according to one embodiment.

[0104] The illustrated process begins at block 1100 and then proceeds to block 1101, which illustrates execution unit 300 of processor core 200 executing TSYNC instruction 408 in instruction sequence 400. In response to execution of TSYNC instruction 408, processor core 200 halts dispatch of any subsequent instructions in the hardware thread (block 1102). As described above, dispatch is halted because Figure 3 In the exemplary embodiment, sidecar logic 322 includes only a single sidecar 324 for each hardware thread of processor core 200 , which means that at most one TLBIE or TSYNC request per thread can be active at a time.

[0105] At block 1104, a TSYNC request corresponding to the TSYNC instruction 408 is generated and issued to the L1STQ 304. The TSYNC request may include, for example, a transaction type indicating the type of request (i.e., TSYNC) and an indication of the initiating processor core 200 and hardware thread that issued the TSYNC request. Processing of the request in the L1 STQ 304 continues, and the TSYNC request is eventually moved from the L1 STQ 304 to the L2 STQ 320 via bus 318, as indicated at block 1106. The process then proceeds to block 1108, which illustrates that the initiating processor core 200 continues to refrain from dispatching instructions within the initiating hardware thread until it receives a TSYNC_ACK signal from the memory subsystem via bus 325 indicating that processing of the TSYNC request by the initiating processor core 200 is complete. (Generation of the TSYNC_ACK signal is discussed below with reference to Figure 12 ) Note again that because instruction dispatch within the initiating thread is stalled, there can be no contention for another TLBIE request to the initiating hardware thread's sidecar 324, since only one of the two requests can exist at a time in the L2 STQ 320 and sidecar logic 322 for any given thread.

[0106] In response to determining at block 1108 that the TSYNC_ACK signal has been received, the process proceeds to block 1110, which illustrates processor core 200 resuming instruction dispatch in the initiating thread; thus, releasing the thread at block 1110 allows processing of the HWSYNC instruction 410 (which is the next instruction in instruction sequence 400) to begin. Thereafter, Figure 11 The process ends at box 1112.

[0107] Now refer to Figure 12, depicts a high-level logic flow diagram of an exemplary method for the sidecar logic 324 to process a TSYNC request according to one embodiment. The process begins at block 1200 and then proceeds to block 1202, which depicts the sidecar logic 324 monitoring for notification that a TSYNC request has been queued in the L2 STQ 320 via interface 321. In response to receiving notification that a TSYNC request has been queued in the L2 STQ 320 via interface 321, the sidecar logic 322 moves the TSYNC request to the sidecar 324 that initiated the hardware thread via interface 321 (block 1204). In response to receiving the TSYNC request, the sidecar 324 issues the TSYNC request on the system fabric 110, 114 via interface 326 (block 1206) and then monitors the coherence responses to the TSYNC request to determine if any TSN machine 346 has provided a retry coherence response, as previously described with respect to Figure 7 724 (block 1208). As described above, if the TSN engine is still in the TLBIE_active state and is waiting for its snooping processor core 200 to complete processing of a previous TLBIE request for the same initiating processor core 200 and hardware thread, the TSN engine 346 provides a retry coherence response. It can be noted that by the time the TSYNC request is issued, the TSN engine 346 of the issuing processing unit will no longer be in the TLBIE_active state and will not issue a retry coherence response because the TLBCMPLT_ACK signal reset the TSN engine 346 of the issuing processor core to the inactive state at block 1006 before the TLBCMPLT_ACK signal was issued to the initiating processor core 200 at block 1010. The receipt of the TLBCMPLT_ACK signal by the processor core 200 causes the initiating processor core 200 to resume dispatching instructions after the TLBIE instruction 406 and thus execute the TSYNC instruction 408 to generate the TSYNC request. However, the initiating processor core 200 may complete processing the TLBIE request long before the snooping processor core 200 has completed invalidating its translation entries and draining memory-reference instructions that are marked as dependent on or may depend on the invalidated translation entries. Therefore, the TSYNC request ensures that the invalidation of translation entries at the snooping processor core 200 and the draining of memory-reference instructions that depend on the invalidated translation entries are completed before the initiating processor core 200 executes the HWSYNC instruction 410.

[0108] Once all snooping processor cores 200 have completed their processing of the TLBIE request, the TSYNC request will eventually complete without a retry coherence response. In response to the TSYNC request completing without a retry coherence response at block 1208, the sidecar 324 issues a TSYNC_ACK signal to the initiating processor core 200 via bus 325 (block 1210). As described above with reference to block 1108, in response to receiving the TSYNC_ACK signal, the initiating processor core 200 executes the HWSYNC instruction 410, which completes the ordering requirement of the initiating thread with respect to the newer memory-referenced instruction. After block 1210, the sidecar 324 removes the TSYNC request (block 1212), and the process returns to block 1202, which has already been described.

[0109] Now reference Figure 5-12 Describes in detail Figure 4A The instruction sequence 400 and associated processing are now referred to Figure 4B , which illustrates an alternative code sequence 420 that reduces the number of instructions (particularly synchronization instructions) in a translation invalidation sequence. As shown, the instruction sequence 420 includes one or more store instructions 422 that invalidate the PTE 222 in the page frame table 220, an HWSYNC instruction 424, and one or more TLBIE instructions 426 that invalidate the cached translation entries for the specified effective address in all processor cores 200. Thus, instructions 422-426 correspond to Figure 4A The instruction sequence 420 further includes a PTESYNC instruction 430 immediately following the TLBIE instruction 426. The PTESYNC instruction 430 sets Figure 4A The work group executed by TSYNC instruction 408 and HWSYNC instruction 410 of instruction sequence 400 is synthesized into a single instruction. That is, execution of PTESYNC instruction 430 generates a PTESYNC request that is broadcast to all processing units 104 of data processing system 100 to ensure that the TLBIE request generated by TLBIE instruction 426 is completed system-wide (as is the case with the TSYNC request generated by executing TSYNC instruction 408) and to enforce instruction ordering with respect to newer memory-referencing instructions (as is the case with the HWSYNC request generated by executing HWSYNC instruction 410).

[0110] Given the similarity between instruction sequences 420 and 400, the processing of instruction sequence 420 is similar to Figure 5-12 The processing of the instruction sequence 400 given in FIG. 4 is the same, except for the processing associated with the PTESYNC request generated by executing the PTESYNC instruction 430, which will be referred to below. Figure 13-15 describe.

[0111] Now refer to Figure 13 , illustrates a high-level logic flow diagram of an exemplary method for processing core 200 to handle a page table synchronization (e.g., PTESYNC) instruction 430 according to one embodiment. As described above, PTESYNC instruction 430 and the PTESYNC request generated by its execution have two functions, namely, to ensure that the TLBIE request generated by TLBIE instruction 426 is completed system-wide and to enforce instruction ordering with respect to newer memory-referencing instructions.

[0112] The illustrated process begins at block 1300 and then proceeds to block 1301, which illustrates the processor core 200 generating a PTESYNC request by executing a PTESYNC instruction 430 in an instruction sequence 420 in an execution unit(s) 300. The PTESYNC request may include, for example, a transaction type indicating the type of request (i.e., PTESYNC) and an indication of the initiating processor core 200 and hardware thread issuing the PTESYNC request. In response to executing the PTESYNC instruction 430, the processor core 200 halts dispatching any newer instructions in the initiating hardware thread (block 1302). As described above, dispatching is halted because Figure 3 In the exemplary embodiment of , sidecar logic 322 includes only a single sidecar 324 for each hardware thread of processor core 200 , which means that in this embodiment, at most one TLBIE or PTESYNC request per thread at a time can be active.

[0113] After block 1302, Figure 13 The process proceeds in parallel to block 1303 and blocks 1304-1312. Block 1303 represents the initiating processor core 200 performing the load ordering function for the PTESYNC request by waiting for all appropriate older load requests from all hardware threads (i.e., those hardware threads architecturally required by HWSYNC in order to receive the data they requested before processing of the HWSYNC request is completed) to be drained from LMQ 306. By waiting for these load requests to be satisfied at block 1303, it is guaranteed that the set of load requests identified at block 906 will receive data from the correct memory page (even if the target address was on the reallocated memory page) rather than the reallocated memory page.

[0114] In parallel with block 1303, the processor core 200 also issues a PTESYNC request corresponding to the PTESYNC instruction 430 to the L1 STQ 304 (block 1304). From block 1304, the process proceeds to block 1308, which illustrates the processor core 200 performing a store ordering function for the PTESYNC request by waiting until all appropriate older store requests for all hardware threads (i.e., those hardware threads architecturally required for HWSYNC to be drained from the L1 STQ 304) are drained from the L1 STQ 304. Once the store ordering performed at block 1308 is complete, the PTESYNC request is issued from the L1 STQ 304 to the L2 STQ 320 via bus 318, as shown in block 1310.

[0115] The process then proceeds from block 1310 to block 1312, which illustrates the initiating processor core 200 monitoring to detect receipt of a PTESYNC_ACK signal from the storage subsystem via bus 325, indicating completion of processing of the PTESYNC request by the initiating processor core 200. (See below for details.) Figure 14 (Note that block 1410 depicts the generation of the PTESYNC_ACK signal.) Note again that because instruction dispatch within the initiating hardware thread remains stalled, there is no contention for another TLBIE request to the initiating hardware thread's sidecar 324, since for any given thread, only one of a TLBIE request or a PTESYNC request can exist in the L2 STQ 320 and sidecar logic 322 at a time.

[0116] Only in response to affirmative determinations at both blocks 1303 and 1312, Figure 13 The process proceeds to block 1314, which illustrates that the processor core 200 resumes instruction dispatch in the initiating thread; therefore, the thread release at block 1314 allows processing of instructions following the PTESYNC instruction 430 to begin. Thereafter, Figure 13 The process ends at box 1316.

[0117] Now refer to Figure 14 , depicts a high-level logic flow diagram of an example method for handling PTESYNC requests by the L2 STQ 320 and sidecar logic 322 of the processing unit 104, according to one embodiment. Figure 14 The process begins at block 1400 and proceeds to block 1402, which depicts the L2 STQ 320 monitoring the receipt of a PTESYNC request from the L1 STQ 304, as described above with reference to FIG. Figure 13 As described in block 1310 of Figure 4BIn the second embodiment, in response to receipt of a PTESYNC request, the L2STQ 320 and the sidecar logic 324 cooperate to perform two functions, namely, (1) memory ordering of memory requests within the L2 STQ 320, and (2) ensuring completion of TLBIE requests at all other processing cores 200. Figure 14 In the embodiment shown, these two functions are executed in parallel along the two paths shown at blocks 1403, 1405 and blocks 1404, 1406, and 1408, respectively. In an alternative embodiment, these functions may instead be serialized by first performing the ordering functions shown at blocks 1403 and 1405 and then ensuring completion of the TLBIE requests at blocks 1404, 1406, and 1408. (It should be noted that attempting to serialize the ordering of these functions by ensuring completion of the TLBIE requests before performing the store ordering can result in deadlock.)

[0118] Referring now to blocks 1403-1405, the L2 STQ 320 performs memory sorting for PTESYNC requests by ensuring that all appropriately older memory requests within the L2 STQ 320 have been drained from the L2 STQ 320. The set of memory requests sorted at block 1403 includes a first subset that may have had their target addresses translated by translation entries invalidated by earlier TLBIE requests. This first subset corresponds to those identified at block 906. Additionally, the set of memory requests sorted at block 1403 includes a second subset that includes those architecturally defined memory requests to be sorted by HWSYNC. Once all such memory requests have been drained from the L2 STQ 320, the L2 STQ 320 removes the PTESYNC request from the L2 STQ 320 (block 1405). Removal of the PTESYNC request allows memory requests newer than the PTESYNC request to flow through the L2 STQ 320.

[0119] Referring now to block 1404, the sidecar logic 322 detects the presence of a PTESYNC request in the L2 STQ 320 and copies the PTESYNC request to the appropriate sidecar 324 via the interface 321 before removing the PTESYNC request from the L2 STQ 320 at block 1405. The process then proceeds to the loop shown at blocks 1406 and 1408, where the sidecar logic 322 continues to issue PTESYNC requests on the system fabric 110, 114 until no processor core 200 responds with a retry coherence response (i.e., until previous TLBIE requests for the same processor core and hardware thread have been completed by all snooping processor cores 200).

[0120] In response only to the completion of the two functions depicted at blocks 1403, 1405 and blocks 1404, 1406, and 1408, the process proceeds to block 1410, which illustrates the sidecar logic 322 issuing a PTESYNC_ACK signal to the attached processor core via bus 325. The sidecar logic 322 then removes the PTESYNC request from the sidecar 324 (block 1412), and the process returns to block 1402, which has already been described.

[0121] Now refer to Figure 15 , shows a high-level logic flow diagram of an exemplary method for the TSN engine 346 to process TLBIE requests, TLBCMPT_ACK signals, and PTESYNC requests according to one embodiment. As shown by the same reference numerals, except for block 1522, Figure 15 With the aforementioned Figure 7 The same. Block 1522 illustrates that while in the TLBIE_active state established at block 704, the TSN engine 346 monitors to determine whether a PTESYNC request has been detected that specifies an initiating processor core and hardware thread that matches its TLBIE request. If not, the process continues to iterate at the loop including blocks 720 and 1522. However, in response to the TSN engine 346 detecting a PTESYNC request that specifies a processor core and initiating hardware thread that matches its TLBIE request while in the TLBIE_active state, the TSN engine 346 provides a retry coherence response, as shown in block 724. As described above, a retry coherence response by any TSN engine 346 processing a TLBIE request for an initiating hardware thread forces the PTESYNC request to be retried and prevents the initiating hardware thread from executing any memory-referencing instructions newer than the PTESYNC instruction 430 until the PTESYNC request completes without a retry coherence response.

[0122] As briefly discussed above, avoiding ping-pong livelock is a design concern of data processing system interconnect architectures that allow for retry of multicast requests. In accordance with the disclosed embodiments, ping-pong livelock can be avoided by implementing a central request broker 120 and an associated communication protocol for managing the distribution of specific multicast requests subject to snooper retries. In the following discussion, reference is made to Figure 18 An example of a suitable design for the central request agent 120 is described. Figure 19-23 The flow of an exemplary multicast request (eg, TLBIE request) distributed to snoopers via central request agent 120 according to a first embodiment is described. In this first embodiment, it is assumed that the broadcast of a multicast request by initiating master device 1600 requires only a single operational occupancy on the system architecture. Figure 24-29The flow of a multicast request in the second embodiment is depicted, wherein the broadcast of a multicast request by the master device 1600 may require multiple operation occupancy periods on the system architecture.

[0123] Now refer to Figure 18 , depicts a high-level block diagram of an exemplary embodiment of central request agent 120, according to one embodiment. In this example, central request agent 120 includes dispatch logic 1802 coupled to receive multicast requests of predetermined transaction types (ttypes) on system fabric 1800 (which includes system interconnect 110 and local interconnect 114). In some embodiments, these multicast requests include TLBIE requests as described above; in other embodiments, these multicast requests may include multicast requests of alternative or additional transaction types. Non-multicast requests and multicast requests other than predetermined ttypes are ignored by dispatch logic 1802 of central request agent 120.

[0124] Central request broker 120 also includes a plurality of request forwarding machines (RFMs) 1804 that manage the forwarding of multicast requests accepted by dispatch logic 1802 to snoopers distributed within data processing system 100. In a preferred embodiment, each snooper associated with a given ttype of multicast requests distributed by central request broker 120 has a plurality of snoopers corresponding in number and identifiers to the snoopers implemented in central request broker 120. Thus, for example, central request broker 120 may implement eight (8) RFMs 1804 for handling the distribution of TLBIE requests, meaning that each L2 cache 230 implements eight TSNs 346, each TSN 346 uniquely and individually corresponding to a respective RFM in RFMs 1804 and identifiable by a common machine identifier (e.g., which may be specified in RFM ID field 1704 of the forwarded multicast request). Each RFM 1804 has an associated request buffer 1806 for buffering at least the address field 1702, ttype field 1710, and tag field 1720 of the multicast request currently assigned to be processed. In some embodiments, each RFM 1804 also has an optional associated epoch indication 1808 (e.g., a single bit) that can be used to track chronological epochs in the operation of data processing system 100.

[0125] Now turn Figure 19 , shows a high-level logic flow chart of an exemplary method for the initiating processing unit 104 or the central request agent 120 to issue a translation entry invalidation request (eg, a TLBIE request) on the system architecture according to the first embodiment. Figure 6 At block 606 , the process is performed by the sidecar logic 322 .

[0126] Figure 19 The process begins at block 1900 and proceeds to block 1902, which illustrates the master sending a request to the system architecture 1800 with Figure 17 . When the TLBIE request is issued by the sidecar logic 322 of the processing unit 104, the forwarding (F) field 1726 is reset (e.g., reset to b'0'); when issued again by the central request agent 120 on behalf of the original initiating master device, the forwarding field 1726 is set (e.g., set to b'1'). As shown in box 1904, the master device monitors the receipt of Cresp for the TLBIE request issued in box 1902. In response to receiving Cresp, the host determines at box 1906 whether Cresp indicates a retry. If so, the process returns to box 1902 and subsequent boxes, which illustrate the master device reissuing the TLBIE request on the system structure 1800. However, if at box 1906, the master device determines that Cresp does not indicate a retry (i.e., Cresp indicates success), then Figure 19 The process ends at box 1908.

[0127] As will become apparent from the following discussion, if the central request agent 120 does not have an RFM 1804 available to handle the forwarding of the TLBIE request to the final snooper (e.g., L2 cache 230), the TLBIE request initially issued on the system fabric by the sidecar logic 322 receives a Cresp indicating a retry. As long as an RFM 1804 is available within the central request agent 120 to handle the forwarding of the TLB request, a Cresp indicating a retry is not provided. However, the TLBIE request forwarded on the system fabric by the RFM 1804 of the central request agent 120 will receive a Cresp indicating a retry until all relevant snoopers (e.g., L2 cache 230) have been able to successfully allocate the TSN 346 specified by the TLBIE request to service the TLBIE request.

[0128] Now refer to Figure 20 , which is a high-level logic flow diagram of an exemplary method according to a first embodiment for a central request agent 120 to receive a translation entry invalidation request (e.g., a TLBIE request) from an originating processing unit 104. The illustrated process is performed for each TLBIE request received by the central request agent 120 via the system fabric with the forwarding (F) field 1726 reset.

[0129] Figure 20The process begins at block 2000 in response to receiving a TLBIE request on the system fabric 1800 and then proceeds to block 2004, which depicts the dispatch logic 1802 determining whether the RFM 1804 is available for allocation to process the TLBIE request. If not, the dispatch logic 1802 provides a partial response indicating a retry on the system fabric 1800 (block 2010). The retry partial response will cause the response logic 1622 to generate a Cresp indicating a retry, as described above with reference to Figure 19 As discussed in block 1906 of FIGURE 19 , this will cause the TLBIE request to be reissued on the system fabric 1800. However, if the dispatch logic 1802 determines at block 2004 that an RFM 1804 of the central request agent 120 is available for allocation to process the TLBIE request, the dispatch logic 1802 assigns the TLBIE request to the available RFM 1804, records the TLBIE request in the associated request buffer 1806, and marks the TLBIE request as completed (i.e., ready for forwarding) (block 2006). Additionally, at block 2008, the dispatch logic 1802 provides an empty partial response on the system fabric 1800, which will allow the TLBIE request to successfully complete without a retry, as described above with reference to FIGURE 19 . Figure 19 After block 2008 or block 2010, Figure 20 The process ends at block 2012.

[0130] Now refer to Figure 21 , illustrates a high-level logical flow diagram of an exemplary method by which the central request agent 120 selects a translation invalidate entry request (eg, a TLBIE request) for issuance according to a first embodiment. Figure 21 The process begins at block 2100 and then proceeds to block 2102, which illustrates the central request agent 120 determining whether any RFM 1804 has been assigned to process a TLBIE request that has completed but not yet been forwarded to the final snooper. If not, the process simply repeats at block 2102. However, if the central request agent 120 makes a positive determination at block 2102, the central request agent 120 selects one of the one or more RFMs 1804 identified at block 2102 and signals the selected RFM 1804 to utilize Figure 22 The process forwards the TLBIE request of the selected RFM 1804 on the system structure 1800 (block 2104). After block 2104, Figure 21 The process returns to block 2102 which has already been described.

[0131] Now refer to Figure 22, depicts a high-level logic flow diagram of an exemplary method for a selected RFM 1804 of a central request agent 120 to issue a translation entry invalidation request (eg, a TLBIE request) according to the first embodiment. Figure 21 Block 2104 calls the described process.

[0132] Figure 22 Processing begins at block 2200 and proceeds to block 2202, which illustrates the selected RFM 1804 modifying the tag 1720 of its buffered TLBIE request by setting the forwarding field 1726 to indicate that the request is being forwarded from the central request agent 120. In addition, the RFM 1804 preferably updates the TLBIE request to include its unique RFM identifier (ID). For example, in Figure 17 In the embodiment of the present invention, the RFM ID can be conveniently specified in the lower bits of the address field 1702, which are not used for TLBIE requests because TLBIE requests operate on logical memory pages with a minimum page size of, for example, 4kB. The selected RFM 1804 is then allocated using the previously described Figure 19 The process issues (and, if necessary, reissues) modified TLBIE requests on the system fabric until the TLBIE request is accepted by the TSN 346 in each processing unit 104 corresponding to the specified RFM ID (block 2204). Thereafter, the selected RFM 1804 is released for reallocation (block 2206), and Figure 22 The process ends at box 2208.

[0133] Now refer to Figure 23 , illustrates a high-level logic flow diagram of an exemplary method for processing unit 104 to receive a translation entry invalidation request (e.g., a TLBIE request) from central request agent 120 according to a first embodiment. The illustrated process is performed by processing unit 104 for each snooped TLBIE request having a forwarding field 1726 set to indicate central request agent 120 as the initiating master. Other requests are simply ignored by the depicted process.

[0134] Figure 23The process begins at block 2300 in response to the processing unit 104 snooping a TLBIE request issued by the central request agent 120, and then proceeds to block 2302, which illustrates the processing unit 104 determining whether the TLBIE request is the same request as the request currently being processed or the request most recently processed by the TSN 346 corresponding to the RFM ID 1704 specified in the TLBIE request. In response to a positive determination at block 2302, the process proceeds to block 2308, which will be described below. However, if the processing unit 104 makes a negative determination at block 2302, the processing unit 104 additionally determines at block 2304 whether the specified TSN 346 (i.e., the TSN 346 corresponding to the RFMID 1704) is currently active, meaning that the specified TSN 346 is still busy processing a previously received TLBIE request. In response to a positive determination at block 2304, the processing unit 104 issues a retry partial response on the system fabric 1800, which will result in the generation of a retry Cresp that will cause the central request agent 120 to reissue the TLBIE request (block 2310).

[0135] Returning to block 2304, in response to determining that the designated TSN 346 is not currently active, the processing unit 104 assigns the TLBIE request to the designated TSN 346 for processing and marks the TLBIE request as complete (i.e., ready for processing by the TSN 346) (block 2306). The steps shown at block 2306 cause the TSN 346 to set TLBIE_active (as described above in Figure 7 ), and causing the TSN 346 to forward the TLBIE request to the associated processor core 200 (as discussed above in Figure 8 (as discussed in boxes 802-804 of FIG. 1 ). Figure 23 The process then passes to block 2308, which illustrates the processing unit 104 issuing an empty partial response on the system structure 1800. After block 2308 or 2310, Figure 23 The process ends at box 2312.

[0136] exist Figure 23 (and similar Figure 28), the processing units 104 can safely provide empty partial responses to previously snooped TLBIE requests reissued on the system fabric 1800 because the central request agent 120 is configured not to send a new TLBIE request from a given RFM 1804 until the central request agent 120 receives a successful (i.e., non-retry) Cresp for the previous TLBIE request, where the successful Cresp indicates that the previous TLBIE request was accepted by all relevant processing units 104. Any TSN 346 that completes processing of a given TLBIE request earlier than the corresponding TSN 346 in the other processing units 104 will simply remain idle, waiting for the other TSNs to complete their processing of the TLBIE requests and allow the corresponding RFM 1804 to issue a new TLBIE request.

[0137] In reference Figure 19-23 In the description of the first embodiment of the present invention, it is assumed by default that multicast requests (e.g., TLBIE requests) can be sent via the system structure 1800 within a single operation occupancy period. In some embodiments, this is not the case for all multicast requests due to the relative size of the multicast requests relative to the width of the system structure 1800. However, the disclosed techniques for distributing multicast requests using a central request agent can be extended for use in such environments, as now described with reference to FIG. Figure 24 In the second embodiment, exemplary multicast requests include those requiring only a single operation occupancy period (referred to as TLBIE_OP1_only) and those requiring two operation occupancies (referred to as TLBIE_OP1 and TLBIE_OP2).

[0138] Now refer to Figure 24 , depicts a high-level logic flow diagram of an exemplary method for an initiating processing unit 104 or a central request agent 120 to issue a translation entry invalidation request (eg, a TLBIE request) according to a second embodiment. For the initiating processing unit 104, Figure 24 The process is performed by the sidecar logic 322 Figure 6 Executed at block 606.

[0139] Figure 24 The process starts at block 2400 and then proceeds to blocks 2402 and 2410 in parallel. Block 2402 shows the master device issuing a Figure 17. When issued by the sidecar logic 322 of the processing unit 104, the forwarding (F) field 1726 is reset (e.g., reset to b'0'); when again issued by the central request agent 120 on behalf of the original initiating master device, the forwarding field 1726 is set (e.g., set to b'1'). As shown in box 2404, the master device monitors the receipt of Cresp for the multicast request issued in box 2402. In response to receiving Cresp, the master device determines at box 2406 whether Cresp indicates a retry. If so, the process returns to box 2402 and subsequent boxes, which illustrate the master device reissuing the TLBIE_OP1_only and TLBIE_OP1 requests. However, if the master device determines at box 2406 that Cresp does not indicate a retry but indicates success, then Figure 24 The process continues to join point 2422.

[0140] Referring now to block 2410, the master device further determines whether the transaction type of the multicast request indicates that only a single operation occupancy period is required. In various embodiments, this information may be conveyed, for example, in the ttype field 1710 or the operation (O) field 1730. In response to a positive determination at block 2410, the process continues to connection point 2422. However, if the master device makes a negative determination at block 2410, meaning that the TLBIE request includes both a TLBIE_OP1 request and a TLBIE_OP2 request, the master device processes the multiple operation occupancy periods of the multicast request according to implementation-specific ordering requirements, if any. That is, if a given implementation requires that a TLBIE_OP1 request be issued before a corresponding TLBIE_OP2 request, as indicated by a positive determination at block 2412, the master device defers issuing the TLBIE_OP2 request until processing of the TLBIE_OP1 request has reached connection point 2422 (block 2414). However, if a given implementation does not require the issuance of a TLBIE_OP1 request before the corresponding TLBIE_OP2 request, as indicated by a negative result at block 2412, the master device issues a TLBIE_OP2 request on the system fabric independent of the issuance timing of the corresponding TLBIE_OP1 request (block 2416). As described above, when issued by sidecar logic 322, forward (F) field 1726 is reset (e.g., to b'0'), and when again issued by central request agent 120 on behalf of the original initiating master device, forward field 1726 is set (e.g., to b'1'). As indicated at block 2418, the master device monitors for receipt of Cresp for the TLBIE_OP2 request issued at block 2416. In response to receiving Cresp, at block 2420, the master device determines whether Cresp indicates a retry. If so, the process returns to block 2416 and subsequent blocks, which illustrate the master device reissuing the TLBIE_OP2 request. However, if at block 2420 the host determines that Cresp does not indicate a retry, then Figure 24 The process proceeds to the junction 2422. Once both branches of the process have reached the junction 2422, Figure 24 The process ends at box 2424.

[0141] As mentioned above Figure 19As described above, a TLBIE_OP1_only, TLBIE_OP1, or TLBIE_OP2 request initially issued on the system fabric by the sidecar logic 322 receives a Cresp indicating a retry only when the central request agent 120 does not have an RFM 1804 available to process the forwarding of the multicast request to the final snooper (e.g., L2 cache 230). As long as an RFM 1804 is available within the central request agent 120 to process the forwarding of the TLBIE request, a successful Cresp will be provided. However, multicast requests forwarded on the system fabric by the RFM 1804 of the central request agent 120 will receive a Cresp indicating a retry until all relevant snoopers (e.g., L2 cache 230) have been able to successfully receive the multicast request and have been assigned a designated TSN 346 to process the multicast request.

[0142] Now refer to Figure 25 , illustrates a high-level logic flow diagram of an exemplary method for the central request agent 120 to receive a translation entry invalidation request (e.g., a TLBIE_OP1_Only, TLBIE_OP1, or TLBIE_OP2 request) from an originating processing unit 104 according to a second embodiment. The illustrated process is performed for each TLBIE_OP1_Only, TLBIE_OP1, or TLBIE_OP2 request whose forwarding (F) field 1726 is reset that the central request agent 120 receives via the system architecture 1800. The central request agent 120 simply ignores other types of requests.

[0143] Figure 25 The process begins at block 2500 in response to central request agent 120 receiving a TLBIE_OP1_Only, TLBIE_OP1, or TLBIE_OP2 request on system architecture 1800. The process then proceeds from block 2500 to block 2502, which depicts dispatch logic 1802 determining whether the received request is a TLBIE_OP1_Only request. In response to a positive determination at block 2502, the process proceeds to block 2504 and subsequent blocks. In response to a negative determination at block 2502, meaning the received request is a TLBIE_OP1 or TLBIE_OP2 request, the process proceeds to block 2520 and subsequent blocks.

[0144] Referring to block 2504, the dispatch logic 1802 determines whether the RFM 1804 is currently available for allocation to process the TLBIE_OP1_Only request. If not, the dispatch logic 1802 provides a partial response indicating a retry on the system structure 1800 (block 2510). The retry partial response will cause the response logic 1622 to generate a Cresp indicating a retry, such as Figure 24, which will cause the TLBIE_OP1_Only request to be reissued on the system fabric 1800. However, if the dispatch logic 1802 determines at block 2504 that at least one RFM 1804 of the central request broker 120 is available for allocation to process the TLBIE_OP1_Only request, the dispatch logic 1802 assigns the TLBIE_OP1_Only request to the available RFM 1804, records the TLBIE_OP1_Only request in the associated request buffer 1806, and marks the TLBIE_OP1_Only request as completed (i.e., ready for forwarding) (block 2506). Additionally, at block 2508, the dispatch logic 1802 provides an empty partial response on the system fabric 1800, which will allow the TLBIE_OP1_Only request to complete without a retry, as described above with reference to FIG. Figure 24 After box 2508 or box 2510, Figure 25 The process ends at box 2530.

[0145] Referring now to block 2520, the dispatch logic 1802 determines whether any RFM 1804 holds an outstanding TLBIE_OP2 or TLBIE_OP1 in its request buffer 1806 corresponding to the TLBIE_OP1 or TLBIE_OP2 request, respectively, received via the system architecture 1800. The corresponding request will have a matching tag 1720, except that the operation type field 1730 is set to indicate a different type of request (e.g., TLBIE_OP1 or TLBIE_OP2). In response to a positive determination at block 2520, the dispatch logic 1802 assigns the received TLBIE_OP1 or TLBIE_OP2 request received at block 2502 to the RFM 1804 that buffered the corresponding TLBIE_OP2 or TLBIE_OP1 request and marks both the TLBIE_OP1 and TLBIE_OP2 requests as completed (block 2522). Thereafter, the process proceeds to block 2529, described below.

[0146] Returning to block 2520, in response to determining that no RFM 1804 is currently buffering an outstanding TLBIE_OP2 or TLBIE_OP1 corresponding to the TLBIE_OP1 or TLBIE_OP2 request received via the system architecture 1800, the dispatch logic 1802 determines at block 2524 whether the RFM 1804 is available for allocation to process the TLBIE_OP1 or TLBIE_OP2 request. If not, the dispatch logic 1802 provides a partial response indicating a retry on the system architecture 1800 (block 2526). The retry partial response will cause the response logic 1622 to generate a Cresp indicating a retry, such as Figure 242406 or 2420, which will cause the TLBIE_OP1 or TLBIE_OP2 request to be reissued on the system fabric 1800. However, if the dispatch logic 1802 determines at block 2524 that the RFM 1804 of the central request agent 120 is available for allocation to process the received TLBIE_OP1 or TLBIE_OP2 request, the dispatch logic 1802 assigns the TLBIE_OP1 or TLBIE_OP2 request to the available RFM 1804 and records the request in the associated request buffer 1806 (block 2528). Additionally, at block 2529, the dispatch logic 1802 provides an empty partial response on the system fabric 1800, which will allow the TLBIE_OP1 or TLBIE_OP2 request to complete without a retry, as described above with reference to FIG. Figure 24 After box 2526 or box 2529, Figure 25 The process ends at box 2530.

[0147] Now refer to Figure 26 , depicts a high-level logical flow diagram of an exemplary method by which the central request agent 120 selects a translation entry invalidation request (eg, a TLBIE_OP1_Only or a TLBIE_OP1 / TLBIE_OP2 request) for issuance according to a second embodiment. Figure 26 The process begins at block 2600 and then proceeds to block 2602, which illustrates the central request agent 120 determining whether any RFM 1804 has been assigned to handle a multicast request that has completed but not yet been forwarded to the final snooper (e.g., a TLBIE_OP1_Only or TLBIE_OP1 / TLBIE_OP2 request). If not, the process simply repeats at block 2602. However, if the central request agent 120 makes a positive determination at block 2602, the central request agent 120 selects one of the one or more RFMs 1804 identified at block 2602 and, based on the RFM 1804, selects the RFM 1804 that has been assigned to handle the multicast request that has completed but not yet been forwarded to the final snooper (e.g., a TLBIE_OP1_Only or TLBIE_OP1 / TLBIE_OP2 request). Figure 27 The process of signaling the selected RFM 1804 to forward its one or more multicast requests on the system structure (block 2604). After block 2604, Figure 26 The process returns to block 2602 which has already been described.

[0148] Now refer to Figure 27 , shows a high-level logic flow chart of an exemplary method for a central request agent to issue a conversion entry invalidation request on a system architecture 1800 according to the second embodiment. Figure 26 Block 2604 calls the process shown.

[0149] Figure 27The process begins at block 2700 and then proceeds to block 2702, which illustrates the selected RFM 1804 modifying the tag 1720 of a buffered TLBIE_OP1_Only, TLBIE_OP1, or TLBIE_OP2 request to be forwarded by setting the forwarding field 1726 to indicate the central request agent 120 as the source. In addition, the RFM 1804 preferably updates the TLBIE_OP1_Only, TLBIE_OP1, or TLBIE_OP2 request to include its unique RFMID 1704, for example, in the lower bits of the address field 1702. If the age indication 1808 is implemented in the central request agent 120, the selected RFM 1804 also configures the age field 1728 of the tag 1720 with the value of its associated age indication 1808. Following the tag modification shown in block 2702, the selected RFM 1804 then utilizes the previously described Figure 24 The process issues (and, if necessary, reissues) modified multicast requests on the system structure 1800 until the multicast request is accepted by the TSN 346 in each processing unit 104 corresponding to the specified RFM ID 1704 (block 2704). At optional block 2706, the selected RFM 1804 switches the value of its associated time period indication 1808. Updating the time period indication 1808 at least between issuing each pair of TLBIE_OP1 and TLBIE_OP2 requests allows the processing unit 104, which may be powered off and on asynchronously with the issuance of the multicast request, to detect whether the snooped TLBIE_OP1 and TLBIE_OP2 form a corresponding pair. After block 2704 or optional block 2706, the selected RFM 1804, if any, is released for reallocation (block 2708), and Figure 27 The process ends at box 2710.

[0150] Now refer to Figures 28-29 , depicts a high-level logic flow diagram of an exemplary method for processing a translation entry invalidation request (e.g., TLBIE_OP1_Only, TLBIE_OP1, or TLBIE_OP2) received by a processing unit 104 according to a second embodiment. The illustrated process is performed by each processing unit 104 for each snooped TLBIE_OP1_Only, TLBIE_OP1, or TLBIE_OP2 request that has a forwarding field 1726 set in its tag 1720 to indicate that the central request agent 120 is the initiating master. Other requests are Figures 28-29 The process is simply ignored.

[0151] Figure 28The process begins at block 2800 in response to processing unit 104 snooping a TLBIE_OP1_Only, TLBIE_OP1, or TLBIE_OP2 request issued by central request agent 120. The process then proceeds to block 2801, which illustrates processing unit 104 determining whether the snooped request is a TLBIE_OP1_Only request. If so, the process proceeds to Figure 28 If not, meaning the snooped request is a TLBIE_OP1 or TLBIE_OP2 request, the process reaches the page connector A through Figure 29 Box 2902 and subsequent boxes.

[0152] Referring now to block 2802, processing unit 104 determines whether the snooped TLBIE_OP1_Only request is the same request as the one currently being processed or the request most recently processed by TSN 346 corresponding to RFM ID 1704 specified in the snooped request. In response to a positive determination at block 2802, the process proceeds to block 2808, described below. However, if processing unit 104 makes a negative determination at block 2802, processing unit 104 additionally determines at block 2804 whether TSN 346 corresponding to specified RFM ID 1704 is currently active, meaning that specified TSN 346 is still busy processing a previously snooped multicast request. In response to a positive determination at block 2804, processing unit 104 issues a retry partial response on system fabric 1800, which results in the generation of a retry Cresp, which in turn causes central request agent 120 to reissue the TLBIE_OP1_Only request (block 2810).

[0153] Returning to block 2804, in response to determining that the TSN 346 corresponding to the RFM ID 1704 specified in the snooped TLBIE_OP1_Only request is not currently active, the processing unit 104 assigns the snooped TLBIE_OP1_Only request to the specified TSN 346 for processing and marks the TLBIE_OP1_Only request as complete (i.e., ready for processing by the TSN 346) (block 2806). The steps shown at block 2806 cause the TSN 346 to set the TLBIE_active bit (as described above in Figure 7 ), and causing the TLBIE_OP1_Only request to be forwarded to the associated processor core 200 (as discussed above in Figure 8 (as discussed in boxes 802-804 of FIG. 1 ). Figure 28The process then passes to block 2808, which illustrates the processing unit 104 issuing an empty partial response on the system structure 1800. After block 2808 or 2810, Figure 28 The process ends at box 2812.

[0154] Now refer to Figure 29 At block 2902, the processing unit 104 determines whether the snooped TLBIE_OP1 or TLBIE_OP2 request is the same request as the request currently being processed or the request most recently processed by the TSN 346 corresponding to the RFM ID 1704 specified in the snooped request. If the use of epochs is implemented, the determination made at block 2902 includes determining whether the epoch field 1728 of the snooped request matches the epoch field 1728 of the currently processed or immediately previously processed request. In response to a positive determination at block 2902, the process proceeds to block 2920, described below. However, if the processing unit 104 makes a negative determination at block 2902, the snooping processing unit 104 optionally additionally determines at block 2904 whether the specified TSN 346 is currently assigned an outstanding request that specifies an epoch different from the epoch specified in the epoch field 1728 of the snooped request. In response to a positive determination at block 2904, the processing unit 104 cancels the outstanding prior multicast request and discards the outstanding prior multicast request from the designated TSN 346 (block 2906). Following a negative determination at block 2906 or block 2904, the process proceeds to block 2910.

[0155] Block 2910 depicts the processing unit 104 determining whether the TSN 346 corresponding to the specified RFM ID 1704 is currently assigned an outstanding TLBIE_OP2 or TLBIE_OP1 corresponding to the TLBIE_OP1 or TLBIE_OP2 request, respectively, that was snooped via the system fabric 1800. The corresponding request would have a matching tag 1720, except that the operation type field 1730 would be set to indicate a request of a different type (e.g., TLBIE_OP1 or TLBIE_OP2). In response to a positive determination at block 2910, the processing unit 104 assigns the snooped TLBIE_OP2 or TLBIE_OP1 to the specified TSN 346 and marks the TLBIE_OP1 and TLBIE_OP2 requests as completed and thus ready for processing by the TSN 346 (block 2912). The TLBIE_OP1 and TLBIE_OP2 requests are marked as shown at block 2912 so that the TSN 346 sets the TLBIE_active setting (as described above in Figure 7), and causing the TLBIE_OP1 and TLBIE_OP2 requests to be forwarded to the associated processor core 200 (as discussed above in Figure 8 The process then moves to block 2920, described below.

[0156] Returning to block 2910, in response to a negative determination, the process proceeds to block 2914, which illustrates the processing unit 104 determining whether the designated TSN 346 is active, meaning that the designated TSN 346 is still busy processing a previously snooped request. In response to a positive determination at block 2914, the processing unit 104 issues a retry partial response on the system fabric 1800, which results in the generation of a retry Cresp, which results in the central request agent 120 reissuing the TLBIE_OP1 or TLBIE_OP2 request, as described above with reference to Figure 24 as described in boxes 2406 and 2420 (box 2918).

[0157] Returning to block 2914, in response to determining that the designated TSN 346 is not currently active, the processing unit 104 assigns the snooped TLBIE_OP1 or TLBIE_OP2 request to the designated TSN 346 for processing and marks the request as incomplete (block 2916). Figure 29 The process proceeds to block 2920, which illustrates processing unit 104 issuing an empty partial response on system architecture 1800. After block 2918 or 2920, Figure 29 The process ends at box 2922.

[0158] refer to Figure 30, depicts a block diagram of an exemplary design flow 3000 used, for example, in semiconductor IC logic design, simulation, testing, layout, and manufacturing. Design flow 700 includes processes, machines, and / or mechanisms for processing a design structure or device to generate a logical or other functionally equivalent representation of the design structure and / or device described above. The design structure processed and / or generated by design flow 3000 may be encoded on a machine-readable transmission or storage medium to include data and / or instructions that, when executed or otherwise processed on a data processing system, generate a logical, structural, mechanical, or other functionally equivalent representation of a hardware component, circuit, device, or system. Machines include, but are not limited to, any machine used in an IC design process, such as designing, manufacturing, or simulating a circuit, component, device, or system. For example, machines may include: a photolithography machine, a machine and / or device for generating masks (e.g., an electron beam writer), a computer or device for simulating a design structure, any device used in a manufacturing or testing process, or any machine for programming a functionally equivalent representation of a design structure into any medium (e.g., a machine for programming a programmable gate array).

[0159] The design flow 3000 may vary depending on the type of representation being designed. For example, the design flow 3000 for building an application specific IC (ASIC) may be different from the design flow 3000 for designing a standard component or from the design flow 3000 for instantiating a design into a programmable array, such as a Company or The company offers programmable gate arrays (PGAs) or field programmable gate arrays (FPGAs).

[0160] Figure 30A plurality of such design structures are shown, including an input design structure 3020 that is preferably processed by design process 3010. Design structure 3020 can be a logic simulation design structure generated and processed by design process 3010 to produce a logically equivalent functional representation of a hardware device. Design structure 3020 can also or alternatively include data and / or program instructions that, when processed by design process 3010, generate a functional representation of the physical structure of the hardware device. Whether representing functional and / or structural design features, design structure 3020 can be generated using electronic computer-aided design (ECAD), such as implemented by a core developer / designer. When encoded on a machine-readable data transmission, gate array, or storage medium, design structure 3020 can be accessed and processed by one or more hardware and / or software modules within design process 3010 to simulate or otherwise functionally represent an electronic component, circuit, electronic or logic module, device, apparatus, or system, such as those shown herein. Thus, design structure 3020 may include files or other data structures, including human and / or machine readable source code, compiled structures, and computer executable code structures, which, when processed by a design or simulation data processing system, functionally simulate or otherwise represent a circuit or other level of hardware logic design. Such data structures may include hardware description language (HDL) design entities or other data structures that conform to and / or are compatible with low-level HDL design languages (such as Verilog and VHDL) and / or high-level design languages (such as C or C++).

[0161] Design process 3010 preferably employs and incorporates hardware and / or software modules for synthesizing, converting, or otherwise processing the design / simulation functional equivalents of the components, circuits, devices, or logic structures described herein to generate a netlist 3080 that may include a design structure such as design structure 3020. Netlist 3080 may include, for example, a compiled or otherwise processed data structure representing a list of wires, discrete components, logic gates, control circuits, I / O devices, models, and the like that describe connections to other elements and circuits in the integrated circuit design. Netlist 3080 may be synthesized using an iterative process, wherein netlist 3080 is resynthesized one or more times depending on the device's design specifications and parameters. As with other design structure types described herein, netlist 3080 may be recorded on a machine-readable storage medium or programmed into a programmable gate array. The medium may be a non-volatile storage medium, such as a magnetic or optical disk drive, a programmable gate array, a compact flash memory, or other flash memory. Additionally or alternatively, the medium may be system or cache memory, or buffer space.

[0162] The design process 3010 may include hardware and software modules for processing various input data structure types, including a netlist 3080. Such data structure types may, for example, reside within a library element 3030 and include a set of commonly used components, circuits, and devices, including models, layouts, and symbolic representations for a given manufacturing technology (e.g., different technology nodes, 30nm, 45nm, 90nm, etc.). Data structure types may also include design specifications 3040, characterization data 3050, verification data 3060, design rules 3070, and test data files 3085, which may include input test patterns, output test results, and other test information. The design process 3010 may also include, for example, standard mechanical design processes, such as stress analysis, thermal analysis, mechanical event simulation, process simulation for operations such as casting, molding, and compression molding, etc. Without departing from the scope and spirit of the present invention, one of ordinary skill in the art of mechanical design will understand the extent of possible mechanical design tools and applications used in the design process 3010. The design process 3010 may also include modules for performing standard circuit design processes such as timing analysis, verification, design rule checking, place and route operations, and the like.

[0163] Design process 3010 employs and combines logical and physical design tools, such as HDL compilers and simulation model building tools, to process design structure 3020 and some or all of the depicted supporting data structures, along with any additional mechanical design or data (if applicable), to generate a second design structure 3090. Design structure 3090 resides on a storage medium or programmable gate array in a data format for data exchange of mechanical devices and structures (e.g., information stored in IGES, DXF, Parasolid XT, JT, DRG, or any other suitable format for storing or representing such mechanical design structures). Similar to design structure 3020, design structure 3090 preferably includes one or more files, data structures, or other computer-coded data or instructions that reside on a transmission or data storage medium and, when processed by an ECAD system, generate a logical or other functionally equivalent form of one or more embodiments of the present invention described herein. In one embodiment, design structure 3090 may include a compiled, executable HDL simulation model that functionally simulates one or more devices described herein.

[0164] Design structure 3090 may also employ a data format and / or symbol data format for exchanging layout data of integrated circuits (e.g., information stored in GDSII (GDS2), GL1, OASIS, a drawing file, or any other suitable format for storing such a design data structure). Design structure 3090 may include information such as symbol data, drawing files, test data files, design content files, manufacturing data, layout parameters, wiring, metal levels, vias, shapes, data for routing through a production line, and any other data required by a manufacturer or other designer / developer to produce the devices or structures described above and shown herein. Design structure 3090 may then proceed to stage 3095, where, for example, design structure 3090 is: taped out, delivered to manufacturing, delivered to a mask shop, sent to another design shop, sent back to a customer, etc.

[0165] As described above, in at least one embodiment, a data processing system includes a master device, a central request agent, and a plurality of snoopers communicatively coupled to a system fabric for transmitting requests to be retried. The master device issues a multicast request on the system fabric intended for the plurality of snoopers. The central request agent receives the multicast request on the system fabric, assigns the multicast request to a specific state machine among a plurality of state machines in the central request agent, and provides a coherence response to the master device indicating successful completion of the multicast request. The central request agent repeatedly issues the multicast request on the system fabric in association with a machine identifier identifying the specific state machine until the coherence response indicates that the multicast request was successfully received by all of the plurality of snoopers.

[0166] While embodiments have been described in which various tag fields (e.g., forwarding, period, and action fields) are explicitly illustrated as separate tag fields for ease of explanation, in other embodiments, other means may be utilized to convey the same information. For example, in some embodiments, the number of code points in the encoded machine ID may be expanded for those state machines that issue multicast requests distributed via a central request agent.

[0167] Although aspects of the claimed invention have been described with reference to embodiments in which TLBIE requests are used as examples of multicast requests that can be distributed by a central request broker, those skilled in the art will appreciate that the disclosed techniques can also be applied to other types of multicast requests. For example, the disclosed techniques can be applied to ICBI (Instruction Cache Block Invalidate) requests, which request the invalidation of specified instructions in a non-coherent instruction cache.

[0168] It should also be understood that not all multicast requests need to be distributed via the central request broker. For example, in the described example, PTESYNC and TSYNC commands are not distributed to the snooper via the central request broker because the snooper (e.g., L2 cache) does not accept and process TSYNC or PTESYNC requests, and these commands are therefore not subject to ping-pong livelock.

[0169] Although various embodiments have been shown and described in detail, it will be understood by those skilled in the art that various changes in form and details may be made thereto without departing from the spirit and scope of the appended claims, and that such alternative implementations all fall within the scope of the appended claims. For example, although various aspects have been described with respect to a computer system that executes program code that directs the functions of the present invention, it will be understood that the present invention may alternatively be implemented as a program product comprising a computer-readable storage device storing program code that can be processed by a processor of a data processing system to cause the data processing system to perform the described functions. The computer-readable storage device may include volatile or non-volatile memory, optical or magnetic disks, etc., but does not include non-specified subject matter such as the propagating signal itself, the transmission medium itself, and the energy form itself.

[0170] As an example, the program product may include data and / or instructions that, when executed or otherwise processed on a data processing system, generate a logically, structurally, or otherwise functionally equivalent representation (including a simulation model) of the hardware components, circuits, devices, or systems disclosed herein. Such data and / or instructions may include hardware description language (HDL) design entities or other data structures that conform to and / or are compatible with low-level HDL design languages (such as Verilog and VHDL) and / or high-level design languages (such as C or C++). In addition, the data and / or instructions may also be in a data format and / or symbolic data format for exchanging layout data of integrated circuits (e.g., information stored in GDSII (GDS2), GL1, OASIS, a drawing file, or any other suitable format for storing such design data structures).

Claims

1. A method of distributing multicast requests in a data processing system comprising a master device, a central request broker, and a plurality of snoopers communicatively coupled to a system structure for transmitting requests to be retried, the method comprising: issuing, by the master device, a multicast request on the system fabric intended for the plurality of snoopers; receiving, by the central request agent on the system fabric, the multicast request, assigning the multicast request to a particular state machine among a plurality of state machines in the central request agent, and providing a coherent response to the master device indicating successful completion of the multicast request; as well as The multicast request is repeatedly issued by the central request agent on the system fabric in association with a machine identifier identifying the particular state machine until a consistency response indicates that the multicast request is successfully received by all of the plurality of snoopers.

2. The method according to claim 1, wherein The multicast request includes a translation entry invalidation request.

3. The method according to claim 1, wherein: Each of the plurality of snoopers includes a plurality of snoop machines corresponding in number to the plurality of state machines in the central request agent; as well as The method further comprises: Based on the machine identifier, each of the plurality of snoopers distributes the multicast request received from a central request agent to a particular snooper among the plurality of snoopers that corresponds to the particular state machine.

4. The method according to claim 1, further comprising: Prior to the reissuing, the central requesting agent modifies the multicast request to indicate that the multicast request is forwarded from the central requesting agent.

5. The method according to claim 1, further comprising: In response to the central requesting agent repeatedly issuing the multicast request, one or more of the plurality of snoopers provide a null consistency response to indicate successful receipt of the multicast request previously issued by the central requesting agent.

6. The method according to claim 1, wherein The multicast request includes a plurality of operation tenures on the system architecture.

7. The method according to claim 6, wherein: The receiving includes: the central request agent receiving at least a first operation occupancy period and a corresponding second operation occupancy period; The allocating includes: the central request agent allocating the first operation tenure and the second operation tenure to the specific state machine; and The method also includes marking, by the central request agent, the multicast request as available for distribution to the plurality of snoopers based on both the first operational tenure and the second operational tenure being assigned to the particular state machine.

8. The method according to claim 6, wherein: The repeatedly issuing includes: the central request agent issuing the first operation occupancy of the multicast request on the system structure with a time indication; and The method also includes, based on the mismatched epoch indication, a snooper of the plurality of snoopers discarding a previously snooped multicast request.

9. A data processing system comprising: A central request agent communicatively coupled to a system fabric to which a plurality of snoopers and a master device are communicatively coupled, wherein the central request agent includes a plurality of state machines, and wherein the central request agent is configured to perform: receiving, from the master device, a multicast request intended for the plurality of snoopers over the system fabric; In response to receiving the multicast request, assigning the multicast request to a particular state machine among the plurality of state machines and providing a conforming response indicating successful completion of the multicast request; and The multicast request is repeatedly issued on the system fabric in association with a machine identifier identifying the particular state machine until a consistency response indicates that the multicast request is successfully received by all of the plurality of snoopers.

10. The data processing system according to claim 9, wherein: The multicast request includes a translation entry invalidation request.

11. The data processing system of claim 9, further comprising the plurality of snoopers, wherein: Each of the plurality of snoopers includes a plurality of snoop machines corresponding in number to the plurality of state machines in the central request agent; as well as Each of the plurality of snoopers is configured to distribute the multicast request received from the central request agent to a particular snooper among the plurality of snoopers corresponding to the particular state machine based on the machine identifier.

12. The data processing system according to claim 9, wherein: The central request broker is configured to perform: Prior to the repeating, the multicast request is modified to indicate that the multicast request is forwarded from the central requesting agent.

13. The data processing system according to claim 9, further comprising the plurality of snoopers, wherein: The plurality of snoopers are configured to perform: An empty conformance response is provided to indicate successful receipt of the multicast request previously issued by the central requesting agent.

14. The data processing system according to claim 9, wherein: The multicast request includes a plurality of operation tenures on the system architecture.

15. The data processing system according to claim 14, wherein: The multicast request includes at least a first operation occupancy period and a corresponding second operation occupancy period; The allocating includes: the central request agent allocating the first operation tenure and the second operation tenure to the specific state machine; and The central request broker is configured to perform: The multicast request is marked as available for distribution to the plurality of snoopers based on both the first operational tenure and the second operational tenure being assigned to the particular state machine.

16. The data processing system according to claim 14, wherein: The repeatedly issuing includes the central request agent issuing the first operational occupancy of the multicast request on the system structure with a time period indication.

17. The data processing system according to claim 9, further comprising: the system structure; the master device; as well as The plurality of snoopers.

18. A design structure tangibly embodied in a machine-readable storage device for use in designing, manufacturing, or testing an integrated circuit, the design structure comprising: A central request agent communicatively coupled to a system fabric to which a plurality of snoopers and a master device are communicatively coupled, wherein the central request agent includes a plurality of state machines, and wherein the central request agent is configured to perform: receiving, from the master device, a multicast request intended for the plurality of snoopers over the system fabric; In response to receiving the multicast request, assigning the multicast request to a particular state machine among the plurality of state machines and providing a conforming response indicating successful completion of the multicast request; and The multicast request is repeatedly issued on the system fabric in association with a machine identifier identifying the particular state machine until a consistency response indicates that the multicast request is successfully received by all of the plurality of snoopers.

19. The design structure according to claim 18, wherein: The multicast request includes a translation entry invalidation request.

20. The design structure according to claim 18, wherein: The central request broker is configured to perform: Prior to the repeating, the multicast request is modified to indicate that the multicast request is forwarded from the central requesting agent.

21. The design structure according to claim 18, wherein: The multicast request includes a plurality of operation tenures on the system architecture.

22. The design structure according to claim 21, wherein: The multicast request includes at least a first operation occupancy period and a corresponding second operation occupancy period; The allocating includes: the central request agent allocating the first operation tenure and the second operation tenure to the specific state machine; and The central request broker is configured to perform: The multicast request is marked as available for distribution to the plurality of snoopers based on both the first operational tenure and the second operational tenure being assigned to the particular state machine.

23. The design structure of claim 21, wherein: The repeatedly issuing includes the central request agent issuing the first operational occupancy of the multicast request on the system structure with a time period indication.