Cache-forbidden write operation

By broadcasting cache-inhibited write requests and data on a shared address interconnect, the inefficiencies in cache-inhibited operations are addressed, enhancing system performance through optimized resource utilization and reduced latency.

JP7702194B2Active Publication Date: 2025-07-03INTERNATIONAL BUSINESS MACHINE CORPORATION
View PDF 10 Cites 0 Cited by

Patent Information

Application Number
JP2022516628
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2019-09-30
Filing Date
2020-08-20
Publication Date
2025-07-03
Estimated Expiration
2040-08-20

AI Technical Summary

Technical Problem

Existing cache-inhibited write operations in shared memory multiprocessor systems are inefficient due to non-productive waiting of system resources and high latency in delivering write data, particularly in cache-inhibited operations.

Method used

Implementing a broadcast address interconnect for cache-inhibited write requests and data transmission, allowing simultaneous broadcasting of write data and requests on the address interconnect, reducing the need for separate data packets and optimizing resource usage.

Benefits of technology

Enhances the efficiency of cache-inhibited write operations by minimizing resource consumption and reducing latency, thereby improving overall system performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007702194000002
    Figure 0007702194000002
  • Figure 0007702194000003
    Figure 0007702194000003
  • Figure 0007702194000004
    Figure 0007702194000004
Patent Text Reader

Abstract

A data processing system includes a plurality of processing units coupled to a system interconnect including a broadcast address interconnect and a data interconnect. The processing units include a processor core that executes memory access instructions and a cache memory coupled to the processor core and configured to store data for access by the processor core. The processing units are configured to broadcast, over the address interconnect, a cache-inhibit write request and write data to a destination device coupled to the system interconnect. In various embodiments, the initial cache-inhibit request and write data can be communicated in the same or different requests on the address interconnect.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention generally relates to a data processing system, and more specifically to a write operation in a data processing system. Even more specifically, the present invention relates to cache-inhibited write operations in a data processing system.

Background Art

[0002] In a shared memory multi-processor (MP) data processing system, each of a plurality of processors in the system can generally access and modify data stored in the shared memory. To reduce the access latency to data stored in the shared memory, a processor typically includes a high-speed local cache that buffers data fetched from the shared memory that is likely to be accessed by the processor. A consistent view of the data held in the various local caches of the processors is maintained by the implementation of a coherence protocol.

[0003] In such a shared memory MP system, for example, it is common to specify a particular address in the address space as cache-inhibited (non-cacheable) by appropriate setting of the page table in the shared memory. Data associated with these cache-inhibited addresses is not cached in the local cache of the processor. By restricting the caching of data for cache-inhibited addresses, maintaining a single consistent view of the contents of the corresponding memory location is greatly simplified in that changing the data associated with a cache-inhibited address does not require invalidation of any old or multiple copies of the data resident in the local cache of the processor.

Summary of the Invention

[0004] In one or more embodiments, writing to a cache - inhibited address is facilitated by broadcasting write data on an address interconnect of a data processing system.

[0005] For example, in at least one embodiment, a data processing system includes a plurality of processing units coupled to a system interconnect that includes a broadcast address interconnect and a data interconnect. The processing units include a processor core that executes memory access instructions and a cache memory coupled to the processor core and configured to store data for access by the processor core. The processing units are configured to broadcast cache - inhibited write requests and write data on the address interconnect to a destination device coupled to the system interconnect. In various embodiments, the first cache - inhibited request and write data can communicate in the same or different requests on the address interconnect.

Brief Description of the Drawings

[0006]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6A

Figure 6B

Figure 6C

Figure 7

Figure 8

Figure 9A

Figure 9B

Figure 10

Figure 11

Figure 12A

Figure 12B

Figure 12C

Figure 12D

Figure 13

Best Mode for Carrying Out the Invention

[0007] Referring now to the figures, and in particular to FIG. 1, a high-level block diagram of a data processing system 100 according to one embodiment is shown. As shown, data processing system 100 includes a plurality of processing units 102 (including at least processing units 102a-102b) for processing data and instructions. Processing units 102 are coupled to communicate with a system interconnect (interconnection) 104 to transfer address, data, and control information between attached devices. In a preferred embodiment, system interconnect 104 includes a branched address interconnect and a data interconnect. As will be further described below with reference to FIG. 3, the address interconnect, which transfers requests and coherence responses, is more preferably a broadcast interconnect such as an address bus that transfers all requests and coherence responses to all attached devices. In contrast, it is preferred that the data interconnect be a point-to-point interconnect such as a switch that supports direct communication of data from a source to a destination.

[0008] In the illustrated embodiment, the devices coupled to the system interconnect 104 include not only the processing unit 102, but also a memory controller 106 that provides an interface to the shared system memory 108, and one or more host bridges 110 that each provide an interface to a respective Mezzanine bus 112. The Mezzanine bus 112 also provides slots for attaching additional devices (not shown) that can include network interface cards, I / O adapters, non-volatile memory, non-volatile storage device adapters, additional bus bridges, and the like.

[0009] As further shown in FIG. 1, each processing unit 102, which can be implemented as a single integrated circuit, includes one or more processor cores 120 (only one of which is explicitly shown) for processing instructions and data. Each processor core 120 includes an instruction sequencing unit (ISU) 122 that fetches and directs instructions for execution, one or more execution units 124 for executing instructions dispatched from the ISU 122, and a set of registers 123 for temporarily buffering data and control information. Instructions executed by the execution units 124 include memory access instructions such as load and store instructions that read and write data associated with cacheable and non-cacheable addresses.

[0010] Each processor core 120 further includes an L1 store queue (STQ) 127 and a load unit 128 for managing the completion of store and load requests corresponding to the executed store and load instructions, respectively. In one embodiment, the L1 STQ 127 can be implemented as a first-in first-out (FIFO) queue that includes a plurality of queue entries. Thus, a store request is loaded into the "top" entry of the L1 STQ 127 at the time of execution of the corresponding store instruction to determine the target address, and is initiated when the store request reaches the "bottom" or "commit" entry of the L1 STQ 127.

[0011] It is important to note that this application distinguishes between "instructions" such as load and store instructions and "requests". Load and store "instructions" are defined herein as inputs to an execution unit that includes an opcode that identifies the type of instruction and one or more operands that specify the data being accessed or its address or both. Load and store "requests" are defined herein as the data or signals or both that are generated following instruction execution and that specify at least the target address of the data being accessed. Thus, load and store requests can be transmitted from the processor core 120 to the memory system to initiate a data access, but load and store instructions cannot be transmitted.

[0012] The operations of the processor core 120 are supported by a multi-level volatile memory hierarchy that has the shared system memory 108 at its lowest level and, in the illustrated embodiment, two or more levels of cache memory including the L1 cache 126 and the L2 cache 130 above it. As in other shared memory multiprocessor data processing systems, the cacheable content of the memory hierarchy can generally be accessed and modified by an execution thread executed by any processor core 120 in any processing unit 102 of the data processing system 100.

[0013] According to one embodiment, the L1 cache 126 is implemented as a store-through cache, which means that the cache coherence point for other processor cores 120 is placed below the L1 cache 126 and, in the illustrated embodiment, is placed in the store-in L2 cache 130. Thus, the L1 cache 126 only maintains valid / invalid bits without maintaining the true cache coherence state (e.g., Modified, Exclusive, Shared, Invalid, etc.) for its cache lines. Since the L1 cache 126 is implemented as a store-through cache, store requests are first completed for the associated processor core 120 within the L1 cache 126 and then completed for other processing units 102 at the system-wide coherence point, which is the L2 cache 130 in the illustrated embodiment.

[0014] As further shown in FIG. 1, the L2 cache 130 includes a storage array and directory 140 that stores cache lines of instructions and data, associated with their respective memory addresses and coherence states. The L2 cache 130 also includes a number of read claim (RC) state machines 142a - 142n that independently and simultaneously service cacheable memory access requests received from associated processor cores 120. The RC machines 142 receive core load requests from the LD unit 128 within the processor core 120 via the load bus 160, in-order L2 load queue (LDQ) 161, and command bus 162. Similarly, the RC machines 142 receive core store requests from the L1 STQ 127 within the processor core 120 via the store bus 164, in-order L2 store queue (STQ) 166, and command bus 162.

[0015] The L2 cache 130 further includes several snooping (SN) state machines 144a - 144n to service cacheable memory accesses and other requests received from other processing units 102 via the system interconnect 104 and the snooping bus 170. The SN machines 144 and the RC machines 142 are each connected to a back - invalidation bus 172, whereby any of the SN machines 144 or RC machines 142 can signal the processor core 120 to invalidate a cache line.

[0016] In at least one embodiment, the L2 cache 130 is constructed such that at most one of the RC machine 142 and the SN machine 144 can actively service requests targeting a given target cache line address at any given time. As a result, if a second request targeting the same cache line is received while a first request is already being serviced by an active RC machine 142 or SN machine 144, the second request needs to be queued or rejected until the service of the first request is complete and the active state machine returns to the idle state.

[0017] The L2 cache 130 further includes a non - cacheable unit (NCU) 146 to service cache - inhibit (CI) memory access requests received from the processor core 120. The NCU 146 includes a number of NCU store (NST) state machines 150a - 150n to independently and simultaneously service memory access requests received from the associated processor core 120 targeting non - cacheable addresses. The NST machines 150 receive core CI write requests from the L1 STQ 127 within the processor core 120 via the store bus 164 and the in - order NCU store queue (NSQ) 148. Further, the NCU 146 includes a number of NCU load (NLD) state machines 152a - 152n that receive core CI read requests from the LD unit 128 within the processor core 120 via the load bus 160.

[0018] One skilled in the art will further understand that the data processing system 100 of FIG. 1 can include many additional components not shown, such as an interconnect bridge, non-volatile storage, ports for connecting to a network or attached devices, etc. Since such additional components are not necessary for understanding the described embodiments, they are not shown in FIG. 1 and are not further discussed herein. However, it should be understood that the extended functions described herein are applicable to data processing systems of various architectures and are in no way limited to the architecture of the generalized data processing system shown in FIG. 1. For example, FIG. 1 shows an embodiment that includes a separate NCU to service cache inhibit requests, but other architectures can use other logic, such as cache logic within the L2 cache 130, to service cache inhibit requests such as CI writes.

[0019] Referring now to FIG. 2, a more detailed view of a destination device 200 according to one embodiment is shown. The destination device 200 can be any device coupled to a system interconnect 104 that can function as a destination (target) for a CI write request. For example, in the embodiment of FIG. 1, the destination device 200 can be the host bridge 110 or the memory controller 106.

[0020] In the illustrated embodiment, the destination device 200 is coupled to both an address interconnect 202 and a data interconnect 204 that form the system interconnect 104. As described above, the address interconnect 202 is preferably a broadcast interconnect such as a bus, meaning that each attached device has visibility to all requests and coherence messages transmitted on the address interconnect 202. On the other hand, the data interconnect 204 is preferably implemented as a point-to-point interconnect, meaning that while data packets can be routed to a plurality of (generally not all) snooping devices, only the destination of the operation receives and processes the data packets transmitted on the data interconnect 204. In the illustrated embodiment, the destination device 200 includes, among other components, a request queue 210 having a plurality of queue entries 212a - 212n for buffering requests received on the address interconnect 202. Each queue entry 212 includes at least an address field 214 for buffering the destination (target) address of the request, a request field 216 for buffering an indication of the request type (e.g., read, write, etc.), and a data field 218 for buffering data associated with the request. Further, as shown in FIG. 2, the destination device 200 also includes selection logic 220 (represented as a multiplexer) that can select information placed into the data field 218 from both the address interconnect 202 and the data interconnect 204.

[0021] Referring now to FIG. 3, a spacetime diagram of exemplary operations on system interconnect 104 of data processing system 100 according to one embodiment is shown. An operation begins when a master 300, such as RC machine 142 or NST machine 150, issues a request 302 on address interconnect 202 of data processing system 100. Request 302 preferably includes a transaction type indicating the type of access desired and a resource identifier (e.g., a target physical address) indicating the resource to be accessed by the request. Preferred types of requests preferably include those shown in Table I below. [Table 1]

[0022] Request 302 is received by destination devices 200 (e.g., host bridge 110 and memory controller 106) and snoopers 304 such as SN machine 144 of L2 cache 130. Generally, request 302 is transmitted on address interconnect 202 only when it cannot be internally serviced by processing unit 102, so there are some exceptions, but SN machine 144 within the same L2 cache 130 as the master that initiated request 302 does not snoop that request 302 (i.e., generally there is no self-snooping).

[0023] In response to receiving request 302, snooper 304 can provide respective partial responses (Presp) 306 on address interconnect 202, each Presp 306 representing at least its coherence response of snooper 304 to request 302. Destination device 200 determines its partial response (if any) based on, for example, whether destination device 200 yields the requested address and whether destination device 200 has resources available to service the request. L2 cache 130 can determine partial response 306 based on, for example, the availability of L2 storage array and directory 140, the availability of resources for processing the request (including available SN machines 144), and the cache state associated with the requested address in L2 storage array and directory 140.

[0024] The partial responses of various snoopers 304 are logically combined, either step - by - step or all at once, by one or more instances of response logic 308 to determine a system - wide combined response (Cresp) 310 to request 302. Response logic 308 provides combined response 310 to master 300 and snooper 304 via address interconnect 202 to indicate a system - wide response (such as Success, Retry, etc.) to request 302. If combined response 310 indicates success of request 302, combined response 310 can indicate, for example, the destination (if applicable) for write data, the cache state (if applicable) in which the requested memory block is cached by master 300, or whether a "cleanup" operation to invalidate the requested memory block in one or more caches 126, 130 is required (if applicable) or a combination thereof.

[0025] In response to receiving the combined response 310, one or more of the master 300 and the snooper 304 generally perform one or more operations to service the request 302. These operations can include supplying data to the master 300, invalidating or otherwise updating the data cached in one or more caches 126, 130, performing a cast-out operation, writing data to the system memory 108 or an I / O device, and the like. As will be described further below, data can be transmitted between the master 300 either before or after the generation of the combined response 310 by the response logic 210. Generally, for most operations, the data associated with the operations on the system interconnect 104 is transmitted via the data interconnect 204. However, as will be further described herein, in at least some embodiments, in some operations, the data is transmitted via the address interconnect 202.

[0026] The partial response provided by snooper 304 in response to a request, and the operation of executing the snooper in response to a request, or a combined response thereof or a combination of them, may depend on whether the snooper is at the Highest Point of Coherency (HPC), the Lowest Point of Coherency (LPC), or neither, with respect to the target address specified by the request. The LPC is defined herein as a memory device or I / O device that functions as a repository for a memory block. If there is no HPC for a memory block, the LPC holds the true image of the memory block and has the authority to permit or deny requests to generate additional cached copies of the memory block. In a typical request in the data processing system of FIGS. 1 and 2, the LPC is the memory controller 106 of the system memory 108 that functions as a repository for the referenced memory block, or the host bridge 110 that provides a memory-mapped I / O address. The HPC is defined herein as a uniquely identified device that caches the true image of the memory block (which may or may not match the corresponding memory block at the LPC) and has the authority to permit or deny requests to modify the memory block. Descriptively, the HPC can also provide a shared copy of the memory block to the requesting side in response to an operation that does not modify the memory block. Thus, in a typical request in the embodiment of the data processing system of FIGS. 1 and 2, the HPC, if it exists, is the L2 cache 230. Other metrics can be used to specify the HPC of a memory block, but the preferred embodiment of the present invention uses a selected cache state within the directory of the L2 cache 130 to specify the HPC (if it exists) of a memory block.

[0027] Referring further to FIG. 3, if present, the HPC for the memory block referenced within request 302, or if no HPC exists, the LPC of the memory block, preferably has the responsibility of preventing the transfer of coherence ownership of the memory block in response to request 402 during protection window 312a. In the exemplary scenario shown in FIG. 3, snooper 304, which is the HPC for the memory block specified by the request address of request 302, prevents, as needed, the transfer of ownership of the requested memory block to master 300 during protection window 312a (and possibly thereafter) that continues from the time snooper 304 determines its partial response 306 until snooper 304 receives the combined response 310. During protection window 312a, snooper 304 prevents the transfer of ownership by providing a partial response 306 (e.g., retry Presp) to other requests specifying the same request address, and this partial response 306 prevents other masters from acquiring coherence ownership until such ownership is successfully transferred to master 300. Master 300 similarly starts protection window 312b to protect the ownership of the memory block requested within request 302 after receiving the combined response 310.

[0028] Since all Snoopers 304 have limited resources for processing the above CPU and I / O requirements, several different levels of partial responses and corresponding combined responses are possible. For example, if the memory controller 106 responsible for the requested memory block has queue entries available for processing the request, the memory controller 106 can respond with a partial response indicating that it can act as an LPC for the request. On the other hand, if the memory controller 106 has no queue entries available for processing the request, the memory controller 106 can respond with a partial response indicating that it is an LPC for the memory block, but it cannot service the request at the current time. Similarly, the L2 cache 130 may request the available SN machines 144 to process the snooped request 302 and may access the L2 storage array and directory 140. If access to any (or both) of these resources is unavailable, it results in a partial response (and corresponding CR) indicating that the request cannot be serviced due to lack of necessary resources (e.g., retry).

[0029] Referring now to FIG. 4, a high-level logic flowchart of a prior art method in which a non-cacheable unit (NCU) issues a cache inhibit (CI) write request on the system interconnect of a data processing system is shown. For purposes of facilitating understanding, the prior art method of FIG. 4 will be described with reference to the exemplary data processing system 100 of FIG. 1 and the exemplary operations shown in FIG. 3.

[0030] The process of FIG. 4 begins at block 400 and then proceeds to block 402, which indicates that the NST machine 150 of the NCU 146 receives a CI write request from the associated processor core 120 via the store bus 164 and the NSQ 148. The CI write request is generated, for example, by the execution of a corresponding store instruction by the processor core 120 targeting a region of cache-inhibited memory and includes at least a target address and write data.

[0031] In response to receiving a CI write request, the NST machine 150 issues a CI write request on the address interconnect 202 (block 404). As shown in FIG. 6A, the CI write request 600 includes at least a request type field 602 that identifies the request as a CI write request, a master tag field 604 that identifies the processing unit 102, the NCU 146, and the NST machine 150 that issue the CI write request, and an address field 606 that specifies the target address provided by the associated processor core 120.

[0032] As described above with reference to FIG. 3, in response to a CI write request, the snooper 304 within the data processing system 100 provides respective partial responses (Presp) 306. These partial responses 306 are logically combined, and an instance of response logic 308 (e.g., resident within the issuing processing unit 102) forms a combined response (Cresp) 310 to the CI write request based on the partial responses 306. The cache snooper typically generates a null partial response to a cache invalidate request and does not perform additional processing on the cache invalidate request.

[0033] Figure 6B shows an example of the format of the coherence response 610 that can be used for both Presp306 and Cresp310. The coherence response 610 includes a master tag field 612 that identifies the processing unit 102, the NCU 146, and the NST machine 150 that issues the CI write request. Further, the coherence response 610 includes a destination tag field 614 that indicates the snooper 304 (more precisely, the destination device 200 and the queue entry 212), which is the destination of the CI write data (e.g., the memory controller 106 or the host bridge 110) in the Presp306 and Cresp310 of the CI write request. In Presp306, the destination tag field 614 is not changed by the snooper 304 that is not the final destination of the CI write data. The coherence response 610 also includes a response field 616 that indicates the accumulated coherence response of the snooper 304 in Presp306 (i.e., the partial responses of the snooper 304 to the request can be logically ORed with each other within the response field 616), and in Cresp310, based on the partial response 306, indicates the system-wide coherence response to the CI write request. The possible coherence responses indicated within the response field 616 for Cresp310 of the CI write request include, for example, success or retry.

[0034] Returning to FIG. 4, at block 406, the NST machine 150 that issued the CI write request monitors the associated Cresp310 (identified by the master tag field 612 of the coherence response 610) and determines whether the response field 616 of Cresp310 indicates success of the CI write request. If not (e.g., Cresp310 indicates a retry in the response field 616), the process returns to block 404 where the NST machine 150 reissues the CI write request on the address interconnect 202. At block 406, instead, in response to the NST machine 150 determining that Cresp310 indicates success of the CI write request, the NST machine 150 extracts the destination tag that identifies the associated destination device 200 and queue entry 212 from the destination tag field 614 of Cresp310 (block 408). The NST machine 150 first forms a data packet for the CI write (block 410). As shown in FIG. 6C, the data packet 620 includes a destination tag field 622 that contains the destination tag information extracted from the destination tag field 614 of Cresp310, and a data field 624 that contains the write data received from the processor core 120. Further as shown in block 410, the NST machine 150 then transmits the data packet to the associated destination device 200 via the point-to-point data interconnect 204. Thereafter, the process of FIG. 4 ends at block 412.

[0035] Referring now to FIG. 5, a high-level logic flowchart of a prior art method in which a destination device 200 (e.g., memory controller 106 or host bridge 110) services a CI write request is shown. The process of FIG. 5 begins at block 500 and then proceeds to block 502, which shows that the destination device 200 receives a CI write request 600 that specifies a target address for which the destination device 200 is responsible within address field 606. In response to receiving CI write request 600, the destination device 200 determines whether queue entry 212 in request queue 210 is available for assignment to the CI write request (block 504). If not available, the destination device 200 provides a Presp306 indicating a retry in response field 616 (block 506). This Presp306 causes the associated instance of response logic 308 to generate a Cresp310 also indicating a retry in response field 616. Following block 506, the process of FIG. 5 ends at block 520.

[0036] Returning to block 504, in response to a determination that entry 212 in request queue 210 is available for a CI write request, destination device 200 provides Presp306 indicating success in response field 616, loads the address specified in address field 606 of CI write request 600 into address field 214, and loads the display of the contents of request type field 602 into request field 216, thereby enqueuing the CI write request in available entry 212 of request queue 210 (block 510). As shown in block 512, next, destination device 200 monitors for receipt of data packet 620 of the CI write request via data interconnect 204. In response to receiving data packet 620 of the CI write request, destination device 200, identified by destination tag field 622, extracts data from data field 624 and loads the data into data field 218 (block 514). Next, destination device 200 processes the CI write request, for example, by storing the data in system memory 108 or writing the data to an attached I / O device (block 516). Thereafter, destination device 200 frees queue entry 212 (block 518), and the process of FIG. 5 ends at block 520.

[0037] In the present disclosure, it is recognized that a prior art process of performing a CI write operation as represented in FIGS. 4-5 is inefficient in that a queue entry 212 assigned to the CI write operation remains occupied without performing useful work between the assignment and the subsequent receipt of data packet 620. In the present disclosure, since system resources such as queue entry 212 are necessarily limited, it is recognized that it is preferable in CI write operations for these limited resources to be used more efficiently by reducing the time such resources are consumed by non-productive waiting. Also in the present disclosure, it is understood that the latency of data delivery for cache-disabled write operations can often determine the critical performance path of a program that can initiate processing by an I / O device, e.g., through a CI write operation. Accordingly, the present application discloses a plurality of techniques for accelerating CI write operations and communicating write data for CI write operations using a broadcast address interconnect 202. A first embodiment of communicating write data using a multi-beat CI write request is disclosed with reference to FIGS. 7, 8, and 9A-9B. A second embodiment of communicating write data using separate requests on the address interconnect 202 is disclosed with reference to FIGS. 10, 11, and 12A-12D.

[0038] Referring now to FIG. 7, a high-level logical flowchart of an exemplary method by which an NCU issues a CI write request on a system interconnect of a data processing system, according to a first embodiment, is shown. For purposes of facilitation of understanding, the method of FIG. 7 is described with reference to the exemplary data processing system 100 of FIG. 1 and the exemplary operations shown in FIG. 3.

[0039] The process of FIG. 7 starts at block 700 and then proceeds to block 702, which indicates that the NST machine 150 of the NCU 146 receives a CI write request from the associated processor core 120 via the store bus 164 and the NSQ 148. The CI write request is generated, for example, by the execution of a corresponding store instruction by the processor core 120 targeting an area of cache - inhibited memory and includes at least a target address and write data.

[0040] In response to receiving the CI write request, the NST machine 150 broadcasts the CI write request via the address interconnect 202 to all snoopers 304 coupled to the system interconnect 104 (block 704). As shown in FIG. 9A, the CI write request 900 includes at least a request type field 902 that identifies the request as a CI write request, a master tag field 904 that identifies the processing unit 102, the NCU 146, and the NST machine 150 that issued the CI write request, and an address field 906 that specifies the target address received from the associated processor core 120. Different from the prior - art CI write request 600 of FIG. 6A, the CI write request 900 of FIG. 9A further includes write data for the CI write operation within the data field 910 and, in some embodiments, can communicate it with additional beats on the broadcast address interconnect 202.

[0041] As described above, in response to the CI write request, the snoopers 304 within the data processing system 100 provide respective partial responses (Presp) 306 as described above with reference to FIG. 3. These partial responses 306 are then logically combined by an instance of response logic 308 (e.g., resident within the issuing processing unit 102) to form a combined response (Cresp) 310 to the CI write request. Cache snoopers typically generate a null partial response to cache - inhibit requests and perform no additional processing.

[0042] FIG. 9B shows an example of the format of a coherence response 920 that can be used for both Presp306 and Cresp310 of the CI write operation according to one embodiment. The coherence response 920 includes a master tag field 922 that identifies the processing unit 102, the NCU 146, and the NST machine 150 that issues the CI write request. Further, the coherence response 920 includes a destination tag field 924 that indicates the snooper 304 (more precisely, the destination device 200 and the queue entry 212), which is the destination of the CI write data (e.g., the memory controller 106 or the host bridge 110) in Presp306 and Cresp310 of the CI write request. In Presp306, the destination tag field 924 is not changed by the snooper 304, which is not the final destination of the CI write data. In FIGS. 7-8, it can be seen that this destination tag field 924 is not used to direct the delivery of the write data for the CI write operation. The coherence response 920 also includes a response field 926 that indicates the accumulated coherence responses of the snooper 304 in Presp306 (i.e., the partial responses of the snooper 304 to the request can be logically ORed with each other within the response field 926), and in Cresp310, indicates the system-wide coherence response to the CI write request. Similar to the prior art, the possible coherence responses indicated within the response field 926 for Cresp310 of the CI write request include, for example, success or retry.

[0043] Returning to FIG. 7, at block 706, the NST machine 150 that issued the CI write request monitors the associated Cresp310 (identified by the master tag field 922 of the coherence response 920) and determines whether the response field 926 of Cresp310 indicates success of the CI write request. If not (e.g., Cresp310 indicates a retry in the response field 926), the process returns to block 704 indicating that the NST machine 150 reissues the CI write request on the address interconnect 202. At block 706, instead, in response to the NST machine 150 determining that Cresp310 indicates success of the CI write request, the NST machine 150 is deallocated and the process of FIG. 7 ends at block 708. Thus, in this embodiment, the NST machine 150 does not perform any steps corresponding to blocks 408 - 410 of the prior art method given in FIG. 4.

[0044] Referring now to FIG. 8, a high - level logical flowchart of an exemplary method by which a destination device 200 (e.g., memory controller 106 or host bridge 110) according to the first embodiment services a CI write request is shown. The process of FIG. 8 starts at block 800 and then proceeds to block 802 which indicates that the destination device 200 receives a CI write request 900 that specifies a target address for which the destination device 200 is responsible within the address field 906. In response to receiving the CI write request 900, the destination device 200 determines (block 804) whether a queue entry 212 in the request queue 210 is available for allocation to the CI write request 900. If not available, the destination device 200 provides a Presp306 indicating a retry in the response field 926. This Presp306 causes the response logic 308 to generate a Cresp310 also indicating a retry in the response field 926. Following block 806, the process of FIG. 8 ends at block 820.

[0045] Returning to block 804, in response to the determination that entry 212 in request queue 210 is available for assignment to a CI write request, destination device 200 provides Presp306 indicating success in response field 926, loads the address specified in address field 906 of CI write request 900 into address field 214, loads the content of request type field 902 into request field 216, and loads the display of the content of data field 910 into data field 218 via selection logic 220 to enqueue the CI write request into available entry 212 (block 810). In contrast to the process of FIG. 5 where destination device 200 needs to wait for receipt of a separate data packet 620, in the process of FIG. 8, destination device 200 can immediately process the CI write request, for example, by storing the data from data field 218 into system memory 108 or writing the data from data field 218 to an attached I / O device (block 812). Thereafter, destination device 200 releases queue entry 212 assigned to the CI write request (block 814), and the process of FIG. 8 ends at block 820.

[0046] Referring now to FIG. 10, a high-level logic flowchart of an exemplary method by which an NCU issues a CI write request on a system interconnect of a data processing system according to a second embodiment is shown. For purposes of facilitating understanding, the method of FIG. 10 is described with reference to the exemplary data processing system 100 of FIG. 1.

[0047] The process of FIG. 10 starts at block 1000 and then proceeds to block 1002, which indicates that the NST machine 150 of the NCU 146 receives a CI write request from the associated processor core 120 via the store bus 164 and the NSQ 148. The CI write request is generated, for example, by the execution of a corresponding store instruction by the processor core 120 targeting an area of cache-inhibited memory and includes at least a target address and write data.

[0048] In response to receiving the CI write request, the NST machine 150 broadcasts a first request on the address interconnect 202 (block 1004). As shown in FIG. 12A, in an exemplary embodiment, the first request 1200 includes at least a request type field 1202 that identifies the request as a CI write request, a master tag field 1204 that identifies the processing unit 102, the NCU 146, and the NST machine 150 that issued the CI write request, and an address field 1206 that specifies the target address received from the associated processor core 120. Thus, in this embodiment, the first request 1200 can be the same as or similar to the prior art CI write request 600 of FIG. 6A.

[0049] As described above, in response to the first request 1200, the snooper 304 within the data processing system 100 provides a partial response (Presp) 306 (which may be a null response) as described above with reference to FIG. 3. These partial responses 306 are then logically combined, and an instance of the response logic 308 (e.g., resident within the issuing processing unit 102) forms a combined response (Cresp) 310 to the CI write request based on the Presp 306.

[0050] Figure 12B shows an exemplary format of a first coherence response 1210 that can be used for both Presp306 and Cresp310 generated for a first request 1200. The coherence response 1210 includes a master tag A field 1212 that specifies the content of the master tag A field 1204 of the first request 1200. Further, the first coherence response 1210 includes a destination tag field 1214 that indicates the snooper 304 (more precisely, the destination device 200 and the queue entry 212), which is the destination of the CI write data (e.g., the memory controller 106 or the host bridge 110) for Presp306 and Cresp310 of the first request 1200. In Presp306, the destination tag field 1214 is not changed by the snooper 304 that is not the final destination of the CI write data. In FIGS. 10-11, it is observed that this destination tag field 1214 is not used to direct data delivery for the CI write operation. The coherence response 1210 also indicates, in Presp306, the accumulated coherence responses of the snooper 304 (i.e., the partial responses of the snooper 304 to the request can be logically ORed with each other within the response field 1216), and in Cresp310, indicates the system-wide coherence response to the first request 1200. Similar to the prior art, the possible coherence responses shown within the response field 1216 for Cresp310 of the first request 1200 include, for example, success or retry.

[0051] Returning to FIG. 10, in block 1006, the NST machine 150 broadcasts a second request on the address interconnect 202 without synchronizing with the receipt of Cresp310 of the first request 1200 by the NST machine 150. As shown in FIG. 12C, in an exemplary embodiment, the second request 1220 includes at least a request type field 1222 that identifies the request as a CI write request, a master tag A' field 1224, and a data field 1226 that specifies the write data of the CI write request. In a preferred embodiment, the master tag A' field 1224 preferably identifies the processing unit 102, the NCU 146, and the NST machine 150 that issues the second request 1220, while being different from the content of the master tag A field 1204. Thus, the associated destination device 200 can distinguish between the first request 1200 that provides the target address of the CI write request and the second request 1220 that provides the write data of the CI write request. The address interconnect 202 preferably keeps the first request 1200 and the second request 1220 in order and ensures that the destination device 200 receives the first request 1200 before the second request 1220.

[0052] Block 1008 in FIG. 10 indicates that the NST machine 150 monitors Cresp310 of the first request 1200 (identified by the master tag A field 1212 of the coherence response 1210) and determines whether the response field 1216 of Cresp310 indicates the success of the first request 1200. If not (e.g., Cresp310 indicates a retry in the response field 1216), the process proceeds to block 1010, which indicates that the NST machine 150 waits to receive Cresp310 of the second request 1220 before reissuing the first request 1200 in block 1004. However, if the NST machine 150 determines in block 1008 that Cresp310 of the first request 1200 indicates success, the NST machine 150 waits to receive Cresp310 of the second request 1220, which can take the form of the coherence response 1230 in FIG. 12D (block 1012). In this example, the coherence response 1230 includes a master tag A' field 1232, a destination tag eID field 1234, and a response field 1236, which respectively correspond to the fields 1212, 1214, and 1216 described above. In response to receiving the Cresp of the second request 1220, the NST machine 150 is deallocated and the process in FIG. 10 ends at block 1020.

[0053] Referring now to FIG. 11, a high-level logic flowchart of an exemplary method by which a destination device 200 (e.g., memory controller 106 or host bridge 110) services a CI write request according to a second embodiment is shown. The process of FIG. 11 begins at block 1100 and then proceeds to block 1102, which indicates that the destination device 200 receives a first request 1200 that specifies a target address for which the destination device 200 is responsible within address field 1206. In response to receiving the first request 1200, the destination device 200 determines whether queue entry 212 is available for assignment to the CI write request (block 1104). If not available, the destination device 200 provides a Presp306 indicating a retry in response field 1216 (block 1106). This Presp306 causes response logic 308 to generate a Cresp310 also indicating a retry in response field 1216. Following block 1106, the process of FIG. 11 ends at block 1120.

[0054] Returning to block 1104, in response to a determination that entry 212 is available for the CI write request, the destination device 200 enqueues the CI write request to the available entry 212 by providing a Presp306 indicating success in response field 1216, loading the address specified in address field 1206 of the first request 1200 into address field 214, and loading the contents of request type field 1202 into request field 216.

[0055] As shown in block 1112, the destination device 200 then monitors the receipt of the write data of the CI write request in the second request 1220 on the address interconnect 202. In response to receiving the second request 1220, the destination device 200 provides a null Presp 306, extracts data from the data field 1226, and loads the data into the data field 218 of the associated queue entry 212 via the selection logic 220 (block 1114). The destination device 200 then processes the CI write request, for example, by storing the data from the data field 218 in the system memory 108 or writing the data from the data field 218 to the attached I / O device (block 1116). Thereafter, the destination device 200 releases the queue entry 212 (block 1118), and the process of FIG. 11 ends at block 1120.

[0056] Referring now to FIG. 13, there is shown, for example, a block diagram of an exemplary design flow 1300 used in the design, simulation, testing, layout, and manufacture of semiconductor IC logic. Design flow 1300 includes a process, machine, or mechanism, or a combination thereof, for processing a design structure or device as described above and shown herein to generate a logically or otherwise functionally equivalent representation of the design structure or device, or both, as described above and shown herein. The design structure processed, generated, or both, by design flow 1300 can be encoded on a machine-readable transmission or storage medium so that, when executed on a data processing system or otherwise processed, it includes data, instructions, or both, that generate a logically, structurally, mechanically, or otherwise functionally equivalent representation of a hardware component, circuit, device, or system. The machine can include any machine used in the IC design process, including, but not limited to, machines for designing, manufacturing, or simulating circuits, components, devices, or systems. For example, the machine can include a lithography machine, a machine or apparatus, or both, for generating a mask (e.g., an e-beam writer), a computer or apparatus for simulating a design structure, any apparatus used in a manufacturing process or a testing process, or any machine for programming a functionally equivalent representation of a design structure onto any medium (e.g., a machine for programming a programmable gate array).

[0057] Design flow 1300 may vary depending on the type of representation being designed. For example, the design flow 1300 for building an application specific integrated circuit (ASIC) may be different from the design flow 1300 for designing standard components, or from the design flow 1300 for instantiating a design onto a programmable gate array (PGA), such as a programmable gate array (PGA) or a field programmable gate array (FPGA) provided by Altera® Inc. or Xilinx® Inc.

[0058] FIG. 13 shows a plurality of such design structures, preferably including an input design structure 1020 processed by a design process 1310. The design structure 1320 can be a logical simulation design structure that is generated and processed by the design process 1310 and results in a logically equivalent functional representation of a hardware device. The design structure 1320 can further, or alternatively, include data and / or program instructions, or both, that generate a functional representation of the physical structure of a hardware device when processed by the design process 1310. Whether representing functional or structural or both design features, the design structure 1320 can be generated using electronic computer-aided design (ECAD) as implemented by a core developer / designer. When encoded on a machine-readable data transmission, gate array, or storage medium, the design structure 1320 can be accessed and processed by one or more hardware modules and / or software modules, or both, within the design process 1310 to simulate or otherwise functionally represent electronic components, circuits, electronic or logical modules, devices, apparatuses, or systems as shown herein. Thus, the design structure 1320 can include human and / or machine-readable source code, compiled structures, and computer-executable code structures that functionally simulate or otherwise represent circuit or other levels of hardware logic design when processed by a design or simulation data processing system, including a file or other data structure. Such data structures can include hardware description language (HDL) design entities, or other data structures that conform to and / or are compatible with low-level HDL design languages such as Verilog and VHDL and / or high-level design languages such as C or C++.

[0059] The design process 1310 preferably uses, incorporates, or otherwise processes hardware modules, software modules, or both to generate a netlist 1380 that can include design structures such as design structure 1320, which synthesize, transform, or otherwise process the functional equivalents of the components, circuits, devices, or logical structures shown herein for design / simulation. The netlist 1380 can include, for example, a compiled or otherwise processed data structure representing a list of wiring, individual components, logic gates, control circuits, I / O devices, models, etc., that describe connections to other elements and circuits within an integrated circuit design. The netlist 1380 can be synthesized using an iterative process in which the netlist 1380 is resynthesized one or more times depending on the design specifications and parameters of the device. Similar to the other types of design structures described herein, the netlist 1380 can be recorded on a machine-readable storage medium or programmed into a programmable gate array. The medium can be a non-volatile storage medium such as a magnetic or optical disk drive, a programmable gate array, a compact flash, or other flash memory. Additionally, or alternatively, the medium can be a system or cache memory, or a buffer region.

[0060] The design process 1310 can include hardware and software modules for processing various types of input data structures including the netlist 1380. Such types of data structures can include, for example, elements that reside within the library elements 1330 and include a set of commonly used elements, circuits, and devices including models, layouts, and symbolic representations for a given manufacturing technology (e.g., different technology nodes 32nm, 45nm, 90nm, etc.). The types of data structures can further include the design specifications 1340, the characteristic data 1350, the verification data 1360, the design rules 1390, and the test data file 1385 which can include input test patterns, output test results, and other test information. The design process 1310 can further include standard mechanical design processes such as process simulations for operations such as stress analysis, thermal analysis, mechanical event simulation, casting, molding, die press forming, etc. One skilled in mechanical design can recognize the scope of possible mechanical design tools and applications used in the design process 1310 without departing from the scope of the present invention. The design process 1310 can also include modules for performing standard circuit design processes such as timing analysis, verification, design rule checking, placement, and routing operations.

[0061] The design process 1310 uses and incorporates logical and physical design tools, such as an HDL compiler and a simulation model building tool, to create a second design structure 1390 and processes the design structure 1320, along with some or all of the illustrated support data structures and any additional mechanical design or data (if applicable), to create the second design structure 1390. The design structure 1390 exists on a storage medium or a programmable gate array in a data format used for the exchange of data of mechanical devices and structures (e.g., information stored in IGES, DXF, Parasolid XT, JT, DRG, or any other format suitable for storing or rendering such mechanical design structures). Similar to the design structure 1320, the design structure 1390 preferably includes one or more files, data structures, or other computer-coded data or instructions, which exist on a transmission or data storage medium and, when processed by an ECAD system, generate a functionally equivalent form in one or more logical or other ways of the embodiments of the invention shown herein. In one embodiment, the design structure 1390 can include, for example, a compiled executable HDL simulation model that functionally simulates the devices shown herein.

[0062] The design structure 1390 can also use either a data format or a symbolic data format (e.g., GDSII (GDS2), GL1, OASIS, map file, or information stored in any other suitable format for storing such design data structures) used for the exchange of integrated circuit layout data, or both. The design structure 1390 can include, for example, symbolic data, map files, test data files, design content files, manufacturing data, layout parameters, wiring, metal levels, vias, shapes, data for path specification through a manufacturing line, and any other data such as that required by a manufacturer or other designer / developer to manufacture a device or structure as described above and shown herein. Next, the design structure 1390 can proceed to stage 1395, where, for example, the design structure 1390 can be read onto a tape, released for manufacturing, released to a mask company, sent to another design company, or returned to a customer.

[0063] As described above, in at least one embodiment, the data processing system includes a plurality of processing units coupled to a system interconnect including a broadcast address interconnect and a data interconnect. The processing unit includes a processor core that executes memory access instructions and a cache memory coupled to the processor core and configured to store data for access by the processor core. The processing unit is configured to broadcast cache invalidate write requests and write data to a destination device coupled to the system interconnect on the address interconnect. In various embodiments, the first cache invalidate request and write data can communicate in the same or different requests on the address interconnect.

[0064] Although various embodiments have been specifically shown and described, those skilled in the art will recognize that various changes in form and detail can be made without departing from the scope of the appended claims, and that all such alternative embodiments fall within the scope of the appended claims.

[0065] The flowcharts and block diagrams in the drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in the flowchart or block diagram can represent a module, segment, or portion of instructions that includes one or more executable instructions for implementing the specified logical function. In some alternative implementations, the functions noted in the blocks may occur out of the order noted in the figures. For example, two blocks shown in succession may, in fact, be executed substantially concurrently depending on the functionality involved, or the blocks may sometimes be executed in the reverse order. It should also be noted that each block of the block diagrams or flowchart diagrams, and combinations of blocks in the block diagrams or flowchart diagrams or both, can be implemented by a dedicated hardware-based system that performs the specified functions or operations or a combination of dedicated hardware and computer instructions.

[0066] Furthermore, although aspects have been described with respect to a computer system, it should be understood that the present invention can alternatively be implemented as a program product including a computer-readable storage device storing program code that can be processed by a data processing system. The computer-readable storage device can include volatile or non-volatile memory, optical disks or magnetic disks, etc. However, as used herein, "storage device" is specifically defined to include only statutory manufactured articles and to exclude signal media per se, transient propagated signals per se, and energy per se.

[0067] When executed on a data processing system or otherwise processed, a program product can include data and / or instructions that generate a logical, structural, or otherwise functionally equivalent representation (including a simulation model) of the hardware components, circuits, devices, or systems disclosed herein. Such data and / or instructions can include other data structures that conform to and / or are compatible with a hardware description language (HDL) design entity, or a low-level HDL design language such as Verilog and VHDL, and / or a high-level design language such as C or C++. Additionally, the data and / or instructions can use a data format used for the exchange of integrated circuit layout data, or a symbolic data format (e.g., GDSII (GDS2), GL1, OASIS, map file, or information stored in any other suitable format for storing such design data structures), or both.

Claims

1. A processing unit for a data processing system including a system interconnect having an address interconnect configured to broadcast requests and coherence responses to a plurality of devices, and a data interconnect, the processing unit comprising: A processor core that executes memory access instructions; A cache memory coupled to the processor core and configured to store data for access by the processor core; Including; The processing unit is configured to broadcast a cache inhibit write request and write data of the cache inhibit write request to a destination device among a plurality of devices coupled to the system interconnect on the address interconnect; Processing unit.

2. The processing unit according to claim 1, wherein the processing unit broadcasts the cache inhibit write request on a first beat on the address interconnect and broadcasts the write data on a second beat on the address interconnect.

3. The processing unit according to claim 1, wherein the processing unit is configured to broadcast the write data included in the cache inhibit write request.

4. The cache inhibit write request is a first request, The processing unit according to claim 1, wherein the processing unit is configured to broadcast the write data included in a different second request on the address interconnect.

5. A plurality of processing units including the processing unit according to claim 1; The destination device; The system interconnect that communicatively couples the destination device and the plurality of processing units; Including a data processing system.

6. The data processing system according to claim 5, wherein the destination device includes a memory controller of a system memory of the data processing system.

7. The data processing system according to claim 5, wherein the destination device includes an interconnect bridge.

8. The data processing system according to claim 5, wherein the data interconnect is a point-to-point interconnect that supports direct communication of data from a source to a destination.

9. A data processing method in a processing unit of a data processing system including a plurality of processing units coupled to a system interconnect having an address interconnect configured to broadcast requests and coherence responses to a plurality of devices and a data interconnect, comprising: receiving, by the processing unit, from a processor core of the processing unit, a cache-disable write request and write data for the cache-disable write request; in response to receiving the cache-disable write request, the processing unit broadcasting, on the address interconnect of the data processing system, the cache-disable write request and the write data for the cache-disable write request to a destination device among a plurality of devices coupled to the system interconnect; A method comprising the above.

10. The method according to claim 9, wherein the broadcasting includes broadcasting the cache-disable write request on a first beat on the address interconnect and broadcasting the write data on a second beat on the address interconnect.

11. The method according to claim 9, wherein the broadcasting includes the processing unit broadcasting the write data included in the cache-disable write request.

12. The cache-disable write request is a first request, The method according to claim 9, wherein the broadcasting includes the processing unit broadcasting the write data included in a different second request on the address interconnect.

13. The method according to claim 9, wherein the destination device is one of a set including a memory controller of a system memory of the data processing system and an interconnect bridge.

14. A data structure described in a hardware description language (HDL) configured to simulate the processing unit according to any one of claims 1 to 4 when executed on a data processing system.

Citation Information

Patent Citations

  • Microprocessor

    JP1990309442A

  • Multiprocessor system and network for the same

    JP1997138782A

  • Data memory and processor bus

    JP1997507325A

  • Memory controller

    JP2004127305A

  • A mechanism for reordering transactions in a computer system with a snoop-based cache consistency protocol

    JP2004507800A