Scalable System-on-Chip
The scalable SOC design with a unified memory architecture and flexible interconnect fabric addresses the challenge of design duplication by enabling seamless scaling and software compatibility across varying complexity levels, enhancing design reuse and functionality.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2025-11-13
- Publication Date
- 2026-03-10
AI Technical Summary
Traditional SOC designs are not scalable and require significant duplication of design effort across different applications, lacking opportunities for design reuse and software compatibility across varying complexity levels.
A scalable SOC design featuring a unified memory architecture with a unified address space, allowing heterogeneous agents to share memory access, and a flexible interconnect fabric that supports easy scaling of complexity by integrating processor cores, graphics processing units, and peripheral devices, with independent networks optimized for different traffic types.
Enables seamless scaling of SOC designs from small to large applications, reducing design duplication, facilitating software reuse, and ensuring consistent functionality across different resource versions.
Smart Images

Figure 2026041763000001_ABST
Abstract
Description
[Technical Field]
[0001] FIELD OF THE INVENTION The embodiments described herein relate to digital systems, and more particularly to systems having a unified memory accessible to disparate agents within the system. [Background technology]
[0002] In the design of modern computing systems, it is becoming increasingly common to integrate a wide variety of system hardware components, which were previously implemented as separate silicon components, onto a single silicon die. For example, in the past, a complete computer system might have included a separately packaged microprocessor mounted on a backplane and coupled to a chipset that interfaced the microprocessor to other devices, such as system memory, a graphics processor, and other peripheral devices. In contrast, advances in semiconductor process technology have made possible the integration of many of these discrete devices. The result of such integration is commonly referred to as a "system on a chip" (SOC).
[0003] Traditionally, SOCs for different applications are built, designed, and implemented individually. For example, a SOC for a smartwatch device may have stringent power consumption requirements because the form factor of such a device limits the available battery size and therefore the device's maximum usage time. At the same time, the small size of such a device may limit the number of peripherals the SOC needs to support and the computational requirements of the applications it runs. In contrast, a SOC for a mobile phone application has a larger available battery and therefore a larger power budget, but is also expected to have more complex peripherals and greater graphics and general computational requirements. Therefore, such a SOC is expected to be larger and more complex than designs for smaller devices. This comparison can be extended to other applications as desired. For example, wearable computing solutions, such as augmented and / or virtual reality systems, may be expected to present greater computing requirements than less complex devices, as well as desktop and / or rack-mounted computer systems.
[0004] The traditional individually architected approach to SOCs provides little opportunity for design reuse and design effort is duplicated across multiple SOC implementations.
[0005] The following detailed description refers to the accompanying drawings, which are briefly described below. [Brief explanation of the drawings]
[0006] [Figure 1] FIG. 1 is a block diagram of one embodiment of a system on a chip (SOC).
[0007] [Figure 2] FIG. 1 is a block diagram of a system including one embodiment of multiple networks interconnecting agents.
[0008] [Figure 3] FIG. 1 is a block diagram of an embodiment of a network using a ring topology.
[0009] [Figure 4] FIG. 1 is a block diagram of an embodiment of a network using a mesh topology.
[0010] [Figure 5] FIG. 1 is a block diagram of an embodiment of a network using a tree topology.
[0011] [Figure 6] FIG. 1 is a block diagram of an embodiment of a system on a chip (SOC) having multiple networks for one embodiment.
[0012] [Figure 7] FIG. 7 is a block diagram of one embodiment of a system-on-chip (SOC) illustrating one of the independent networks shown in FIG. 6 for one embodiment.
[0013] [Figure 8] FIG. 7 is a block diagram of an embodiment of a system-on-chip (SOC) illustrating another one of the independent networks shown in FIG. 6 for one embodiment.
[0014] [Figure 9] FIG. 7 is a block diagram of one embodiment of a system on a chip (SOC) illustrating yet another of the independent networks shown in FIG. 6 for one embodiment.
[0015] [Figure 10] FIG. 1 is a block diagram of an embodiment of a multi-die system including two semiconductor dies.
[0016] [Figure 11] FIG. 2 is a block diagram of one embodiment of an input / output (I / O) cluster.
[0017] [Figure 12] FIG. 2 is a block diagram of one embodiment of a processor cluster.
[0018] [Figure 13] 10 is a pair of tables showing the virtual channels and traffic types and networks shown in FIGS. 6-9 used for one embodiment.
[0019] [Figure 14] 1 is a flow chart illustrating one embodiment of initiating a transaction over a network.
[0020] [Figure 15] FIG. 1 is a block diagram of one embodiment of a system including an interrupt controller and multiple cluster interrupt controllers corresponding to multiple clusters of processors.
[0021] [Figure 16] FIG. 16 is a block diagram of an embodiment of a system-on-chip (SOC) that may implement an embodiment of the system shown in FIG. 15.
[0022] [Figure 17] FIG. 2 is a block diagram of an embodiment of a state machine that may be implemented in an embodiment of an interrupt controller.
[0023] [Figure 18] 10 is a flowchart illustrating the operation of an embodiment of an interrupt controller that implements soft or hard iteration of interrupt delivery.
[0024] [Figure 19] 1 is a flowchart illustrating the operation of one embodiment of a cluster interrupt controller.
[0025] [Figure 20] FIG. 2 is a block diagram of one embodiment of a processor.
[0026] [Figure 21] FIG. 2 is a block diagram of one embodiment of a reorder buffer.
[0027] [Figure 22] 21 is a flow diagram illustrating the operation of one embodiment of the interrupt acknowledgement control circuit shown in FIG. 20.
[0028] [Figure 23] FIG. 16 is a block diagram of multiple SOCs that may implement an embodiment of the system shown in FIG. 15.
[0029] [Figure 24] 24 is a flow diagram illustrating the operation of one embodiment of the primary interrupt controller shown in FIG. 23.
[0030] [Figure 25] 24 is a flowchart illustrating the operation of one embodiment of the secondary interrupt controller shown in FIG. 23.
[0031] [Figure 26] 1 is a flow chart illustrating one embodiment of a method for handling interrupts.
[0032] [Figure 27] FIG. 1 is a block diagram of one embodiment of a cache coherent system implemented as a system on a chip (SOC).
[0033] [Figure 28] FIG. 2 is a block diagram illustrating one embodiment of a three-hop protocol for coherent transfer of cache blocks.
[0034] [Figure 29] 1 is a block diagram illustrating one embodiment of managing contention between a fill of one coherent transaction and a snoop of another coherent transaction.
[0035] [Figure 30]2 is a block diagram illustrating one embodiment of managing contention between a snoop of one coherent transaction and an acknowledgment of another coherent transaction.
[0036] [Figure 31] FIG. 2 is a block diagram of a portion of an embodiment of a coherent agent.
[0037] [Figure 32] 10 is a flowchart illustrating the operation of one embodiment of processing a request in a coherence controller.
[0038] [Figure 33] FIG. 10 is a flow diagram illustrating the operation of one embodiment of a coherent agent that has transmitted a request to a memory controller to process a completion associated with the request.
[0039] [Figure 34] FIG. 10 is a flow diagram illustrating the operation of one embodiment of a coherent agent receiving a snoop.
[0040] [Figure 35] FIG. 2 is a block diagram illustrating a chain of competing requests for a cache block, according to one embodiment.
[0041] [Figure 36] 10 is a flow chart illustrating one embodiment of a coherent agent absorbing snoops.
[0042] [Figure 37] FIG. 2 is a block diagram illustrating one embodiment of a non-cacheable request.
[0043] [Figure 38] FIG. 10 is a flow diagram illustrating the operation of one embodiment of a coherence controller that generates snoops based on the cacheable and non-cacheable properties of a request.
[0044] [Figure 39] 1 is a table illustrating multiple cache states according to one embodiment of a coherence protocol.
[0045] [Figure 40] 1 is a table illustrating messages that may be used in one embodiment of a coherency protocol.
[0046] [Figure 41] FIG. 10 is a flow diagram illustrating the operation of one embodiment of a coherence controller for processing changes to exclusive conditional requests.
[0047] [Figure 42] FIG. 1 is a flow diagram illustrating the operation of one embodiment of a coherence controller for reading directory entries and generating snoops.
[0048] [Figure 43] 10 is a flowchart illustrating the operation of one embodiment of a coherence controller for processing an exclusive no data request.
[0049] [Figure 44] FIG. 2 is a block diagram illustrating exemplary elements of a system-on-chip, according to some embodiments.
[0050] [Figure 45] 1 is a block diagram illustrating exemplary elements of interaction between an I / O agent and a memory controller, according to some embodiments.
[0051] [Figure 46A] FIG. 2 is a block diagram illustrating exemplary elements of an I / O agent configured to process write transactions, according to some embodiments.
[0052] [Figure 46B]FIG. 2 is a block diagram illustrating exemplary elements of an I / O agent configured to process read transactions, according to some embodiments.
[0053] [Figure 47] FIG. 1 is a flow diagram illustrating an example of processing a read transaction request from a peripheral component, according to some embodiments.
[0054] [Figure 48] FIG. 1 is a flow diagram illustrating an exemplary method for processing a read transaction request by an I / O agent, according to some embodiments.
[0055] [Figure 49] FIG. 1 shows a block diagram of one embodiment of a system having two integrated circuits coupled together.
[0056] [Figure 50] FIG. 1 illustrates a block diagram of one embodiment of an integrated circuit having an external interface.
[0057] [Figure 51] 1 shows a block diagram of a system having two integrated circuits that utilize interface wrappers to route the pin assignments of their respective external interfaces.
[0058] [Figure 52] FIG. 1 illustrates a block diagram of one embodiment of an integrated circuit having an external interface utilizing pin bundles.
[0059] [Figure 53A] 1 shows two examples of two integrated circuits coupled together using complementary interfaces.
[0060] [Figure 53B] Two additional examples of two integrated circuits coupled together are shown.
[0061] [Figure 54] 1 illustrates a flow diagram of one embodiment of a method for transferring data between two coupled integrated circuits.
[0062] [Figure 55] 1 illustrates a flow diagram of one embodiment of a method for routing signal data between an external interface and an on-chip router within an integrated circuit.
[0063] [Figure 56] FIG. 1 is a block diagram of an embodiment of multiple systems-on-chip (SOCs), where a given SOC includes multiple memory controllers.
[0064] [Figure 57] FIG. 2 is a block diagram illustrating one embodiment of a memory controller and physical / logical layout on a SOC.
[0065] [Figure 58] FIG. 1 is a block diagram of one embodiment of a binary decision tree for determining which memory controller services a particular address.
[0066] [Figure 59] FIG. 2 is a block diagram illustrating one embodiment of a multiple memory location configuration register.
[0067] [Figure 60] 1 is a flowchart illustrating the operation of one embodiment of a SOC during boot / power-on.
[0068] [Figure 61] 1 is a flowchart illustrating the operation of one embodiment of a SOC for routing memory requests.
[0069] [Figure 62] 10 is a flowchart illustrating the operation of one embodiment of a memory controller in response to a memory request.
[0070] [Figure 63] FIG. 10 is a flow diagram illustrating the operation of one embodiment of a monitoring system operation to determine memory folding.
[0071] [Figure 64] 10 is a flowchart illustrating the operation of one embodiment of collapsing a memory slice.
[0072] [Figure 65] 10 is a flowchart illustrating the operation of one embodiment for expanding a memory slice.
[0073] [Figure 66] 1 is a flow chart illustrating one embodiment of a method for memory folding.
[0074] [Figure 67] 1 is a flow chart illustrating one embodiment of a method for hashing a memory address.
[0075] [Figure 68] 1 is a flow chart illustrating one embodiment of a method for forming a compressed pipe address.
[0076] [Figure 69] FIG. 1 is a block diagram of an embodiment of an integrated circuit design that supports whole and partial instantiations.
[0077] [Figure 70] 70A-70C are various embodiments of full and partial instances of the integrated circuit shown in FIG. 69. [Figure 71] 70A-70C are various embodiments of full and partial instances of the integrated circuit shown in FIG. 69. [Figure 72] 70A-70C are various embodiments of full and partial instances of the integrated circuit shown in FIG. 69.
[0078] [Figure 73]FIG. 70 is a block diagram of one embodiment of the integrated circuit shown in FIG. 69 having a local clock source in each sub-region of the integrated circuit.
[0079] [Figure 74] FIG. 70 is a block diagram of one embodiment of the integrated circuit shown in FIG. 69 having local analog pads in each sub-region of the integrated circuit.
[0080] [Figure 75] FIG. 69 is a block diagram of one embodiment of an integrated circuit, with block-out areas at the corners of each sub-region and areas for interconnect "bumps" that exclude areas near the edges of each sub-region.
[0081] [Figure 76] FIG. 2 is a block diagram illustrating one embodiment of a stub and corresponding circuit components.
[0082] [Figure 77] FIG. 1 is a block diagram illustrating one embodiment of a pair of integrated circuits and certain further details of the pair of integrated circuits.
[0083] [Figure 78] FIG. 1 is a flow diagram illustrating an embodiment of an integrated circuit design method.
[0084] [Figure 79] FIG. 1 is a block diagram illustrating a test bench arrangement for testing whole and partial instances.
[0085] [Figure 80] FIG. 1 is a block diagram illustrating a test bench layout for component level testing.
[0086] [Figure 81] FIG. 1 is a flow diagram illustrating one embodiment of a method for designing and manufacturing integrated circuits.
[0087] [Figure 82]1 is a flowchart illustrating one embodiment of a method for manufacturing an integrated circuit.
[0088] [Figure 83] FIG. 1 is a block diagram of one embodiment of a system.
[0089] [Figure 84] FIG. 1 is a block diagram of one embodiment of a computer-accessible storage medium.
[0090] While the embodiments described in this disclosure may be susceptible to various modifications and alternative forms, specific embodiments thereof are shown by way of example in the drawings and are herein described in detail. It should be understood, however, that the drawings and detailed description relating to the drawings are not intended to limit the embodiments to the particular forms disclosed, but rather the intent is to cover all modifications, equivalents, and alternatives falling within the spirit and scope of the appended claims. The headings used herein are for organizational purposes only and are not intended to be used to limit the scope of the description. DETAILED DESCRIPTION OF THE INVENTION
[0091] An SOC may include most of the elements necessary to implement a complete computer system, although some elements (e.g., system memory) may be external to the SOC. For example, an SOC may include one or more general-purpose processor cores, one or more graphics processing units, and one or more other peripheral devices (such as application-specific accelerators, I / O interfaces, or other types of devices) separate from the processor cores and graphics processing units. The SOC may further include one or more memory controller circuits configured to interface with the system memory, as well as an interconnect fabric configured to provide communication between the memory controller circuit(s), the processor core(s), the graphics processing unit(s), and the peripheral device(s).
[0092] The design requirements of a given SOC are often dictated by the power constraints and performance requirements of the particular application the SOC targets. For example, an SOC for a smartwatch device may have stringent power consumption requirements because the form factor of such a device limits the available battery size and therefore the device's maximum usage time. At the same time, the small size of such a device may limit the number of peripherals the SOC needs to support as well as the computational requirements of the applications it runs. In contrast, an SOC for a mobile phone application may have a larger available battery and therefore a larger power budget, but it is also expected to have more complex peripherals and greater graphics and general computational requirements. Therefore, such an SOC is expected to be larger and more complex than designs for smaller devices.
[0093] This comparison can be optionally extended to other applications, for example, wearable computing solutions such as augmented and / or virtual reality systems may be expected to present greater computing requirements than less complex devices, as well as devices for desktop and / or rack-mounted computer systems.
[0094] As systems are built for larger applications, multiple chips can be used together to scale performance and form a "system of chips." We will continue to refer to these systems as "SOCs" herein, whether they are a single physical chip or multiple physical chips. The principles in this disclosure are equally applicable to multi-chip and single-chip SOCs.
[0095] The inventors' insight in this disclosure is that the computational requirements for the various applications described above and the corresponding SOC complexity tend to scale from small to large. If a SOC can be designed to easily scale in physical complexity, the core SOC design can be easily tailored for a wide variety of applications while leveraging design reuse and reducing duplication of effort. Such a SOC also provides a consistent view of functional blocks, such as processing cores or media blocks, making their integration into the SOC easier and further reducing effort. That is, the same functional block (or "IP") design can be used essentially unmodified in SOCs ranging from small to large. Furthermore, if such an SOC design can scale in a manner that is largely or completely transparent to the software running on the SOC, developing software applications that can easily scale across different resource versions of the SOC will be greatly simplified. Applications can be written only once and automatically run correctly on many different systems, again from small to large. When the same software scales across different resource versions, the software presents the same interface to the user, which is an additional benefit of scaling.
[0096] The present disclosure contemplates such scalable SOC designs. In particular, a core SOC design may include a set of processor cores, a graphics processing unit, a memory controller circuit, peripheral devices, and an interconnect fabric configured to interconnect them. Furthermore, the processor cores, the graphics processing unit, and the peripheral devices may be configured to access system memory via a unified memory architecture. The unified memory architecture includes a unified address space, which enables heterogeneous agents in a system (e.g., processors, graphics processing units, peripheral devices) to cooperate in a simple and high-performance manner. That is, rather than dedicating a private address space to the graphics processing unit and requiring data to be copied to and from that private address space, the graphics processing unit, the processor cores, and other peripheral devices can, in principle, share access to any memory address accessible by the memory controller circuit (subject, in some embodiments, to a privilege model or other security features that restrict access to certain types of memory content). Additionally, the unified memory architecture provides the same memory semantics (e.g., a common set of memory semantics) as the complexity of the SOC scales to meet the requirements of different systems. For example, memory semantics may include memory ordering properties, quality of service (QoS) support and attributes, memory management unit definitions, cache coherency capabilities, etc. The unified address space may be a physical address space that is distinct from the virtual address space, or may be a physical address space, or both.
[0097] The architecture remains the same as the SOC is scaled, but various implementation choices are possible. For example, virtual channels may be used as part of QoS support, but if not all QoS can be guaranteed in a given system, a subset of the supported virtual channels may be implemented. Depending on the bandwidth and latency characteristics required in a given system, different interconnect fabric implementations may be used. Additionally, some features may not be necessary in smaller systems (e.g., address hashing to balance memory traffic to various memory controllers may not be needed in a single-memory controller system). A hashing algorithm may not be important when having a small number of memory controllers (e.g., two or four), but will contribute more to system performance when a larger number of memory controllers are used.
[0098] Additionally, some of the components may be designed with scalability in mind, for example, a memory controller may be designed to scale up by adding additional memory controllers to the fabric, each with its own address space, memory cache, and part of the coherency tracking logic.
[0099] More specifically, embodiments of SOC designs are disclosed that can be easily scaled down and up in complexity. For example, in an SOC, processor cores, graphics processing units, fabric, and other devices can be arranged and configured such that the size and complexity of the SOC can be easily reduced prior to manufacturing by "shearing" the SOC along defined axes so that the resulting design includes only a subset of the components defined in the original design. When buses that would otherwise extend to the removed portion of the SOC are properly terminated, a reduced-complexity version of the original SOC design can be obtained with relatively little design and verification effort. A unified memory architecture can facilitate the deployment of applications in the reduced-complexity design, which in some cases can simply operate without substantial modification.
[0100] As previously mentioned, embodiments of the disclosed SOC designs can be configured to scale up in complexity. For example, multiple instances of a single-die SOC design can be interconnected, resulting in a system with resources that are two, three, four, or more times greater than the single-die design. Again, the unified memory architecture and coherent SOC architecture can facilitate the development and deployment of software applications that scale to use the additional computing resources provided by these multi-die system configurations.
[0101] FIG. 1 is a block diagram of one embodiment of a scalable SOC 10 coupled to one or more memories, such as memories 12A-12m. SOC 10 may include multiple processor clusters 14A-14n. Processor clusters 14A-14n may include one or more processors (P) 16 coupled to one or more caches (e.g., cache 18). Processor 16 may include general-purpose processors (e.g., central processing units or CPUs) as well as other types of processors, such as graphics processing units (GPUs). SOC 10 may include one or more other agents 20A-20p. One or more other agents 20A-20p may include, for example, various peripheral circuits / devices and / or bridges, such as input / output agents (IOAs), coupled to one or more peripheral devices / circuits. SOC 10 may include one or more memory controllers 22A-22m, each coupled to a respective memory device or circuit 12A-12m during use. In one embodiment, each memory controller 22A-22m may include a coherency controller circuit (more simply, a "coherency controller" or "CC") coupled to a directory (the coherency controller and directory are not shown in FIG. 1 ). Additionally, die-to-die (D2D) circuitry 26 is shown within SOC 10. Memory controllers 22A-22m, other agents 20A-20p, D2D circuitry 26, and processor clusters 14A-14n may be coupled to interconnect 28 for communication between the various components 22A-22m, 20A-20p, 26, and 14A-14n. As indicated by the names, the components of SOC 10 may, in one embodiment, be integrated onto a single integrated circuit "chip." In other embodiments, the various components may be external to SOC 10 on other chips, or may even be discrete components. Any amount of integrated or discrete components may be used. In one embodiment, a subset of processor clusters 14A-14n and memory controllers 22A-22m may be implemented on one of multiple integrated circuit chips coupled together to form the components shown in SOC 10 of FIG.
[0102] The D2D circuitry 26 may be an off-chip interconnect coupled to an interconnect fabric 28 and configured to couple the interconnect fabric 28 to a corresponding interconnect fabric 28 on another instance of the SOC 10. The interconnect fabric 28 and the off-chip interconnect 26 provide an interface that transparently connects one or more memory controller circuits, processor cores, graphics processing units, and peripheral devices in either a single instance of an integrated circuit or two or more instances of an integrated circuit. That is, via the D2D circuitry 26, the interconnect fabric 28 extends across two or more integrated circuit dies, and communications are routed between sources and destinations transparently to the location of the sources and destinations on the integrated circuit dies. The interconnect fabric 28 extends across two or more integrated circuit dies using hardware circuitry (e.g., the D2D circuitry 26) to automatically route communications between sources and destinations, regardless of whether the sources and destinations are on the same integrated circuit die.
[0103] Thus, D2D circuitry 26 supports scalability of SOC 10 to two or more instances of SOC 10 in a system. When two or more instances are included, the unified memory architecture, including the unified address space, extends across two or more instances of the integrated circuit die transparent to software running on the processor cores, graphics processing units, or peripheral devices. Similarly, in the case of a single instance of an integrated circuit die in a system, the unified memory architecture, including the unified address space, maps to the single instance transparent to software. When two or more instances of the integrated circuit die are included in a system, the system set of processor cores 16, graphics processing units, peripheral devices 20A-20p, and interconnect fabric 28 is distributed across two or more integrated circuit dies, again transparent to software.
[0104] As described above, each processor cluster 14A-14n may include one or more processors 16. The processor 16 may function as the central processing unit (CPU) of the SOC 10. The system's CPU includes one or more processors that execute the system's main control software, such as an operating system. Generally, the software executed by the CPU during use may control other components of the system to achieve the system's desired functionality. The processor may also execute other software, such as application programs. The application programs may provide user functionality and may rely on the operating system for low-level device control, scheduling, memory management, etc. Thus, the processors may also be referred to as application processors. Additionally, the processor 16 in a given cluster 14A-14n may be a GPU, as described above, and may implement a graphics instruction set optimized for rendering, shading, and other operations. The clusters 14A-14n may further include other hardware, such as a cache 18 and / or an interface to other components of the system (e.g., an interface to the interconnect 28). Other coherent agents may include processors that are not CPUs or GPUs.
[0105] Generally, a processor may include any circuitry and / or microcode configured to execute instructions defined in an instruction set architecture implemented by the processor. A processor may include a processor core implemented on an integrated circuit with other components as a system-on-chip (SOC) or other level of integration. A processor may further include a discrete microprocessor, a processor core and / or microprocessor integrated in a multi-chip module implementation, a processor implemented as multiple integrated circuits, etc. The number of processors 16 in a given cluster 14A-14n may differ from the number of processors 16 in another cluster 14A-14n. Generally, one or more processors may be included. Furthermore, the processors 16 may differ in microarchitecture implementation, performance and power characteristics, etc. In some cases, processors may also differ in the instruction set architecture they implement, their functionality (e.g., CPU, graphics processing unit (GPU) processor, microcontroller, digital signal processor, image signal processor, etc.), etc.
[0106] Cache 18 may have any capacity and configuration, such as set associative, direct mapped, or fully associative. Cache block size may be any desired size (e.g., 32 bytes, 64 bytes, 128 bytes, etc.). A cache block may be the unit of allocation and deallocation in cache 18. Furthermore, a cache block may be an address space (e.g., an aligned coherence granule-sized segment of a memory unit) within which coherency is maintained in this embodiment. A cache block may also be referred to as a cache line in some cases.
[0107] The memory controllers 22A-22m may generally include circuitry for receiving memory operations from other components of the SOC 10 and accessing the memories 12A-12m to complete the memory operations. The memory controllers 22A-22m may be configured to access any type of memory 12A-12m. More specifically, the memories 12A-12m may be any type of memory device that can be mapped as random access memory. For example, the memories 12A-12m may be static random access memory (SRAM), double data rate (DRAM) such as synchronous DRAM (SDRAM) including dynamic RAM (DDR, DDR2, DDR3, DDR4, etc.), non-volatile memory, graphics DRAM such as graphics DDR DRAM (GDDR), and high bandwidth memory (HBM). Low-power / mobile versions of DDR DRAM (e.g., LPDDR, mDDR, etc.) may be supported. The memory controllers 22A-22m may include queues for memory operations to order (and potentially reorder) the operations and present them to the memories 12A-12m. The memory controllers 22A-22m may further include data buffers for storing write data awaiting writing to memory and read data awaiting return to the source of the memory operation (if data is not provided from a snoop). In some embodiments, the memory controllers 22A-22m may include a memory cache for storing recently accessed memory data. In SOC implementations, for example, a memory cache can reduce power consumption in the SOC by avoiding re-accessing data from the memories 12A-12m if it is expected to be accessed again soon. In some cases, a memory cache may also be referred to as a system cache, as opposed to a private cache, such as cache 18 or a cache in the processor 16, that serves only a particular component. Furthermore, in some embodiments, the system cache need not be located within the memory controllers 22A-22m.Thus, there may be one or more levels of cache between the processor cores, graphics processing unit, peripheral devices, and system memory. One or more memory controller circuits 22A-22m may include a respective memory cache interposed between the interconnect fabric and the system memory, each memory cache being one of the one or more levels of cache.
[0108] The other agents 20A-20p may generally include various additional hardware functions (e.g., "peripherals," "peripheral devices," or "peripheral circuits") included with SOC C10. For example, peripherals may include video peripherals such as an image signal processor configured to process image capture data from a camera or other image sensor, video encoder / decoders, scalers, rotators, blenders, etc. Peripherals may include audio peripherals such as microphones, speakers, interfaces to microphones and speakers, audio processors, digital signal processors, and mixers. Peripherals may include interface controllers for various interfaces external to SOC 10, including interfaces such as Universal Serial Bus (USB), Peripheral Component Interconnect (PCI), including PCI Express (PCIe), serial and parallel ports, etc. Peripherals may include networking peripherals such as media access controllers (MACs). Any set of hardware may be included. In one embodiment, the other agents 20A-20p may also include bridges to a set of peripherals, such as IOAs, described below. In one embodiment, the peripheral device includes one or more of an audio processing device, a video processing device, a machine learning accelerator circuit, a matrix arithmetic accelerator circuit, a camera processing circuit, a display pipeline circuit, a non-volatile memory controller, a peripheral component interconnect controller, a security processor, or a serial bus controller.
[0109] Interconnect 28 may be any communications interconnect and protocol for communicating between components of SOC 10. Interconnect 28 may be bus-based, including shared bus configurations, crossbar configurations, and hierarchical buses with bridges. Interconnect 28 may be packet-based or circuit-switched, or may be hierarchical with bridges, crossbar, point-to-point, or other interconnects. Interconnect 28, in one embodiment, may include multiple independent communications fabrics.
[0110] In one embodiment, when two or more instances of an integrated circuit die are included in the system, the system may further comprise at least one interposer device configured to couple buses of the interconnect fabric across the two or more integrated circuit dies. In one embodiment, a given integrated circuit die includes a power manager circuit configured to manage the local power state of the given integrated circuit die. In one embodiment, when two or more instances of an integrated circuit die are included in the system, a separate power manager is configured to manage the local power state of the integrated circuit die, and at least one of the two or more integrated circuit dies includes another power manager circuit configured to synchronize the power manager circuits.
[0111] In general, the number of each component 22A-22m, 20A-20p, and 14A-14n may vary from embodiment to embodiment, and any number may be used. As indicated by the "m," "p," and "n" suffixes, the number of components of one type may differ from the number of components of another type. However, the number of a given type may be the same as the number of other types. Furthermore, while the system of FIG. 1 is shown with multiple memory controllers 22A-22m, embodiments having a single memory controller 22A-22m are also contemplated.
[0112] The concept of scalable SOC design is simple to explain but difficult to implement. Numerous innovations have been developed in support of this effort, which are described in more detail below. In particular, FIGS. 2-14 include further details of an embodiment of a communications fabric 28. FIGS. 15-26 illustrate an embodiment of a scalable interrupt structure. FIGS. 27-43 illustrate an embodiment of a scalable cache coherency mechanism that may be implemented among coherent agents in a system, including processor clusters 14A-14n and one or more directory / coherency control circuits. In one embodiment, the directory and coherency control circuits are distributed among multiple memory controllers 22A-22m, with each directory and coherency control circuit configured to manage cache coherency for the portion of the address space mapped to the memory devices 12A-12m to which a given memory controller is coupled. FIGS. 44-48 illustrate an embodiment of an IOA bridge for one or more peripheral circuits. FIGS. 49-55 illustrate further details of an embodiment of a D2D circuit 26. Figures 56-68 illustrate embodiments of a hashing scheme that distributes address space across multiple memory controllers 22A-22m. Figures 69-82 illustrate embodiments of a design methodology that supports multiple tapeouts of scalable SOC 10 for different systems based on the same design database.
[0113] The various embodiments described below and above may be used in any desired combination to form embodiments of the present disclosure. In particular, any subset of embodiment features from any of the embodiments may be combined to form embodiments that include less than all of the features and / or embodiments described in any given embodiment. All such embodiments are contemplated embodiments of the scalable SOC described herein. fabric
[0114] 2-14 illustrate various embodiments of interconnect fabric 28. Based on this description, a system is contemplated that includes multiple processor cores, multiple graphics processing units, multiple peripheral devices distinct from the processor cores and the graphics processing units, one or more memory controller circuits configured to interface with system memory, and an interconnect fabric configured to provide communication between the processor cores, the graphics processing units, and the peripheral devices, the interconnect fabric including at least two networks having heterogeneous operating characteristics. In one embodiment, the interconnect fabric includes at least two networks having heterogeneous interconnect topologies. The at least two networks may include coherent networks interconnecting the processor cores and one or more memory controller circuits. More specifically, the coherent network interconnects coherent agents, and the processor cores may be coherent agents, or the processor clusters may be coherent agents. The at least two networks may include relaxed ordering networks coupled to the graphics processing units and one or more memory controller circuits. In one embodiment, the peripheral devices include a subset of devices, the subset including one or more of a machine learning accelerator circuit or a relaxed-order bulk media device, and the relaxed-order network further couples the subset of devices to one or more memory controller circuits. At least two networks may include an input / output network coupled to interconnect the peripheral devices and the one or more memory controller circuits. The peripheral devices include one or more real-time devices.
[0115] In one embodiment, the at least two networks include a first network that includes one or more characteristics for reducing latency compared to a second network of the at least two networks. For example, the one or more characteristics may include a shorter route on a surface area of the integrated circuit than the second network. The one or more characteristics may include wiring for the first interconnect in a metal layer that provides lower latency characteristics than wiring for the second interconnect.
[0116] In one embodiment, the at least two networks include a first network that includes one or more characteristics for increasing bandwidth relative to a second network of the at least two networks. For example, the one or more characteristics include wider interconnects relative to the second network. The one or more characteristics include wiring in a metal layer that is farther from a surface of a substrate on which the system is implemented than wiring for the second network.
[0117] In one embodiment, the interconnect topology adopted by the at least two networks includes at least one of a star topology, a mesh topology, a ring topology, a tree topology, a fat tree topology, a hypercube topology, or a combination of one or more of the topologies. In another embodiment, the at least two networks are physically and logically independent. In yet another embodiment, the at least two networks are physically separate in a first mode of operation, and a first of the at least two networks and a second of the at least two networks are virtual in a second mode of operation and share a single physical network.
[0118] In one embodiment, a SOC is integrated on a semiconductor die. The SOC includes multiple processor cores, multiple graphics processing units, multiple peripheral devices, one or more memory controller circuits, and an interconnect fabric configured to provide communication between the one or more memory controller circuits and the processor cores, the graphics processing unit, and the peripheral devices. The interconnect fabric includes at least a first network and a second network, where the first network includes one or more characteristics for reducing latency compared to a second of the at least two networks. For example, the one or more characteristics include a route for the first network on the surface of the semiconductor die that is shorter than a route for the second network. In another example, the one or more characteristics include wiring in a metal layer that has lower latency characteristics than a wiring layer used for the second network. In one embodiment, the second network includes one or more second characteristics for increasing bandwidth compared to the first network. For example, the one or more second characteristics may include wider interconnects compared to the second network (e.g., more wires per interconnect than the first network). The one or more second characteristics may include wiring in a metal layer that is denser than the wiring layer used for the first network.
[0119] In one embodiment, a system-on-chip (SOC) may include multiple independent networks. The networks may be physically independent (e.g., having dedicated wires and other circuitry forming the networks) or logically independent (e.g., communications provided by agents in the SOC may be logically defined to be transmitted on selected ones of the networks and may not be affected by transmissions on the other networks). In some embodiments, a network switch may be included to transmit packets on a given network. The network switch may be physically part of the network (e.g., each network may have its own dedicated network switch). In other embodiments, the network switch may be shared between the physically independent networks, thus ensuring that communications received on one of the networks remain on that network.
[0120] By providing physically and logically independent networks, high bandwidth can be achieved through parallel communication on different networks. Furthermore, different traffic can be carried on different networks, and thus a given network can be optimized for a given type of traffic. For example, a processor, such as a central processing unit (CPU) in a SOC, may be sensitive to memory latency and may cache data that is expected to be coherent between the processor and memory. Therefore, a CPU network can be provided in which the CPU and memory controller in the system are agents. The CPU network can be optimized to provide low latency. For example, in one embodiment, there may be virtual channels for low-latency requests and bulk requests. Low-latency requests may be favored over bulk requests in transferring around the fabric and by the memory controller. The CPU network can also support cache coherency using defined messages and protocols to communicate coherently. Another network can be an input / output (I / O) network. This network can be used by various peripheral devices (“peripherals”) to communicate with memory. The network can support the bandwidth required by the peripherals and can also support cache coherency. However, I / O traffic can often have significantly higher latency than CPU traffic. By separating I / O traffic from memory traffic, CPU traffic can be less affected by I / O traffic. The CPU may be included as an agent on the I / O network to manage coherency and communicate with peripherals. Yet another network, in one embodiment, may be a relaxed-ordering network. Both the CPU and I / O networks can support ordering models between communications on those networks that provide ordering expected by the CPU and peripherals.However, the relaxed ordering network may be non-coherent and may not enforce as many ordering constraints. The relaxed ordering network may be used by a graphics processing unit (GPU) to communicate with a memory controller. Thus, the GPU may have dedicated bandwidth in the network and may not be constrained by the ordering required by the CPU and / or peripherals. Other embodiments may employ any subset of the above networks and / or any additional networks, as desired.
[0121] A network switch may be a circuit configured to receive communications on a network and forward the communications on the network in the direction of the communication's destination. For example, a communication provided by a processor may be transmitted to a memory controller that controls a memory mapped to the communication's address. At each network switch, the communication may be transmitted in a forward direction toward the memory controller. If the communication is a read, the memory controller may return the data to the source, and each network switch may forward the data on the network toward the source. In one embodiment, a network may support multiple virtual channels. A network switch may employ dedicated resources (e.g., buffers) for each virtual channel so that communications on the virtual channels may remain logically independent. A network switch may also employ arbitration circuitry to select a communication from among the buffered communications to forward on the network. Virtual channels may be channels that physically share the network but are logically independent on the network (e.g., communications in one virtual channel do not block the progress of communications on another virtual channel).
[0122] An agent may generally be any device (e.g., processor, peripheral, memory controller, etc.) that can source and / or sink communications on a network. A source agent generates (sources) communications and a destination agent receives (sinks) communications. A given agent may be a source agent for some communications and a destination agent for other communications.
[0123] Referring now to the drawings, FIG. 2 is a general diagram illustrating physically and logically independent networks. FIGS. 3-5 are examples of various network topologies. FIG. 6 is an example of an SOC with multiple physically and logically independent networks. FIGS. 7-9 show the various networks of FIG. 6 separately for further clarity. FIG. 10 is a block diagram of a system including two semiconductor dies, illustrating network scalability for multiple instances of an SOC. FIGS. 11 and 12 are example agents shown in more detail. FIG. 13 illustrates various virtual channels and communication types and which networks in FIG. 6 they apply to. FIG. 14 is a flowchart illustrating a method. Further details are provided below with reference to the drawings.
[0124] FIG. 2 is a block diagram of a system including one embodiment of multiple networks interconnecting agents. While FIG. 1 shows agents A10A, A10B, and A10C, in various embodiments, any number of agents may be included. Agents A10A-A10B are coupled to network A12A, and agents A10A and A10C are coupled to network A12B. In various embodiments, any number of networks A12A-A12B may be included. Network A12A includes multiple network switches, including network switches A14A, A14AB, A14AM, and A14AN (collectively network switches A14A). Similarly, network A12B includes multiple network switches, including network switches A14BA, A14BB, A14BM, and A14BN (collectively network switches A14B). Different networks A12A-A12B may include different numbers of network switches A14A, and 12A-A12B include physically separate connections ("wires," "buses," or "interconnects") shown as various arrows in FIG. 2.
[0125] Networks A12A-A12B are physically and logically separate because each network A12A-A12B has its own physically and logically separate interconnect and network switches. Communications on network A12A are not affected by communications on network A12B, and vice versa. Even the bandwidth on the interconnects within each network A12A-A12B is separate and independent.
[0126] Optionally, the agents A10A-A10C may include or be coupled to network interface circuitry (respectively referenced A16A-A16C). Some agents A10A-A10C may include or be coupled to a network interface A16A-A16C, while other agents A10A-A10C may not include or be coupled to a network interface A16A-A16C. The network interfaces A16A-A16C may be configured to transmit and receive traffic over the networks A12A-A12B on behalf of the corresponding agents A10A-A10C. The network interfaces A16A-A16C may be configured to translate or modify communications issued by the corresponding agents A10A-A10C to conform to the protocols / formats of the networks A12A-A12B, remove the modifications, or convert received communications to the protocols / formats used by the agents A10A-A10C. Thus, network interfaces A16A-A16C may be used for agents A10A-A10C that are not specifically designed to interface directly with networks A12A-A12B. In some cases, agents A10A-A10C may communicate on more than one network (e.g., agent A10A communicates on both networks A12A-A12B of FIG. 1). The corresponding network interface A16A may be configured to separate traffic issued by agent A10A to networks A12A-A12B according to which network A12A-A12B each communication is assigned to, and the network interface A16A may be configured to combine traffic received from networks A12A-A12B for the corresponding agent A10A. Any mechanism for determining which network A12A-A12B carries a given communication may be used (e.g., in various embodiments, based on the type of communication, the destination agent A10B-A10C for the communication, the address, etc.).
[0127] Because network interface circuitry is optional and not much is required for an agent to directly support networks A12A-A12B, the network interface circuitry is omitted from the remainder of the drawings for simplicity, but it should be understood that network interface circuitry may be employed by any agent, or a subset of agents, or all of the agents in any of the illustrated embodiments.
[0128] In one embodiment, the system of Figure 2 may be implemented as a SOC, and the components shown in Figure 2 may be formed on a single semiconductor substrate die. The circuitry included in the SOC may include a plurality of agents A10C and a plurality of network switches A14A-A14B coupled to the plurality of agents A10A-A10C. The plurality of network switches A14A-A14B are interconnected to form a plurality of physically and logically independent networks A12A-A12B.
[0129] Because the networks A12A-A12B are physically and logically independent, different networks can have different topologies. For example, a given network may have a ring, mesh, tree, star, a fully connected set of network switches (e.g., switches directly connected to each other within the network), a shared bus with multiple agents coupled to the bus, or a hybrid of any one or more of the topologies. Each network A12A-A12B may employ a topology that provides, for example, desired bandwidth and latency attributes for that network, or any desired attributes for the network. Thus, in general, a SOC may include a first network constructed according to a first topology and a second network constructed according to a second topology that is different from the first topology.
[0130] Figures 3-5 illustrate exemplary topologies. Figure 3 is a block diagram of one embodiment of a network coupling agents A10A-A10C using a ring topology. In the example of Figure 3, the ring is formed from network switches A14AA-A14AH. Agent A10A is coupled to network switch A14AA. Agent A10B is coupled to network switch A14AB, and agent A10C is coupled to network switch A14AE.
[0131] In a ring topology, each network switch A14AA-A14AH may be connected to two other network switches A14AA-A14AH, and the switches form a ring such that any network switch A14AA-A14AH can reach any other network switch in the ring by transmitting a communication in the direction of the other network switches on the ring. A given communication may pass through one or more intermediate network switches in the ring to reach a target network switch. When a given network switch A14AA-A14AH receives a communication from an adjacent network switch A14AA-A14AH on the ring, the given network switch may inspect the communication and determine that the agent A10A-A10C to which the given network switch is coupled is the destination of the communication. If so, the given network switch may terminate the communication and forward the communication to the agent. If not, the given network switch may forward the communication to the next network switch on the ring (e.g., another network switch A14AA-A14AH that is adjacent to the given network switch and is not the adjacent network switch from which the given network switch received the communication). There may be network switches adjacent to a given network switch, and there may be network switches to which the given network switch may directly transmit the communication without the communication traveling through any intermediate network switches.
[0132] FIG. 4 is a block diagram of one embodiment of a network using a mesh topology to couple agents A10A-A10P. As shown in FIG. 4, the network may include network switches A14AA-A14AH. Each network switch A14AA-A14AH is coupled to two or more other network switches. For example, as shown in FIG. 4, network switch A14AA is coupled to network switches A14AB and A14AE, network switch A14AB is coupled to network switches A14AA, A14AF, and A14AC, and so on. Thus, different network switches within the mesh network may be coupled to different numbers of other network switches. Furthermore, while the embodiment of FIG. 4 has a relatively symmetrical structure, other mesh networks may be asymmetrical, depending, for example, on various traffic patterns expected to prevail on the network. At each network switch A14AA-A14AH, one or more attributes of the received communication may be used to determine to which neighboring network switch A14AA-A14AH the receiving network switch A14AA-A14AH should transmit the communication (unless the communication is destined for the agent A10A-A10P to which the receiving network switch A14AA-A14AH is coupled, in which case the receiving network switch A14AA-A14AH may terminate the communication on the network and provide it to the destination agent A10A-A10P.) For example, in one embodiment, the network switches A14AA-A14AH may be programmed at system initialization to route communications based on various attributes.
[0133] In one embodiment, communications may be routed based on the destination agent. The routing may be configured to transport communications through the minimum number of network switches (the "shortest path") between the source agent and the destination agent that can be supported in the mesh topology. Alternatively, different communications from a given source agent to a given destination agent may take different paths through the mesh. For example, latency-sensitive communications may be transmitted over a shorter path, while less critical communications may take a different path to avoid consuming bandwidth on the short path; for example, the different path may be less heavily loaded in use.
[0134] 4 may be an example of a partially connected mesh, where at least some communications may pass through one or more intermediate network switches in the mesh. A fully connected mesh may have a connection from each network switch to every other network switch, so that no communications may be transmitted without traversing any intermediate network switches. In various embodiments, any level of interconnectivity may be used.
[0135] FIG. 5 is a block diagram of one embodiment of a network using a tree topology to couple agents A10A-A10E. Network switches A14A-A14AG are interconnected to form a tree in this example. The tree is a form of hierarchical network in which there are edge network switches (e.g., A14A, A14AB, A14AC, A14AD, and A14AG in FIG. 5) that couple to agents A10A-A10E and intermediate network switches (e.g., A14AE and A14AF in FIG. 5) that couple only to other network switches. A tree network can be used, for example, when a particular agent is often the destination of communications issued by other agents or is often the source agent of communications. Thus, for example, the tree network of FIG. 5 can be used for agent A10E, which is a primary source or destination for communications. For example, agent A10E may be a memory controller, which is often the destination of memory transactions.
[0136] There are many other possible topologies that may be used in other embodiments. For example, a star topology has a source / destination agent at the "center" of the network, and other agents may be coupled to the central agent directly or through a series of network switches. A star topology, like a tree topology, may be used when the central agent is often the source or destination of communications. A shared bus topology may also be used, as may a hybrid of two or more of any of the topologies.
[0137] FIG. 6 is a block diagram of one embodiment of a system-on-chip (SOC) A20 having multiple networks for one embodiment. For example, SOC A20 may be an instance of SOC 10 of FIG. 1. In the embodiment of FIG. 6, SOC A20 includes multiple processor clusters (P clusters) A22A-A22B, multiple input / output (I / O) clusters A24A-A24D, multiple memory controllers A26A-A26D, and multiple graphics processing units (GPUs) A28A-A28D. As implied by the name (SOC), the components shown in FIG. 6 (except for memories A30A-A30D in this embodiment) may be integrated onto a single semiconductor die or "chip." However, other embodiments may employ two or more dies combined or packaged in any desired manner. Additionally, although a particular number of P clusters A22A-A22B, I / O clusters A24-A24D, memory controllers A26A-A26D, and GPUs A28A-A28D are shown in the example of Figure 6, the number and arrangement of any of the above components may vary and may be more or less than that shown in Figure 6. Memories A30A-A30D are coupled to SOC A20, and more specifically, are coupled to memory controllers A26A-A26D, respectively, as shown in Figure 6.
[0138] In the illustrated embodiment, SOC A20 includes three physically and logically independent networks formed from multiple network switches A32, A34, and A36 as shown in FIG. 6, with interconnects therebetween shown as arrows between the network switches and other components. Other embodiments may include more or fewer networks. Network switches A32, A34, and A36 may be instances of network switches similar to network switches A14A-A14B described above with respect to FIGS. 2-5, for example. The multiple network switches A32, A34, and A36 are coupled to multiple P clusters A22A-A22B, multiple GPUs A28A-A28D, multiple memory controllers A26-A25B, and multiple I / O clusters A24A-A24D, as shown in FIG. 6. The P clusters A22A-A22B, the GPUs A28A-A28B, the memory controllers A26A-A26B, and the I / O clusters A24A-A24D may all be examples of agents that communicate over the various networks of the SOC A20. Other agents may also be included as desired.
[0139] In FIG. 6, a central processing unit (CPU) network is formed from a first subset of a plurality of network switches (e.g., network switch A32), with interconnects therebetween shown as short-dashed / long-dashed lines, such as reference numeral A38. The CPU network couples P clusters A22A-A22B and memory controllers 26A-A26D. The I / O network is formed from a second subset of a plurality of network switches (e.g., network switch A34), with interconnects therebetween shown as solid lines, such as reference numeral A40. The I / O network couples P clusters A22A-A22B, I / O clusters A24A-A24D, and memory controllers A26A-A26B. The relaxed ordering network is formed from a third subset of a plurality of network switches (e.g., network switch A36), with interconnects therebetween shown as short-dashed lines, such as reference numeral A42. The relaxed ordering network couples the GPUs A28A-A28D and the memory controllers A26A-A26D. In one embodiment, the relaxed ordering network may also couple selected ones of the I / O clusters A24A-A24D as well. As described above, the CPU network, the I / O network, and the relaxed ordering network are independent of one another (e.g., logically and physically independent). In one embodiment, protocols on the CPU network and the I / O network support cache coherency (e.g., the networks are coherent). The relaxed ordering network may not support cache coherency (e.g., the networks are non-coherent). The relaxed ordering network also has reduced ordering constraints compared to the CPU network and the I / O network. For example, in one embodiment, a set of virtual channels and subchannels within the virtual channels are defined for each network. In the case of the CPU and I / O networks, communications within the same virtual channel and subchannel between the same source agent and destination agent can be ordered. In a relaxed ordering network, communications between the same source and destination agents can be ordered.In one embodiment, only communications to the same address (at a given granularity, such as a cache block) between the same source and destination agents may be ordered. Because less strict ordering is enforced on a relaxed-order network, higher bandwidth may be achieved on average because transactions may be allowed to complete out of order, for example, if a newer transaction is ready to complete before an older transaction.
[0140] The interconnect between network switches A32, A34, and A36 can have any form and configuration in various embodiments. For example, in one embodiment, the interconnect may be a point-to-point, unidirectional link (e.g., a bus or serial link). Packets may be transmitted over the link, and the packet format may include data indicating the virtual channel and subchannel the packet is traveling on, a memory address, source and destination agent identifiers, data (if appropriate), etc. Multiple packets may form a given transaction. A transaction may be a complete communication between a source agent and a target agent. For example, a read transaction, depending on the protocol, may include a read request packet from the source agent to the target agent, one or more coherence message packets between the caching agent and the target agent and / or the source agent if the transaction is coherent, a data response packet from the target agent to the source agent, and possibly a completion packet from the source agent to the target agent. A write transaction may include a write request packet from the source agent to the target agent, one or more coherence message packets, similar to a read transaction if the transaction is coherent, and possibly a completion packet from the target agent to the source agent. In one embodiment, the write data may be included in the write request packet or may be transmitted from the source agent to the target agent in a separate write data packet.
[0141] The placement of agents in FIG. 6 may, in one embodiment, represent the physical placement of agents on a semiconductor die forming SOC A20. That is, FIG. 6 may be viewed in terms of the surface area of the semiconductor die, and the locations of the various components in FIG. 6 may approximate their physical locations by that area. Thus, for example, I / O clusters A24A-A24D may be placed within the semiconductor die area represented by the top of SOC A20 (as oriented in FIG. 6). P clusters A22A-A22B may be placed in the area represented by the portion of SOC A20 below and between the placement of I / O clusters A24A-A24D, as oriented in FIG. 6. GPUs A24A-A28D may be located in the center and extend toward the area represented by the bottom of SOC A20 as oriented in FIG. 6. Memory controllers A26A-A26D may be placed on the areas represented by the right and left of SOC A20 as oriented in FIG. 6.
[0142] In one embodiment, SOC A20 can be designed to couple directly to one or more other instances of SOC A20, logically combining a given network on the instance into one network over which an agent on one die can logically communicate to an agent on a different die via the network in the same way as another agent on the same die communicates. While latency may vary, communication may be performed in the same manner. Thus, as shown in FIG. 6, the network extends to the bottom of SOC A20 as oriented in FIG. 6. Interface circuitry (e.g., serializer / deserializer (SERDES) circuitry) not shown in FIG. 6 may be used to communicate across die boundaries to another die. Thus, the network may be scalable to two or more semiconductor dies. For example, two or more semiconductor dies may be configured as a single system in which the presence of multiple semiconductor dies is transparent to software running on the single system. In one embodiment, delays in die-to-die communication may be minimized as an aspect of software transparency to a multi-die system, such that inter-die communication typically does not incur significant additional latency compared to intra-die communication, hi other embodiments, the network may be a closed network that communicates only within the die.
[0143] As mentioned above, different networks may have different topologies. In the embodiment of FIG. 6, for example, the CPU and I / O networks may implement a ring topology, and the relaxed-ordering network may implement a mesh topology. However, other topologies may be used in other embodiments. FIGS. 7, 8, and 9 illustrate portions of SOC A30 that include different networks: CPU (FIG. 7), I / O (FIG. 8), and relaxed-ordering (FIG. 9). As seen in FIGS. 7 and 8, network switches A32 and A34 each form a ring when coupled to a corresponding switch on another die. If only a single die is used, a connection may be made between the two network switches A32 or A34 at the bottom of SOC A20 as oriented in FIGS. 7 and 8 (e.g., via external connections on pins of SOC A20). Alternatively, the two bottom network switches A32 or A34 may have a link between them, which can be used in a single-die configuration, or the network may operate in a daisy-chain topology.
[0144] 9 illustrates the connections of network switch A36 in a mesh topology between GPUs A28A-A28D and memory controllers A26A-A26D. As previously mentioned, in one embodiment, one or more of I / O clusters A24A-A24D may be coupled to a relaxed-order network that was good. For example, I / O clusters A24A-A24D that include video peripherals (e.g., display controllers, memory scalers / rotators, video encoders / decoders, etc.) may have access to a relaxed-order network for video data.
[0145] A network switch A36 near the bottom of the SOC A30, as oriented in FIG. 9, may include connections that can be routed to another instance of the SOC A30, allowing the CPU to span multiple dies, as described above with respect to the mesh network and I / O network. In a single-die configuration, routes that extend off-chip may not be used. FIG. 10 is a block diagram of a two-die system in which each network extends across two SOC dies A20A-A20B, forming networks that are logically the same even though they span two dies. Network switches A32, A34, and A36 have been removed in FIG. 10 for simplicity, and the relaxed ordering network is simplified to a line, but may be a mesh in one embodiment. The I / O network A44 is shown with a solid line, the CPU network A46 with a dashed-dotted line, and the relaxed ordering network A48 with a dashed line. The ring structure of networks A44 and A46 is also evident in FIG. 10. Although two dies are shown in Figure 10, other embodiments can employ three or more dies. The network, in various embodiments, may be daisy-chained together, fully connected with point-to-point links between teach die pairs, or any other connection structure.
[0146] In one embodiment, physically separating the I / O network from the CPU network may help the system provide low-latency memory access by the processor clusters A22A-A22B because I / O traffic may be offloaded to the I / O network. The networks use the same memory controller to access memory, and therefore the memory controller may be designed to prioritize memory traffic from the CPU network over memory traffic from the I / O network to some extent. The processor clusters A22A-A22B may also be part of the I / O network to access device space within the I / O clusters A24A-A24D (e.g., with programmed input / output (PIO) transactions). However, memory transactions initiated by the processor clusters A22A-A22B may be transmitted over the CPU network. Thus, the CPU clusters A22A-A22B may be examples of agents coupled to at least two of multiple physically and logically independent networks. The agent may be configured to generate a transaction to be transmitted and, based on the type of transaction (e.g., memory or PIO), select one of at least two of a plurality of physically and logically independent networks over which to transmit the transaction.
[0147] Various networks may include different numbers of physical and / or virtual channels. For example, an I / O network may have multiple request and completion channels, while a CPU network may have one request and one completion channel (or vice versa). The requests transmitted on a given request channel, when there are more than one, may be determined in any desired manner (e.g., by type of request, by priority of the request, to balance bandwidth across physical channels, etc.). Similarly, the I / O network and the CPU network may include snoop virtual channels that carry snoop requests, while the relaxed-ordering network may not include a snoop virtual channel because it is non-coherent in this embodiment.
[0148] FIG. 11 is a block diagram of one embodiment of an input / output (I / O) cluster A24A shown in further detail. The other I / O clusters A24B-A24D may be similar. In the embodiment of FIG. 11, the I / O cluster A24A includes peripherals A50 and A52, a peripheral interface controller A54, a local interconnect A56, and a bridge A58. The peripheral A52 may be coupled to an external component A60. The peripheral interface controller A54 may be coupled to a peripheral interface A62. The bridge A58 may be coupled to a network switch A34 (or a network interface coupled to the network switch A34).
[0149] The peripherals A50 and A52 may include any set of additional hardware functions (e.g., other than a CPU, GPU, and memory controller) included in the SOC A20. For example, the peripherals A50 and A52 may include video peripherals such as an image signal processor configured to process image capture data from a camera or other image sensor, a video encoder / decoder, a scaler, a rotator, a blender, a display controller, etc. The peripherals may include audio peripherals such as a microphone, a speaker, an interface to a microphone and a speaker, an audio processor, a digital signal processor, a mixer, etc. The peripherals may include networking peripherals such as a media access controller (MAC). The peripherals may include other types of memory controllers, such as a non-volatile memory controller. Some peripherals A52 may include on-chip and off-chip components A60. The peripheral interface controller A54 may include interface controllers for various interfaces A62 external to the SOC A20, including interfaces such as Universal Serial Bus (USB), Peripheral Component Interconnect (PCI) including PCI Express (PCIe), serial and parallel ports.
[0150] The local interconnect A56 may be the interconnect through which the various peripherals A50, A52, and A54 communicate. The local interconnect A56 may be different from the system-wide interconnect (e.g., CPU, I / O, and relaxed networks) shown in FIG. 6 . The bridge A58 may be configured to convert communications on the local interconnect to communications on the system-wide interconnect and vice versa. In one embodiment, the bridge A58 may be coupled to one of the network switches A34. The bridge A58 may also manage ordering among transactions issued from the peripherals A50, A52, and A54. For example, the bridge A58 may use a cache coherency protocol supported on the network to ensure ordering of transactions on behalf of the peripherals A50, A52, and A54, etc. Different peripherals A50, A52, and A54 may have different ordering requirements, and the bridge A58 may be configured to accommodate the different requirements. Bridge A58, in some embodiments, may also implement various performance-enhancing features. For example, bridge A58 may prefetch data for a given request. Bridge A58 may capture a coherent copy of a cache block to which one or more transactions from peripherals A50, A52, and A54 are directed (e.g., in an exclusive state), allowing the transactions to complete locally and enforcing ordering. Bridge A58 may speculatively capture an exclusive copy of one or more cache blocks targeted by a subsequent transaction, and may use the cache block to complete the subsequent transaction if the exclusive state is successfully maintained until the subsequent transaction can be completed (e.g., after satisfying any ordering constraints with previous transactions). Thus, in one embodiment, multiple requests in a cache block may be serviced from a cached copy.Various details can be found in U.S. Provisional Patent Application Nos. 63 / 170,868, filed April 5, 2021, 63 / 175,868, filed April 16, 2021, and 63 / 175,877, filed April 16, 2021. These patent applications are incorporated herein by reference in their entirety. If any of the incorporated material conflicts with the material expressly set forth herein, the material expressly set forth herein takes precedence.
[0151] Figure 12 is a block diagram of one embodiment of a processor cluster A22A. Other embodiments may be similar. In the embodiment of Figure 12, the processor cluster A22A includes one or more processors A70 coupled to a last level cache (LLC) A72. The LLC A72 may include interface circuitry for interfacing to the network switches A32 and A34 to transmit transactions over the CPU network and the I / O network, as needed.
[0152] Processor A70 may include any circuitry and / or microcode configured to execute instructions defined in an instruction set architecture implemented by processor A70. Processor A70 may have any microarchitecture implementation, performance and power characteristics, etc., for example, the processor may be in-order execution, out-of-order execution, superscalar, superpipelined, etc.
[0153] The LLC A72 and any cache in the processor A70 may have any capacity and configuration, such as set associative, direct-mapped, or fully associative. The cache block size may be any desired size (e.g., 32 bytes, 64 bytes, 128 bytes, etc.). A cache block may be the unit of allocation and deallocation in the LLC A70. Furthermore, a cache block may be the unit in which coherency is maintained in this embodiment. A cache block may also be referred to as a cache line in some cases. In one embodiment, a distributed directory-based coherency scheme may be implemented using a coherency point in each memory controller A26 in the system, where the coherency point applies to memory addresses mapped to the memory controller. The directory can track the state of cache blocks cached in any coherent agent. The coherency scheme may be scalable to many memory controllers, potentially across multiple semiconductor dies.For example, a coherency scheme may have the following features: precise directories for snoop filtering and contention resolution at the coherent and memory agents; ordering points (access order) determined at the memory agents; serialization points moving between the coherent and memory agents; secondary completion (invalidation acknowledgment) collection at the requesting coherent agent, tracked in a completion count provided by the memory agent; fill / snoop and snoop / victim-ack contention resolution handled at the coherent agent via directory state provided by the memory agent; separate primary / secondary shared state to aid contention resolution and restrict flight snoops to the same address / target; absorption of conflicting snoops at the coherent agents to avoid deadlocks without additional ack / conflict / retry messages or actions; and serialization minimization (one additional access mechanism per accessor to transfer ownership through the contention chain). additional message latency), message minimization (direct messages between involved agents, no additional messages to handle conflicts / races (e.g., no messages back to the memory agent), store conditional without excessive invalidation on race failure, minimal data transfer (only if dirty) and exclusive ownership requests intended to modify the entire cache line with the involved cache / directory state, separate snoop-back and snoop-forward message types (e.g., 3-hop and 4-hop protocols) to handle both cacheable and non-cacheable flows may be employed. Further details may be found in U.S. Provisional Patent Application No. 63 / 077,371, filed September 11, 2020, which is incorporated herein by reference in its entirety. If any of the incorporated material conflicts with the material expressly set forth herein, the material expressly set forth herein shall control.
[0154] FIG. 13 is a pair of tables A80 and A82 illustrating the virtual channels and traffic types and networks shown in FIGS. 6-9 used for one embodiment. As shown in Table A80, the virtual channels may include a bulk virtual channel, a low latency (LLT) virtual channel, a real-time (RT virtual channel), and a virtual channel for non-DRAM messages (VCP). The bulk virtual channel may be the default virtual channel for memory accesses. The bulk virtual channel may receive a lower quality of service than the LLT and RT virtual channels, for example. The LLT virtual channel may be used for memory transactions that require low latency for high-performance operation. The RT virtual channel may be used for memory transactions (e.g., video streams) that have latency and / or bandwidth requirements for correct operation. The VCP channel may be used to isolate traffic not directed to memory to prevent interference with memory transactions.
[0155] In one embodiment, bulk virtual channels and LLT virtual channels may be supported on all three networks (CPU, I / O, and relaxed-ordering). RT virtual channels may be supported on the I / O network but not on the CPU or relaxed-ordering networks. Similarly, VCP virtual channels may be supported on the I / O network but not on the CPU or relaxed-ordering networks. In one embodiment, VCP virtual channels may be supported on the CPU and relaxed-ordering networks only for transactions targeting network switches on that network (e.g., for configuration) and therefore may not be used during normal operation. Thus, as Table A80 shows, different networks may support different numbers of virtual channels.
[0156] Table A82 shows various traffic types and which networks carry them. Traffic types may include coherent memory traffic, non-coherent memory traffic, real-time (RT) memory traffic, and VCP (non-memory) traffic. Both the CPU and I / O networks may carry coherent traffic. In one embodiment, coherent memory traffic served by processor clusters A22A-A22B may be carried on the CPU network, and the I / O network may carry coherent memory traffic served by I / O clusters A24A-A24D. Non-coherent memory traffic may be carried on the relaxed-ordering network, and RT and VCP traffic may be carried on the I / O network.
[0157] FIG. 14 is a flowchart illustrating one embodiment of a method for initiating a transaction on a network. In one embodiment, an agent may generate a transaction to be transmitted (block A90). The transaction is transmitted over one of a plurality of physically and logically independent networks. A first network of the plurality of physically and logically independent networks is constructed according to a first topology, and a second network of the plurality of physically and logically independent networks is constructed according to a second topology different from the first topology. One of the plurality of physically and logically independent networks is selected for transmitting the transaction based on a type of transaction (block A92). For example, processor clusters A22A-A22B may transmit coherent memory traffic over a CPU network and PIO traffic over an I / O network. In one embodiment, the agent may select a virtual channel from a plurality of virtual channels supported on the selected network of the plurality of physically and logically independent networks based on one or more attributes of the transaction other than type (block A94). For example, the CPU may select an LLT virtual channel for a subset of memory transactions (e.g., the oldest memory transaction that is a cache miss, or the number of cache misses up to a threshold number after the bulk channel is selected). The GPU may select between an LLT virtual channel and a bulk virtual channel based on the urgency with which the data is needed. A video device may use an RT virtual channel as needed (e.g., a display controller may issue frame data reads on an RT virtual channel). A VCP virtual channel may be selected for transactions that are not memory transactions. The agent may transmit transaction packets on the selected network and virtual channel. In one embodiment, transaction packets in different virtual channels may take different paths through the network.In one embodiment, transaction packets may take different paths based on the type of transaction packet (e.g., request vs. response). In one embodiment, different paths may be supported for both different virtual channels and different types of transactions. Other embodiments may employ one or more additional attributes of transaction packets to determine the path through the network for those packets. Viewed another way, network switches forming the network may route different packets based on virtual channel, type, or any other attribute. A different path may refer to traversing at least one segment between network switches that is not traversed on another path, even if transaction packets using the different path are traveling from the same source to the same destination. Using different paths may provide load balancing in the network and / or reduced latency for transactions.
[0158] In one embodiment, a system includes a plurality of processor clusters, a plurality of memory controllers, a plurality of graphics processing units, a plurality of agents, and a plurality of network switches coupled to the plurality of processor clusters, the plurality of graphics processing units, the plurality of memory controllers, and the plurality of agents. A given processor cluster includes one or more processors. The memory controller is configured to control access to a memory device. A first subset of the plurality of network switches is interconnected to form a central processing unit (CPU) network between the plurality of processor clusters and the plurality of memory controllers. A second subset of the plurality of network switches is interconnected to form an input / output (I / O) network between the plurality of processor clusters, the plurality of agents, and the plurality of memory controllers. A third subset of the plurality of network switches is interconnected to form a relaxed-ordering network between the plurality of graphics processing units, selected agents from the plurality of agents, and the plurality of memory controllers. The CPU network, the I / O network, and the relaxed-ordering network are independent of one another. The CPU network and the I / O network are coherent. The relaxed-ordering network is non-coherent and has reduced ordering constraints compared to the CPU network and the I / O network. In one embodiment, at least one of the CPU network, the I / O network, and the relaxed-ordering network has some physical channels that are different from some physical channels on another one of the CPU network, the I / O network, and the relaxed-ordering network. In one embodiment, the CPU network is a ring network. In one embodiment, the I / O network is a ring network. In one embodiment, the relaxed-ordering network is a mesh network. In one embodiment, a first agent of the plurality of agents includes an I / O cluster including a plurality of peripheral devices. In one embodiment, the I / O cluster further includes a bridge coupled to the plurality of peripheral devices and further coupled to a first network switch in the second subset.In one embodiment, the system further comprises a CPU circuit configured to convert communications from a given agent into communications for a given network among a network interface network, an I / O network, and a relaxed ordering network, the network interface circuit coupled to one of a plurality of network switches in the given network.
[0159] In one embodiment, a system-on-chip (SOC) includes a semiconductor die having a circuit formed thereon. The circuit includes a plurality of agents and a plurality of network switches coupled to the plurality of agents. The plurality of network switches are interconnected to form a plurality of physically and logically independent networks. A first network of the plurality of physically and logically independent networks is constructed according to a first topology, and a second network of the plurality of physically and logically independent networks is constructed according to a second topology different from the first topology. In one embodiment, the first topology is a ring topology. In one embodiment, the second topology is a mesh topology. In one embodiment, coherency is implemented on the first network. In one embodiment, the second network is a relaxed-order network. In one embodiment, at least one of the plurality of physically and logically independent networks implements a first number of physical channels, and another of at least one of the plurality of physically and logically independent networks implements a second number of physical channels, the first number being different from the second number. In one embodiment, the first network includes one or more first virtual channels and the second network includes one or more second virtual channels. At least one of the one or more first virtual channels is different from the one or more second virtual channels. In one embodiment, the SOC further includes a network interface circuit configured to convert communications from a given agent of the plurality of agents into communications for a given network of the plurality of physically and logically independent networks. The network interface circuit is coupled to one of a plurality of network switches in the given network. In one embodiment, a first agent of the plurality of agents is coupled to at least two of the plurality of physically and logically independent networks. The first agent is configured to generate a transaction to be transmitted.The first agent is configured to select one of at least two of a plurality of physically and logically independent networks over which to transmit the transaction based on the type of transaction, in one embodiment, one of the at least two networks is an I / O network over which the I / O transaction is transmitted.
[0160] In one embodiment, the method includes generating a transaction at an agent coupled to a plurality of physically and logically independent networks, wherein a first network of the plurality of physically and logically independent networks is constructed according to a first topology and a second network of the plurality of physically and logically independent networks is constructed according to a second topology different from the first topology, and selecting one of the plurality of physically and logically independent networks over which to transmit the transaction based on a type of the transaction. In one embodiment, the method further includes selecting a virtual channel of a plurality of virtual channels supported by one of the plurality of physically and logically independent networks based on one or more attributes of the transaction other than type. interrupt
[0161] 15-26 illustrate various embodiments of a scalable interrupt structure. For example, in a system including two or more integrated circuit dies, a given integrated circuit die may include a local interrupt distribution circuit for distributing interrupts among processor cores in the given integrated circuit die. At least one of the two or more integrated circuit dies may include a global interrupt distribution circuit, where the local interrupt distribution circuit and the global interrupt distribution circuit implement a multi-level interrupt distribution scheme. In one embodiment, the global interrupt distribution circuit is configured to sequentially transmit interrupt requests to the local interrupt distribution circuits, and the local interrupt distribution circuits are configured to sequentially transmit interrupt requests to local interrupt destinations before responding to interrupt requests from the global interrupt distribution circuit.
[0162] A computing system generally includes one or more processors that function as central processing units (CPUs), along with one or more peripheral devices that implement various hardware functions. The CPU executes control software (e.g., an operating system) that controls the operation of the various peripheral devices. The CPU may also execute applications that provide user functionality within the system. Additionally, the CPU may execute software that interacts with the peripheral devices and performs various services on their behalf. Other processors not used as CPUs within the system (e.g., processors integrated into some peripheral devices) may also execute such software for the peripheral devices.
[0163] Peripheral devices can use interrupts to cause the processor to execute software on their behalf. Generally, peripheral devices issue interrupts by asserting an interrupt signal to an interrupt controller, which typically controls the interrupt going to the processor. The interrupt causes the processor to stop executing its current software task and save the state of the task so that it can be resumed later. The processor can load the state associated with the interrupt and begin executing an interrupt service routine. The interrupt service routine may be driver code for the peripheral device, or may transfer execution to driver code as needed. Generally, driver code is code provided for a peripheral device that is executed by the processor to control and / or configure the peripheral device.
[0164] The latency from assertion of an interrupt to servicing the interrupt can be important to performance and even functionality in a system. Furthermore, efficiently determining which CPU will service the interrupt and delivering the interrupt with minimal perturbation to the rest of the system can be important to both system performance and maintaining low power consumption. As the number of processors in a system increases, scaling interrupt delivery efficiently and effectively becomes even more important.
[0165] 15, a block diagram of one embodiment of a portion of a system B10 is shown, which includes an interrupt controller B20 coupled to a plurality of cluster interrupt controllers B24A-B24n. Each of the plurality of cluster interrupt controllers B24A-B24n is coupled to a respective plurality of processors B30 (e.g., processor clusters). The interrupt controller B20 is coupled to a plurality of interrupt sources B32.
[0166] When at least one interrupt is received by the interrupt controller B20, the interrupt controller B20 may be configured to attempt to deliver the interrupt (e.g., to the processor B30 to service the interrupt by executing software to record the interrupt for further servicing by an interrupt service routine and / or to provide processing requested by the interrupt via the interrupt service routine). In the system B10, the interrupt controller B20 may attempt to deliver the interrupt through the cluster interrupt controllers B24A-B24n. Each cluster controller B24A-B24n may be associated with a processor cluster and attempt to deliver the interrupt to the processor B30 in a respective one of the multiple processors forming the cluster.
[0167] More specifically, the interrupt controller B20 can be configured to attempt to deliver an interrupt in multiple iterations through the cluster interrupt controllers B24A-B24n. The interface between the interrupt controller B20 and each interrupt controller B24A-B24n can include a request / acknowledge (Ack) / not acknowledge (Nack) structure. For example, requests can be identified by soft, hard, and forced iterations in the illustrated embodiment. The first iteration (a "soft" iteration) can be signaled by asserting a soft request. The next iteration (a "hard" iteration) can be signaled by asserting a hard request. The last iteration (a "forced" iteration) can be signaled by asserting a force request. A given cluster interrupt controller B24A-B24n may respond to soft and hard repeats with an Ack response (indicating that a processor B30 in the processor cluster associated with the given cluster interrupt controller B24A-B24n has accepted the interrupt and is processing at least one interrupt) or a Nack response (indicating that a processor B30 in the processor cluster has rejected the interrupt). Forced repeats may not use an Ack / Nack response, but rather may continue to request an interrupt until the interrupt is serviced, as described in more detail below.
[0168] The cluster interrupt controllers B24A-B24n may also use a Request / Ack / Nack structure with the processor B30 to attempt to deliver an interrupt to a given processor B30. Based on a request from the cluster interrupt controllers B24A-B24n, the given processor B30 may be configured to determine whether the given processor B30 can interrupt the current instruction execution within a predetermined period of time. If the given processor B30 can commit the interrupt within that period of time, the given processor B30 may be configured to assert an Ack response. If the given processor B30 cannot commit the interrupt, the given processor B30 may be configured to assert a Nack response. The cluster interrupt controllers B24A-B24n may be configured to assert an Ack response to the interrupt controller B20 if at least one processor asserts an Ack response to the cluster interrupt controller B24A-B24n, and may be configured to assert a Nack response if the processor B30 asserts a Nack response in a given iteration.
[0169] In one embodiment, using a Request / Ack / Nack structure can provide a quick indication of whether an interrupt has been accepted by the receiver of the request (e.g., the cluster interrupt controller B24A-B24n or the processor B30, depending on the interface). The indication may be quicker than a timeout, for example, in one embodiment. Additionally, in one embodiment, the hierarchical structure of the cluster interrupt controllers B24A-B24n and the interrupt controller B20 may be more scalable to a larger number of processors (e.g., multiple processor clusters) in the system B10.
[0170] The iteration through the cluster interrupt controllers B24A-B24n may include attempting to deliver the interrupt through at least a subset of the cluster interrupt controllers B24A-B24n, up to all of the cluster interrupt controllers B24A-B24n. The iteration may proceed in any desired manner. For example, in one embodiment, the interrupt controller B20 may be configured to serially assert interrupt requests to each cluster interrupt controller B24A-B24n, terminated by an Ack response from one of the cluster interrupt controllers B24A-B24n (and, in one embodiment, the absence of additional pending interrupts) or by a Nack response from all of the cluster interrupt controllers B24A-B24n. That is, the interrupt controller may select one of the cluster interrupt controllers B24A-B24n and assert an interrupt request to the selected cluster interrupt controller B24A-B24n (e.g., by asserting a soft request or a hard request, depending on which iteration is being performed). The selected cluster interrupt controller B24A-B24n may respond with an Ack response, thereby terminating the iteration. On the other hand, if the selected cluster interrupt controller B24A-B24n asserts a Nack response, the interrupt controller may be configured to select another cluster interrupt controller B24A-B24n and may assert a soft or hard request to the selected cluster interrupt controller B24A-B24n. The selection and assertion may continue until an Ack response is received or until each of the cluster interrupt controllers B24A-B24n has been selected and asserted a Nack response. Other embodiments may implement the iteration over the cluster interrupt controllers B24A-B24n in other ways.For example, the interrupt controller B20 may be configured to simultaneously assert interrupt requests to a subset of two or more cluster interrupt controllers B24A-B24n, and continue with the other subsets if each cluster interrupt controller B24A-B24n in the subset provides a NACK response to the interrupt request. Such an implementation may cause spurious interrupts if more than one cluster interrupt controller B24A-B24n in the subset provides an Ack response, and therefore code executed in response to an interrupt may be designed to handle the occurrence of spurious interrupts.
[0171] The initial iteration may be a soft iteration, as described above. In a soft iteration, a given cluster interrupt controller B24A-B24n may attempt to deliver an interrupt to a subset of the multiple processors B30 associated with the given cluster interrupt controller B24A-B24n. The subset may be the processors B30 that are powered on, and a given cluster interrupt controller B24A-B24n may not attempt to deliver an interrupt to a processor B30 that is powered off (or is sleeping). That is, powered-off processors are not included in the subset to which the cluster interrupt controller B24A-B24n attempts to deliver an interrupt. Thus, powered-off processors B30 may remain powered off in the soft iteration.
[0172] Based on the Nack responses from each cluster interrupt controller B24A-B24n during the soft iterations, the interrupt controller B20 can perform hard iterations. In hard iterations, powered-off processors B30 in a given processor cluster can be powered on by individual cluster interrupt controllers B24A-B24n, which can attempt to deliver interrupts to each processor B30 in the processor cluster. More specifically, in one embodiment, when a processor B30 is powered on to perform hard iterations, that processor B30 can be quickly available for interrupts and can provide frequent Ack responses.
[0173] If a hard iteration ends with one or more interrupts still pending, or if a timeout occurs before completing the soft and hard iterations, the interrupt controller may initiate a forced iteration by asserting a force signal. In one embodiment, the forced iteration may be performed in parallel with the cluster interrupt controllers B24A-B24n, and Nack responses may not be allowed. In one embodiment, the forced iteration may remain in progress until there are no more pending interrupts.
[0174] A given cluster interrupt controller B24A-B24n may attempt to deliver an interrupt in any desired manner. For example, a given cluster interrupt controller B24A-B24n may serially assert an interrupt request to each processor B30 in a processor cluster, which is terminated by an Ack response from one of the respective processors B30 or by a Nack response from each of the respective processors B30 to which the given cluster interrupt controller B24A-B24n attempts to deliver the interrupt. That is, a given cluster interrupt controller B24A-B24n may select one of the respective processors B30 and assert an interrupt request to the selected processor B30 (e.g., by asserting a request to the selected processor B30). The selected processor B30 may respond with an Ack response, thereby terminating the attempt. On the other hand, if the selected processor B30 asserts a Nack response, a given cluster interrupt controller B24A-B24n can be configured to select another processor B30 and assert an interrupt request to the selected processor B30. The selection and assertion can continue until an Ack response is received or until each of the processors B30 has been selected and asserted a Nack response (excluding powered-off processors in soft iteration). Other embodiments can assert interrupt requests to multiple processors B30 simultaneously or in parallel, potentially resulting in spurious interrupts as described above. A given cluster interrupt controller B24A-B24n can respond to the interrupt controller B20 with an Ack response based on receiving an Ack response from one of the processors B30, or can respond to the interrupt controller B20 with a Nack response if each of the processors B30 has responded with a Nack response.
[0175] In one embodiment, the order in which the interrupt controller B20 asserts interrupt requests to the cluster interrupt controllers B24A-B24n may be programmable. More specifically, in one embodiment, the order may vary based on the source of the interrupt (e.g., interrupts from one interrupt source B32 may result in one order, while interrupts from another interrupt source B32 may result in a different order). For example, in one embodiment, the processors B30 in one cluster may be different from the processors B30 in another cluster. One processor cluster may have processors that are optimized for performance but may be higher power, while another processor cluster may have processors optimized for power efficiency. Interrupts from sources requiring relatively little processing may favor clusters with power-efficient processors, while interrupts from sources requiring considerable processing may favor clusters with higher performance processors.
[0176] The interrupt source B32 may be any hardware circuit configured to assert an interrupt to cause the processor B30 to execute an interrupt service routine. For example, in one embodiment, various peripheral components (peripherals) may be interrupt sources. Examples of various peripherals are described below with respect to FIG. 16. Interrupts are asynchronous to the code being executed by the processor B30 when it receives the interrupt. Generally, the processor B30 may be configured to service the interrupt by stopping execution of the current code, saving the processor context to enable resumption of execution after servicing the interrupt, and branching to a predetermined address to begin execution of the interrupt code. The code at the predetermined address may read state from the interrupt controller to determine which interrupt source B32 asserted the interrupt and the corresponding interrupt service routine to execute based on the interrupt. The code may queue the interrupt service routine for execution (which may be scheduled by the operating system) and provide the data expected by the interrupt service routine. The code may then return execution to the previously executing code (e.g., the processor context may be reloaded and execution may resume at the instruction where execution was halted).
[0177] Interrupts may be transmitted from the interrupt sources B32 to the interrupt controller B20 in any desired manner. For example, dedicated interrupt wires may be provided between the interrupt sources and the interrupt controller B20. A given interrupt source B32 may assert a signal on that dedicated wire to transmit the interrupt to the interrupt controller B20. Alternatively, message-signaled interrupts may be used, in which a message is transmitted over the interconnect used for other communications within the system B10. The message may be in the form of, for example, a write to a specified address. The write data may be a message identifying the interrupt. A combination of dedicated wires from some interrupt sources B32 and message-signaled interrupts from other interrupt sources B32 may be used.
[0178] The interrupt controller B20 can receive interrupts and record them as pending interrupts within the interrupt controller B20. Interrupts from various interrupt sources B32 can be prioritized by the interrupt controller B20 according to various programmable priorities configured by the operating system or other control code.
[0179] Referring now to FIG. 16, a block diagram of one embodiment of a system B10 implemented as a system-on-chip (SOC) B10 coupled to a memory B12 is shown. In one embodiment, the SOC B10 may be an instance of the SOC B10 shown in FIG. 1. As implied by the name, the components of the SOC B10 may be integrated on a single semiconductor substrate as an integrated circuit "chip." In some embodiments, the components may be implemented on two or more separate chips within the system. However, the SOC B10 will be used herein as an example. In the illustrated embodiment, the components of the SOC B10 include multiple processor clusters B14A-B14n, an interrupt controller B20, one or more peripheral components B18 (more simply, "peripherals"), a memory controller B22, and a communications fabric B27. Components B14A-B14n, B18, B20, and B22 may all be coupled to the communications fabric B27. The memory controller B22 may be coupled to the memory B12 during use. In some embodiments, there may be more than one memory controller coupled to a corresponding memory. The memory address space may be mapped across the memory controllers in any desired manner. In the illustrated embodiment, the processor clusters B14A-B14n may include a respective plurality of processors (P) B30 and a respective cluster interrupt controller (IC) B24A-B24n, as shown in FIG. 16. The processors B30 may form the central processing unit (CPU(s)) of the SOC B10. In one embodiment, one or more of the processor clusters B14A-B14n may not be used as a CPU.
[0180] The peripherals B18, in one embodiment, may include peripherals that are examples of interrupt source BB32. Thus, one or more peripherals B18 may have dedicated wires to the interrupt controller B20 to transmit interrupts to the interrupt controller B20. Other peripherals B18 may use message signal interrupts transmitted over the communications fabric B27. In some embodiments, one or more off-SOC devices (not shown in FIG. 16) may also be interrupt sources. The dotted line from the interrupt controller B20 to off-chip indicates a possible off-SOC interrupt source.
[0181] The hard / soft / forced Ack / Nack interfaces between the cluster ICs B24A-B24n shown in Figure 15 are shown in Figure 16 via arrows between the cluster ICs B24A-B24n and the interrupt controller B20. Similarly, the Req Ack / Nack interfaces between the processor B30 and the cluster ICs B24A-B24n in Figure 1 are shown by arrows between the cluster ICs B24A-B24n and the processor B30 in each cluster B14A-B14n.
[0182] As mentioned above, processor clusters B14A-B14n may include one or more processors B30 that may function as the CPU of the SOC B10. The CPU of the system includes a processor or processors that execute the system's main control software, such as an operating system. Generally, during use, the software executed by the CPU may control other components of the system to achieve the desired functionality of the system. The processors may also execute other software, such as application programs. The application programs may provide user functionality and may rely on the operating system for low-level device control, scheduling, memory management, etc. Thus, the processors may also be referred to as application processors.
[0183] Generally, a processor may include any circuitry and / or microcode configured to execute instructions defined in an instruction set architecture implemented by the processor. A processor may include a processor core implemented on an integrated circuit along with other components as a system-on-chip (SOC B10) or other level of integration. A processor may also include a separate microprocessor, a processor core and / or microprocessor integrated in a multi-chip module implementation, a processor implemented as multiple integrated circuits, etc.
[0184] The memory controller B22 may generally include circuitry for receiving memory operations from other components of the SOC B10 and for accessing the memory B12 to complete the memory operations. The memory controller B22 may be configured to access any type of memory B12. For example, the memory B12 may be static random access memory (SRAM), dynamic RAM (DDR, DDR2, DDR3, DDR4, etc.), double data rate (DRAM), such as synchronous DRAM (SDRAM), including DRAM. Low-power / mobile versions of DDR DRAM (e.g., LPDDR, mDDR, etc.) may be supported. The memory controller B22 may include queues for memory operations to order (and possibly reorder) the operations and present them to the memory B12. The memory controller B22 may further include data buffers for storing write data waiting to be written to memory and read data waiting to be returned to the source of the memory operation. In some embodiments, the memory controller B22 may include a memory cache for storing recently accessed memory data. In an SOC implementation, for example, a memory cache can reduce power consumption in the SOC by avoiding re-accessing data from memory B12 if it is expected to be accessed again soon. In some cases, a memory cache may also be referred to as a system cache, as opposed to a private cache that serves only a specific component, such as an L2 cache or a cache within a processor. Furthermore, in some embodiments, the system cache need not be located within the memory controller B22.
[0185] Peripherals B18 may be any set of additional hardware functions included in SOC B10. For example, peripherals B18 may include video peripherals such as an image signal processor configured to process image capture data from a camera or other image sensor, a GPU, a video encoder / decoder, a scaler, a rotator, a blender, and a display controller. Peripherals may include audio peripherals such as a microphone, a speaker, an interface to a microphone and a speaker, an audio processor, a digital signal processor, and a mixer. Peripherals may include interface controllers for various interfaces external to SOC B10, including interfaces such as Universal Serial Bus (USB), Peripheral Component Interconnect (PCI) including PCI Express (PCIe), serial and parallel ports, and the like. Interconnects to external devices are indicated by dashed arrows in FIG. 15 extending outside of SOC B10. Peripherals may also include network peripherals such as a Media Access Controller (MAC). Any set of hardware may be included.
[0186] The communications fabric B27 can be any communications interconnect and protocol for communicating between components of the SOC B10. The communications fabric B27 can be bus-based, including shared bus configurations, crossbar configurations, and hierarchical buses with bridges. The communications fabric B27 can also be packet-based or hierarchical with bridges, crossbar, point-to-point, or other interconnects.
[0187] Note that the number of components of SOC B10 (and the number of subcomponents for those shown in FIG. 16, such as processor B30 in each processor cluster B14A-B14n) may vary from embodiment to embodiment. Additionally, the number of processors B30 in one processor cluster B14A-B14n may differ from the number of processors B30 in another processor cluster B14A-B14n. There may be more or fewer of each component / subcomponent than shown in FIG. 16.
[0188] 17 is a block diagram illustrating an embodiment of an interrupt controller that may be implemented by the interrupt controller B20 in one embodiment. In the illustrated embodiment, the states include an idle state B40, a soft state B42, a hard state B44, a force state B46, and a wait drain state B48.
[0189] In the idle state B40, there may be no pending interrupts. Generally, the state machine can return to the idle state B40 from any of the other states, such as those shown in FIG. 17, any time there are no pending interrupts. When at least one interrupt is received, the interrupt controller B20 can transition to the soft state B42. The interrupt controller B20 can also initialize a timeout counter to begin counting a timeout interval that can cause the state machine to transition to the forced state B46. The timeout counter can be initialized to zero, incremented, and compared to a timeout value to detect a timeout. Alternatively, the timeout counter can be initialized to a timeout value and decremented until it reaches 0. The increment / decrement can occur every clock cycle of the clock for the interrupt controller B20 or in response to another clock (e.g., a fixed frequency clock from a piezoelectric oscillator, etc.).
[0190] In the soft state B42, the interrupt controller B20 may be configured to begin soft iterations attempting to deliver the interrupt. If one of the cluster interrupt controllers B24A-B24n transmits an Ack response during the soft iterations and there is at least one pending interrupt, the interrupt controller B20 may transition to a wait drain state B48. The wait drain state B48 may be provided because a given processor may service an interrupt but may actually take multiple interrupts from the interrupt controller and queue them for their respective interrupt service routines. In various embodiments, the processor may continue to drain interrupts until all interrupts have been read from the interrupt controller B20, or may read up to a certain maximum number of interrupts and return to processing, or may read interrupts until a timer expires. If the timer times out and there are still pending interrupts, the interrupt controller B20 may be configured to transition to a force state B46 and begin force iterations to deliver the interrupt. If the processor stops draining interrupts and there is at least one pending interrupt or a new interrupt is pending, the interrupt controller B20 may be configured to return to soft state B42 and continue soft iteration.
[0191] When the soft iterations are complete with a Nack response from each cluster interrupt controller B24A-B24n (and at least one interrupt remains pending), the interrupt controller B20 may be configured to transition to a hard state B44 and may begin a hard iteration. If the cluster interrupt controllers B24A-B24n provide an Ack response during the hard iterations and there is at least one pending interrupt, the interrupt controller B20 may transition to a wait drain state B48 similar to the description above. When the hard iterations are complete with a Nack response from each cluster interrupt controller B24A-B24n and there is at least one pending interrupt, the interrupt controller B20 may be configured to transition to a force state B46 and may begin a force iteration. The interrupt controller B20 may remain in the force state B46 until there are no more pending interrupts.
[0192] FIG. 18 is a flowchart illustrating the operation of one embodiment of the interrupt controller B20 when performing soft or hard iterations (e.g., when in state B42 or B44 of FIG. 17). For ease of understanding, the blocks are shown in a particular order, but other orders may be used. The blocks may be implemented in parallel in combinational logic within the interrupt controller B20. The blocks, combinations of the blocks, and / or the flowchart as a whole may be pipelined over multiple clock cycles. The interrupt controller B20 may be configured to implement the operations shown in FIG. 18.
[0193] The interrupt controller may be configured to select a cluster interrupt controller B24A-B24n (block B50). Any mechanism for selecting a cluster interrupt controller B24A-B24n from multiple interrupt controllers B24A-B24n may be used. For example, a programmable order of the clusters of the interrupt controllers B24A-B24n may indicate which cluster of the interrupt controllers B24A-B24n is selected. In one embodiment, the order may be based on the interrupt source of a given interrupt (e.g., there may be multiple orders available, and a particular order may be selected based on the interrupt source). Such an implementation may allow different interrupt sources to favor a given type of processor (e.g., performance-optimized or efficiency-optimized) by first attempting to deliver the interrupt to a processor cluster of a desired type before moving to a processor cluster of a different type. In another embodiment, a least-recently delivered algorithm may be used to distribute interrupts across different processor clusters by selecting the most-recently delivered cluster interrupt controller B24A-B24n (e.g., the cluster interrupt controller B24A-B24n that generated the least-recently acknowledged response to the interrupt). In another embodiment, a most recently delivered algorithm may be used to select a cluster interrupt controller (e.g., the cluster interrupt controller B24A-B24n that most recently generated an Ack response for the interrupt) to take advantage of the possibility that the interrupt code or state is still cached in the processor cluster. Any mechanism or combination of mechanisms may be used.
[0194] The interrupt controller B20 may be configured to transmit an interrupt request (hard or soft, depending on the current iteration) to the selected cluster interrupt controller B24A-B24n (block B52). For example, the interrupt controller B20 may assert a hard or soft interrupt request signal to the selected cluster interrupt controller B24A-B24n. If the selected cluster interrupt controller B24A-B24n provides an Ack response to the interrupt request (decision block B54, "Yes" branch), the interrupt controller B20 may be configured to transition to a wait drain state B48 to allow a processor B30 in the processor cluster B14A-B14n associated with the selected cluster interrupt controller B24A-B24n to service one or more pending interrupts (block B56). If the selected cluster interrupt controller provides a Nack response (decision block B58, "Yes" branch) and there is at least one cluster interrupt controller B24A-B24n that has not been selected in the current iteration (decision block B60, "Yes" branch), the interrupt controller B20 may be configured to select the next cluster interrupt controller B24A-B24n according to the implemented selection mechanism (block B62) and return to block B52 to assert the interrupt request to the selected cluster interrupt controller B24A-B24n. Thus, the interrupt controller B20, in this embodiment, may be configured to attempt to distribute the interrupt controller serially to the multiple cluster interrupt controllers B24A-B24n during iterations over the multiple cluster interrupt controllers B24A-B24n.If the selected cluster interrupt controller B24A-B24n provides a Nack response (decision block B58, "Yes" branch) and there are no more cluster interrupt controllers B24A-B24n left to select (e.g., all cluster interrupt controllers B24A-B24n have been selected), the cluster interrupt controller B20 may be configured to transition to the next state in the state machine (e.g., to hard state B44 if the current iteration is a soft iteration, or to forced state B46 if the current iteration is a hard iteration) (block B64). If a response to the interrupt request has not yet been received (decision blocks B54 and B58, "No" branch), the interrupt controller B20 may be configured to continue waiting for a response.
[0195] As mentioned above, there may be a timeout mechanism that may be initialized when the interrupt delivery process begins. If a timeout occurs during any state, in one embodiment, the interrupt controller B20 may be configured to transition to the force state B46. Alternatively, timer expiration may only be considered in the wait drain state B48.
[0196] FIG. 19 is a flow diagram illustrating the operation of one embodiment of a cluster interrupt controller B24A-B24n based on an interrupt request from the interrupt controller B20. While the blocks are shown in a particular order for ease of understanding, other orders may be used. The blocks may be implemented in parallel within combinational logic within the cluster interrupt controller B24A-B24n. The blocks, combinations of blocks, and / or the overall flowchart may be pipelined over multiple clock cycles. The cluster interrupt controllers B24A-B24n may be configured to implement the operations shown in FIG. 19.
[0197] If the interrupt request is a hard request or a forced request (decision block B70, "yes" branch), the cluster interrupt controllers B24A-B24n may be configured to power up any powered-down (e.g., sleeping) processors B30 (block B72). If the interrupt request is a forced interrupt request (decision block B74, "yes" branch), the cluster interrupt controllers B24A-B24n may be configured to interrupt all processors B30 in parallel (block B76). Ack / Nack may not apply in the forced case, so the cluster interrupt controllers B24A-B24n may continue to assert the interrupt request until at least one processor takes the interrupt. Alternatively, the cluster interrupt controllers B24A-B24n may be configured to receive an Ack response from the processor indicating that they will take the interrupt, terminate the forced interrupt, and transmit the Ack response to the interrupt controller B20.
[0198] If the interrupt request is a hard request (decision block B74, “No” branch) or a soft request (decision block B70, “No” branch), the cluster interrupt controller may be configured to select a powered-on processor B30 (block B78). Any selection mechanism similar to the mechanism described above for selecting a cluster interrupt controller B24A-B24n by the interrupt controller B20 may be used (e.g., programmable order, least recently interrupted, most recently interrupted, etc.). In one embodiment, the order may be based on processor IDs assigned to the processors in the cluster. The cluster interrupt controller B24A-B24n may be configured to assert an interrupt request to the selected processor B30 and transmit the request to the processor B30 (block B80). If the selected processor B30 provides an Ack response (decision block B82, “Yes” branch), the cluster interrupt controller B24A-B24n may be configured to provide an Ack response to the interrupt controller B20 (block B84) and terminate the attempt to deliver the interrupt within the processor cluster. If the selected processor B30 provides a Nack response (decision block B86, "Yes" branch) and there is at least one powered-on processor B30 that has not yet been selected (decision block B88, "Yes" branch), the cluster interrupt controller B24A-B24n may be configured to select the next powered-on processor (e.g., according to the selection mechanism described above) (block B90) and assert an interrupt request to the selected processor B30 (block B80). Thus, the cluster interrupt controller B24A-B24n may attempt to serially deliver the interrupt to the processors B30 in the processor cluster. If there are no more powered-on processors to select (decision block B88, "No" branch), the cluster interrupt controller B24A-B24n may be configured to provide a Nack response to the interrupt controller B20 (block B92).If the selected processor B30 has not yet provided a response (decision blocks B82 and B86, "No" branch), the cluster interrupt controllers B24A-B24n may be configured to wait for a response.
[0199] In one embodiment, in a hard iteration, if processor B30 is powered on from a powered-off state, processor B30 may be available for interrupts quickly because it has not yet been assigned a task by the operating system or other control software. The operating system may be configured to unmask interrupts in processor B30 powered on from a powered-off state as soon as possible after initializing the processor. The cluster interrupt controllers B24A-B24n may select the most recently powered-on processor first in the selection order to improve the likelihood that the processor will provide an Ack response to an interrupt.
[0200] 20 is a more detailed block diagram of one embodiment of a processor B30. In the illustrated embodiment, the processor B30 includes a fetch and decode unit B100 (including an instruction cache or I-cache, B102), a map dispatch rename (MDR) unit B106 (including a processor interrupt acknowledge (Int Ack) control circuit B126 and a reorder buffer B108), one or more reservation stations B110, one or more execution units B112, a register file B114, a data cache (D-cache) B104, a load / store unit (LSU) B118, a reservation station for the load / store units (RS) B116, and a core interface unit (CIF) B122. The fetch and decode unit B100 is coupled to the MDR unit B106, which is coupled to the reservation station B110, the reservation station B116, and the LSU B118. The reservation station B110 is coupled to the execution units B28. The register file B114 is coupled to the execution units B112 and the LSUs B118. The LSUs B118 are also coupled to the D-cache B104, which is coupled to the CIF B122 and the register file B114. The LSUs B118 include a store queue B120 (STQ B120) and a load queue (LDQ B124). The CIF B122 is coupled to the processor Int Ack control circuit B126 to transmit asserted interrupt requests (Int Req) to the processor B30 and to transmit Ack / Nack responses from the processor Int Ack control circuit B126 to the interrupt request source (e.g., cluster interrupt controllers B24A-B24n).
[0201] The processor Int Ack control circuit B126 may be configured to determine whether the processor B30 may accept an interrupt request transmitted to the processor B30 and may provide Ack and Nack indications to the CIF B122 based on the determination. If the processor B30 provides an Ack response, the processor B30 has committed to servicing the interrupt (and starting execution of interrupt code to identify the interrupt and the interrupt source) within a specified period of time. That is, the processor Int Ack control circuit B126 may be configured to generate an acknowledgement (Ack) response to a received interrupt request based on a determination that the reorder buffer B108 will retire an instruction operation to an interruptible point and the LSU B118 will complete a load / store operation to an interruptible point within a specified period of time. If the determination is that at least one of the reorder buffer B108 and the LSU B118 will not (or may not) reach an interruptible point within a specified period of time, the processor Int Ack control circuit B126 may be configured to generate a non-acknowledgement (Nack) response to the interrupt request. For example, the specified period of time may be approximately 5 microseconds in one embodiment, but may be longer or shorter in other embodiments.
[0202] In one embodiment, the processor Int Ack control circuit B126 may be configured to examine the contents of the reorder buffer 108 to make an initial Ack / Nack decision. That is, there may be one or more cases in which the processor Int Ack control circuit B126 may be able to determine that a Nack response will be generated based on conditions within the MDR unit B106. For example, if the reorder buffer B108 contains one or more instruction operations that have not yet executed and have a potential execution latency greater than a certain threshold, the processor Int Ack control circuit B126 may be configured to determine that a Nack response will be generated. The execution latency is referred to as “potential” because some instruction operations may have variable execution latencies that may be data-dependent, memory-latency-dependent, etc. Thus, the potential execution latency may be the longest execution latency that could occur, even if it does not always occur. In other cases, the potential execution latency may be the longest execution latency that occurs beyond a certain probability, etc. Examples of such instructions may include certain cryptographic acceleration instructions, certain types of floating-point instructions, vector instructions, etc. If an instruction is not interruptible, it may be considered to have a potentially long latency, i.e., the non-interruptible instruction must complete execution after it begins execution.
[0203] Another condition that may be considered when generating an Ack / Nack response is the state of interrupt masking in the processor 30. When an interrupt is masked, the processor B30 is prevented from servicing the interrupt. A Nack response may be generated if the processor Int Ack control circuit B126 detects that an interrupt is masked in the processor (which, in one embodiment, may be state-maintained in the MDR unit B106). More specifically, in one embodiment, the interrupt mask may have an architected current state corresponding to the most recently retired instruction, and one or more speculative updates to the interrupt mask may be queued as well. In one embodiment, a Nack response may be generated if the architected current state is that the interrupt is masked. In another embodiment, a Nack response may be generated if the architected current state is that the interrupt is masked, or if any of the speculative states indicate that the interrupt is masked.
[0204] Other cases may similarly be considered Nack response cases in the processor Int Ack control circuit B126. For example, a Nack response may be generated if there is a pending redirection related to exception handling in the reorder buffer (e.g., no microarchitectural redirection such as a branch misprediction). Certain debug modes (e.g., single-step mode) and high-priority internal interrupts may be considered Nack response cases.
[0205] If the processor Int Ack control circuit B126 does not detect a Nack response based on examining the processor state in the reorder buffer B108 and the MDR unit B106, the processor Int Ack control circuit B126 can interface with the LSU B118 to determine whether there are any long latency load / store instructions issued (e.g., to the CIF B122 or external to the processor B30) and coupled to the reorder buffer and load / store unit that have not yet completed. For example, loads and stores to device space (e.g., loads and stores that are mapped to peripherals instead of memory) can have a potentially long latency. If the LSU B118 responds that there is a long latency load / store instruction (e.g., potentially greater than a threshold that may be different from or the same as the above-mentioned threshold used internally in the MDR unit B106), the processor Int Ack control circuit B126 can determine that the response is a Nack. Other potentially long latency operations can be, for example, synchronization barrier operations.
[0206] In one embodiment, if the determination is not a Nack response for the above case, the LSU B118 may provide to the reorder buffer B108 a pointer identifying the oldest load / store instruction that the LSU B118 has committed to completion (e.g., launched from the LDQ B124 or STQ B120, or otherwise non-speculative in the LSU B118). The pointer may be referred to as a "true load / store (LS) non-speculative (NS) pointer." The MDR B106 / reorder buffer B108 may attempt to intercept at the LS NS pointer, and if that is not possible within a specified period of time, the processor Int Ack control circuit B126 may determine that a Nack response is generated. Otherwise, an Ack response may be generated.
[0207] The fetch and decode unit B100 may be configured to fetch instructions for execution by the processor B30 and decode the instructions into ops for execution. More specifically, the fetch and decode unit B100 may be configured to cache previously fetched instructions from memory (via the CIF B122) in the I-cache B102 and to fetch speculative paths of instructions for the processor B30. The fetch and decode unit B100 may implement various prediction structures to predict the fetch path. For example, a next fetch predictor may be used to predict a fetch address based on previously executed instructions. Various types of branch predictors may be used to verify the next fetch prediction or to predict the next fetch address if a next fetch predictor is not used. The fetch and decode unit 100 may be configured to decode instructions into instruction operations. In some embodiments, a given instruction may be decoded into one or more instruction operations depending on the complexity of the instruction. In some embodiments, particularly complex instructions may be microcoded. In such an embodiment, the microcode routine for the instruction may be coded in an instruction operation. In other embodiments, each instruction in the instruction set architecture implemented by the processor B30 may be decoded into a single instruction operation, and thus an instruction operation may be essentially synonymous with an instruction (although the format may be modified by the decoder). The term "instruction operation" may be more succinctly referred to herein as an "op."
[0208] The MDR unit B106 may be configured to map ops to speculative resources (e.g., physical registers) to enable out-of-order and / or speculative execution and may dispatch ops to the reservation stations B110 and B116. Ops may be mapped from architectural registers used by the corresponding instruction to physical registers in the register file B114. That is, the register file B114 may implement a set of physical registers that may be larger in number than the architectural registers specified by the instruction set architecture implemented by the processor B30. The MDR unit B106 may manage the mapping of architectural registers to physical registers. In one embodiment, there may be separate physical registers for different operand types (e.g., integer, medium, floating point, etc.). In other embodiments, physical registers may be shared across operand types. The MDR unit B106 may also be responsible for tracking speculative execution and retiring ops or flushing misspeculated ops. The reorder buffer B108 may be used to track the program order of ops and manage retirement / flushing. That is, the reorder buffer B108 may be configured to track multiple instruction operations corresponding to instructions that are fetched by the processor and not retired by the processor.
[0209] An op can be scheduled for execution when its source operands are ready. In the illustrated embodiment, decentralized scheduling is used for each of the execution units B28 and LSUs B118, e.g., at reservation stations B116 and B110. Other embodiments may implement a centralized scheduler if desired.
[0210] The LSU B118 may be configured to execute load / store memory ops. Generally, a memory operation (memory op) may be an instruction operation that specifies access to memory (although the memory access may be completed in a cache, such as the D-cache B104). A load memory operation may specify a transfer of data from a memory location to a register, and a store memory operation may specify a transfer of data from a register to a memory location. A load memory operation may be referred to as a load memory op, a load op, or a load, and a store memory operation may be referred to as a store memory op, a store op, or a store. In one embodiment, a store op may be executed as a store address op and a store data op. A store address op may be defined to generate the address of the store, probe the cache for an initial hit / miss determination, and update the store queue with the address and cache information. Thus, a store address op may have an address operand as a source operand. A store data op may be defined to deliver store data to the store queue. Thus, a store data op may not have an address operand as a source operand, but may have a store data operand as a source operand. In many cases, the address operand of a store may be available before the store data operand, and therefore the address may be determined and made available earlier than the store data. In some embodiments, for example, if a store data operand is provided before one or more of the store address operands, it may be possible for a store data op to be executed before the corresponding store address op. A store op may be executed as a store address and store data op in some embodiments, although other embodiments may not implement the store address / store data split. While the remainder of this disclosure often uses a store address op (and a store data op) as examples, implementations that do not use the store address / store data optimization are also contemplated. The address generated through execution of a store address op may be referred to as the address corresponding to the store op.
[0211] Load / store ops may be received at a reservation station B116, which may be configured to monitor the operation's source operands to determine when they are available and then issue the operation to the load or store pipeline, respectively. Some source operands may be available when the operation is received at the reservation station B116, which may be indicated in data received by the reservation station B116 from the MDR unit B106 for the corresponding operation. Other operands may become available through execution of operations by other execution units B112, or even through execution of a previous load op. The operands may be gathered by the reservation station B116 or may be read from the register file B114 as issued from the reservation station B116, as shown in FIG. 20.
[0212] In one embodiment, the reservation station B116 may be configured to issue load / store ops out of order (from the original order in the code sequence being executed by the processor B30, referred to as “program order”) as operands become available. To ensure there is space in the LDQ B124 or STQ B120 for older operations that are bypassed by newer operations in the reservation station B116, the MDR unit B106 may include circuitry that pre-allocates entries in the LDQ B124 or STQ B120 for operations transmitted to the load / store unit B118. If there are no available LDQ entries for a load being processed in the MDR unit B106, the MDR unit B106 may stall dispatch of the load op and subsequent ops in program order until one or more LDQ entries become available. Similarly, if there are no available STQ entries for a store, the MDR unit B106 may stall op dispatch until one or more STQ entries become available. In other embodiments, the reservation station B116 may issue operations in program order, and the LRQ B46 / STQ B120 allocation may occur upon issuance from the reservation station B116.
[0213] The LDQ B124 can track loads from their initial execution by the LSU B118 to their retirement. The LDQ B124 can be responsible for ensuring that memory ordering rules are not violated (between loads executed out of order, and between loads and stores). If a memory ordering violation is detected, the LDQ B124 can signal a redirection for the corresponding load. The redirection can cause the processor B30 to flush the load and subsequent ops in program order and refetch the corresponding instructions. The speculative state for the load and subsequent ops may be discarded, and the ops may be refetched by the fetch and decode unit B100 and reprocessed to be executed again.
[0214] When a load / store address op is issued by the reservation station B116, the LSU B118 can be configured to generate the address accessed by the load / store and to translate the address from an effective or virtual address created from the address operand of the load / store address op to a physical address actually used to address memory. The LSU B118 can be configured to generate accesses to the D-cache B104. For a load operation that hits in the D-cache B104, data can be speculatively forwarded from the D-cache B104 to the load operation's destination operand (e.g., a register in the register file B114) unless the address hits an earlier operation in the STQ B120 (i.e., an older store in program order) or the load is replayed. Data can also be speculatively scheduled and forwarded to a dependent op residing in the execution unit B28. In this case, the execution unit B28 may bypass the forwarded data in place of data output from the register file B114. If the store data is available for forwarding at the time of the STQ hit, the data output by the STQ B120 may be forwarded instead of the cache data. Cache misses and STQ hits where the data cannot be forwarded may be reasons for replay; load data may not be forwarded in those cases. Cache hit / miss status from the D-cache B104 may be logged in the STQ B120 or LDQ B124 for later processing.
[0215] The LSU B 118 may implement multiple load pipelines. For example, in one embodiment, three load pipelines (“pipes”) may be implemented, although other embodiments may implement more or fewer pipelines. Each pipeline may execute a different load independently and in parallel with other loads. That is, the RS B 116 may issue any number of loads up to the number of load pipes in the same clock cycle. The LSU B 118 may also implement one or more store pipes, and in particular, may implement multiple store pipes. However, the number of store pipes need not equal the number of load pipes. In one embodiment, for example, two store pipes may be used. The reservation station B 116 may issue store address ops and store data ops independently and in parallel to the store pipes. The store pipes may be coupled to the STQ B 120, which may be configured to hold executed but uncommitted store operations.
[0216] The CIF B122 may be responsible for communicating with the rest of the system, including the processor B30, on behalf of the processor B30. For example, the CIF B122 may be configured to request data for D-cache B104 misses and I-cache B102 misses. Once the data is returned, the CIF B122 may signal a cache fill to the corresponding cache. In the case of a D-cache fill, the CIF B122 may also notify the LSU B118. The LDQ B124 may attempt to schedule recycled loads awaiting a cache fill so that the recycled load can forward the fill data when the fill data is provided to the D-cache B104 (referred to as a fill forwarding operation). If the recycled load is not successfully reclaimed during the fill, the recycled load may subsequently be scheduled and reclaimed through the D-cache B104 as a cache hit. The CIF B122 may also write back modified cache lines evicted by the D-cache B104, merge store data for non-cacheable stores, etc.
[0217] Execution unit B 112 may, in various embodiments, include any type of execution unit. For example, execution unit B 112 may include integer, floating-point, and / or vector execution units. The integer execution unit may be configured to execute integer ops. Generally, an integer op is an op that performs a defined operation (e.g., arithmetic, logical, shift / rotate, etc.) on integer operands. Integers may be numeric values whose values correspond to mathematical integers. The integer execution unit may include branch handling hardware for processing branch ops, or there may be a separate branch execution unit.
[0218] A floating-point execution unit may be configured to execute floating-point ops. Generally, a floating-point op may be an op defined to operate on floating-point operands. A floating-point operand is an operand expressed as a base raised to the power of an exponent multiplied by a mantissa (or mantissa). The exponent, the sign of the operand, and the mantissa / mantissa may be explicitly represented in the operand, or the base may be implicit (e.g., base 2 in one embodiment).
[0219] A vector execution unit may be configured to execute vector ops. Vector ops may be used, for example, to process media data (e.g., image data such as pixels, audio data, etc.). Media processing may be characterized by performing the same operation on a substantial amount of data, each of which is a relatively small value (e.g., 8 or 16 bits, compared to 32 to 64 bits for integers). Thus, a vector op includes single instruction multiple data (SIMD) or vector operations on operands representing multiple media data.
[0220] Thus, each execution unit B 112 may comprise hardware configured to perform a defined operation on an op that the particular execution unit is defined to process. Execution units may generally be independent of one another in the sense that each execution unit may be configured to operate on an op issued to it without relying on other execution units. Viewed another way, each execution unit may be an independent pipe for executing ops. Different execution units may have different execution latencies (e.g., different pipe lengths). Furthermore, different execution units may have different latencies relative to the pipeline stage where bypassing occurs, and thus the clock cycles over which speculative scheduling of a dependent op based on a load op occurs may vary based on the type of op and the execution unit B 28 executing the op.
[0221] It should be noted that any number and type of execution unit B 112 may be included in various embodiments, including embodiments with one execution unit and embodiments with multiple execution units.
[0222] A cache line may be the unit of allocation / deallocation in a cache. That is, data in a cache line may be allocated / deallocated in the cache as a unit. Cache lines may vary in size (e.g., 32-byte, 64-byte, 128-byte, or larger or smaller cache lines). Different caches may have different cache line sizes. The I-cache B102 and D-cache B104 may each be caches having any desired capacity, cache line size, and configuration. In various embodiments, there may be more additional levels of cache between the D-cache B104 / I-cache B102 and main memory.
[0223] At various times, load / store operations are referred to as being newer or older than other load / store operations. A first operation may be younger than a second operation if it follows the second operation in program order. Similarly, a first operation may be older than a second operation if it precedes the second operation in program order.
[0224] 21 is a block diagram of one embodiment of reorder buffer B 108. In the illustrated embodiment, reorder buffer 108 includes multiple entries. Each entry may correspond to an instruction, an instruction operation, or a group of instruction operations, in various embodiments. Various state associated with an instruction operation may be stored in the reorder buffer (e.g., target logical and physical registers for updating the architectural register map, exceptions or redirections detected during execution, etc.).
[0225] Multiple pointers are shown in FIG. 21. A retired pointer B130 may point to the oldest unretired op in the processor B30. That is, ops prior to the op in retired B130 have been retired from the reorder buffer B108, the architectural state of the processor B30 has been updated to reflect the execution of the retired op, etc. A resolved pointer B132 may point to the oldest op where a preceding branch instruction has been resolved as correctly predicted and a preceding op that could cause an exception has been resolved so that it does not cause an exception. Ops between the retired pointer B130 and resolved pointer B132 may be committed ops in the reorder buffer B108. That is, execution of the instruction that generated the op is completed up to the resolved pointer B132 (barring external interrupts). A newest pointer B134 may point to an op recently fetched and dispatched from the MDR unit B106. Ops between the resolved pointer B132 and newest pointer B134 are speculative and may be flushed due to exceptions, branch mispredictions, etc.
[0226] The true LS NS pointer B136 is the true LS NS pointer described above. The true LS NS pointer can only be generated when an interrupt request is asserted and other tests for a Nack response are negative (e.g., an Ack response is indicated by those tests). The MDR unit B106 can attempt to return the resolved pointer B132 to the true LS NS pointer B136. There may be committed ops in the reorder buffer B108 that cannot be flushed (e.g., once they are committed, they must be completed and retired). Some groups of instruction operations may not be interruptible (e.g., microcode routines, certain non-interruptible exceptions, etc.). In such cases, the processor Int Ack controller B126 may be configured to generate a Nack response. There may be ops or combinations of ops that are too complex to "undo" in the processor B30, and the presence of such ops between the resolved pointer and the true LS NS pointer B136 may cause the processor Int Ack controller B126 to generate a Nack response. If the reorder buffer B108 successfully returns the resolution pointer to the true LS NS pointer B136, the processor Int Ack control circuit B126 may be configured to generate an Ack response.
[0227] FIG. 22 is a flow diagram illustrating the operation of one embodiment of the processor Int Ack control circuit B126 upon receipt of an interrupt request by the processor B30. For ease of understanding, the blocks are shown in a particular order, although other orders may be used. The blocks may be implemented in parallel in combinational logic within the processor Int Ack control circuit B126. The blocks, combinations of the blocks, and / or the overall flowchart may be pipelined over multiple clock cycles. The processor Int Ack control circuit B126 may be configured to implement the operations shown in FIG. 22.
[0228] The processor Int Ack control circuit B126 may be configured to determine whether any Nack conditions have been detected in the MDR unit B106 (decision block B140). For example, potentially long latency operations such as not completing or interrupts being masked may be Nack conditions detected in the MDR unit B106. If so (decision block B140, "Yes" branch), the processor Int Ack control circuit B126 may be configured to generate a Nack response (block B142). If not (decision block B140, "No" branch), the processor Int Ack control circuit B126 may communicate with the LSU to request a Nack condition and / or a true LS NS pointer (block B144). If the LSU B118 detects a Nack condition (decision block B146, "Yes" branch), the processor Int Ack control circuit B126 may be configured to generate a Nack response (block B142). If the LSU B118 does not detect a Nack condition (decision block B146, "No" branch), the processor Int Ack control circuit B126 may be configured to receive the true LS NS pointer from the LSU B118 (block B148) and may attempt to change the resolution pointer in the reorder buffer B108 back to the true LS NS pointer (block B150). If the move is not successful (e.g., there is at least one instruction operation between the true LS NS pointer and the resolution pointer that cannot be flushed) (decision block B152, "No" branch), the processor Int Ack control circuit B126 may be configured to generate a Nack response (block B142). Otherwise (decision block B152, "Yes" branch), the processor Int Ack control circuit B126 may be configured to generate an Ack response (block B154). The processor Int Ack control circuitry B126 may be configured to freeze the resolution pointer to the true LS NS pointer and retire ops until the retirement pointer reaches the resolution pointer (block B156). The processor Int Ack control circuitry B126 may then be configured to take an interrupt (block B158).That is, processor B30 may begin fetching the interrupt code (eg, from a predetermined address associated with the interrupt according to the instruction set architecture implemented by processor B30).
[0229] In another embodiment, SOC B10 may be one of the SOCs in the system. More particularly, in one embodiment, multiple instances of SOC B10 may be employed. Other embodiments may have asymmetric SOCs. Each SOC may be a separate integrated circuit chip (e.g., implemented on a separate semiconductor substrate or "die"). The dies may be packaged and connected to each other via an interposer, a package-on-package solution, etc. Alternatively, the dies may be packaged in a chip-on-chip packaging solution, a multi-chip module, etc.
[0230] FIG. 23 is a block diagram illustrating one embodiment of a system including multiple instances of SOC B10. For example, SOC B10A, SOC B10B, etc. through SOC B10q may be coupled together in the system. Each SOC B10A-B10q includes an instance of an interrupt controller B20 (e.g., interrupt controller B20A, interrupt controller B20B, and interrupt controller B20q in FIG. 23). One interrupt controller, in this example, interrupt controller B20A, can function as the primary interrupt controller for the system. The other interrupt controllers B20B-B20q can function as secondary interrupt controllers.
[0231] The interface between the primary interrupt controller B20A and the secondary controller B20B is shown in more detail in FIG. 23, although interfaces between the primary interrupt controller B20A and other secondary interrupt controllers, such as the interrupt controller B20q, may be similar. In the embodiment of FIG. 23, the secondary controller B20B is configured to provide interrupt information as Int B160 identifying an interrupt issued from an interrupt source on the SOC B10B (or an external device coupled to the SOC B10B, not shown in FIG. 23). The primary interrupt controller B20A is configured to signal hard repeat, soft repeat, and forced repeat to the secondary interrupt controller B20B (reference numeral B162) and to receive Ack / Nack responses from the interrupt controller B20B (reference numeral B164). The interface may be implemented in any manner. For example, dedicated wires may be coupled between the SOC B10A and the SOC B10B to implement reference numerals B160, B162, and / or B164. In another embodiment, messages may be exchanged between the primary interrupt controller B20A and the secondary interrupt controllers B20B-B20q via a generic interface between the SOCs B10A-B10q that is also used for other communications. In one embodiment, programmed input / output (PIO) writes may be used with interrupt data, hard / soft / force requests, and Ack / Nack responses, respectively, as data.
[0232] The primary interrupt controller B20A may be configured to collect interrupts from various interrupt sources that may be on the SOC B10A, one of the other SOCs B10B-B10q, which may be off-chip devices, or any combination thereof. The secondary interrupt controllers B20B-B20q may be configured to transmit interrupts to the primary interrupt controller B20A (Int in FIG. 23) and identify the interrupt source to the primary interrupt controller B20A. The primary interrupt controller B20A may also be responsible for ensuring delivery of the interrupts. The secondary interrupt controllers B20B-B20q may be configured to receive instructions from the primary interrupt controller B20A, receive soft, hard, and forced iteration requests from the primary interrupt controller B20A, and perform iterations via the cluster interrupt controllers B24A-B24n implemented on the corresponding SOCs B10B-B10q. Based on the Ack / Nack responses from the cluster interrupt controllers B24A-B24n, the secondary interrupt controllers B20B-B20q can provide Ack / Nack responses. In one embodiment, the primary interrupt controller B20A can attempt to deliver interrupts serially through the secondary interrupt controllers B20B-B20q in soft and hard iterations, and can deliver in parallel to the secondary interrupt controllers B20B-B20q in forced iterations.
[0233] In one embodiment, the primary interrupt controller B20A may be configured to perform a given iteration on a subset of cluster interrupt controllers integrated on the same SOC B10A as the primary interrupt controller B20A before performing a given iteration on a subset of cluster interrupt controllers on other SOCs B10B-B10q (with the assistance of secondary interrupt controllers B20B-B20q). That is, the primary interrupt controller B20A may attempt to deliver the interrupt serially through the cluster interrupt controllers on the SOC B10A, which may then communicate to the secondary interrupt controllers BB20B-B20q. Delivery attempts through the secondary interrupt controllers B20B-B20q may similarly be performed serially. The order of attempts through the secondary interrupt controllers BB20-B20q may be determined in any desired manner (e.g., programmable order, least recently accepted order, most recently accepted order, etc.), similar to the embodiments described above for the cluster interrupt controllers and processors within a cluster. Thus, the primary interrupt controller B20A and secondary interrupt controllers B20B-B20q can largely insulate software from the presence of multiple SOCs B10A-B10q. That is, the SOCs B10A-B10q can be configured as a single system that is largely transparent to software running on the single system. During system initialization, some embodiments may be programmed to configure the interrupt controllers B20A-B20q as described above, but the interrupt controllers B20A-B20q can otherwise manage the delivery of interrupts across potentially multiple SOCs B10A-B10q, each on a separate semiconductor die, without software assistance or specific software visibility into the multi-die nature of the system. For example, delays due to inter-die communication may be minimized in the system. Thus, during execution after initialization, the single system may appear to software as a single system, and the multi-die nature of the system may be transparent to the software.
[0234] It should be noted that the primary interrupt controller B20A and the secondary interrupt controllers B20B-B20q may operate in a manner also referred to by those skilled in the art as "master" (i.e., primary) and "slave" (i.e., secondary). Although the term primary / secondary is used herein, it is expressly intended that the terms "primary" and "secondary" be interpreted to encompass these corresponding terms.
[0235] In one embodiment, each instance of a SOC B10A-B10q may have both a primary and a secondary interrupt controller circuit implemented within its interrupt controller B20A-B20q. One interrupt controller (e.g., interrupt controller B20A) may be designated as primary during system manufacturing (e.g., via fuses on the SOC B10A-B10q or pin-straps on one or more pins of the SOC B10A-B10q). Alternatively, the primary and secondary designations may be made during system initialization (or boot) configuration.
[0236] FIG. 24 is a flow diagram illustrating the operation of one embodiment of the primary interrupt controller B20A based on receipt of one or more interrupts from one or more interrupt sources. While the blocks are shown in a particular order for ease of understanding, other orders may be used. The blocks may be implemented in parallel within combinational logic within the primary interrupt controller B20A. Blocks, combinations of blocks, and / or the overall flowchart may be pipelined over multiple clock cycles. The primary interrupt controller B20A may be configured to implement the operations shown in FIG. 24.
[0237] The primary interrupt controller B20A may be configured to perform a soft repeat on the cluster interrupt controller integrated on the local SOC B10A (block B170). For example, the soft repeat may be similar to the flowchart of FIG. 18. If the local soft repeat results in an Ack response (decision block B172, "Yes" branch), the interrupt may be delivered normally, and the primary interrupt controller B20A may be configured to return to the idle state B40 (assuming there are no more pending interrupts). If the local soft repeat results in a Nack response (decision block B172, "No" branch), the primary interrupt controller B20A may be configured to select one of the other SOCs B10B-B10q using any desired order as described above (block B174). The primary interrupt controller B20A may be configured to assert a soft repeat request to the secondary interrupt controllers B20B-B20q on the selected SOC B10B-B10q (block B176). If the secondary interrupt controller B20B-B20q provides an Ack response (decision block B178, "Yes" branch), the interrupt may be delivered normally, and the primary interrupt controller B20A may be configured to return to the idle state B40 (assuming there are no more pending interrupts). If the secondary interrupt controller B20B-B20q provides a Nack response (decision block B178, "No" branch) and there are more SOCs B10B-B10q that have not yet been selected in the soft repeat (decision block B180, "Yes" branch), the primary interrupt controller B20A may be configured to select the next SOC B10B-B10q according to the implemented ordering mechanism (block B182), transmit the soft repeat request to the secondary interrupt controller B20B-B20q on the selected SOC (block B176), and continue processing. On the other hand, when each SOC B10B-B10q is selected, the soft iterations can be completed because the serial attempts to deliver interrupts through the secondary interrupt controllers B20B-B20q have been completed.
[0238] Based on completing soft iterations for the secondary interrupt controllers B20B-B20q without successful interrupt delivery (decision block B180, “No” branch), the primary interrupt controller B20A may be configured to perform hard iterations for the local cluster interrupt controllers integrated on the local SOC B10A (block B184). For example, the soft iterations may be similar to the flowchart of FIG. 18. If the local hard iterations result in an Ack response (decision block B186, “Yes” branch), the interrupt may be successfully delivered, and the primary interrupt controller B20A may be configured to return to the idle state B40 (assuming there are no more pending interrupts). If the local hard iterations result in a Nack response (decision block B186, “No” branch), the primary interrupt controller B20A may be configured to select one of the other SOCs B10B-B10q using any desired order as described above (block B188). The primary interrupt controller B20A may be configured to assert a hard repeat request to the secondary interrupt controller B20B-B20q on the selected SOC B10B-B10q (block B190). If the secondary interrupt controller B20B-B20q provides an Ack response (decision block B192, "Yes" branch), the interrupt may be delivered successfully and the primary interrupt controller B20A may be configured to return to the idle state B40 (assuming there are no more pending interrupts). If the secondary interrupt controller B20B-B20q provides a Nack response (decision block B192, "No" branch) and there are further SOCs B10B-B10q that have not yet been selected in the hard iteration (decision block B194, "Yes" branch), the primary interrupt controller B20A may be configured to select the next SOC B10B-B10q according to the implemented ordering mechanism (block B196), transmit a hard iteration request to the secondary interrupt controller B20B-B20q on the selected SOC (block B190), and continue processing.On the other hand, if each SOC B10B-B10q is selected, the hard iterations can be completed because the serial attempts to deliver interrupts through the secondary interrupt controllers B20B-B20q have been completed (decision block B194, "No" branch). The primary interrupt controller B20A may be configured to proceed to a forced iteration (block B198). The forced iterations may be performed locally, or may be performed in parallel or serial across the local SOC B10A and the other SOCs B10B-B10q.
[0239] As mentioned above, there may be a timeout mechanism that may be initialized when the interrupt delivery process begins. If a timeout occurs during any state, in one embodiment, the interrupt controller B20 may be configured to go into a forced repeat. Alternatively, timer expiration may only be considered in the wait drain state B48, as also mentioned above.
[0240] FIG. 25 is a flowchart illustrating the operation of one embodiment of the secondary interrupt controllers B20B-B20q. While the blocks are shown in a particular order for ease of understanding, other orders may be used. The blocks may be implemented in parallel within combinational logic within the secondary interrupt controllers B20B-B20q. The blocks, combinations of blocks, and / or the overall flowchart may be pipelined over multiple clock cycles. The secondary interrupt controllers B20B-B20q may be configured to implement the operations shown in FIG. 25.
[0241] If an interrupt source within (or coupled to) a corresponding SOC B10B-B10q provides an interrupt to a secondary interrupt controller B20B-B20q (decision block B200, "Yes" branch), the secondary interrupt controller B20B-B20q may be configured to transmit the interrupt to the primary interrupt controller B20A for processing along with other interrupts from other interrupt sources (block B202).
[0242] If the primary interrupt controller B20A transmits a repetition request (decision block B204, "Yes" branch), the secondary interrupt controllers B20B-B20q may be configured to implement the requested repetition (hard, soft, or forced) on the cluster interrupt controllers in the local SOCs B10B-B10q (block B206). For example, the hard and soft repetition may be similar to FIG. 18, and the forced repetition may be implemented in parallel on the cluster interrupt controllers in the local SOCs B10B-B10q. If the repetition results in an Ack response (decision block B208, "Yes" branch), the secondary interrupt controllers B20B-B20q may be configured to transmit an Ack response to the primary interrupt controller B20A (block B210). If the repetition results in a Nack response (decision block B208, "No" branch), the secondary interrupt controllers B20B-B20q may be configured to transmit a Nack response to the primary interrupt controller B20A (block B212).
[0243] 26 is a flowchart illustrating one embodiment of a method for handling interrupts. While the blocks are shown in a particular order for ease of understanding, other orders may be used. The blocks may be performed in parallel within combinational logic within the systems described herein. Blocks, combinations of blocks, and / or the flowchart as a whole may be pipelined over multiple clock cycles. The systems described herein may be configured to implement the operations illustrated in FIG. 26.
[0244] The interrupt controller B20 may receive an interrupt from an interrupt source (block B220). In an embodiment having primary and secondary interrupt controllers B20A-B20q, an interrupt may be received at any of the interrupt controllers B20A-B20q and provided to the primary interrupt controller B20A as part of receiving the interrupt from the interrupt source. The interrupt controller B20 may be configured to perform a first iteration (e.g., soft iteration) that serially attempts to deliver the interrupt to the multiple cluster interrupt controllers (block B222). Individual cluster interrupt controllers of the multiple cluster interrupt controllers are associated with individual processor clusters that include multiple processors. A given cluster interrupt controller of the multiple cluster interrupt controllers may be configured to attempt to deliver the interrupt to a subset of the respective multiple processors that are powered on in the first iteration, but not to attempt to deliver the interrupt to some of the respective multiple processors that are not included in the subset. If an ACK response is received, the iteration may be terminated by the interrupt controller B20 (decision block B224, “Yes” branch and block B226). On the other hand (decision block B224, "No" branch), based on non-acknowledgement (Nack) responses from the multiple cluster interrupt controllers in the first iteration, the interrupt controller may be configured to perform a second iteration (e.g., a hard iteration) across the multiple cluster interrupt controllers (block B228). A given cluster interrupt controller may be configured to power on some of the respective multiple processors that are powered off in the second iteration and attempt to deliver interrupts to the respective multiple processors. If an Ack response is received, the iteration may be terminated by the interrupt controller B20 (decision block B230, "Yes" branch and block B232).On the other hand (decision block B230, "No" branch), based on a non-acknowledgement (Nack) response from multiple cluster interrupt controllers in the second iteration, the interrupt controller may be configured to perform a third iteration (e.g., a forced iteration) across the multiple cluster interrupt controllers (block B234).
[0245] According to the present disclosure, a system may include a plurality of cluster interrupt controllers and an interrupt controller coupled to the plurality of cluster interrupt controllers. Each of the plurality of cluster interrupt controllers may be associated with a respective processor cluster including a plurality of processors. The interrupt controller may be configured to receive an interrupt from a first interrupt source, perform a first iteration across the plurality of cluster interrupt controllers to attempt to deliver the interrupt based on the interrupt, and perform a second iteration across the plurality of cluster interrupt controllers based on a non-acknowledgement (Nack) response from the plurality of cluster interrupt controllers in the first iteration. A given cluster interrupt controller of the plurality of cluster interrupt controllers may be configured to, in the first iteration, attempt to deliver the interrupt to a subset of the plurality of processors in a powered-on individual processor cluster without attempting to deliver the interrupt to some of the respective plurality of processors in the respective cluster that are not included in the subset. In the second iteration, the given cluster interrupt controller may be configured to power on any powered-off processors of the respective plurality of processors and attempt to deliver the interrupt to the respective plurality of processors. In one embodiment, during an attempt to distribute the interrupt across the multiple cluster interrupt controllers, the interrupt controller may be configured to assert a first interrupt request to a first cluster interrupt controller of the multiple cluster interrupt controllers, and based on a Nack response from the first cluster interrupt controller, the interrupt controller may be configured to assert a second interrupt request to a second cluster interrupt controller of the multiple cluster interrupt controllers. In one embodiment, during an attempt to distribute the interrupt across the multiple cluster interrupt controllers, based on a second Nack response from the second cluster interrupt controller, the interrupt controller may be configured to assert a third interrupt request to a third cluster interrupt controller of the multiple cluster interrupt controllers.In one embodiment, during an attempt to deliver an interrupt through multiple cluster interrupt controllers, the interrupt controller may be configured to terminate the attempt based on an acknowledgement (Ack) response from a second cluster interrupt controller and a lack of additional pending interrupts. In one embodiment, during an attempt to deliver an interrupt through multiple cluster interrupt controllers, the interrupt controller may be configured to assert an interrupt request to a first cluster interrupt controller of the multiple cluster interrupt controllers, and based on an acknowledgement (Ack) response from the first cluster interrupt controller and a lack of additional pending interrupts, the interrupt controller may be configured to terminate the attempt. In one embodiment, during an attempt to deliver an interrupt across multiple cluster interrupt controllers, the interrupt controller may be configured to serially assert interrupt requests to one or more cluster interrupt controllers of the multiple cluster interrupt controllers, terminated by an acknowledgement (Ack) response from the first cluster interrupt controller of the one or more cluster interrupt controllers. In one embodiment, the interrupt controllers may be configured to assert serially in a programmable order. In one embodiment, the interrupt controller may be configured to assert interrupt requests serially based on a first interrupt source. A second interrupt from a second interrupt source may result in a different order of serial assertion. In one embodiment, during an attempt to deliver an interrupt through the multiple cluster interrupt controllers, the interrupt controller may be configured to assert an interrupt request to a first cluster interrupt controller of the multiple cluster interrupt controllers, and the first cluster interrupt controller may be configured to serially assert processor interrupt requests to the multiple processors in the respective processor clusters based on the interrupt request to the first cluster interrupt controller. In one embodiment, the first cluster interrupt controller is configured to terminate the serial assertion based on an acknowledgement (Ack) response from a first processor of the multiple processors.In one embodiment, the first cluster interrupt controller may be configured to transmit an Acknowledgement response to the interrupt controller based on an Acknowledgement response from the first processor. In one embodiment, the first cluster interrupt controller may be configured to provide a Nack response to the interrupt controller based on Nack responses from multiple processors in individual clusters during serial assertion of processor interrupts. In one embodiment, the interrupt controller may be included on a first integrated circuit on a first substrate including a first subset of the multiple cluster interrupt controllers. A second subset of the multiple cluster interrupt controllers may be implemented on a second integrated circuit on a second, separate semiconductor substrate. The interrupt controller may be configured to serially assert interrupt requests to the first subset before attempting delivery to the second subset. In one embodiment, the second integrated circuit includes a second interrupt controller, and the interrupt controller may be configured to communicate the interrupt request to the second interrupt controller in response to the first subset rejecting the interrupt. The second interrupt controller may be configured to attempt to deliver the interrupt to the second subset.
[0246] In one embodiment, the processor includes a reorder buffer, a load / store unit, and control circuitry coupled to the reorder buffer and the load / store unit. The reorder buffer may be configured to track multiple instruction operations corresponding to instructions fetched by the processor and not retired by the processor. The load / store unit may be configured to execute the load / store operation. The control circuitry may be configured to generate an acknowledge (Ack) response to an interrupt request received by the processor based on a determination that the reorder buffer retires the instruction operation to an interruptible point and that the load / store unit completes the load / store operation to the interruptible point within a specified period of time. The control circuitry may be configured to generate a not acknowledge (Nack) response to the interrupt request based on a determination that at least one of the reorder buffer and the load / store unit does not reach the interruptible point within a specified period of time. In one embodiment, the determination may be a Nack response based on the reorder buffer having at least one instruction operation with a potential execution latency greater than a threshold. In one embodiment, the determination may be a Nack response based on the reorder buffer having at least one instruction operation that causes an interrupt to be masked. In one embodiment, the determination is a Nack response based on the load / store unit having at least one load / store operation to the device address space outstanding.
[0247] In one embodiment, the method includes receiving an interrupt from a first interrupt source in an interrupt controller. The method may further include performing a first iteration of serially attempting to deliver the interrupt to a plurality of cluster interrupt controllers. In the first iteration, an individual cluster interrupt controller of a plurality of cluster interrupt controllers associated with an individual processor cluster comprising a plurality of processors may be configured to attempt to deliver the interrupt to a subset of the plurality of processors in the individual processor cluster that are powered on without attempting to deliver the interrupt to some of the plurality of processors in the individual processor cluster that are not included in the subset. The method may further include performing a second iteration by the interrupt controller to the plurality of cluster interrupt controllers based on a non-acknowledgement (Nack) response from the plurality of cluster interrupt controllers in the first iteration. In the second iteration, a given cluster interrupt controller may be configured to power on some of the plurality of processors that are powered off in the individual processor cluster and attempt to deliver the interrupt to the plurality of processors. In one embodiment, the attempt to serially deliver the interrupt to the plurality of cluster interrupt controllers is terminated based on an acknowledgement from one of the plurality of cluster interrupt controllers. Coherency
[0248] 27-43, various embodiments of cache coherency mechanisms that may be implemented in embodiments of SOC 10 are shown. In one embodiment, the coherency mechanism may include multiple directories configured to track the coherency state of subsets of the unified memory address space. The multiple directories are distributed within the system. In one embodiment, the multiple directories are distributed across memory controllers. In one embodiment, a given memory controller of one or more memory controller circuits includes a directory configured to track multiple cache blocks corresponding to data within a portion of system memory to which the given memory controller interfaces, the directory configured to track which of multiple caches within the system has cached a given cache block of the multiple cache blocks, and the directory is accurate with respect to memory requests ordered and processed in the directory, even if the memory requests have not yet completed within the system. In one embodiment, a given memory controller is configured to issue one or more coherency maintaining commands for the given cache block based on a memory request for the given cache block, the one or more coherency maintaining commands including a cache state for the given cache block in a corresponding one of the plurality of caches, the corresponding cache being configured to delay processing of the given coherency maintaining command based on a cache state in the corresponding cache not matching the cache state in the given coherency maintaining command. In one embodiment, a first cache is configured to store the given cache block in a primary shared state and a second cache is configured to store the given cache block in a secondary shared state, and the given memory controller is configured to cause the first cache to transfer the given cache block to a requestor based on the memory request and the primary shared state in the first cache.In one embodiment, a given memory controller is configured to issue one of a first coherency maintaining command and a second coherency maintaining command to a first cache of the plurality of caches based on a type of the first memory request, the first cache is configured to transfer the first cache block to a requestor that issued the first memory request based on the first coherency maintaining command, and the first cache is configured to return the first cache block to the given memory controller based on the second coherency maintaining command.
[0249] A scalable cache coherency protocol is provided for a system including multiple coherent agents coupled to one or more memory controllers. A coherent agent may generally include a cache for caching memory data, or may otherwise include any circuitry that may take ownership of one or more cache blocks and potentially modify the cache blocks locally. The coherent agents participate in the cache coherency protocol to ensure that modifications made by one coherent agent are visible to other agents that subsequently read the same data, and that modifications made in a particular order by two or more coherent agents (as determined by an ordering point in the system, such as a memory controller for a memory that stores the cache blocks) are observed in that order at each of the coherent agents.
[0250] A cache coherency protocol may specify a set of messages or commands that can be transmitted between an agent and a memory controller (or a coherency controller within a memory controller) to complete a coherent transaction. Messages may include request, snoop, snoop response, and completion. A "request" is a message that initiates a transaction and specifies the requested cache block (e.g., having the address of the cache block) and the state (or minimal state, in some cases, a more permissive state can be provided) in which the requestor receives the cache block. As used herein, a "snoop" or "snoop message" refers to a message transmitted to a coherent agent to request a state change in a cache block, and may also request that the cache block be provided by the coherent agent if the coherent agent has an exclusive copy of or is otherwise responsible for the cache block. A snoop message may be an example of a coherency maintenance command, which may be any command transmitted to a particular coherence agent to cause a change in the coherent state of a cache line within the particular coherent agent. Another term that is an example of a coherency maintenance command is a probe. A coherency-maintaining command does not refer to a broadcast command sent to all coherency agents, as is often used, for example, in shared bus systems. The term "snoop" is used below as an example, but it should be understood that this term generally refers to a coherency-maintaining command. A "completion" or "snoop response" may be a message from a coherent agent indicating that a state change has occurred and, if applicable, providing a copy of a cache block. In some cases, completion may be provided by the request source of a particular request.
[0251] "State" or "cache state" may generally refer to a value indicating whether a copy of a cache block is valid in a cache, and may also indicate other attributes of a cache block. For example, the state may indicate whether the cache block is modified with respect to the copy in memory. The state may indicate the level of ownership of the cache block (e.g., whether the agent with the cache is permitted to modify the cache block, whether the agent is responsible for providing the cache block or returning the cache block to a memory controller if evicted from the cache, etc.). The state may also indicate the possible presence of the cache block in other coherent agents (e.g., a "shared" state may indicate that a copy of the cache block may be stored in one or more other cacheable agents).
[0252] Various embodiments of cache coherency protocols may include various features. For example, memory controller(s) may each implement a coherency controller and a directory for cache blocks corresponding to the memory controlled by that memory controller. The directory may track the state of cache blocks in multiple cacheable agents, allowing the coherency controller to determine which cacheable agents to snoop to change the state of the cache block and possibly provide a copy of the cache block. That is, a snoop need not be broadcast to all cacheable agents based on a request received at the cache controller; rather, the snoop may be transmitted to an agent that has a copy of the cache block affected by the request. When a snoop is generated, the directory may be updated to reflect the state of the cache block in each coherent agent after the snoop is processed and the data is provided to the source of the request. Thus, the directory may be accurate for the next request processed for the same cache block. Snooping may be minimized, reducing traffic on the interconnect between the coherent agents and the memory controller, for example, compared to a broadcast solution. In one embodiment, a "three-hop" protocol may be supported in which one of the caching coherent agents provides a copy of the cache block to the source of the request, or if there is no caching agent, the memory controller provides the copy. Thus, data is provided in three "hops" (or messages transmitted over the interface): a request from the source to the memory controller, a snoop to the coherent agent responding to the request, and completion with the cache block of data from the coherent agent to the source of the request. If there is no cached copy, there may be two hops: a request from the source to the memory controller and completion of the data from the memory controller to the source.There may be additional messages (e.g., completions from other agents indicating the requested state change has occurred when there are multiple snoops on a request), but the data itself may be provided in three hops. In contrast, many cache coherency protocols are four-hop protocols in which the coherent agent responds to a snoop by returning a cache block to the memory controller, which then forwards the cache block to the source. In one embodiment, in addition to three-hop flows, four-hop flows may be supported by the protocol.
[0253] In one embodiment, a request for a cache block may be processed by the coherency controller, and the directory may be updated when a snoop (and / or a completion from the memory controller if there is no cached copy) is generated. Another request for the same cache block may then be serviced. Thus, requests for the same cache block may not be serialized, as is the case in some other cache coherence protocols. Because a message associated with a subsequent request may arrive at a given coherent agent before a message associated with a preceding request ("subsequent" and "previous" refer to requests ordered at the coherency controller within the memory controller), there may be various race conditions that arise when there are multiple outstanding requests for a cache block. To allow the agents to sort the requests, messages (e.g., snoops and completions) may include the expected cache state at the receiving agent, as indicated by the directory at the time the request was processed. Thus, if the receiving agent does not have the cache block in the state indicated in the message, the receiving agent may delay processing the message until the cache state changes to the expected state. The change to the expected state may be made via a message related to the previous request. Additional discussion of race conditions and using predicted cache states to resolve them is provided below with respect to FIGS. 29-30 and 32-34.
[0254] In one embodiment, the cache state may include a primary shared state and a secondary shared state. The primary shared state may apply to a coherent agent responsible for transmitting a copy of a cache block to a requesting agent. The secondary shared agent may not even need to snoop while processing a given request (e.g., reading a cache block that is allowed to be returned in the shared state). Further details regarding the primary and secondary shared states are described with respect to Figures 40 and 42.
[0255] In one embodiment, at least two types of snoops may be supported: snoop forward and snoop back. A snoop forward message may be used to cause a coherent agent to forward a cache block to a requesting agent, and a snoop back message may be used to cause a coherent agent to return a cache block to the memory controller. In one embodiment, a snoop invalidate message may also be supported (and may include forward and back variants to specify the destination of the completion). A snoop invalidate message causes a caching coherent agent to invalidate a cache block. Supporting snoop forward and snoop back flows, for example, can provide both cacheable (snoop forward) and non-cacheable (snoop back) behavior. Because a caching agent stores a cache block and can potentially use the data therein, snoop forward can be used to minimize the number of messages when a cache block is provided to the caching agent. On the other hand, a non-coherent agent may not store the entire cache block, and therefore copy-back to memory can ensure that the complete cache block is retrieved into the memory controller. Thus, the variants or types of snoop forward and snoop back may be selected based on the capabilities of the requesting agent (e.g., based on the identity of the requesting agent) and / or based on the type of request (e.g., cacheable or non-cacheable). Further details regarding snoop forward and snoop back messages are provided below with respect to Figures 37, 38, and 40. Various other features are shown in the remaining figures and described in more detail below.
[0256] FIG. 27 is a block diagram of an embodiment of a system including a system-on-chip (SOC) C10 coupled to one or more memories, such as memories C12A-C12m. The SOC C10 may be, for example, an instance of the SOC C10 shown in FIG. 1. The SOC C10 may include multiple coherent agents (CAs) C14A-C14n. A coherent agent may include one or more processors (P) C16 coupled to one or more caches (e.g., cache C18). The SOC C10 may include one or more non-coherent agents (NCAs) C20A-C20p. The SOC C10 may include one or more memory controllers C22A-C22m, each coupled to a respective memory C12A-C12m during use. Each memory controller C22A-C22m may include a coherency controller circuit C24 (more simply, a "coherency controller" or "CC") coupled to a directory C26. The memory controllers C22A-C22m, non-coherent agents C20A-C20p, and coherent agents C14A-C14n may be coupled to an interconnect C28 for communication between the various components C22A-C22m, C20A-C20p, and C14A-C14n. As indicated by the name, the components of SOC C10 may, in one embodiment, be integrated onto a single integrated circuit “chip.” In other embodiments, the various components may be external to SOC C10 on other chips or even discrete components. Any amount of integrated or discrete components may be used. In one embodiment, a subset of the coherent agents C14A-C14n and memory controllers C22A-C22m may be implemented on one of multiple integrated circuit chips coupled together to form the components shown in SOC C10 of FIG. 27.
[0257] The coherency controller C24 may implement the memory controller portion of a cache coherency protocol. Generally, the coherency controller C24 may be configured to receive requests from the interconnect C28 (e.g., through one or more queues, not shown, in the memory controllers C22A-C22m) targeting cache blocks mapped to the memories C12A-C12m to which the memory controllers C22A-C22m are coupled. The directory may include multiple entries, each of which may track the coherency state of an individual cache block in the system. The coherency state may include, for example, the cache state of the cache block in the various coherent agents C14A-C14N (e.g., in the cache C18 or in other caches, such as a cache in the processor C16, not shown). Thus, based on the directory entry for the cache block corresponding to a given request and the type of a given request, the coherency controller C24 may be configured to determine which coherent agent C14A-C14n should receive the snoop and the type of snoop (e.g., snoop invalidate, snoop share, change to shared, change to owned, change to invalid, etc.). The coherency controller C24 may also independently determine whether a snoop forward or snoop back is transmitted. The coherent agents C14A-C14n may receive the snoop, process the snoop, update the cache block state within the coherent agent C14A-C14n, and (if specified by the snoop) provide a copy of the cache block to the requesting coherent agent C14A-C14n or memory controller C22A-C22m that transmitted the snoop. Further details are provided further below.
[0258] As described above, the coherent agents C14A-C14n may include one or more processors C16. The processor C16 may function as the central processing unit (CPU) of the SOC C10. The system's CPU includes one or more processors that execute the system's main control software, such as an operating system. Generally, the software executed by the CPU during use may control other components of the system to achieve the system's desired functionality. The processors may also execute other software, such as application programs. The application programs may provide user functionality and may rely on the operating system for low-level device control, scheduling, memory management, etc. Thus, the processors may also be referred to as application processors. The coherent agents C14A-C14n may further include other hardware, such as a cache C18 and / or an interface to other components of the system (e.g., an interface to the interconnect C28). Other coherent agents may include processors that are not CPUs. Additionally, other coherent agents may not include a processor (e.g., fixed function circuits such as a display controller or other peripheral circuits, fixed function circuits with processor support via one or more embedded processors, etc. may be coherent agents).
[0259] Generally, a processor may include any circuitry and / or microcode configured to execute instructions defined in an instruction set architecture implemented by the processor. A processor may include a processor core implemented on an integrated circuit with other components as a system-on-chip (SOC) or other level of integration. A processor may further include a discrete microprocessor, a processor core and / or microprocessor integrated in a multi-chip module implementation, a processor implemented as multiple integrated circuits, etc. The number of processors C16 in a given coherent agent C14A-C14n may differ from the number of processors C16 in another coherent agent C14A-C14n. Generally, one or more processors may be included. Furthermore, the processors C16 may differ in microarchitectural implementation, performance and power characteristics, etc. In some cases, processors may also differ in the instruction set architecture they implement, their functionality (e.g., CPU, graphics processing unit (GPU) processor, microcontroller, digital signal processor, image signal processor, etc.), etc.
[0260] Cache C18 may have any capacity and configuration, such as set associative, direct mapped, or fully associative. Cache block size may be any desired size (e.g., 32 bytes, 64 bytes, 128 bytes, etc.). A cache block may be the unit of allocation and deallocation in cache C18. Furthermore, a cache block may be an address space (e.g., an aligned coherence granularity-sized segment of a memory unit) within which coherency is maintained in this embodiment. A cache block may also be referred to as a cache line in some cases.
[0261] In addition to the coherency controller C24 and the directory C26, the memory controllers C22A-C22m may generally include circuitry for receiving memory operations from other components of the SOC C10 and accessing the memories C12A-C12m to complete the memory operations. The memory controllers C22A-C22m may be configured to access any type of memory C12A-C12m. For example, the memories C12A-C12m may be static random access memory (SRAM), double data rate (DRAM) such as synchronous DRAM (SDRAM) including dynamic RAM (DDR, DDR2, DDR3, DDR4, etc.), DRAM, non-volatile memory, graphics DRAM such as graphics DDR DRAM (GDDR), and high bandwidth memory (HBM). Low-power / mobile versions of DDR DRAM (e.g., LPDDR, mDDR, etc.) may be supported. The memory controllers C22A-C22m may include queues for memory operations to order (and possibly reorder) the operations and present them to the memories C12A-C12m. The memory controllers C22A-C22m may further include data buffers for storing write data waiting to be written to memory and read data waiting to be returned to the source of the memory operation (if the data is not provided by a snoop). In some embodiments, the memory controllers C22A-C22m may include a memory cache for storing recently accessed memory data. In SOC implementations, for example, a memory cache can reduce power consumption in the SOC by avoiding re-accessing data from the memories C12A-C12m if it is expected to be accessed again soon. In some cases, a memory cache may be referred to as a system cache, as opposed to a private cache, such as cache C18, or a cache in the processor C16 that serves only certain components. Furthermore, in some embodiments, the system cache need not be located within the memory controllers C22A-C22m.
[0262] The non-coherent agents C20A-C20p may generally include various additional hardware functions (e.g., "peripherals") included in the SOC C10. For example, the peripherals may include video peripherals such as an image signal processor configured to process image capture data from a camera or other image sensor, a GPU, a video encoder / decoder, a scaler, a rotator, a blender, etc. The peripherals may include audio peripherals such as a microphone, a speaker, an interface to a microphone and a speaker, an audio processor, a digital signal processor, a mixer, etc. The peripherals may include interface controllers for various interfaces external to the SOC C10, including interfaces such as Universal Serial Bus (USB), Peripheral Component Interconnect (PCI) including PCI Express (PCIe), serial and parallel ports, etc. The peripherals may include networking peripherals such as a Media Access Controller (MAC). Any set of hardware may be included. In one embodiment, the non-coherent agents C20A-C20p may also include a bridge to the set of peripherals.
[0263] Interconnect C28 can be any communications interconnect and protocol for communicating between components of SOC C10. Interconnect C28 can be bus-based, including shared bus configurations, crossbar configurations, and hierarchical buses with bridges. Interconnect C28 can be packet-based or circuit-switched, and can be hierarchical with bridges, crossbars, point-to-point, or other interconnects. Interconnect C28, in one embodiment, can include multiple independent communications fabrics.
[0264] In general, the number of components C22A-C22m, C20A-C20p, and C14A-C14n may vary from embodiment to embodiment, and any number may be used. As indicated by the "m," "p," and "n" postfixes, the number of components of one type may differ from the number of components of another type. However, the number of a given type may be the same as the number of other types. Furthermore, while the system of FIG. 27 is shown with multiple memory controllers C22A-C22m, embodiments having a single memory controller C22A-C22m are also contemplated and may implement the cache coherency protocols described herein.
[0265] Referring now to FIG. 28, a block diagram illustrates multiple coherent agents C12A-C12D and a memory controller C22A implementing a coherent transaction for a cacheable read exclusive request (CRdEx) according to one embodiment of a scalable cache coherency protocol. A read exclusive request may be a request for an exclusive copy of a cache block, so that any other copies are invalidated by coherent agents C14A-C14D, and the requestor has only one valid copy when the transaction is completed. A memory C12A-C12m having a memory location assigned to the cache block has data at the location assigned to the cache block within memory C12A-C12m, but that data becomes "invalid" even if the requestor modifies the data. A read exclusive request may be used, for example, to allow a requestor the ability to modify a cache block without transmitting an additional request in a cache coherency protocol. If an exclusive copy is not required, other requests may be used (e.g., if a writable copy is not necessarily required by the requestor, a read share request CRdSh may be used). The "C" in the "CRdEx" label may refer to "cacheable." Other transactions may be issued by non-coherent agents (e.g., agents C20A-C20p in FIG. 27), and such transactions may be labeled "NC" (e.g., NCRd). Additional discussion of request types and other messages in transactions is provided further below with respect to FIG. 40 for one embodiment, and further discussion of cache states is provided further below with respect to FIG. 39 for one embodiment.
[0266] In the example of FIG. 28, coherent agent C14A may initiate a transaction by transmitting a read exclusive request to memory controller C22A (which controls the memory location assigned to the address in the read exclusive request). Memory controller C22A (more specifically, coherency controller C24 within memory controller C22A) may read the entry in directory C26 and determine that coherent agent C14D has a cache block in the primary shared state (P) and may therefore be the coherent agent that should provide the cache block to the requesting coherent agent C14D. Coherency controller C24 may generate a snoop forward (SnpFwd[st]) message to coherent agent C14D and issue the snoop forward message to coherent agent C14D. Coherency controller C24 may include an identifier of the current state in the coherent agent that received the snoop according to directory C26. For example, in this case, according to directory C26, the current state in coherent agent C14D is "P." Based on the snoop, coherent agent C14D can access the cache storing the cache block and generate a fill completion (Fill in FIG. 28) with data corresponding to the cache block. Coherent agent C14D can transmit the fill completion to coherent agent C14A. Thus, the system implements a "three-hop" protocol for delivering data to the requestor: CRdEx, SnpFwd[st], and Fill. As indicated by the "[st]" in the SnpFwd[st] message, the snoop forward message may also be coded with the state of the cache block to transition to after the coherent agent processes the snoop. In various embodiments, there may be different variations of the message, or the state may be carried as a field in the message.In the example of Figure 28, the new state of the cache block in the coherent agent may be invalid because the request is a read exclusive request. Other requests may be granted the new shared state.
[0267] Furthermore, the coherency controller C24 can determine from the directory entry for the cache block that coherent agents C14B-C14C have a cache block in the secondary shared state (S). Therefore, it can issue a snoop to each coherent agent that (i) has a cached copy of the cache block and (ii) the state of the block within the coherent agent changes based on the transaction. Because coherent agent C14A has obtained an exclusive copy, the shared copy will be invalidated, and therefore the coherency controller C24 can generate a snoop invalidate (SnpInvFw) message for coherent agents C14B-C14C and issue a snoop to coherent agents C14B-C14C. The snoop invalidate message includes an identifier indicating that the current state of coherent agents C14B-C14C is shared. Coherent agents C14B-C14C can process the snoop invalidate request and provide an acknowledgement (Ack) completion to coherent agent C14A. Note that in the illustrated protocol, messages from the snooping agent to the coherency controller C24 are not implemented in this embodiment. The coherency controller C24 can update the directory entry based on the issuance of the snoop and can process the next transaction. Therefore, as mentioned above, in this embodiment, transactions to the same cache block cannot be serialized. The coherency controller C24 can allow additional transactions to the same cache block to begin and can identify which snoop belongs to which transaction based on the current state indication in the snoop (e.g., the next transaction to the same cache block will detect the cache state corresponding to the completed previous transaction). In the illustrated embodiment, the snoop invalidation message is a SnpInvFw message because the completion is sent to the initiating coherent agent C14A as part of the three-hop protocol.In one embodiment, a four-hop protocol is also supported for a particular agent. In such an embodiment, the SnpInvBk message may be used to indicate that the snooping agent should return completion to the coherency controller C24.
[0268] Thus, the cache state identifier in the snoop allows the coherent agent to resolve conflicts between messages forming different transactions to the same cache block. That is, messages may be received in a different order than the order in which the corresponding requests were processed by the coherency controller. The order in which the coherency controller C24 processes requests to the same cache block via the directory C26 can define the order of the requests. That is, the coherency controller C24 can be the ordering point for transactions received at a given memory controller C22A-C22m. Meanwhile, message serialization can be managed within the coherent agent C14A-C14n based on the current cache state corresponding to each message and the cache state within the coherent agent C14A-C14n. A given coherent agent can be configured to access the cache block within the coherent agent based on the snoop and compare the cache state specified in the snoop with the cache state currently in the cache. If the states do not match, the snoop belongs to the transaction ordered after another transaction that changes the cache state within the agent to the state specified in the snoop. Thus, the snooping agent may be configured to delay processing of the snoop based on the first state not matching the second state until the second state is changed to the first state in response to a different communication related to a different request than the first request. For example, the state may change based on fill completions received by the snooping agent from different transactions, etc.
[0269] In one embodiment, the snoop may include a completion count (Cnt) indicating the number of completions corresponding to the transaction, so the requester can determine when all of the completions associated with the transaction have been received. The coherency controller C24 may determine the completion count based on the state indicated in the cache block's directory entry. The completion count may be, for example, the number of completions minus 1 (e.g., 2 in the example of FIG. 28 because there are three completions). This implementation may allow the completion count to be used as initialization of the transaction's completion counter when the first completion of the transaction is received by the requesting agent (e.g., it has already been decremented to reflect the receipt of a completion carrying the completion count). Once the count is initialized, further completions of the transaction may cause the requesting agent to update the completion counter (e.g., decrement the counter). In other embodiments, an actual completion count may be provided and may be decremented by the requester to initialize the completion count. In general, the completion count may be any value that identifies the number of completions the requester should observe before the transaction is fully completed. That is, the requesting agent may complete the request based on the completion counter.
[0270] 29 and 30 illustrate exemplary race conditions that may occur for transactions to the same cache block and the use of a given agent's current cache state (also referred to as the "expected cache state") as reflected in the directory when the transaction is processed by the memory controller and the given agent's current cache state (e.g., as reflected in the given agent's cache(s) or buffers that may temporarily store cache data). In FIGS. 29 and 30, coherent agents are listed as CA0 and CA1, and the memory controller associated with the cache block is shown as MC. Vertical lines 30, 32, and 34 in CA0, CA1, and MC indicate the sources (bases of the arrows) and destinations (tips of the arrows) of various messages corresponding to the transaction. In FIGS. 29 and 30, time progresses from top to bottom. A memory controller may be associated with a cache block if the memory to which the memory controller is coupled contains memory locations assigned to the cache block's address.
[0271] Figure 29 illustrates a race condition between a fill completion for one transaction and a snoop for a different transaction on the same cache block. In the example of Figure 29, CA0 initiates a read exclusive transaction with a CRdEx request to MC (arrow 36). CA1 also initiates a read exclusive transaction with a CRdEx request (arrow 38). The CA0 transaction is processed first by MC, establishing the CA0 transaction as ordered before the CA1 request. In this example, the directory indicates that there is no cached copy of the cache block in the system, so MC responds to the CA0 request by filling in an exclusive state (FillE, arrow 40). MC updates the cache block's directory entry with an exclusive state for CA0.
[0272] The MC selects the CRdEx relationship from CA1 for processing and detects that CA0 has the cache block in an exclusive state. Therefore, the MC can generate a snoop forward request to CA0 (SnpFwdI) requesting that CA0 invalidate the cache block in its cache(s) and provide it to CA1. The snoop forward request also includes an identifier for the E state of CA0's cache block, since this is the cache state reflected in CA0's directory. The MC can issue a snoop (arrow 42) and update the directory to indicate that CA1 has the exclusive copy and that CA0 no longer has a valid copy.
[0273] The snoop and fill completion may arrive at CA0 in either time order. Messages may travel in different virtual channels and / or other delays in the interconnect may allow the messages to arrive in either order. In the illustrated example, the snoop arrives at CA0 before the fill completion. However, CA0 may delay processing the snoop because the expected state in the snoop (E) does not match the current state of the cache block in CA0 (I). The fill completion can then arrive at CA0. CA0 may write the cache block to the cache and set the state to exclusive (E). CA0 may also be permitted to perform at least one operation on the cache block to support forward progress of tasks in CA0, which may change the state to modified (M). In a cache coherence protocol, directory C26 may not track the M state separately (e.g., it may be treated as E) but may match the E state as the expected state in the snoop. CA0 can issue a fill completion to CA1 in the modified state (FillM, arrow 44). Thus, the race condition between the snoop and fill completion of the two transactions has been handled correctly.
[0274] In the example of Figure 29, the CRdEx request is issued by CA1 following the CRdEx request from CA0, but the CRdEx request may also be issued by CA1 before the CRdEx request from CA0, and the CRdEx request from CA0 may still be ordered by the MC before the CRdEx request from CA1 because the MC is the ordering point for the transaction.
[0275] Figure 30 illustrates a race condition between a snoop for one coherent transaction and the completion of another coherent transaction for the same cache block. In Figure 30, CA0 initiates a writeback transaction (CWB) to write a modified cache block to memory (arrow 46), but the cache block may actually be tracked as exclusive in the directory as described above. The CWB may be transmitted, for example, when CA0 evicts the cache block from its cache but the cache block is in a modified state. CA1 initiates a read share transaction (CRdS) for the same cache block (arrow 48). The CA1 transaction is ordered before the CA0 transaction by the MC, and the MC reads the directory entry for the cache block and determines that CA0 has the cache block in an exclusive state. The MC issues a snoop forward request to CA0, requesting a change to the secondary shared state (SnpFwdS, arrow 50). An identifier in the snoop indicates the current cache state of exclusive (E) in CA0. MC updates the directory entry to indicate that CA0 has the cache block in the secondary shared state and CA1 has a copy in the primary shared state (since the previous exclusive copy has been provided to CA1).
[0276] The MC processes the CWB request from CA0 and re-reads the directory entry for the cache block. The MC issues an Ack Complete (arrow 52) indicating that the current cache state is Secondary Shared (S) in CA0, along with an identifier for the cache state in the Ack Complete. Based on the fact that the expected state of Secondary Shared does not match the modified current state, CA0 can delay processing the Ack Complete. Processing the Ack Complete allows CA0 to discard the cache block and will not have a copy of the cache block to provide to CA1 in response to a later-arriving SnpFwdS request. When the SnpFwdS request is received, CA0 can provide a Fill Complete (arrow 54) to CA1, bringing the cache block to the Primary Shared state (P). CA0 can also change the state of the cache block in CA0 to Secondary Shared (S). The state change matches the expected state of the Ack Complete, so CA0 can invalidate the cache block and complete the CWB transaction.
[0277] 31 is a block diagram illustrating one embodiment of a portion of one embodiment of coherent agent C14A in more detail. Other coherent agents C14B-C14n may be similar. In the illustrated embodiment, coherent agent C14A may include request control circuitry C60 and request buffers C62. Request buffers C62 are coupled to request control circuitry C60, and both request buffers C62 and request control circuitry C60 are coupled to cache C18 and / or processor C16, as well as interconnect C28.
[0278] The request buffer C62 may be configured to store multiple requests generated by the cache C18 / processor C16 for coherent cache blocks. That is, the request buffer C62 may store requests to initiate transactions on the interconnect C28. FIG. 31 shows one entry of the request buffer C62, but other entries may be similar. An entry may include a valid (V) field C63, a request (Req.) field C64, a count valid (CV) field C66, and a completion count (CompCnt) field C68. The valid field C63 may store a valid indication (e.g., a valid bit) indicating whether the entry is valid (e.g., stores an outstanding request). The request field C64 may store data defining the request (e.g., request type, address of the cache block, tag or other identifier for the transaction, etc.). The count valid field C66 may store a valid indication for the completion count field C68, indicating that the completion count field C68 has been initialized. The request control circuit C68 can use the count valid field C66 when processing completions received from the interconnect C28 for a request to determine whether the request control circuit C68 should initialize the field with the completion count contained in the completion (count field not valid) or whether it should update the completion count, such as decrementing the completion count (count field valid). The completion count field C68 can store the current completion count.
[0279] The request control circuit C60 may receive requests from the cache 18 / processor 16 and may allocate request buffer entries in the request buffer C62 to the requests. The request control circuit C60 may track the requests in the buffer C62, cause the requests to be transmitted on the interconnect C28 (e.g., according to any type of arbitration scheme), track received completions of the requests, complete the transaction, and transfer the cache block to the cache C18 / processor C16.
[0280] Referring now to FIG. 32, a flow diagram illustrating the operation of one embodiment of a coherency controller C24 within a memory controller C22A-C22m upon receiving a request to be processed is shown. The operations of FIG. 32 may be performed when a request is selected from among the received requests for service at the memory controller C22A-C22m via any desired arbitration algorithm. While the blocks are shown in a particular order for ease of understanding, other orders may be used. The blocks may be performed in parallel in combinational logic within the coherency controller C24. Blocks, combinations of blocks, and / or the entire flowchart may be pipelined over multiple clock cycles. The coherency controller C24 may be configured to implement the operations shown in FIG. 32.
[0281] The coherency controller C24 may be configured to read a directory entry from the directory C26 based on the address of the request. The coherency controller C24 may be configured to determine which snoop to generate based on the type of request (e.g., the state requested for the cache block by the requestor) and the current state of the cache block in the various coherent agents C14A-C14n indicated in the directory entry (block C70). The coherency controller C24 may also generate a current state to be included in each snoop based on the current state of the coherent agents C14A-C14n receiving the snoop as indicated in the directory. The coherency controller C24 may be configured to insert the current state into the snoop (block C72). The coherency controller C24 may also be configured to generate a completion count and insert the completion count into each snoop (block C74). As previously mentioned, the completion count, in one embodiment, may be the number of completions minus one or the total number of completions. The completion number may be the number of snoops, or, if the memory controller C22A-C22m provides the cache block, it may be a fill completion from the memory controller C22A-C22m. In most cases where there is a snoop for a cacheable request, one of the snooped coherent agents C14A-C14n can provide the cache block, and therefore the completion number may be the number of snoops. However, if the coherent agents C14A-C14n do not have a copy of the cache block (no snoop), for example, the memory controller may provide a fill completion. The coherency controller C24 may be configured to queue the snoop for transmission to the coherent agents C14A-C14n (block C76). Upon successful queuing of the snoop, the coherency controller C24 may be configured to update the directory entry to reflect the completion of the request (block C78).For example, the update may change the cache state tracked in the directory entry to match the cache state requested by the snoop, change the agent identifier indicating which agent should provide a copy of the cache block to coherent agents C14A-C14n that will place the cache block in an exclusive, modified, owned, or primary shared state upon completion of the transaction, etc.
[0282] Referring now to FIG. 33, a flow diagram illustrates the operation of one embodiment of a request control circuit C60 in a coherent agent C14A-C14n upon receiving completion of an outstanding request in a request buffer C62. While the blocks are shown in a particular order for ease of understanding, other orders may be used. The blocks may be implemented in parallel in combinational logic within the request control circuit C60. Blocks, combinations of blocks, and / or the entire flowchart may be pipelined over multiple clock cycles. The request control circuit C60 may be configured to implement the operations illustrated in FIG. 33.
[0283] The request control circuit C60 may be configured to access the request buffer entry in the request buffer C62 associated with the request with which the received completion is associated. If the count valid field C66 indicates that the completion count is valid (decision block C80, "yes" branch), the request control circuit C60 may be configured to decrement the count in the request count field C68 (block C82). If the count is zero (decision block C84, "yes" branch), the request is completed, and the request control circuit C60 may be configured to forward an indication of the completion (and, if applicable, the received cache block) to the cache C18 and / or processor C16 that generated the request (block C86). The completion may update the state of the cache block. If the new state of the cache block after the update matches the state expected in the pending snoop (decision block C88, "yes" branch), the request control circuit C60 may be configured to process the pending snoop (block C90). For example, the request control circuit C60 may be configured to pass the snoop to the cache C18 / processor C16 and generate a completion corresponding to the pending snoop (and change the state of the cache block as indicated by the snoop).
[0284] A new state may match an expected state if the new state is the same as the expected state. Additionally, a new state may match an expected state if the expected state is the state tracked by the directory C26 for new states. For example, in one embodiment, the modified state is tracked as an exclusive state in the directory C26, and therefore the modified state matches an expected state of exclusive. The new state may be modified, for example, if the state is provided in a fill completion transmitted by another coherent agent C14A-C14n that has the cache block exclusive and has locally modified the cache block.
[0285] If the count valid field C66 indicates that the completion count is valid (decision block C80) and the completion count is not zero after decrementing (decision block C84, "no" branch), the request has not completed and therefore remains pending in the request buffer C62 (and any pending snoops waiting for the request to complete may remain pending). If the count valid field C66 indicates that the completion count is not valid (decision block C80, "no" branch), the request control circuit C60 may be configured to initialize the completion count field C68 with the completion count provided in the completion (block C92). The request control circuit C60 may still be configured to check that the completion count is 0 (e.g., if there is only one completion for the request, the completion count may be 0 in the completion) (decision block C84), and processing may continue as described above.
[0286] FIG. 34 is a flowchart illustrating the operation of one embodiment of coherent agents C14A-C14n upon receipt of a snoop. For ease of understanding, the blocks are shown in a particular order, although other orders may be used. The blocks may be performed in parallel in combinational logic within coherent agents 14CA-C14n. Blocks, combinations of blocks, and / or the entire flowchart may be pipelined over multiple clock cycles. Coherent agents 14CA-C14n may be configured to implement the operations shown in FIG.
[0287] Coherent agents C14A-C14n may be configured to check the predicted state in the snoop against the state in cache C18 (decision block C100). If the predicted state does not match the current state of the cache block (decision block C100, "no" branch), a completion is pending that changes the current state of the cache block to the predicted state. The completion corresponds to a transaction ordered before the transaction corresponding to the snoop. Thus, coherent agents C14A-C14n may be configured to suspend the snoop, delaying processing of the snoop until the current state changes to the predicted state indicated in the snoop (block C102). In one embodiment, the suspended snoop may be stored in a buffer dedicated to suspended snoops. Alternatively, the suspended snoop may be absorbed into an entry in request buffer C62 that stores a conflicting request, as described in more detail below with respect to FIG. 36.
[0288] If the expected state matches the current state (decision block C100, "yes" branch), the coherent agent C14A-C14n may be configured to process the state change based on the snoop (block C104). That is, the snoop may indicate a desired state change. The coherent agent C14A-C14n may be configured to generate a completion (e.g., a fill if the snoop is a snoop forward request, a copyback snoop response if the snoop is a snoop back request, or an acknowledgement (forward or back based on the snoop type) if the snoop is a state change request). The coherent agent may be configured to generate a completion with a completion count from the snoop (block C106) and queue the completion for transmission to the requesting coherent agent C14A-C14n (block C108).
[0289] Using the cache coherency algorithms described herein, cache blocks can be transmitted from one coherent agent C14A-C14n to another through a chain of competing requests with low message bandwidth overhead. For example, FIG. 35 is a block diagram illustrating the transmission of a cache block among four coherent agents CA0-CA3. As in FIGS. 29 and 30, the coherent agents are listed as CA0-CA3, and the memory controller associated with the cache block is designated as MC. Vertical lines 110, 112, 114, 116, and 118 for CA0, CA1, CA2, CA3, and MC indicate the sources (arrow bases) and destinations (arrow tips) of various messages corresponding to the transaction, respectively. Time progresses from top to bottom in FIG. 35. At the point corresponding to the top of FIG. 35, coherent agent CA3 has the cache block involved in the transaction in a modified state (tracked as exclusive in directory C26). The transactions in FIG. 35 are all for the same cache block.
[0290] Coherent agent CA0 initiates a read exclusive transaction with a CRdEx request to the memory controller (arrow 120). Coherent agents CA1 and CA2 also initiate read exclusive transactions (arrows 122 and 124, respectively). As indicated by the tips of arrows 120, 122, and 124 on line 118, memory controller MC orders the transaction as CA0, then CA1, and finally CA2. In the exclusive state, the directory state of the transaction from CA0 is CA3, so a snoop forward and invalidate (SnpFwdI) is transmitted with the current cache state as exclusive (arrow 126). Coherent agent CA3 receives the snoop and forwards a FillM completion with the data to coherent agent CA0 (arrow 128). Similarly, the directory state of the transaction from CA1 is coherent agent CA0 in exclusive state (from the preceding transaction to CA0), so the memory controller MC issues SnpFwdI to coherent agent CA0 with a current cache state of E (arrow 130), and the directory state of the transaction from CA2 is coherent agent CA1 with a current cache state of E (arrow 132). When coherent agent CA0 has an opportunity to perform at least one memory operation on the cache block, coherent agent CA0 responds with a FillM completion to coherent agent CA1 (arrow 134). Similarly, when coherent agent CA1 has an opportunity to perform at least one memory operation on the cache block, coherent agent CA1 responds to its snoop by returning a FillM completion to coherent agent CA2 (arrow 136). The order and timing of the various messages may vary (e.g., similar to the race conditions shown in Figures 29 and 30), but generally, cache blocks may move from agent to agent with one extra message (FillM Complete) once the conflicting requests are resolved.
[0291] In one embodiment, due to the race condition described above, a snoop may be received before the fill completion it is supposed to snoop (as detected by the snoop carrying the expected cache state). Furthermore, a snoop may be received before the Ack completion is collected and the fill completion can be processed. The Ack completion is due to the snoop and therefore depends on the progress of the virtual channel carrying the snoop. Thus, a conflicting snoop (delayed wait on the expected cache state) may fill internal buffers and backpressure the fabric, which can cause deadlock. In one embodiment, coherent agents C14A-C14n may be configured to absorb one snoop forward and one snoop invalidate into outstanding requests in the request buffer rather than allocating separate entries. Non-conflicting snoops, or conflicting snoops that reach a point where they can be processed without further interconnect dependency, may then flow around the conflicting snoop, avoiding deadlock. When a snoop forward occurs, one snoop forward and one snoop invalidation absorption may be sufficient because the responsibility for the forwarding is transferred to the target. Therefore, another snoop forward will not occur again until the requester completes its current request and issues another new request after the previous snoop forward completes. When a snoop invalidation occurs, the requester will not receive another invalidation again until it is invalid according to the directory, processes the previous invalidation, requests the cache block again, and obtains a new copy.
[0292] Thus, coherent agents C14A-C14n may be configured to help ensure forward progress and / or prevent deadlock by detecting a snoop received by a coherent agent to a cache block for which the coherent agent has an outstanding request ordered before the snoop. The coherent agent may be configured to absorb the second snoop into the outstanding request (e.g., into a request buffer entry that stores the request). The coherent agent may process the absorbed snoop after completing the outstanding request. For example, if the absorbed snoop is a snoop forward request, the coherent agent may be configured to forward the cache block to another coherent agent indicated in the snoop forward snoop (and may change the cache state to the state indicated by the snoop forward request) after completing the outstanding request. If the absorbed snoop is a snoop invalidate request, the coherent agent may update the cache state to invalid and transmit an acknowledge completion after completing the outstanding request. Absorbing a snoop into a conflicting request may be implemented, for example, by including additional storage in each request buffer entry for data describing the absorbed snoop.
[0293] FIG. 36 is a flowchart illustrating the operation of one embodiment of coherent agents C14A-C14n upon receipt of a snoop. While the blocks are shown in a particular order for ease of understanding, other orders may be used. The blocks may be performed in parallel in combinational logic within coherent agents C14A-C14n. The blocks, combinations of blocks, and / or the entire flowchart may be pipelined over multiple clock cycles. Coherent agents C14A-C14n may be configured to implement the operations shown in FIG. 36. For example, the operations shown in FIG. 36 may be part of detecting a snoop that has a predicted cache state that does not match the predicted cache state and is deferred (decision block C100 and block C102 of FIG. 34).
[0294] Coherent agents C14A-C14n may be configured to compare the address of the snoop that is deferred due to lack of a consistent cache state with the addresses of outstanding requests (or pending requests) in request buffer C62. If an address conflict is detected (decision block C140, “yes” branch), request buffer C62 may absorb the snoop into a buffer entry assigned to the pending request for which the address conflict was detected (block C142). If there is no address conflict with the pending request (decision block C140, “no” branch), coherent agents C14A-C14n may be configured to allocate a separate buffer location for the snoop (e.g., in request buffer C62 or another buffer in coherent agents C14A-C14n) and store data describing the snoop in the buffer entry (block C144).
[0295] As previously mentioned, the cache coherency protocol, in one embodiment, can support both cacheable and non-cacheable requests while maintaining coherency of associated data. Non-cacheable requests may be issued, for example, by non-coherent agents C20A-C20p, which may not have the ability to coherently store cache blocks. In one embodiment, coherent agents C14A-C14n may also be able to issue non-cacheable requests, and the coherent agents may not cache data provided in response to such requests. Thus, for example, if the data requested by a given non-coherent agent C20A-C20p is in a modified cache block in one of the coherent agents C14A-C14n and the modified cache block is forwarded to the given non-coherent agent C20A-C20p in anticipation of being saved by the given non-coherent agent C20A-C20p, a snoop forward request for a non-cacheable request is not appropriate.
[0296] To support coherent non-cacheable transactions, one embodiment of a scalable cache coherency protocol may include multiple types of snoops. For example, in one embodiment, snoops may include snoop forward requests and snoop back requests. As previously mentioned, snoop forward requests can cause a cache block to be forwarded to a requesting agent, while snoop back requests can cause a cache block to be returned to the memory controller. In one embodiment, snoop invalidate requests may be supported to invalidate a cache block (with forward and back versions to indicate completion).
[0297] More specifically, the memory controller C22A-C22m (and even more specifically, the coherency controller C24 within the memory controller C22A-C22m) that receives the request may be configured to read from the directory C26 an entry corresponding to the cache block identified by the address in the request. The memory controller C22A-C22m may be configured to issue a snoop to a given one of the coherent agents C14A-C14m that has a cached copy of the cache block according to the entry. The snoop indicates that the given agent should transmit the cache block to the source of the request based on the first request being of a first type (e.g., a cacheable request). The snoop indicates that the given agent should transmit the first cache block to the memory controller based on the first request being of a second type (e.g., a non-cacheable request). The memory controller C22A-C22n may be configured to respond with a completion to the source of the request based on receiving the cache block from the given agent. Additionally, similar to other coherent requests, memory controllers C22A-C22n may be configured to update entries in directory C26 to reflect completion of non-cacheable requests based on issuing multiple snoops for the non-cacheable requests.
[0298] Figure 37 is a block diagram illustrating an example of a non-cacheable transaction that is managed coherently in one embodiment. Figure 37 may be an example of a four-hop protocol for passing snooped data through a memory controller to a requestor. The non-coherent agent is listed as NCA0, the coherent agent is listed as CA1, and the memory controller associated with the cache block is listed as MC. Vertical lines 150, 152, and 154 in NCA0, CA1, and MC indicate the sources (arrow bases) and destinations (arrow tips) of various messages corresponding to the transaction. Time progresses from top to bottom in Figure 37.
[0299] At the time corresponding to the top of Figure 37, coherent agent CA1 has the cache block in an exclusive state (E). NCA0 issues a non-cacheable read request (NCRd) to MC (arrow 156). MC determines from directory 26 that CA1 has the cache block containing the data requested by NCRd in an exclusive state and generates a snoop-back request (SnpBkI(E)) to CA1 (arrow 158). CA1 provides a copy-back snoop response (CpBkSR) with the cache block of data to MC (arrow 160). If the data was modified, MC can update memory with the data and provide the data for the non-cacheable read request to NCA0 in a non-cacheable read response (NCRdRsp) (arrow 162), completing the request. In one embodiment, there can be two or more types of NCRd requests: a request to invalidate the cache block in the snooped coherent agent and a request to allow the snooped coherent agent to retain the cache block. The above discussion indicates invalidation. In other cases, the snooped agent may keep the cache block in the same state.
[0300] A non-cacheable write request may be similarly implemented, using a snoop-back request to obtain the cache block and modifying the cache block with the non-cacheable write data before writing it to memory. A non-cacheable write response may still be provided to notify the non-cacheable agent (NCA0 in Figure 37) that the write has completed.
[0301] FIG. 38 is a flowchart illustrating the operation of one embodiment of a memory controller C22A-C22m (more specifically, in one embodiment, a coherency controller 24 within the memory controller C22A-C22m) in response to a request, illustrating cacheable and non-cacheable operations. The operations illustrated in FIG. 38 may be, for example, a more detailed view of a portion of the operations shown in FIG. 32. For ease of understanding, the blocks are shown in a particular order, but other orders may be used. The blocks may be performed in parallel in combinational logic within the coherency controller C24. The blocks, combinations of the blocks, and / or the entire flowchart may be pipelined over multiple clock cycles. The coherency controller C24 may be configured to implement the operations illustrated in FIG. 38.
[0302] The coherency controller C24 may be configured to read the directory based on the address in the request. If the request is a directory hit (decision block C170, "yes" branch), the cache block is present in one or more caches in the coherent agents C14A-C14n. If the request is non-cacheable (decision block C172, "yes" branch), the coherency controller C24 may be configured to issue a snoop-back request to the coherent agents C14A-C14n responsible for providing copies of the cache block (and, if applicable, snoop invalidation requests to the shared agents (back variant)—block C174). The coherency controller C24 may be configured to update the directory to reflect that the snoop has completed (e.g., invalidate the cache block in the coherent agents C14A-C14n—block C176). The coherency controller C24 may be configured to wait for a copyback snoop response (decision block C178, "Yes" branch) as well as any Ack snoop responses from the shared coherent agents C14A-C14n, and may be configured to generate a non-cacheable completion to the requesting agent (NCRdRsp or NCWrRsp as appropriate) (block C180). Data may also be written to memory by the memory controllers C22A-C22m if the cache block was modified.
[0303] If the request is cacheable (decision block C172, "No" branch), the coherency controller C24 may be configured to generate a snoop forward request (block C182) to the coherent agent C14A-C14n responsible for transferring the cache block, as well as other snoops to other caching coherent agents C14A-C14n, if necessary. The coherency controller C24 may update the directory C24 to reflect completion of the transaction (block C184).
[0304] If the request is not a hit in directory C26 (decision block C170, "No" branch), there is no cached copy of the cache block in coherent agent C14A-C14n. In this case, no snoop may be generated, and memory controller C22A-C22m may be configured to generate a fill completion (for a cacheable request) or a non-cacheable completion (for a non-cacheable request) to provide the data or complete the request (block C186). For a cacheable request, coherency controller C24 may update directory C26 to create an entry for the cache block and may initialize the requesting coherent agent C14A-C14n as having a copy of the cache block in the cache state requested by coherent agent C14A-C14n (block C188).
[0305] FIG. 39 is a table C190 illustrating exemplary cache states that may be implemented in one embodiment of a coherent agent C14A-C14n. Other embodiments may employ different cache states, subsets of the cache states shown and other cache states, supersets of the cache states shown and other cache states, etc. The modified state (M) or “dirty exclusive” state may be a state within a coherent agent C14A-C14n that has only a cached copy of the cache block (the copy is exclusive) and the data in the cached copy has been modified with respect to the corresponding data in memory (e.g., at least one byte of the data differs from the corresponding byte in memory). Modified data is sometimes referred to as dirty data. The owned state (O) or “dirty shared” state may be a state within a coherent agent C14A-C14n that has a modified copy of the cache block but may be sharing a copy with at least one other coherent agent C14A-C14n (although the other coherent agent C14A-C14n may have subsequently evicted the shared cache block). Other coherent agents C14A-C14n place the cache block in a secondary shared state. An exclusive state (E) or "clean exclusive" state may be a state within a coherent agent C14A-C14n that has only a cached copy of the cache block, but the cached copy has the same data as the corresponding data in memory. An exclusive no data (EnD) state, or "clean exclusive, no data" state may be a state within a coherent agent C14A-C14n that is similar to the exclusive (E) state, except that the cache block of data has not been distributed to the coherent agent. Such a state may be used when a coherent agent C14A-C14n is to modify every byte within a cache block and therefore there is no benefit or coherency reason to supply the previous data in the cache block. The EnD state may be an optimization to reduce traffic on the interconnect C28 and may not be implemented in other embodiments.The primary shared (P) state or “clean shared primary” state may be a state within a coherent agent C14A-C14n that has a shared copy of a cache block but is also responsible for forwarding the cache block to another coherent agent based on a snoop forward request. The secondary shared (S) state, or “clean shared secondary” state, may be a state within a coherent agent C14A-C14n that has a shared copy of a cache block but is not responsible for providing the cache block if another coherent agent C14A-C14n has the cache block in the primary shared state. In some embodiments, if a coherent agent C14A-C14n does not have the cache block in the primary shared state, the coherency controller C24 may select a secondary shared agent to provide the cache block (and may send a snoop forward request to the selected coherent agent). In other embodiments, if no coherent agents C14A-C14n are in the primary shared state, the coherency controller C24 can cause the memory controllers C22A-C22m to provide the cache block to the requestor. The invalid state (I) can be a state in a coherent agent C14A-C14n that does not have a cached copy of the cache block. A coherent agent C14A-C14n in the invalid state may not have previously requested a copy, or may have a copy and have invalidated it based on a snoop or eviction of the cache block to cache a different cache block.
[0306] Figure 40 is a table C192 illustrating various messages that may be used in one embodiment of a scalable cache coherence protocol. In other embodiments, there may be alternative messages, a subset of the illustrated messages and additional messages, a superset of the illustrated messages and additional messages, etc. The messages may carry a transaction identifier that links messages from the same transaction (e.g., initial request, snoop, completion). The initial request and snoop may carry the address of the cache block affected by the transaction. Some other messages may also carry an address. In some embodiments, all messages may carry an address.
[0307] A cacheable read transaction can be initiated with a cacheable read request message (CRd). There can be various versions of the CRd request to request different cache states. For example, CRdEx may request an exclusive state, CRdS may request a secondary shared state, etc. The cache state actually provided in response to a cacheable read request can be at least as permissive as the requested state, or even more permissive. For example, CRdEx may receive the cache block in an exclusive or modified state. CRdS can receive the block in a primary shared, exclusive, owned, or modified state. In one embodiment, a timely CRd request can be implemented to provide the most permissive state possible (without invalidating other copies of the cache block) (e.g., exclusive if no other coherent agent has a cache copy, owned or primary shared if a cache copy is present, etc.).
[0308] A change to exclusive (CtoE) message may be used by a coherent agent that has a copy of a cache block in a state that does not allow modification (e.g., owned, shared primarily, shared secondary), and the coherent agent is attempting to modify the cache block (e.g., the coherent agent needs exclusive access to change the cache block to modified). In one embodiment, a conditional CtoE message may be used for a store-conditional instruction. A store-conditional instruction is part of a load-reserve / store-conditional pair in which a load obtains a copy of a cache block and sets a reservation for that cache block. A coherent agent C14A-C14n can monitor accesses to the cache block by other agents and conditionally perform the store based on whether the cache block has not been modified by another coherent agent C14A-C14n between the load and the store (store successfully if the cache block is not modified, do not store if the cache block is modified). Further details are provided below.
[0309] In one embodiment, if a coherent agent C14A-C14n modifies an entire cache block, it can use a cache read exclusive data only (CRdE-Donly) message. If the cache block has not been modified in another coherent agent C14A-C14n, the requesting coherent agent C14A-C14n can use the EnD cache state to modify all bytes of the block without transferring previous data in the cache block to the agent. If the cache block has been modified, the modified cache block can be transferred to the requesting coherent agent C14A-C14n, and the requesting coherent agent C14A-C14n can use the M cache state.
[0310] Non-cacheable transactions can be initiated using non-cacheable read and non-cacheable write (NCRd and NCWr) messages.
[0311] Snoop forward and snoop back (SnpFwd and SnpBk, respectively) may be used for snoops as previously described. There may also be messages requesting various states (e.g., invalid or shared) within the receiving coherent agent C14A-C14n after processing the snoop. There may also be snoop forward messages for CRdE-Donly requests, which request a transfer if the cache block is modified, but not otherwise, and are invalidated at the receiver. In one embodiment, there may also be invalidation-only snoop forward and snoop back requests (e.g., snoops that cause the receiver to invalidate and acknowledge the requestor or memory controller, respectively, without returning data), shown as SnpInvFw and SnpInvBk in table C192.
[0312] Completion messages may include fill messages (Fill) and acknowledgement messages (Ack). The fill message may specify the state of the cache block that the requester should assume upon completion. A cacheable write-back (CWB) message may be used to transfer a cache block to a memory controller C22A-C22m (e.g., based on evicting the cache block from the cache). A copy-back snoop response (CpBkSR) may be used to transfer a cache block to a memory controller C22A-C22m (e.g., based on a snoop-back message). A non-cacheable write completion (NCWrRsp) and a non-cacheable read completion (NCRdRsp) may be used to complete non-cacheable requests.
[0313] FIG. 41 is a flow diagram illustrating the operation of one embodiment of the coherency controller C24 upon receipt of a conditionally exclusive change (CtoECond) message. For example, FIG. 41 may be a more detailed description of a portion of block C70 of FIG. 32, in one embodiment. For ease of understanding, the blocks are shown in a particular order, but other orders may be used. The blocks may be performed in parallel in combinational logic within the coherency controller C24. The blocks, combinations of the blocks, and / or the entire flowchart may be pipelined over multiple clock cycles. The coherency controller C24 may be configured to implement the operations shown in FIG. 41.
[0314] A CtoECond message may be issued by a coherent agent C14A-14n ("source") based on the execution of a conditional store instruction. If the source loses its copy of the cache block (e.g., the copy is no longer valid) before the store-conditional instruction, the store-conditional instruction may fail locally at the source. If the source still has a valid copy (e.g., in a secondary or primary shared state, or an owned state), when the conditional store instruction is executed, it is still possible that another transaction may be ordered before an exclusive-modify message from the source that causes the source to invalidate its cached copy. The same transaction that invalidates the cached copy also causes the store-conditional instruction to fail at the source. To avoid invalidating the cache block and transferring it to the source, which would cause the store-conditional instruction to fail, a CtoECond message is provided and may be used by the source.
[0315] The CtoECond message can be defined to have at least two possible outcomes when ordered by the coherency controller C24. If, once the CtoECond message has been ordered and processed, the source still has a valid copy of the cache block, as indicated in the directory C26, the CtoECond can proceed to issue a snoop and obtain exclusive status for the cache block, similar to a non-conditional CtoE message. If the source does not have a valid copy of the cache block, the coherency controller C24 can fail the CtoE transaction and return an Ack Complete to the source with an indication that the CtoE failed. The source can terminate the CtoE transaction based on the Ack Complete.
[0316] As shown in FIG. 41, the coherency controller C24 may be configured to read the directory entry for the address (block C194). If the source holds a valid copy of the cache block (e.g., in a shared state) (decision block C196, "yes" branch), the coherency controller C24 may be configured to generate a snoop (e.g., a snoop to invalidate the cache block so that the source can change it to an exclusive state) based on the cache state in the directory entry (block C198). If the source does not hold a valid copy of the cache block (decision block C196, "no" branch), the cache controller C24 may be configured to transmit an acknowledge completion to the source indicating failure of the CtoECond message (block C200). Thus, the CtoE transaction may be terminated.
[0317] Referring now to FIG. 42, a flow diagram is shown illustrating the operation of one embodiment of a coherency controller C24 (e.g., in one embodiment, at least a portion of block C70 of FIG. 32) to read directory entries and determine snoops. For ease of understanding, the blocks are shown in a particular order, although other orders may be used. The blocks may be performed in parallel in combinational logic within the coherency controller C24. Blocks, combinations of blocks, and / or the entire flowchart may be pipelined over multiple clock cycles. The coherency controller C24 may be configured to implement the operations shown in FIG. 42.
[0318] As shown in FIG. 42, the coherency controller C24 may be configured to read the directory entry for the address of the request (block C202). Based on the cache state in the directory entry, the coherency controller C24 may be configured to generate a snoop. For example, based on the cache state in at least one of the primary shared agents (decision block C204, “Yes” branch), the coherency controller C24 may be configured to transmit an SnpFwd snoop to the primary shared agent indicating that the primary shared agent should transmit the cache block to the requesting agent. For other agents (e.g., in the secondary shared state), the coherency controller C24 may be configured to generate an invalidate-only snoop (SnpInv) indicating that the other agents will not transmit the cache block to the requesting agent (block C206). In some cases (e.g., a CRdS request requesting a shared copy of the cache block), the other agents do not need to receive the snoop because they do not need to change state. An agent may have a cache state that is at least first-order shared if that cache state is at least as permissive as first-order shared (e.g., in the embodiment of FIG. 39, first-order shared, owned, exclusive, or modified).
[0319] If there is no agent with a cache state that is at least primary shared (decision block C204, "No" branch), the coherency controller C24 may be configured to determine whether one or more agents have a cache block in a secondary shared state (decision block C208). If so (decision block C208, "Yes" branch), the coherency controller C24 may be configured to select one of the agents with the secondary shared state and transmit an SnpFwd request instruction to the requesting agent for the selected agent to transfer the cache block. The coherency controller C24 may be configured to generate an SnpInv request for other agents in the secondary shared state, indicating that the other agents will not transmit the cache block to the requesting agent (block C210). As described above, if the other agents do not need to change their state, an SnpInv message need not be generated and transmitted.
[0320] If there are no agents with secondary shared cache state (decision block C208, "No" branch), the coherency controller C24 may be configured to generate a fill completion and cause the memory controller to read the cache block for transmission to the requesting agent (block C212).
[0321] FIG. 43 is a flow diagram illustrating the operation of one embodiment of the coherency controller C24 (e.g., in one embodiment, at least a portion of block C70 of FIG. 32) to read a directory entry and determine a snoop in response to a CRdE-Donly request. For ease of understanding, the blocks are shown in a particular order, but other orders may be used. The blocks may be performed in parallel in combinational logic within the coherency controller C24. The blocks, combinations of the blocks, and / or the entire flowchart may be pipelined over multiple clock cycles. The coherency controller C24 may be configured to implement the operations shown in FIG. 43.
[0322] As described above, the CRdE-Donly request may be used by coherent agents C14A-C14n that modify all bytes in a cache block. Thus, the coherency controller C24 can cause other agents to invalidate the cache block. If an agent has the modified cache block, it can supply the modified cache block to the requesting agent. Otherwise, the agent does not have to supply the cache block.
[0323] The coherency controller C24 may be configured to read the directory entry for the address of the request (block C220). Based on the cache state in the directory entry, the coherency controller C24 may be configured to generate a snoop. More specifically, if a given agent may have a modified copy of the cache block (e.g., the given agent has a cache block in an exclusive or primary state) (block C222, “Yes” branch), the cache controller C24 may generate a snoop forward dirty only (SnpFwdDonly) to the agent to transmit the cache block to the requesting agent (block C224). As described above, the SnpFwdDonly request may cause the receiving agent to transmit the cache block if the data has been modified, or not transmit the cache block otherwise. In either case, the receiving agent may invalidate the cache block. The receiving agent may transmit a Fill completion if the data has been modified and provide the modified cache block. Otherwise, the receiving agent may transmit an Ack completion. If no agents have modified copies (decision block C222, "No" branch), the coherency controller C24 may be configured to generate a snoop invalidate (SnpInv) for each agent that has a cached copy of the cache block (block C226). In another embodiment, the coherency controller C24 may not request a transfer of data even if the cache block is modified because the requestor modifies the entire cache block. That is, the coherency controller C24 may cause agents that have modified copies to invalidate the data without transferring it.
[0324] Based on the present disclosure, a system may include a plurality of coherent agents, where a given agent of the plurality of coherent agents includes one or more caches for caching memory data. The system may further include a memory controller coupled to the one or more memory devices, the memory controller including a directory configured to track which of the plurality of coherent agents cache copies of a plurality of cache blocks in the memory device and the states of the cached copies in the plurality of coherent agents. Based on a first request for a first cache block by a first agent of the plurality of coherent agents, the memory controller may be configured to read an entry corresponding to the first cache block from the directory, issue a snoop to a second agent of the plurality of coherent agents having a cached copy of the first cache block according to the entry, and include in the snoop an identifier of a first state of the first cache block in the second agent. Based on the snoop, the second agent may be configured to compare the first state with a second state of the first cache block within the second agent and delay processing of the snoop based on the first state not matching the second state until the second state is changed to the first state in response to a different communication associated with a different request than the first request. In one embodiment, the memory controller may be configured to determine a completion count indicating a number of completions the first agent will receive for the first request, the determination being based on the state from the entry, and include the completion count in multiple snoops issued based on the first request, including the snoop issued to the second agent. The first agent may be configured to initialize the completion counter with the completion count based on receiving an initial completion from one of the multiple coherent agents, update the completion counter based on receiving a subsequent completion from another one of the multiple coherent agents, and complete the first request based on the completion counter.In one embodiment, the memory controller may be configured to update a state in an entry of the directory to reflect completion of the first request based on issuing multiple snoops based on the first request. In one embodiment, the first agent may be configured to detect a second snoop received by the first agent for the first cache block, and the first agent may be configured to absorb the second snoop into the first request. In one embodiment, the first agent may be configured to process the second snoop after completing the first request. In one embodiment, the first agent may be configured to transfer the first cache block to a third agent indicated in the second snoop after completing the first request. In one embodiment, the third agent may be configured to generate a conditional change to exclusive state request based on a conditional store instruction to the second cache block that is in a valid state at the third agent. The memory controller may be configured to determine whether a third agent holds a valid copy of the second cache block based on a second entry in the directory associated with the second cache block, and the memory controller may be configured to transmit a completion indicating failure to the third agent and terminate the conditional modification to the exclusive request based on a determination that the third agent no longer holds a valid copy of the second cache block. In one embodiment, the memory controller may be configured to issue one or more snoops to other coherent agents of the plurality of coherent agents indicated by the second entry based on a determination that the third agent holds a valid copy of the second cache block. In one embodiment, the snoops indicate that the second agent should transmit the first cache block to the first agent based on the first state being primary shared, and the snoops indicate that the second agent should not transmit the first cache block based on the first state being secondary shared.In one embodiment, the snoop indicates that the second agent transmits the first cache block even if the first state is secondary shared.
[0325] In another embodiment, a system includes a plurality of coherent agents, a given agent of the plurality of coherent agents including one or more caches for caching memory data. The system further includes a memory controller coupled to the one or more memory devices. The memory controller may include a directory configured to track which of the plurality of coherent agents cache copies of a plurality of cache blocks in the memory device and the state of the cached copies in the plurality of coherent agents. Based on a first request for a first cache block by a first agent of the plurality of coherent agents, the memory controller may be configured to read an entry corresponding to the first cache block from the directory and issue a snoop to a second agent of the plurality of coherent agents that has a cached copy of the first cache block according to the entry. The snoop may indicate that the second agent should transmit the first cache block to the first agent based on the entry indicating that the second agent has the first cache block in at least a primary shared state. The snoop indicates that the second agent should not transmit the first cache block to the first agent based on a different agent having the first cache block in at least a primary shared state. In one embodiment, if the different agent is in the primary shared state, the first agent is in a secondary shared state with respect to the first cache block. In one embodiment, the snoop indicates that the first agent should invalidate the first cache block based on a different agent having the first cache block in at least a primary shared state. In one embodiment, the memory controller is configured to not issue a snoop to the second agent based on a different agent having the first cache block in a primary shared state and the first request being a request for a shared copy of the first cache block.In one embodiment, the first request may be for an exclusive state for the first cache block, and the first agent is to modify the entire first cache block. The snoop may indicate that the second agent should transmit the first cache block if the first cache block is in a modified state at the second agent. In one embodiment, the snoop indicates that the second agent should invalidate the first cache block if the first cache block is not in a modified state at the second agent.
[0326] In another embodiment, a system includes a plurality of coherent agents, a given agent of the plurality of coherent agents including one or more caches for caching memory data. The system further includes a memory controller coupled to the one or more memory devices. The memory controller may include a directory configured to track which of the plurality of coherent agents cache copies of a plurality of cache blocks in the memory device and the state of the cached copies in the plurality of coherent agents. Based on a first request for a first cache block, the memory controller may be configured to read an entry corresponding to the first cache block from the directory and issue a snoop to a second agent of the plurality of coherent agents that has a cached copy of the first cache block according to the entry. The snoop may indicate that the second agent should transmit the first cache block to a source of the first request based on an attribute associated with the first request having a first value, and the snoop indicates that the second agent should transmit the first cache block to the memory controller based on an attribute having a second value. In one embodiment, the attribute is a type of request, the first value being cacheable, and the second value being non-cacheable. In another embodiment, the attribute is a source of the first request. In one embodiment, the memory controller may be configured to respond to the source of the first request based on receiving the first cache block from the second agent. In one embodiment, the memory controller is configured to update a state in an entry of the directory to reflect completion of the first request based on issuing multiple snoops based on the first request. IOA
[0327] 44-48 illustrate various embodiments of an input / output agent (IOA) that may be employed in various embodiments of a SOC. The IOA may be inserted between a given peripheral device and the interconnect fabric. The IOA agent may be configured to enforce the coherency protocol of the interconnect fabric for the given peripheral device. In one embodiment, the IOA uses the coherency protocol to ensure ordering of requests from the given peripheral device. In one embodiment, the IOA is configured to couple a network of two or more peripheral devices to the interconnect fabric.
[0328] Computer systems often implement data / cache coherency protocols in which a coherent view of data is guaranteed within the computer system. As a result, changes to shared data are usually propagated throughout the computer system in a timely manner to ensure a consistent view. Computer systems also typically include or interface with peripherals, such as input / output (I / O) devices. However, these peripherals are not configured to understand or efficiently use the cache coherency protocols implemented by the computer systems. For example, peripherals often use specific ordering rules (discussed further below) for their transactions that are stricter than cache coherency protocols. Many peripherals also do not have caches, i.e., they are not cacheable devices. As a result, it may take a significant amount of time for peripherals to receive completion acknowledgments for their transactions because they are not completed in their local caches. The present disclosure addresses these technical issues, among other things, with respect to peripherals that cannot properly use cache coherency protocols and do not have caches.
[0329] This disclosure describes various techniques for implementing an I / O agent configured to bridge peripheral devices to a coherent fabric and implement a coherency mechanism for processing transactions associated with those I / O devices. In various embodiments described below, a system-on-chip (SOC) includes a memory coupled to a peripheral device, a memory controller, and an I / O agent. The I / O agent is configured to receive read and write transaction requests from the peripheral device that target specified memory addresses where data may be stored in a cache line of the SOC. (A cache line is sometimes referred to as a cache block.) In various embodiments, specific ordering rules of the peripheral device impose that read / write transactions be completed consecutively (e.g., not out of order with respect to the order in which they are received). As a result, in one embodiment, the I / O agent is configured to complete a read / write transaction according to its execution order before initiating a next occurring read / write transaction. However, to perform those transactions in a more performant manner, in various embodiments, the I / O agent is configured to obtain exclusive ownership of the targeted cache lines so that the data in those cache lines is not cached in a valid state by other caching agents (e.g., processor cores) of the SOC. Instead of waiting for the first transaction to complete before beginning work on the second transaction, the I / O agent can preemptively obtain exclusive ownership of the cache line(s) targeted by the second transaction. As part of obtaining exclusive ownership, in various embodiments, the I / O agent receives the data in those cache lines and stores the data in the I / O agent's local cache.Once the first transaction is complete, the I / O agent can then submit a request for the data in those cache lines and complete a second transaction in its local cache without having to wait for the data to be returned. As described in more detail below, the I / O agent can obtain exclusive read ownership or exclusive write ownership depending on the type of transaction involved.
[0330] In some cases, an I / O agent may lose exclusive ownership of a cacheline before the I / O agent can perform the corresponding transaction. For example, the I / O agent may receive a snoop that causes the I / O agent to relinquish exclusive ownership of the cacheline, including invalidating data stored at the I / O agent for the cacheline. As used herein, a “snoop” or “snoop request” refers to a message transmitted to a component to request a state change of a cacheline (e.g., invalidating the data of the cacheline stored in the component's cache); the message may also request that the cacheline be provided by a component if that component has an exclusive copy of the cacheline or is otherwise responsible for the cacheline. In various embodiments, an I / O agent may regain exclusive ownership of a cacheline if a threshold number of outstanding transactions directed to the cacheline remain. For example, if there are three outstanding write transactions targeting a cacheline, the I / O agent may regain exclusive ownership of the cacheline. This may prevent unduly slow serialization of remaining transactions targeting a particular cacheline. In various embodiments, a greater or lesser number of outstanding transactions may be used as the threshold.
[0331] These techniques may be advantageous over conventional approaches because they allow peripheral ordering rules to be preserved while partially or completely negating the adverse effects of those rules by implementing coherency mechanisms. In particular, the paradigm of performing transactions in a specific order according to ordering rules, in which a transaction is completed before work on a subsequent transaction begins, can be unduly slow. As an example, reading a cache line of data into a cache can take more than 500 clock cycles. Thus, if a subsequent transaction does not begin until the previous transaction has completed, each transaction would take at least 500 clock cycles to complete, resulting in a high number of clock cycles used to process a set of transactions. As disclosed in this disclosure, the high number of clock cycles for each transaction can be avoided by preemptively obtaining exclusive ownership of the associated cache lines. For example, when an I / O agent is processing a set of transactions, the I / O agent may preemptively begin caching data before the first transaction completes. As a result, the data for a second transaction can be cached and made available when the first transaction completes, allowing the I / O agent to complete the second transaction shortly thereafter. Thus, some of the transactions may each require no more than, for example, 500 clock cycles to complete. An exemplary application of these techniques will now be described with reference to Figure 44.
[0332] Referring now to FIG. 44, a block diagram of an exemplary system-on-chip (SOC) D100 is shown. In one embodiment, SOC D100 may be an embodiment of SOC 10 shown in FIG. 1. As implied by the name, the components of SOC D100 are integrated on a single semiconductor substrate as an integrated circuit "chip." However, in some embodiments, the components are implemented on two or more separate chips within a computing system. In the illustrated embodiment, SOC D100 includes a caching agent D110, memory controllers D120A and D120B coupled to memories D130A and D130B, respectively, and an input / output (I / O) cluster D140. The components D110, D120, and D140 are coupled to each other via interconnect D105. Also as shown, caching agent D110 includes processor D112 and cache D114, and I / O cluster D140 includes I / O agent D142 and peripherals D144. In various embodiments, SOC D100 is implemented differently than that shown. For example, SOC D100 may include a display controller, power management circuitry, etc., and memories D130A and D130B may be included on SOC D100. As another example, I / O cluster D140 may have multiple peripherals D144, one or more of which may be external to SOC D100. Therefore, it should be noted that the number of components (and also the number of subcomponents) of SOC D100 may vary between embodiments. Each component / subcomponent may have more or fewer components than shown in FIG. 44.
[0333] A caching agent D110, in various embodiments, is any circuitry that includes a cache for caching memory data or that potentially controls cache lines and potentially updates the data in those cache lines locally. The caching agent D110 can participate in a cache coherency protocol to ensure that updates to data made by one caching agent D110 are visible to other caching agents D110 that subsequently read the data, and that updates made by two or more caching agents D110 in a particular order (determined at an ordering point within the SOC D100, such as memory controllers D120A-B) are observed in that order by the caching agent D110. The caching agent D110 can include, for example, a processing unit (e.g., a CPU, a GPU, etc.), fixed-function circuitry, and fixed-function circuitry with processor support via one or more embedded processors. Because the I / O agent D142 includes a set of caches, the I / O agent D142 can be considered a type of caching agent D110. However, I / O agent D142 differs from other caching agents D110 at least because I / O agent D142 functions as a cacheable entity configured to cache data for other separate entities that do not have their own caches (e.g., peripherals such as displays, USB-connected devices, etc.) In addition, I / O agent D142 may temporarily cache a relatively small number of cache lines to improve peripheral memory access latency, but may aggressively retire cache lines once transactions are complete.
[0334] In the illustrated embodiment, the caching agent D110 is a processing unit having a processor D112 that can function as the CPU of the SOC D100. The processor D112, in various embodiments, includes any circuitry and / or microcode configured to execute instructions defined in an instruction set architecture implemented by the processor D112. The processor D112 may include one or more processor cores implemented on an integrated circuit with other components of the SOC D100. The individual processor cores of the processor D112 may share a common last-level cache (e.g., an L2 cache) while including their own respective caches (e.g., an L0 cache and / or an L1 cache) for storing data and program instructions. The processor D112 may execute the system's main control software, such as an operating system. Generally, the software executed by the CPU controls the other components of the system to achieve the desired functionality of the system. The processor D112 may also execute other software, such as application programs, and thus may be referred to as an application processor. Caching agent D110 may further include hardware configured to interface caching agent D110 to other components of SOC D100 (eg, an interface to interconnect D105).
[0335] The cache D114, in various embodiments, is a storage array containing entries configured to store data or program instructions. Thus, the cache D114 may be a data cache or an instruction cache, or a shared instruction / data cache. The cache D114 may be an associative storage array (e.g., fully associative or set associative, such as 4-way content) or a direct-mapped storage array and may have any storage set associative scheme. In various embodiments, a cache line (or alternatively, a "cache block") is the unit of allocation and deallocation within the cache D114 and may be of any desired size (e.g., 32 bytes, 64 bytes, 128 bytes, etc.). During operation of the caching agent D110, information may be pulled into the cache D114 from other components of the system and used by the processor cores of the processor D112. For example, as the processor core progresses through its execution path, the processor core may cause program instructions to be fetched from memory D130A-B to cache D114, and the processor core may then fetch them from cache D114 and execute them. Also, during operation of the caching agent D110, information may be written from cache D114 to memory (e.g., memory D130A-B) via memory controllers D120A-B.
[0336] The memory controller D120, in various embodiments, includes circuitry configured to receive memory requests (e.g., load / store requests, instruction fetch requests, etc.) from other components of the SOC D100 to perform memory operations, such as accessing data from the memory D130. The memory controller D120 may be configured to access any type of memory D130. The memory D130 may be implemented using a variety of different physical memory media, such as hard disk storage, floppy disk storage, removable disk storage, flash memory, random access memory (RAM-SRAM, EDO RAM, SDRAM, DDR SDRAM, RAMBUS RAM, etc.), read-only memory (PROM, EEPROM, etc.), etc. However, the memory available to the SOC D100 is not limited to primary storage such as the memory D130. Rather, the SOC D100 may further include other forms of storage, such as cache memory (e.g., L1 cache, L2 cache, etc.) within the caching agent D110. In some embodiments, the memory controller D120 includes a queue for storing and ordering memory operations to be presented to the memory D130. The memory controller D120 may also include data buffers for storing write data waiting to be written to the memory D130 and read data waiting to be returned to the source of the memory operation, such as the caching agent D110.
[0337] As described in more detail with respect to FIG. 45 , the memory controller D120 may include various components for maintaining cache coherency within the SOC D100, including components that track the location of cache line data within the SOC D100. Thus, in various embodiments, requests for cache line data are routed through the memory controller D120, which can access data from other caching agents D110 and / or memories D130A-B. In addition to accessing data, the memory controller D120 can cause the caching agents D110 and I / O agents D142 that store the data in their local caches to issue snoop requests. As a result, the memory controller D120 can cause those caching agents D110 and I / O agents D142 to invalidate and / or evict the data from their caches to ensure coherency within the system. Thus, in various embodiments, the memory controller D120 processes exclusive cache line ownership requests, and the memory controller D120 grants component exclusive ownership of the cache line while using snoop requests to ensure that the data is not cached in other caching agents D110 and I / O agents D142.
[0338] I / O cluster D140, in various embodiments, includes one or more peripheral devices D144 (or simply peripherals D144) that may provide additional hardware functionality and I / O agents D142. Peripherals D144 may include, for example, video peripherals (e.g., GPUs, blenders, video encoders / decoders, scalers, display controllers, etc.) and audio peripherals (e.g., microphones, speakers, interfaces to microphones and speakers, digital signal processors, audio processors, mixers, etc.). Peripherals D144 may include interface controllers for various interfaces external to SOC D100 (e.g., universal serial bus (USB), peripheral component interconnect (PCI) and PCI express (PCIe), serial and parallel ports, etc.). Interconnects to external components are indicated in FIG. 44 by dashed arrows extending outside of SOC D100. Peripherals D144 may also include networking peripherals such as media access controllers (MACs). Although not shown, in various embodiments, the SOC D100 includes multiple I / O clusters D140 having respective sets of peripherals D144. As an example, the SOC D100 may include a first I / O cluster D140 having an external display peripheral D144, a second I / O cluster D140 having a USB peripheral D144, and a third I / O cluster D140 having a video encoder peripheral D144. Each of those I / O clusters D140 may include its own I / O agent D142.
[0339] The I / O agent D142, in various embodiments, includes circuitry configured to bridge its peripherals D144 to the interconnect D105 and implement a coherency mechanism for processing transactions associated with those peripherals D144. As described in more detail with respect to FIG. 45 , the I / O agent D142 may receive transaction requests from the peripherals D144 to read and / or write data to cache lines associated with the memories D130A-B. In response to those requests, in various embodiments, the I / O agent D142 communicates with the memory controller D120 to obtain exclusive ownership of the target cache line. Thus, the memory controller D120 may grant exclusive ownership to the I / O agent D142, which may involve providing the cache line data to the I / O agent D142 and sending snoop requests to other caching agents D110 and the I / O agent D142. After obtaining exclusive ownership of the cache line, I / O agent D142 may begin completing the transaction targeting the cache line. In response to completing the transaction, I / O agent D142 may send an acknowledgment to the requesting peripheral device D144 that the transaction is completed. In some embodiments, I / O agent D142 does not obtain exclusive ownership for relaxed-ordered requests that do not have to be completed in a specified order.
[0340] The interconnect D105, in various embodiments, is any communications-based interconnect and / or protocol for communicating between components of the SOC D100. For example, the interconnect D105 may enable the processor D112 in the caching agent D110 to interwork with the peripherals D144 in the I / O cluster D140. In various embodiments, the interconnect D105 is bus-based, including shared bus configurations, crossbar configurations, and hierarchical buses with bridges. The interconnect D105 may be packet-based, hierarchical with bridges, crossbars, point-to-point, or other interconnects.
[0341] Referring now to FIG. 45, a block diagram of exemplary elements of interaction including a caching agent D110, a memory controller D120, an I / O agent D142, and a peripheral D144 is shown. In the illustrated embodiment, the memory controller 120 includes a coherency controller D210 and a directory D220. In some cases, the illustrated embodiment may be implemented differently than shown. For example, there may be multiple caching agents D110, multiple memory controllers D120, and / or multiple I / O agents D142.
[0342] As noted, the memory controller D120 can maintain cache coherency within the SOC D100, including tracking the locations of cache lines within the SOC D100. Thus, the coherency controller D210, in various embodiments, is configured to implement the memory controller portion of a cache coherency protocol. The cache coherency protocol can specify messages or commands that can be transmitted between the caching agent D110, the I / O agent D142, and the memory controller D120 (or the coherency controller D210) to complete a coherent transaction. These messages can include a transaction request D205, a snoop D225, and a snoop response D227 (or alternatively, a “completion”). The transaction request D205, in various embodiments, is a message that initiates a transaction and specifies the requested cache line / block (e.g., having the address of the cache line) and the state in which the requestor receives the cache line (or, in various cases, a minimum state such that a more permissive state can be provided). The transaction request D205 may be a write transaction in which the requestor attempts to write data to a cache line, or a read transaction in which the requestor attempts to read data from a cache line. For example, the transaction request D205 may specify a non-relaxed-ordered dynamic random access memory (DRAM) request. The coherency controller D210, in some embodiments, is also configured to issue memory requests D222 to the memory D130 to access data from the memory D130 on behalf of components of the SOC D100, and to receive memory responses D224, which may include the requested data.
[0343] As shown, the I / O agent D142 receives transaction requests D205 from a peripheral device D144. The I / O agent D142 may receive a series of write transaction requests D205, a series of read transaction requests D205, or a combination of read and write transaction requests D205 from a given peripheral device D144. For example, within a set time interval, the I / O agent D142 may receive four read transaction requests D205 from peripheral device D144A and three write transaction requests D205 from peripheral device D144B. In various embodiments, the transaction requests D205 received from the peripheral device D144 must be completed in a particular order (e.g., they must be completed in the order they were received from the peripheral device D144). Instead of waiting until a transaction request D205 completes before beginning work on the next transaction request D205 in the sequence, in various embodiments, the I / O agent D142 performs work on the later request D205 by preemptively obtaining exclusive ownership of the target cache line. Accordingly, the I / O agent D142 can issue an exclusive ownership request D215 to the memory controller D120 (particularly the coherency controller D210). In some examples, a set of transaction requests D205 can target cache lines managed by different memory controllers D120, and therefore the I / O agent 142 can issue an exclusive ownership request D215 to the appropriate memory controller D120 based on those transaction requests D205. In the case of a read transaction request D205, the I / O agent D142 can obtain exclusive read ownership. In the case of a write transaction request D205, the I / O agent D142 can obtain exclusive write ownership.
[0344] The coherency controller D210, in various embodiments, is circuitry configured to receive requests (e.g., exclusive ownership requests D215) from the interconnect D105 (e.g., via one or more queues included in the memory controller D120) that target cache lines mapped to the memory D130 to which the memory controller D120 is coupled. The coherency controller D210 can process those requests and generate responses (e.g., exclusive ownership responses D217) with the data of the requested cache line while also maintaining cache coherency within the SOC D100. To maintain cache coherency, the coherency controller D210 can use a directory D220. The directory D220, in various embodiments, is a storage array having a set of entries, each of which can track the coherency state of an individual cache line in the system. In some embodiments, the entries also track the location of the cache line's data. For example, an entry in the directory D220 may indicate that data for a particular cache line is cached in the cache D114 of the caching agent D110 in a valid state. (Although exclusive ownership is discussed, in some cases, a cache line may be shared among multiple cache-capable entities (e.g., caching agents D110) for read purposes, and thus shared ownership may be provided.) To provide exclusive ownership of a cache line, the coherency controller D210 may ensure that the cache line is not stored in a valid state outside of the memory D130 and the memory controller D120. Thus, based on the directory entry associated with the cache line targeted by the exclusive ownership request D215, in various embodiments, the coherency controller D210 determines which component (e.g., the caching agent D110, the I / O agent D142, etc.) should receive the snoop D225 and the type of snoop D225 (e.g., invalidation, change to ownership, etc.).For example, memory controller D120 may determine that caching agent 110 stores the cache line of data requested by I / O agent D142 and therefore may issue snoop D225 to caching agent D110, as shown in Figure 45. In some embodiments, coherency controller D210 does not target a particular component, but instead broadcasts snoop D225 which is observed by many of the components of SOC D100.
[0345] In various embodiments, at least two types of snoops are supported: snoop forward and snoop back. A snoop forward message may be used to cause a component (e.g., a caching agent D110) to forward a cache line's data to a requesting component, and a snoop back message may be used to cause a component to return a cache line's data to the memory controller D120. Supporting snoop forward and snoop back flows can enable both three-hop (snoop forward) and four-hop (snoop back) behavior. For example, snoop forward may be used to minimize the number of messages when a cache line is provided to a component, since the component stores the cache line and may potentially use the data therein. On the other hand, non-cacheable components may not store the entire cache line, and therefore, copy back to memory can ensure that the full cache line data is retrieved by the memory controller D120. In various embodiments, the caching agent D110 receives a snoop D225 from the memory controller D120, processes the snoop D225 to update the cache line state (e.g., invalidate the cache line), and (if specified by the snoop D225) returns a copy of the cache line's data to the initial ownership requestor or the memory controller D120. The snoop response D227 (or "completion"), in various embodiments, is a message indicating that the state change has occurred and, if applicable, provides a copy of the cache line data. When a snoop forward mechanism is used, data is provided to the requesting component in three hops over the interconnect D105: a request from the requesting component to the memory controller D120, a snoop from the memory controller D120 to the caching component, and a snoop response by the caching component to the requesting component.When the snoop-back mechanism is used, four hops can occur: a request and snoop as in a three-hop protocol, a snoop response by the caching component to the memory controller D120, and data from the memory controller D120 to the requesting component.
[0346] In some embodiments, the coherency controller D210 may update the directory D220 when the snoop D225 is generated and transmitted instead of when the snoop response D227 is received. Once the requested cache line is reclaimed by the memory controller D120, in various embodiments, the coherency controller D210 grants exclusive read (or write) ownership to the ownership requestor (e.g., the I / O agent D142) via an exclusive ownership response D217. The exclusive ownership response D217 may include data for the requested cache line. In various embodiments, the coherency controller D210 updates the directory D220 to indicate that the cache line has been granted to the ownership requestor.
[0347] For example, the I / O agent D142 may receive a series of read transaction requests D205 from a peripheral device D144A. For a given one of those requests, the I / O agent D142 may send an exclusive read ownership request D215 for data associated with a particular cache line to the memory controller D120 (or, if the cache line is managed by another memory controller D120, the exclusive read ownership request D215 is sent to the other memory controller D120). The coherency controller D210 may determine, based on entries in the directory D220, that the caching agent D110 currently stores the data associated with a particular cache line in a valid state. Thus, the coherency controller D210 sends a snoop D225 to the caching agent D110, causing the caching agent D110 to relinquish ownership of the cache line and return a snoop response D227, which may include the cache line data. After receiving the snoop response D227, the coherency controller D210 can generate and send an exclusive ownership response D217 to the I / O agent D142, providing the I / O agent D142 with the cache line data and exclusive ownership of the cache line.
[0348] After receiving exclusive ownership of the cacheline, in various embodiments, the I / O agent D142 waits until the corresponding transaction can be completed (according to the ordering rules). That is, until the corresponding transaction becomes the most senior transaction and an ordering dependency resolution exists for the transaction. For example, the I / O agent D142 may receive a transaction request D205 from the peripheral device D144 to perform write transactions A-D. The I / O agent D142 may obtain exclusive ownership of the cacheline associated with transaction C, but transactions A and B may not have completed. As a result, the I / O agent D142 waits until transactions A and B have completed before writing the associated data of the cacheline associated with transaction C. After completing a given transaction, in various embodiments, the I / O agent D142 provides a transaction response D207 to the transaction requestor (e.g., the peripheral device D144A) indicating that the requested transaction has been performed. In various cases, I / O agent D142 may acquire exclusive read ownership of a cache line, perform a set of read transactions on the cache line, and then release exclusive read ownership of the cache line without performing any writes to the cache line while exclusive read ownership was held.
[0349] In some cases, the I / O agent D142 may receive multiple transaction requests D205 targeting the same cache line (within a reasonably short period of time), resulting in the I / O agent D142 being able to perform bulk reads and writes. As an example, two write transaction requests D205 received from a peripheral device D144A may target a lower and upper portion of the cache line, respectively. The I / O agent D142 may therefore obtain exclusive write ownership of the cache line and retain the data associated with the cache line until at least both write transactions are completed. Thus, in various embodiments, the I / O agent D142 may transfer executive ownership between transactions targeting the same cache line. That is, the I / O agent D142 need not send an ownership request D215 for each individual transaction request D205. In some cases, I / O agent D142 can transfer executive ownership from a read transaction to a write transaction (or vice versa), but in other cases, I / O agent D142 transfers executive ownership only between transactions of the same type (e.g., from a read transaction to another read transaction).
[0350] In some cases, the I / O agent D142 may lose exclusive ownership of a cacheline before the I / O agent D142 performs an associated transaction on the cacheline. As an example, while waiting for a transaction to become most senior so that it can be performed, the I / O agent D142 may receive a snoop D225 from the memory controller D120 as a result of another I / O agent D142 attempting to acquire exclusive ownership of the cacheline. After relinquishing exclusive ownership of the cacheline, in various embodiments, the I / O agent D142 determines whether to regain ownership of the lost cacheline. If the lost cacheline is associated with one pending transaction, the I / O agent D142 often does not regain exclusive ownership of the cacheline, but in some cases, if the pending transaction is behind a configured number of transactions (and therefore not attempting to become a senior transaction), the I / O agent D142 may issue an exclusive ownership request D215 for the cacheline. However, in various embodiments, if there is a threshold number of pending transactions directed to the cache line (e.g., two pending transactions), the I / O agent D142 regains exclusive ownership of the cache line.
[0351] Referring now to FIG. 46A, a block diagram of exemplary elements associated with an I / O agent D142 that processes write transactions is shown. In the illustrated embodiment, the I / O agent D142 includes an I / O agent controller D310 and a coherency cache D320. As shown, the coherency cache D320 includes a fetched data cache D322, a merged data cache D324, and a new data cache D326. In some embodiments, the I / O agent D142 is implemented differently than shown. As one example, the I / O agent D142 may not include separate caches for data retrie...
Claims
1. 1. A system comprising: Multiple processor cores; a plurality of graphics processing units; a plurality of peripheral devices different from the processor core and the graphics processing unit; one or more memory controller circuits configured to interface with the system memory; an interconnect fabric configured to provide communication between the one or more memory controller circuits, the processor cores, the graphics processing unit, and the peripheral devices; The system, wherein the processor core, the graphics processing unit, the peripheral device, and the memory controller are configured to communicate via a unified memory architecture.
2. 2. The system of claim 1, wherein the processor core, the graphics processing unit, and the peripheral device are configured to access any address within a unified address space defined by the unified memory architecture.
3. 3. The system of claim 2, wherein the unified address space is a virtual address space that is distinct from the physical address space provided by the system memory.
4. The system of claim 1 , wherein the unified memory architecture provides a common set of semantics for memory accesses by the processor core, the graphics processing unit, and the peripheral devices.
5. The system of claim 4 , wherein the semantics include memory ordering properties.
6. The system of claim 4 or 5, wherein the semantics include quality of service attributes.
7. The system of claim 4 , wherein the semantics include memory coherency.
8. 8. The system of claim 1, wherein the one or more memory controller circuits include respective interfaces to one or more random access memory mappable memory devices.
9. The system of claim 8 , wherein the one or more memory devices include dynamic random access memory (DRAM).
10. The system of claim 1 , further comprising one or more levels of cache between the processor core, the graphics processing unit, the peripheral devices and the system memory.
11. 11. The system of claim 10, wherein the one or more memory controller circuits include a respective memory cache interposed between the interconnect fabric and the system memory, the respective memory cache being one of the one or more levels of cache.
12. The system of claim 1 , wherein the interconnect fabric includes at least two networks having heterogeneous interconnect topologies.
13. The system of claim 1 , wherein the interconnect fabric includes at least two networks having disparate operating characteristics.
14. 14. The system of claim 12 or 13, wherein the at least two networks include a coherent network interconnecting the processor cores and the one or more memory controller circuits.
15. 15. The system of claim 12, wherein the at least two networks include a relaxed ordering network coupled to the graphics processing unit and the one or more memory controller circuits.
16. 16. The system of claim 15, wherein the peripheral devices include a subset of devices, the subset including one or more of a machine learning accelerator circuit or a relaxed-order bulk media device, and the relaxed-order network further couples the subset of devices to the one or more memory controller circuits.
17. 17. The system of claim 12, wherein the at least two networks include an input / output network coupled to interconnect the peripheral devices and the one or more memory controller circuits.
18. The system of claim 17 , wherein the peripheral devices include one or more real-time devices.
19. 19. The system of claim 12, wherein the at least two networks include a first network that includes one or more characteristics for reducing latency compared to a second network of the at least two networks.
20. 20. The system of claim 19, wherein the one or more characteristics include a shorter route than the second network.
21. 21. The system of claim 19 or 20, wherein the one or more characteristics include the wiring in a metal layer closer to a surface of a substrate on which the system is implemented than wiring for the second network.
22. 22. The system of claim 12, wherein the at least two networks include a first network that includes one or more characteristics for increasing bandwidth compared to a second network of the at least two networks.
23. 23. The system of claim 22, wherein the one or more characteristics include a wider interconnect compared to the second network.
24. 24. The system of claim 22 or 23, wherein the one or more characteristics include wiring in a metal layer that is further from a surface of a substrate on which the system is implemented than the wiring for the second network.
25. 25. The system of claim 12, wherein an interconnect topology adopted by the at least two networks comprises at least one of a star topology, a mesh topology, a ring topology, a tree topology, a fat tree topology, a hypercube topology, or a combination of one or more of the interconnect topologies.
26. 26. The system of claim 12, wherein the operational characteristics adopted by the at least two networks include at least one of strong-order memory coherence or relaxed-order memory coherence.
27. 27. The system of claim 12, wherein the at least two networks are physically and logically independent.
28. 28. The system of claim 12, wherein the at least two networks are physically separate in a first mode of operation, and a first of the at least two networks and a second of the at least two networks are virtual in a second mode of operation and share a single physical network.
29. 29. The system of claim 1, wherein the processor cores, the graphics processing unit, the peripheral devices, and the interconnect fabric are distributed across two or more integrated circuit dies.
30. 30. The system of claim 29, wherein a unified address space defined by the unified memory architecture extends across the two or more integrated circuit dies in a manner transparent to software executing on the processor cores, the graphics processing unit, or the peripheral devices.
31. 31. The system of claim 29 or 30, wherein the interconnect fabric is configured to route, the interconnect fabric extending across the two integrated circuit dies, and communications routed between the source and the destination transparent to the location of the source and destination on the integrated circuit dies.
32. 32. The system of claim 29, wherein the interconnect fabric extends across the two integrated circuit dies using hardware circuitry to automatically route communications between sources and destinations, regardless of whether the sources and destinations are on the same integrated circuit die.
33. 33. The system of any one of claims 29 to 32, further comprising at least one interposer device configured to couple buses of the interconnect fabric across the two or more integrated circuit dies.
34. 34. The system of claim 1, wherein a given integrated circuit die includes local interrupt distribution circuitry for distributing interrupts among processor cores within the given integrated circuit die.
35. 35. The system of claim 34, comprising two or more integrated circuit dies including respective local interrupt distribution circuits, at least one of the two or more integrated circuit dies including a global interrupt distribution circuit, the local interrupt distribution circuit and the global interrupt distribution circuit implementing a multi-level interrupt distribution scheme.
36. 36. The system of claim 35, wherein the global interrupt distribution circuitry is configured to transmit interrupt requests to the local interrupt distribution circuitry in sequence, and the local interrupt distribution circuitry is configured to transmit the interrupt requests to local interrupt destinations in sequence before responding to the interrupt requests from the global interrupt distribution circuitry.
37. 37. The system of claim 1, wherein a given integrated circuit die includes a power manager circuit configured to manage a local power state of the given integrated circuit die.
38. 38. The system of claim 37, comprising two or more integrated circuit dies including respective power manager circuits configured to manage the local power state of the integrated circuit die, at least one of the two or more integrated circuit dies including another power manager circuit configured to synchronize the power manager circuits.
39. 39. The system of claim 1, wherein the peripheral device comprises one or more of an audio processing device, a video processing device, a machine learning accelerator circuit, a matrix arithmetic accelerator circuit, a camera processing circuit, a display pipeline circuit, a non-volatile memory controller, a peripheral component interconnect controller, a security processor, or a serial bus controller.
40. 40. The system of claim 1, wherein the interconnect fabric interconnects coherent agents.
41. 41. The system of claim 40, wherein each one of the processor cores corresponds to a coherent agent.
42. 41. The system of claim 40, wherein a cluster of processor cores corresponds to a coherent agent.
43. 43. The system of claim 1, wherein a given one of the peripheral devices is a non-coherent agent.
44. 44. The system of claim 43, further comprising an input / output agent interposed between the given peripheral device and the interconnect fabric, the input / output agent configured to enforce a coherency protocol of the interconnect fabric with respect to the given peripheral device.
45. 45. The system of claim 44, wherein the input / output agent uses the coherency protocol to ensure ordering of requests from the given peripheral device.
46. 46. The system of claim 44 or 45, wherein the input / output agent is configured to couple a network of two or more peripheral devices to the interconnect fabric.
47. 47. The system of any one of claims 1 to 46, further comprising a hashing circuit configured to distribute memory request traffic across system memory according to a selectively programmable hashing protocol.
48. 48. The system of claim 47, wherein at least one programming of the programmable hashing protocol distributes a sequence of memory requests evenly across multiple memory controllers in the system due to a wide variety of memory requests in the sequence of memory requests.
49. 30. The system of claim 29, wherein at least one programming of the programmable hashing protocol distributes adjacent requests in memory space to physically separate memory interfaces with a specified granularity.
50. 50. The system of claim 1, further comprising a plurality of directories configured to track coherency states of subsets of a unified memory address space specified by the unified memory architecture, the plurality of directories being distributed within the system.
51. 51. The system of claim 50, wherein the multiple directories are distributed across the memory controller.
52. 52. The system of claim 1, wherein a given memory controller of the one or more memory controller circuits includes a directory configured to track a plurality of cache blocks corresponding to data in a portion of the system memory to which the given memory controller interfaces, the directory configured to track which of a plurality of caches in the system has cached a given cache block of the plurality of cache blocks, and the directory is accurate with respect to memory requests ordered and processed in the directory even if the memory request has not yet completed within the system.
53. 53. The system of claim 52, wherein the given memory controller is configured to issue one or more coherency maintaining commands for the given cache block based on a memory request for the given cache block, the one or more coherency maintaining commands including a cache state for the given cache block in a corresponding cache of the plurality of caches, and the corresponding cache is configured to delay processing of the given coherency maintaining command based on the cache state in the corresponding cache not matching the cache state in the given coherency maintaining command.
54. 54. The system of claim 52 or 53, wherein a first cache is configured to store the given cache block in a primary shared state, a second cache is configured to store the given cache block in a secondary shared state, and the given memory controller is configured to cause the first cache to transfer the given cache block to a requestor based on the memory request and the primary shared state in the first cache.
55. the given memory controller is configured to issue one of a first coherency maintaining command and a second coherency maintaining command to a first cache of the plurality of caches based on a type of a first memory request, the first cache being configured to transfer a first cache block to a requestor that issued the first memory request based on the first coherency maintaining command; 55. The system of any one of claims 52 to 54, wherein the first cache is configured to return the first cache block to the given memory controller based on the second coherency-maintaining command.
56. 1. An integrated circuit comprising: Multiple processor cores; a plurality of graphics processing units; a plurality of peripheral devices different from the processor core and the graphics processing unit; one or more memory controller circuits configured to interface with the system memory; an interconnect fabric configured to provide communication between the one or more memory controller circuits and the processor cores, the graphics processing unit, and the peripheral devices; an off-chip interconnect coupled to the interconnect fabric and configured to couple the interconnect fabric to a corresponding interconnect fabric on another instance of the integrated circuit, the interconnect fabric and the off-chip interconnect providing an interface that transparently connects the one or more memory controller circuits, the processor cores, the graphics processing unit, and the peripheral devices in either a single instance of the integrated circuit or two or more instances of the integrated circuit.
57. 1. A method comprising: A method comprising communicating via an interconnect fabric in a system between a plurality of processing cores, a plurality of graphics processing units, a plurality of peripheral devices different from the plurality of processor cores and the plurality of graphics processing units, and one or more memory controller circuits via a unified memory architecture.